Saturday, 28 July 2007

Bug: Warning messages parsing native CSV import

When parsing the first (``{\tt Epoch}'') line of a native CSV database import, two warning messages would be generated because this record contains only two fields, while the code which deletes embedded blanks for log entry records assumed all records to have four fields or more. I made each of these statements conditional upon the field's being defined.

Feature: Decimal character preference in XML and native CSV export formats

Added a new <decimal-character> container to the preferences section of the XML output format whose content is the user's choice for decimal separator character (period or comma). Note that regardless of this setting, decimal numbers within the XML file itself always use period as the decimal separator. Appropriate declarations for this item were added to the XML Document Type Definition and CSS style sheet of this {\tt DOCTYPE}.

Added a sixth field at the end of the ``{\tt Preferences}'' header item in our native CSV export for the decimal character. Natually, this field will be quoted if the decimal character is set to comma.

Bug: Decimal place display in XML log import listing

The synthetic listing generated for log items imported from an XML file did not show decimal places due to an incorrect format code in the {\tt sprintf} which generated the output. (The records were imported correctly; only the listing was affected.) I fixed the format code.

Bug: Setting log weight unit in XML import

Import of an XML database set the weight unit of the log only from the {\tt log-unit} in the preferences, and did not allow the {\tt weight-unit} in an individual monthly log (which might be different) to override the default. I added code to set the unit for a given month from its own {\tt weight-unit}. Note that logs created in the process of importing CSV or XML data are always created using the user's current log unit setting; the log unit in the data imported is used to convert weight values (if necessary) from the unit in the imported log.

Friday, 27 July 2007

Bug: Rounding of diet duration different in JavaScript and Perl CGI application

The JavaScript live update for the diet calculator rounded the diet duration in weeks differently from the Perl code: JavaScript truncated to the next lower integer, while Perl rounded to the nearest integer. This could result in a one-week discrepancy in diet duration between the value shown immediately and that which appeared after the user saved the diet calculator results. I modified the JavaScript code to round the same way as the Perl code does. (Reported by Jim Hollcraft.)

Bug: Dates after 2038-01-19 in Diet Calculator and elsewhere

Arghhh! Perl 5.8 relies upon the underlying C library's {\tt gmtime} function for the Perl {\tt gmtime} function. This means that on a 32-bit platform the Perl function is limited to dates between the start of 1970 and ``doomsday'', 2038-01-19. (I understand that this problem does not exist on native 64-bit platforms and will be fixed in Perl 6.) Even though we have some time to go until the tick of doom, it is easily possible to generate dates beyond 2038 by entering small calorie balance values in the diet calculator. I added a new \verb+Julian::unix_time_to_civil_date_time+ function which uses the Julian day functions to convert a Unix {\tt time()} value to a list of year, month, day of month, hour, minute, seconds (actual values, not the crazy offsets returned by {\tt gmtime}, so this is not a drop-in replacement). I replaced all references to {\tt gmtime} in the program to calls on \verb+unix_time_to_civil_date_time+, which corrects the original problem reported in the diet calculator. There are a few calls on {\tt localtime} left in the code, but these are all in {\tt describe} methods for various objects (used only for administrator debugging output, and all representing times close to the present) and in the generation of log entries, which are also obviously in the present. Since these won't break for more than thirty years, it's likely we'll be on a version of Perl with the truncation fixed before then or, failing that, there's plenty of time to fix them before the dawn of the dreaded day. (Reported by Jim Hollcraft.)

Accessibility: Label wrappers on check boxes in CSV/XML Import and Diet Calculator

Wrapped the checkboxes and labels for the ``Allow overwrite'' and ``List imported records'' options in the Import CSV/XML page and the ``Plot plan in chart'' item in the Diet Calculator with <label> containers so that the labels as well as the checkboxes can be clicked.

Thursday, 26 July 2007

Bug: Diet plan end date beyond current year plus one

If the start or end date of a diet plan in the diet calculator were outside the union of the range of years in the database and the current year plus one, the start and/or end year of the diet plan would not be included in the start and end date selection boxes, resulting in an incorrect date appearing in the form. This most often manifested itself when a long-term diet extends past the end of the year after the present. I added logic to make sure that the range of years included in the start and end date boxes includes the least of the current year and the first year of the diet plan minus one and the greatest of the next year and the last year of the diet plan plus one. (Reported by Jim Hollcraft.)

Bug: Julian date constant definitions

The ``Julian date constant definitions'' macro had an incorrect name. Its name had accidentally been left the same as the support functions macro, which worked fine since the references to the two macros are consecutive. I corrected the name to get rid of a harmless warning message.

Feature/admin: Stack trace for terminated transactions

Added a handler for the {\tt INT} signal which, when received, prints a stack trace to {\tt STDERR} (which will thus appear in the HTTP server error log) and terminates. This simplifies the task of debugging CPU hang or other problems which lead to a CGI program timeout and the resulting 500 response to the requester.

Feature/admin: Cluster synchronisation production mode

Added configuration parameters which allow {\tt ClusterSync.pl}, if started as super-user, to change to a designated group and user identity. Running a Perl program under an assumed identity turns on the ``taint'' mechanism, so input from the transaction directory and the files within it is sanitised before being used in potentially dangerous ways (even though it should, in fact, only be coming from the CGI application, never the ``outside'').

Added much more stringent validation to {\tt ClusterSync.pl} transaction processing. Every file name submitted must begin with the ``Database Directory'' path name, and may not contain abusive (shell-interpreted) characters or sequences such as ``{\tt ..}''. In addition, all input from transaction files is single quoted when used on {\tt system()} commands to prevent attack by overlooked shell escapes. Finally, an invalid transaction type in a transaction file causes an immediate abort. Now, since we're basically using the transaction directory as an interprocess communication channel, this might be deemed paranoia, but ``you can't be too careful''. Besides, one can imagine an attack where somebody manages to hijack another CGI application and trick it into adding bogus transactions to the directory which cause {\tt ClusterSync} to do its dirty work for it.

Added an SHA1 signature as an additional line in cluster synchronisation transaction files. This signature incorporates the content of the transaction as well as our site-secret ``Confirmation signature encoding suffix'', without which it is unlikely in the extreme an attacker will be able to spoof transactions. Signature failure crashes {\tt ClusterSync}, alerting the administrator that something untoward is underway and thwarting an attacker who contemplates a brute-force search for the suffix.

Modified the {\tt Makefile} {\tt publish} and {\tt production} targets to install the cluster synchronisation program as an executable named {\tt ClusterSync} in the {\tt server}{\em n}{\tt /bin/hackdiet} directory. This allows it to work without modification with our standard {\tt /server/init} mechanism, in particular a new {\tt /server/init/hackdiet} script which starts and stops the cluster synchronisation process.

Moved the process ID file for the cluster synchronisation process to {\tt /server/run/ClusterSync/ClusterSync.pid} to conform with our standard structure in the {\tt /server} partition.

Feature/admin: Cluster synchronisation log improvements

Added date and time to the first line of the log items written to standard output by {\tt ClusterSync.pl} when \verb+$verbose+ is set and set standard output to ``autoflush'' mode so log items are written immediately regardless of redirection.

Implemented a proper log file for {\tt ClusterSync.pl}. The full path name for the log file is configured with ``Cluster Synchronisation Log File''. If the null string, logging is disabled. Otherwise, the specified file is opened for appending, and items are appended for each transaction. When logging is active, the program listens for the {\tt HUP} signal and, upon receiving it, closes and re-opens the log file to permit it to be rotated by renaming it and then sending the signal. The format of the log file identical to the information written to {\tt STDOUT} when \verb+$verbose+ is nonzero. The default location for the log file is {\tt /server/log/hackdiet/ClusterSync.log}.

Replaced all of the parallel calls in {\tt ClusterSync.pl} to write output to standard output in verbose mode and to the log file when logging with calls on a new {\tt logmsg} function which writes its arguments to the appropriate destinations according to the global option variables.

Tuesday, 24 July 2007

Bug: CPU loop searching for trend in logs with no weight entries

If a log had an unspecified trend carry-forward and the previous log in the database was present but had no weight entries whatsoever, ``Fill in trend carry-forward from most recent previous log, if required'' would hang in a CPU loop due to a backwards-coded loop termination test. If there were any log entries, the loop would bail out due to a {\tt last}, but for an empty log the termination when the beginning of the log was reached would never occur and the program would crash when the CGI time limit expired. I corrected the loop termination test and verified that the hang no longer occurs for a blank previous log. (Reported by Andres Kievsky.)

Sunday, 22 July 2007

Feature/admin: Cluster file system support for server farms

Completed implementation and began production test of cluster file system synchronisation support for server farm architectures such as Fourmilab's. Cluster support is implemented in the new {\tt Cluster} module, through functions such as {\tt clusterCopy}, {\tt clusterDelete}, {\tt clusterMkdir}, etc. When a database file or directory is modified, immediately after the modification is made, (for example, after the {\tt close()} when writing back a file), the corresponding cluster function is called with the full path name of the modified file. This then calls {\tt enqueueClusterTransaction} with the specified operation and path name, which creates one or more synchronisation transaction files in the {\tt ClusterSync} directory, within subdirectories bearing the names of the servers defined in ``Cluster Member Hosts''. (Transactions are never queued for the server executing the transaction, nor for servers named as cluster members for which no server subdirectory exists. This allows you to have identical directory structures on all servers, or to exercise fine-grained control over which servers are updated automaticallly [for example, if you wish to reserve one server for testing new releases and not have changes made on it propagated back to the production server]).

Synchronisation transaction files are named with the current date and time to the microsecond, a journal sequence number which is incremented for each transaction generated during a given execution of the CGI application (to preserve transaction order in case the time does not advance between two consecutive transactions), and for easy examination of the synchronisation directory, the operation and path name, the latter with slashes translated to underscores. The contents of the transaction file is a version number, the operation, and the full path name.

Actual synchronisation is accomplished by a separate, stand-alone program, {\tt ClusterSync.pl}, which runs under group and user {\tt apache}, which is the owner of the {\tt ClusterSync} transaction directory and its contents. This program is started automatically from the {\tt init} script and runs as a daemon, saving its process ID in a {\tt ClusterSync.pid} file in the {\tt ClusterSync} directory.

When a synchronisation transaction is queued, the CGI program sends a {\tt SIGUSR1} signal to the {\tt ClusterSync.pl} process, which then traverses the server subdirectories, sorting the transactions into time and journal number order, and attempts to perform the operations they request. Synchronisation operations are performed by executing {\tt scp} and {\tt ssh} commands directed at the designated cluster host, which must be configured to permit public key access by user {\tt apache} without a password. If the synchronisation operation fails with a status indicating that the destination host is down or unreachable, the host is placed in a \verb+%failed_hosts+ hash with a timeout value of ten minutes from the time of failure. Synchronisation operations for that host will not be attempted until the timeout has expired, which prevents flailing away in vain trying to contact a down host over and over, possibly delaying synchronisation of other cluster members which are accessible. In the absence of a signal indicating newly-queued transactions, {\tt ClusterSync.pl} sweeps the transaction directory every five minutes to check for transactions queued for failed hosts which should now be retried due to expiry of the timeout.

All of the directory names, signal, and timeout values given above are specified by items in the ``Host System Properties'' section of the configuration; I have given the default settings, which should be suitable in most circumstances.

You can check whether two cluster hosts are synchronised by logging into one host, say {\tt server1}, and then running a command like:

    rdist -overify -P /usr/bin/ssh -c \
        /server/pub/hackdiet \
        server0:/server/pub/hackdiet

This will report any discrepancies between the database directory trees on the two servers. If the servers are synchronised, you should see only a ``need to update'' message for the {\tt ClusterSync/ClusterSync.pid}, plus any synchronisation transactions queued for failed servers awaiting retry. This operation is non-destructive and requires only read access to the database directory.

Monday, 2 July 2007

Bug: History records for CSV/XML data import

History records for CSV/XML import transactions were not being generated because the test for a read-only session was backwards---fixed.

Beta test concluded

The beta test program was concluded today with the release of version 1.0, including source code. An invitation code is no longer required to create an account, and the online documentation has been updated accordingly.

Bug: Convert trend to log unit when it differs from the unit of the most recent log

When creating a new monthly log, the code which fills in the trend carry-forward from the last trend value in the most recent existing log failed to convert the trend value from the log unit of that log to the log unit of the new one. This resulted in wild variances if a user changed the log unit and then entered data in a new log. I added unit conversion for this case, which was already handled correctly for the case of complete recalculation of trend carry-forwards (and hence can be used to correct any existing problems due to this bug). (Reported by Eric Carr.)

Security: Disable browsing public accounts from demonstration accounts

Added code to disallow browsing of publicly-visible account by users logged into read-only demonstration accounts. This restricts access to public accounts to those users who have gone to the trouble of creating an account of their own. This is enforced not only by removing the ``Browse public user accounts'' item from the Utilities menu for read-only accounts, but also aborting transactions ginned up from a read-only login with the transaction codes for public account access.

Bug: Embedded spaces in Palm CSV import

A Palm Eat Watch CSV record which included leading or trailing spaces in the Date, Weight, Rung, or Flag fields would be ignored or, in the case of the Flag field, interpreted incorrectly. I added code to discard all white space in these fields, as none should be present. (Reported by Reed Lipman.)

Feature: Feedback form in non-beta builds

Made the feedback form not conditional upon beta test (but left the test in the code, commented out). We'll leave the feedback form in for the nonce as we transition from beta to production. If spammers start creating accounts just to use the feedback form to send me junk mail, I'll turn it off, but for now and ideally from now on, it remains