Monday, 10 March 2008

Feature: Recovery from transient cluster synchronisation failures

Implemented recovery from transient and permanent cluster synchronisation transaction failures. Previously, if any error occurred reading, verifying, or executing a cluster synchronisation transaction, the {\tt ClusterSync.pl} program would crash, suspending cluster synchronisation until it was restarted. Unfortunately, there were a number of circumstances in which such errors could occur, the most common being cases where a race condition between queueing the transaction and {\tt ClusterSync}'s processing of it caused an incomplete file to be read (transient), and those where a crash of the process queueing the transaction caused an incomplete file to be written to the transaction directory (persistent).

When a cluster sync transaction fails, for whatever reason, it is placed into a failed transaction hash whose key is the transaction file name and whose value is an array containing the number of times the transaction has been tried and the next time the transaction should be retried. On subsequent passes through the transaction directory, failed transactions are skipped unless their retry time has arrived, whereupon they are retried and, if they fail, their try count is incremented and the next attempt count updated.

If the transaction eventually succeeds, it is closed out normally and removed from the failed transaction hash. If the transaction fails again, its try count is increment and if it has reached the limit, the transaction is deleted from the transaction directory and the failed transaction hash. Failre to delete the transaction from the transaction directory remains fatal to the {\tt ClusterSync} program.

The intervals between retries of a failed transaction and the number of failures which cause a transaction to be abandoned are set by configuration parameters.

No comments: