Monday, 10 March 2008
Feature: Recovery from transient cluster synchronisation failures
Implemented recovery from transient and permanent cluster synchronisation
transaction failures. Previously, if any error occurred reading, verifying, or
executing a cluster synchronisation transaction, the {\tt ClusterSync.pl}
program would crash, suspending cluster synchronisation until it was
restarted. Unfortunately, there were a number of circumstances in which
such errors could occur, the most common being cases where a race condition
between queueing the transaction and {\tt ClusterSync}'s processing of it
caused an incomplete file to be read (transient), and those where a crash
of the process queueing the transaction caused an incomplete file to be
written to the transaction directory (persistent).
When a cluster sync transaction fails, for whatever reason, it is placed
into a failed transaction hash whose key is the transaction file name
and whose value is an array containing the number of times the
transaction has been tried and the next time the transaction
should be retried. On subsequent passes through the transaction
directory, failed transactions are skipped unless their retry time
has arrived, whereupon they are retried and, if they fail, their
try count is incremented and the next attempt count updated.
If the transaction eventually succeeds, it is closed out normally
and removed from the failed transaction hash. If the transaction
fails again, its try count is increment and if it has reached
the limit, the transaction is deleted from the transaction
directory and the failed transaction hash. Failre to delete
the transaction from the transaction directory remains fatal to the
{\tt ClusterSync} program.
The intervals between retries of a failed transaction and the number of
failures which cause a transaction to be abandoned are set by configuration
parameters.
Subscribe to:
Post Comments (Atom)
No comments:
Post a Comment