Rekeying: Encryption
Overview
Key Rotation switched new writes to a new key but left existing data exactly where it was: still encrypted under the old one. This stage is phase 2 - sweeping already-persisted records onto the new key - and it uses mango4j-crypto's own production RekeyScheduler, not a hand-rolled loop. Facilitators should pair this with Rekeying: Encryption, and set expectations up front: this is a heavier API than anything earlier in the workshop - a background scheduled job, asynchronous completion, and a forced System.exit() - because that's genuinely what the production mechanism looks like, not a workshop simplification.
This stage comes as two projects:
starter/- what you work in. It compiles and runs, but the old key is never marked for retirement, soRekeySchedulerfinds nothing to do and the wait just times out after 10 seconds. Look for the// TODOcomment.complete/- the finished reference, where the scheduler actually sweeps all three records.
Follow along
cd stages/10-Rekeying-Encryption/starter
stages/10-Rekeying-Encryption/starter as its own project.
Triggering a rekey: mark the old key KEY_OFF
/**
* Marks a key for retirement: RekeyScheduler picks this up on its next
* cycle and sweeps every record still using it onto whichever key is
* "current" for this stage's shield.
*/
public void markForRetirement(String keyId) {
getById(keyId).setRekeyMode(CryptoKey.RekeyMode.KEY_OFF);
}
CryptoKey has a rekeyMode field for exactly this (KEY_ON/KEY_OFF, see the framework's CryptoKey javadoc). Marking a key KEY_OFF tells RekeyScheduler "sweep everything using this key onto whichever key is current, then tell me it's safe to delete."
Your turn: in starter/.../Main.java, mark "workshop-encryption-key" for retirement: provider.markForRetirement("workshop-encryption-key").
RekeyScheduler also supports the mirror image - marking a new key KEY_ON to pull every record onto it, regardless of which old key each one is currently on. This stage uses KEY_OFF instead, verified working end to end.
Wiring up a RekeyService
package ie.bitstep.mango.workshop;
import ie.bitstep.mango.crypto.core.domain.CryptoKey;
import ie.bitstep.mango.crypto.keyrotation.RekeyEvent;
import ie.bitstep.mango.crypto.keyrotation.RekeyService;
import java.util.List;
import java.util.concurrent.CountDownLatch;
/**
* The bridge between RekeyScheduler and this workshop's in-memory store.
* RekeyScheduler calls findRecordsUsingCryptoKey()/save() over and over,
* in batches, until an empty list comes back - so both "which records
* still need this?" methods have to inspect the record's *current* state
* each time, not a snapshot taken once.
*/
public class PaymentCardRekeyService implements RekeyService<PaymentCardEntity> {
private final PaymentCardStore store;
private final CountDownLatch rekeyFinishedLatch;
public PaymentCardRekeyService(PaymentCardStore store, CountDownLatch rekeyFinishedLatch) {
this.store = store;
this.rekeyFinishedLatch = rekeyFinishedLatch;
}
@Override
public Class<PaymentCardEntity> getEntityType() {
return PaymentCardEntity.class;
}
@Override
public List<PaymentCardEntity> findRecordsNotUsingCryptoKey(CryptoKey cryptoKey) {
return store.all().stream()
.filter(record -> !isOnKey(record, cryptoKey))
.toList();
}
@Override
public List<PaymentCardEntity> findRecordsUsingCryptoKey(CryptoKey cryptoKey) {
return store.all().stream()
.filter(record -> isOnKey(record, cryptoKey))
.toList();
}
@Override
public void save(List<?> records) {
// Our records are already updated in place (no serialization round
// trip in this workshop); a real RekeyService persists them here.
}
@Override
public void notify(RekeyEvent rekeyEvent) {
if (rekeyEvent.getType() == RekeyEvent.Type.REKEY_FINISHED) {
rekeyFinishedLatch.countDown();
}
}
private static boolean isOnKey(PaymentCardEntity record, CryptoKey cryptoKey) {
return record.getEncryptedData().contains("\"cryptoKeyId\":\"" + cryptoKey.getId() + "\"");
}
}
RekeyScheduler doesn't know how to query your store - RekeyService<T> is the interface it calls to find records and save them back. findRecordsUsingCryptoKey() is called repeatedly, in batches, until it returns empty, so it has to check each record's current state every time, not a snapshot taken once. notify() is the only way to know the job finished - there's no return value to wait on, since the scheduler runs asynchronously.
Wiring up key deletion
package ie.bitstep.mango.workshop;
import ie.bitstep.mango.crypto.core.domain.CryptoKey;
import ie.bitstep.mango.crypto.keyrotation.RekeyCryptoKeyManager;
import java.util.concurrent.CountDownLatch;
/**
* RekeyScheduler calls this once it's confirmed nothing is using a
* KEY_OFF-marked key anymore - the automatic version of what earlier
* stages had to check for manually before removing an old key.
*/
public class RetiringKeyManager implements RekeyCryptoKeyManager {
private final CountDownLatch keyRetiredLatch;
public RetiringKeyManager(CountDownLatch keyRetiredLatch) {
this.keyRetiredLatch = keyRetiredLatch;
}
@Override
public void markKeyForDeletion(CryptoKey retiredKey) {
InMemoryCryptoKeyProvider.remove(retiredKey.getId());
keyRetiredLatch.countDown();
}
}
Once every record using the retired key has been swept, RekeyScheduler calls RekeyCryptoKeyManager.markKeyForDeletion() on its own - the manual "has the sweep reached everything yet?" bookkeeping earlier versions of this stage did by hand is now the framework's job.
Starting the scheduler
RekeySchedulerConfig config = RekeySchedulerConfig.builder()
.withCryptoShield(cryptoShield)
.withRekeyServices(List.of(new PaymentCardRekeyService(store, rekeyFinishedLatch)))
.withRekeyCryptoKeyManager(new RetiringKeyManager(keyRetiredLatch))
.withObjectMapper(new ObjectMapper())
.withClock(Clock.systemUTC())
.withCryptoKeyCachePeriod(Duration.ZERO)
.withRekeyCheckInterval(0, 1, TimeUnit.SECONDS)
.build();
// RekeyScheduler polls on its own background thread with no
// shutdown()/close() method exposed anywhere - once started, it
// never stops on its own. See System.exit() below.
new RekeyScheduler(config);
RekeySchedulerConfig is a real production configuration surface - cache duration (how long application instances might still be caching old key data, to avoid rekeying onto a key some instances don't know about yet), batch interval, failure tolerance, and the poll interval itself. This stage sets a 1-second poll interval and zero cache duration purely to make the demo finish fast; a real deployment would set these to match its own caching and traffic patterns.
Waiting for an asynchronous job to finish
boolean rekeyed = rekeyFinishedLatch.await(10, TimeUnit.SECONDS);
boolean retired = rekeyed && keyRetiredLatch.await(10, TimeUnit.SECONDS);
System.out.println("rekey finished within timeout? " + rekeyed);
System.out.println("old key retired within timeout? " + retired);
for (PaymentCardEntity record : store.all()) {
System.out.println(" " + record.getEncryptedData());
}
Every previous stage's "sweep" was a synchronous method call: RekeySweep.run(...) returned once it was done. RekeyScheduler doesn't work that way - it's a background job polling on its own schedule, so Main has no method call to block on. PaymentCardRekeyService.notify() and RetiringKeyManager.markKeyForDeletion() are the only signals available, so this stage uses two CountDownLatches to turn "the scheduler will eventually tell us, asynchronously" into something a single-shot Main can wait on with a timeout.
RekeyScheduler also has no shutdown()/close() method - its background thread pool is non-daemon and polls forever once started, so nothing in Main returning normally would actually end the process. The System.exit(0) at the very end of main() is there on purpose, not a shortcut.
Running it
starter/ before your change: nothing is marked for retirement, so every scheduler cycle logs "No re-keying needed" and does nothing. Both waits time out after 10 seconds each:
rekey finished within timeout? false
old key retired within timeout? false
{"cryptoKeyId":"workshop-encryption-key",...}
{"cryptoKeyId":"workshop-encryption-key",...}
{"cryptoKeyId":"workshop-encryption-key",...}
After your change, the scheduler finds the marked key on its very first cycle:
rekey finished within timeout? true
old key retired within timeout? true
{"cryptoKeyId":"workshop-encryption-key-v2",...}
{"cryptoKeyId":"workshop-encryption-key-v2",...}
{"cryptoKeyId":"workshop-encryption-key-v2",...}
You'll also see RekeyScheduler's own log lines above those - "Re-key complete", "All records (3) using deprecated encryption key ... have been keyed onto the current encryption key ...", and "Notifying the application to mark the following Crypto key as deleted". A recurring "0 HMAC keys were found ... skipping" line is expected noise: the scheduler always checks for HMAC rekey work too, on every cycle, whether or not this stage's entity uses any.