Field notes from Santala Research after running this setup on the sapphire-1 and pearl-1 testnets (August to September 2026) and validating the gno.land Horcrux fork against stock gnoland v1.5.0 on onyx-1 (September 2026). Written for operators who already know how to run a plain gnoland node. Everything here is what we actually ran; the mistakes are ours too.

Why bother

A plain validator keeps priv_validator_key.json on one box. Whoever gets that box can sign as you; and if you ever start a second copy “just in case”, you double-sign. Horcrux fixes both: the key is split into 3 shards on 3 machines, any 2 must agree to sign, no machine ever holds the full key, and the cosigners share a sign-state through Raft so a stale node cannot make them sign a lower height. We ran it with gno.land's built-in tmkms listener, which speaks the upstream privval protocol Horcrux expects.

Architecture we ran

                  internet
                     │
              ┌──────┴──────┐
              │  sentry(s)  │  public p2p, private_peer_ids = validator
              └──────┬──────┘
          WireGuard  │  (validator: pex=false, persistent_peers = sentries only)
              ┌──────┴──────┐  tmkms_listener on the WireGuard address, port 5556
              │  validator  │◄────────────── cosigner #1 (same host or another)
              └─────────────┘◄──── WireGuard ── cosigner #2
                              ◄──── WireGuard ── cosigner #3
  • All cosigner-to-cosigner (Raft, gRPC) and cosigner-to-validator (signer) traffic goes over a private WireGuard mesh; the host firewall allows those ports only on the wg0 interface.
  • Cosigners are tiny (about 20 MB RAM, negligible CPU). They can share a host with a sentry or a monitoring box; one shard alone cannot sign.
  • Put the validator within about 15 ms of at least two cosigners. Every vote is a round trip through the threshold signer. We measured roughly 350 ms per vote with the signing node in Asia and the cosigners in Europe, and about 50 ms once co-located. That difference decides whether you make the commit on a fast round.

0. Prerequisites: use gno.land's Horcrux fork, not upstream

  • Use gnolang/horcrux (tag v3.3.2-gno.4 or newer), a hardened fork of strangelove's archived v3.3.2 maintained by the gno core team. Upstream v3.3.2 has two compatibility problems with gno.land's tmkms listener and, worse, critical security bugs that the fork fixes: a nonce-handling flaw that could expose key material, a residual double-sign path, and an unauthenticated cluster admin surface. Do not run upstream on a validator.
  • Three fork features make stock gnoland (v1.5.0 and later) work without any patch:
    1. horcrux create-conn-key gives each cosigner a persistent connection identity, so you can pin its hex pubkey in the node's allowed_kms_pubkeys. Upstream regenerates the identity on every start, which cannot be pinned.
    2. thresholdMode.leaderOnlyChainNodeConnections: true: only the Raft leader dials the node. gno.land's listener holds a single signer slot and churns if all cosigners dial.
    3. A spec-compliant sign response (the full vote echoed back), which gno.land validates strictly.
  • Build the fork with cgo enabled (CGO_ENABLED=1, needs gcc). The ECIES cosigner keys use secp256k1, and a static build panics with ScalarMult is not available when secp256k1 is built without cgo.
git clone --branch v3.3.2-gno.4 --depth 1 https://github.com/gnolang/horcrux.git && cd horcrux
CGO_ENABLED=1 go build -o /usr/local/bin/horcrux ./cmd/horcrux

gnoland itself: build from the release tag exactly as the network docs say, record the SHA-256, no patches.

1. Create the key and the shards (once, on the validator host)

gnoland secrets init -data-dir ./gnoland-data   # priv_validator_key.json, priv_validator_state.json, node_key.json

Gotcha (v1.5.0): secrets init writes the three files into the data-dir root, while the node reads them from data-dir/secrets/. Move them:

mkdir -p gnoland-data/secrets && mv gnoland-data/{priv_validator_key,priv_validator_state,node_key}.json gnoland-data/secrets/
gnoland secrets get validator_key -data-dir gnoland-data/secrets   # point at secrets/, not the data-dir, or it panics

Back the key up now, encrypted, off the host (we use openssl enc -aes-256-cbc -pbkdf2, passphrase on paper; test the decrypt on another machine).

Gotcha: Horcrux reads the CometBFT key format, not gno.land's ("address": "g1…", "@type": "/tm.PubKeyEd25519"), so create-ed25519-shards fails with encoding/hex: invalid byte: 'g'. Convert to a temporary CometBFT-format file first. The key bytes are identical; only the envelope differs:

python3 - <<'PY'
import json,base64,hashlib
d=json.load(open("gnoland-data/secrets/priv_validator_key.json"))
pub=base64.b64decode(d["pub_key"]["value"])
json.dump({"address":hashlib.sha256(pub).digest()[:20].hex().upper(),
           "pub_key":{"type":"tendermint/PubKeyEd25519","value":d["pub_key"]["value"]},
           "priv_key":{"type":"tendermint/PrivKeyEd25519","value":d["priv_key"]["value"]}},
          open("pvk-cometbft.json","w"))
PY
chmod 600 pvk-cometbft.json
horcrux create-ed25519-shards --chain-id <chain-id> --key-file pvk-cometbft.json --threshold 2 --shards 3 --out ./shards
horcrux create-ecies-shards --shards 3 --out ./ecies      # one ecies_keys.json per cosigner
shred -u pvk-cometbft.json

You get shards/cosigner1..3/<chain-id>_shard.json. Copy each to its cosigner over SSH (never chat or email), then delete the full key from every host:

shred -u gnoland-data/secrets/priv_validator_key.json

The node no longer needs it: with the tmkms listener enabled it never reads a local key.

Shard files are named per chain (<chain-id>_shard.json). Moving to a new testnet with the same key is a file copy, plus a Raft wipe (see section 5).

2. Cosigner setup (each of the 3 hosts)

On each cosigner host, as the cosigner user:

horcrux create-conn-key --home ~/.horcrux     # prints a 64-hex public key; collect all three for the node's allowlist

Then ~/.horcrux/config.yaml:

signMode: threshold
thresholdMode:
  threshold: 2
  cosigners:
  - shardID: 1
    p2pAddr: tcp://<wg-ip-of-cosigner-1>:2222
  - shardID: 2
    p2pAddr: tcp://<wg-ip-of-cosigner-2>:2223
  - shardID: 3
    p2pAddr: tcp://<wg-ip-of-cosigner-3>:2224
  grpcTimeout: 2000ms
  raftTimeout: 2000ms
  leaderOnlyChainNodeConnections: true     # required for gno.land
chainNodes:
- privValAddr: tcp://<wg-ip-of-validator>:5556
connKeyFile: conn_key.json                 # the persistent identity from create-conn-key

Plus, per host: its own <chain-id>_shard.json and ecies_keys.json. Permissions matter: the home and state/ directories must be 0700. A 0600 directory makes the cosigner fail its own sign-state check, and the node reports remote signer error … checking file existence. Run it as a dedicated user under systemd with Restart=always and a memory cap.

Firewall: allow 2222–2224 (Raft/gRPC) and 5556 (signer) only on wg0. Nothing else on the public interface except SSH and WireGuard UDP.

3. Validator config (tmkms listener)

In config.toml:

[consensus.priv_validator.tmkms_listener]
listen_addr = "tcp://<wg-ip-of-validator>:5556"   # WireGuard address only
chain_id = "<chain-id>"
allowed_kms_pubkeys = ["<hex-conn-pubkey-cosigner-1>", "<hex-conn-pubkey-cosigner-2>", "<hex-conn-pubkey-cosigner-3>"]
protocol_version = "v0.34"

Set them with gnoland config set in this order, chain_id and allowed_kms_pubkeys first and listen_addr last, or validation rejects the intermediate state.

Start the node, then start the three cosigners together. Followers log Not the cluster leader, deferring connection to chain node; the leader logs Connected to Sentry; the node obtains its validator pubkey through the signer and, once in the valset, logs Signed and pushed vote every round. You can delete the local priv_validator_key.json now, the node no longer reads it. Check /status on the node: validator_info.address must be your validator's address even though the host has no key file.

4. Verify on chain, not in logs

Logs say “signed”; only the chain says “counted”. Check that your validator address appears in last_commit.precommits of recent blocks:

for h in $(seq $((TIP-12)) $((TIP-1))); do
  curl -s "$RPC/block?height=$h" | jq -r --arg a "$ADDR" \
    '[.result.block.last_commit.precommits[]? | select(. != null) | .validator_address] | index($a) != null'
done

On pearl-1 this is how we found that a node reporting 12 sign events per minute was landing 0 of 12 on chain. It was one block behind and precommitting nil. Nothing in its own logs said so.

5. Things that bit us (in order of pain)

  1. Raft remembers addresses and chain-ids. Changing a cosigner's address, moving to a new chain, or renaming a shard file without wiping ~/.horcrux/raft/ on all cosigners leaves the cluster arguing with ghosts (“Error loading sign state during raft replication”). Procedure: stop all three, rm -rf ~/.horcrux/raft on each, fix configs, start all three together.
  2. Two live signing nodes make each other slower, not safer. We tried a second gnoland pointing at the same cosigners for redundancy. Horcrux does prevent double-signing, but the two nodes race and produce constant step regression errors, adding latency. Use a cold standby instead (data synced, listener disabled) and switch by hand.
  3. Never use the signer port as a liveness probe. Our alerting TCP-connected to 5556 every minute; the listener accepts one connection at a time, the accept backlog filled to 4096, and we got false “OFFLINE” alerts for days. Probe p2p or WireGuard instead.
  4. A validator can never be more current than the peers feeding it. Sentries on shared vCPU executed heavy blocks slower than our validator, so the validator sat one block behind and missed every commit under load, with 92% of one core idle. Size sentries at least as fast as the validator, or keep a few outbound-only persistent peers to the network's core nodes as a safety net.
  5. Sentries lock out their own validator when their inbound table is full. Keep max_num_inbound_peers headroom on sentries and list the validator in the sentry's persistent_peers so the sentry dials the validator (sentry-initiated dials do not depend on free inbound slots).
  6. Provider images can hurt you. A stock hosting image ran echo 1 > /proc/sys/vm/drop_caches and fstrim from /etc/cron.hourly. We missed blocks at minute 17 of every hour until we found it. Audit cron on every new host.
  7. Geography is latency. See “Why bother”: keep the signing node next to at least two cosigners.
  8. Log level. At info a public sentry writes millions of lines per hour; a validator's GC pressure came mostly from logging and peer gossip, not from the VM. Use warn on sentries, and profile before tuning.

6. Failover, the safe way

  1. Declare the incident; alerting stays on.
  2. Prove the old node is dead, not unreachable: stop gnoland and its cosigner, or power the host off from the provider console. Confirm no precommit from you in the next commits and no signer connection on the remaining cosigners.
  3. The cosigners' shared Raft sign-state already holds the last signed height, round and step; on the new node, carry priv_validator_state.json forward or write one at the current height.
  4. Only now point the cosigners' chainNodes at the new node's listener and start it.
  5. Verify /status and the on-chain check in section 4.

Our real one (a host migration on 2026-09-16) took about three minutes without signing and no double-sign.

7. Monitoring notes specific to gno.land

gno.land does not expose a CometBFT-style /metrics endpoint. The node pushes OTLP (the [telemetry] section) to a collector; missed-block and validator-set metrics come from RPC-based tools such as samouraiworld/gnomonitoring (Prometheus exporter and alerts) and gnoverse/gnockpit (live valset dashboard). Alert on your own misses (blocks where at least two thirds of the set signed and you did not), not on network-wide drops.

Questions or corrections: find us on the gno.land Discord (Santala-Research).

Keep going: why a single key on a single machine is the weak point of any setup is covered from first principles in Wallets, keys and seed phrases, and how validators and finality fit together in Proof of Work vs Proof of Stake. Our longer investigations live under Research.

Frequently asked questions

What does Horcrux actually protect against? Two things: key theft, because no single machine ever holds the full validator key, and accidental double-signing, because the cosigners keep a shared sign-state in Raft and refuse to sign a height, round or step they have already signed.

Do I need to patch gnoland to use Horcrux? No. Use the gno core team's fork, gnolang/horcrux v3.3.2-gno.4 or newer, with a stock gnoland v1.5.0 or later. The fork adds a persistent connection key you can pin in allowed_kms_pubkeys, a leader-only connection mode, and a spec-compliant sign response, which together remove the need for any node-side changes.

Why not run upstream strangelove Horcrux v3.3.2? Besides being incompatible with gno.land's tmkms listener, upstream v3.3.2 is archived and carries security bugs that the gno.land fork fixes: a nonce-handling flaw that could expose key material, a residual double-sign path, and an unauthenticated cluster admin surface.

Why does create-ed25519-shards fail with invalid byte 'g'? Horcrux expects the CometBFT key file format, while gno.land writes a bech32 address starting with g1 and a different type envelope. Convert the file to CometBFT format first, as shown in section 1; the key bytes themselves are identical.

How do I know my validator is really signing? Do not trust the signer's logs. Query recent blocks over RPC and check that your validator address appears in last_commit.precommits. A node can log a successful signature every round while its votes never make it into a commit.