This will be the story of another bug in Ethereum, this time an actual protocol bug and not a client’s one. I’ll interleave it with the history of how I got involved in Ethereum core protocol development as this bug is intimately intertwined with my interests in blockchain. The account of events told here is rather a reflection of how I remember them and not how they actually happened. My memory is not a shade of what it used to be a couple of decades ago.

Devconnect, Amsterdam, 2022

It was April 2022. Raúl Jordan had bought me direct flight tickets to Amsterdam. I arrived slightly late for the opening of the core dev workshop. I had never met anyone in person, wasn’t on anyone’s list so had to start messaging anyone I knew in the organization to please let me in. I don’t remember if it was Tim Beiko or Hsiao-Wei Wang that allowed me in. I was pretty uncomfortable, I had gained about 30Kg, or almost 50% of my body mass, during the last couple of years in the pandemic, since I traded about 20 hours a week sitting on a bike saddle to sitting in front of a computer. I did not have any appropriate clothing for the weather and I had to buy new pants for the trip. As I was getting in, a small crowd was gathered around a group of postdocs/grad students from Stanford, they were accompanying David Tse. I am uncertain of who was presenting at the time, I think one of them was Joaquim Neu. They were explaining how an arbitrary validator could cause some reorgs at epoch boundaries because of unexplored rough boundaries between Casper, the friendly finality gadget (FFG) and Latest Message Driven Greedy Heaviest Observed SubTree (LMD-GHOST) . As they are making their exposition someone stands up and grabs the marker (it was later pointed out to me that it was Michael Sproul with whom I had had tons of online interactions and it would later be years to meet again in person). Michael goes “Oh yeah, we’re also aware of this other type of reorg” and proceeds to explain how any validator could reorg about 11 slots (1/3 of an epoch) every single time they were the first one proposing in the epoch. Jaws dropped. I could not believe this. Michael was surprised this wasn’t known to more of the audience. This seemed to have been a bug that himself and Paul Hauner from Lighthouse had known for a while now.

It was decided that a small group would gather later to try to discuss possible solutions to this problem. I had conveniently finished rewriting all of Prysm’s forkchoice from scratch in the previous weeks, so I had this rather niche topic fresh in my mind and it made sense for me to participate. Prysm was back then by far the largest CL client by usage, and an active client diversity campaign was in place to reduce its usage. A bug in forkchoice could create a split in which Prysm could be responsible for finalizing a bad chain—this did not let people sleep back then, ahead of the merge.

Between sessions, I met Vitalik Buterin for the first time. Back then I was trying to find some interesting problem to work on. I think officially I was an EPF fellow in the pilot cohort organized by Piper Merriam. That cohort mostly worked on the portal network, understandably, or on EVM fuzzing. I was thinking about enshrined Proposer-Builder separation (ePBS) since Vitalik had just posted a two-slot approach to ePBS and Justin Drake had written on the topic before. The biggest drawback of Vitalik’s approach was that it would reduce Ethereum’s throughput in half because of the two rounds of voting needed per execution payload. I had in mind an interleaved slot pipelining that ended up being the basis for EIP-7732 which is scheduled for inclusion in the Glamsterdam Fork. Danny Ryan had politely dismissed the idea while encouraging me to continue working on it. Vitalik on the other hand was brutally honest as only himself can be: it took him 30 milliseconds to destroy the core of the idea and back to the drawing board. I met Barnabé Monot during that conversation, he was also interested in similar mechanisms.

The whole workshop promised to be quite an interesting place to meet people, surprisingly not so different than the much smaller workshops of long-time-friends that I am used to from mathematics. I saw an interesting exchange of slasher T-shirts between Raúl Jordan and Michael Sproul: they had written the (still) unique two implementations of Ethereum slashers, and the detail of them sharing a T-shirt was touching. Those shirts should be now some valuable piece of Ethereum memorabilia, perhaps they are already coined somewhere.

Getting back to the forkchoice bug. I joined Aditya Asgaonkar and Paul Hauner in a room for the remainder of the next couple of days. I don’t remember doing anything else (although there’s physical proof that I did attend other talks and I did have drinks and made friends with my team). From time to time some people entered that room—I also met Francesco D’Amato there—and Danny Ryan would pass by between his numerous other tasks to attend. At that time the consensus layer was a benevolent dictatorship and nothing got in without his scrutiny.

The Bug

There are numerous detailed descriptions of the bug that was surreptitiously fixed by coordination of client devs. The account here will focus less on the technical details and more on the historical ones from my (very narrow) point of view. To understand the bug, you need to know that Ethereum’s attestations take care of two conceptually different things at once:

  • Finality.
  • Forkchoice.

Forkchoice

Forkchoice is the mechanism that lets honest nodes chose their view of the head of the chain. For example if blocks are coming in regularly:

graph RL F["92"]:::green E["..."]:::lightblue C["65"]:::lightblue B["64"]:::orange D["..."]:::lightblue G["32"]:::red H["..."]:::lightblue F --> E E --> C C --> B B --> D D --> G G --> H classDef lightblue fill:#ADD8E6 classDef green fill:#90EE90 classDef orange fill:#FFA500 classDef red fill:#FF0000

then the head of the chain is simple to guess. Here, during slot 60, this node’s head is the block marked in green at slot 92. The block marked in orange is its checkpoint (if you need a refresher on checkpoints, take a look at my last post about Prysm’s bug during the Fulu fork). The block marked in red here is another checkpoint, this would be the justified checkpoint typically and we will talk more about them below. If the next proposer, during slot 93, were to propose

graph RL I["93"]:::white_dashed F["92"]:::green E["..."]:::lightblue C["65"]:::lightblue B["64"]:::orange D["..."]:::lightblue G["32"]:::red H["..."]:::lightblue F --> E I --> E I ~~~ F E --> C C --> B B --> D D --> G G --> H classDef white_dashed fill:#FFFFFF,stroke:#5F9EA0,stroke-width:2px,stroke-dasharray: 5 5 classDef lightblue fill:#ADD8E6 classDef green fill:#90EE90 classDef orange fill:#FFA500 classDef red fill:#FF0000

The head of the chain would remain being 92 and the next proposer will simply continue advancing the chain as follows:

graph RL J["94"]:::green I["93"]:::white_dashed F["92"]:::lightblue E["..."]:::lightblue C["65"]:::lightblue B["64"]:::orange D["..."]:::lightblue G["32"]:::red H["..."]:::lightblue I --> E I ~~~ F J --> F J ~~~ I F --> E E --> C C --> B B --> D D --> G G --> H classDef white_dashed fill:#FFFFFF,stroke:#5F9EA0,stroke-width:2px,stroke-dasharray: 5 5 classDef lightblue fill:#ADD8E6 classDef green fill:#90EE90 classDef orange fill:#FFA500 classDef red fill:#FF0000

Finality

I won’t go here over any formal definition of finalization in Ethereum. The above cited article on Casper is a great resource. It suffices to say that there are some checkpoints that the node tracks. Some are justified and some are finalized. The following rules are more than what we need to know about them at this moment:

  • A node is able to track contending branches and decide among them. In the example above, the honest node has synced blocks at slots 93 and 94. The branch containing 92 and 94 diverges from that containing 93. We say this is a fork and the node is able to sync blocks on either branch and track them both, while knowing that in this case, the canonical chain is that containing 94 and not 93. When the canonical chain switches from one branch to a contending one, we say that a reorg happened.
  • A node is also able to sync and track contending justified checkpoints. However, nodes always choose their head chain descending from their latest justified checkpoint. It suffices to say that this rule makes it very very hard for a node to actually reorg a justified checkpoint. Such a situation would necessarily require asynchrony in the network and multiple messages being delayed.
  • A node will only track one latest finalized checkpoint and it will never reorg it. Even if 100% validators were dishonest and were to finalize contending branches, honest nodes will never even see the second branch. Honest nodes only consider blocks that descend from their latest finalized checkpoint and will never even sync, gossip or even consider a block that does not descend from the latest finalized checkpoint.

When thinking about finalization it is important to consider the difference between information that is on-chain and information that is off-chain. When choosing between the possible heads at slots 93 or 92 in the above example, the proposer of 94 had to count attestations for both branches to decide what is its head view. Those attestations aren’t necessarily included in blocks (in particular any attestation for block 93 couldn’t have been included in any block yet) and different branches will have blocks that included different sets of attestations. To choose between branches to decide the head of the chain, the node considers all attestations it has seen (I’m neglecting equivocations) regardless of whether these attestations were or not included in blocks. We say that the node considers off-chain information to choose its head.

On the other hand, finalization is entirely on-chain. Only attestations included in blocks are considered, and from the point of view of the chain, any block that is not part of the canonical chain, might as well not exist. So, if any attestation was included in block 93 but not in any of the blocks of the canonical chain containing 94 above, then those attestations would not count for finalization in the canonical chain.

Attestations indicate a head and a checkpoint (the green and the orange blocks above). During an epoch, nodes keep count of attestations that were included in blocks and, if during the epoch 2/3 of the validators (counted by stake) have attested for the same checkpoint, that checkpoint is considered justified. In the typical situation, this would upgrade the last justified checkpoint into a finalized one.

Crucially, this process happens only at epoch boundaries. Even though enough votes are typically already obtained after the first 2/3 of the blocks in the epoch have been processed.

So for example, after block 95 has been processed, and nodes move onto slot 96, (the first slot of Epoch 3), we have the following situation

graph RL L["95"]:::green_dashed K["95"]:::lightblue J["94"]:::lightblue I["93"]:::white_dashed F["92"]:::lightblue E["..."]:::lightblue C["65"]:::lightblue B["64"]:::red D["..."]:::lightblue G["32"]:::black H["..."]:::lightblue L -.-> K K --> J I --> E I ~~~ F J --> F J ~~~ I F --> E E --> C C --> B B --> D D --> G G --> H classDef white_dashed fill:#FFFFFF,stroke:#5F9EA0,stroke-width:2px,stroke-dasharray: 5 5 classDef green_dashed fill:#90EE90,stroke:#008000,stroke-width:2px,stroke-dasharray: 5 5 classDef lightblue fill:#ADD8E6 classDef green fill:#90EE90 classDef orange fill:#FFA500 classDef red fill:#FF0000 classDef black fill:#a0a0a0

Even though block 96 hasn’t arrived yet, the head of the chain is still block 95, but nodes have processed to slot 96, the justified checkpoint is now marked red (it was the previous orange checkpoint) and the previous justified checkpoint is now in gray denoting it is finalized. Notice however that this information about justification is not yet on-chain! there is no block in Epoch 3 that commits to this justification. When block 96 arrives on top of 95 this information will be committed.

Typically however, when processing blocks, the chain is working normally at 98% of participation, thus by slot 22 into the epoch, there are already enough votes to justify the current checkpoint:

graph RL L["95"]:::green_dashed K["95"]:::lightblue E["..."]:::lightblue O["This block has already
enough votes to
justify 64"] M["86"]:::lightblue N["..."]:::lightblue C["65"]:::lightblue B["64"]:::red D["..."]:::lightblue G["32"]:::black H["..."]:::lightblue O ~~~ H O --> M L -.-> K K --> E E --> M M --> N N --> C C --> B B --> D D --> G G --> H classDef white_dashed fill:#FFFFFF,stroke:#5F9EA0,stroke-width:2px,stroke-dasharray: 5 5 classDef green_dashed fill:#90EE90,stroke:#008000,stroke-width:2px,stroke-dasharray: 5 5 classDef lightblue fill:#ADD8E6 classDef green fill:#90EE90 classDef orange fill:#FFA500 classDef red fill:#FF0000 classDef black fill:#a0a0a0

This means that any chain that contains the block in slot 86, will justify, in the epoch transition to Epoch 3, the checkpoint at slot 64.

We call this unrealized justification, that is, at slot 86, there is already enough information on chain to justify the current checkpoint 64, we know for a fact that regardless what happens on chain, as long as the chain advances from descendants of this block, the checkpoint at 64 will be justified when we do the epoch transition into slot 96.

This is the crucial flaw that could be exploited by the proposer of 96, if they decide to build on top of 86 as follows:

graph RL L["96"]:::green K["95"]:::lightblue E["..."]:::lightblue M["86"]:::lightblue N["..."]:::lightblue C["65"]:::lightblue B["64"]:::red D["..."]:::lightblue G["32"]:::black H["..."]:::lightblue L --> M L ~~~ K K --> E E --> M M --> N N --> C C --> B B --> D D --> G G --> H classDef white_dashed fill:#FFFFFF,stroke:#5F9EA0,stroke-width:2px,stroke-dasharray: 5 5 classDef green_dashed fill:#90EE90,stroke:#008000,stroke-width:2px,stroke-dasharray: 5 5 classDef lightblue fill:#ADD8E6 classDef green fill:#90EE90 classDef orange fill:#FFA500 classDef red fill:#FF0000 classDef black fill:#a0a0a0

The whole chain from 87 until 95 would be reorged! This is because that chain never realized the justification of slot 64. That chain did have the votes, but it was never continued into Epoch 3. The proposer of 96 however, has produced a block whose associated beacon state does carry the justified checkpoint of slot 64. Thus honest nodes, will only consider for head nodes descending from 96 because it carries a higher justification than that of 95 (which still has 32 as its justified checkpoint). This bug therefore allowed the proposer of 96 to reorg about 11 slots trivially.

Pulled tips

During those closed doors sessions in Amsterdam, the fix that was devised was to simply pull all tips from previous epochs. That is, when the attacker’s block like 96 arrived, we would also process the states coming from all other possible branches (in this case the branch containing 95). As this branch, when advancing, also contains block 86, it would also justify checkpoint 64, and all the votes in this branch would make it win over the attacker’s block 96.

Many different approaches to pulling tips were discussed and analyzed, all with their pros and cons, but we agreed that the general structure of the fix would come along these lines. Aditya started working on a private repository with spec changes, testvectors, multiple options, etc. Clients were rushing to fix this ahead of the merge as this was considered a merge blocker.

Pre Genesis

During the COVID-19 pandemic I started reading about crypto, something that I had put out because of lack of time many times throughout the years. Bought some bitcoin and some Ether. I knew enough not to trust any custodian, and ran my own nodes at home. I thought to myself that a good way of forcing myself not to sell in the market swings would be to deposit in the ETH2 staking contract as soon as it was deployed. But then the Medalla incident happened. I was already committed to my plan and I am stubborn, but the more I read about the people in charge of this merge the more I realized that they were a bunch of kids playing with my money (and billions from others). In retrospect, it was quite irresponsible of me to commit a sizeable chunk of our savings to crypto (we had just bought an apartment and had essentially no cash), but anyway, I was already in that boat. I got into Ethstaker’s Discord and through them I landed in the Prysm Discord server. Discussions there were technical and the whole team was 24/7 directly helping users and getting contributions from power users. As a user testing their software and getting my feedback directly taken into account, it was an empowering feeling.

My first PR on Prysm landed in October 2020. I had coded since 5 years old, but had never written nor read a word of golang. Of course my PR broke something, in this case the Windows runtime. Preston Van Loon reverted it immediately but instead of being mad at me, he tipped me on the fix for Windows the next week. Looking at my commit history it seems I did a bunch of work back then—I landed 14 commits pre-genesis. I don’t know why my starting date in Protocol Guild is June 2021; I should ping Trenton van Eps.

I was originally mostly interested in learning about the protocol, monitoring my node and making sure that I didn’t lose money. By the end of 2021 Terence Tsao contacted me and offered me to be part of the team, no strings attached and no time commitments. Had a call with Raúl and I was onboard. This opened an entire new world for me. On the one hand I had access to the internal parts of the server, where all the actual technical aspects of the client were being discussed. On the other hand it allowed me to code on more serious aspects of the client instead of simply monitoring tools.

I took over rewriting the implementation of the forkchoice tree. One of the problems of having an executable spec, is that implementation details are leaked from it. The forkchoice tree is, no surprise here, a tree, with the finalized node at its root and every block that the beacon client knows about is a node in this tree. We mentioned above that nodes do not even consider blocks that do not descend from the finalized checkpoint. Otherwise we would have a forkchoice forest. However, all clients implemented this tree structure as an array, in fact called a protoarray in honour of protolambda that implemented it this way in the spec repo. This didn’t make any sense to me, trees are more naturally implemented as linked lists, and in this case as a doubly linked one, with each node tracking its parent and all children.

I already mentioned that I was fortunate to have had rewritten forkchoice from scratch at that time, since I had it very fresh to be involved in the pulled tips discussions. But the very first exposure to protocol bugs I had was with regards to balancing attacks. I am not going to explain that one, but just give the historical context relevant to the pulled tips scenario. In order to prevent these kinds of attacks, and certain ex-ante reorgs, core devs decided to change forkchoice by adding a boost to the proposer of a timely block. This, being a forkchoice change, is inherently not a hard fork. We naively thought there wasn’t a strong need to be closely coordinated on when we would release this change. So clients worked these changes silently on hidden branches and released them at roughly the same time.

Problem is, roughly is not exactly. Thus, about 25% of nodes were not boosted as node operators were still upgrading nodes when a 7 block reorg happened on mainnet. This was a combination of clients not coordinating on the release and a very unlikely situation of 7 proposers in a row being on the not upgraded minority.

graph RL F["82"]:::green E["81"]:::lightblue C["..."]:::lightblue B["75"]:::lightblue D["74"]:::lightblue G["73"]:::lightblue H["..."]:::lightblue F --> D F ~~~ E E --> C C --> B B --> G B ~~~ D D --> G G --> H classDef lightblue fill:#ADD8E6 classDef green fill:#90EE90 classDef orange fill:#FFA500 classDef red fill:#FF0000

Disclosure ethics

The experience with the proposer boost fiasco was ingrained in many minds. The episode was traumatic enough that people were really scared about any changes to forkchoice (a fear that remains today). There was a palpable feeling in that we hadn’t disclosed the proposer boost situation and fix correctly. The pulled tips fix was taking longer than expected and client devs were increasingly uncomfortable knowing that we had such a serious bug live on mainnet and we were not releasing a fix. People started setting deadlines by which we would disclose the bug regardless of the status of the fix. This was completely new to me. At the time, I took this as a hobby, some alternative form of puzzle-solving. But people that had invested their last few years of their lives were already feeling the pressure of being responsible for the health of the chain. Geth had a bug exploited on mainnet not long before that and the merge had been delayed multiple times—this was looking like another potential delay.

A race started to get a complete spec and testvectors before we met that disclosure deadline. We needed to coordinate client fixes and run devnets, etc. At the time I was still trying to fit in the process, but then something happened that made me more involved in the issue. I realized that a particular insidious attack was enabled by the proposed pulled tips fix. We had somehow noticed this attack back in Amsterdam, but dismissed it as highly unlikely. Recall the situation above, that typically 22 slots into the epoch, there’s enough information to justify the checkpoint (typically the first slot in the epoch). A staker with enough stake could manipulate this, by withholding attestations it could make this happen later than 22 slots into the epoch, at the same time it could wait until they held all the latest slot in the epoch. For example a 15% stake holder, could delay justification until 27 slots into the epoch, and if it happened to hold the last 5 proposals in the epoch, it could withhold those 5 proposals, preventing thus from including onchain the information that everyone already has: that the epoch is about to be justified. This leads to a family of justification withholding attacks, some of which could still happen for very large concentrations of stake. One would tend to think that if you hold 15% of the stake, it would be very unlikely that you would get to propose the last 5 slots of an epoch. But the law of large numbers is often counter-intuitive. I ran the numbers and it turned out that it was highly likely to happen on a relatively short period of time: expected time to attack It would become rather critical if a 20% staker wanted to exploit this attack. This kind of situation allowed attackers to reorg an entire epoch worth of slots. Moreover, it allowed them to set up the attack before these blocks even appeared.

I convinced Danny Ryan that this was not to be neglected, and to make matters worse, Danny realized that this same type of attack was even possible with the vanilla forkchoice we had live on mainnet! Danny’s comments This is probably my only serious contribution to Ethereum: having been stubborn enough to bug Danny into considering this was serious enough to be fixed. This issue determined what was the right flavor of the pulled tips options we had, and eventually we settled on a mechanism by which clients would be continuously checking when the unrealized justified checkpoint changed—that is, when they had enough votes included on blocks to be certain that, during the next epoch transition, the chain would definitely update the justified checkpoint no matter what.

Devnets and testing

We quickly patched clients and coordinated devnets with some malicious clients with enough stake to trigger any of the attacks. Simple enough, we immediately found they were trivially triggerable: Enrico’s tool This is an image of an epoch long reorg that was caused by a withholding attack. One of the great things that we got from that testing time was testing tooling. This image was taken with a forkchoice viewer coded by Enrico Del Fante. It is the precursor to forky from ethPandaOps. Kurtosis was just starting to be used back then. Visualization tools like these have dramatically changed the game on how we triage and diagnose bugs quickly.

Devnets were an immediate success, we made sure every client team was patched and released jointly this time. Fears of proposer boost were still fresh. Luckily this was not an issue on mainnet. However, these reorgs were happening routinely by chance on Goerli, and even after the fixes went up, some related reorgs still continued to happen: Goerli Epoch reorg The success, the due process, and the spirit of multiple team collaboration that led to the fixing of this bug was without a doubt what made Ethereum my home. Every time that I feel down about Ethereum governance taking turns that I consider contrary to our ethos, I recall those days in which devs from 5 different teams, researchers and devops engineers, got together and prioritized a months-long process to properly analyze, fix, deploy and disclose a serious bug that could have affected the merge.