Dear all, The below is intended as a detailed write-up for future reference and not directed at anyone specific. I share this because the topic of "why not have the RIRs also use the RPKI for negative attestations" continues to come up from time to time. On Mon, Jul 27, 2026 at 05:32:38PM +0000, James Bensley wrote:
Thanks for providing more info that was helpful.
I don’t agree with quite a lot of what has been written (simply because there seems to be more emotional that strong/clear logical reasoning), summarising some of the points across the archives:
I'm really just trying to help! :-)
* "It’s a lot of effort" → maybe it was in 2021, to put together a pipeline today that generates an AS0 ROA when a prefix changes state of unallocated, and removing it when it loses that state, should take a couple of days to churn out in today, maximum 1 week. So I don’t buy the effort argument at all. If it really is a serious amount of work, that just raises more question about the RIPE NCC.
I wonder what one would base the "this feature shouldn't take more than a week" assertion on? In my mind such an assertion raises questions about the method of estimating development effort rather than anything about the RIPE NCC. :-) FWIW, I think the brunt of the work is NOT in simply encoding a list of prefixes into ASN.1 BIT STRING conforming to the ROA profile (encoding transformations obviously are the simple part). I think some of the substantial cost is in both cross-departmental and non-technical work, e.g, alignment with CA audit procedures, alignment with registration services, revising certification practise statements, updating of educational materials, QA and scale testing, legal impact analysis, insurance & liability assessment, impact on regulatory compliance posture, impact on sanctions compliance posture, preparations for obtaining executive board approval, etcetera. You don't need to take my word for it, but for the sake of this discussion I'd recommend to assume a budget of 2 FTE for a period of 12 to 18 months to get to a production-grade AS0 implementation and then 0.5 FTE going forward to maintain the result. Perhaps needless to say, but overspending as a result from underbudgeting is a common failure in organisations and in this context the membership would foot the bill. I base these estimates on my own experience managing multi-year/multi-FTE RPKI projects in hopes of helping clarify where I'm coming from.
* "We’re failing closed" → this is a logical fallacy. The general approach with RPKI based validation is “if there is a signal to reject, then reject, otherwise accept”, this is fail open logic. That logic is not being changed. What is being changed is that there are no signals for these resources we’re discussing, so for those resources specifically we hit the default accept clause in our routing policies. We’re taking about adding in signals for those resources so that we hit the “there is a signal to reject” clause in the if statement. The suggestion that the logic would be changed to fail closed is simply not correct.
I disagree with this narrow reading. RFC 6811 outlines a validation process that has as outcome a trinary state: "valid", "invalid", and "not-found". The latter state is is the 'fail open' that facilitates incremental deployment, but also is where resources end up for which there is a contractual lapse or some temporary administrative issue. Deregistrations and disputes do happen from time to time, as do software defects (I'll expand on both). In a world where all unassigned/unallocated/reclaimed space is covered by AS0 RPKI ROAs, for BGP announcements related to those type of space the third state ('not-found') is fully subsumed into the 'invalid' state. This is what I mean with 'fail closed': many kinds of process failures can result in AS0 listing, with the end result likely being reachability issues. ... which brings me to an seemingly underappreciated topic: RISK. In reasoning about the construction of new common infrastructure we shouldn't merely look at "what does it cost to build this?" (~ 2 FTE), and "once it is build, what will it gain us?" (detection of few tens of discrepant prefixes), but also important - "what is the risk of this thing even existing? What new risks come into existence?". Risk management is a complicated topic because we know risk is real but often risk is hard to qualify or quantify. There may even be unforeseen risks! One might retort "Job, but fear is the mindkiller!", however such a platitude wouldn't mitigate the real risks. (I'm not trying to put words in anyone's mouth but to get ahead of blanket dismissal of risk as a tangible factor to account for.) I see a few risks: * Administrative process failure -- In the current model, a payment dispute (just an example of a problem of an adminitrative nature) lands things in the 'not-found' bucket, while the creation of a new CA issuing AS0 objects causes more severe impact. This could lead to an increase in claimed damnages in court cases. This example of a risk is not theoretical: I know of real world examples in other regions of the world where a prefix accidentally ended up on an AS0 and this caused material damage to all involved. * Software defects -- While it is formally specified that software defects MUST NOT exist (RFC9225), research across a multitude of codebases suggests that bugs are introduced at a rate of roughly one bug per 1000 CLOC. Mitigations include separation of privileges (isolation), reduction of privileges when executing, embracing safe coding practises, but also a much less obvious one: avoidance creation of certain software to begin with. This may sound like a nihilist argument, but my point is that we just need to be _very_ careful materializing everything we wish for. Some lived experience: in the 2019/2020 era with the global deployment of RPKI-ROV an incredible number of software defects was uncovered in BGP router implementations, RPKI validators, and RPKI signer/publication software. It was painful. Bug hunting was like shooting fish in a barrel. We almost didn't make it. However in my mind the juice was worth the squeeze because the then ongoing avalanche of BGP routing incidents (related to allocated resources) also carried a tremendous collective cost. With the advent of ASPA we are about to start the cycle anew: while I know that many lessons have already been taken to heart (e.g., one outcome is that ASPA will not suffer from some classes of race conditions that were disovered in 2019/2020). However, there is a strong probability that initial ASPA implementations will contain severe bugs. Time will tell whether the market ends up embracing ASPA or disabling it wholesale. Additionally, there already is a documented case where a software bug related to in contractual status in a database system bled through into the existence / disappearance of RPKI ROAs. (Note: this is different from an administrative process failure.) From what I heard it took the RIR a fair a bit of effort to audit and strengthen their systems by adding controls and improve workflows to prevent reoccurance. The bug surface is real and from my perspective the practise of negative attestation would certainly increase the bug surface. * Administrative failures & software bugs can erode community trust -- While there is _some_ room to absorb fallout from problems (nothing is perfect), how many outages / problems is 'one too many'? We all know the infamous mantra "if you have network problems just disable IPv6!". It takes more energy to establish credibility of an infrastructure than to lose it. The current state of RPKI infrastructure & deployment is the fruit of years of hard work and investment. We all paid for it. I consider expansion of the certification scope towards (purportedly!) unallocated/unassigned/reclaimed resources to carry a substantial risk of devaluing the existing investment. We'd be in a fine pickle if people wholesale stop using RPKI due to fallout from incidents related to negative attestation. * Risks of fallout related to sanctions -- Developments in recent years certainly posed challenges for our collective ability to operate in compliance with newly imposed sanctions. I believe the concept of negative attestation may be somewhat like Chekhov's gun. This might be a good read: https://www.techpolicy.press/towards-the-multistakeholder-imposition-of-inte... The section titled "Manipulation of routing security attestations" ties into the aforementioned risk of community distrust. * Technical scaling issues -- Already today, "staying ahead of the demand curve" is a challenge in the RPKI ecosystem. It might not be apparent to outsiders (i.e., to people who do not develop RPKI protocols and software as part of their daily work), but a lot of time and effort is spend on optimising the efficiency of RPKI (e.g., in the RPKI encoding and transport layers, understanding publication best practises, but also BGP RIB lookup optimisations in routers). Expanding the certification scope to additionally cover the *complement* of allocated resources is no small matter! In systems design, it seems there may be general principle that it is more economical/reliable/efficient to compose sets of instructions describing what you DO want to happen - rather than disseminating the inverse possibility space. Constructing and distributing both at the same time obviously is a substantial increase in burden. I suspect there is something architecturally unsound with the concept of negative attestations in context of the global routing system. An analogy (undoubtly severely flawed): nation states usually do not proactively jam unlicensed/unallocated radio spectrum. As a strategy it is just easier to passively measure and apply tactical intervention rather than to try and make the radio spectrum as a whole unusable for those without license. One might say "but negative attestation ROAs and regular ROAs are intended to be ships in the night!", but there _is_ some fate-sharing, at the very least the information streams will coalesce in BGP routers. To put this in numbers: looking at today's APNIC & LACNIC data for their AS0 certification, a validator's output results a 40% increase (!) in number of payloads compared to the 'regular' certification trees. I suspect that this number differs from Jeroen's estimate (which was RIPENCC-centric) because the opportunity to efficiently aggregate may be different from region to region due to (historic) assignment strategies. And, knowing that RIR assignment policies and allocation strategies do change over time, there might be profoundly ill scaling effects down the road resulting from that. Opportunity to aggregate data is not a given. In other words - I'd much rather retain this fictious "40% budget increase" as breathing room to accommodate growth of the existing regular certification trees for allocated resources or spend such "headroom" towards novel 'high yield' features (e.g., ASPA).
* "It’s a slippery slope" → this is another logical fallacy. If we decide to do something because it's easy ($old_thing we did in the past makes $new_thing easier to achieve than it would have been, had we not done $old_this already), rather than deciding to do something because it's the right thing to do; then the problem is that we are doing something which isn’t the right thing to be doing, we doing something which is easy. We can decide to do many things because they are easy, they could be easy without any past work laying a foundation, that still doesn’t me we should do those things. We should always do things because we decided they are the right thing to do. If we are evaluating things properly, the easiness to achieve the goal has no bearing (only in extreme scenarios where $effort >= some absurd value).
I wholeheartedly agree with all of the above, but I fail to reconcile the above with what I wrote. :-)
Having thought about this, I think a more water tight argument I can get behind is specifically the risk that comes from the extremely low propagation times for the RPKI; Even though I hate prefix lists, one of the few benefits is that there is a 24 hour window (the industry standard interval IMO) afforded to operators between making a mistake in your IRR data, and getting it fixed before the poop decorates the walls. And it’s a risk I have been musing on recently regarding the RPKI and what to do about it. I would guess that a new ROA is in the control plane of the major of routers in the global DFZ within about 30 minutes. So I think a serious risk is; if something goes wrong with RIPE’s automated process, *there is no time buffer to correct anything* (on it’s own, some sort of bug or mistake is not enough of a reason for me, we’re not talking about launching a rocket to mars, but it’s specifically the combination of a bug /mistake + near real time propagation).
All very good points, and you make an excellent case for strigent reliable fast-paced monitoring of the world's RPKI data. Improving propagation speed of RPKI ROAs is very high on my list of things! :) Kind regards, Job