Public RPKI RTR Validators
Hello everyone, I'm looking for recommendations for reliable public RPKI RTR (RFC 8210) servers, excluding Cloudflare. I know the common recommendation is to run my own validator, but in my particular case I'd prefer to use one or more independently operated public RTR servers instead. The main reason is that I don't want my routers to depend on the reachability of infrastructure that I operate myself. I'd rather have RTR connectivity provided by infrastructure that is operationally independent of my own network. Does anyone know of operators that intentionally provide public RTR servers suitable for production use? Thank you, Stefan
Hi Stefan, For production I would deploy a private instance of routinator or FORT, and use a public service like Cloudflare's or Level66's only as secondary. Besides Cloudflare the only other one I know is Level66: rpki.level66.services:3323 Best regards, Salvador El jue, 30 jul 2026 a la(s) 12:47 p.m., Stefan-Gabriel Lungu via routing-wg (routing-wg@ripe.net) escribió:
Hello everyone,
I'm looking for recommendations for reliable public RPKI RTR (RFC 8210) servers, excluding Cloudflare.
I know the common recommendation is to run my own validator, but in my particular case I'd prefer to use one or more independently operated public RTR servers instead.
The main reason is that I don't want my routers to depend on the reachability of infrastructure that I operate myself. I'd rather have RTR connectivity provided by infrastructure that is operationally independent of my own network.
Does anyone know of operators that intentionally provide public RTR servers suitable for production use?
Thank you, Stefan ----- To unsubscribe from this mailing list or change your subscription options, please visit: https://mailman.ripe.net/mailman3/lists/routing-wg.ripe.net/ As we have migrated to Mailman 3, you will need to create an account with the email matching your subscription before you can change your settings. More details at: https://www.ripe.net/membership/mail/mailman-3-migration/
-- Salvador Bertenbreiter (+51) 947 145 352 (+57) 323 802 9009
Hi, https://rpki-validator.ripe.net/ui/ is a thing. However, I would stick with Routinator locally, does not take much to host it. Cheers Max Emig On 30.07.26 21:03, Salvador Bertenbreiter wrote:
Hi Stefan, For production I would deploy a private instance of routinator or FORT, and use a public service like Cloudflare's or Level66's only as secondary.
Besides Cloudflare the only other one I know is Level66: rpki.level66.services:3323
Best regards,
Salvador
El jue, 30 jul 2026 a la(s) 12:47 p.m., Stefan-Gabriel Lungu via routing-wg (routing-wg@ripe.net) escribió:
Hello everyone,
I'm looking for recommendations for reliable public RPKI RTR (RFC 8210) servers, excluding Cloudflare.
I know the common recommendation is to run my own validator, but in my particular case I'd prefer to use one or more independently operated public RTR servers instead.
The main reason is that I don't want my routers to depend on the reachability of infrastructure that I operate myself. I'd rather have RTR connectivity provided by infrastructure that is operationally independent of my own network.
Does anyone know of operators that intentionally provide public RTR servers suitable for production use?
Thank you, Stefan ----- To unsubscribe from this mailing list or change your subscription options, please visit: https://mailman.ripe.net/mailman3/lists/routing-wg.ripe.net/ As we have migrated to Mailman 3, you will need to create an account with the email matching your subscription before you can change your settings. More details at: https://www.ripe.net/membership/mail/mailman-3-migration/
-- Salvador Bertenbreiter (+51) 947 145 352 (+57) 323 802 9009
----- To unsubscribe from this mailing list or change your subscription options, please visit:https://mailman.ripe.net/mailman3/lists/routing-wg.ripe.net/ As we have migrated to Mailman 3, you will need to create an account with the email matching your subscription before you can change your settings. More details at:https://www.ripe.net/membership/mail/mailman-3-migration/
I know that you have specifically said that you do not want your routers to depend on the reachability of infrastructure that you operate yourself, but you really should actually just run a RPKI Validator yourself. 1) You almost certainly do not have any contractual agreement with cloudflare that they are going to operate a service that will stay online and correct (in a way that does not damage your business!) 2) Running such infrastructure is typically quite easy, especially in the case of Routinator or rpki-client + StayRTR (the latter I'm pretty sure being what cloudflare uses anyway), I /personally/ wouldn't recommend FORT 3) A outage on all RTR sessions should ideally not impact you in a meaningful operational way, RPKI "unknown"/"not-founds" should fail open, otherwise you are at the mercy of many other possible problems I do know that this is not the question you're asking but the existence of this email is provoking further questions about what your infrastructure is configured to do and what your models of reliability risk you are running on Regards Ben On Thu, 30 Jul 2026 at 18:47, Stefan-Gabriel Lungu via routing-wg <routing-wg@ripe.net> wrote:
Hello everyone,
I'm looking for recommendations for reliable public RPKI RTR (RFC 8210) servers, excluding Cloudflare.
I know the common recommendation is to run my own validator, but in my particular case I'd prefer to use one or more independently operated public RTR servers instead.
The main reason is that I don't want my routers to depend on the reachability of infrastructure that I operate myself. I'd rather have RTR connectivity provided by infrastructure that is operationally independent of my own network.
Does anyone know of operators that intentionally provide public RTR servers suitable for production use?
Thank you, Stefan ----- To unsubscribe from this mailing list or change your subscription options, please visit: https://mailman.ripe.net/mailman3/lists/routing-wg.ripe.net/ As we have migrated to Mailman 3, you will need to create an account with the email matching your subscription before you can change your settings. More details at: https://www.ripe.net/membership/mail/mailman-3-migration/
3) A outage on all RTR sessions should ideally not impact you in a meaningful operational way, RPKI "unknown"/"not-founds" should fail open, otherwise you are at the mercy of many other possible problems
From downstream customers, RPKI unknowns are dropped. Thanks, Stefan. Sent from Proton Mail for iOS. -------- Original Message -------- On Thursday, 07/30/26 at 22:15 Ben Cartwright-Cox <ripencc@benjojo.co.uk> wrote: I know that you have specifically said that you do not want your routers to depend on the reachability of infrastructure that you operate yourself, but you really should actually just run a RPKI Validator yourself. 1) You almost certainly do not have any contractual agreement with cloudflare that they are going to operate a service that will stay online and correct (in a way that does not damage your business!) 2) Running such infrastructure is typically quite easy, especially in the case of Routinator or rpki-client + StayRTR (the latter I'm pretty sure being what cloudflare uses anyway), I /personally/ wouldn't recommend FORT 3) A outage on all RTR sessions should ideally not impact you in a meaningful operational way, RPKI "unknown"/"not-founds" should fail open, otherwise you are at the mercy of many other possible problems I do know that this is not the question you're asking but the existence of this email is provoking further questions about what your infrastructure is configured to do and what your models of reliability risk you are running on Regards Ben On Thu, 30 Jul 2026 at 18:47, Stefan-Gabriel Lungu via routing-wg <routing-wg@ripe.net> wrote:
Hello everyone,
I'm looking for recommendations for reliable public RPKI RTR (RFC 8210) servers, excluding Cloudflare.
I know the common recommendation is to run my own validator, but in my particular case I'd prefer to use one or more independently operated public RTR servers instead.
The main reason is that I don't want my routers to depend on the reachability of infrastructure that I operate myself. I'd rather have RTR connectivity provided by infrastructure that is operationally independent of my own network.
Does anyone know of operators that intentionally provide public RTR servers suitable for production use?
Thank you, Stefan ----- To unsubscribe from this mailing list or change your subscription options, please visit: https://mailman.ripe.net/mailman3/lists/routing-wg.ripe.net/ As we have migrated to Mailman 3, you will need to create an account with the email matching your subscription before you can change your settings. More details at: https://www.ripe.net/membership/mail/mailman-3-migration/
Yeah so don't do that. You are going to run a policy like that. Then you should absol utely run your own validator, hardcode the customers that you are expecting into a SLURM file, or ensure that your customers are not allowed to single home with you It is a very bad idea to rely on rpki data always returning valid (vs being okay with not found state). The immediate thing that comes to mind is that if any of your customers are on ARIN signed space, ARIN, for some reason, Will occasionally just turn stuff off to try and see what happens (I am slightly exaggerating. They have a bit of a more scientific reasoning but (in my professional opinion) it's not that much better.) On Thu, Jul 30, 2026, 20:22 Stefan-Gabriel Lungu <hi@lungustefan.com> wrote:
3) A outage on all RTR sessions should ideally not impact you in a meaningful operational way, RPKI "unknown"/"not-founds" should fail open, otherwise you are at the mercy of many other possible problems
From downstream customers, RPKI unknowns are dropped.
Thanks, Stefan.
Sent from Proton Mail for iOS.
-------- Original Message -------- On Thursday, 07/30/26 at 22:15 Ben Cartwright-Cox <ripencc@benjojo.co.uk> wrote: I know that you have specifically said that you do not want your routers to depend on the reachability of infrastructure that you operate yourself, but you really should actually just run a RPKI Validator yourself.
1) You almost certainly do not have any contractual agreement with cloudflare that they are going to operate a service that will stay online and correct (in a way that does not damage your business!)
2) Running such infrastructure is typically quite easy, especially in the case of Routinator or rpki-client + StayRTR (the latter I'm pretty sure being what cloudflare uses anyway), I /personally/ wouldn't recommend FORT
3) A outage on all RTR sessions should ideally not impact you in a meaningful operational way, RPKI "unknown"/"not-founds" should fail open, otherwise you are at the mercy of many other possible problems
I do know that this is not the question you're asking but the existence of this email is provoking further questions about what your infrastructure is configured to do and what your models of reliability risk you are running on
Regards
Ben
On Thu, 30 Jul 2026 at 18:47, Stefan-Gabriel Lungu via routing-wg <routing-wg@ripe.net> wrote:
Hello everyone,
I'm looking for recommendations for reliable public RPKI RTR (RFC 8210)
servers, excluding Cloudflare.
I know the common recommendation is to run my own validator, but in my
particular case I'd prefer to use one or more independently operated public RTR servers instead.
The main reason is that I don't want my routers to depend on the
reachability of infrastructure that I operate myself. I'd rather have RTR connectivity provided by infrastructure that is operationally independent of my own network.
Does anyone know of operators that intentionally provide public RTR
servers suitable for production use?
Thank you, Stefan ----- To unsubscribe from this mailing list or change your subscription
options, please visit: https://mailman.ripe.net/mailman3/lists/routing-wg.ripe.net/
As we have migrated to Mailman 3, you will need to create an account with the email matching your subscription before you can change your settings. More details at: https://www.ripe.net/membership/mail/mailman-3-migration/
Hi, On Thu, Jul 30, 2026 at 07:22:48PM +0000, Stefan-Gabriel Lungu via routing-wg wrote:
3) A outage on all RTR sessions should ideally not impact you in a meaningful operational way, RPKI "unknown"/"not-founds" should fail open, otherwise you are at the mercy of many other possible problems
From downstream customers, RPKI unknowns are dropped.
That sounds like all the more reason to use a reliable source (= your own infrastructure that you can control and fix), not a free volunteer service "somewhere in the cloud" which might or might be unavailable at times... We run this on two FreeBSD VMs, one with routinator, one with rpki-client + stayrtr, and the maintenance effort has been very low. Gert Doering -- NetMaster -- have you enabled IPv6 on something today...? SpaceNet AG Vorstand: Sebastian v. Bomhard, Karin Schuler, Sebastian Cler Joseph-Dollinger-Bogen 14 Aufsichtsratsvors.: Dr. Frank Thiäner D-80807 Muenchen HRB: 136055 (AG Muenchen) Tel: +49 (0)89/32356-444 USt-IdNr.: DE813185279
On Thu, 30 Jul 2026 at 21:14, Ben Cartwright-Cox via routing-wg <routing-wg@ripe.net> wrote:
I know that you have specifically said that you do not want your routers to depend on the reachability of infrastructure that you operate yourself, but you really should actually just run a RPKI Validator yourself.
1) You almost certainly do not have any contractual agreement with cloudflare that they are going to operate a service that will stay online and correct (in a way that does not damage your business!)
2) Running such infrastructure is typically quite easy, especially in the case of Routinator or rpki-client + StayRTR (the latter I'm pretty sure being what cloudflare uses anyway), I /personally/ wouldn't recommend FORT
Running rpki-client + StayRTR reliably is non-trivial, because when rpki-client stops completing validation for whatever reason, StayRTR keeps serving VRP's. Only since StayRTR v0.6.3 is a local json file discarded after 24 hours and the stale VRPs are withdrawn. That's a good safeguard, hopefully everyone rewriting a RTR server from scratch will implement this as well (and not rely only on per VRP expiration). In older releases, because everything apparently keeps working, this is often not noticed. I'm not talking about not noticed for a couple of days, but it could be weeks or months. And this leads to missing routes in your routing table, because people do change ROA configurations and sometimes remove ROAs. This is especially true during prefix transfers or M&A which happen all the time., I was immediately able to find 6 missing BGP prefixes in a Tier 1 network (present in every other Tier 1 network) that I suspected had a hung validation with StayRTR, which was indeed the case. It took a couples of days to find the right person, and I don't think I found the issue just when rpki-client stopped validating, so it was probably hung for a while (likely months). There are many reasons rpki-client could stop validating. Connectivity issues, upstream FW changes, read-only FS, no enough memory, not enough disk space, bugs, permission problems, "it's a container so nobody understands what is going on, but the lights are all green". All in one software stacks will crash and tear down RTR sessions with them in most of those situations. With StayRTR v0.6.3+ your RTR count will go to 0, the RTR connection will keep running. Also something you'd wanna know, and it will probably not generate syslogs on your router. Therefor this setup requires monitoring that triggers when the validation does not complete successfully (rpki-client exit code with a dead man switch). And additionally it requires monitoring the RTR endpoint, triggering when RTR serial stays the same for too long [1]. Serving stale VRP's has an impact. A different validator/RTR servers keeps uptodate VRPs on the router - great, but that doesn't help the folks that had to remove the ROA due to a transfer and can't enable ROAs yet for whatever reason. Every RTR server should be monitored against a never changing RTR serial, but the all in one validation + RTR packages are a lot less [1] https://github.com/lukastribus/rtrcheck
Email went out prematurly ... On Thu, 30 Jul 2026 at 22:52, Lukas Tribus <lukas@ltri.eu> wrote:
Every RTR server should be monitored against a never changing RTR serial, but the all in one validation + RTR packages are a lot less
... painful to deal with in these situations. Lukas
On 30 Jul 2026, at 22:52, Lukas Tribus <lukas@ltri.eu> wrote: [..] There are many reasons rpki-client could stop validating. Connectivity issues, upstream FW changes, read-only FS, no enough memory, not enough disk space, bugs, permission problems, "it's a container so nobody understands what is going on, but the lights are all green".
Disk full is such a standard error case. If you do not watch for that you are just not doing ops correctly. Any software will die horrible death if done so and often silently as messages cannot get out as there is no space to store them etc. Oh and a reminder: Debian Trixie changed /tmp to tmpfs, thus instead of the terabytes of / filesystem you had it might now be a few GBs, thus it will fill, be aware of that one too now. I dump my rpki-client output into JSON, publish it on an internal rsync server and then rsync that from the routers. Each of my router (as they are debian/bird, debian/frr and openbsd/openbgpd) then has the following simple bash script in cron next to a diskspace check: 8<-------------------------- #!/bin/bash F="/rpki/rpki-client.json" if [ ! -f ${F} ]; then echo "ERROR: Missing RPKI JSON: ${F}" exit 1 fi TS=$(cat ${F} | jq -r .metadata.buildtime) if [ -z "${TS}" ]; then echo "ERROR: RPKI misses timestamp in ${F}" exit 1 fi THEN=$(date -d "${TS}" +%s) if [ -z "${THEN}" ]; then echo "ERROR: RPKI timestamp did not convert: ${TS}" exit 1 fi NOW=$(date +%s) MAX=$((4 * 60 * 60)) AGE=$((NOW - THEN)) if [ ${AGE} -gt ${MAX} ]; then echo "ERROR: RPKI more than ${AGE} seconds, max: ${MAX}, timestamp: ${TS}" exit 1 fi exit 0 --------->8 As such, spam will reach me when it is too much. Alternatively, you could poll the file remotely from a central system (so that you know your disk is not full etc, or at least monitoring fails), or do many other things. One can easily monitor the buildtime stamp. Of course, that is just to ensure that the rpki-client.json is updated recently. StayRTR is easier, you can monitor them with rtrmon that is included. (See https://github.com/bgp/stayrtr/blob/master/cmd/rtrmon/index.html.tmpl ) which even has prometheus exports for those that use that. Otherwise do similar to the above and check the metadata....
On Thu, Jul 30, 2026 at 10:52:28PM +0200, Lukas Tribus wrote:
There are many reasons rpki-client could stop validating. Connectivity issues, upstream FW changes, read-only FS, no enough memory, not enough disk space, bugs, permission problems, "it's a container so nobody understands what is going on, but the lights are all green".
This seems weird to pick on rpki-client this. What software is resistent against hardware failures or execution in a unfitting environment? For exactly these reasons, a sane program will error out with a non-zero exit code when problems are detected. Whether the program is one-shot or long running is irrelevant in this context: the operator must monitor the process execution. Rpki-client even warns the operator on STDERR when it suspects there won't be enough inodes or disk space.
All in one software stacks will crash and tear down RTR sessions with them in most of those situations.
This doesn't match my experience: I've discovered catastrophic silent bugs in basically every RPKI validator projects. And not even all of those bugs have been solved when I last took stock! I do agree that programs which as a general rule 'crash hard and fast' are easier to manage than programs that just 'limp on'. In my experience it is easier to construct reliable setups when you can monitor each individual component in the pipeline (fwiw, both rpki-client and StayRTR support OpenMetrics/Grafana, which helps me with faster fault identification). This is also why I love programs that properly set exit codes. To pivot to a more constructive line: in the realm of dead man's switches, I've grown fond of https://healthchecks.io/ (up to 20 monitors is free) linked with Pushover ($5 one-time purchase). Very cheap way to keep an eye on whether backups or rpki-client invocations are humming along nicely. Alternatively, OpenBSD's crontab implementation has built-in 'cronic' functionality. In systemd do this: https://wiki.archlinux.org/title/Systemd#Notifying_with_e-mail
With StayRTR v0.6.3+ your RTR count will go to 0, the RTR connection will keep running. Also something you'd wanna know, and it will probably not generate syslogs on your router.
Wouldn't any validator's count go to zero if there are connectivity issues to the rest of the Internet (but not with the BGP routers it is serving)? There is no method of teleporting the ROAs into airgapped validator instances. The RTR protocol even has a "No Data" Error Report PDU: https://datatracker.ietf.org/doc/html/draft-ietf-sidrops-8210bis#name-cache-...
Every RTR server should be monitored against a never changing RTR serial,
Yes.
but the all in one validation + RTR packages are a lot less painful to deal with in these situations.
On the other hand, separation of the 'validation' and 'RTR distribution' function allows for seamless upgrades of the security-sensitive component WITHOUT flapping the BGP router-facing RTR sessions. Different deployment models bring different benefits to the table. Kind regards, Job
Hi Job, On Fri, 31 Jul 2026 at 15:44, Job Snijders via routing-wg <routing-wg@ripe.net> wrote:
On Thu, Jul 30, 2026 at 10:52:28PM +0200, Lukas Tribus wrote:
There are many reasons rpki-client could stop validating. Connectivity issues, upstream FW changes, read-only FS, no enough memory, not enough disk space, bugs, permission problems, "it's a container so nobody understands what is going on, but the lights are all green".
This seems weird to pick on rpki-client this. What software is resistent against hardware failures or execution in a unfitting environment?
FWIW I don't feel like I'm picking on rpki-client at all :)
exactly these reasons, a sane program will error out with a non-zero exit code when problems are detected. Whether the program is one-shot or long running is irrelevant in this context: the operator must monitor the process execution. Rpki-client even warns the operator on STDERR when it suspects there won't be enough inodes or disk space.
I agree completely. I'm just disagreeing with the notion that it is "easy to setup", as suggested in this thread. A single software stack is easier to monitor than 2 software stacks that have to talk to each other. It's is only natural that this architecture has more moving parts, and with more moving parts comes more external monitoring requirements.
All in one software stacks will crash and tear down RTR sessions with them in most of those situations.
This doesn't match my experience: I've discovered catastrophic silent bugs in basically every RPKI validator projects. And not even all of those bugs have been solved when I last took stock!
I do agree that programs which as a general rule 'crash hard and fast' are easier to manage than programs that just 'limp on'. In my experience it is easier to construct reliable setups when you can monitor each individual component in the pipeline (fwiw, both rpki-client and StayRTR support OpenMetrics/Grafana, which helps me with faster fault identification). This is also why I love programs that properly set exit codes.
That's exactly my point, it needs monitoring. It depends only on the administrator/operator that is setting up those services, which is what this thread is about (how easy or hard it is to properly run a RPKI RP and RTR Server in a reliable way). This thread and my points are not about software quality.
To pivot to a more constructive line: in the realm of dead man's switches, I've grown fond of https://healthchecks.io/ (up to 20 monitors is free) linked with Pushover ($5 one-time purchase). Very cheap way to keep an eye on whether backups or rpki-client invocations are humming along nicely.
Agreed, I'm using the very same 2 services as well.
Alternatively, OpenBSD's crontab implementation has built-in 'cronic' functionality. In systemd do this: https://wiki.archlinux.org/title/Systemd#Notifying_with_e-mail
I'm using chronic in regular cronjobs on Linux. Same here: the person setting those services up needs to know this and set it up accordingly. I'm not worried about my clients or your clients. I'm worried about the guy that has to setup all of this, without the time to properly analyze all of those possible operational problems. Like the original poster in this thread. I have written rtrcheck as a simple nagios compatible plugin and written a RIPE labs post about it in 2020 and after 6 years I am not aware of a single user of rtrcheck besides my own clients. We all like to think that at least the big SPs are tracking and monitoring those metrics, and then we find out that the reality is not not great. I also feel like there is an amount of hindsight bias involved in these discussions. I always hear that everything is so obvious. And yet people don't do and don't really know how to do proper monitoring. Perhaps a proper BCP document about these operational aspects would be good idea. But then I will probably just get the usual "Why? Everything is so obvious!" ...
With StayRTR v0.6.3+ your RTR count will go to 0, the RTR connection will keep running. Also something you'd wanna know, and it will probably not generate syslogs on your router.
Wouldn't any validator's count go to zero if there are connectivity issues to the rest of the Internet (but not with the BGP routers it is serving)?
Yes, I'm just advocating for monitoring for RTR serial monitoring here.
Every RTR server should be monitored against a never changing RTR serial,
Yes.
but the all in one validation + RTR packages are a lot less painful to deal with in these situations.
On the other hand, separation of the 'validation' and 'RTR distribution' function allows for seamless upgrades of the security-sensitive component WITHOUT flapping the BGP router-facing RTR sessions. Different deployment models bring different benefits to the table.
That is absolutely a big advantage. *I* would use rpki-client/stayrtr every day of the week. I'm just not recommending it to someone that I know will think of monitoring "at some point in the future", I'd rather this operator uses a All-in-One stack if he can't invest the time for proper operation *before* going into production, because that is likely a less severe problem than the alternative. Nothing about this is about software quality. Best regards, Lukas
participants (8)
-
Ben Cartwright-Cox -
Gert Doering -
Jeroen Massar -
Job Snijders -
Lukas Tribus -
ripe@emigm.ax -
Salvador Bertenbreiter -
Stefan-Gabriel Lungu