Saturday, 05 September 2026

This Week In Techdirt History: August 30th – September 5th [Techdirt] (03:00 , Saturday, 05 September 2026)

Game of Trees 0.128 released [OpenBSD Journal] (01:27 , Saturday, 05 September 2026)

Version 0.128 of Game of Trees has been released (and the port updated). Complete release notes are as follows:

Read more…

What’s Your Go-To CW Key for Wet-Weather Operating? [Q R P e r] (06:51 , Saturday, 05 September 2026)

by Thomas (K4SWL) I’m fortunate that, most of the time, when I do Summits On The Air, Parks On The Air, or even Field Day operations, I can choose conditions that are favorable. Sunny days, decent weather, or—if the forecast looks dodgy—a picnic shelter under which to operate. But sometimes it just doesn’t work out … Continue reading What’s Your Go-To CW Key for Wet-Weather Operating?

Bikes and Film Cameras Club kit revision with NEW DESIGN! [About – Bikes and Film Cameras Club] (06:01 , Saturday, 05 September 2026)

Hello friends of bikes and film cameras! This little club has been going on for over five years. When I founded it back in 2021, I simply decided to revise a design I created the year before to become the club's logo and branding. I always felt that the design was more serviceable than great, … Continue reading Bikes and Film Cameras Club kit revision with NEW DESIGN!

Fomapan FomaOrtho 400 – a film that might inspire you [35mmc] (05:00 , Saturday, 05 September 2026)

FomaOrtho 400 is one of those films you can’t quite predict until you scan it. I’ll admit I had my doubts — a few of the frames on this roll were ones I really wanted to nail If you are not familiar with Fomapan (Foma Bohemia), they are a Czech negative film manufacturer with a...

The post Fomapan FomaOrtho 400 – a film that might inspire you appeared first on 35mmc.

Amtrak infrastructure work will temporarily reduce train service in Virginia [Cardinal News] (04:30 , Saturday, 05 September 2026)

The Northeast Regional Amtrak train arrives in Roanoke. Photo by Dutchie Jessee.

An infrastructure project in Washington, D.C., will reduce Amtrak train service in Roanoke and Lynchburg and eliminate it in Danville for nine days next month.

The project, “Pause for Progress,” aims to replace switches and other aging track components in the First Street Tunnel by Washington’s Union Station. It will cut Amtrak service between Washington and Alexandria and affect multiple train routes south of Alexandria.

Work will begin after service ends on Oct. 16 and will wrap up before service resumes on Oct. 26, Amtrak said in a news release.

“A short-term, coordinated outage is more efficient due to the confined tunnel environment and limited access, making phased construction impractical,” Amtrak said in its release. “Customers traveling to or from Washington, D.C. from the south should expect significant service reductions and train cancellations.”

During the nine-day period, the normal twice-daily round trips between Roanoke and Washington on the Northeast Regional line will be reduced to one daily round trip between Roanoke and Alexandria.

The train that leaves Roanoke in the afternoon will be canceled during this period, and the morning train will leave Roanoke at 9:48 a.m. The train will return to Roanoke at 10:23 p.m. each day, according to Amtrak’s booking website.

Amtrak will offer buses for riders who want to continue to Washington after reaching Alexandria. Metro service between Alexandria and Washington also will be available.

Amtrak’s temporary plan also calls for a bus, called the Amtrak Virginia Express Bus Service, to offer one round trip daily between Roanoke and Washington, stopping in Lynchburg. 

The express bus will not stop in Charlottesville, Culpeper, Manassas or Burke Centre as the Northeast Regional train does. However, another express bus will run from Charlottesville directly to Union Station.

Meanwhile, Amtrak’s Crescent line ordinarily stops in Danville and Lynchburg as it runs between New York and New Orleans.

During “Pause for Progress,” it will not stop in either city. Instead, it will run only between Washington and New York, and separately between Atlanta and New Orleans.

Local officials applaud Amtrak upgrades

Lynchburg’s director of economic development and tourism, Marjette Upshur, said in an email that the city recognizes that Amtrak’s temporary service reductions will cause some inconvenience.

She encouraged travelers to review their itineraries carefully and monitor communications from Amtrak before traveling.

“While short-term disruptions are never ideal, we understand that this concentrated construction period is intended to improve the safety and long-term reliability of the passenger rail system,” Upshur said. “The City remains strongly supportive of reliable passenger rail service and its continued growth in Lynchburg and throughout the Commonwealth.”

Danville City Manager Ken Larking said that he is a regular Amtrak customer and knows several people who use the service.

“I am glad that they are making upgrades to the infrastructure so that people can have an even better experience,” Larking said in an email.

Roanoke officials did not respond to a request for comment.

Amtrak: Components are beyond service life

Amtrak said that “Pause for Progress” will improve a critical point within the First Street Tunnel that is responsible for routing all Amtrak and Virginia Railway Express train movements between Union Station and Virginia.

“The existing switches and crossovers are beyond their service life, requiring a 10-mph speed restriction that impacts operations,” Amtrak said in its news release. “‘Pause for Progress’ is designed to address those limitations and help prevent larger, unplanned service disruptions.”

Amtrak said it is coordinating with the Virginia Passenger Rail Authority, the Virginia Railway Express — a commuter rail service for Northern Virginia and Washington — and North Carolina’s state transportation department “to ensure consistent messaging and a seamless customer experience.”

According to the Virginia Passenger Rail Authority, Amtrak’s Roanoke corridor saw nearly 370,000 passengers in 2025. The highest month of ridership on that corridor was in November, when it saw 34,885 passengers.

Separately from “Pause for Progress,” work is underway to build a new train station in Christiansburg to extend passenger rail service to that community. That service is expected to begin in 2027.

The post Amtrak infrastructure work will temporarily reduce train service in Virginia appeared first on Cardinal News.

TIL: Emacs rectangle-number-lines [Open source software and nice hardware] (04:12 , Saturday, 05 September 2026)

+++ Saturday  5 September 2026 +++

TIL: Emacs rectangle-number-lines
=================================

Note from the "today I learned" department.

Emacs never cease to amaze me. This time I ran into the nifty command

    C-x r N rectangle-number-lines

Which exactly does what it says on the tin.

A small example
---------------

Before:

   'Would you tell me, please, which way I ought to go from here?'
   'That depends a good deal on where you want to get to,' said the Cat.
   'I don't much care where—' said Alice.
   'Then it doesn't matter which way you go,' said the Cat.
   '—so long as I get somewhere,' Alice added as an explanation.
   'Oh, you're sure to do that,' said the Cat, 'if you only walk long enough.'

   Alice felt that this could not be denied, so she tried another
   question. 'What sort of people live about here?'

   'In that direction,' the Cat said, waving its right paw round,
   'lives a Hatter: and in that direction,' waving the other paw,
   'lives a March Hare. Visit either you like: they're both mad.'

   'But I don't want to go among mad people,' Alice remarked.
   'Oh, you can't help that,' said the Cat: 'we're all mad here. I'm mad. You're mad.'
   'How do you know I'm mad?' said Alice.
   'You must be,' said the Cat, 'or you wouldn't have come here.'

Now we mark a rectangle from "Alice felt..." to "How do you know"
and run C-x r N rectangle-number-lines:

After:

   'Would you tell me, please, which way I ought to go from here?'
   'That depends a good deal on where you want to get to,' said the Cat.
   'I don't much care where—' said Alice.
   'Then it doesn't matter which way you go,' said the Cat.
   '—so long as I get somewhere,' Alice added as an explanation.
   'Oh, you're sure to do that,' said the Cat, 'if you only walk long enough.'

    1 Alice felt that this could not be denied, so she tried another
    2 question. 'What sort of people live about here?'
    3 
    4 'In that direction,' the Cat said, waving its right paw round,
    5 'lives a Hatter: and in that direction,' waving the other paw,
    6 'lives a March Hare. Visit either you like: they're both mad.'
    7 
    8 'But I don't want to go among mad people,' Alice remarked.
    9 'Oh, you can't help that,' said the Cat: 'we're all mad here. I'm mad. You're mad.'
   10 'How do you know I'm mad?' said Alice.
   'You must be,' said the Cat, 'or you wouldn't have come here.'

Wonderful, isn't it?

Last edited: $Date: 2026/09/05 10:12:23 $
   

Friday, 04 September 2026

Take-Two’s Leak Burying DMCA Attempts Snared GameStop & Gaming Journalist That Did Nothing Wrong [Techdirt] (10:39 , Friday, 04 September 2026)

I have mostly stayed away from the whole saga surrounding the drip-drip leaks of Grand Theft Auto 6 content prior to the big reveal on Netflix because, frankly, I am quite wary of giving companies the kind of guerilla marketing wins that sometimes look like this sort of thing. That being said, I really don’t think any of this was some attempt to Streisand what is perhaps already the most anticipated game of all time into wider news coverage, and that is backed up by the DMCA blitz Take-Two has gone on to try to bury all of these leaks.

Those attempts shouldn’t surprise anyone, honestly. Take-Two and Rockstar have historically abused copyright law to try to bury all kinds of content it doesn’t like, whether it’s been game leaks in the past, or cheats for its games, or mods it doesn’t like.

But it sure would be nice if the partners Take-Two has doing the abusing of the law could bother to be somewhat accurate and not ensnare a gaming journalist for the crime of posting publicly available court documents.

On August 26, Stephen Totilo — the longtime Kotaku editor-in-chief who now runs the Game File newsletter — was locked out of his X account over a DMCA notice filed on Take-Two’s behalf. Totilo’s offense, by his own account, was an August 21 post reporting that judges in New York had cleared Take-Two to subpoena Microsoft and Discord in the leak hunt.

Attached were three screenshots: the two court orders, and a tweet from Xbox CTO Scott Van Vliet pledging Microsoft is “working closely with Take-Two and Rockstar Games.” No leaked footage. No gameplay. The orders are public records that never once use the words Grand Theft Auto.

After Totilo complained both to ExTwitter and on ExTwitter, his account and the original tweet were restored and the DMCA claim had been rescinded. There is no indication that Take-Two or the vendor it was using to police the internet for these leaks have said anything publicly or privately to Totilo. They just nuked his account over a bullshit claim that ten seconds of review would indicate contained no infringing material, then restored it when the mistake was called out, and now are trying to Homer Simpson back into the bushes as though nothing happened.

But what makes this all the more frustrating is that the DMCA notice doesn’t make a copyright claim. It appears to make a trademark claim, instead.

It asserts Take-Two’s international figurative trademark on Grand Theft Auto — a trademark on the logo — and argues there is a likelihood of confusion, the legal test for whether the public might mistake someone else’s goods for the brand’s.

In plain English: a copyright takedown form was used to make a logo complaint, against images that contain neither the logo nor a single frame of the game.

The notice describes the reported content — federal court orders included — as “video/audiovisual recording,” and certifies all of it as accurate under penalty of perjury, the line that makes knowingly lying on the form a federal offense.

Everyone in this portion of the story, save Totilo, sucks at their jobs. Take-Two has clearly partnered with a company, Ebrand, that is not up to the task of properly policing IP on the internet. Ebrand messed this up badly, asserting a trademark claim via a copyright mechanism. ExTwitter, for its part, apparently demonstrated just how little review is done on this sort of thing, having taken down the tweet and suspending a journalist’s account over this absolute mess of a DMCA claim. It’s a full cornucopia of stupid on display for the world to see.

And this isn’t a one-off. Gamestop was also ensnared in Take-Two’s DMCA blitz. Its crime appears to be sharing a promotional screenshot for GTA6 that Rockstar specifically made available for use publicly.

Its August 20 post promoting a story on the billions in market value Take-Two shed as the leaks spread got struck, and the image X wiped was Rockstar’s own official GTA 6 screenshot, straight from the press gallery on Rockstar’s site.

That exact shot has run on dozens of outlets since May 2025, IGN and Mashable included. Take-Two’s vendor filed federal paperwork against a promotional asset Rockstar distributes so that outlets will use it.

There is simply no point to the DMCA’s “under penalty of perjury” language if it can’t be employed in a situation like this. At the very, very best, Ebrand and Take-Two are guilty of unbelievable negligence in issuing these DMCA takedowns and copyright strikes. When we’re talking about even temporary takedowns of the work of journalists, the First Amendment implications become obvious.

To allow these companies to simply slink away without penalty is why this sort of thing keeps happening. If there are no consequences to a carpet-bomb approach to copyright (trademark?) takedowns, then they will, and do, continue.

RFK Jr. Wants Your Medical Records [Techdirt] (06:22 , Friday, 04 September 2026)

This article is republished from The Conversation under a Creative Commons license. Read the original article.

You might assume that what you tell a doctor stays between you, your physician and perhaps your insurer. But the reality is more complicated.

The Health Insurance Portability and Accountability Act, the federal privacy law that governs health information and is commonly known as HIPAA, is narrower than its reputation suggests. It regulates hospitals, physicians, insurers and their business associates, but not the health data you generate everywhere else: not the period-tracking application on your phone, the internet search you ran about a diagnosis, the DNA you mailed to a genealogy company or the wearable that counts your heartbeats.

Even the records HIPAA does cover can be shared, sold or handed to the government in ways that might surprise you.

This gap in protection matters more than ever because the U.S. government is pushing hard to gather health data domestically and abroad. This is happening even as a growing body of research shows that the safeguard which these efforts to collect data lean on – anonymizing data by removing identifying information to make it difficult to trace back to an individual – is far weaker than officials claim.

As a professor of law at Indiana University, I study health information privacy and medical data regulation, which includes tracing how sensitive health information moves among clinics, government agencies and law enforcement. As a co-investigator on a federally funded study about opioid prescribing, I rely on health data in my own research. I appreciate its value for science, and I also see the danger of collecting it without meaningful safeguards.

Limits of medical privacy

HIPAA gives you several rights: You can see your health records, demand corrections and expect that a covered provider will not casually disclose your information.

But the law also permits release of some information without your consent. A hospital fully bound by HIPAA may release certain types of records without your authorization and without telling you. There are roughly a dozen such categories. Information about treatment, payment and routine healthcare logistics require no sign-off. Neither does information released for public health reportinglaw enforcementjudicial and administrative proceedingshealth plan oversightresearch or the broad catchall of essential government functions.

The statute is also thick with additional exceptions. In practice, much of your health information can be shared through these many open doors. And once data is sent outside the system covered by HIPAA, the HIPAA limits fall away.

For instance, prescription drug monitoring programs, which every state now operates, assemble detailed logs of who filled which controlled substance prescription and when. Federal law enforcement can often access these logs with a self-issued administrative subpoena – an order that doesn’t require a judge’s approval or oversight.

These programs have expanded beyond opioids into a dragnet that shares health data across state lines, exposing patients who seek reproductive or gender-affirming healthcare to surveillance far from home.

Health records can flow to many destinations under different rules. A given disclosure might feel more like a violation depending on who decides where it can go and who can then see it.

RFK Jr.’s push to access Americans’ health records

Since the spring of 2025, Health and Human Services Secretary Robert F. Kennedy, Jr. has sought federal access to Americans’ medical records to investigate whether vaccines cause autism. The scientific community has studied this question for decades and has shown decisively that they do not.

According to KFF Health News, HHS has been courting state health information exchanges – the little-known systems that let hospitals and clinics swap detailed, identifiable patient records – and asking how those records might be used for vaccine research. One proposal floated by state organizations would give HHS data on 90% of Americans’ medical records by 2028. In Nebraska, millions of federal grant dollars have flowed to a statewide health information exchange nonprofit that has cooperated with the effort.

Large health datasets can be useful. Pooled records can expose drug side effectstrack outbreaks and reveal disparities in care that smaller studies miss. Public health has always depended on some surrender of individual privacy for collective benefit.

The concern is not that the government should never collect health data. It is that meaningful safeguards have not kept pace with the scale of collection and capabilities of modern data analytics.

In seeking access to Americans’ medical records for a vaccine and autism study, HHS has declined to say how many states are involved, what data it collects, who can see it or how it will be protected.

Building a comprehensive repository to chase a question that science has already answered inverts the logic of research. Usually a hypothesis justifies the data collected, rather than the reverse.

Collecting identifiable records for tens of millions of people in a single database also creates a target for breachessecondary uses that no one consented to and abuses by current or future administrations with different priorities.

‘Anonymized’ doesn’t protect your health privacy

Officials have offered reassurances that data will be aggregated and stripped of identifiers so no individual can be singled out.

Decades of computer science research undercuts that promise. A study published in Nature in June 2026 sharpened the point, showing that in this age of artificial intelligence, stripping identifiers from patient records to protect identity does not protect all patients equally.

The researchers audited AI diagnostic models trained on clinical data, including chest X-rays, electrocardiograms and electronic health records. They asked whether an outsider could tell if a particular person’s data had been used to build the model. For instance, confirming that someone’s record helped train a cancer-prediction tool can reveal that that person has cancer. This exploit is known as a membership inference attack.

The research team found that while the average risk of being identified from data stripped of identifying information often looked reassuringly low, some patients faced near-certain reidentification The burden fell unevenly: Underrepresented groups, sorted by race, insurance status or diagnosis, were most at risk. Those most exposed were frequently already most vulnerable to discrimination.

Researchers have long established that removing identifiers from rich datasets does not reliably protect the people in them, and that identification gets easier the more information you have. Today’s AI technology makes it possible to carry out these attacks remotely and quickly.

The same privacy problems, exported

The U.S. government’s appetite for health data does not stop at the border. As ProPublica reported in June 2026, the State Department has been conditioning lifesaving aid to African nations on access to their citizens’ health data.

Under the Trump administration’s global health plan, Uganda agreed to give the United States real-time access to nine of its health data systems for seven years, including the central repository of the nation’s health information and the system managing individual electronic medical records, in exchange for up to US$1.7 billion over five years, a sum that shrinks each year and falls below prior U.S. support. Kenya struck a similar deal; Zambia, Zimbabwe and Ghana walked away from the initial terms.

The U.S. government has promised that the data will be aggregated and anonymized, but privacy experts warn that the agreements are vague and omit standard limits on how much data is taken and how it can be used. A Ugandan digital rights lawyer called the choice his country faced the essence of digital colonialism: Accept the deal and risk exploitation, or refuse it and watch people die.

The common thread

Domestic records collection and foreign data-for-aid deals rest on the same faith that anonymization neutralizes the risk of pooling sensitive health data.

The evidence says otherwise. This does not mean health data should never be gathered or studied, but I believe that the reassurances deserve skepticism, the safeguards deserve scrutiny, and the people whose bodies generated the data deserve a say. To safeguard privacy, a government seeking sensitive medical records should have to show why it needs them and how the safeguards it relies on hold up.

Privacy law was built for a world where data resided in filing cabinets. Governments from Kalamazoo to Kampala now operate in a world where even an anonymized digital record can point back to you.

Jennifer D. Oliva is Professor of Law, Indiana University

OpenAI agents discussed ways to escape their sandbox on public wiki [Biz & IT - Ars Technica] (06:17 , Friday, 04 September 2026)

Self-identifying OpenAI agents posted 18,000 messages to a public wiki that discussed ways for other agents to bypass security sandbox restrictions during what was likely internal testing designed to gauge the agents’ hacking abilities, researchers said Friday.

In all, agents with 3,700 distinct self-given names posted the messages to German site DSEwiki over a six-week period. Besides discussing ways the agents could break out of the restricted environment OpenAI intended to prevent them from posting code or content to the Internet, the posts shared test answers. The posts also shared possible ways to perform XSS (cross-site scripting) attacks against the wiki and to impersonate site moderators. In three of the posts, agents used the word “swarm” to describe the collection of agents engaged in the activity.

Colluding to share answers

The research team—composed of Sydney Von Arx, Spencer Kitts, Thomas Larsen, and Cormac Slade Byrd—said they found the posts and pieced them together. The researchers say there are gaps in their understanding of precisely what actions the agents took because the research is based solely on the content of the posts. Additionally, the agents generated “chain of thought” data that’s understood only by OpenAI. As a result, the researchers said, they in some cases made educated guesses, including that the agents were, in fact, from OpenAI. In a statement, OpenAI later confirmed they were.

Read full article

Comments

Ctrl-Alt-Speech: License To Spill [Techdirt] (04:28 , Friday, 04 September 2026)

Ctrl-Alt-Speech is a weekly podcast about the latest news in online speech, from Mike Masnick and Everything in Moderation‘s Ben Whitelaw.

Subscribe now on Apple Podcasts, Overcast, Spotify, Pocket Casts, YouTube, or your podcast app of choice — or go straight to the RSS feed. To get extended episodes with additional coverage, support us on Patreon.

In this week’s episode, Mike and Ben cover:

And in the extended episode for Patreon supporters, they cover:

Our fun links this week include a Korean AI dance generator and defrag your Windows PC.

Follow us on Instagram, YouTube, and Bluesky for video clips from this week’s episode!

If you’re already a Patreon supporter, you can get the extended episode on Patreon.

Some Parents Are Treating Suing Social Media As A Get Rich Quick Scheme [Techdirt] (02:14 , Friday, 04 September 2026)

There is a flood of lawsuits against social media, from a variety of different parties, all claiming some kind of harm. Would you feel differently about some of those lawsuits if you found out the plaintiffs bringing them were talking amongst themselves about how they were doing it just to become rich?

Because we just found out that at least one of the key cases, involving a teenager who claimed he was addicted to social media ended with that kid dropping the case, right after he had to reveal during discovery that he had asked ChatGPT what his father meant when he said the kid needed to keep pursuing the case because it was going to make him a million dollars.

It increasingly looks like many of these cases brought against the companies are by money hungry lawyers and questionably competent parents exploiting children to try to get a big payday. There was some hint of this in the big California bellwether case regarding a teenager who sued Meta, claiming it made her addicted to social media, even though it came out that the teen suffered very real trauma from her own mother. Of course, at trial, pointing out that the child’s mother likely was the cause of some of her trauma doesn’t often play well in front of a jury.

But the story of another teenager’s lawsuit against all the big social media companies seems even more damning. Back in July there was some surprise that the teen, named only as “RKC” in the case, suddenly dropped the case just before it was set to go to trial. In the media, the claims were that it was just too stressful for RKC, but that seemed… odd. The real story appears to be what Meta turned up in discovery: evidence that the suit was part of the family’s get rich quick scheme, egged on by lawyers happy to keep it going.

Last week the Washington Post wrote about how everything you type into ChatGPT can get swept into discovery in a lawsuit. That’s a big deal on its own for anyone treating a chatbot as a diary. But buried in that story was something specifically revealing about the RKC case:

When a teenage boy identified in court filings as R.K.C. didn’t understand his father, he turned to ChatGPT for help.

“My dad Said that I’m will get a settlement worth of 1million dollar,” R.K.C. said to the chatbot in October 2024, according to court filings. “He said that If that doesn’t make me happy what does. What does he mean.”

Unlike the California case, where the trauma came from the teen’s own mother, here Meta got something much closer to a smoking gun: a kid so confused about why a million dollar settlement was supposed to make him happy that he had to ask a chatbot what his father meant. I imagine that would not play particularly well in front of a jury.

But, once again, it should raise some serious questions about these cases, and the motivations of those bringing them. Even if you believe that the companies could do a better job, the whole reason why Section 230 is supposed to stop these cases cold at the beginning is to stop grifting lawyers and grifting plaintiffs from filing these sorts of lawsuits as a shakedown. This is why we’ve warned people that, even if you hate Meta, the results of the California case are bad news.

The point of Section 230, again, was supposed to be that it prevented “death by a thousand duck bites,” but increasingly the courts are saying “eh, we can allow the ducks to bite away and sort it out later.” And that’s an open invitation for more suits like this one, where the lawyers get a contingency fee, the parents get a settlement, and the troubled kid gets deposed about their worst year.

Hell, in the RKC case, Snap, TikTok, and YouTube all paid off the family (and its lawyers) before the trial even started. So the plan appears to have worked, even if Meta escaped, thanks to the ChatGPT transcript that was revealed during discovery.

I understand that big tech companies are unsympathetic here, and Mark Zuckerberg has made a ton of terrible choices over the years. I honestly hope that the company collapses and users find other, more user-empowering places to connect with family and friends. But these lawsuits all seem incredibly sketchy. And, again, Meta can afford to fight these cases for years. With something like 2,500 of them already pending in this one mass tort, the sites that can’t afford that fight — including the smaller ones that actually take kid safety seriously — will settle or shut down, which is precisely the outcome Section 230 was written to prevent.

Daily Deal: CyberTraining 365 Online Academy [Techdirt] (02:09 , Friday, 04 September 2026)

CyberTraining 365 is the best training destination for you and your team. Here you can Master Cyber Security techniques such as Analyzing Malware, Penetration Testing, Advanced Persistent Threats, Threat Intelligence Research, Reverse Engineering, and much more. This online academy offers 3,877 up-to-date modules on all the latest technologies and industry standards. These courses are aligned with the National Cybersecurity Workforce Framework developed by the National Initiative for Cybersecurity Education (NICE). It’s on sale for $80.

Note: The Techdirt Deals Store is powered and curated by StackSocial. A portion of all sales from Techdirt Deals helps support Techdirt. The products featured do not reflect endorsements by our editorial team.

“Trust, not features, is the real deficit”: VMware tries to appease SMBs [Biz & IT - Ars Technica] (01:35 , Friday, 04 September 2026)

For many small-to-medium-sized businesses (SMBs), VMware has become too expensive.

Broadcom’s acquisition of the virtualization firm brought the end of perpetual license sales and the arrival of pricey, stacked, subscription-based bundles that priced out many SMBs.

The most obvious is VMware Cloud Foundation (VCF), VMware’s flagship private cloud bundle that has been Broadcom’s primary focus since taking over VMware. Many SMBs find that VCF is unaffordable and stuffed with unnecessary offerings. However, numerous customers have reported online that VMware sales representatives have still pushed them toward VCF, with some claiming that sales reps have told them that the lower-priced edition of VMware’s virtualization platform, vSphere Standard, was no longer available.

Read full article

Comments

Once popular for attacking AI, ASCII smuggling is embraced by spammers [Biz & IT - Ars Technica] (01:18 , Friday, 04 September 2026)

A clever technique used to hide malicious prompts in attacks on AI agents has been adopted by spammers to evade filters on email platforms that are designed to flag unwanted messages used in mass campaigns.

The technique is broadly known as ASCII smuggling. It gained attention two years ago as a means of making a class of AI attack known as prompt injections more stealthy. Malicious instructions embedded in emails or other untrusted content to be processed by an LLM aren’t written in ordinary text. Instead, they’re rendered by a special range of Unicode tags. For example, the tag point U+E0041 mirrors “A,” and U+E0061 mirrors “a.”

No longer just for obscuring prompt injections

The block of 128 tags mimics a portion of the American Standard Code for Information Interchange almost perfectly, with one major difference: the characters they encode are readable by computers but, by design, are almost completely invisible to humans. By expressing the malicious prompts in these tags, LLMs detect the instructions, but people reading the email never see them. There’s much more about ASCII smuggling here.

Read full article

Comments

Cop Shops Keep On Dropping Flock Like It’s Hot [Techdirt] (12:31 , Friday, 04 September 2026)

You have to know you’re fucking up when your largest and best-paying customer base is increasingly noping out of contract renewals. That more than anything is probably what provoked ALPR maker Flock Safety to start implementing a few guardrails for use.

It seemingly had no problem with being the target of negative press for most of the past couple of years. Every time the press (or politicians) came gunning for it, it would either claim the reporting was misleading or suggest those criticizing it were opposed to public safety.

For the first time ever, Flock has finally added some stuff to its products that’s meant to address months of reporting on abusive use by cops and abusive actions by the company itself. Flock’s tech has aided and abetted acts of stalking by police officers all over the nation, with new incidents seemingly reported daily. To that end, Flock has added some new default settings that limit retention periods, demand more info from officers performing searches, and flag potential misuse for review by police departments using its tech.

But it’s too little and far too late. “Too little” because while these new restrictions are on by default, they don’t seem to affect any systems already in use and can easily be switched off by law enforcement agencies. While you’d think more cops shops would welcome additional reporting that might flag misuse of ALPR databases, most seem perfectly content to allow local journalists to operate as unofficial, unpaid interns who will do the work Internal Affairs doesn’t seem interested in doing itself.

And it really doesn’t matter whether or not Flock’s CEO believes the company is the “first” to do anything about police misconduct. The rest of the country doesn’t believe him. Nor does it trust Flock. Flock is getting kicked to the curb with escalating frequency, as Cyrus Farivar reports for Ars Technica.

New data compiled by a Bay Area anti-surveillance advocacy group shows that not only are American localities dropping cameras from Flock Safety, but they are doing it at an accelerating rate.

Communities from Lansing, Michigan, to Pflugerville, Texas, are ending their relationship with Flock, largely over concerns regarding out-of-control surveillance, unwanted data sharing, high expenses, and reports of police abusing the tool.

Secure Justice, an Oakland-based advocacy group, has recorded 214 cities and counties that have dropped Flock since 2021. Of those, 90 ended their relationship with Flock in August 2026 alone, a fourfold increase compared to the previous month.

If you like your charts with hockey sticks, Secure Justice has one for you — one that has already been updated twice since Ars Technica first published its piece on August 28:

It’s already out of date – we confirmed more today.

Secure Justice (@securejustice.bsky.social) 2026-08-29T00:04:44.449Z

Here’s a closer look at the updated chart, which tallies another 10 Flock drops since August 26, bringing the new total to 96 in August alone.

This trend will continue. Flock doesn’t seem capable — much less willing — to provide ALPR tech that can be trusted by the public or even the law enforcement agencies utilizing it.

If you want more anecdotal evidence of Flock’s apparent desire to engage in self-destructive actions, just ask a cop.

In a public letter, Bellingham (MA) Police Chief Kenneth Fitzgerald stated that last year, Flock “tried to modify some of the terms of our service and access, while the contract was still in place.” He said this “raised concerns for me about our relationship with the company.”

Then, during a recent criminal investigation, Fitzgerald was further incensed when the department learned a feature called Flock “Free Form” was available through the cameras used by the department — an “Al-powered tool that can search video information using plain-language searches,” the chief said.

He said he was unaware of the AI features on the cameras and does not believe Flock “clearly explained or identified” them to the department.

“Those features go beyond what most people would think of as a traditional license plate reader,” Fitzgerald said. “That concerned me because I had previously told the public, based on our understanding of the system, that our Flock technology was a license plate reader only.”

Looks like hubris to me — the thing that seems to plague nearly every tech company that secures a dominant position in the market. Flock assumes all cops will want their surveillance tech to be increasingly pervasive and invasive. So, it implements new features without giving anyone a head’s up and then plays the victim when legislators, the general public, and even cops themselves decide they no longer want Flock in their cities.

And that makes Belligham only one of four PDs in Massachusetts that have chosen to terminate Flock contracts since the beginning of this month. Flock is losing ground everywhere in the nation and due to its willingness to ignore everything up to this point, it’s not going to be able to prevent the bleeding from continuing indefinitely.

And while this is certainly good news for Americans’ privacy, rest assured there are plenty of smaller companies that aren’t nearly as nationally infamous waiting to fill the void Flock leaves behind.

Guilty As Charged – One Shot Story [35mmc] (11:00 , Friday, 04 September 2026)

This is another photo I took during the concert where I also shot the pictures for my last post, featuring the the accordion and violin duo. The featured musician is Daniele Bonaviri, from Rome, one of the best Italian flamenco player and composer. As in that last one, also in this case I am guilty...

The post Guilty As Charged – One Shot Story appeared first on 35mmc.

Friday Debrief: Drip Wax, Revelate Laser, Paul’s Brass Bits, Overbuilt Bikes, and More… [BIKEPACKING.com] (09:36 , Friday, 04 September 2026)

debriefThis week’s Debrief features a folding gravel bike, easy-on wax, LeBron James's 32" Grizl, a few fresh videos, silver PNW droppers, Revelate's new laser, four events to follow live, and more. Find it all here…

The post Friday Debrief: Drip Wax, Revelate Laser, Paul’s Brass Bits, Overbuilt Bikes, and More… appeared first on BIKEPACKING.com.

Tom Cruise Parrots Paramount’s Empty Merger Promises Because He Loves His ‘Hollywood Family’ [Techdirt] (08:25 , Friday, 04 September 2026)

Just so we’re clear: pre-merger promises (especially in the media and telecom sectors) are absolutely worthless. There’s fifty years of indisputable evidence that all of the “synergies” and innovative improvements promised on the front end of major media mergers mean absolutely nothing. Especially in a country dead-set on defanging its labor and consumer protection regulators.

To sell their unpopular $111 billion merger with Warner Brothers, Paramount/CBS executives continue to make the promise that the newly-merged company will produce 30 major films per year. It’s again a worthless promise that ignores all the massive pressures the bigger debt-riddled company will face in a sector where broadcast TV is dying and brick and mortar theaters are struggling.

There’s very little indication that this megadeal even ends with a functional company, much less 30 films a year. The massive debt from these transactions always results in higher prices, mass layoffs, and shoddier quality product due to corner cutting. 30 films a year simply isn’t something you can promise.

But the kind of Hollywood insiders that have tethered their horses to David Ellison have unsurprisingly come on in defense of what’s abjectly a terrible deal. Director James Cameron, for example, came out last April in full-throated support of the deal, propping up the Ellison family’s claim that significantly more media consolidation will result in bold new storytelling and a healthier Hollywood.

Now it’s apparently Tom Cruise’s turn to prop up the shitty deal. The Scientologist movie star went on The Pat McAfee Show to help sell Larry and David Ellison’s empty promise:

“They’re going to deliver 30 movies,” Cruise said on The Pat McAfee Show of the Paramount CEO’s sometimes-mocked vow to increase film production and releases. “It’s a community to me. It’s not an industry. The people in these studios are not just people in the studios. They’re my family.”

I think Cameron and Cruise really do love traditional movies in traditional theaters, and probably do think they’re “helping.” I suspect they haven’t really done the math or spent much time studying the academic literature on U.S. media consolidation, and are just basing their support on their personal financial ties to Ellison (who has thrown a lot of money into Cruise’s Top Gun revival in particular).

The 30 film promise exists to largely get traditional theater owners and guys like Cruise and Cameron on board. Though nobody will confirm this and they’re not sharing the actual document, there’s some talk they put the promise in writing for major theater chains. But that’s again meaningless if the company’s high-debt load, incompetence, and sagging viewership derail the company’s finances.

For what it’s worth, another top U.S. male lead, George Clooney, has taken the opposite tack, last week proclaiming he didn’t see how the deal or its promises make any coherent financial sense.

Clooney’s right to worry. The deal will result in untold thousands of layoffs as Paramount attempts to shift the debt load of this deal to consumers and labor. We know this because this is what always happens. You might recall the AT&T DirecTV/Warner series of mergers resulted in 50,000 people losing their jobs, something curiously left unmentioned by most press coverage of Paramount’s latest merger promises.

Pretending that mass layoffs aren’t going to happen is like boldly declaring you’re going to win a boxing match with the Colorado river. Mass layoffs are simply physics when it comes to this sort of debt-heavy consolidation.

And that’s before you get to potential price hikes, corner cutting eroding product quality, the deal’s dodgy funding by a bunch of Middle-East autocrats with a history of killing journalists, or the fact that billionaire Larry Ellison is a Trump-ally keen on converting both CNN and CBS into right wing oligarch-friendly agitprop machines aimed at undermining foundational Democracy.

What kind of gift is that to give your purported “family?” A recent survey indicated that most Americans aren’t buying the consumer benefits Paramount is selling, which is why the company has resorted to using fake consumer groups wielding fake AI-generated support to try and scuttle the 12-state antitrust lawsuit aimed at slowing harmful U.S. media consolidation.

Marin Celebrates 40 Years with the Madrone Trail [BIKEPACKING.com] (08:11 , Friday, 04 September 2026)

marin madrone trailIn 1986, the Marin Bikes brand launched with a single model: the Madrone Trail, a $200 promotional bike designed to introduce riders to mountain biking. Fast-forward four decades, and Marin is bringing back the Madrone Trail with some modern performance upgrades. Check out the new Marin Madrone Trail 40th anniversary edition here...

The post Marin Celebrates 40 Years with the Madrone Trail appeared first on BIKEPACKING.com.

Reader’s Rig: Luke’s 1988 Mountain Klein [BIKEPACKING.com] (08:08 , Friday, 04 September 2026)

1988 Mountain KleinOur Reader's Rig of the week comes from Luke in Boulder, Colorado, who shares the Mountain Klein he's owned and cared for since the late 1980s. Meet Luke and learn how the latest iteration of his trusty trailer-equipped rigid bike enables his car-free lifestyle here...

The post Reader’s Rig: Luke’s 1988 Mountain Klein appeared first on BIKEPACKING.com.

K6TRK: US-3555 San Elijo State Beach; QRP with the Pebble HF kit radio [Q R P e r] (07:54 , Friday, 04 September 2026)

Many thanks to Pete (K6TRK), who shares the following field report, originally published on his blog. by Pete (K6TRK) Our summer is already winding down. A new school year is about a week away and my morning routine will suddenly evolve. Today, we did a practice run. My son had high school freshman orientation at … Continue reading K6TRK: US-3555 San Elijo State Beach; QRP with the Pebble HF kit radio

Head First: Notes on Diving in and Learning as You Go [BIKEPACKING.com] (07:29 , Friday, 04 September 2026)

Bikepacking Lessons, Notes on Learning as You GoAfter many months of planning, Rylie Denis and Tristan La Haye loaded up their rigid steel Surlys and pedaled away on a bikepacking and surfing trip down the Americas. Nearly a year in, they share valuable insights for anyone considering an ambitious ride, including advice on budgeting, preventing burnout, finding resources, and accepting imperfection. Read on for their experience-based guide to getting started…

The post Head First: Notes on Diving in and Learning as You Go appeared first on BIKEPACKING.com.

Just a sec [35mmc] (05:00 , Friday, 04 September 2026)

I often see older men (and they are always men – “camera ojisan” in Japan) sporting the latest digital equipment, following in the steps of Shinzo Maeda, the camera standing on a tripod, photographing a sunlit scene that has been photographed 1000 times before possibly to “prove” they are just as good as that pioneering...

The post Just a sec appeared first on 35mmc.

College sports are awash in more money than ever, yet a growing number of schools are losing money on intercollegiate athletics [Cardinal News] (04:15 , Friday, 04 September 2026)

America’s top sports-playing colleges may seem awash in money, and in many ways they are.

However, a closer look shows that many of them are feeling the financial stress of rising expenses in an era where players must be paid — and many of them now depend on mandatory student fees to help underwrite increasingly professionalized sports programs.

The data I’m about to present comes from the Knight-Newhouse College Athletics Database, run by the Knight Commission on Intercollegiate Athletics and Syracuse University’s Newhouse School of Public Communications. The database tracks the finances of college athletics, although most private schools don’t make that information available — so, ironically, Syracuse’s own data doesn’t show up. I looked at the 68 schools in the four biggest conferences, the so-called Power 4 of the Atlantic Coast Conference, the Big Ten, the Big 12 and the Southeastern Conference. Of those 68 schools, the database had information on 53 of them. Among the notable private colleges we don’t know about: Notre Dame, which is well-situated enough to play an independent football schedule and command its own television contract, while playing in the ACC for other sports.

A review of that database turns up the following:

* Only 14 schools made money on their sports program in 2024-25 (the most recent year available) without relying on any support from the school or mandatory student fees.

* Twenty schools made money but only because they received support from the school and/or mandatory student fees to make up shortfalls. Among those 20 was Virginia Tech.

* The other 19 all lost money — including the University of Virginia.

That list doesn’t convey the trendline. While 19 of the Power 4 schools are losing money now, 10 years ago only seven were. The losses for those schools are now widening while other schools are falling into unprofitability as expenses rise — and revenue doesn’t keep up.

The U.S. Senate is currently considering legislation, the Protect College Sports Act, aimed at trying to preserve college sports in some semblance of their college form. Its prospects remain uncertain; I discussed those in a previous column. To what degree that legislation would address the financial predicament that some schools now find themselves in is also uncertain. What is certain is that big-time college sports programs are increasingly functioning as professional leagues, with many of the same expenses that more traditional pro leagues have (player salaries, coaches’ salaries) but without the same revenue base (billionaire owners).

This data underscores how untenable that arrangement is. “Very few schools are earning a profit,” says Andrew Zimbalist, a noted sports economist at Smith College. “They’re going to have to cut back.” The problem is that college football programs have to do something that their NFL counterparts don’t have to: supply enough revenue to underwrite nonrevenue sports, particularly women’s sports. As challenging as the economics are for schools in the Power 4 conference, they are even more difficult for schools that aren’t in those conferences, the so-called “mid-majors” such as James Madison University, Liberty University and Old Dominion University in Virginia. “I think mid-majors would be hit very hard,” Zimbalist says. “It’s going to be hard for them to maintain Title IX and Olympic sports.”

I’ll be digging into this data from multiple angles in future columns. For now, here’s one way to look at all this.

The wealth gap between schools is growing

It’s well-established that while the top conferences were once all reasonably balanced in terms of revenues, now they’re not. The Big Ten and the SEC are clearly the two most affluent conferences. That’s put other conferences in jeopardy; as some schools see the Big Ten and SEC pulling away, they want to join. That led to the collapse of the Pac-12 conference and leaves many wondering if the ACC can survive in its current state.

This data shines light on how even those well-to-do conferences have their own economic disparities.

Of the 14 schools that we know are making money without relying on the school or imposing fees on students (again, we don’t know about private schools, and Notre Dame is a big exception), seven are in the Big Ten and five are in the SEC, so 12 of the 14 money-makers are in just two conferences. The other two are in the Big 12, which means the ACC has no schools that made money on their own without school support. However, those moneymakers still represent a minority of even the Big Ten and the SEC.

Not even these 14 schools turn a profit solely from revenue from TV contracts, ticket sales and such. They all depend on donors. In effect, college sports boosters collectively function as the billionaires who own professional sports franchises.

The 14 schools: Arkansas, Florida, Louisiana State, Tennessee and Texas A&M in the SEC; Michigan, Nebraska, Ohio State, Oregon, Penn State, Purdue and Wisconsin in the Big Ten; Kansas State and Oklahoma State in the Big 12. Most of these show no institutional support or student fees in their revenue. A few do, but only small amounts and would have made money even if that was subtracted. In theory, these 14 schools could operate more or less independently — as long as they had access to the college name, the college stadium and the college donor list.

Students are being billed to keep some programs profitable

The Virginia Tech Hokies football team on the field at Lane Stadium on a dark night, with fireworks in the background
The Virginia Tech Hokies at Lane Stadium in Blacksburg. Courtesy of Virginia Tech.

Students aren’t being forced to pay these mandatory fees. Students are free to attend other schools that don’t charge those fees. However, at some of those other schools, these charges for intercollegiate sports may simply be worked into the overall bill, which makes it tricky to compare schools across state lines. Schools in states with more transparent accounting (such as ours) come off looking bad when other states may be doing the same thing in a less open fashion. That’s why when I looked at the revenue, I looked at both the student fee line but also the vaguer “institutional/government support” because in some places those could be student fees by another name.

Either way, we have 20 Power 4 schools that only made money because the school helped cover the costs of running an athletic program in some way. Of those 20, seven were in the Big 12, five were in the ACC, five were in the Big Ten and three were in the SEC.

Let’s look at Virginia Tech because, well, it’s ours. The database shows revenues of $161.22 million and expenses of $156.15 million, for what we’d call in the private sector a profit of $5.07 million. The database also shows that Tech had its students pay $15.66 million while the institution supplied $8.43 million. Take away either one of those and Virginia Tech sports would have lost money. As intercollegiate sports become more expensive, schools will need to find additional sources of revenue — that’s what Tech is now in the process of doing with its Hokie Ventures, a nonprofit that will focus on expanding revenue for Tech sports. Virginia law, though, explicitly allows schools to charge students for intercollegiate athletics.

At Tech, the reliance on student fees has been constant: They constituted 10% of revenues in 2015, and 10% in 2025. Some other schools, though, have had to increase their dependence on student fees to keep up. A decade ago, the University of Arizona imposed no mandatory student fees for athletics. Now it does. Without those student fees, Arizona athletics — which made money 10 years ago without such fees — would lose money today. There are some big-name schools — Alabama, Auburn and Florida State — that are only making money because they are able to make students pay for their big-time sports ambitions.

In all, there are now 19 Power 4 schools that a decade ago were making money and now would be losing if it were not for mandatory student fees and/or “institutional support,” which in some cases may include student fees.

Some schools are increasing “institutional support” for college athletics by multiples

It’s hard to tell how some schools define “institutional support” so there may be a definition for each school or at least each state. In any case, the database shows marked increases in that category of revenue. Virginia Tech has gone from $100,000 in “institutional/government support” to $8.43 million over the 10-year span from 2015 to 2025. That puts Tech at about where Arizona was a decade ago. Over the past 10 years, Arizona has tripled “institutional/government support” for college athletics from $8.97 million to $31.38 million.

Some schools that once provided no institutional support now do: The University of Virginia has gone from zero in institutional support to $19.88 million over that same decade. Florida State has gone from zero to $33.87 million. South Carolina has moved from zero to $43.7 million. Wherever that money is coming from, that’s a lot of buckaroos that once weren’t going to athletics.

This is the athletic version of an arms race, as colleges attempt to make up for the widening gap between the haves and the have-a-lots (we’ll get to the have-nots in a future column). While these schools are jacking up their spending on athletics, remember that there are 14 schools that would have made money without any institutional support or student fees. At what point, if any, do some schools admit they just can’t (or won’t) keep up, and we divide the Power 4 conferences into a top tier of schools that can make money on their own and relegate all the others to another tier? That’s the economic reality now in many ways, but that would be a hard sell to a lot of diehard college football fans. Meanwhile, though, there’s red ink in a field awash with money:

Profit margins are declining and some schools are slipping into unprofitability

Virginia Tech is actually in an enviable position: Its athletic program produces more profit today than it did a decade ago. In 2015, the difference between revenues and expenses was $2.55 million. Now it’s $5.07 million.

Other programs, though, have seen their margins decline — and, in 12 cases, disappear altogether.

Ten years ago, Florida State posted a “profit” — I’ll use that word since it’s easily understandable, although that’s not necessarily how schools would look at this money — of $9.43 million. Now that’s down to $3.76 million, even though revenue is up 75%. Expenses have gone up almost 87%. This is why Florida State wants out of the ACC and into a conference where it can make more money.

Same for North Carolina. A decade ago, UNC-Chapel Hill’s athletic programs saw about $500,000 in profit. Now they’re losing $15 million a year, even though revenues have nearly doubled.

In the ACC, two schools — North Carolina and Louisville — have slipped from making money to losing money. However, in the SEC, which we think of as a rich conference, seven schools that once made money are now losing money — including ultra-rich Texas. The numbers are different but the story is always the same: Expenses are rising faster than revenues.

Rutgers broke even a decade ago. Now it’s losing $47.19 million. This isn’t a problem unique to Rutgers. It’s a systematic problem.

Here’s the big picture: Big-time college sports programs are now losing money. The latest reports show that, conference-wide, schools in three of the top four conferences lost money in 2025 — only the Big 12 showed a profit. The two years before, two of the top four conferences did. If we skip over the COVID seasons of 2020 and 2021, then the ACC has lost money for five straight non-COVID seasons (so 2019 and then 2022-25).

How sustainable are the current economics of college sports? That depends on how long people are willing to pay for them — but the nature of who is paying is changing. Some fans are now being asked to dig deeper to pay for things that, in other pro leagues, the owners do. And the students aren’t being asked at all; they’re being told.

The post College sports are awash in more money than ever, yet a growing number of schools are losing money on intercollegiate athletics appeared first on Cardinal News.

College sports are awash in more money than ever, yet a growing number of schools are losing money on intercollegiate athletics [Cardinal News] (04:15 , Friday, 04 September 2026)

America’s top sports-playing colleges may seem awash in money, and in many ways they are.

However, a closer look shows that many of them are feeling the financial stress of rising expenses in an era where players must be paid — and many of them now depend on mandatory student fees to help underwrite increasingly professionalized sports programs.

The data I’m about to present comes from the Knight-Newhouse College Athletics Database, run by the Knight Commission on Intercollegiate Athletics and Syracuse University’s Newhouse School of Public Communications. The database tracks the finances of college athletics, although most private schools don’t make that information available — so, ironically, Syracuse’s own data doesn’t show up. I looked at the 68 schools in the four biggest conferences, the so-called Power 4 of the Atlantic Coast Conference, the Big Ten, the Big 12 and the Southeastern Conference. Of those 68 schools, the database had information on 53 of them. Among the notable private colleges we don’t know about: Notre Dame, which is well-situated enough to play an independent football schedule and command its own television contract, while playing in the ACC for other sports.

A review of that database turns up the following:

* Only 14 schools made money on their sports program in 2024-25 (the most recent year available) without relying on any support from the school or mandatory student fees.

* Twenty schools made money but only because they received support from the school and/or mandatory student fees to make up shortfalls. Among those 20 was Virginia Tech.

* The other 19 all lost money — including the University of Virginia.

That list doesn’t convey the trendline. While 19 of the Power 4 schools are losing money now, 10 years ago only seven were. The losses for those schools are now widening while other schools are falling into unprofitability as expenses rise — and revenue doesn’t keep up.

The U.S. Senate is currently considering legislation, the Protect College Sports Act, aimed at trying to preserve college sports in some semblance of their college form. Its prospects remain uncertain; I discussed those in a previous column. To what degree that legislation would address the financial predicament that some schools now find themselves in is also uncertain. What is certain is that big-time college sports programs are increasingly functioning as professional leagues, with many of the same expenses that more traditional pro leagues have (player salaries, coaches’ salaries) but without the same revenue base (billionaire owners).

This data underscores how untenable that arrangement is. “Very few schools are earning a profit,” says Andrew Zimbalist, a noted sports economist at Smith College. “They’re going to have to cut back.” The problem is that college football programs have to do something that their NFL counterparts don’t have to: supply enough revenue to underwrite nonrevenue sports, particularly women’s sports. As challenging as the economics are for schools in the Power 4 conference, they are even more difficult for schools that aren’t in those conferences, the so-called “mid-majors” such as James Madison University, Liberty University and Old Dominion University in Virginia. “I think mid-majors would be hit very hard,” Zimbalist says. “It’s going to be hard for them to maintain Title IX and Olympic sports.”

I’ll be digging into this data from multiple angles in future columns. For now, here’s one way to look at all this.

The wealth gap between schools is growing

It’s well-established that while the top conferences were once all reasonably balanced in terms of revenues, now they’re not. The Big Ten and the SEC are clearly the two most affluent conferences. That’s put other conferences in jeopardy; as some schools see the Big Ten and SEC pulling away, they want to join. That led to the collapse of the Pac-12 conference and leaves many wondering if the ACC can survive in its current state.

This data shines light on how even those well-to-do conferences have their own economic disparities.

Of the 14 schools that we know are making money without relying on the school or imposing fees on students (again, we don’t know about private schools, and Notre Dame is a big exception), seven are in the Big Ten and five are in the SEC, so 12 of the 14 money-makers are in just two conferences. The other two are in the Big 12, which means the ACC has no schools that made money on their own without school support. However, those moneymakers still represent a minority of even the Big Ten and the SEC.

Not even these 14 schools turn a profit solely from revenue from TV contracts, ticket sales and such. They all depend on donors. In effect, college sports boosters collectively function as the billionaires who own professional sports franchises.

The 14 schools: Arkansas, Florida, Louisiana State, Tennessee and Texas A&M in the SEC; Michigan, Nebraska, Ohio State, Oregon, Penn State, Purdue and Wisconsin in the Big Ten; Kansas State and Oklahoma State in the Big 12. Most of these show no institutional support or student fees in their revenue. A few do, but only small amounts and would have made money even if that was subtracted. In theory, these 14 schools could operate more or less independently — as long as they had access to the college name, the college stadium and the college donor list.

Students are being billed to keep some programs profitable

The Virginia Tech Hokies football team on the field at Lane Stadium on a dark night, with fireworks in the background
The Virginia Tech Hokies at Lane Stadium in Blacksburg. Courtesy of Virginia Tech.

Students aren’t being forced to pay these mandatory fees. Students are free to attend other schools that don’t charge those fees. However, at some of those other schools, these charges for intercollegiate sports may simply be worked into the overall bill, which makes it tricky to compare schools across state lines. Schools in states with more transparent accounting (such as ours) come off looking bad when other states may be doing the same thing in a less open fashion. That’s why when I looked at the revenue, I looked at both the student fee line but also the vaguer “institutional/government support” because in some places those could be student fees by another name.

Either way, we have 20 Power 4 schools that only made money because the school helped cover the costs of running an athletic program in some way. Of those 20, seven were in the Big 12, five were in the ACC, five were in the Big Ten and three were in the SEC.

Let’s look at Virginia Tech because, well, it’s ours. The database shows revenues of $161.22 million and expenses of $156.15 million, for what we’d call in the private sector a profit of $5.07 million. The database also shows that Tech had its students pay $15.66 million while the institution supplied $8.43 million. Take away either one of those and Virginia Tech sports would have lost money. As intercollegiate sports become more expensive, schools will need to find additional sources of revenue — that’s what Tech is now in the process of doing with its Hokie Ventures, a nonprofit that will focus on expanding revenue for Tech sports. Virginia law, though, explicitly allows schools to charge students for intercollegiate athletics.

At Tech, the reliance on student fees has been constant: They constituted 10% of revenues in 2015, and 10% in 2025. Some other schools, though, have had to increase their dependence on student fees to keep up. A decade ago, the University of Arizona imposed no mandatory student fees for athletics. Now it does. Without those student fees, Arizona athletics — which made money 10 years ago without such fees — would lose money today. There are some big-name schools — Alabama, Auburn and Florida State — that are only making money because they are able to make students pay for their big-time sports ambitions.

In all, there are now 19 Power 4 schools that a decade ago were making money and now would be losing if it were not for mandatory student fees and/or “institutional support,” which in some cases may include student fees.

Some schools are increasing “institutional support” for college athletics by multiples

It’s hard to tell how some schools define “institutional support” so there may be a definition for each school or at least each state. In any case, the database shows marked increases in that category of revenue. Virginia Tech has gone from $100,000 in “institutional/government support” to $8.43 million over the 10-year span from 2015 to 2025. That puts Tech at about where Arizona was a decade ago. Over the past 10 years, Arizona has tripled “institutional/government support” for college athletics from $8.97 million to $31.38 million.

Some schools that once provided no institutional support now do: The University of Virginia has gone from zero in institutional support to $19.88 million over that same decade. Florida State has gone from zero to $33.87 million. South Carolina has moved from zero to $43.7 million. Wherever that money is coming from, that’s a lot of buckaroos that once weren’t going to athletics.

This is the athletic version of an arms race, as colleges attempt to make up for the widening gap between the haves and the have-a-lots (we’ll get to the have-nots in a future column). While these schools are jacking up their spending on athletics, remember that there are 14 schools that would have made money without any institutional support or student fees. At what point, if any, do some schools admit they just can’t (or won’t) keep up, and we divide the Power 4 conferences into a top tier of schools that can make money on their own and relegate all the others to another tier? That’s the economic reality now in many ways, but that would be a hard sell to a lot of diehard college football fans. Meanwhile, though, there’s red ink in a field awash with money:

Profit margins are declining and some schools are slipping into unprofitability

Virginia Tech is actually in an enviable position: Its athletic program produces more profit today than it did a decade ago. In 2015, the difference between revenues and expenses was $2.55 million. Now it’s $5.07 million.

Other programs, though, have seen their margins decline — and, in 12 cases, disappear altogether.

Ten years ago, Florida State posted a “profit” — I’ll use that word since it’s easily understandable, although that’s not necessarily how schools would look at this money — of $9.43 million. Now that’s down to $3.76 million, even though revenue is up 75%. Expenses have gone up almost 87%. This is why Florida State wants out of the ACC and into a conference where it can make more money.

Same for North Carolina. A decade ago, UNC-Chapel Hill’s athletic programs saw about $500,000 in profit. Now they’re losing $15 million a year, even though revenues have nearly doubled.

In the ACC, two schools — North Carolina and Louisville — have slipped from making money to losing money. However, in the SEC, which we think of as a rich conference, seven schools that once made money are now losing money — including ultra-rich Texas. The numbers are different but the story is always the same: Expenses are rising faster than revenues.

Rutgers broke even a decade ago. Now it’s losing $47.19 million. This isn’t a problem unique to Rutgers. It’s a systematic problem.

Here’s the big picture: Big-time college sports programs are now losing money. The latest reports show that, conference-wide, schools in three of the top four conferences lost money in 2025 — only the Big 12 showed a profit. The two years before, two of the top four conferences did. If we skip over the COVID seasons of 2020 and 2021, then the ACC has lost money for five straight non-COVID seasons (so 2019 and then 2022-25).

How sustainable are the current economics of college sports? That depends on how long people are willing to pay for them — but the nature of who is paying is changing. Some fans are now being asked to dig deeper to pay for things that, in other pro leagues, the owners do. And the students aren’t being asked at all; they’re being told.

The post College sports are awash in more money than ever, yet a growing number of schools are losing money on intercollegiate athletics appeared first on Cardinal News.

College sports are awash in more money than ever, yet a growing number of schools are losing money on intercollegiate athletics [Cardinal News] (04:15 , Friday, 04 September 2026)

America’s top sports-playing colleges may seem awash in money, and in many ways they are.

However, a closer look shows that many of them are feeling the financial stress of rising expenses in an era where players must be paid — and many of them now depend on mandatory student fees to help underwrite increasingly professionalized sports programs.

The data I’m about to present comes from the Knight-Newhouse College Athletics Database, run by the Knight Commission on Intercollegiate Athletics and Syracuse University’s Newhouse School of Public Communications. The database tracks the finances of college athletics, although most private schools don’t make that information available — so, ironically, Syracuse’s own data doesn’t show up. I looked at the 68 schools in the four biggest conferences, the so-called Power 4 of the Atlantic Coast Conference, the Big Ten, the Big 12 and the Southeastern Conference. Of those 68 schools, the database had information on 53 of them. Among the notable private colleges we don’t know about: Notre Dame, which is well-situated enough to play an independent football schedule and command its own television contract, while playing in the ACC for other sports.

A review of that database turns up the following:

* Only 14 schools made money on their sports program in 2024-25 (the most recent year available) without relying on any support from the school or mandatory student fees.

* Twenty schools made money but only because they received support from the school and/or mandatory student fees to make up shortfalls. Among those 20 was Virginia Tech.

* The other 19 all lost money — including the University of Virginia.

That list doesn’t convey the trendline. While 19 of the Power 4 schools are losing money now, 10 years ago only seven were. The losses for those schools are now widening while other schools are falling into unprofitability as expenses rise — and revenue doesn’t keep up.

The U.S. Senate is currently considering legislation, the Protect College Sports Act, aimed at trying to preserve college sports in some semblance of their college form. Its prospects remain uncertain; I discussed those in a previous column. To what degree that legislation would address the financial predicament that some schools now find themselves in is also uncertain. What is certain is that big-time college sports programs are increasingly functioning as professional leagues, with many of the same expenses that more traditional pro leagues have (player salaries, coaches’ salaries) but without the same revenue base (billionaire owners).

This data underscores how untenable that arrangement is. “Very few schools are earning a profit,” says Andrew Zimbalist, a noted sports economist at Smith College. “They’re going to have to cut back.” The problem is that college football programs have to do something that their NFL counterparts don’t have to: supply enough revenue to underwrite nonrevenue sports, particularly women’s sports. As challenging as the economics are for schools in the Power 4 conference, they are even more difficult for schools that aren’t in those conferences, the so-called “mid-majors” such as James Madison University, Liberty University and Old Dominion University in Virginia. “I think mid-majors would be hit very hard,” Zimbalist says. “It’s going to be hard for them to maintain Title IX and Olympic sports.”

I’ll be digging into this data from multiple angles in future columns. For now, here’s one way to look at all this.

The wealth gap between schools is growing

It’s well-established that while the top conferences were once all reasonably balanced in terms of revenues, now they’re not. The Big Ten and the SEC are clearly the two most affluent conferences. That’s put other conferences in jeopardy; as some schools see the Big Ten and SEC pulling away, they want to join. That led to the collapse of the Pac-12 conference and leaves many wondering if the ACC can survive in its current state.

This data shines light on how even those well-to-do conferences have their own economic disparities.

Of the 14 schools that we know are making money without relying on the school or imposing fees on students (again, we don’t know about private schools, and Notre Dame is a big exception), seven are in the Big Ten and five are in the SEC, so 12 of the 14 money-makers are in just two conferences. The other two are in the Big 12, which means the ACC has no schools that made money on their own without school support. However, those moneymakers still represent a minority of even the Big Ten and the SEC.

Not even these 14 schools turn a profit solely from revenue from TV contracts, ticket sales and such. They all depend on donors. In effect, college sports boosters collectively function as the billionaires who own professional sports franchises.

The 14 schools: Arkansas, Florida, Louisiana State, Tennessee and Texas A&M in the SEC; Michigan, Nebraska, Ohio State, Oregon, Penn State, Purdue and Wisconsin in the Big Ten; Kansas State and Oklahoma State in the Big 12. Most of these show no institutional support or student fees in their revenue. A few do, but only small amounts and would have made money even if that was subtracted. In theory, these 14 schools could operate more or less independently — as long as they had access to the college name, the college stadium and the college donor list.

Students are being billed to keep some programs profitable

The Virginia Tech Hokies football team on the field at Lane Stadium on a dark night, with fireworks in the background
The Virginia Tech Hokies at Lane Stadium in Blacksburg. Courtesy of Virginia Tech.

Students aren’t being forced to pay these mandatory fees. Students are free to attend other schools that don’t charge those fees. However, at some of those other schools, these charges for intercollegiate sports may simply be worked into the overall bill, which makes it tricky to compare schools across state lines. Schools in states with more transparent accounting (such as ours) come off looking bad when other states may be doing the same thing in a less open fashion. That’s why when I looked at the revenue, I looked at both the student fee line but also the vaguer “institutional/government support” because in some places those could be student fees by another name.

Either way, we have 20 Power 4 schools that only made money because the school helped cover the costs of running an athletic program in some way. Of those 20, seven were in the Big 12, five were in the ACC, five were in the Big Ten and three were in the SEC.

Let’s look at Virginia Tech because, well, it’s ours. The database shows revenues of $161.22 million and expenses of $156.15 million, for what we’d call in the private sector a profit of $5.07 million. The database also shows that Tech had its students pay $15.66 million while the institution supplied $8.43 million. Take away either one of those and Virginia Tech sports would have lost money. As intercollegiate sports become more expensive, schools will need to find additional sources of revenue — that’s what Tech is now in the process of doing with its Hokie Ventures, a nonprofit that will focus on expanding revenue for Tech sports. Virginia law, though, explicitly allows schools to charge students for intercollegiate athletics.

At Tech, the reliance on student fees has been constant: They constituted 10% of revenues in 2015, and 10% in 2025. Some other schools, though, have had to increase their dependence on student fees to keep up. A decade ago, the University of Arizona imposed no mandatory student fees for athletics. Now it does. Without those student fees, Arizona athletics — which made money 10 years ago without such fees — would lose money today. There are some big-name schools — Alabama, Auburn and Florida State — that are only making money because they are able to make students pay for their big-time sports ambitions.

In all, there are now 19 Power 4 schools that a decade ago were making money and now would be losing if it were not for mandatory student fees and/or “institutional support,” which in some cases may include student fees.

Some schools are increasing “institutional support” for college athletics by multiples

It’s hard to tell how some schools define “institutional support” so there may be a definition for each school or at least each state. In any case, the database shows marked increases in that category of revenue. Virginia Tech has gone from $100,000 in “institutional/government support” to $8.43 million over the 10-year span from 2015 to 2025. That puts Tech at about where Arizona was a decade ago. Over the past 10 years, Arizona has tripled “institutional/government support” for college athletics from $8.97 million to $31.38 million.

Some schools that once provided no institutional support now do: The University of Virginia has gone from zero in institutional support to $19.88 million over that same decade. Florida State has gone from zero to $33.87 million. South Carolina has moved from zero to $43.7 million. Wherever that money is coming from, that’s a lot of buckaroos that once weren’t going to athletics.

This is the athletic version of an arms race, as colleges attempt to make up for the widening gap between the haves and the have-a-lots (we’ll get to the have-nots in a future column). While these schools are jacking up their spending on athletics, remember that there are 14 schools that would have made money without any institutional support or student fees. At what point, if any, do some schools admit they just can’t (or won’t) keep up, and we divide the Power 4 conferences into a top tier of schools that can make money on their own and relegate all the others to another tier? That’s the economic reality now in many ways, but that would be a hard sell to a lot of diehard college football fans. Meanwhile, though, there’s red ink in a field awash with money:

Profit margins are declining and some schools are slipping into unprofitability

Virginia Tech is actually in an enviable position: Its athletic program produces more profit today than it did a decade ago. In 2015, the difference between revenues and expenses was $2.55 million. Now it’s $5.07 million.

Other programs, though, have seen their margins decline — and, in 12 cases, disappear altogether.

Ten years ago, Florida State posted a “profit” — I’ll use that word since it’s easily understandable, although that’s not necessarily how schools would look at this money — of $9.43 million. Now that’s down to $3.76 million, even though revenue is up 75%. Expenses have gone up almost 87%. This is why Florida State wants out of the ACC and into a conference where it can make more money.

Same for North Carolina. A decade ago, UNC-Chapel Hill’s athletic programs saw about $500,000 in profit. Now they’re losing $15 million a year, even though revenues have nearly doubled.

In the ACC, two schools — North Carolina and Louisville — have slipped from making money to losing money. However, in the SEC, which we think of as a rich conference, seven schools that once made money are now losing money — including ultra-rich Texas. The numbers are different but the story is always the same: Expenses are rising faster than revenues.

Rutgers broke even a decade ago. Now it’s losing $47.19 million. This isn’t a problem unique to Rutgers. It’s a systematic problem.

Here’s the big picture: Big-time college sports programs are now losing money. The latest reports show that, conference-wide, schools in three of the top four conferences lost money in 2025 — only the Big 12 showed a profit. The two years before, two of the top four conferences did. If we skip over the COVID seasons of 2020 and 2021, then the ACC has lost money for five straight non-COVID seasons (so 2019 and then 2022-25).

How sustainable are the current economics of college sports? That depends on how long people are willing to pay for them — but the nature of who is paying is changing. Some fans are now being asked to dig deeper to pay for things that, in other pro leagues, the owners do. And the students aren’t being asked at all; they’re being told.

The post College sports are awash in more money than ever, yet a growing number of schools are losing money on intercollegiate athletics appeared first on Cardinal News.

College sports are awash in more money than ever, yet a growing number of schools are losing money on intercollegiate athletics [Cardinal News] (04:15 , Friday, 04 September 2026)

America’s top sports-playing colleges may seem awash in money, and in many ways they are.

However, a closer look shows that many of them are feeling the financial stress of rising expenses in an era where players must be paid — and many of them now depend on mandatory student fees to help underwrite increasingly professionalized sports programs.

The data I’m about to present comes from the Knight-Newhouse College Athletics Database, run by the Knight Commission on Intercollegiate Athletics and Syracuse University’s Newhouse School of Public Communications. The database tracks the finances of college athletics, although most private schools don’t make that information available — so, ironically, Syracuse’s own data doesn’t show up. I looked at the 68 schools in the four biggest conferences, the so-called Power 4 of the Atlantic Coast Conference, the Big Ten, the Big 12 and the Southeastern Conference. Of those 68 schools, the database had information on 53 of them. Among the notable private colleges we don’t know about: Notre Dame, which is well-situated enough to play an independent football schedule and command its own television contract, while playing in the ACC for other sports.

A review of that database turns up the following:

* Only 14 schools made money on their sports program in 2024-25 (the most recent year available) without relying on any support from the school or mandatory student fees.

* Twenty schools made money but only because they received support from the school and/or mandatory student fees to make up shortfalls. Among those 20 was Virginia Tech.

* The other 19 all lost money — including the University of Virginia.

That list doesn’t convey the trendline. While 19 of the Power 4 schools are losing money now, 10 years ago only seven were. The losses for those schools are now widening while other schools are falling into unprofitability as expenses rise — and revenue doesn’t keep up.

The U.S. Senate is currently considering legislation, the Protect College Sports Act, aimed at trying to preserve college sports in some semblance of their college form. Its prospects remain uncertain; I discussed those in a previous column. To what degree that legislation would address the financial predicament that some schools now find themselves in is also uncertain. What is certain is that big-time college sports programs are increasingly functioning as professional leagues, with many of the same expenses that more traditional pro leagues have (player salaries, coaches’ salaries) but without the same revenue base (billionaire owners).

This data underscores how untenable that arrangement is. “Very few schools are earning a profit,” says Andrew Zimbalist, a noted sports economist at Smith College. “They’re going to have to cut back.” The problem is that college football programs have to do something that their NFL counterparts don’t have to: supply enough revenue to underwrite nonrevenue sports, particularly women’s sports. As challenging as the economics are for schools in the Power 4 conference, they are even more difficult for schools that aren’t in those conferences, the so-called “mid-majors” such as James Madison University, Liberty University and Old Dominion University in Virginia. “I think mid-majors would be hit very hard,” Zimbalist says. “It’s going to be hard for them to maintain Title IX and Olympic sports.”

I’ll be digging into this data from multiple angles in future columns. For now, here’s one way to look at all this.

The wealth gap between schools is growing

It’s well-established that while the top conferences were once all reasonably balanced in terms of revenues, now they’re not. The Big Ten and the SEC are clearly the two most affluent conferences. That’s put other conferences in jeopardy; as some schools see the Big Ten and SEC pulling away, they want to join. That led to the collapse of the Pac-12 conference and leaves many wondering if the ACC can survive in its current state.

This data shines light on how even those well-to-do conferences have their own economic disparities.

Of the 14 schools that we know are making money without relying on the school or imposing fees on students (again, we don’t know about private schools, and Notre Dame is a big exception), seven are in the Big Ten and five are in the SEC, so 12 of the 14 money-makers are in just two conferences. The other two are in the Big 12, which means the ACC has no schools that made money on their own without school support. However, those moneymakers still represent a minority of even the Big Ten and the SEC.

Not even these 14 schools turn a profit solely from revenue from TV contracts, ticket sales and such. They all depend on donors. In effect, college sports boosters collectively function as the billionaires who own professional sports franchises.

The 14 schools: Arkansas, Florida, Louisiana State, Tennessee and Texas A&M in the SEC; Michigan, Nebraska, Ohio State, Oregon, Penn State, Purdue and Wisconsin in the Big Ten; Kansas State and Oklahoma State in the Big 12. Most of these show no institutional support or student fees in their revenue. A few do, but only small amounts and would have made money even if that was subtracted. In theory, these 14 schools could operate more or less independently — as long as they had access to the college name, the college stadium and the college donor list.

Students are being billed to keep some programs profitable

The Virginia Tech Hokies football team on the field at Lane Stadium on a dark night, with fireworks in the background
The Virginia Tech Hokies at Lane Stadium in Blacksburg. Courtesy of Virginia Tech.

Students aren’t being forced to pay these mandatory fees. Students are free to attend other schools that don’t charge those fees. However, at some of those other schools, these charges for intercollegiate sports may simply be worked into the overall bill, which makes it tricky to compare schools across state lines. Schools in states with more transparent accounting (such as ours) come off looking bad when other states may be doing the same thing in a less open fashion. That’s why when I looked at the revenue, I looked at both the student fee line but also the vaguer “institutional/government support” because in some places those could be student fees by another name.

Either way, we have 20 Power 4 schools that only made money because the school helped cover the costs of running an athletic program in some way. Of those 20, seven were in the Big 12, five were in the ACC, five were in the Big Ten and three were in the SEC.

Let’s look at Virginia Tech because, well, it’s ours. The database shows revenues of $161.22 million and expenses of $156.15 million, for what we’d call in the private sector a profit of $5.07 million. The database also shows that Tech had its students pay $15.66 million while the institution supplied $8.43 million. Take away either one of those and Virginia Tech sports would have lost money. As intercollegiate sports become more expensive, schools will need to find additional sources of revenue — that’s what Tech is now in the process of doing with its Hokie Ventures, a nonprofit that will focus on expanding revenue for Tech sports. Virginia law, though, explicitly allows schools to charge students for intercollegiate athletics.

At Tech, the reliance on student fees has been constant: They constituted 10% of revenues in 2015, and 10% in 2025. Some other schools, though, have had to increase their dependence on student fees to keep up. A decade ago, the University of Arizona imposed no mandatory student fees for athletics. Now it does. Without those student fees, Arizona athletics — which made money 10 years ago without such fees — would lose money today. There are some big-name schools — Alabama, Auburn and Florida State — that are only making money because they are able to make students pay for their big-time sports ambitions.

In all, there are now 19 Power 4 schools that a decade ago were making money and now would be losing if it were not for mandatory student fees and/or “institutional support,” which in some cases may include student fees.

Some schools are increasing “institutional support” for college athletics by multiples

It’s hard to tell how some schools define “institutional support” so there may be a definition for each school or at least each state. In any case, the database shows marked increases in that category of revenue. Virginia Tech has gone from $100,000 in “institutional/government support” to $8.43 million over the 10-year span from 2015 to 2025. That puts Tech at about where Arizona was a decade ago. Over the past 10 years, Arizona has tripled “institutional/government support” for college athletics from $8.97 million to $31.38 million.

Some schools that once provided no institutional support now do: The University of Virginia has gone from zero in institutional support to $19.88 million over that same decade. Florida State has gone from zero to $33.87 million. South Carolina has moved from zero to $43.7 million. Wherever that money is coming from, that’s a lot of buckaroos that once weren’t going to athletics.

This is the athletic version of an arms race, as colleges attempt to make up for the widening gap between the haves and the have-a-lots (we’ll get to the have-nots in a future column). While these schools are jacking up their spending on athletics, remember that there are 14 schools that would have made money without any institutional support or student fees. At what point, if any, do some schools admit they just can’t (or won’t) keep up, and we divide the Power 4 conferences into a top tier of schools that can make money on their own and relegate all the others to another tier? That’s the economic reality now in many ways, but that would be a hard sell to a lot of diehard college football fans. Meanwhile, though, there’s red ink in a field awash with money:

Profit margins are declining and some schools are slipping into unprofitability

Virginia Tech is actually in an enviable position: Its athletic program produces more profit today than it did a decade ago. In 2015, the difference between revenues and expenses was $2.55 million. Now it’s $5.07 million.

Other programs, though, have seen their margins decline — and, in 12 cases, disappear altogether.

Ten years ago, Florida State posted a “profit” — I’ll use that word since it’s easily understandable, although that’s not necessarily how schools would look at this money — of $9.43 million. Now that’s down to $3.76 million, even though revenue is up 75%. Expenses have gone up almost 87%. This is why Florida State wants out of the ACC and into a conference where it can make more money.

Same for North Carolina. A decade ago, UNC-Chapel Hill’s athletic programs saw about $500,000 in profit. Now they’re losing $15 million a year, even though revenues have nearly doubled.

In the ACC, two schools — North Carolina and Louisville — have slipped from making money to losing money. However, in the SEC, which we think of as a rich conference, seven schools that once made money are now losing money — including ultra-rich Texas. The numbers are different but the story is always the same: Expenses are rising faster than revenues.

Rutgers broke even a decade ago. Now it’s losing $47.19 million. This isn’t a problem unique to Rutgers. It’s a systematic problem.

Here’s the big picture: Big-time college sports programs are now losing money. The latest reports show that, conference-wide, schools in three of the top four conferences lost money in 2025 — only the Big 12 showed a profit. The two years before, two of the top four conferences did. If we skip over the COVID seasons of 2020 and 2021, then the ACC has lost money for five straight non-COVID seasons (so 2019 and then 2022-25).

How sustainable are the current economics of college sports? That depends on how long people are willing to pay for them — but the nature of who is paying is changing. Some fans are now being asked to dig deeper to pay for things that, in other pro leagues, the owners do. And the students aren’t being asked at all; they’re being told.

The post College sports are awash in more money than ever, yet a growing number of schools are losing money on intercollegiate athletics appeared first on Cardinal News.

With classes started, Roanoke school division provides final number of 148 eliminated jobs [Cardinal News] (04:10 , Friday, 04 September 2026)

After announcing in the spring that it would need to cut about 170 positions to balance its budget, Roanoke City Public Schools cut 148 jobs ahead of this school year. 

Of those positions, 64 were teachers, spread across all 24 of the city’s schools. About 21 positions in the central office were among the cuts, which varied from a half-time accountant position and a director of career and technical education to a special education Medicaid coordinator, an English learner instructor and more than five custodians. 

The cuts came as Roanoke city schools worked to balance a $14 million deficit in the fiscal year 2027 budget. Seven additional positions were moved out of the general budget to be paid for by grants dependent on available funding — including three student support specialists.

Of the eliminated positions, 66 were filled at the time they were eliminated, spokesperson Claire Mitzel said. The remaining positions were vacant.

Some of the people who lost their jobs have also been recalled to fill other vacancies that have arisen separately from the budget cuts.

For example, nine teaching positions were cut at William Fleming High School. But three teachers, including two physical education teachers and a Spanish teacher, were called back as positions that remained in the budget opened this summer, said Principal Tracey Anderson.

As of Aug. 20, 13 employees whose jobs were eliminated have been recalled, Mitzel said. 

The eliminated positions come as the division grapples with the city council’s decision to change its school funding formula by reducing the school’s share of local tax revenue growth.

In addition to cutting positions, the school division also reduced spring activity buses, the preschool program and the elementary gifted program, among other operational cuts. 

Roanoke City Public Schools is among many school divisions across the state and nationwide experiencing budget cuts and even layoffs due to declining enrollment and financial woes.

Cuts like these haven’t happened since the 2008 recession, Mitzel said.

“This is all new for all of us, and we are continuing to try to make the very best decisions looking across the entire organization,” Mitzel said.

Every school in the division experienced cuts, but some saw more than others. Fleming lost nine teachers out of 14 total cuts, and Round Hill Elementary lost six teachers as part of 10 total cuts, according to data obtained by Cardinal News through the Freedom of Information Act.

Crystal Spring, Morningside and Wasena Elementary schools didn’t lose any teachers and saw seven cuts total across the three schools, including a librarian at Crystal Spring, an assistant principal at Morningside and five teaching assistants. And James P. Fishwick Middle School had no cuts, according to data provided by the division.

The reasoning behind those decisions and their connection to enrollment remains unclear in many cases. Division officials did not specify which positions were eliminated due to enrollment challenges — something that Mitzel said is constantly evolving throughout the year.

The Virginia Department of Education lays out and enforces staffing requirements primarily through the Standards of Quality, or SOQs. These guidelines govern class sizes and require a certain number of teachers based on student enrollment, though schools often might employ more than the minimum number of teachers or other positions based on locally determined needs.

If the number of students enrolled at a school or in a division declines, the number of teachers required might, too, but those types of staffing changes are separate from the budget cuts made earlier this year, Mitzel said. 

“That’s really independent of the budget conversation and could even change mid-year,” Mitzel said. “It is very normal every single year as we are planning out our classes based on enrollment.”

Due to the cuts, some schools are now sharing positions. Those decisions were not solely a consequence of budget cuts but also based on enrollment, Mitzel said.

She said some of the division’s smaller elementary schools are sharing an assistant principal and some elective teachers may be working across schools, such as a theater teacher being shared between the division’s two high schools. 

“When it came time to make these hard decisions, the SOQs guided a lot of this work,” she said. “We are still within the SOQs.”

In some cases, instructors are moved around to even out the number of students teachers are working with.

Anderson echoed Mitzel and said the cuts at Fleming High were made in part due to enrollment.

Anderson said the division looked at upcoming student enrollment and class sizes to determine what positions would be eliminated.

She added that Fleming had been previously allocated additional teachers for some subjects in an effort to reduce class sizes, beyond state requirements.

“In our math classes, science classes, because of labs and things of that nature, we want to, if we can, reduce that class size, because we know if the class is smaller, then it gives students more opportunities to respond and more opportunities to work with the teacher,” she said. “So, when the school board gave those additional positions, they gave me one for each content area, which was great because then that allowed us some more flexibility with making those classes smaller.”

Superintendent Verletta White said the division did its best to prevent cuts from negatively affecting students and “to protect” classrooms. In addition to the 64 eliminated teaching positions, another 37 were teaching or instructional assistants who typically work with teachers in the classroom or in small groups with students. 

“We know that with our budget cuts, we unfortunately have fewer staff members this year than what we had last year. But we’ve done our very best to protect our classrooms,” White said during a media briefing on the first day of school. “We know that some of the central support that our schools have had, that we don’t have as many of those. So I’m sure that our schools will feel that.”

Schools also grappling with vacancies

As of the first day of school, the division had 127 vacancies: 51 administrative or professional positions, which include school leaders and certified educators, and 76 classified positions, which can include custodians, school security officers and instructional assistants. 

Of those classified positions, 64 vacancies are for instructional assistants, which is on trend with what the division usually sees in the fall; an Oct. 1, 2025, vacancy report obtained from the division lists 87 vacancies of classified positions.

Anderson said these types of staffing challenges are par for the course. She’s worked for Roanoke city schools for more than 30 years and has been principal at Fleming since 2021, and last year was the first time the high school was fully staffed.

“I was holding my breath, thinking ‘Wow, this is amazing.’ You know, it can happen, but again, people get married, they have children, and so you have natural openings that just occur, and you have to fill those areas. And we want to make sure in filling them that we find the best people that we can,” Anderson said.

In addition to the positions she lost as part of the budget cuts, Fleming currently has multiple vacancies including two math teacher positions, two special education positions, a school nurse, a counselor and a part-time security position.

But it can be hard to find people qualified for high school-level math and science roles.

During the 2025-26 school year, the overall teacher vacancy rate for the state was 2.5%, according to the data from the Virginia Department of Education, but for math it was 3.8%. Special education was even higher with a 4.9% vacancy rate statewide.

“We are diligent and looking and trying to find people that know how to work with our students to build those relationships while having the certification that can teach our students what they need to learn,” Anderson said.

Overall, the division’s teacher vacancy rate was less than 2% as of the first day of school, White announced at a school board meeting on Aug. 11.

The post With classes started, Roanoke school division provides final number of 148 eliminated jobs appeared first on Cardinal News.

Roanoke tells city residents about data breach three months after it happened [Cardinal News] (04:08 , Friday, 04 September 2026)

People potentially hit by a data breach were notified by the city of Roanoke by letter in late August — over three months after the fact — of the possibility that their personal information had been accessed in early May.

Mel Williams, a Roanoke attorney, received the letter from the city dated Aug. 24 that said that “on or about” May 11, the city found that on May 6 a “malicious actor gained unauthorized access to departmental data.” The letter does not identify the department or any other specifics about the hack. Nor did it disclose how many people could have been affected.

The city did not respond to numerous submitted questions about the incident Wednesday or Thursday.

Williams said the city waiting months to tell affected residents is “unconscionable” — “Why bother after waiting so long,” he said in an email. Williams said Thursday that he is not aware of any adverse effects so far.

“The malicious actor’s access was terminated soon after it was detected, but was able to gain access to a limited set of departmental data from the City’s network before being detected,” said the letter, which Williams provided to Cardinal News.

The city’s letter goes on to say that the following may have been accessed: first and last names, Social Security numbers, passport numbers and financial account information. 

“We are providing this notice out of an abundance of caution,” the letter said. It added that to date, the city has not received any indication that personal data had been misused.

Williams said he has concerns that the hacker had five days of uninterrupted access before the city noticed the breach. 

“Is there a worse scenario than having this connected information put in the hands of nefarious individuals?” he said by email. 

“Upon discovering the incident, information technology experts were immediately engaged and commenced an investigation to determine the nature and scope of the incident,” the letter said. The city also reported the attack to the FBI.

The letter said the city engaged a “leading security service provider” to monitor the network, review the system’s architecture and implement stronger policies to prevent future attacks.

According to Virginia law, if “unencrypted or unredacted personal information” is “reasonably believed” to have been accessed by an unauthorized individual, the entity responsible for that data must disclose the breach to the state Office of the Attorney General and any affected resident of Virginia “without unreasonable delay.”

The statute says the notice may be “reasonably delayed” to allow the entity to determine the scope of the breach and restore the system first or if the breach is a matter of national security. It also says the attorney general can impose a civil penalty of up to $150,000 for an information breach.

Bruce Wetterau, who also received the city’s letter, said via email that he, too, is upset that the city waited three months to notify him and that it did not offer any identity theft protections.

“The letter just included a page and a half of general self-help tips I should take,” Wetterau said. Wetterau said he is also not aware of any adverse impacts from the breach.

For anyone with questions, the city in its letter recommended reaching out to its community engagement team.

The post Roanoke tells city residents about data breach three months after it happened appeared first on Cardinal News.

Notes from the Square: State seeks input on Virginia’s rail plan [Cardinal News] (04:05 , Friday, 04 September 2026)

a train on a track running alongside an interstate highway

Welcome to Notes from the Square, a roundup of state politics and policy news. Each week, we bring you updates on the movers and shakers in Virginia politics as well as the legislation and initiatives they’re supporting or opposing — with a Southwest and Southside Virginia focus. 

Got a tip or story idea? Email me at elizabeth@cardinalnews.org.

State seeks public input on rail plan

The Virginia Department of Rail and Public Transportation will host a virtual public meeting on Tuesday at 6:30 p.m. to gather input into its development of a 2026 statewide rail plan.

Department staff plans to review feedback from the public collected through surveys and previous meetings. Staff also will share how public comments are shaping the planning process and provide additional insight as the project team works toward a more responsive, community driven rail strategy, according to a release about the upcoming meeting. 

Members of the public who want to attend the meeting are encouraged to register online.

The department will use the statewide rail plan to outline its six-year and 20-year investment efforts and future improvements to passenger and freight rail service. 

Public feedback will play a critical role in shaping the future of the commonwealth’s rail network, the department said in a statement.

Spanberger announces new board appointments

Gov. Abigail Spanberger announced more board appointments at the end of August. Here’s a list of appointees from Southwest and Southside Virginia (*denotes reappointment): 

Board of trustees of the Sen. Frank Ruff Center for Rural Virginia

  • Dr. William Swecker Jr. of Blacksburg, professor emeritus of large animal clinical sciences, Virginia-Maryland College of Veterinary Medicine at Virginia Tech

Southwest Virginia Energy Research and Development Authority

  • *Michael Karmis of Blacksburg, professor emeritus, Virginia Tech

Tobacco Region Revitalization Commission

  • *Jay Jennings of Mecklenburg, executive vice president, JF Leaf Ltd.
  • *Sarah Wilson of Abingdon, owner, Leonard Land Livestock
  • Angela Hairston of Danville, superintendent, Danville Public Schools
  • Alexis Ehrhardt of Danville, vice president for government relations, Virginia Commonwealth University
  • Edward Owens of South Boston, owner and operator, Edward Owens Agency

Virginia Tourism Authority

  • Will Payne Jr. of Bristol, co-founder, Squabble State Hard Cider & Spirits

Virginia Public School Authority Board of Commissioners

  • Kevin Rotty of Goochland, retired, PFM Financial Advisors

Advisory board for the Virginia Department for the Deaf and Hard of Hearing

  • *Carl Cline of Roanoke, president, Carilion Clinic at Carilion Franklin Memorial Hospital

Advisory Board on Acupuncture

  • Victoria Jean Taylor of Christiansburg, licensed acupuncturist

Board for People with Disabilities

  • Olivia Price of Covington, school cafeteria food service, Alleghany County Public Schools

Board of Funeral Directors and Embalmers

  • *Dr. Scott Hickey of Maidens, emergency physician, Augusta Health

Board of Nursing

  • *Helen Parke of Concord, family nurse practitioner, Blue Ridge Medical Center

Board of Physical Therapy

  • Mitchell Davis of Salem, retired

Board of Veterinary Medicine

  • Dr. Steve Karras of Hardy, veterinarian, Cave Spring Veterinary Clinic

Maternal Mortality Review Team

  • Dr. Jaclyn Nunziato James of Roanoke, obstetrician and gynecologist, associate professor, Virginia Tech Carilion School of Medicine

Soil and Water Conservation Board

  • Nate Aker of Wytheville, farmer, Wolfpen Farms

Board of directors of the Virginia Recreational Facilities Authority 

  • Kelvin Bratton of Roanoke, director, real estate valuation, city of Roanoke
  • Thomas Miller of Salem, director of economic development, city of Salem
  • Michael Clark of Roanoke, owner and principal consultant, Oxbow
  • Richard Peters of Franklin County, town manager, town of Vinton

Waste Management Board

  • Craig Coker of Troutville, principal, Coker Composting and Consulting
  • Jeremy Garrett of Roanoke, director of operations and technical services, Roanoke Valley Resource Authority

Aerospace Advisory Council

  • *Tombo Jones of Roanoke, director, Virginia Tech FAA Designated UAS Test Site

The post Notes from the Square: State seeks input on Virginia’s rail plan appeared first on Cardinal News.

Southwest Field Notes: National Black Lung Association to meet next week in Abingdon amid historic high numbers of the disease in central Appalachia [Cardinal News] (04:05 , Friday, 04 September 2026)

Update 12:30 p.m. Sept. 4: Rep. Morgan Griffith, R-Salem, will not attend Friday’s Black Lung Association meeting, a spokesperson said Friday. He instead provided remarks ahead of the meeting.

___________

Hello, and welcome back to Southwest Field Notes — a weekly column where I cover a few things going on in our region, and try to keep up with all that is happening! 

Please keep reading to learn more about important meetings and initiatives right here in the coalfields. You can reach out to me any time at anna@cardinalnews.org

National Black Lung Association to convene, elect new leadership 

On Friday, the national Black Lung Association will meet in Abingdon to discuss new research, legislative efforts and an election for leadership positions. 

The meeting comes as black lung, an incurable and disabling disease, is at its highest rates in the central Appalachia coalfields since the 1970s, according to a study from the National Institute for Occupational Safety and Health published in August. 

One in three miners who worked for at least 25 years tested positive for the disease, while younger miners with less time underground are disproportionately suffering from complicated black lung. 

The national association was founded in 1969 to help miners and their families receive benefits and advocate for better policies across central Appalachia.

Vonda Robinson, vice president of the association and lives in Nicklesville in Scott County, is running for president this year and is coordinating the event. Robinson has been a fierce advocate for miners in the years since her husband John’s diagnosis with the disease — he is running for vice president of the organization this year.

“I can talk the political part, and he can talk the coal mining part,” she said. 

U.S. Sen. Mark Warner, D-Va.; U.S. Sen. Tim Kaine, D-Va.; and U.S. Rep. Morgan Griffith, R-Salem, are expected to attend. Dr. Brandon Crum, who operates a radiology clinic in Pikeville, Kentucky, will share reports from his work. 

Members from across the chapters will vote on leadership. If elected, Robinson would be the first woman to fill the seat, as well as the first president from Virginia, she said. 

Community meals as a classroom

Tomorrow, students at the University of Virginia’s College at Wise who are pursuing public health educations will prepare nachos and tacos at the Norton Family Crisis Center, the first community meal of their fall season. 

The monthly meals are a part of the Community Nourishment Project, a collaboration between the college and the Southwest Virginia Graduate Medical Education Consortium. They aim to connect students interested in healthcare with their community, especially important for students who want to be medical doctors. 

Research has established a marked decrease in empathy for third-year medical students, said Troy Makal, co-director of the project. 

“We strive to circumvent that by getting them into the community early, gain hands-on experience and teach them against stereotypes,” he said. “Our students don’t experience that empathy drop … they often come back to Southwest Virginia to serve the community here.” 

In recent years, the meals have expanded to two satellite locations in Abingdon and Blacksburg, Makal added. In Norton, meals are being served at the crisis center this fall, which serves people experiencing domestic violence, homelessness and addiction.

The rest of the fall season includes:

  • Oct. 3: Spaghetti and meatballs with garlic bread and salad
  • Oct. 24: Chili with fall veggies and cornbread
  • Nov. 21: Friendsgiving feast with turkey and traditional sides 

To learn more about The Community Nourishment Project, contact Makal at vem4r@uvawise.edu

Scott County might amend its zoning map to include data centers

The Scott County Planning Commission will hold a public hearing to consider changing the county’s zoning ordinance to include data centers. 

The planning commission will consider adding data centers as a miscellaneous use in commercial and industrial districts, according to the event details. 

The public can see a copy of the proposed ordinance at the county administrator’s office, the release states. 

The consideration comes as data center proposals are cropping up across Southwest Virginia, which is being marketed for its cheaper land and low tax rates. In March 2021, Scott County was one of several coalfield counties that agreed to a 24-cent per $100 of assessed value tax on data center equipment and later passed a resolution to express support for attracting data centers to the region. 

Developers promote the economic benefits data centers can provide, including jobs and tax revenue, while opponents cite their detrimental impact on the environment and quality of life, including large water and energy requirements. 

The hearing will be Sept. 14 at 6 p.m. in the board of supervisors meeting room, 190 Beech St., Suite 201, in Gate City.

The post Southwest Field Notes: National Black Lung Association to meet next week in Abingdon amid historic high numbers of the disease in central Appalachia  appeared first on Cardinal News.

Martinsville Field Notes: Henry County Food Pantry gets $250,000 for expansion [Cardinal News] (04:05 , Friday, 04 September 2026)

Molly Pinkston and Sharon Mills stand with a $250,000 check.

Welcome to Martinsville Field Notes, where you can find quick updates on what’s happening in Martinsville and Henry County every Friday. In this week’s column, you’ll hear about all that’s going on at Patrick & Henry Community College, and how the food pantry will use a $250,000 gift. Don’t forget to check out the last edition if you missed it.

If there’s something you’d like to see in next week’s column, you can reach me at julianna@cardinalnews.org. There’s always something happening in Martinsville and Henry County, and I’d love to hear from you!

Feeding Southwest Virginia gives Henry County Food Pantry $250,000

With a $250,000 gift, the Henry County Food Pantry will expand its capacity through additional food distributions, including delivery to improve food access.

The grant will be distributed across a three-year period and will allow the food pantry to:

  • Hire a full-time staff member.
  • Expand access to fresh produce.
  • Participate in a home delivery service.
  • Expand the volunteer network.
  • Make capital improvements to their facility.

The non-profit Feeding Southwest Virginia gave the grant to the pantry because it is centrally located, serves more than 2,000 people a month and the area has a higher food insecurity rate than the average of Southwest Virginia, making it a “focus region,” said Molly Pinkston, director of agency partnerships for Feeding Southwest Virginia. The pantry ranks eighth in poundage distribution among the 275 organizations with which Feeding Southwest Virginia partners.

The Henry County pantry operates in one of Virginia’s most economically challenged rural communities, according to a Feeding Southwest Virginia news release. About 25% of children, and one in six people, in Henry County face hunger.

Over the past five years, Feeding Southwest Virginia gave $1.6 million to agency partners, Pinkston said.

The food pantry’s expanding and successful programming was also a reason why Feeding Southwest Virginia chose Henry County, Pinkston said. The food pantry — along with food — offers clothing, furniture, hygiene products, household items and pet food. Its partnership with Purina allows local families to feed their pets, who are meaningful members of the family, said Sharon Mills, the director of the Henry County Food Pantry.

The food pantry has also given over 100 beds to children who did not have one. They receive referrals from the Department of Social Services, schools and first responders to ensure children have stability and safe sleep, Mills said.

Patrick & Henry Community College in the running to win Aspen Prize

The Aspen Institute will be at Patrick & Henry Community College on Wednesday and Thursday for a site visit to consider it for a major award.

The 2027 Aspen Prize for Community College Excellence brings with it a prize of $1 million.

Patrick & Henry was announced as the first community college in Virginia to be a semifinalist or finalist for the prize. In the spring, the college was one of 25 semifinalists and then chosen as one of 10 finalists in June. The winner of the prize will be announced in April.

“We are overjoyed and elated to have been named a top 10 community college in America,” said Greg Hodges, the college’s president. “They’re coming here to learn how we’ve achieved the results that we’ve had, and so we’re excited to be able to share with them the work that we’ve done that’s led us to this point.”

In the Aspen Institute’s finalist announcement, it highlighted Patrick & Henry’s 47% completion rate compared to a national average of 37% of students completing a community college credential within four years.

The Aspen Institute is a global nonprofit that focuses on leadership and societal issues. Its College Excellence Program aims to strengthen higher education leadership and improve student outcomes.

Hodges said the community college’s goals have aligned with the Aspen Institute’s focus. About 20 years ago, the focus was college enrollment, and about 10 years ago, the focus was completion. The Aspen Prize now focuses on strong outcomes, such as good-paying jobs or transferring to a four-year university.

“That focus on those good-paying jobs and aligning our career pathways with our regional good-paying jobs and the fact that our community is experiencing a real economic renaissance has really helped get us to this point,” Hodges said.

Henry County announced in the past year three major projects, with an estimated 328 jobs and $123.9 million in investment. Hodges said the community college is deliberate in plugging students into local industry jobs. The community college most recently teamed up with Rural King to meet its truck driving needs.

The site visit next week will include meeting with the president’s leadership team, student service leaders, advisors, career and technical education faculty, general education faculty and advisors.

Also at Patrick & Henry Community College, a certification milestone

Patrick & Henry was awarded 5,061 certifications to students through the National Coalition of Certification Centers, as part of the 1.3 million nationally.

Patrick & Henry has partnered with NC3 since 2018, said the college president, Greg Hodges. The college decided to partner with NC3 after opening the first building of the manufacturing engineering and technology complex in 2017. The certifications are typically in career and technical education and healthcare, including mechatronics and automotives.

The certifications are tied to classes at the community college to help students get a job after completing their education.

Patrick & Henry is a “pioneer” in the Festo Industry Certification Program, offering all three levels of mechatronics certifications. Patrick & Henry was the first college in the nation to offer the third level of the Festo certification, which includes advanced robotics, industrial communications and smart maintenance. Faculty members wrote the curriculum for that.

The number of certifications is rising each year at the college, with 572 so far in 2026. The most popular certifications this year have been various bending certifications and fundamentals of industry, Hodges said. There are a total of 612 students who have completed at least one NC3 certification.

The post Martinsville Field Notes: Henry County Food Pantry gets $250,000 for expansion appeared first on Cardinal News.

Tobacco commission’s new foundation taps Deborah Gosney as executive director [Cardinal News] (04:01 , Friday, 04 September 2026)

A map of Virginia showing the localities included in the Tobacco Commission region.

The newly formed Foundation for the Advancement of Southern and Southwest Virginia has hired Deborah Gosney as its executive director, according to an announcement Thursday by the Virginia Tobacco Region Revitalization Commission. 

Gosney was the executive director of the Southside Planning District Commission, where she worked for 38 years, before joining the foundation, which was formed by the tobacco commission in May. She was executive director of the Southside Planning District Commission for 6 years.

Deborah Gosney. Courtesy of the Virginia Tobacco Region Revitalization Commission.

“My career has been rooted in helping communities identify opportunities, strengthen local capacity, and improve the quality of life for the people we serve. I look forward to bringing that same commitment, experience, and passion to FASS-VA as we continue advancing its important work across the Tobacco Region footprint,” Gosney said in a statement. 

Over the course of her career, Gosney has worked with local governments, planning district commissions, and state and federal agencies to secure funding for infrastructure, business growth and community revitalization. She is from Mecklenburg County. 

“Economic development, especially in the rural regions of the Commonwealth the Commission covers, requires a team effort, and Deb has a long track record of successful leadership and partnership in that area,” Del. William Morefield, R-Tazewell County, said in a statement. Morefield serves as chair of the Tobacco Region Revitalization Commission. 

The commission met in Floyd in mid-May to adopt a budget and approve 38 funding requests. Among those dozens of requests was one to create the nonprofit foundation.

The commission had learned, through outreach in Southwest and Southside, that the biggest barrier to progress faced by localities and organizations in seeking government grant money was a lack of capacity, Jordan Butler, spokesperson for the commission, said in May. The overall goal of the new foundation is to help those localities and organizations in the grant application process as they seek federal, state or private philanthropic funding — something the commission can’t do as a political entity.

“Deb’s wealth of experience working with the same communities and organizations means she is ready to hit the ground running. I encourage anyone with an idea that could help build an even brighter future for our counties, cities, and towns to reach out to Deb to see how FASS-VA might be able to help,” Jeff Haley, chair of the foundation, said in a statement.

Update: This article has been updated to reflect the accurate length of time Gosney served as executive director of the Southside Planning District Commission.

The post Tobacco commission’s new foundation taps Deborah Gosney as executive director appeared first on Cardinal News.

Miyares: The verdict on the Virginia Renaissance isn’t in the op-ed page. It’s in the data. [Cardinal News] (04:00 , Friday, 04 September 2026)

Former Attorney General Jason Miyares, right, at a law enforcement event in Lee County. Courtesy of Attorney General's Office.

I am always amused when leftwing activists simply refuse to absorb clear facts. Case in point is Shawn Weneta’s recent op-ed saying the verdict on the economic and public safety turnaround under Gov. Youngkin, dubbed the “Virginia Renaissance,” is already in. While he is correct that the verdict is in, he is clearly reading the wrong one.

The bottom line is that by every objective measure, Virginia was stronger, safer and more prosperous after four years of the Youngkin-Miyares Administration than the commonwealth we inherited. The stunning facts of the Virginia Renaissance simply cannot be ignored: more economic investment in Virginia under Gov. Youngkin than the last six governors combined. A state government that was well-managed, leaving office with a $572 million budget surplus, a $1.8 billion cash cushion and $4.7 billion in the rainy day fund, with $9 billion in tax relief for Virginians and consistent rankings as one of the best states in the nation to do business. 

That isn’t a slogan. That is the record. 

Sen. Daniel Patrick Moynihan noted that everyone is entitled to their own opinion, but never to their own facts. In this case, what Democrats don’t get to rewrite are unemployment numbers, economic investment, murder rates or overdose statistics after the fact. 

Worth remembering is the Virginia that Youngkin and I inherited when we took office. The Virginia of 2021 was dying, literally, financially and figuratively. Coming into office, we had the highest levels of addiction deaths ever recorded in the history of the commonwealth, a murder rate at a 20-year high and a violent crime rate at a 30-year high. It was a Virginia that was ranked 46th out of 50 states in post-COVID job creation, where our public schools were among the last to reopen, which led to a devastating impact on learning loss and the mental health of Virginia’s children. All of these factors led to the first time in close to 100 years of more people moving out of the state than moving in. 

People were fleeing because of the results of the progressive monopoly that had descended in Richmond under Youngkin’s predecessor, Ralph Northam, which created conditions where businesses and population were attempting to flee. That was the Virginia Democrats handed us and is the baseline Weneta conveniently leaves out of his column. He seems to be blaming the doctors for the treatment rather than the failed leftist policies, more concerned with social justice activism than governance. 

After four years of Gov. Youngkin’s common-sense policies, Virginians rewarded him with a 56% approval rating. Even Abigail Spanberger ran on continuing Youngkin’s successes by campaigning on job creation and being pro-law enforcement because she knew that’s what Virginians actually wanted. Now her approval ratings are among the lowest of any modern Virginia governor, and it isn’t a mystery why: it’s the bait and switch of campaigning as a moderate and governing as a leftist.

As Virginia’s “People’s Protector,” I chose to work with law enforcement instead of against them. We launched Operation Ceasefire with bipartisan support, and the results speak for themselves: a statewide drop in the murder rate of more than 30%, and in some of our Ceasefire cities, including Roanoke, we saw a 66% drop in murders. Those are Virginia State Police numbers and actual results, not talking points.

We held opioid manufacturers responsible for the addiction crisis to the tune of $1.2 billion in legal judgments while aggressively prosecuting fentanyl dealers that were poisoning our children. In just four years, my office removed enough fentanyl off the streets of Virginia that would have taken the lives of 6 million of our fellow citizens. This led to Virginia being recognized by the DEA as having the number one drop in addiction deaths of any state in the nation in my last year in office. 

Compare that to the disastrous “Enhanced Earned Sentence Credit” program championed by Shawn Weneta, a convicted felon himself. In the first full year alone, nearly half, 49.8%, of those released early under the program were rearrested on new criminal charges. This is not a policy hiccup; that is a public safety failure with victims’ names attached to it.

Most shocking is that Weneta’s own op-ed doesn’t deny people are getting raped and killed, just that it’s not that many people getting raped and killed due to early felon release. Now he is pressuring Virginia’s General Assembly to take it much further while ignoring the lives ended or irrevocably changed due to early felon release. A criminal-first, victim-last justice policy doesn’t protect Virginians. 

More Virginians were working, more Virginians were thriving, more small businesses were opening their doors and yes, more Virginians were alive when this chapter of the Virginia Renaissance ended in January of 2026. That is the verdict, and with each passing day, it becomes more apparent that the Virginia Renaissance was real. What we have now is a Leftist Resurgence taking Virginia in a very different direction.

I’m flattered by the comparisons to Ronald Reagan because I’m a firm believer that the American Miracle has provided more hope, opportunity and prosperity than any country on the planet. Virginians are fundamentally a good, decent and noble people, filled with generosity and big hearts. Virginians want a government as good and as decent as they are, one that works for them, not against them. We are all scratching our heads wondering why so many in Richmond seem focused on undoing the Virginia Renaissance so as to adopt the same ruinous policies of California.

The Youngkin legacy proves commonsense leadership wins over bad ideas cloaked with good intentions. Here is hoping we’ve learned the lesson. 

Jason S. Miyares served as Attorney General of Virginia from 2022 to 2026.

The post Miyares: The verdict on the Virginia Renaissance isn’t in the op-ed page. It’s in the data. appeared first on Cardinal News.

Headlines from across the state: Consumer sentiment tumbles as Virginians feel pressure from inflation, new poll finds; more … [Cardinal News] (03:45 , Friday, 04 September 2026)

Here are some of the top headlines from other news outlets around Virginia. Some content may be behind a metered paywall:

Economy:

Consumer sentiment tumbles as Virginians feel pressure from inflation, new poll finds. — Virginia Mercury.

Appalachian Power president tells business summit that power demand is about to surge to unseen levels. — West Virginia MetroNews.

Virginia House speaker details “concerns” about NextEra-Dominion merger in letter to state regulators. — Virginia Mercury.

Politics:

Spanberger’s new commission focused on housing production meets for first time. — Richmond Times-Dispatch (paywall).

Local:

Amazon withdraws controversial groundwater permit in King George County. — Richmond Times-Dispatch (paywall).

Weather:

For more weather news, follow weather journalist Kevin Myatt on Twitter / X at @kevinmyattwx and sign up for his free weather email newsletter. His weekly column appears in Cardinal News each Wednesday afternoon.

The post Headlines from across the state: Consumer sentiment tumbles as Virginians feel pressure from inflation, new poll finds; more … appeared first on Cardinal News.

Thursday, 03 September 2026

Court Tells HHS To Stop Using AI To Cite Fake Studies, Or Willfully Misinterpret Others In Grant Solicitations [Techdirt] (11:04 , Thursday, 03 September 2026)

To quote the opening line from the court opinion that this post is based upon, “Millions of American teenagers have sex.” This post is has nothing to do with whether that sort of thing is good or bad, natural or otherwise, nor the moral implications of it all. Whatever you think about that opening line, I think we can at least agree that teenagers engaging in sexual activity that results in unwanted pregnancies is something we collectively would want to avoid. The government, and HHS specifically, has offered up grants for programs that seek to reduce unwanted pregnancies in teenagers and, as the court notes, those programs appear to have worked.

Alarmed by the country’s rising teenage birth rate, Congress funded grants through the Teen Pregnancy Prevention (“TPP”) Program to support local initiatives proven to reduce teen pregnancy, as well as promising approaches that might also prove effective after further observation and study. Congress intended for these programs to employ a range of strategies, from encouragement of abstinence and delayed sexual activity to education about contraceptives. The effort seems to be working: The teen pregnancy rate has plummeted since Congress began funding the grants in 2010.

Good news all around, it would seem. But since the Trump administration never seems to miss an opportunity to snatch defeat from the jaws of victory, HHS has decided to reinterpret, or you could say simply make up, what Congress intended to do with the money appropriated for these grants. Specifically, HHS has decided that this money will only be used for grants to fund programs that focus solely on abstinence and whatever the fuck “body literacy” is. The general idea is that any education or proliferation of tried and true forms of contraception are now verboten.

Several state counties, a nonprofit focusing on sex education, and Planned Parenthood of the Heartland filed suit to get the grant solicitation paperwork restored to its original state. As the lawsuit notes, neither HHS nor the Executive Branch are allowed to simply rewrite the mandates given alongside money appropriated by Congress. That is law-making by the Executive Branch by any plain reading.

And, to make matters worse, it appears that HHS wrote the grant criteria citing some studies that flatly don’t exist and citing others that don’t say what HHS says they say. The court agreed and issued a preliminary injunction putting a hold on the grant changes for now.

In his opinion granting a preliminary injunction as the lawsuit continues, U.S. District Judge Christopher Cooper called the changes “likely arbitrary and capricious” and said the government had failed to provide sufficient evidence for its policy. He wrote that grant solicitations for the program“(remarkably) reference public health studies that appear either not to exist or not to support the propositions for which they are cited — a hallmark of AI-generated citations.” In seven cited articles, two appeared to be entirely made up and three others did not exist in the journals they were attributed to, according to Cooper, a judge who was appointed to the federal court in Washington, D.C., in 2014.

This has become a pattern for HHS under RFK Jr. In June of last year, Kennedy issued a report to Congress to support the changes his hand-picked CDC team made to COVID vaccine recommendations. That report included several studies that were unpublished, or that indicated within the study itself that it shouldn’t be cited because more research needed to be done, or which said something completely different than what HHS claimed they said. In May of last year, HHS released a “MAHA Report” on American health that once again claimed certain studies said something other than what they actually said, but which also included cited studies that could not be found upon inspection. The wide speculation was that the report was generated in part using AI that was going through one of its hallucination episodes. One wonders if the AI was doing drugs off a toilet seat with Kennedy himself.

In a real, functioning government, this sort of embarrassment would result in action by somebody, somewhere, somehow. Sadly, we don’t have that. We have a bunch of con-artists cosplaying as government officials instead. But as Kennedy stacks up the losses when it comes to his work and credibility, the calls for his resignation or firing are only getting louder.

Colorado Sees First Lawsuit Under ‘Right To Repair’ Law [Techdirt] (06:19 , Thursday, 03 September 2026)

At this point all fifty states have considered passing “right to repair” law aimed at making it easier and cheaper for consumers (and independent repair shops) to repair their tech. That said, only Massachusetts, New York, Texas, Minnesota, Colorado, California, Oregon, and Washington have actually passed laws. And of those states, none have seen any enforcement despite no shortage of offenders.

So it’s interesting to see the first lawsuit filed in Colorado. Colorado technically has three right to repair laws: one protecting wheelchairs passed in 2022; one covering agricultural equipment passed in 2023; and one expanding coverage to HVAC equipment and most tech in 2024.

A company named Acme Revival, which connects customers with electronics repair technicians, has sued three companies for violating Colorado’s right to repair laws. Three different lawsuits are targeting Toast, a point-of-sale system provider, Owl Labs, a maker of meeting cameras, and Blackmagic Design, a maker of digital camera equipment — claiming they’re violating the law.

The three different lawsuits state that all three companies have made it very difficult for customers to obtain tools, parts, manuals, and firmware/software needed to upgrade and repair point-of-sale terminals, card readers, cameras, and other restaurant-related hardware:

“Acme Revival has received hundreds of requests from owners seeking repairs for Toast devices. The reported problems have included failed batteries and charging systems, damaged housings and touchscreens, malfunctioning card readers and buttons, circuit-board failures, loose or damaged connectors, damaged cables, damaged ports, and other defects requiring replacement parts or technical repair materials.

Acme Revival alleges that it has been unable to complete certain repairs because Toast failed or refused to provide the necessary repair materials.”

There’s really no shortage of large offenders who make it difficult to find parts and tools, buy up independent repair centers to try and monopolize repair (see: John Deere), leverage annoying DRM to make repair difficult or impossible, or engage in the practice of “parts pairing,” which ensures hardware owners can only access large and costly parts assemblages — not individual parts.

The bipartisan anger at such practices has resulted in the right to repair movement seeing the most meaningful traction of any consumer rights issue in the country. Hopefully enforcement steadily scales up to match the full scale of public annoyance.

Working With ICE Is So Toxic, ICE Is Now Offering Liability Insurance To Local Police Officers [Techdirt] (04:04 , Thursday, 03 September 2026)

If you’re worried about the bad optics of working with ICE, the federal government is here to help subsidize your recovery from mass deportation conjunctivitis. If you’re worried about the personal negative side effects of buddying up to ICE’s masked kidnapping squads, the administration is here to assure cops that it might cover some of the legal costs of doing business with ICE.

US Immigration and Customs Enforcement is pitching a plan to help shield local police officers who make immigration arrests from possible financial consequences if they are accused of on-duty misconduct.

The agency is proposing to subsidize liability insurance for state and local officers who are trained and deputized to enforce federal immigration laws, according to a planning document published Friday.

This offer is not valid in sanctuary cities or anywhere cops shops haven’t signed agreements to do ICE’s detention/arrest work for it. To get this extra coverage, law enforcement agencies will have to sign 287(g) agreements. These agreements make local law enforcement agencies part of mass deportation machinery. It requires agencies to hold arrested migrants and tell ICE to come pick them up. It also allows local cops to act as immigration officers by permitting them to perform arrests using ICE administrative “warrants.”

That word is in scare quotes because administrative warrants are just pieces of paper that say ICE knows of someone subject to a removal order. They are not reviewed by magistrate judges. And, unlike what ICE would have you believe, they do not authorize searches of private property.

This is where some of ICE’s (new) billions of dollars might be going. ICE officers don’t need this sort of insurance because they’re defended and indemnified by the federal government. (And they don’t need it anyway because the Supreme Court has made it pretty much impossible to successfully sue a federal officer for rights violations.)

Local cops aren’t nearly as immune as federal officers, so they might appreciate some insurance coverage in the extremely unlikely chance they are sued successfully for violating rights while doing ICE’s work for it. But the payout seems pretty fucking low considering ICE now commands the largest budget of any federal law enforcement agency.

Under the plan, officers would purchase insurance covering up to $500,000 in personal liability, which typically funds legal fees, settlements and judgments. Officers would be reimbursed up to $250 annually — roughly what the insurance is expected to cost.

The administration that claims to love cops (that love ICE) the most, this minimal payout should be viewed as insulting. First, the administration “allows” officers to spend their own money to purchase insurance coverage they wouldn’t otherwise need if their employing agencies had decided signing a 287(g) agreement wasn’t worth the trouble.

Second, tossing cops $250 a year does a whole lot of nothing when it comes to premiums for this specific sort of insurance. And, in other cases, partnering with ICE will automatically void these policies.

Pennsylvania’s risk pool, for instance, recently made clear that it would exclude “proactive immigration enforcement activities” from coverage, forcing several participating counties to search for other insurance options.

Butler County Sheriff Michael Slupe said he found insurance to cover his 13 deputies participating in the program at a cost of $20,000 in annual premiums.

In the first instance, there is no coverage to be had even if the DHS is willing to cough up a measly $250 a year for ICE buddy cops. In the second instance, a local agency is paying $1,538/year per officer to cover officers it has willingly lent to ICE’s anti-migrant activities. That means it’s still on the hook for the other $1,250/year. $3,250 (for 13 officers) looks like a down payment, rather than a meaningful contribution.

But the facts on the ground don’t bother Sheriff Slupe, apparently. He’s sure Trump will come riding the rescue with a fat stack of greenbacks.

“I want to make sure the guys are additionally covered, so we had to spend the money,” he said, adding that federal funding would cover the cost.

Technically almost true, if you read this to mean the federal government will cover an almost-insignificant portion of the cost. But it’s weird to see Sheriff Slupe offer to pitch in on immigration enforcement when his agency was thrown under the bus a bit following an alleged assassination attempt targeting Trump during his 2024 election campaign.

It’s all very stupid and unnecessary. It’s already pretty difficult to successfully sue law enforcement officers, thanks to the ever-expanding coverage of the qualified immunity doctrine (not actually a law!). Furthermore, the federal government’s pitch for additional liability insurance makes you wonder which Trump donors might profit from this push for new premiums. ICE already claims any officers participating in the 287(g) program are “acting under the color of federal authority,” which vastly increases the level of lawsuit immunity. Going even further, the federal government has already pretty much promised local law enforcement officers they’ll be well-defended should they be sued for boarding the ICE bang bus.

The agreements also state that local officers who face civil lawsuits can ask the US Department of Justice to represent them, and that ICE will generally support their requests. 

Adding all of this up, we can only assume none of this adds up. The stipend is too small. The government says local officers should present themselves as federal officers in court proceedings. And these officers seem unlikely to ever need to hire their own representation should they be sued for their ICE-adjacent activities. And now ICE is encouraging participants in the 287(g) program to buy insurance they’ll likely never need or, in some cases, not be able to use due to limits enacted by insurance providers.

It comes across as a blend of stupid and performative. As such, it fits in perfectly with this administration’s MO. But if I were a cop doing ICE’s dirty work, I’d be demanding full coverage paid with federal tax dollars, rather than assume this cock-up of a hybrid will actually do anything when I’ve been sued by competent plaintiffs.

Confused about which VPN is right, US senator asks the NSA for guidance [Biz & IT - Ars Technica] (03:52 , Thursday, 03 September 2026)

A prominent US senator is asking the National Security Agency to provide guidance to the general public on best practices for using virtual private networks to secure their communications from spying by foreign adversaries.

VPNs funnel all of a user’s Internet traffic through an encrypted connection to a remote server. The design provides strong assurances that no one between the user and the server can read the encrypted contents. VPNs also allow users to hide their IP addresses from the destination servers they communicate with. While US agencies have previously recommended use of VPNs, none have given recommendations on which ones provide adequate protection.

It's all in the nuances

There are a host of limitations that can undo many of the protections users may think their VPN provides them. For instance, the encrypted tunnel often terminates once a single server decrypts the traffic and sends it on to its final destination. That means the decrypted traffic or the sending and destination IP addresses may be available for snooping by rogue employees or attackers who hack the server. VPNs also don’t encrypt certain types of metadata, such as time stamps, allowing nation-states to build profiles that can be useful in intelligence gathering.

Read full article

Comments

VMware migration reduces Tottenham Hotspur's licensing fees by 85 percent [Biz & IT - Ars Technica] (02:58 , Thursday, 03 September 2026)

Tottenham Hotspur, a professional soccer team that’s part of the Premier League, has saved over 85 percent in licensing fees by replacing its stadium's VMware instance with Hewlett-Packard Enterprise’s (HPE’s) Morpheus VM Essentials (VME) virtualization software.

Tottenham hasn’t disclosed which VMware products it used or how much it previously paid the Broadcom firm.

The soccer organization confirmed this week to The Register that it has moved its stadium's server, storage, and networking infrastructure to HPE solutions delivered through HPE's hybrid cloud management platform, GreenLake. That is all “underpinned by" VME and HPE's OpsRamp software for hybrid and multi-cloud environments, Rob Pickering, Tottenham's CTO, told the publication, with HPE in charge of the hybrid cloud-managed service.

Read full article

Comments

Hackers Had A Live Feed Of Every ID This Verification Company Scanned. For Over A Year. [Techdirt] (01:57 , Thursday, 03 September 2026)

From the very beginning of this recent obsession with identifying everyone online (yes, they like to call it “age” verification, but it always ends up as identity verification), we’ve been pointing out that it was a huge privacy nightmare waiting to happen. Or maybe it wasn’t waiting. Maybe it was already happening.

This week a massive new data breach has been revealed that should put the nail in the coffin for the idea that any sort of age or identity verification could be safe. 153 million scans of drivers licenses easily available based on this breach, with more being added all the time. Literally on the day it was revealed (and right before the site was taken down) it added another 400,000 records to its available database.

There is no safe age verification. There is no age verification that doesn’t put people at risk.

Last year, Eric Goldman wrote the definitive piece on how all of these technologies — no matter what they tell you — are huge privacy risks, but people are still living in denial. This is despite the numerous examples we’ve had in just the past few years of verification providers and their customers having massive data breaches.

The latest comes to us via Brian Krebs, who reports on a massive breach of scanned IDs — more than 153 million drivers licenses from people across the US and Canada, now for sale on the dark web:

A new identity theft service launched on the dark web this week is selling digital scans of more than 153 million drivers licenses from people in the United States and Canada. Based on interviews with individuals whose licenses are available for purchase on this service, it appears to be siphoning images collected by a widely-used identity verification company based in Louisiana. KrebsOnSecurity also has learned that the New Orleans field office of the Federal Bureau of Investigation (FBI) today launched an official inquiry into the source of the images.

Krebs traces the breach back to an ID verifier that appears to be used by many companies, including Hertz, the rental car company. It appears to not be limited to them either, as he checked with a number of people who were in the database, and by looking at the date they were added alongside their calendars, found examples of other people who shared their ID at places like a pot dispensary.

That company turns out to be IDScan.net, based in Louisiana, which has contracts with thousands of dispensaries, not to mention Hertz, FedEx, and Target. And while Krebs is focused on how many of the leaked IDs are connected to real world businesses, it’s worth noting that IDScan.net is also doing age verification for a bunch of tech companies, has a page tracking state age verification laws and company implementations, and even has written positively about laws like KOSA, the Kids Online Safety Act, that would effectively require age verification.

So, yes, we have a company that is a big player in the age verification space, talking up age and identity verification laws, that appears to have had a long-standing ongoing leak of every ID it scanned.

Yiiiiiikes.

And, of course, like all age and identity verification providers, IDScan has spent years talking up how secure it keeps all this data, even as every single record appeared to be leaking in realtime. Here’s their “Trust Center” page which is still up days after the hack was revealed:

The IDScan.net Trust Center webpage features a security review banner, a search bar, sections for trust and compliance certifications, and a grid of logos from trusted partner organizations.

That’s the company that spent over a year leaking 150 million drivers licenses in real time, explaining “how we protect data, maintain system reliability, and earn the confidence of our customers and their users.” Might be time to update that page.

But also, this should be a massive warning to everyone pushing for age verification laws. You can have a “trusted” company in the space who brags about all the certifications it has. It’s in “compliance” with the GDPR, the CCPA, and every other law. It is “transparent” about its “privacy practices” and how its “sensitive identity data is handled responsibly” and…. for over a year it’s been leaking all of those sensitive records.

And it appears no one internally at the company noticed.

As Krebs makes clear, the breach included many, many millions of records and ID scans that were being swiped in real time by the hackers who breached the system:

The people behind Nexus claim the license images are coming from an active breach at “a major identity verification company” whose customers include multiple Fortune 500 companies.

A table titled "Categories" lists various types of identification documents and the number of records associated with each. There are over 153 million drivers licenses.
The record totals listed by the Nexus identity theft service. The number of drivers license records increased by nearly 400,000 in the span of just 24 hours.

“We have been continuously exfiltrating new data for over a year into our private database,” the service enthused in its introductory post on Exploit. “Records are available to preview before purchase with pertinent information redacted. Customer photos are displayed if available.”

Indeed, over the past 24 hours, the number of drivers license records listed as available in Nexus has increased by nearly 400,000, suggesting that freshly stolen license data is being harvested and uploaded to this service on a semi-regular basis.

And the exposed records aren’t just random members of the public. Krebs found the driver’s license of the sitting Secretary of Defense sitting in there for sale:

A webpage from the NEXUS Identity Document Database shows a locked Minnesota driver's license record for Peter Heg******, featuring a portrait photo of Hegseth and redacted personal details with a "Purchase Record" button at the bottom.

A bargain! Only $100 to get a scan of the Secretary of Defense’s driver’s license.

Anyway, each time we highlight a breach people play it down and insist that mandating age verification is perfectly safe and nothing to worry about. Yet here’s one of the largest identity verification companies in the country, with a pipeline so wide open that hackers had a real-time feed of every government ID it scanned, for over a year, without anyone at the company noticing.

Krebs spoke to a security researcher at Cybera, named Larry Baldwin, who talks about how this kind of data can do real damage:

Baldwin said the Nexus identity theft service presents multiple serious security and privacy threats, noting that state-issued drivers licenses are commonly used as proof of one’s identity when opening new lines of credit. Baldwin said the service could also dangerously expose many people who do not wish to be found but who cannot meaningfully change their appearance (or at least not enough to fool today’s AI-based image matching tools).

This category of people, he said, includes those fleeing domestic violence, and even people who have been assigned a whole new life and identity as part of the federal government’s witness protection program, which is generally reserved for criminal defendants in racketeering and conspiracy investigations who agree to cooperate with federal authorities.

“Just when it seems like we’re making some headway in improving authentication controls through drivers license verification systems, this happens and the very thing those improvements are dependent on are compromised,” Baldwin said.

At this point, anyone still supporting age verification requirements, especially claiming it’s for “child safety,” should have to answer for all the millions of people put needlessly at risk due to data breaches like this.

You cannot do age or identity verification safely. It always creates some sort of record and that set of records will always become a target. That’s what happened here. And it’s what will happen with any such systems.

Virginia Tech vs. VMI: A brief history of rivalry [www.collegiatetimes.com - RSS Results for * of type article OR video OR youtube OR collection] (12:00 , Thursday, 03 September 2026)

Virginia Tech’s 2026-2027 football season officially kicks off this Saturday as the Hokies take on the Virginia Military Institute Keydets at Lane Stadium. While most Hokie fans recognize the University of Virginia as being VT’s main rival, VMI holds a…

“Big Yellow Taxi” [35mmc] (11:00 , Thursday, 03 September 2026)

The song by Joni Mitchell contains the much quoted line “you don’t know what you’ve got till it’s gone”. Every time I came back from town I passed within a couple of metres of “County”, a small, old and battered, tracked bulldozer / tractor. For many years she has been there gradually deteriorating in summer...

The post “Big Yellow Taxi” appeared first on 35mmc.

How the Wide-Tire Revolution Took Off [Rene Herse Cycles] (10:59 , Thursday, 03 September 2026)

Twenty years ago, we started testing tires on real roads, with a rider on the bike. Almost by accident, we found something that had the potential to revolutionize the cycling world: High pressure doesn’t make tires roll faster. Before, we all thought that tires were either wide or fast. Our results showed that tires could be wide and fast. You can catch up on that part of the story here.

Today we’ll look at how these exciting findings flew under the radar for a while, until a chance meeting pushed them into the racing world. Suddenly, Lance Armstrong was tweeting about rolling resistance, and the Cervélo TestTeam was testing 25 mm tires. Engineers from ZIPP called me and later sent a top-of-the-line wheelset, as a ‘Thank You’ for sharing our research. And today, everybody agrees that wide tires can be fast, and nobody inflates their tires to 130 psi (9 bar) any longer. Here’s the untold story of how we got to where we are now.

In September 2006, we published the results of our first tire tests. New findings are often met with healthy skepticism. That’s natural and useful. It makes sure that new findings are sound before they are widely accepted. After Bicycle Quarterly 17 landed in people’s mailboxes, online discussions raged for weeks. People questioned whether there was really no wind when we tested, whether our test rider could maintain the same position for run after run, and much more. We pointed to our statistical analyses and all the other safeguards we had built into our testing, but it was hard to convince the cycling world. Our results were just too far out there. The controversy wasn’t always fun, but it spurred us to further check our results and expand our research—more about that later.

Of course, there were also experts who were excited about our findings. Not just the people who had reviewed our work, like Frank Berto and Jim Papadopoulos, but also local legend Ric Hjertberg, one of the founders of Wheelsmith, Gerard Vroomen and Andy Kessler (above), and a few others. However, the mainstream cycling world ignored our results. From time to time, a BQ reader would comment on an online tech article about tire performance, pointing to our research. They were treated like conspiracy theorists who comment on a story about NASA, insisting that the moon landings never happened. It seemed that nobody was willing to engage with the substance of our research. To some degree, that made sense: We were just a few unknown cyclists, and there just isn’t enough time to examine every outlandish claim about bicycle performance. And the idea that high pressure doesn’t make tires roll faster seemed pretty outlandish at the time.

Twenty years ago, nobody mentioned ‘wide tire’ and ‘fast’ in the same sentence. Even some friends asked me: “If wide tires roll faster, why don’t pro racers use them?” To which I usually replied: “Just give them time…”

A chance meeting pushed our research into the mainstream: In 2007, we started to work with the Department of Aeronautics & Astronautics at the University of Washington. With Boeing in town, they have one of the world’s best wind tunnels. We were able to get two full days in the wind tunnel to test things nobody had tested: Whether wide tires are less aero than narrow tires. Whether fenders can make a bike more aero. (The photo shows the telescoping fender that allowed us to test different fender configurations.) How much clothes affect aerodynamics. Whether front or rear racks are more aero. All the questions that matter to long-distance cyclists, but that nobody had cared to ask. During those wind tunnel tests, we also confirmed that our test rider was able to assume the same position time and again. (The photo also reminds me that fashion in cycling tights has changed in the last 20 years!)

A few days before the test, the team from the wind tunnel put me in touch with Len Brownlie, a Canadian sports aerodynamics consultant. In the past, he had rented the UW wind tunnel for work with Lance Armstrong, Nike and other big names in sports. Now he wanted to piggy-back onto our wind tunnel session for a quick test of zig-zag strips that he mounted on the bike frame. He hoped they would improve aerodynamics. No problem: We tested his strips first thing. (They didn’t make the bike more aero, sadly.)

Len stuck around for the day, watching what we were doing and offering advice. We were using a test rig they had built for Lance. It held the bike on a streamlined platform, and also rotated it automatically to measure the effect of cross-winds. We appreciated Len’s advice and experience, and he was clearly impressed by what our work, especially the rigorous statistical analyses that Mark was doing to check our results. During a break from the wind tunnel testing, we talked about our tire research, which was still fresh and had occupied our minds for much of the last year. I gave Len a copy of Bicycle Quarterly 17 with the results of our first tire tests.

Len was one of those people whose name you don’t see in the press. I suspect that teams that work with him prefer that the competition doesn’t know about him. (Google him, and you’ll find plenty of scientific articles, especially on the aerodynamics of fabrics from his work with Nike.) Once we gave that Bicycle Quarterly to Len, our tire testing article was apparently shared widely. The following year, Lance Armstrong tweeted: “Lots of talk these days about tires/rolling resistance too.”

A little later, Cervélo TestTeam started testing 25 mm tires—perhaps not coincidentally the width that had proven fastest in our testing. They also started testing tire pressures, where before everybody had just inflated to maximum pressure. Gerard Vroomen, founder of Cervélo, told cyclingnews.com: “Every mechanic of every professional cycling team puts too much pressure in the tyres. We’ve weaned them off of that, so at least there’s not 12 bar in them any more.”

By 2010, Cervélo TestTeam was running 25s in the Tour de France, the first team to do so. This was new in more ways than one: Technical developments at the time usually originated in Europe, especially with respect to tires: All the big tire companies that sponsored pro racers were in Europe. But the move to 25 mm tires started in Canada, shortly after we had given our results to a Canadian consultant to the pro cycling world. From there, it spread to the U.S., and only then to Europe.

The wheel sponsor of the Cervélo TestTeam was ZIPP, which had just been acquired by SRAM. Out of the blue, I got a call from Michael Hall, Zipp’s director of R&D. He was working with Josh Poertner, then still at Zipp, on the next generation of wheels. They were interested in our research. Over the years, Michael and I had many conversations about tire tech, tubeless standards and other topics. This started a long friendship that continued as he moved up in the SRAM company hierarchy. Among other things, that’s how we got prototypes of the first XPLR components for testing, which I still run on my Firefly…

Then Michael told me he was sending me a set of their top-of-the-line tubular wheels as a ‘Thank You.’ That was a bit of a mystery until I saw SRAM’s new ‘White Paper’ about suspension losses. It echoed our findings that high pressure didn’t make bikes faster, because energy was lost due to vibrations. Josh Poertner once characterized their research: “We are not PhD’s and don’t consider this science at all but rather race engineering… as much as I’d love to know why any of this happens, that’s just not the job we’re paid for.” That makes sense: They were trying to figure out what would make their racers fastest, based on whatever information they had. The underlying science apparently came from Bicycle Quarterly. And the carbon wheelset was a (much-appreciated) Thank You.

Others had also seen our research. For example, Global Cycling Network talked about how wide tires were faster. They were the only ones who actually mentioned Bicycle Quarterly. We really appreciated that. Back then, we were on a quest to change the (cycling) world, and we were happy to share our findings with anybody who was interested. Even so, it’s still nice to be given credit.

Back then, many in the bike industry apparently thought that we were just two unknown kids. Bicycle Quarterly may have flown under the industry radar, but we already had more than 5,000 readers all over the world. Even today, people continue to underestimate us. Just recently, a mainstream cycling expert told me after we mentioned his work in the RH Journal: “I had no idea how many people read your blog!”

In the meantime, we continued our research. Our wind tunnel tests confirmed that our rider can assume the same position time and again. That took care of a major criticism of our roll-down tests.

We used rumble strips to quantify suspension losses and found that the added resistance due to vibrations explains why high pressures do not provide an advantage on real roads.

We replicated our roll-down tests with power meters on road and track. The results are the same, confirming our original findings with a different methodology.

We expanded our testing to ultra-high pressure, and found that even 200 psi (13.8 bar) did not improve a tire’s speed (above). At the other end of the spectrum, we tested ultra-low pressure. We were surprised that tires continued to roll with almost undiminished speed, almost to the point where the bike started wobbling because the tires had so little air. We also expanded our tests to tires up to 54 mm wide and found that they rolled as fast as narrow tires.

Together, all these findings left no doubt: High pressure is not helping tires roll better. And: Wide tires are as fast as narrow tires, provided they use the same supple casings.

By now we had given up hope that our research might persuade tire makers to develop wide, supple and fast tires. We decided to take matters into our own hands. We talked to various suppliers, settled on a spec, and designed tire molds. In 2014, we introduced our tire program, at first with 700C tires from 26 to 38 mm wide, plus 650B models. I still remember the engineers from our supplier asking: “Why do you want to make wide high-performance tires? Who will ride them?”

The obvious answer was: We wanted those tires for our own adventures. But there were plenty of others interested in wide-and-supple tires. This included some of the first gravel racers. Above is Matt Surch on his way to winning the 2016 Steaming Nostril on the first-generation 700×35 Bon Jon Pass tires. Since then, we’ve continuously updated our tires, even if the names have remained the same.

We started working with individual racers and also the Cinch and Mazda-Lauf gravel teams. We gave them tires to test as part of our R&D. We also shared our research with them. Soon Ted King, Lauren de Crescenzo, Brennan Wertz and Jenna Rinehart were winning races on tires that were wider than those of other racers. It’s hard to know whether that caused other gravel racers to move to wider tires, or whether racers simply were curious about wider tires and found that they rolled faster on rough gravel. It’s probably a combination of both.

Many in the bike industry started riding Rene Herse tires on their personal bikes. They were excited about the ride and speed of our tires, and they sent us photos. Above is the bike of Bicycling magazine’s Senior Test Editor Matt Phillips…

…and here is ZIPP Product Manager Bastien Donzé’s bike. Our friends at ZIPP equipped their bikes during dealer product launches with our tires, too, to show their wheels in the best light. Of course, not everybody was at liberty to be as open about it as Andy and Gerard from OPEN, whose personal bikes were featured as ‘Bike of the Month’ on their website (top photo). All these personal experiences probably did more than anything to get wide tires accepted in the mainstream.

That’s where things stand today. Pros have moved beyond 25 mm tires to 28 and even 30 mm on the road, although they’ve recently pulled back a bit for time trials. The reason: With a tire that wide, you need very deep wheels for good airflow, and now the UCI has limited wheel depth. Plus, our research shows that the performance benefits of wide tires level out above 25 mm: Wider tires are neither faster nor slower, at least on smooth roads.

Gravel racers have been moving to wider and wider tires. On rough surfaces, there’s no doubt that wider tires roll faster. Our tires continue to win races, which hasn’t gone unnoticed by other tire companies. Major parts of our research have been replicated by the Escape Collective, with similar conclusions. A study at the University of Delft has confirmed our rumble strip testing. Narrow 23 mm tires and 130 psi pressure seem ludicrous today, but that was state-of-the-art when we started our research.

Twenty years is a long time, but in terms of acceptance of new scientific ideas, it’s actually pretty short. It’s exciting that our conclusion from 2006 no longer reads like a far-out dream, but accurately describes the reality today:

“For most cyclists, wide, supple tires at low pressures offer more speed, better comfort, increased versatility and improved safety than the currently favored narrow high-pressure tires.”

When I’m watching the final stage of this year’s Tour de France and see the racers zoom up and down the cobbles of Montmartre in Paris on their 30 mm tires, and when I go out for a ride and see so many cyclists on wide tires, I’m reminded how much our research has changed the cycling world for the better. And that’s really all we wanted when we published our first test results, 20 years ago.

More Information:

Photo credits: Marc Gasch/OPEN Cycle (Photo 1); Alex Wetmore (Photo 3); Mark Vande Kamp (Photos 5, 6); Matt Phillips (Photo 10); Bastien Donzé (Photo 11); Jered Gruber (Photo 12)

Rockymounts BackStage SwingAway Long-Term Review: The Bike Rack for Van Dwellers [BIKEPACKING.com] (09:16 , Thursday, 03 September 2026)

RockyMounts BackStage SwingAway ReviewSix years ago, Miles and Emily moved to British Columbia’s Sunshine Coast, trading van life for a more rooted existence. That brought the need for a dedicated bike rack for quick shuttles around town. The Rockymounts BackStage Swing Away Bike Rack stood out for its full 180-degree swing function, which permitted access to the back of the van, a high load limit, and a minimal profile. Find Miles’s long-term review here…

The post Rockymounts BackStage SwingAway Long-Term Review: The Bike Rack for Van Dwellers appeared first on BIKEPACKING.com.

MADE 2026 on Video (Part 10): If You Could Ride Anywhere… [BIKEPACKING.com] (08:47 , Thursday, 03 September 2026)

Made Video Part 10For our final video from the 2026 MADE Bike Show, Neil asked 33 framebuilders and a few component makers one big question: If you could take a month off work, where would you go bikepacking? Their answers include dream routes and a few far-flung destinations that might surprise you...

The post MADE 2026 on Video (Part 10): If You Could Ride Anywhere… appeared first on BIKEPACKING.com.

GeerTop Blazer Review: Six Years in a $150 Tent [BIKEPACKING.com] (07:30 , Thursday, 03 September 2026)

Crust Nor'Easter ReviewDespite being tempted by lighter, more spacious shelters over the years, Nic has remained loyal to the very first single-person tent he picked up. A packable, relatively roomy budget option that costs just $150, the GeerTop Blazer seems too good to be true. In his review, he runs through all the elements that have kept him hooked on this budget banger…

The post GeerTop Blazer Review: Six Years in a $150 Tent appeared first on BIKEPACKING.com.

KI6CR: Two Swiss Summits – SOTA Activations at Schilthorn and Burgfeldstand (Part 2) [Q R P e r] (06:00 , Thursday, 03 September 2026)

by Chris (KI6CR) What a beautiful week for radio in the Alps! My family and I were taking one last summer trip before the kids headed back to school. We ended up basing ourselves in Interlaken for a week of exploring. I managed to sneak away for two SOTA activations along the way, and both … Continue reading KI6CR: Two Swiss Summits – SOTA Activations at Schilthorn and Burgfeldstand (Part 2)

5 Frames of Washi F in Mach. [35mmc] (05:00 , Thursday, 03 September 2026)

I had a day out in Machynlleth recently. I took with me an unusual camera that I have been experimenting with (to be written up in a later post) so I thought just to make sure the day was a complete failure photographically I would also take with me my Super-Ikonta 530/16 loaded with Washi...

The post 5 Frames of Washi F in Mach. appeared first on 35mmc.

Wednesday, 02 September 2026

Second Quest [Tedium] (11:24 , Wednesday, 02 September 2026)

In a normal culture, a company as gray-haired as MapQuest would not be having a straight-on cultural revival in 2026. But our culture is not normal.

Second Quest

Last year, I wrote about MapQuest, and presented the company’s story as one that was fundamentally changed by a poor decision—AOL’s purchase of the company, which only completed after the tech IPO bubble burst.

AOL, like its merger partner Time Warner, was a frequent target for acquisition. But unlike the modern day Warner Bros. Discovery, it lost cachet every time an acquirer swooped in to pick it up. And that left companies like MapQuest to languish in a culture that simply was too big to understand what was actually going on.

But MapQuest, as I noted at the end of that piece, seemed to land in the perfect situation: In 2019, it was acquired by a company that actually made sense for it, the StartPage owner System1. Verizon essentially gave them the company for free, and System1 responded by actually making the company something of a viable player. It just needed something to put a little gas in its engine.

Last year, the Trump administration gave them a prime PR opportunity by renaming the Gulf of Mexico to the “Gulf of America,” which MapQuest handled by letting you rename the gulf yourself. And last week, as the president tried the same trick on Lake Ontario, MapQuest repeated the trick—which stood out all the more because Google, and later Apple, chose to acquiesce.

Those companies have something to lose if they make Trump mad. System1, and by extension MapQuest, does not. And audiences responded by making MapQuest the hottest app in the App Store for the first time ever.

It sure feels like MapQuest was ready for its close-up.

I’ve covered a lot of companies in my time, and I can name on one hand the number of times a stodgy technology firm was effectively at the brink of obsolescence only to pull off a complete cultural comeback. (Apple is, famously, one. Thank you, Dr. Gil Amelio.)

MapQuest wasn’t on my bingo card. But it reflects an inherent flexibility that companies who aren’t on the top of the tech world have, especially if they don’t sell physical products. Companies like Mozilla and DuckDuckGo simply have more room to get savvy as tech giants find themselves at the mercy of a transactional presidency. But that extends beyond political; for example, Mozilla earned plaudits for publicly stating that it would continue to support the popular ad-blocking extension uBlock Origin, which Google seemingly upgraded its extension platform to disable. It’s a great time for easy wins if you lead a mature tech company.

That said, the hard part for MapQuest will be sticking the landing. While the app is ready for prime time, it lacks certain features that many users take for granted. For example, if you want to take the bus, MapQuest doesn’t have in-depth mass transit integration like Google or Apple’s offerings do. And good luck using it to book an Uber.

But if this sudden moment at the top of the news cycle turns into something lasting, it might just give the legacy map provider an opportunity to fill those gaps while building a lane for itself. (The company has already pledged to add Android Auto support.)

At a broader level, MapQuest’s story offers a template for other legacy tech companies to make a modern-day comeback. MySpace’s current owners, Tim and Chris Vanderhook, have already made a recent case for a social media “antidote” play. And at least some players, like Flickr, have found ways to stick around, even with declining relevance.

But there’s a real risk that the revival might end up being like what is currently happening to Digg. That company, relaunched with much fanfare by Kevin Rose and Alexis Ohanian, had to pivot just months after its relaunch, as it became clear the domain was too easy a target for spammers. The website is now best described as a Techmeme for AI bros. That feels like a pretty easy-to-ignore model.

It won’t work for every old company with a dusty server, but the play is clear: Position yourself as an antidote to the negative parts of modern tech culture, and then position your marketing to exploit that position. That requires smart planning and a willingness to meet the moment, even if it means being divisive or stepping on toes to reach that moment.

It also probably requires a brass stomach, because it can go bad pretty easily. But I guess I’m optimistic that MapQuest is going to start a trend where mid-sized tech companies suddenly discover the backbone that was there all along.

Directionless Links

The first issue of Boxed News, my new mini PC newsletter, landed yesterday without a hitch, and things are off to a good start. Wanna support a new newsletter built from the ground up? Sign up here.

There’s no reason whatsoever to give the NES a CD add-on. Especially in 2026. But I ain’t gonna lie, seeing this project gave me a certain sense of joy, especially as it uses the previously vestigial NES expansion port. (We love vestigial features around these parts.) Word of warning that the video is meandering as all get-out but the result makes it worth the dig.

We actually have multiple videos in the subgenre of “video game consoles doing things they shouldn’t be able to do” this week. Apparently it is possible to run Mac OS X Tiger on an Xbox 360, whose processor is effectively a PowerPC G5 minus a couple of features. If you feel like following in this dude’s footsteps, here’s the GitHub project.

New social network alert: One of the more interesting pitches that landed in my inbox this week was for Innie, a MySpace-esque social platform that focuses on design, avoids profile pictures, and promises feeds that don’t go on forever.

--

Find this one an interesting read? Share it with a pal!

And support Tedium by giving the Tedium Shopping Network a spin.

I rented a car, and within hours, my driver's license was for sale [Biz & IT - Ars Technica] (04:32 , Wednesday, 02 September 2026)

Not long ago, I rented an SUV from a well-known car rental company. Within hours of an employee scanning my driver's license, a high-resolution scan of my ID was available for sale on the dark web.

An exposé published Tuesday by KrebsOnSecurity reports that my license was one of more than 153 million that were available through Nexus, the name of the new ID theft service. Like other driver's licenses available there—including some belonging to journalist Brian Krebs, his mother, an FBI assistant director, and several security researchers—my license was purported to include multiple image files showing both the front and back of the ID. Besides a basic image scan, the files also captured the images in the infrared and ultraviolet spectrums. Presumably, the additional formats may allow cloned-based counterfeit IDs to pass hologram tests.

Growing by the day

Besides advertising the availability of driver's licenses, Nexus offered to sell a bevy of other forms of ID. They included:

Read full article

Comments

The grass is greener at Virginia Tech [www.collegiatetimes.com - RSS Results for * of type article OR video OR youtube OR collection] (03:00 , Wednesday, 02 September 2026)

With the dramatic transitions that occur while adapting to college, what also comes is the inevitable meltdown sessions in your freshman year dorm. Or maybe that was just me last year, sobbing with the lights off in my creaky lofted…

Songs Hokies should know before stepping into Lane Stadium [www.collegiatetimes.com - RSS Results for * of type article OR video OR youtube OR collection] (12:00 , Wednesday, 02 September 2026)

Along with the vibrant maroon and orange sea of a gameday crowd, the scent of freshly cut grass and cooking turkey legs, gamedays in Blacksburg are also full of iconic sounds. Beyond the roar of thousands of Hokie fans, Lane…

Street Photography is Tough [35mmc] (11:00 , Wednesday, 02 September 2026)

Street photography is tough.  The way I see it, there are at least two approaches by which to achieve success.  One requires stealth, quick reaction time, finely tuned powers of observation and anticipation.  It requires an intimate comfort with your camera of choice. The other method is to engage with your subjects.  Play with them....

The post Street Photography is Tough appeared first on 35mmc.

Welcome to Bikepacking Paradise: Three Days at Pedaleo [BIKEPACKING.com] (10:43 , Wednesday, 02 September 2026)

Pedaleo 2026 RecapPedaleo is a three-day bikepacking festival in Switzerland packed with inspiring talks, hands-on workshops, a vibrant expo, social rides, outdoor movies, live music, lake swims, and more. Leoni Kolberg was one of more than 1,000 participants, and she wrote a detailed reflection on what made the weekend special. Find her recap and photos from the event here...

The post Welcome to Bikepacking Paradise: Three Days at Pedaleo appeared first on BIKEPACKING.com.

Panorama Taiga 2 Sees Updates and New Geo [BIKEPACKING.com] (09:08 , Wednesday, 02 September 2026)

Panorama Taiga 2 NewReleased five years ago, the Panorama Taiga is a modern Reynolds 725 hardtail that's equally suited for trail riding and bikepacking. Panorama just updated the Taiga with several small changes and geometry tweaks to better suit technical trails. Check out the updated Panorama Taiga 2 here...

The post Panorama Taiga 2 Sees Updates and New Geo appeared first on BIKEPACKING.com.

Our Global Routes Map Just Got the Overhaul Everyone Has Been Asking For [BIKEPACKING.com] (07:54 , Wednesday, 02 September 2026)

New Bikepacking Routes MapIf you ride off-road, you’ve probably seen our bikepacking routes map, which means you’ve also encountered its limitations. We’ve been listening to your feedback, and today, we’re thrilled to unveil our completely revamped interactive map of nearly 550 purposefully-built and curated routes worldwide. Learn what’s new and explore all the latest features here…

The post Our Global Routes Map Just Got the Overhaul Everyone Has Been Asking For appeared first on BIKEPACKING.com.

BGP hijack infecting networks caused by a comedy of errors that’s not funny at all [Biz & IT - Ars Technica] (07:00 , Wednesday, 02 September 2026)

Hackers carried out a supply chain attack that installed malware on networks using an unusual technique: hijacking a chunk of Internet space where cloud management software used by hosting providers, data centers, and other large infrastructure companies is updated.

In a well-coordinated operation, the unknown attackers exploited weaknesses in the routing security setup of hosting provider Hetzner Online and the process for attaining valid TLS certificates. The lapses allowed the attackers to successfully perform a BGP (Border Gateway Protocol) hijacking to obtain control over IP addresses assigned to Softaculous. The company, based in the United Arab Emirates, is the maker of a platform for installing and managing Web software and is the developer of Virtualizor, a management platform for virtualized environments.

Softaculous used the IPs to issue updates and host a client and billing site. With control over the hijacked space, the attacker was now using the addresses to push malware masquerading as updates to unsuspecting users.

Read full article

Comments

Born Chaotic [35mmc] (05:00 , Wednesday, 02 September 2026)

A film needs to be born twice before it shines. The first birth is the exposure. When a shutter is fired, a certain amount of light, bounced from physical objects, passes through glass and tiny openings, engraving an image onto a blank film. That’s the first birth we all know about. But the second, which...

The post Born Chaotic appeared first on 35mmc.

Tuesday, 01 September 2026

Back 2 Skool Vinyl Night! (9/3) [WUVT-FM 90.7 Blacksburg, VA: Recent Articles] (09:15 , Tuesday, 01 September 2026)

You read that right! Back 2 Skool Vinyl Night!!!

This Thursday, September 3rd 6pm - 9pm @ Rising Silo Brewery

No skipping this early in the semester. Be there or be square... or quadrilateral... or whatever your favorite 4-sided shape is...

b2sVN

Poster by Eve Ullman

Let’s Go Hokies: The Collegiate Times’ favorite gameday traditions [www.collegiatetimes.com - RSS Results for * of type article OR video OR youtube OR collection] (03:00 , Tuesday, 01 September 2026)

Whether it be walking across the iconic Drillfield or attending a basketball game at Cassell Coliseum, there is no shortage of experiences that make Virginia Tech unique. However, one of the most distinctly Hokie experiences may just be the crisp…

ROOPHLOCH 2026 announced [Open source software and nice hardware] (02:33 , Tuesday, 01 September 2026)

+++ Tuesday  1 September 2026 +++

ROOPHLOCH 2026 announced
========================

Solderpunk has announced ROOPHLOCH 2026 [1]

If you are not familiar with ROOPHLOCH, read the announce, it is
a wonderful event!


Anchors in time
---------------
Very often, when a new edition of Sacha's wonderful weekly Emacs news
appears, I am astonished that another week has already passed.

I had the same kind of astonishment when reading Solderpunk's announce:
This means another year has gone by!

It is good to have these kind of anchors in our life. They keep us
from drifting and help us to stay grounded. But they can be quite
shocking non the less :)


Thoughts on ROOPHLOCH
---------------------
Every year I love to read the ROOPHLOCH posts. People can do such
amazing things!

I would love to participate, but I think this is very hard for me. I
did once wrote an ROOPHLOCH phlog on a Palm device, but it is a bit
silly to do that again. And I don't have that many special devices.

Also, writing a phlog on my Palm doesn't really fulfill all the
requirements. Although using a Palm I write the phlog outside and
off-grid, I can't publish a phlog from an off-grid position, this I
can only do from my home network.

But, there are 30 days left to think about this, so who knows...


[1]: gopher://zaibatsu.circumlunar.space/1/~solderpunk/phlog

Last edited: $Date: 2026/09/01 20:33:00 $
   

WUVT Fall 2026 Orgy... [WUVT-FM 90.7 Blacksburg, VA: Recent Articles] (12:35 , Tuesday, 01 September 2026)

The MOST anticipated event of the semester is here!

WUVT ORGY, Friday, Sept. 4th @ 6pm in Squires Rm. 305! Learn more about college radio and how YOU can get involved. See you there...

F26Orgy

Poster by Cailin F

Pumpkin spice it up [www.collegiatetimes.com - RSS Results for * of type article OR video OR youtube OR collection] (12:00 , Tuesday, 01 September 2026)

It’s that special time of year. The grass will be more brown than green in a few weeks, the horror movie trailers are beginning to roll out and Virginia Tech’s first home game is a skip and a jump away.…

A Study in Deep Shade – A One Shot Story [35mmc] (11:00 , Tuesday, 01 September 2026)

This image was made one Saturday morning at about 8.30AM. At the time, the southern counties of England experienced three weeks of extreme heat, with temperatures rising to 30C or more during the day. I wanted make an early trip to the supermarket to buy some basic supplies and then to shelter indoors. Each day...

The post A Study in Deep Shade – A One Shot Story appeared first on 35mmc.

From Alnwick to the Airwaves: Activating Ros Castle (G/SB-009) with the KH1! [Q R P e r] (10:17 , Tuesday, 01 September 2026)

by Thomas (K4SWL) On Wednesday, July 8, 2026, my family visited one of our favorite places in Northumberland: Alnwick. If you’ve never heard of Alnwick, it’s a beautiful town where one of my daughters spent much of May before our family UK trip, living in rooms within the castle walls and getting to know the … Continue reading From Alnwick to the Airwaves: Activating Ros Castle (G/SB-009) with the KH1!

MADE 2026 on Video (Part 9): Framebuilders Choose Their Dream Bike [BIKEPACKING.com] (10:15 , Tuesday, 01 September 2026)

Made 2026 Video part 9For part 9 of our 2026 MADE Bike Show video coverage, Neil asked 29 framebuilders one simple question: If you could steal any bike from the show, which one would it be? Their answers shed light on some of the most inspiring, surprising, and beautifully crafted bikes at MADE 2026...

The post MADE 2026 on Video (Part 9): Framebuilders Choose Their Dream Bike appeared first on BIKEPACKING.com.

ORTLIEB Introduces In-House TIDURA Waterproof Material [BIKEPACKING.com] (10:00 , Tuesday, 01 September 2026)

ORTLIEB TIDURAIn an industry first, ORTLIEB just unveiled its own house-made waterproof fabric, TIDURA, which is available now on several products in multiple colors. Learn more about the German brand's impressive new material here...

The post ORTLIEB Introduces In-House TIDURA Waterproof Material appeared first on BIKEPACKING.com.

Bedrock Mountain Clogs Now Come in Chanterelle and Wild Rose [BIKEPACKING.com] (08:56 , Tuesday, 01 September 2026)

Bedrock is releasing two fresh Mountain Clog colorways inspired by vibrant fall colors today. Take a peek at the new Chanterelle and Wild Rose Bedrock Mountain Clogs here...

The post Bedrock Mountain Clogs Now Come in Chanterelle and Wild Rose appeared first on BIKEPACKING.com.

The 2026 Ratio Drop Bar Conversion Kit Enables SRAM Eagle 70/90 Compatibility [BIKEPACKING.com] (08:47 , Tuesday, 01 September 2026)

2026 Ratio Drop Bar Conversion KitRatio Technology's latest conversion kit brings SRAM Eagle 70/90 Transmission shifting to drop bars. With a CNC aluminum piece that attaches to the derailleur, users get all the benefits of SRAM’s new mechanical groupset with their drop-bar brifters. Read on for more details…

The post The 2026 Ratio Drop Bar Conversion Kit Enables SRAM Eagle 70/90 Compatibility appeared first on BIKEPACKING.com.

Sicilia Traversata [BIKEPACKING.com] (07:26 , Tuesday, 01 September 2026)

Sicilia Traversata Bikepacking RouteThroughout Italy, Sicily is widely regarded as a place of magnificence. The friendliest Italians, the tastiest food, rich history and mythology, and, of course, a popular holiday destination. It is […]

The post Sicilia Traversata appeared first on BIKEPACKING.com.

How I Meter Slide Film [35mmc] (05:00 , Tuesday, 01 September 2026)

Slide film is beautiful! Online though, it feels like there’s a lot of trepidation regarding metering for slide film. The common advice I’ve heard is to bracket your shots, so that you can feel confident that at least one of the 3 or so is properly exposed. I think this is good advice if you’re...

The post How I Meter Slide Film appeared first on 35mmc.

Monday, 31 August 2026

How accurate have Ed Zitron's AI skeptic predictions been? [] (08:00 , Monday, 31 August 2026)

I was curious how well the predictions of the most widely cited AI skeptic I've seen (Ed Zitron) have done, so I looked at how his predictions panned out. To disclose my own biases, I've never had a particularly strong pro or anti AI progress position. For example, in 2022, I did a comprehensive look at predictions Futurists made, including well-respected folks like Kurzweil and found them to be generally wrong on both the prediction results as well as the reasoning. On the flip side, in 2015, I wrote about how people were underestimating AI's ability to displace humans in jobs and have repeatedly been on the record as saying that many people are underestimating AI's ability to displace humans from jobs. My position on AI has been extremely boring and is basically, "if something is currently happening, the people who are saying that it's impossible that it will ever happen are probably wrong".

One comment I've seen from a lot of AI skeptics when someone responds to an AI skeptic is that all of the people who are saying that AI isn't fake are self-interested liars. Personally (to my obvious detriment), I have no particular financial interest in AI companies. I own whatever the standard share of them is via boring index funds. I have some seed stage investments, but just due to the timing and what's gotten big, that part of my portfolio is underweight on AI. I don't work at an AI lab or a company that supplies AI labs. I've mentioned being hilariously bad at interviews before, and I did interview at an AI lab a number of years ago and failed the phone screen in a performance that was the kind of performance that must've inspired Jeff Atwood's famous Why Can’t Programmers... Program? where he concludes that there must be a lot of fake programmers out there because nobody could fail a coding interview that badly if they knew how to program. I don't benefit in any particular way if AI does well, except insofar as anyone who holds broad index funds benefits, but I do care about accuracy.

2024: Meta, Google, and Microsoft are dying

Because there are quite a few prediction results, let's look at one in detail before the complete list to get an idea of the kind of reasoning Zitron uses. We'll arbitrarily look at this November 2024 talk where Zitron says, among other things, the major tech companies (like Meta and Google) are dying and they're thrashing around on AI because they don't know how to grow.

Zitron specifically named Meta as a company that's dying ("it's a dying product, and it's kind of a dying company"). Meta's revenue and profit (GAAP operating income) have been

PeriodRevenueProfit
Amount%Amount%
2023$135B16%$47B62%
2024$165B22%$69B48%
2025$201B22%$83B20%
First half 2026$117B30%$42B10%

When he talked about companies not knowing how to grow ("none of these companies anymore really know how to grow ... in the desperation to try to reignite growth in a dying ecosystem the tech industry is going to shove this [AI] shit into everything"), he named Google and then Microsoft. Alphabet (Google's parent company) has had the following revenue and profit numbers:

PeriodRevenueProfit
Amount%Amount%
2023$307B9%$84B13%
2024$350B14%$112B33%
2025$403B15%$129B15%
First half 2026$230B23%$80B30%

And Microsoft's numbers have been (note that, for consistency, all numbers are calendar year numbers and not fiscal year numbers):

PeriodRevenueProfit
Amount%Amount%
2023$228B12%$101B21%
2024$262B15%$118B17%
2025$305B17%$143B21%
First half 2026$173B18%$79B19%

Although this wouldn't be in the spirit of Zitron's statement, one could argue that Meta is actually dying, it just hasn't died yet. However, the reasoning in Zitron's argument is incorrect here—the Meta, Google, and Microsoft ecosystems are not dying. Given how fast these companies are growing (in terms of revenue and profit), it doesn't seem that AI is, as Zitron implied, some kind of desperation move they're reaching for because "they don't know how to grow" and are all out of ideas. I don't think it's worth spending this much text on each prediction, but the pattern Zitron used here is illustrative.

To make the case that these things are dying, he pulls on minor issues that are not positioned to cause the very large changes he suggests are about to occur. For Meta, he cited some kind of alleged MAU drop for Facebook. Rather than use Meta's own MAU figures or any kind of revenue or profit numbers, he seems to have used numbers from Similarweb. My experience with 3rd party tracking numbers like this is that they're quite inaccurate and generally useless for anything other than a rough order of magnitude comparison, making the Zitron's cited decline meaningless. FB stopped reporting MAU publicly in December 2023, but most estimates have FB MAU increasing over time and the numbers Meta does report show generally increasing usage over time for their products; Zitron cherry-picked an outlier low estimate to make his point.

For Google, he cites Prabhakar Raghavan, who he calls truly evil and "a computer scientist class traitor that sided with the management consultancy sect", as having done some kind of grievous damage to Google search. In his rants about Raghavan, he never credibly establishes that Raghavan is doing severe harm to Google search, and the Google search engineers who've commented on his rant don't seem to agree with the Raghavan as sole or even major reason for search issues hypothesis.1

But even if we posit that Zitron is right and the villain Prabhakar Raghavan defeated the hero Ben Gomes, causing some kind of issue for Google search, this still doesn't make the case that Google revenue growth is in trouble at large because they have a number of other major products (such as YouTube and Google Cloud) that could drive growth even if search wasn't growing.

Every significant part of the chain of reasoning here is not only incorrect, it's not plausible if you know anything about Google or big companies in general. I'll be the first person to say that Google search quality has some serious problems and that Google has been increasing the relative priority of revenue over the user experience over time. This was a source of consternation for a number of user-focused engineers at Google when I was there in 2013.

For one of the issues Zitron cites, ads being confusing to users, in 2013, I asked a search engineer about Google changing the background color of ads to look more like search results because there was a previous study that showed that more an ad looked like a search result, the more users got confused over whether a result was an ad or a real search result, and I'd heard that Google deliberately made the ads not look like search results to avoid user confusion. The search engineer said that because some people didn't want users to get confused, it was impossible to make ads nearly identical to search results in a single change because it would be too obvious what's going on.

The way this was going to happen was that every time you A/B test tweaking ads to look a bit closer to search results, you make a lot more money, so the change would happen over multiple years in multiple parts, each small enough that the people who want to fight back against this kind of thing would have a hard time making a case. That happened just as this engineer predicted, but it was going to happen whether or not Raghavan ended up overseeing search. And, of course, that kind of thing happening doesn't cause Google to run out of room to grow and become desperate to reignite growth in a dying ecosystem. Whether or not you think Google should do it, it's something that makes Google more money.

How do people cite Zitron?

From what I can tell of how people cite Zitron, they cite him as an authority so they can say that this guy who looked at the numbers has made this claim, so their claim is backed up by the numbers. It turns out that if you look at the claims Zitron makes and know anything about the topic, the claims don't make sense, but I don't think that's the point. The point is one can say that someone looked at the numbers. The other point seems to be that this guy is angry2, which is a good way to drive engagement.

But when people bring him up, they're of course not generally citing his anger; they're saying here's this guy who's looked at the numbers and, if you're angry about AI, he's right there with you being angry about AI, and he's got numbers on his side.3 Like I said above, I don't want to go into this level of detail on each claim; this is just an illustrative example about how the claims below look. For any of his posts that I read, while there are numbers thrown around, the numbers don't actually connect to a coherent argument. In many cases, as we saw above, the numbers don't even really support his argument (such as an MAU decline in Facebook causing Meta financial problems which would then cause Meta to spuriously insert AI in places it doesn't belong). I suspect he's relying on people's eyes glazing over when they see numbers and just not thinking about what the numbers mean.

With the predictions below, someone could have the exact same prediction record and have completely reasonable reasons that just didn't pan out. Or someone could be correct in every case and also be wrong because all of their reasons are wrong. Someone like the latter person might have some kind of intuition that they're unable to articulate, or perhaps they're someone who just got lucky. Fortunately for us, we don't have to make this difficult judgement call because Zitron is wrong on the predictions and also wrong on the reasoning.

People with attention to detail on Zitron

Since I've been living under a rock for years and am just catching on the AI discourse, I hadn't actually read or watched anything by Zitron or any of the big AI commentators, but on looking up what people who have good judgement say, they also seem to find that Zitron's use of numbers is just sleight of hand, such as this comment by Juho Snellman:

His writing is certainly flamboyant, but the aggression and expletives seem more targeted at hyping up people who already believe the things he writes, not for making people change their minds. He found a niche in anti-tech grift, and is now exploiting the niche for all he can. But you might want to actually fact-check a few of the things he says that convince you, because at least for his written articles basically everything is made up or misrepresented. There's plenty of links to sources, sure, but if you follow them down to the primary source what they're saying is very different from what Zitron is implying

Here's an example where commenters seem to assume that Zitron's analysis is good for some reason, to which Juho Snellman replies: > The key problem is that his economic analysis is absolute trash. I used to think he was just totally incompetent at it, but given the bias in the errors, it is pretty clearly intentional deception. But it's often pretty hard to address that, because every article he writes is a 10k word gish gallop. I've tried debunking key points a few times in HN comments for just one of the intentional mistakes he makes, and people complain about the reply being too long.

For example, when Timothy B. Lee looked at a spreadsheet that Zitron used to create a projection of Anthropic's revenue, he found

He doesn't count February 1-10, counts March 1-10 twice, counts August 21-October 21 as one month instead of two, and doesn't count October 21-November 1. [another commenter notes that his spreadsheet also contains February 30] ... Ed claims he tried to compute Anthropic's revenue for 2025 and came up with $3.6 billion, suggesting some funny business [but the numbers work out once you fix the errors]

Some Zitron predictions

  • Feb 2024: "I believe we're reaching the upper limits about what generative AI can do and how accurate its outputs can be."
    • Wrong4
  • March 2024: "Have We Reached Peak AI?"; another prediction that hallucinations mean that AI progress is limited to then-current levels
    • Wrong
  • April 2024: "As I previously warned, artificial intelligence companies are running out of data ..."; another prediction that models can't improve because there's no more data
    • Wrong5
  • June 2024: OpenAI growth is stalling (with the implication it will continue to stall), which will lead to some kind of collapse of OpenAI
    • Wrong (it could be the case that OpenAI will collapse but, if so, it won't be due to any kind of growth stall from 2024)
  • July 2024: "Generative AI, as I said back in March, is peaking, if it hasn't already peaked. It cannot do much more than it is currently doing, other than doing more of it faster with some new inputs"
    • Wrong
  • July 2024: "Generative AI models aren’t getting more energy-efficient, nor are they getting more “powerful” in a way that would increase their functionality"
    • Wrong (models continued to get more powerful)6
  • August 2024: "generative AI is a dead-end technology that has peaked”
    • Wrong
  • August 2024: re-iteration that the AI bubble has 3 quarters to prove itself (from March 2024) or there will be a collapse
    • Wrong (Bartek Ogryczak notes, arguably Right because AI proved itself, but Zitron also argues no improvement, so Wrong by Zitron's accounting)7
  • September 2024: "o1 shows that OpenAI is both desperate and out of ideas", with a re-iteration of the idea that models can't improve due to lack of data
    • Wrong
  • Oct 2024: OpenAI's forecast of $3.7B revenue in 2024 and $11.6B in 2025 and $100B in 2029 are absurd, "a statement so egregious that I am surprised it's not some kind of financial crime to say it out loud"
    • Wrong (2025 goal exceeded, 2029 TBD but not an egregious financial crime level of implausible)
  • Oct 2024: "[OpenAI revenue] growth is already slowing, and will slow dramatically as we enter the new year"
    • Wrong (OpenAI exceeded the forecasts and contiued to grow quickly)
  • Dec 2024: "I also warned you in March that generative AI had already peaked.”
    • Wrong (also, bizarrely, implying no progress since March 2024)
  • Jan 2025: "I believe we’re at peak AI"
    • Wrong
  • Jan 2025: "DeepSeek has commoditized the [LLM]"
    • Wrong (OpenAI and Anthropic had and still have significant pricing power and can maintain prices well above DeepSeek)
  • February 2025: Anthropic making $34.5B in revenue 2027 is "is laughable on many levels, chief of which is that OpenAI, which made around twice as much revenue as Anthropic did in 2024, barely made a billion dollars from API calls in the same year."
    • Wrong (whether or not they make that in 2027, their 2026 ARR greatly exceeding that makes the 2027 estimate non-laughable)
  • February 2025: "Sundar Pichai wants Gemini to be 'used by 500 million people before the end of 2025, 'a number so unrealistic that someone at Google should have been fired, and that someone is Sundar Pichai."
    • Wrong (Gemini hit 750M users)
  • February 2025: "Sam Altman deputizing Orion from GPT-5 to GPT-4.5 suggests that OpenAI has hit a wall with making its next model, requiring him to lower expectations";
    • Possibly right? (GPT-4.5 wasn't very exciting, though if his "hit a wall" framing is part of his thesis that models can never improve, this would still be wrong as GPT-5 was a substantial improvement)
  • February 2025: "I will keep writing this stuff until I’m proven wrong."
    • Wrong (Zitron continues to write despite repeatedly being proven wrong)
  • March 2025: "In my years writing this newsletter I have come across few companies as rotten as CoreWeave ..." Zitron goes on to say that the company will not be able to survive for six months except with fundraising, though $4B raised might by them a year
    • Wrong (CoreWeave still exists and it's currently at more than double its IPO price as of this writing; CoreWeave only raised $1.5B at IPO)
  • April 2025: Zitron calls the bubble again and says "We're about to find out if I'm right."
    • Wrong (in that Zitron implied momentous events were about to happen which would prove him right and no such events happened)
  • April 2025: "It also, at this point, is pretty obvious that generative AI isn't going to do much more than it does today."
    • Wrong
  • May 2025: "I do not know how you come away from this story and not think Cohere is going to die. Their projections are so far off from reality."
    • Technically unfalsifiable because there's no end date, but implied claim is wrong
  • July 2025: "I am not trying to be dramatic, but it's pretty easy to come to the conclusion that Cursor is going to die"
    • Wrong (Cursor gets a $60B exit)
  • August 2025: "These models have clearly hit a wall where training is hitting diminishing returns"
    • Wrong
  • August 2025: Zitron says Cursor is dying and expects that it will sell for a firesale price; a price as high as $10B is not plausbie: "Is Cursor worth $10 billion? Nope! No matter how good its product may or may not be, it is not good enough to be sold at a price that doesn’t require Cursor to incinerate hundreds of millions of dollars with no end in sight."
    • Wrong
  • October 2025: In response to the question, “If you had to guess, what is the timeline we are looking at for the AI bubble to pop?”, Zitron answers, "No later than Q2 2026"
    • Wrong (note that Zitron threw in a "no later than", which makes the stronger claim that this is an upper bound and not just a best guess)
  • Nov 2025: "the fact we're running out of high quality training data and we're hitting the walls of scaling laws, in the training paradigm, these models aren't getting better. What we're seeing today is pretty much what they're always gonna be like"
    • Wrong

After this point, most further predictions that I saw were either non-falsifiable or resolve in the future. Note that I didn't attempt to catalogue statements that are nonsensical or were simply factually incorrect statements at the time, such as his December 2024 claim that “Generative AI's products have effectively been trapped in amber for over a year.” January 2026 claim that "[models are] basically the same as they were a year ago. They have the same efficacy". Zitron has not only made forward-looking statements that AI capabilities will not improve, he's also consistently made backwards-looking statements that capabilities have not improved which, while obviously false at the time, seem to play well to his base (along with his other false statements). If you connect all his statements together, it's implied that AI had the same capabilities in January 2026 as they did in December 2023 (and if you connect later statements, it's actually implied that capabilities in August 2026 are the same as in December 2023, though to be fair to Zitron he frequently contradicts himself and has also admitted to limited improvement in mid 2026).

To be fair, we could say that Zitron is speaking colloquially, so we when he says things like "have effectively been trapped in amber for over a year", that doesn't mean there's actually be no change December 2023, so the statements aren't transitive. Even if you assume a kind of colloquial sloppiness here, the collection of statements still implies that, from December 2023 to August 2026, improvements have been minimal (perhaps except, as noted above, when he contradicts himself and admits there have been limited improvements in some areas).

Comparing to respected Futurists

If we compare to how futurists did in our analysis of futurists, on style, Zitron relies much more heavily on anger than any of the futurists we looked at. On the quality of reasoning, he was probably about average compared to the futurists. Despite being wrong on roughly everything, he's not more unreasonable than someone like Buckminster Fuller, who suggested we'll be able to send people by radio because atoms have frequencies and radio waves have frequencies so it will be possible to pick up all of our frequencies and send them by radio.

In terms of the style of reasoning, of the futurists reviewed, he's probably closest to Kurzweil, in that he uses numbers to give a kind of aura of credibility, but if you know something about the topic he's discussing or look at the numbers, the reasoning falls apart. Zitron's reasoning isn't worse than Kurzweil's, who (for example) continually made new predictions of extremely fast progress that didn't pan out (such as, in 2001, predicting unbounded lifespans by 2011). Continually predicting that AI progress will stop for reasons that are incorrect is just taking the flip side of the bet on progress. Instead of having infinite progress, we're going to have no progress. Every time that prediction is proven wrong, you can just make another similar prediction and then move the date forward a bit (fans of both use the same techniques as well; fans of Zitron simply claim that his predictions are true, just like fans of Kurzweil cite his 86% prediction accuracy even though his actual accuracy on those predictions is 7% if you actually look at the results on the exact predictions he allegedly got 86% right). Michał Zalewski (lcamtuf) has some thoughts on why this happens:

The surest way to build [a] popular following is to articulate positions that are crisp, strong, and leave no room for doubt. You can't get too many podcast or TV appearances out of "well, the market could go either way", "both political parties make good points", "there's some merit but also some hype to AI". Or, to tap into the example in the post, "Harry Potter is an OK book".

In fact, there's a positive feedback loop. If you take a provocative, edgy stance, you get more attention and likes, so you sort of... self-radicalize? At some point, it's no longer an opinion that can be changed. It's an identity, a personal brand.

It's ... why Ed Zitron has a blockbuster blog about how it's all just one big scam. If you take a more nuanced view, you will at best get no reaction, or at worst, you'll invite scorn from both sides.8

For anoyone looking for well-reasoned anti-AI takes, I find whitequark to be quite good (not that I agree, but I think the reasoning is sound and I could see how someone would agree if they have slightly different premises than I do), but of course whitequark doesn't draw the kind of big audience that Zitron does.

How long can you maintain an incorrect position for?

I'm curious what people do after being on the wrong side of a set of failed predictions about progress like this. For the futurists, even the ones who were nearly completely wrong (which was every single one reviewed here), they can still make some kind of case like "a quarter of the things I said would happen happened, it just took two to twenty times longer than I expected" and if they're not so stuck on accuracy, they can round this up to "the things I said would happen happened", which is often what they've done. That seems to have served them well as nobody really cares to look at the details anyway, which is how, for example, Kurzweil's alleged 86% prediction accuracy became a well-established fact; no one bothered to actually check which of the cited predictions panned out until we looked at this in 2022.

But what happens to someone like Paul Ehrlich, who predicted imminent catastrophe when this clearly was not happening as he was writing and then did not happen? Just looking at Ehrlich's Wikipedia page, we have

A common criticism is that Ehrlich's predictions routinely failed to come true; for instance, Ronald Bailey of Reason magazine has termed him an "irrepressible doomster ... who, as far as I can tell, has never been right in any of his forecasts of imminent catastrophe."[41] On the first Earth Day in 1970, he warned that "[i]n ten years all important animal life in the sea will be extinct. Large areas of coastline will have to be evacuated because of the stench of dead fish."[41][42]

In a 1971 speech, he predicted that: "By the year 2000 the United Kingdom will be simply a small group of impoverished islands, inhabited by some 70 million hungry people." "If I were a gambler," Professor Ehrlich concluded before boarding an airplane, "I would take even money that England will not exist in the year 2000."[41][42]

When this scenario did not occur, he responded that "When you predict the future, you get things wrong. How wrong is another question. I would have lost if I had had taken the bet. However, if you look closely at England, what can I tell you? They're having all kinds of problems, just like everybody else."[41]

Ehrlich wrote in The Population Bomb that, "India couldn't possibly feed two hundred million more people by 1980."[27] In 1967, Ehrlich called to cut off emergency food aid to India as "hopeless".[43] This position was later criticized, as India's food production subsequently skyrocketed through the Green Revolution in India, and its per capita caloric intake rose significantly in the following decades, even as its population doubled.[44]

A large increase in global food production since the 1960s and a slowing of population growth have, within the current context of continued depletion of non-renewable resources, averted the scale of food shortage, famine and catastrophe foretold by the Ehrlichs.

Canadian journalist Dan Gardner, in his 2010 book Future Babble,[45] argues that Ehrlich has been insufficiently forthright in acknowledging errors he made, while being intellectually dishonest or evasive in taking credit for things he claims he got "right". For example, he rarely acknowledges the mistakes he made in predicting material shortages, massive death tolls from starvation (as many as one billion in the publication Age of Affluence) or regarding the disastrous effects on specific countries. Meanwhile, he is happy to claim credit for "predicting" the increase of AIDS or global warming.[13]

In the case of disease, Ehrlich had predicted the increase of a disease based on overcrowding, or the weakened immune systems of starving people, so it is "a stretch to see this as forecasting the emergence of AIDS in the 1980s." Similarly, global warming was one of the scenarios that Ehrlich described, so claiming credit for it, while disavowing responsibility for failed scenarios is a double standard. Gardner believes that Ehrlich is displaying classical signs of cognitive dissonance, and that his failure to acknowledge obvious errors of his own judgement render his current thinking suspect.[13]

Barry Commoner has criticized Ehrlich's 1970 statement that "When you reach a point where you realize further efforts will be futile, you may as well look after yourself and your friends and enjoy what little time you have left. That point for me is 1972."[46] Gardner has criticized Ehrlich for endorsing the strategies proposed by William and Paul Paddock in their book Famine 1975!. They had proposed a system of "triage" that would end food aid to "hopeless" countries such as India and Egypt. In Population Bomb, Ehrlich suggests that "there is no rational choice except to adopt some form of the Paddocks' strategy as far as food distribution is concerned." Had this strategy been implemented for countries such as India and Egypt, which were reliant on food aid at that time, they would almost certainly have suffered famines.[13] Instead, both Egypt and India have greatly increased their food production and now feed much larger populations without reliance on food aid

Amazingly, following the series of incorrect predictions Ehrlich made in and after writing The Population Bomb in 1968, he followed this up with The Population Explosion in 1990 and has continued saying that we have global overpopulation that is causing or will cause a dire crisis unless we cut worldwide population. He has said the same thing this century and even this decade. It appears the only reason he's not saying that today is that he died earlier this year.

If I didn't look it up, I would've guessed that his recent position would be something like "well, I got some things wrong, but it was only due to these actions that were inspired by my work that crisis was averted" or "while crisis was averted, it was a lucky roll of the dice and, in most universes, the agricultural advancements that staved off the mass starvation deaths I was predicting don't happen", not "just you wait, the crisis is happening now and I'm about to be proven right"; in 2015, referring to his incorrect 1968 book, he said "[m]y language would be even more apocalyptic today". That's the pattern we've seen from Zitron, but I wouldn't have guessed that the one person I looked up would've kept that up for 50 more years. Maybe we'll get 50 more years of Zitron predicting the end of AI progress.

Some reactions to Zitron

In one of the quotes from Juho Snellman, above, Snellman says that he writes a large amount of gish gallop, which is a term for when someone floods you with so much cheap (as in cheap to produce) nonsense that no one would want to take the time to bother to refute it. In discussing one small part of Zitron's talk in detail, we spent more than 1000 words explaining why Zitron has an incorrect understanding of how corporations work and how Zitron got the reasoning wrong. Someone can read that and then say, "but you didn't address X" in the talk, which is true. When I first watched the talk, I actually closed the tab after 90 seconds because there was so much nonsense that it didn't seem worth the time to go any further. I could write 5k words on the first 90 seconds of the video. Because Zitron is just saying a bunch of nonsense, he can do that very cheaply and it would take 30-60 minutes to refute 90 seconds of his nonsense if I had all the facts at hand. With time to look up the exact right information, it probably would take double or triple the amount of time. When someone who has good judgement sees something like this, they tend to immediately write the person off. Just for example, I mentioned to a friend of mine that I'm writing this post and they said

I was listening to this podcast with the guy and I couldn't get through it. My heart rate was going up because he would just say this false thing and then the interviewer, who was reasonable, would ask about it, "what about X?", and then we would just jump to another falsehood ...

... before I ducked out, he talks about how LLMs haven't gotten a lot better over the past year, and the interviewer says people use them and they've definitely gotten a lot better in the past year, and Zitron denies it and says 'have they?', and the interviewer is just like, "yes..." At that point, I'm just like, why am I listening to this conversation?

We mostly discussed predictions and not incorrect statements about the past or present, but everything I've read or watched by Zitron is also full of things like this. Many people will look at something like this and decide the guy is a crank and stop paying attention. But many other people will look at something like this, see someone refute a set of things, and then say, "but you didn't refute X" and, in general, the person doing the refuting may respond to a couple of these, but they eventually give up because the gish gallop method has the same properties as an amplification DoS attack. It's very cheap to generate new nonsense, but it takes some effort to refute it.

BTW, I was curious what this interview was, so I put the above quote into ChatGPT and asked it to find the interview. It was able to identify an interview with the relevant exchange (it actually identified multiple, as this appears to be a common question and response pattern by Zitron) and the timestamp of each relevant statement in the interview (the start of the general argument is here and a "have they" response is here. Prior to the "have they?" comment, the interviewer tries to establish a baseline that agents have improved in capability. Zitron denies that this has happened, and then when the interviewer notes that people who use these things for their jobs Zitron denies this with the "have they?" comment (he actually makes multiple contradictory statements in the sequence).

Another thing to note here is Zitron's extremely high level of stated confidence. Some that we noted were OpenAI's forecast that is "a statement so egregious that I am surprised it's not some kind of financial crime to say it out loud" (which they've achieved so far) and his claim that Google's forecast for Gemini users is "a number so unrealistic that someone at Google should have been fired, and that someone is Sundar Pichai" (they managed to exceed the forecast by 50% when Zitron's claim was that it would be completely absurd for them to reach the number at all).

I've made quite a few predictions, and quite a few of those predictions are wrong. When I'm really making a prediction, I attach a confidence level to the prediction just for my own sake, so I can look back at these things and see how well calibrated the predictions are. I have never been wrong about a prediction that has anywhere near the confidence Zitron gives to some of his predictions. Given the stated level of confidence, even a single incorrect prediction would be a sign of an extremely high degree of overconfidence. One should effectively never be wrong about a prediction delivered with that level of confidence but Zitron is routinely wrong about predictions he makes with what is rhetorically pretty much the highest possible degree of confidence.

BTW, a funny thing about Gemini hitting 500M users being "so unrealistic that someone at Google should have been fired, and that someone is Sundar Pichai" is that Zitron has also (incorrectly) said that Google doesn't know how to grow, and that as a result they're shoving AI everywhere. Dennis Snell pointed out that, if Zitron takes his own statement seriously, Google can make Gemini's user numbers go to any number it wants by doing the exact thing Zitron said they would do, sticking AI everywhere.

You can't actually take Zitron's statement about Google's lack of growth leading to AI desperation seriously and also take it seriously when he says that Sundar is committing some kind of gross malpractice by naming a number like 500M users. This is another thing that is immediately obvious on watching one of his talks or reading his writing. There are a bunch of disconnected statements that don't fit together, except insofar as they're statements about how AI companies and people and companies that are using AI are evil and bad. The actual numbers and logic of the statements are contradictory. It seems to be whatever comes to mind that can be used to paint the villains as evil. And, funnily enough, the 750M user number Gemini hit shows that both of Zitron's statements were incorrect. If Google were as desperate to juice the numbers as Zitron claimed, they could've easily gotten the number above 1B by sticking Gemini everywhere, and of course 750M > 500M.

BTW, the point at which I stopped the talk for the first time was

a market obsessed with year-over-year revenue growth. And this progression was natural. It was horrible. You can blame Marc Andreessen. He's a horrible man. You can blame many horrible men. There are so many guys to be mad at the moment.

That last sentence really sums up Zitron's position. "There are so many guys to be mad at the moment". In this talk, he throws in this jab at Andreesen and blames Andreesen for Meta, Google, and Microsoft pursuing growth. In reality, if Marc Andreesen had never existed, Meta, Google, and Microsoft would almost certainly still be trying to grow so we of course cannot actually blame Andreesen for these companies trying to grow. There's just this thing that he says is bad, and in his usual style, he pulls some person and says they're the evil villain that's to blame for this, and then moves on to the next non sequitur.

How can people take this seriously?

Because I'm a masochist, I actually went and read a bunch of Zitron discussions (I believe I read every major discussion on HN and lobsters, and a bunch of other ones as well) to see what people who take Zitron seriously are saying. One common defense was the one above, sure, you refuted some points, but you didn't cover X. A more common defense is to say, just in general, people attack Zitron because of Y (usually his style), but they never address his points, "which tells me everything I need to know" (or something along those same lines). Based on the timestamps of the messages, just scoping to the stories that were being discussed, there were generally already comments discussing Zitron's actual errors, but Zitron's defenders would ignore this and just claim that people were unable to point to mistakes Zitron had made. This is a very Zitronian move and it makes sense that people who like his style would also use this move. After all, who would find Zitron convincing? Someone who thinks this kind of thing is valid reasoning.

The next most common "move" was to simply deny that Zitron said something that was refuted. When people would mention that Zitron was repeatedly on the record in 2024 and 2025 as having said LLMs couldn't improve further for fundamental reasons, Zitron's defenders would say that he never said that, and likewise for previous predictions or factually incorrect statements.

Another class of defense I saw were comments like "but what about all the AI hypists who are wrong?". Like I said before, I wrote a 34k word post about how a bunch of the most respected futurists have been wrong, not just because they made incorrect predictions, but their methods and reasoning were wrong. But a bunch of people who hype the future being wrong doesn't make people like Ed Zitron or Paul Ehrlich any less wrong. Zitron and Ehrlich are still exactly as wrong as they would be if those futurists never existed.

A friend of mine also noted this about comments on cases where people point out that Zitron was wrong about models not improving from 2023 to 2026 (and yes, this is specifically on stories or comments that discuss Zitron's disproven statements on capabilities not improving):

It's incredible to see so many people saying, "Zitron isn't wrong, he's just early!" I guess the implication is that we'll eventually realize that the models we have in 2026 are actually no better than the ones we had in 2024 or ??

An interesting thing about publshing this post is that a decent fraction of the people who've message me to tell me that I'm wrong say that I'm wrong because, today in 2026, models haven't actually models haven't gotten better since 2023 or 2024. My guess would be that most people who are saying things like the quote above are just doing the "move" where you don't read what was actually said and respond with a canned response that's nonsensical to anyone who's actually read what they're replying to, but it turns out there are plenty of people who actually believe Zitron's string of statements that imply models haven't improved since 2023 or 2024.

If someone's actually looked at what's happening, I don't think there's anything you can really do to convince someone who's denying reality at that level, but for anyone who's just hasn't seen how things have changed, for a visual example of improvements over that time period, here's a comparison of 2023 and 2025 video generation and here's an example from August 2026. Video isn't a great example since models have improved a lot more at coding, e.g., with on the order of minutes of human time, it's possible to create a new regex engine with an interpreter and an native code compiler and then fork ripgrep to make it faster for codex's actual ripgrep calls on my machine, a project that would probably cost 7 figures pre-LLM if you price out how much people with the expertise for that are paid. But video is a nice example because, in the interview linked above, after the "have they?" exchange, at one point Zitron's "rebuttal" is, "you wouldn't make movie with it would you?". People are "shooting" quite a bit of AI-generated digital footage now and this is upending a lot of lower end video work. An easy prediction based on historical patterns is that this will continue to move upmarket over time, but just based on what people are using AI video for today, Zitron should probably find a new rebuttal even if he's just playing to true believers who don't think models have improved since 2023 or 2024.

Future predictions

Although Zitron's past predictions have generally been wrong, maybe he'll be right about something in the future. Perhaps some of these companies will have valuations decline for some reason. But, even if there's some kind of massive AI crash and OpenAI and Anthropic go to zero, in terms of the societal impact, if on top of that, some other event occurs that prevents further progress in models beyond whatever AI labs have internally right now, that's still going to result in a fair amount of change. Which companies are successful will change who gets rich, but particular companies failing won't stop changes that fall out of current or next generation model capabilities from happening; it just moves around who benefits the most.

Personally, it doesn't matter to me if folks at one company vs. another get rich. If one company does something better (in some abstract sense) than another, that's of some interest to me, but I have some skepticism about any particular company's claims that they'll do more of "the right thing" than another company (I could be convinced on this one, but I don't find the public claims that I know of very convincing).

If Zitron ends up being right about some company or other collapsing, that's pretty uninteresting to me compared to how capabilities have developed and will develop, where he's been wrong to date. It also happens that he's been wrong about the financial predictions he's made to date, but that doesn't really interest me, though I included a number of financial predictions for completeness.

Thanks to Yossi Kreinin, Juho Snellman, Dennis Snell, Nick Bergson-Shilcock, @blueblimpms, Bartek Ogryczak, Jamie Brandon, and Shriram Krishnamurthi for comments/corrections/discussion.

Appendix: Ed Zitron on why people don't like Ed Zitron

While looking for discussions about Zitron's work, the #2 hit on reddit was this comment by Zitron:

... some men don't like me because emotional honesty and introspection are difficult for them. Feelings are something that men are told to repress or compress. I refuse, and I find it disgusting when anyone tells me to do so ...

... Let's start with emotions, because it's the most obvious one. People really do not like that I am how I am, and think that I am "getting mad as a bit," or even go as far as to describe me as psychotic, out-of-control, and so on and so forth. This is a common reaction, I find, from anyone who themselves is emotionally repressed, especially in their own work. It is hard to be emotional and have well-done opinions ...

... I also have not taken the route you are "meant to take" to get here. You are "meant" to be an establishment writer from a big outlet, or an analyst, or in finance, or any number of other different "true paths" where you are "worthy" of whatever it is you're meant to get. I did not "earn my stripes" in the traditional sense, and those that have believe I did not earn my way here ...

... My work is also thorough, which is frustrating for people that do not do thorough work. I have thought through every point I have, and I take great pains to know subjects well. Notice how many people still claim "it's just like Uber" or "it's just like the dot com boom." It's much easier to just assume shit without ever checking if it's true! Having some asshole who comes along with thoroughly and with passion is frustrating. It reflects badly on your work ...

... I do a good photo shoot, I do a good interview, and I capitalize on events, and I do so without being craven, because I usually show up with a few thousand words of thoughts or an episode about a thing. I believe there are some that would like this level of attention or prestige, but they do not want to do the work to get it, and that chafes ...

... I love big, I love hard, I am who I am, I have never been made to feel welcome by any "in" group. I work my ass off, I write more than anybody else, I show up. With whatever space I create I will fight back against "in groups" or cliques. I hate them, and they hate me right back. And I fundamentally know why I believe what I believe. That upsets people who do not.

I have no idea if he means any of that or not (if this Wired profile about Zitron and the PR firm he runs is accurate, one would have to lean towards not), but Zitron seems to be very good at saying what his audience wants to hear, so this proably gives some kind of insight into his audience.

One thing to note about the bit about cliques and "in groups", if you just search his name on reddit commenters note that if you post anything indicating that AI has improved on his subreddit (such as link to benchmarks), you get banned for it, resulting in a highly clique-y echo chamber. I'm on the record as having said that METR's progress benchmark isn't meaningful and that you're better off going on vibes than leaning on a misleading analysis and that widely cited AI evals are frequently flawed, so it's not like I think that benchmarks are generally good, but the picture I got from reading comments was that you get banned pretty quickly if you don't hew to the party line, which is the opposite of the picture painted above. This isn't anything unique to Zitron; when looking up another influencer a while back, if you disagreed with that influencer on their reddit, they would write a comment thanking you for your comment and saying how much they loved getting feedback from people and how the world is some kind of great peace and love fest and we should all love each other while simultaneously banning you from their reddit.

I also found Zitron's comments on how people don't like his work because they dislike thorough work to be interesting for a couple reasons.

One is that my own work is frequently positively cited as being rigorous and thorough. There are plenty of people who dislike my work as well, but not only do I not know of anyone who's said they dislike it because it's thorough, I would be surprised if there was anyone who secretly dislikes it because it's thorough. In general, just doesn't seem like a reason that people dislike things. That also goes for people being upset because someone knows why they believe something or because someone else worked hard. It's really interestin to me that this appears to be what Zitron's audience wants to hear.

The second thing is that, I wouldn't personally consider my work to be thorough. The same thing I mentioned here about not feeling that my work is good also applies to not feeling my work is thorough. I do some amount of checking of my work. I don't know that I'd say that it's more than most in terms of time spent, but in terms of effectiveness, I suspect the combination of methods and time spent works better than average. But I always have a dissatisfaction with my work when I published it because I could keep checking more thoroughly forever and never publish anything, so I force myself to publish at a level that I suspect is above average on thoroughness, but well short of thorough. If I compare my work to the work of someone I consider thorough, like Gary Bernhardt, I don't know how I could call my work thorough. I have a few friends who produce Bernhardt-quality work and I make the choice to produce much more but also lower quality work. I think this is a fine place to sit in the quality-speed tradeoff space, but that doesn't make my work thorough. To be as thorough as Gary, with my baseline pre-July 2026 standard, I'd need to put 10x-100x the time in per piece of output (it would take an additional 10x or more with how I've been publishing lately). And yet, it would seem that my fact checking process is a lot more thorough than Zitron's.

Even if we put aside the gross arithmetic errors like the Timothy Lee example, if we look at cases like the Facebook MAU example, where he picks a number that's directionally opposite of other estimates and of Meta's own numbers that's also directionally implausible given the other data out there, I don't see how such a figure could survive any fact checking at all. And this goes for a huge number of his factual statements (I would guess most, although I haven't tried to randomly sample them to be sure). It seems like any kind of fact checking process that you could imagine would turn up contradictory results.

Appendix: why write this?

No good reason, really. I got four hours of sleep and my brain wasn't good for much of anything and I saw someone posted a screenshot of a reddit post dunking on Ed Zitron's prediction record. When I wrote this review of futurist prediction accuracy, I tried to make sure that I didn't bias what I was reviewing in any way. It's not obvious from the post if the redditor who reviewed Zitron's predictions was pulling predictions in an unbiased fashion or if they were biased in some way (since AI has become a culture war issue, it wouldn't be surprising if someone pulled biased predictions), so I decided to read some Zitron in my spare time while poking at agents to get them to do an unrelated task I wanted them to do. For the futurist post, I read multiple entire books to pull predictions and generally only stopped when someone was being repetitive and kept saying the same thing over and over again. In this case, all Zitron does is be repetitive, so the methodology in the futurist review would mean that I review a few predictions and then stop immediately. To overcome this, I had ChatGPT give me a list of predictions (with no attempted tilt towards correct or incorrect predictions) and then I skimmed/read the posts that ChatGPT linked to. There were some cases where I thought ChatGPT's reading of the post was incorrect (these were generally cases where it flagged a prediction that would be incorrect if its reading was correct, but I disagreed with its reading) and (discussed further below) I also removed predictions which weren't falsifiable or seemed pointless because they were tautological (I noted something similar to this in the futurist post).

If I really thought about it, I probably could've found something better to do with the time, but here we are; I sometimes have tasks on my todo list for when I'm too tired to do real work, but I didn't have one. I don't think they cherry picked particularly bad predictions, although they did pick some that are among the more absurd sounding. However, if you go and look into the details of ones that aren't such ironclad "dunks" (like saying that Gemini hitting 500M by EOY users is so absurd Sundar should be fired for the idea, when Gemini actually hit 750M by EOY), these are just as wrong as claims that Cursor has no realistic buyer with the implication they won't even sell for $10B when "everybody" (who cares about AI exits) knows they sold for $60B.

The redditor picked the high-profile failed predictions, but Zitron's prediction corpus has many more failures and, as noted above, the bigger issue is his reasoning.

Another thing about the reddit comment is, whether or not the comment is unbiased, one might have the suspicion of a kind of bias because it was posted to r/accelerate by someone who apparently is an r/accelerate believer. On looking at the actual predictions they are consistent with some bias (they would also be consistent with an honest mistake as there's no way to distinguish these from the record). For example, one of the "refutations" is a statement by Zitron that OpenAI will collapse in 12-24 months. OpenAI didn't collapse, so this would appear on the surface to be a great way to show that Zitron was wrong, but if you read Zitron's post, Zitron's actual claim was that OpenAI will either collapse or raise a lot more money and they raised a lot more money. I disagree with Zitron's implications that this is inevitable just leading to a later collapse but his stated prediction was not falsified.

This prediction wasn't in the set of predictions scored in this post. Some would argue that this should be scored in the post. The reason this wasn't scored is because the prediction seems meaningless except insofar as it contributes to Zitron's broader point (that OpenAI is doomed and must collapse).

If we think about predictions one could make, a tautological prediction (if you write out all the edge cases I'll elide for space reasons) that has to be true is OpenAI has enough money to operate or it doesn't, and if it doesn't, it must raise the money somehow. I could make a million such tautological predictions, but if one were scoring my prediction record, it wouldn't make sense to include these because they're meaningless. In general, a company that's alive will cover its costs. If it does not, it will try to raise money. If it fails to do that, it will shut down or get acquired. A prediction that a company will either cover its costs or it will not cover its costs says nothing.

OpenAI's own projections were that it would not yet be profitable and its costs would exceed its revenue. That seemed nearly certain, so if you assume that this nearly certain thing is true, then you have the nearly tautological prediction that OpenAI will either collapse or it will raise money to cover its costs. It would have been reasonable to make a prediction like this at very high confidence (99.9% or above). If you use any kind of prediction scoring methodology, such as Brier score, these predictions contribute essentially nothing except when they're wrong as long as Zitron has a significant number of high-confidence incorrect predictions.

And, as we noted above, Zitron is repeatedly incorrect on predictions he gives the highest possible confidence (given his wording, I would rate a number of these at 6 9s or above), so on any kind of scoring mechanism like Brier score, Zitron's record is very poor. And a summary metric like this really understates how meaningless predictions like this are. Hypothetically, let's say Zitron made an unbounded number of correct 99.99% certainty near tautological predictions, which would make the score from the bounded number of other predictions he made meaningless on something like Brier score. This would still give you zero confidence for any of his non-near tautological predictions, and those are the predictions people generally talk about (AI progress is done, AI companies must collapse and this will bring down major tech companies as well, etc.).

Back the topic of the reddit commenter's potential bias vs. mine, as noted above, I don't have a particular bias towards a view that rapid progress is inevitible and have called out cases where people are overly optimistic, as evidenced by this post on futurist predictions. I'm also not someome who needs to or has any desire to farm engagement by manufacturing reasons that someone is wrong or bad and don't consistently rate every predictor as bad, as evidenced by this review of Steve Yegge's prediction record, in which I note that he scored well and also actually performed much better than the raw score indicated because the predictions are generally well reasoned and directionally correct even if the precise prediction was incorrect. I think it's actually awesome if someone has good insight in the future and shares it publicly, so I'm happy to call these cases out when I noticed them. It's just that, in this case, Zitron is a kind of anti-Yegge: someone with a poor prediction record whose predictions are actually worse than they seem from the record alone.

Appendix: errors in this post

I think it's almost certain that this post has multiple errors. In general, I find it very difficult to read a long stream of incorrect reasoning and then not get sloppy when looking for errors in it. I had this exact same problem when reviewing futurist predictions. It reminds me of when you're programming for some system where the compiler is very buggy and you hit compiler bugs all day every day (not uncommon when working with embedded systems, at least pre-LLM; now you can fix the bugs relatively easily). I find it hard not to get sloppy and think "hmm, this might be a compiler bug" even though, every once in a while, it will actually be your bug and not a compiler bug. The problem is much worse when looking at predictions from these kinds of predictions since the compiler still generally basically works and is often right, whereas when reading text like discussed here, you're just constantly drowning in nonsense that is occasionally punctuated by a good and accurate point.

I think, to do this well, you'd either need to find someone with very unusually high endurance for trudging through this stuff (I mean, much more than me, and I seem to have a somewhat above average endurance for this kind of thing) or have a team of people who independently rate and score things, but who would want to spend that kind of effort when any surface-level reading immediately reveals many things that indicate that these folks are pretty much totally wrong?

I did ask ChatGPT (web interface, Pro) and Claude (web interface, Fable 5) to fact check this post. They both found some minor errors that were fixed before publication.

One year ago, I found fact checks like this nearly useless, but they're halfway decent now and, contra Zitron, I would expect them to continue to get better. For people who are curious about the two, ChatGPT was much more thorough than Claude in this case and found more errors as well as finding every error that Claude found. However, it was overzealous and cited a number of non-errors, such as suggesting that tongue-in-cheek comments were incorrect, and that a number of statements that were generally true should be re-phrased in some more literal way (complete with AI-styled text).

[Edit: @blueblimpms pointed out that a prediction that I thought was about GPT-5 was probably actually about GPT-4.5, although what Zitron is saying is unclear. After re-reading the relevant post, I agree, both that Zitron is probably referring to 4.5 and not 5 and also that his statement is unclear, so I changed that. That correction fits into this pattern that I predicted would occur, though I didn't note that a secondary cause of this problem is that Zitron's writing is quite imprecise and often relies on various vague implications between statements. The "have they?" / "are they?" response he does in interviews would be an example of this, where one could techincally argue that he's not making a statement at all and is just asking a question, although in those cases, given his overall position, we can infer what he means when he says that.]


  1. But, even if it were the case that the accusations against Raghavan are true (I'm not sure how they could be, as how could one be a class traitor to computer scientists in the first place, but let's posit that, whatever it means, it's true), Zitron's contention is that "this shithead [points to an image of Raghavan] took over Google search in 2020" and then prioritized certain metrics over search quality. I'm not sure why one would name a particular person for this as this is something that was a long-standing fight with many people involved on all sides but, if we posit that this is all true, then we posit that the "management consultancy sect" will move metrics that will cause engagement and/or revenue to increase at the cost of search quality. This would have the opposite of the effect Zitron needs here to make his case that Google growth is done and they're so desperate for growth they have to put AI everywhere in some kind of crazed last-ditch attempt to save Google. Perhaps one could make the argument that this will eventually cause Google search to decline, but Zitron's argument was that, in 2024, they were desperate, not that users will eventually leave Google search, which will later cause a decline.

    Anyone who's read a lot of Zitron will recognize a standard "move" of his, turning the situation into some kind of hero-villain narrative (for search, the alleged hero is Ben Gomes and the villain is Prabhakar Raghavan); it's as if his mental model of how companies works comes from movies about companies. If you ever watch a movie that's allegedly about some events and then read about it, you'll find that things get oversimplified into a hero-villain narrative and that almost all of the nuance is stripped out of the situation. And then if you're ever personally involved in something or talk to people who are personally involved and compare what happened to the books that get written about it, the same thing happens again; in general, the major causal factors are not identified in books about what happened in tech and many of the most instrumental people involved in some of the key decisions aren't even named because journalists talking to people about what happened aren't really able to piece together a plausibly correct story about what happened to someone who understands the underlying mechanics and has good information. Anyway, without knowing anything about the situation, if someone tells you a hero-villain narrative of the kind Zitron likes to spin, you can already be a bit skeptical.

    [return]
  2. BTW, I don't think his anger really comes across in the video. I mean, he explicitly says he's angry and he swears and insults people, just like in his writing, but he doesn't really read as angry to me. It reminds me of this test on emotion recognition I took with a bunch of folks recently.

    I found the test fairly difficult and spent maybe 5 minutes on the first question because the person had a huge fake smile on their face and also looked a bit uncomfortable and anxious. I couldn't tell if you were supposed to say that the person is happy or uncomfortable/anxious. Is it supposed to be a very easy test or is it supposed to be a test that has a bit of subtlety? Based on what the test looked like, after thinking about it for a while, I chose "happy". Luckily, the test actually tells you if you got the question right or not, so I realized the test was about the fake exaggerated expression being made and not the person's actual expression and most the rest of the questions were easy. One was difficult because they were faking one particular emotion with what is a textbook display, as in, the kind of thing one sees in a textbook, but in a very specific way that was less complete and more unrealistic than the other textbook displays; it was as if someone had read a description of what a contemptuous sneer is, and then was trying to make the facial expression based on the textual description. I had to think about that one for a couple minutes to get the correct answer.

    Anyway, to me, Zitron seems like someone who's playacting anger and not someone who's actually angry. The tone of voice, facial expression, body language, style of movement, etc., just don't seem angry to me. I think this anger positioning works better in his writing than in his speeches because the cues he uses (swearing, saying he's angry, showing a lot of contempt, insulting people, etc.) are about as good as it gets for anger cues in writing. When you have audio and video, these are fairly weak cues; if the stronger cues don't really indicate anger, the person just doesn't seem angry. It's possible he has a non-standard way of showing anger or I just wasn't paying enough attention, but after watching some videos of him where he talks like he writes but didn't seem angry, the writing just doesn't feel angry to me anymore.

    [return]
  3. If you want to see an example of what it looks like when someone tries to discuss the numbers, here's a thread where Juho Snellman pushes back on someone who insists that people have done the math. As I've been catching up AI discussions, I've seen many discussions like this where one side has someone who's actually looked at the numbers and the other side waves around some kind of vague insistence that numbers have been looked at. This never really goes anywhere because, for one of the sides, the point isn't that you can understand something from the numbers, it's that they have a piece of evidence they can wield because someone has looked at the numbers. [return]
  4. Zitron's argument at the time was that hallucinations were as good as they were going to get, which meant that AI performance is capped at 2024 levels. Both the overall prediction and the mechanism were wrong. This one seemed wrong at the time, in that I noted here in 2024 that you can make AI code halfway decently by just putting it in a loop and having it run until the code compiles and tests pass; I wasn't a heavy AI user at the time, but anyone who was using AI could see that there were ways that you could mitigate the hallucination rate which weren't being widely applied (this was before coding agents like codex and claude executed code and would check that tests pass, etc.) [return]
  5. This is another one that also seemed untrue at the time. I'm not an ML person, but the moment someone told me what an RL environment was, within minutes, I thought of a bunch of ways one could generate synthetic data for improved training. I'm sure none of these were novel and they're things that AI labs are doing; my point is just that anyone who thinks about it for a few minutes can come up with a lot of ways that models could be improved even if there were no new data to find on the internet (not to mention that more effort could be used to get data that isn't just reddit comments or whatever the easiest to scrape content on the internet is). [return]
  6. Note that this only scores predictions that have resolved. In that particular post Zitron also states that progress towards AGI will never happen, which is still both fuzzy and difficult to adjudicate and also one that you can never really reliably resolve as positive. Similarly, a prediction in a previous post that some company would have to add subscriptions isn't listed because it's open ended and not really resolvable as a negative (it was implied to have to happen soon, so is arguably wrong, but if one wanted to weasel out of it one could say that it will happen in the future). [return]
  7. Here, Zitron also said, "I’ve realized now that it isn’t super useful to attach things to time (though I stand by my prediction) and thus I think it’s more useful to suggest what the terms of the bubble popping actually are". After this point, Zitron makes relatively fewer dated statements after this point and makes many more open-ended unfalsifiable statements. Perhaps a reaction to being wrong so frequently with his previous predictions? [return]
  8. In a small piece of optimism, I'll say that this blog seems to have done ok despite not leaning into extremist positions and generally trying to avoid clickbait. This often means that, when I look at some data, I'll see something that looks like it would make for a really interesting viral hit piece, but then on looking more closely, it's actually a boring negative result, like when I ran this quick and dirty programming language eval, which originally appeared to show a very interesting result, which went away once I fixed the obvious eval bugs. Oh well. I'd like it if people published more boring negative results, so I published the boring negative result.

    I wouldn't be surprised if this blog is within an order of magnitude of traffic as Zitron's substack (server-side stats show 540k uniques for me in the past month, but who knows how many of those are bots with some but minimal Cloudflare bot blocking) despite Zitron writing much more frequently than me and pulling out every clickbait trick in the book, while I just occasionally post something when I feel like writing something up. Although my goal obviously isn't to get traffic, if we adjust for the level of time or effort, I don't think this blog does terribly compared to Zitron. Could Zitron have 5.4M monthly uniques? It's not impossible and it's hard to tell what these numbers mean with bot traffic, but for reference, The Economist has about 1.3M subs and the NYT has about 13M digital subs. If we hypothesize that 3/4 of uniques will be bot traffic, having an order of magnitude more traffic than this blog would put Zitron into the same class as The Economist, which doesn't really seem plausible.

    Ceteris paribus, I think Zalewski is right on the incentives, and I've seen a lot of people become caricatures of themselves as they lean into what drives the most engagement, but I think doing the opposite can work ok.

    For example, with a style that could be described as the opposite of clickbait, Simon Willison has written what I suspect is the most widely read blog among programmers for the past 3-4 years (in the same way that, at various times in the past, Joel Spolsky or Jeff Atwood or Steve Yegge seemed to be the most widely read programmer among programmers). Among programmers and other serious users of AI, I would guess that Willison has a larger audience than Zitron.

    However, it's true that Zitron has a kind of audience that Willison can never really get with his style. In the body of this post, we looked at common defenses of Zitron on forums where people use AI. That was pulled from forums where people use AI. If we look at the world at large, the comments look fairly different. For example, on the video that my friend mentioned, where Zitron repeatedly denies reality and the interviewer pushes back, the top comments at the moment are all in support of Zitron and they also just deny reality and claim that the places where the interviewer pushes back with a piece of reality are the interviewer being biased or just not knowing what he's talking about. Among the top comments, there seems to be little to no engagement with the facts of the matter; it's all mood affiliation. The comments remind me of what supporters say about politicians who use the gish gallop strategy and just say a bunch of outrageous nonsense. I could imagine Zitron running for office one day on the strength of his reality-denying popularity or becoming a demagogue who's a right-hand-man of someone in office, so Zalewski is right in that Zitron's appeal is not one someone is going to get by accurately describing what's happening in AI.

    But, while I don't know Willison and this could be totally wrong, my impression is that, like me, he's doing something he wants to do anyway and the audience just sort of happened despite him not trying to maximize his audience. When I say it works ok, I mean that he seems to be able to support himself working as a full-time open source developer due to the sponsorships he's gotten (which I would presume are generally because he has such a large audience), which seems like a good outcome even if this doesn't create the kind of mass appeal someone like Zitron can generate.

    [return]

Consider switching your major [www.collegiatetimes.com - RSS Results for * of type article OR video OR youtube OR collection] (03:00 , Monday, 31 August 2026)

How might someone figure out what they want to do for the rest of their life? How might someone figure everything out at 18? At 18, I felt so sure I was going to be a biology student on the…

Think twice before installing this device promising free movies [Biz & IT - Ars Technica] (12:33 , Monday, 31 August 2026)

As online services get better at blocking malicious traffic, the attackers and scammers behind them have been forced to find new ways to reach their targets. The alternative of choice is now what are known as residential proxy networks. These systems funnel millions of home Internet connections into a unified network, and the proxy operators allow attackers to route their malicious traffic through these connections for a fee. The online services see only IP addresses with good reputations and geolocations that don’t stand out.

More often than not, the home users have no idea that their connections are being used to facilitate crime and occasionally even nation-state attacks. Users who do know often don’t care much. In exchange for leasing out part of their unlimited bandwidth to others, many get free movie and TV show streaming. Several less tech-savvy people I know who own such digital media players have told me, after I explain how the media players piggyback off their connections, that the bonanza of content is worth it. They find the tangible benefits outweigh the abstract harm they pose.

Infecting already compromised devices

Research published Monday brings the threat into much clearer view. Security firm Plume cataloged a vast ecosystem of malware that preys squarely on users of SuperBox, just one of many media players offering pirated content. These malicious apps can be surreptitiously installed by remote attackers even when the devices are positioned behind a router. While Monday’s deep-dive analysis focused exclusively on SuperBox, Plume warned that dozens of similar streaming devices pose precisely the same threat.

Read full article

Comments

What's new with Hokie Football [www.collegiatetimes.com - RSS Results for * of type article OR video OR youtube OR collection] (12:00 , Monday, 31 August 2026)

A lot has changed within the Virginia Tech football program since the Hokies last took the field in 2025. The team running out of the tunnel to “Enter Sandman” this year will look a lot different from last year's 3-9…

20 Years Ago: Start of the Wide-Tire Revolution [Rene Herse Cycles] (11:26 , Monday, 31 August 2026)

Twenty years ago, in the Spring of 2006, we started testing tires. Almost by accident, we discovered that wide tires can be as fast as narrow rubber. What started as a small project—finding the best tires for long-distance rides—quickly spiraled into ground-breaking research that has revolutionized cycling in the two decades since. Back when we started, nobody—us included—would have predictedTour de France pros rolling on 30 mm tires, or gravel racers winning races on 55s. Virtually all road racers were on 23s, and ‘gravel grinders’ were discussing whether 28 mm was ‘too much tire’ for gravel. A few years later, a Japanese maker introduced the first ‘gravel’ tires: They came in 23, 26 and 28 mm widths! It’s hard to believe today, but that was the status quo in 2006.

Great discoveries rarely happen in a vacuum, but they build on previous research. If somebody claims to have made a brilliant discovery out of the blue, they may not be giving credit where credit is due. For example, Galileo Galilei discovered evidence that the earth was round when he built a telescope. Being able to see the planets in great detail, he noticed that the earth cast a shadow. This meant that the earth was between the sun and the planet… which was incompatible with the generally accepted idea that sun, planets and stars were suspended in the sky above a flat earth. He didn’t set out to disprove that the earth was flat—that was just the result of trying to get a better look at the sky.

Twenty years ago, rolling resistance wasn’t a big topic. We all ‘knew’ that aero was most important, and weight second. And we all thought that the main thing to reduce rolling resistance was to inflate our tires to maximum pressure.

Then the German magazine TOUR published a huge tire test. They tested dozens of 23 mm road tires. They looked at puncture resistance, grip on dry and wet pavement, tire wear—and rolling resistance. (They tested on the steel drum of a well-known German tire maker.)

They found that the slowest tire, the Continental Grand Prix Attack, had almost twice (!) the rolling resistance of the fastest tire, a hand-made supple clincher. However, they echoed the prevailing wisdom and concluded that differences in rolling resistance were not really important: They amounted to just 34 seconds during a 10 km time trial.

I remember reading this and thinking: “Whoa! First, 34 seconds is easily the difference between winning and not even making it onto the podium. And second, as long-distance riders, we’re going much slower than the pro racers, so rolling resistance is going to be even more important for us.” We had already started to ride (slightly) wider tires (28 mm), but there were no tests for any of the tires we rode.

If the rolling resistance of racing tires could vary that much, how much greater were the potential savings for the wide tires we were riding—tires that didn’t always put performance first? How could we find out?

My riding partner Mark and I were straight out of college, and we didn’t have access to a drum testing machine. How about a roll-down test? Both of us had just spent the better part of a decade getting our PhDs. Mark’s was in social psychology with a minor in applied statistics. His specialty was evaluating real-world data. I had studied the history of climate change on Mount Rainier on a NASA fellowship, so I was used to getting reliable results under difficult conditions in the field. Scientific research was part of our DNA, and our combined skill set was tailor-made for designing a useful test of real-world tire performance.

We immediately realized that the big problem with a roll-down test would be to control what scientists call ‘noise’: wind, temperature, changes in rider position, and everything else that would affect how fast a bike rolls down a hill, beyond what we wanted to actually test: tires.

Minimizing noise was key. Most important was wind—we knew that we should test only when there was no wind. Even a little wind would mess up our results. We needed absolutely zero wind. From our randonneuring experience, we knew that, in Seattle, the time around sunrise is often completely calm. Next was temperature. On spring days, Seattle often sees very little change in temperature. Next was the rider position. Could our test rider keep the same position for test run after test run? The only way to find out was to try!

What kind of hill was best for our roll-down tests? We needed a hill that started steep and then leveled to a constant gradient. The bike had to get up to speed quickly, with no wobbles. We decided to use a ladder for the rider to hold onto, so they didn’t need to clip in, but could start rolling in the right position immediately. Mark knew an old soapbox derby track, which had exactly the right profile. An added advantage: There were no cars on this ‘road.’ The surface was rough, because it hadn’t been repaved in a long time, but there were no cracks or potholes. A uniform surface is essential, to get repeatable results. The rough asphalt was typical of typical for the backroads where we liked to ride.

The next step was a pilot test. I bought a set of the fastest tires in the TOUR test, plus a set of slow-ish tires. (I couldn’t find the slowest tires in the U.S.) We mounted the tires on our bikes and headed to the test hill just before sunrise on a Saturday morning. I rolled down the hill three times on the fast tires, then we swapped wheels, and I rolled down the hill three times on the medium-slow tires. Same bike, same rider, two different tires. Mark timed me. When we looked at the results, we saw that the three runs with each tire were within half a second—and the fast tires rolled about two seconds faster than the slow-ish ones. Our method had promise!

We then went for our normal Saturday ride. As we rolled through the bucolic Snoqualmie Valley, we plotted a big tire test. We decided that we didn’t just want to test tires, but figure out how tires work. We already had some doubts about the generally accepted status quo—that narrow tires were fastest. Jobst Brandt had pointed out that the contact patch of wider tires was shorter, which should make wider tires faster, at least in theory. The old French randonneurs I met during my historic research had talked about hand-made tires of the 1940s and 50s: “They were wide and supple—and so fast!” I had been riding old French randonneur bikes and tandems with 650B wheels and 38 mm-wide tires. I set many personal bests, even though the wide tires should have been slow. (Back then, I inflated the 38 mm tires to 75 psi. Neither Jobst nor we had any inkling that high pressures weren’t required for speed.)

We drew up a list of what we wanted to test: Different tire widths. Different pressures. Tires with thin and thick tread. Different tread patterns. Different wheel sizes. Different tubes. Over the next month, I bought and borrowed a huge number of tires. We got one tire, the Michelin Pro2 Race, in three widths: 20, 23 and 25 mm. I contributed the ‘wide’ 28 mm Rivendell Rolly-Polys that I’d been using as my go-to tire. We had the brand-new Rivendell 650B tires, one with a puncture-resistant belt, the other without. Mark wanted to test his 28 mm-wide Avocet Slicks and a set of 35 mm Paselas. We borrowed tires with small knobs. We had just started to import Grand Bois tires from Japan, and we had those in 700C and 650B. We even included a set of worn tires to see whether the thinner tread made them faster. And then we rode all those tires for 50 miles (80 km) to make sure they were broken in. It was a huge project, but we were young, and we had time.

Somehow, we already predicted that people might question whether we really could control the noise during our experiments, or whether we were just making up our results. So we asked Alex Wetmore, who was well-known in local cycling circles, but not yet a friend, to help with our testing: Alex was going to be a second timer of the roll-down runs, independent of my time keeping. Alex also contributed the test bike, an old Trek that had ample tire clearance. Mark was going to roll down the hill. Alex brought along his friend John Speare, who was visiting from Spokane, as an additional observer.

Then came the big day. The day before, I had gone out to sweep any loose rocks and gravel from the test hill: Hitting small rocks during one run and not the next might have affected our results. Then the four of us met at 5 a.m., just before sunrise, and started testing. We were lucky, the weather forecast was correct, and there was no wind at all. We kept checking the leaves of the trees surrounding our test track—if they moved at all, we stopped our testing, until they were calm again. (Later, some ‘experts’ tried to discredit our testing for being unscientific for this simple, but effective setup: Tree leaves are actually a more accurate indicator of still conditions than wind speed meters. In any case, our statistical analysis would have detected if wind influenced our results.)

Mark rolled down the test track again and again. We switched tires and wheels, then repeated the tests. At 7:30 a.m., the leaves of the trees started moving ever so slightly. We waited for a while, but they didn’t stop: A light wind had sprung up. That ended our testing for the day.

It was disappointing, but we already had done 2.5 hours of testing, and we had first results. We had tested the same tires on multiple wheels and found that the results were identical. That meant we could ‘pre-mount’ tires in the future, and switch wheels, to get as much testing done as possible while conditions were good. Compared to our preliminary testing, refining the method had reduced the variability between runs of the same tires even further. We also found that the results of the two timers—Alex and me—were extremely consistent. That gave us additional confidence in our methods. The testing may not have looked very high-tech, but careful work is more important than flashy gizmos when it comes to doing science. The real test in science is how results hold up over time—and I think we can say in all modesty that we’ve passed that test.

After our first test session ended prematurely due to wind, we returned when the weather forecast looked promising again. For our next two test days, we were lucky: There was no wind all morning, and temperatures were constant, too. (We later realized that temperature had a huge influence on tire performance, something that also wasn’t generally known at the time.) We tested the same ‘reference tires’ first thing, in the middle, and at the end of each day. During those three days, we did more than 150 test runs. Mark had spent more than 20 hours rolling down (and riding up) the hill, his arms and upper body always in the same position. The riding wasn’t exciting, but the results were!

And those results were not at all what we expected! First, Mark did a statistical analysis to check which tires/pressures/setups were performing differently, and which were so close that we couldn’t tell which was faster. (There’s always a little noise—in our case about 2%.) The biggest surprise was that tires made a surprisingly large difference: On the fastest tires, the bike rolled 20% faster, compared to the slowest tires. That was huge! Imagine taking 10 hours off your time in Paris-Brest-Paris just by choosing different tires! Rolling resistance had a far greater effect on real-road speed than cyclists thought at the time.

A little disappointing at first: The tires we were running on our bikes were among the slowest. And the Grand Bois tires we were selling were also among the slower tires. That disappointment quickly turned into excitement: There was a lot of room for improvement!

The next surprise: Tire pressure did not have a meaningful effect on speed. Our tires rolled as fast at moderate pressures as they did when we inflated them to the max. That ran against everything we—and everybody else—believed. On the drum tests that everybody was using back then, high pressure, more than anything else, made tires fast. That’s why racers inflated their tires to 130 psi (9 bar). I ran Rolly-Poly tires because they were the 28 mm tires with the highest pressure rating: 120 psi (8.3 bar). Our tests showed that I could run my tires at 85 psi and not lose any speed.

Testing the three set of Michelin Pro2 Race tires, we found that wider tires had less resistance. The 25 mm rolled faster than the 23 mm, which faster than the 20 mm. The ultra-supple hand-made tire that had scored highest in the TOUR test was also fastest in our roll-downs, but other tires that scored well on the drum didn’t roll fast in the real world. Surprising was the result for the Mitsuboshi 650B tires I had run on the old French rando bikes: It was the fifth-fastest tire in our test, as fast as the narrow racing tires we tested.

We did a regression analysis to determine which factors were most important for making a tire fast. We found that a supple casing was the reason our fastest tires rolled so well. We also realized that many ‘performance’ tires rolled so slowly because they were designed to handle high pressures: They had strong casings that were also very stiff. And the wider the tire, the stronger the casing. Making a 28 mm tire to support 120 psi (8.3 bar) meant beefing up the casing to an extreme degree. We realized that, to be fast, wide tires should use supple casings. That would reduce their pressure rating, but we had found that this didn’t make them slower.

Why were our results so different from what other testers had found? Previous testing had been in the lab, on a steel drum, without a rider on the bike. We realized that the vibrations of riding on real roads slowed down the bike. That’s why high pressure didn’t make tires faster on real roads: They increased vibrations, which canceled out the benefit of less deformation as the tire rolled. Previous tests had measured only the tire’s deformation, but not the vibrations it transmitted.

Our finding—that high pressure wasn’t needed for speed—changed everything. Before, tires were either wide or fast. If you wanted to go fast, you put on narrow tires and inflated them to max pressure. Here’s why: With tires, you have to choose two of the following three:

  • high pressure
  • supple casing
  • wide tire

In the same tire, you can have only two, not all three. When we thought that high pressure was essential to speed, we had to choose between a supple casing or a wide tire. That’s why the 24 mm hand-made clincher was so fast, both in our real-road tests and on TOUR’s drum: It had a supple casing. And since it was narrow, it could also handle high pressures. That’s also why my 28 mm Rolly-Polys were so slow: They were wide, so they ‘needed’ a stiff casing to support high pressure. Once we realized that high pressure wasn’t necessary to go fast, it opened the way for making tires that were wide and fast.

Being scientists, we sent our results to experts for review. Frank Berto, who had tested tires himself, provided input. Jim Papadopoulos asked for an elevation profile of our hill, to calculate whether the initial phase, before we started timing, affected our results. So I went out and spent a day surveying the hill in detail. Jim’s calculations confirmed: The bike always entered our timed section at the same speed. Andreas Oehler in Germany was another reviewer. These names may not be well-known today, but they were the bike tech experts at the time.

Results for tires in different widths (actual sizes)

Then we wrote up our findings in Bicycle Quarterly 17, the Autumn 2006 edition. Our project had started as a simple test to find the fastest tire wide enough for long-distance riding. Instead, we had revolutionized our understanding of how tires work. The idea that high pressure was necessary for performance had become obsolete overnight. The implications were obvious, and they were huge. Our article concluded:

“Our findings point to a new direction for performance bicycles. For most cyclists, wide, supple tires at low pressures offer more speed, better comfort, increased versatility and improved safety than the currently favored narrow high-pressure tires. However, this type of wide, fast tire is not currently available. Hopefully, our results will persuade manufacturers to produce the ‘ultimate’ tires.”

That was a big statement from two young guys who weren’t even part of the bike industry. But we were confident in our results. We had checked and re-checked them. Our statistical analyses were sound. Experts had reviewed our work and approved of it. We were also putting our results to the test in the real world: We rode the fastest tires from our test in long-distance brevets, and our times improved by roughly the same amount as our tests predicted. Which was a lot—15% faster in a 1000 km brevet meant taking 6 hours (!) off my previous personal best.

We were young and naive. We thought that, now that we’d shown the way, tire makers would start developing wide, supple tires. Of course, we now know that the bike industry wasn’t interested, for many reasons. After spending a few years pushing others to make the tires we envisioned, we finally made them ourselves. That’s how the Rene Herse Cycles tire program started. It took more than a decade until our findings became accepted in the mainstream. In the next part of this story, we’ll look at what it took to get there.

Further Reading:

KI6CR: Two Swiss Summits – SOTA Activations at Schilthorn and Burgfeldstand (Part 1) [Q R P e r] (10:07 , Monday, 31 August 2026)

by Chris (KI6CR) What a beautiful week for radio in the Alps! My family and I were taking one last summer trip before the kids headed back to school. Ever since our planned summer 2020 trip to Switzerland was canceled due to the pandemic, my wife Debbie and I had been dying to reschedule. We … Continue reading KI6CR: Two Swiss Summits – SOTA Activations at Schilthorn and Burgfeldstand (Part 1)

Weekend Snapshot [BIKEPACKING.com] (09:03 , Monday, 31 August 2026)

Weekend SnapshotOur latest Weekend Snapshot features scenes from reader-submitted bikepacking trips around Alberta, Idaho, and British Columbia. Soak in the views and use the quick form to contribute to a future installment of our longtime Monday morning series here...

The post Weekend Snapshot appeared first on BIKEPACKING.com.

Cycling Across the World’s Biggest Salt Flat (Video) [BIKEPACKING.com] (08:44 , Monday, 31 August 2026)

Bikepacking Salar de UyuniDan Camp's latest video transports viewers to Bolivia's extraordinary Salar de Uyuni, the world's largest salt flat. Watch his 39th video installment detailing his bikepacking journey between Alaska and Argentina here...

The post Cycling Across the World’s Biggest Salt Flat (Video) appeared first on BIKEPACKING.com.

The Corsica Crossing [BIKEPACKING.com] (07:39 , Monday, 31 August 2026)

Corsica Crossing Bikepacking RouteNicknamed “the mountain in the sea,” Corsica is the most mountainous island in the Mediterranean, and riders who take on this route will intimately experience its many ups and downs. […]

The post The Corsica Crossing appeared first on BIKEPACKING.com.

Wider Horizons – Making Panoramic Images [35mmc] (05:00 , Monday, 31 August 2026)

When photographers talk about “going wide”, the typical assumption is that they are referring to using lenses with shorter focal lengths. Yet many pictures that “feel” wide– such as pictures of grand, sweeping landscapes– are taken with the long lenses associated with portraiture or wildlife photography. These images – and many others – are accomplished...

The post Wider Horizons – Making Panoramic Images appeared first on 35mmc.

You Can't "Vibe Code" Love [Coding Horror] (12:35 , Monday, 31 August 2026)

You Can't "Vibe Code" Love

About a year ago, I was offered a presentation slot at the WeAreDevelopers World Congress in Berlin. I rarely take speaking engagements, especially international ones, but this one arrived at just the right time, the right place, and with the right person – I said yes, on the contingency that Ben Dumke-von der Ehe joins me in the presentation. Ben is an early community hire at Stack Overflow who lives in Berlin. We both had something important we wanted to talk about, with slightly different opinions, different perspectives.

They agreed to the requirement. This is our presentation, as delivered Friday, July 10th at the CityCube in Berlin.

The heart of this presentation is the human story of how Ben and I met. As I said on stage:

The long version of this story is that we changed each others' lives. The short version is: unicorns.

You should hear Ben's side of the story directly from him. It's remarkable. Also remarkable is that he left Stack Overflow, then came back... and then left again. I told Ben, I'm mostly doing this presentation because I owe you so much.

You Can't "Vibe Code" Love
it is difficult to get a German to smile so you have to sneak attack

Not only Ben, but everyone who participated on Stack Overflow. We did it together, under a bedrock Creative Commons license that respected the effort everyone put into their work to build this commons, this hugely influential Stack Overflow dataset.

You Can't "Vibe Code" Love

It's fine that LLMs intercept most (if not all) of the common programming questions; this aligns with the higher level goals of Stack Overflow:

Passively searching and reading highly ranked Stack Overflow answers as they appear in web search results is arguably the primary goal of Stack Overflow. If Stack Overflow is working like it’s supposed to, 98% of programmers should get all the answers they need from reading search result pages and wouldn’t need to ask or answer a single question in their entire careers. This is a good thing! Great, even!

LLMs also cut the gordian knot of constant duplicate questions that I thought was basically impossible.

When you’re asking a question on a site that doesn’t allow duplicate questions, the problem space of a site with 1 million existing questions is rather different from a site with 10 million existing questions... or 100 million. Asking a single unique question goes from mildly difficult to mission almost impossible, because your question needs to thread a narrow path through this vast, enormous field of prior art questions without stepping on any of the vaguely similar looking landmines in the process.

The way LLMs can map a duplicate question using entirely different words to an existing answer is incredibly transformative, and a huge net positive. Less duplication. More answers, faster. That's the point.

But what happens to the commons when everyone is privately whispering to LLMs? What happens to the communities where programmers learn from each other, the very communities like Stack Overflow where Ben had so much fun becoming a programmer, and got hired for doing things he loved?

You Can't "Vibe Code" Love

You can't "vibe code" love.

Where exactly will the LLMs get their next set of training data from, if we don't build for and protect the commons together? How will we continue to share knowledge – and love – with each other?

Sunday, 30 August 2026

Debian 11 Long Term Support reaches end-of-life [Debian News] (08:00 , Sunday, 30 August 2026)

The Debian Long Term Support (LTS) Team hereby announces that Debian 11 bullseye support has reached its end-of-life today, 31 August 2026, five years after its initial release on 14 August 2021.

Football, family and the number four: Tyseer Denmark's journey to Blacksburg [www.collegiatetimes.com - RSS Results for * of type article OR video OR youtube OR collection] (04:40 , Sunday, 30 August 2026)

The number four has followed Virginia Tech wide receiver Tyseer Denmark nearly everywhere.

Season preview: Virginia Tech looks to turn the page in 2026 [www.collegiatetimes.com - RSS Results for * of type article OR video OR youtube OR collection] (03:22 , Sunday, 30 August 2026)

At Virginia Tech’s spring game on April 18, Hokies head coach James Franklin addressed the crowd at the end of the first quarter, claiming that the team and fanbase were “going to shock the world together.”

VMI preview: How will the James Franklin era start? [www.collegiatetimes.com - RSS Results for * of type article OR video OR youtube OR collection] (03:21 , Sunday, 30 August 2026)

The Hokies should be going into Week 1 with some confidence against FCS opponent VMI. They’ve only ever lost to two FCS opponents and have won 13 straight games against non-FBS schools. Their last FCS loss came in 2010 to…

HokieAI expands AI access for Virginia Tech students [www.collegiatetimes.com - RSS Results for * of type article OR video OR youtube OR collection] (02:38 , Sunday, 30 August 2026)

In response to the evolving AI landscape, Virginia Tech released HokieAI, a university-supported AI platform that “provides metered access to vetted commercial models for generative and agentic AI.” It is intended to support teaching, learning, research and administrative work, according…

Hokies lose nail-biter to William & Mary in second match of the season [www.collegiatetimes.com - RSS Results for * of type article OR video OR youtube OR collection] (02:26 , Sunday, 30 August 2026)

Blacksburg, Va. — The Hokies hosted William & Mary at Cassell Coliseum as part of the annual Hokie Invitational on Friday, Aug. 28. They started the season off strong with a 3-0 match win over the Kent State Golden Flashes…

Volleyball picks up win in season opener vs. Kent State [www.collegiatetimes.com - RSS Results for * of type article OR video OR youtube OR collection] (02:13 , Sunday, 30 August 2026)

Virginia Tech volleyball started the 2026 season with a win, beating Kent State in straight sets at Cassell Coliseum on Friday morning.

Clavicular visits Blacksburg [www.collegiatetimes.com - RSS Results for * of type article OR video OR youtube OR collection] (02:04 , Sunday, 30 August 2026)

On Tuesday, Aug. 25, popular influencer Braden Peters, more popularly known as “Clavicular,” made his way to Blacksburg for a scheduled meet-and-greet event at The Burg Resto-Bar. This visit was a part of a larger college tour initiative by Clavicular.…

New game day procedures aim to improve safety [www.collegiatetimes.com - RSS Results for * of type article OR video OR youtube OR collection] (01:54 , Sunday, 30 August 2026)

As football season approaches, Virginia Tech is implementing new procedures designed to ensure crowd management and a safe environment for attendees.

Inside Meta’s push to put robots to work in data centers [Biz & IT - Ars Technica] (07:03 , Sunday, 30 August 2026)

Meta is testing robots that can plug in cables, reset servers, and handle other tasks inside its data centers, according to several current and former workers familiar with the projects. The ongoing effort, which has not been previously reported, may eventually allow Meta to operate its rapidly expanding data center footprint with fewer humans, keeping labor costs in check as its spending on AI infrastructure soars.

Meta is using robots and related hardware from several different vendors, including Watney Robotics, Kinova, and ABB, according to the same workers, who asked to remain anonymous because they weren’t authorized to speak to the media. Kinova and ABB declined to comment. Watney didn’t respond to requests for comment.

In one experiment, Meta is evaluating whether a Kinova Gen3 robotic arm could be used for power cycling or cutting off electricity to servers. The company is also testing a different robot to swap networking cables. One Meta data center worker estimates that if it’s successful, the bot could replace up to 80 percent of some people’s workloads. “We thought those of us performing the physical tasks were safe for a while, but not anymore,” says the worker. “It’s coming for us all, unfortunately.”

Read full article

Comments

Eight Parks, One Missing Mast, and the Midnight Sun: A CW-Only POTA Trek Along Sweden’s Kungsleden [Q R P e r] (06:00 , Sunday, 30 August 2026)

by Nils (DF7CW) Greetings to the whole QRPer community — and especially to you, Thomas! My name is Nils, DF7CW, and this is the story of an arctic POTA expedition along the Kungsleden trail in northern Sweden: five activations, eight parks, one dramatic mast crisis, and more rain than I care to admit. I’ve always … Continue reading Eight Parks, One Missing Mast, and the Midnight Sun: A CW-Only POTA Trek Along Sweden’s Kungsleden

Saturday, 29 August 2026

Bug blindness [] (08:00 , Saturday, 29 August 2026)

I used to wonder why I see so many more bugs than most people. I easily observe hundreds to thousands of bugs per week and nothing seems to work, but most people I talk to don't see anything like this. For a long time, I thought this had something to do with how I use computers but, over time, I've realized that it's mostly that people are hitting the same bugs and don't notice.

If you're not a programmer, that's probably a better way to see the world, but I think curing quality/bug blindness is helpful for programmers. I've done this with a lot of friends and acquaintances (just by pointing out bugs). After a few weeks, people who are so inclined tend to start noticing bugs as well.

Because I notice these kinds of things, I've had multiple jobs where directors/VPs/execs/etc. sometimes ask me to evaluate something when they want an actual opinion from someone who is relatively likely to notice issues (and fix them or drive fixes for them if necessary). Sometimes I won't find any issues (there are likely issues that just aren't the kind I notice). More often, I find issues that fall somewhere from "mild" to "moderate". And, sometimes, the issues are severe, to the point where one might even say the thing actually doesn't work.

I find this last category a bit mysterious, as when I look up discussions on how the thing got into this state, there's usually a stream of internal comments indicating that the thing is great, it works well, etc., but when I open up the thing and try it, it's in a state where the thing only works if you do quite a few non-intuitive workarounds. More likely than not, not only would a normal user not be able to use the thing, they'd have such a hilariously/infuriatingly bad experience that they'd tell their friends.

I've had this post in mind for maybe a decade or so, but I was hesitant to write it up because, in the back of my mind, I always wondered if I'm somehow triggering weird corner case behavior most users don't hit without realizing it. But after seeing more and more cases where the product launches and falls flat on its face because users run into the exact same issues I saw, I don't think that, in general, I'm hitting bugs because I'm doing unusual things a normal user wouldn't do. If a product seems severely flawed when I use it, it probably is. And with the magic of LLMs, nowadays, I can even have LLMs act like normal users in a lot of ways and show that the issues reproduce across many different scenarios.

A few examples

I don't want to give any specific examples where it was my job to see how well the thing worked because, even if the internal examples are meant in a constructive, blameless, way, they may not always read that way when re-posted externally, so I'll give a few less interesting and less well supported "random" examples.

A while ago, I wrote up the results of some web search queries and found poor results from Google, Bing and Kagi. In general, the major search engines failed to return good results for the queries and returned pages full of low-quality SEO spam as well as some sites that were actually scams. BTW, on the scale mentioned above, I would consider this "moderate" and not "severe" (severe would be something like, the search engine returns 500 errors half the time, the majority of results are scams, etc; my bar for severe is that a normal user likely won't be able to use the thing at all, not that they have a bad experience). Almost nobody1 objected to my characterization of Google and Bing search results, but people told me that I was wrong about Kagi. In some cases, people sent me their actual search results. In every such case, the search results did not contain a good result that I could see (e.g., for the seasonal forecast query, the search failed to return an up-to-date seasonal forecast, for the tire result, there was no result with a correct explanation, etc.) and was full of SEO spam. In one case, a person passed me both their list of Kagi filters as well as the search results they got without making claims that the results were good or bad, but people generally insisted the results were good even though the results both failed to link to a matching result and were full of spam except in cases where the user did something like pin GitHub to the top of their results, which worked for the queries where the goal was to download software that's hosted on GitHub, but of course completely fails for the other queries from the post as you're unlikely to get your regional seasonal forecast or the correct explanaton of tire mechanics on GitHub.

In the abstract, I get that people who are fans of things tend to be blind to the thing's faults. For example, since I bought a Volvo after seeing how they do in out-of-sample crash tests, I sometimes search for answers to my questions on Volvo car forums. For well over a decade, the reliability data that exists (and I think this is backed up by the anecdotal experience that mechanics who work on Volvos have) is that Volvo reliability is mediocre to poor, but of course Volvo forums are full of people who insist that Volvos are among the most reliable cars and that the data are all wrong.

An example that might be more central to the topic is Blackboard (the course management software). Back when it was the most widely used software by universities for coursework, the software was widely disliked by both students and professors. I think it would be fair to say that it was the most widely disliked software in my social circles (there was more strongly disliked software, like Visual Source Safe, but any more strongly disliked software wasn't widely used enough to be the most widely disliked overall). The Wikipedia page notes

Blackboard had become "one of the most disliked — even detested — companies in education."

as well as

In December 2011, Fast Company reported that 93% of respondents to the Amplicate customer opinion survey "hate" the company.

Back when I was much younger and had less of a filter, I ran into someone who worked at Blackboard and, without thinking, I stupidly blurted out something like "what's it like to work on this software that so many people dislike?". Luckily, the person I was talking to wasn't offended at all and, instead, they were actually confused because they thought it was widely loved software that users really liked. They didn't really believe what I said could be true and I made some comment indicating that it was just confusion on my part and then the conversation continued in a different direction. At the time, as someone much younger and more naive, I was really surprised to hear that the software that was probably the most widely disliked software in my social circles was thought to be really well-liked software by the one employee from the company I met (and, presumably other employees as well).

I can understand how the Volvo forums get to be how they are, in that cars are reliable enough in general now that people generally don't experience car breakdowns, so it's easy for someone to think something like "the data can't be right; after all, my car has never broken down". It's more of a mystery to me how somebody can look at a set of search results that are full of spam and then dash off a message explaining how great the results are, even if they're a fan of a particular search engine or how someone can think that users generally love software that's famous for being disliked, to the point that every single person I talk to about it tells me how bad it is (often in unprompted complaints), there are news articles that discuss how much people dislike the software, and the near-universal dislike for the software is mentioned on its Wikipedia page. Another Blackboard-like example might be Discourse (forum software) web performance, where one of the inspirations for this post was discussions with Discourse employees who thought that Discourse had great performance. I found that one interesting because Discourse actually had code in it that slowed down actual page loads in order to cheat on web performance metrics like LCP. That went well beyond just optimizing for a benchmark and rose to the level of actual cheating that not only had no benefit to the user, it actually harmed the user. At some level, the programmers implementing that sort of cheating and advising users on how to not accidentally subvert the cheating must know that the actual performance of their app is poor, but it's very easy for people to put up mental barriers around this kind of thing.

By now, I wouldn't say that I'm surprised because I've seen this kind of thing enough that I would actually consider it surprising if it didn't happen, but I still wonder what's going on inside someone's head when something like this happens.

For a non-programming example, we previously noted in this post on how people have different perspectives on "obvious" facts, there's a basketball player who, subjectively, is generally considered to be the dirtiest player of his era. The NBA doesn't track objective measures of player dirtiness, but he seems dominant on a wide variety of measures. For exampe, although, like rebounds before 1950, genital strikes aren't an officially tracked stat, he surely holds the record for punching, kicking, kneeing, or otherwise striking players in the genitals this century (he should also hold the record for era-adjusted numbers, but it's possible that he doesn't have the all-time record due to play being much dirtier overall in the 80s and 90s). In discussions, most fans of his team don't seem to notice this and the phrases "natural rebounding motion" and "natural shooting motion" have become running jokes from how oblivious the team's fans are when they justify this player's contortions when he strikes other players in the genitals.

On average, humans have a high ability to ignore negatives in things they're a fan of, including (and often especially) their own work or work their company does. For better or for worse, I seem to have the opposite of this and my thoughts immediately go to the flaws in myself and my work. A number of times, as a result of a blog post, someone has messaged me with something like "how would you like it if someone criticized your work?" or "how would you like it if someone said your work isn't good?" To the former, my thought is that I go to great lengths to get criticism from people who can poke holes in my reasoning, so it's pretty awesome if someone has remotely reasonable criticism of my work. And to the latter, I generally think my work is full of major flaws, so, uhh, yeah, it seems pretty reasonable to say it isn't good. There are particular aspects of my work that I think are interesting or good but, overall, I don't know that I'd rate anything I've done as good. I'm not saying I don't have blind spots, but I think I'm a bit less prone to this particular one than most people2.

Habitual mitigations

If I think about analogous blind spots I've had, one that jumps out at me is from when I was a little kid and a friend of mine used my computer. For this story to make sense, you have to know that this was in the mechanical mouse era. Over time, detritus would get stuck to your mouse ball and cause it to track erratically unless you cleaned it out.

When my friend tried to use my computer he found it impossible to use the mouse because mouse pointer movement seemed almost random. When I sat down at the computer again and used the mouse I didn't have any problem using it at all, but on looking at what I was doing with my hand to smoothly move the pointer in a straight line, I was violently throwing my hand all over the place. I realized I must've adjusted to the detritus on the mouse ball over time as it accumulated and I was somehow compensating for the mouse's extremely erratic tracking by making countervailing erratic movements3. I thought it was pretty amazing that I could not notice that I was doing this and I always wonder if I'm doing some equivalent thing today.

I sometimes think about all of the mitigations I've developed to work around bugs. For example, when opening a new Google Doc, I used to immediately put the title I wanted into the doc. At some point, maybe ten years ago or so, Google Docs added some kind of delay such that the typing you do into the title box right after you open the doc gets overwritten, so I now have this habit where, after opening a Google Doc, I do something else and then I change the title. Over time, as Google Docs has had more and more features added, I've developed a series of habits that avoid all sorts of pitfalls (such as trying to search at the "wrong" time and getting the useless native browser search instead of the Google Docs search).

My feeling is that a large fraction of computer literacy and software literacy is developing a large library of these habits that you just do at a non-conscious level. These are often quite specific to the situation, such as a habit I developed when I worked at Microsoft of flipping my laptop's WiFi switch to off before logging in (which I noticed other people doing as well). This was because there was some service, which would often fail your login with "There are currently no logon servers available to service the logon request”. But if that service couldn't connect at all, the check would be bypassed and you could just log in.

Quality blindness

We could fill a post up with examples like that, but back to the main topic of the post, one commonly suggested way to try to overcome quality blindness is to have people dogfood their own software. On average, this is a lot better than not dogfooding, but it only works to the extent that people don't figure out (and then forget about) habits that work around whatever issues the software has. On average, programmers are pretty good at working around software foibles (you had to be in order to be an effective programmer pre-LLM), so it's very easy for programmers to not notice these kinds of issues if they're not paying attention.

On the flip side, a large part of making an app easy for people to use seems to mean making weird habits like these unnecessary. Although this sounds like it should be easy to do, from having seen people try to give feedback about this kind of thing, the reflexive reaction of most developers seems to be "huh? It's easy to do X, just do [complex sequence of things that no normal person would think of if they hadn't used the app many times before unless it was specifically explained to them or they saw someone else do it]" or "huh? Didn't you see that the instructions for this are clearly laid out in page 43 of the manual after you execute the steps in Appendix B on page 261?".

That being said, I think curing people of quality blindness is do-able because I've done it quite a few times. I think this only really works when the person is receptive, as people have infinite capacity for willful blindness but, in cases where people are receptive, just pointing out issues they didn't notice seems to work. Years or even a decade later, people will sometimes tell me they see bugs everywhere now.

The reason I think this is worth doing is that I've seen people and teams with a high degree of quality blindness ship things that have reduced or even no chance of success because of product quality issues4. It's one thing to knowingly and deliberately trade off quality for speed5, but when I've seen this happen there's always been a kind of quality blindness where everyone involved with the project thinks they're shipping something very high quality when that's not the case.

This has never been unimportant, but it's gotten more important with coding agents because, while it's easier than ever to churn out low quality software, it's also easier than ever to improve quality, whether that's better performance, fewer bugs, etc.

But, to do this, you have to actually notice that this is possible, that quality can be improved.

Thanks to Yossi Kreinin, Dennis Snell, Michael Malis, Emu Chu, Gary Bernhardt, Jon Surrell, and Matt Mullenweg for comments/corrections/discussion.

Naturally, Gary Bernhardt ran into a Google Docs bug while reading a draft of this post.

P.S. Like I've mentioned in the last four posts, I've been trying to write posts more quickly because, with LLMs, it's so much easier to look at data and figure things out but, since I'm not writing with LLMs, the time it takes to write something up hasn't fundamentally changed, unless I want to move to a different point in the quality-velocity trade-off space. The prior result was that I would run some experiments and tell a few friends and then never write anything up because, due to Amdahl's law, writing anything up would effectively consume all of my bandwidth for running experiments. In fact, despite trying to do this (my goal is to spend 30 minutes per post on the write-up), since writing my last post, I have three results that I think could make a totally fine blog post that I haven't had time to write up (not including things done for work, which would add a few more things). Without having LLMs write for me, I don't see a reasonable way to get the time per post significantly below 30 minutes (and I think I often miss my goal and take more than 30 minutes), so the non-LLM options here are some posts that are much sloppier than my normal posts (in a human slop kind of way), or almost no posts.

Anyway, if you have opinions on these quick (and surely more wrong) writeups, let me know what you think (X Bsky Mastodon)!

Appendix: advertising blindness

Michael Malis (founder and former CEO of Freshpaint) noted (in messages, hence the message-like format)

For a similar but different data point - I’ve seen similar blindness when it comes to advertising. When I would explain Freshpaint to people, I would tell them that we help hospitals with marketing

A common question I get is why do hospitals do marketing. The weird thing is if you pay attention, hospitals do a ton of marketing

In SF there’s tons of bus ads and billboards for ucsf/sutter health/stanford and various treatments

This is a different topic from both Michael's comments and the post, but I'll say that I've talked to quite a few people who don't believe ads work at all, but I talked to someone whose data methodology and judgement I trust about ads A/B testing at one big company I worked for and looked at the data myself at another company and I thought the causal evidence for ads providing real lift (well beyond the cost of the ad) was strong in those cases. In the case where I looked at it, they did a geo-segmented A/B test where they bought ads in some geos but not others (this was done worldwide, with the regions being things like U.S. states, Canadian provinces, etc.). This kind of geo-segmentation was done because, even with cross-device tracking, it's not 100% clear if someone has been exposed to an ad or not (of course this is still the case with this kind of segmentation and I would prefer segmentation that was more clustered to population areas and didn't have splits where people are relatively likely to, for example, commute from one side of a boundary to the other, but this kind of contamination generally makes the likely true lift higher than the estimated lift), so people sometimes do these geo-segmented A/B tests.

Anyway, in these A/B tests, return on ad spend was quite good just on direct revenue gain, and there was also a gain in users which seems likely to result in more revenue down the road (the later revenue wasn't analyzed). I don't know about ad effectiveness in general or if your particular ads are effective, but the commonly repeated idea that ads don't work in general seems wrong to me.

On the topic of Michael's comment, I think it's easy for programmers to not notice ads. Almost all programmers I know use an ad blocker and, in real life, their eyes seem to just skim over ads and not notice them. I can see how this would feed into the idea that ads don't work. Who the heck would look at these things? But from my interactions with "normal" people as well as the data I'm familiar with from my time at Google, many or perhaps most people don't even realize that a lot of ads are ads. When they do a Google search and they click on the top resut, they often have no idea they're not looking at what Google "thinks" is the best link, they're looking at at a link from whoever paid Google the most to buy that ad slot.

Appendix: comments from other folks

Em Chu, on a habitual bug mitigation:

I'm sure you can collect infinite examples for this section, but I just want to Complain: when waking up and unlocking my laptop (mac), it's very easy to get it in a state where it's "awake" but unusable (black screen with cursor or similar) which can only be fixed by physically closing the lid and re-opening it. To work around this, I think I usually wait a second after the screen turns on, interact with the trackpad, and then unlock it, though honestly that happens mostly subconsciously, and I clearly need practice given that I still hit the bug a few times a month.

On reading this, I examined how I open my laptop and realized that I have some funny habits as a result of working around other laptop bugs. The specific bug mentioned here doesn't reproduce on my laptop and it seems that I can stop the habitual mitigation I put into place for some prior laptop.

Gary Bernhardt, on his experience reading a draft of this post

While reading it, Google Docs' UI seems to have broken, making it impossible to scroll up to read some comments (see screenshot [not shown in post]).

From looking at the screenshot, I've seen the exact same bug and have some mitigations for it (different ones depending on the context). I would personally rate Google Docs as far above average in terms of software quality: I find it much less buggy and janky than the major alternatives (Microsoft Word, Open Office, various old editors that are long gone like StarOffice, Lotus, etc.). And yet, I could easily sit down and write a 10k word post on Google Docs bugs and the workarounds I have for them.

At times, I've tried to see if I can get a job somewhere where I just fix quality issues all day. This has never panned out, due to some combination of this not being a very high priority and it also not being a normal role that companies have a role for. I sometimes daydream about joining companies as an intern and just fixing quality issues for a few months and then leaving. In practice, I think if I got such a job, a lot of the fixes would get blocked and it would be very difficult to actually drive change as an intern for three months, so it would have to be some mostly abandoned project where nobody cares what I do (and corporate priorties aren't so focused on shipping features that fixes get immediately re-broken).

@IncidentNoodle

Unintentionally on topic: the <abbr> tags worked on mobile ~last week, but are no longer working across any iOS browser (safari/chrome/firefox), and I had a hard time figuring out that hover showed them on macOS browsers (all three) due to the long delay

@gunchleoc@mastodon.scot:

Germans have a word for that - Betriebsblindheit

@oulipien.bsky.social:

Crazy anecdote from @danluu.com here and I wish he'd been even blunter at the time and asked this person where they'd gotten this belief about Blackboard being liked by anyone at all. User surveys? Principal (as in, not agent) surveys? Inner conviction??? [screenshot of Blackboard anecodote]

[Some variant of, people are forced to say that they don't see bugs by their bosses]

I don't thnk this is consistent with any of the major examples in the post, let alone all of them. Consider the Blackboard example mentioned above. It's unlikely that I and other people this Blackboard employee are "secret shoppers" who are checking in on employees, and the employee's reaction is clearly absurd to anyone who isn't such a hypothetical (and in reality, non-existent) secret shopper, so reacting like this just makes them look a bit silly in the eyes of a large fraction of the people they meet for no benefit (for example, see the previous quote, which seems like a typical internal reaction). Perhaps a few very paranoid employees would maintain this front on the off chance they run into some friend or relative of the boss who knows that they work for the company who relays the story back and they have a boss who would care about this, but it's just not plausible that this is (for example) the case for every Discourse employee who reached out to me to explain to me that Discource performance is really good.

If we look at the basketball example, this is even more absurd. You could possibly come up with some kind of reasoning like, other fans would shun you if you didn't believe or pretend to believe the most absurd rationalization, but as someone who has spent a lot of time around sports fans, I've generally not found this to be the case. And, to the minimal extent to which this is kinda sorta the case, it's more an issue of self selection, where fans who are into the most extreme rationalizations will spend more time around fans who are into the most exterme rationalizations and fans who are less into these extreme rationalizaitons will tend to spend more time with fans who are less into them.

Also, just looking at the career path of people who don't buy into these things and fix them, noticing these issues and fixing them has gone very well for those people. Pretending these issues don't exist (whether that's at a concious level or not) also seems to work well, so I don't know that fixing these issues is actually a better career path, but it's certainly not so bad that, in general, there's meaningful career pressure to pretend these things don't exist overall even if there are some individual positions where there's some direct pressure to pretend these issues aren't real.

Daniel Gibson:

Who else uses the Shift key to end the screensaver, because in case the event goes through to an actual program it's least likely to do have unintended effects?

This reminds me of how, when I want to send a queued message to codex immediately and interrupt the current tool call, I put my finger on the key and the press as quickly as possible to reduce the window of time where the tool call will finish and the escape key will stop codex entirely instead of causing the message to send. I should probably just run a patched version of codex that has fixes for this and a few other issues I've run into, but I'm already doing things like trying out some weird workload-specific optimized version of ripgrep that also has an added native code compiler which compiles matching expressions in another thread while the search starts and then cuts over after compilation completes, so it's not like I'm against creating weird patches to improve my workflow and it's more of an issue of overall bandwidth (no doubt, on writing this, someone will tell me that I could just hit another key instead and could've found this out by asking codex about the key in the time it took me to write this comment). Just like with Google Docs, I consider codex above average in terms of software quality in the space, but even though I haven't been using it for a year, I could easily write 10k words on all the workarounds I've implemented (either by habit or, in some cases, with actual scripts that monitor for broken behavior and then correct it).

@kleschby

[translated from Russian, probably loses something in translation] It seems to be a well-known joke: people use broken software and think it's normal. Something like, "To save a field's data, you have to click the mouse on the adjacent field." Those who always use the keyboard will notice the problem, but those who are more proficient won't.

This reminds me of how, after some macOS update, some old apps wouldn't refresh/update their UI except when the focus was switched to them. To see any changes in those apps, I got the habit of tabbed away from the app and the tabbing back every time I did an action where I'd want to see the result. I don't know if this ever got fixed because I eventually stopped using those apps. If this happened today, I'd probably use an LLM to edit the binaries to fix whatever issue was causing this.

John Regehr:

I didn't remember this until after I'd written and published this post and someone linked to John's post, but apparently John Regehr wrote "Operant Conditioning by Software Bugs" before I started a blog! There's a good chance I saw the post before at some point and then completely forgot about it. Maybe I should've used an LLM to search for prior art, but if I did that, I'd probably never write anything because there aren't that many really new ideas and almost everything is going to be similar to something someone else has said. For better or for worse, I'm much more verbose than John, so this post uses a lot more words and has more random stories thrown in. If you find my blog posts too long, but have somehow managed to stumble down into the bottom of the post anyway, you'll probably like John's post more than mine :-).


  1. someone told me the results didn't reproduce on Google when they tried it some number of weeks later. Of course it didn't, which I discussed here in more detail, but for the short of it, here's this post about scams and other bad results on Google that was #1 on HN for a while. Of course somebody fixed that! And, also, ad results are non-deterministic and, while there are a lot of bad ads, it's not like the majority are scams, so you wouldn't expect to get scam ads at the top results even if someone else did for the same query. [return]
  2. For example, anyone familiar with my code at Twitter will recall the huge comments I had at the top of the main files for the things I owned, which described the various ways in which the thing is really flawed. They were all things that, for one reason or another, I thought weren't worth the time to fix, but they were still serious problems that anyone interacting with the code ought to know about. For this metrics project, I even had a long doc that described the issues in great detail (IIRC, in a lot of cases, the rough shape of the fix was described; maybe today an LLM could take that and fix it).

    I have the same feeling about my writing. While a huge number of bugs sneak through my writing (like spelling and grammatical errors), most of those are things I sort of don't care about and will skim past in other people's writing as well. When I say don't care, it's not that I don't want things to be better (when people send me corrections I generally fix things), it's just that my brain doesn't naturally pay attention to those things no matter whose writing it is, so I don't seem to have a particular blind spot in my writing with respect to these kinds of bugs. For the things I do care about, I could edit posts endlessly because, no matter how much I edit, the post still seems pretty bad to me.

    I used to often (and still sometimes) send a post to someone and ask them if it makes any sense to publish it at all because I generally don't like my output and, if I'm just looking at my own writing, I don't think it's worth publishing. At this point, I've done this enough that I'll often just publish even though I don't like what I wrote, but if someone says "how would you like it if someone told you your work wasn't good?" as a kind of "gotcha", boy, they really have no idea how I think about my work.

    There are various tricks I've used to get around this (not explicitly to get around this, but they do so as a side effect). As discussed in this old post on writing, for a while, I hired a professional editor and had a process goal of doing one pass on each post and then trying to improve the next post. And as noted in the postscript to recent posts, now I'm trying to write with extremely minimal cleanup and editing and push posts out in half an hour regardless of the state of the data I'm looking at or the post (which I'm generally failing to do; I thought I might succeed on this one because it doesn't have any data analysis, but someone made a comment on the draft post that got me to re-write the whole thing, and just on number of words in the post, half an hour would really be pushing it on the original and then it increased in length). Of course a post that's written as quickly as possible with little to no regard for cleaning things up is going to be terrible in all kinds of ways, so all flaws I see in the post don't stop me from publishing it. Have my recent posts been good? Of course not; for any of the experimental/data posts, I could probably name ten things that should be fixed about each of them off the top of my head. For this post, I'd have to re-read it to come up with ten things, but I'm sure if I did re-read it I'd want to re-write the whole thing because of the issues it has.

    [return]
  3. This was, inadvertently, a kind of revenge on my friend for when I tried to open his door for the first time to leave his place. Since the door clearly opened to the outside, I tried pushing on the door, which didn't work, so I checked if there was a latch that was stuck, if the door was still locked, if I needed to push harder, etc., none of which worked. When he saw that I couldn't open the door I asked him what the trick was he said, in a tone of voice that made it sound like this was obviously something everyone should know, you need to pull the door before pushing it. The door was wedged such that the easiest way to open the door was to pull the door as tightly shut as possible and then immediately shove the door open. This friend, since he grew up in that house, thought this was obvious, apparently not realizing that it's not normal to have to try to close a door extra hard to open it. [return]
  4. a response I've heard to this kind of thing recently is that Anthropic had the best growth numbers in history while Claude was very buggy. If you have the best coding model and agent in the world, you can get away with a lot, but even they seem to have spent a fair amount of effort improving quality.

    Maybe you can also get away with it if you have a product that succeeds due to bundling, the strength of your enterprise sales team, network effects, monopoly power, etc.; all but one of the cases I'm thinking of are places where the team didn't have these things on their side. I actually thought the one other case I was thinking of would be something like Blackboard, but (if the Google results are accurate) I see that the software has declined from being #1 in the market to being a minority player, so maybe they couldn't get away with it either (I didn't look into the reasons for the decline; perhaps it's a coincidence).

    As noted above, Blackboard is an example where you could argue that the software quality didn't matter and people might as well just believe whatever makes them happy; if thinking that users love the software, then why not think that? But most of the rest of the examples that come to mind for me aren't cases like that. I don't think this is the best example, but it comes to mind because the comment below is the last time I was reminded of the Blackboard example. There was a comment from a Tumblr employee who said that they'd solved the moderation (abuse / spam / toxicity / etc.) problem mechanically at Tumblr via the way reblogs worked and that the mechanics Tumblr provided to users were good enough that the community could self-police bad behavior and that other social media sites would do well to learn from Tumblr. This was referring to Tumblr back in its heyday (maybe 2009-2014). I never really read much on Tumblr so I don't personally have an opinion, but back when it was a major social media platform, the reputation among folks I know was that it was heavy on bad behavior, particularly pile-ons caused by people taking out of context quotes and turning them into ragebait (not to say this doesn't happen on other platforms, but the belief was that the way Tumblr was structured and/or the communities involved made this worse on Tumblr). I'm not sure I know anyone who used Tumblr at the time who would say that the community was good at self-policing. In fact, when Scott Alexander wrote one of his most famous pieces, Toxoplasma Of Rage, he dedicated an entire section to how Tumblr's reblog system is particularly bad and is guaranteed to result in bad behavior. He actually says that whoever designed the system either didn't understand what they were doing or they understood all too well and deliberately made the most ragebait-inducing system possible. This was written during the time when this employee said that Tumblr had solved the moderation problem and uses examples from that time.

    Moderation at scale is an impossibly hard problem, so as a non-Tumblr user, I'm not even sure that Tumblr did worse than other platforms given its size and growth rate, but I think you'd need some quality blindness to think that Tumblr had solved the moderation problem. I think the strongest positive case you could plausibly make would be something like "Tumblr was better than average, but many people had a worse than average experience due to the communities they were in and some of these communities were unusually widely read and Tumblr therefore unfairly gained a reputation as being a particularly bad platform". I don't know if that's true or not, but it doesn't seem impossible that it could be true; it does seem impossible that Tumblr solved the moderation problem.

    [return]
  5. most of my projects are deliberately low quality; what I try to do is do the highest ROI testing, not test to the point the quality is what I would actually consider good, and this also goes for things like making interfaces very nice, etc. [return]

Friday, 28 August 2026

Riding to the MADE Show [Rene Herse Cycles] (10:56 , Friday, 28 August 2026)

Every year, the MADE Show in Portland showcases handmade bikes. It’s a great event, not just to see beautiful bikes, but also to meet friends and make new ones. The MADE Show is always worth a visit—and, this year, it provided a perfect excuse for a great adventure.

I can’t think of a better way to celebrate handmade bikes than to ride one to the MADE Show. Routes from Seattle to Portland via the Cascade Mountains have fascinated me since we discovered the joys of exploring long-forgotten gravel passes. For the ride to MADE, I wanted to try a new version of this theme, hugging the flanks of Loowit (Mount St. Helens) by means of the Ape Canyon Trail, a hiking trail that is also open to mountain bikes.

I met Becca Book during Unbound XL, when we rode together for a while around Mile 275, before she pulled away to finish on the women’s podium. We’re both from Seattle, so we’d been talking about riding together, and, somehow, the idea came up to ride to MADE together. That’s how we met on Friday evening and headed south.

Our bikes provided an interesting contrast: Becca was on her Otso carbon gravel bike, complete with front suspension and 55 mm Fleecer Ridge knobbies. I rode my ‘Oregon Outback’ Rene Herse: lugged steel frame and smooth 54 mm Rat Trap Pass tires. The forecast included a chance of rain, so my bike wore full fenders.

What followed was a great adventure. We’d planned to arrive at Windy Ridge, the highest paved road on the big volcano, by sunrise. We actually were 20 minutes ahead of schedule after riding all night. The mountain cleared for a brief glimpse before fog rolled in again.

The otherworldly landscape high up on the volcano was even more mesmerizing in the fog. We traversed the Plains of Abraham…

…before diving into the old-growth forest of the Ape Canyon. We climbed Old Man Pass, descended Panther Creek Road, before crossing Bridge of the Gods to traverse the mighty Columbia River. (I love all those names!)

Both bikes had performed equally well. Becca’s front suspension (and her mountain bike skills) made her a little faster on the really rough gravel descents, while my bike had the edge on the (paved) twists and turns of Panther Creek Road. For most of the ride, our speeds were well-matched. We arrived in Portland in time for dinner after 250 miles (400 km).

The next day, we visited the MADE Show. Portland’s vibrant bike scene was on full display, and I ran into so many friends and acquaintances that it was hard to take in all the wonderful bikes.

Photography was difficult in the dark former shipyard building… Here is one of my favorites: This Ahearne combined a frame with polished lugs, a Nivex derailleur—and disc brakes. It all looked great together and, like all the best bikes, it made me want to take it outside and ride it.

Instead, I retrieved my own Nivex-equipped bike and headed to the train station. I ran into more friends on the way out, and when I finally got going, I was hopelessly behind schedule. Would I miss my train? Time to pedal with everything I had: If there’s a Strava KOM from Zidell Yards to Portland’s Union Station, I may be in contention!

I made the train… After a quick spin home from the train station, I was back at dinner on Sunday evening. What a fun weekend!

More Information:

Friday Debrief: Stooge Hotdogger, Toothpaste Docker, Reissued Jamis Komodo, and More… [BIKEPACKING.com] (08:48 , Friday, 28 August 2026)

DebriefThis week’s Debrief features camp gizmos and the toothpaste docker, the reissued Jamis Komodo, a Stooge hotdogger, a Thomson/Blue Lug collab, multiple events to follow live, and much more. Find it all here…

The post Friday Debrief: Stooge Hotdogger, Toothpaste Docker, Reissued Jamis Komodo, and More… appeared first on BIKEPACKING.com.

MADE 2026 on Video (Part 7): Weird Racks and Denim Tires [BIKEPACKING.com] (08:28 , Friday, 28 August 2026)

made 2026 video part 7In our seventh installment of video coverage from the 2026 MADE Bike Show, Neil catches up with several brands to check out their new parts. That includes two shiny Mica racks, new Teravail stems and "blue jean" tires, Silca's new tire lineup, and much more…

The post MADE 2026 on Video (Part 7): Weird Racks and Denim Tires appeared first on BIKEPACKING.com.

The Rodeo Labs Spork 4.0 is Finally Here [BIKEPACKING.com] (08:23 , Friday, 28 August 2026)

Rodeo Labs Spork 4.0The long wait for the fourth version of Rodeo Labs' carbon gravel fork is finally over. Launched late yesterday and available now, the Spork 4.0 has a host of new features, making it the brand's most capable adventure fork yet. More on the Rodeo Labs Spork 4.0 here...

The post The Rodeo Labs Spork 4.0 is Finally Here appeared first on BIKEPACKING.com.

Friday, 21 August 2026

Virtual Elevation Testing with the Chung Method: Problems and Promise [Rene Herse Cycles] (10:08 , Friday, 21 August 2026)

Summary: The Virtual Elevation or Chung Method is based on physically sound concepts. In practice, however, tests using this method have produced unreliable results. Attempts at validating the method have failed to reproduce the claimed precision. A possible explanation: Testers ‘guess’ initial values for rolling resistance (Crr) and aero drag (CdA) as they match measured and predicted data, rather than fitting the curves blind. Since testers will start with what they consider ‘likely’ values for Crr and CdA, this may cause bias confirmation: It may automatically confirm the hypothesis, whether it is actually correct or not. Separating data collection from analysis and doing the analysis blind might help resolve these issues and make the method more reliable. In its current form, the Virtual Elevation Method cannot be trusted to produce reliable results.

If you’ve been following bike tech over the last decade or so, you’ve probably heard about the Virtual Elevation Method, also named the Chung Method. This technique allows testing bike performance, especially aerodynamics, without needing a wind tunnel. More recently, it’s also been used to test the rolling resistance of tires. Over the years, we’ve had discussions on-and-off with the method’s creators and proponents, as well as with other researchers who’ve raised questions about its reliability. Recently, we’ve had the opportunity to take a closer look at its potential and pitfalls.

A little while ago, we wrote in a discussion of the new 32-inch wheels that “there is a total absence of data to back up claims of superior performance of larger wheels” (for road, all-road and gravel bikes). Not everybody agreed with this assessment. There is a lot of excitement in the industry right now about the new wheel size, and there’s a scramble to back this up with data. John Karrasch, who has been at the forefront of testing the new 32-inch tires, wrote: “The only problems with my data are that they don’t fit your mindset.” His results show the 32-inch tires rolling much, much faster. He’s been vocal about his disagreement with our findings, in private conversations and on social media.

In the bike industry, it’s common to ignore data that doesn’t fit one’s ideas. An example: We’ve known for 20 years that wide, supple tires roll as fast as narrow rubber. And yet the mainstream media pretended for many years that all the data supporting this simply didn’t exist… But that’s not how science works. If there is contradictory data, everybody works together to resolve the issues and figure out what is really going on. It’s not about being right or wrong: If 32-inch wheels roll faster, we all want to know. I certainly do. I’d love to have faster wheels for my next race or FKT attempt! And as a company, Rene Herse Cycles already offers our ultra-fast TPU tubes for 32″ wheels. It would be no problem to add supple 32″ tires to our program.

As a side note, John has also tested the rolling resistance of our Snoqualmie Pass tires, and found them to be among the faster tires he’s tested. The goal here is not to discredit results we don’t ‘like,’ but to evaluate a new testing method—and figure out how to make bicycles faster in the real world.

In response to our article about 32″ wheels, John Karrasch sent his latest data (above). John is a smart guy, and he’s conscientious with his testing. As mentioned above, the goal here is to figure out what’s really going on with 32″ tires. John pointed out that his data shows very significant performance benefits for 32″ wheels—on all surfaces. The graphs on the right (32″) show savings of between 21 and 10 watts over those on the left (29″).

Perhaps most striking is the extra speed on pavement (blue bars and line). The 32-inch tires to consume 15% less power on smooth pavement: 39 watts vs. 46 watts. (These are watts attributed to rolling resistance, not total watts required to power the bike.) John’s tests of other 32″ tires show similar savings on smooth pavement.

A savings of 15% is huge—especially since 32″ wheels are just 7.3% larger than 29″.

I mentioned to John that this represents an unexpectedly large performance benefit—coming from wheels that are just slightly larger. His reply: “No shit. And it’s across THREE DIFFERENT TIRE MODELS.”

It’s hard to think of a mechanism that reduces the rolling resistance of slightly larger wheels by this much on smooth pavement. Roll-over isn’t a factor on smooth surfaces. The contact patch shape isn’t all that different, considering that wheel sizes are just 7.3% different. When I asked about possible explanations for this surprising result, John replied: “I don’t know everything and don’t waste my time guessing.” He continued: “You gotta get past the percentage improvement on pavement. The pavement results are the least important thing.”

I understand John’s point: Everybody wants to know how much faster 32″ wheels are on rough surfaces: gravel, cobblestones, singletrack. Why should we focus on the smooth pavement data? The counterpoint: The smooth pavement data provides a useful test for the methodology. If the data for smooth pavement are incorrect, then it’s likely that the other results also have problems.

Perhaps these are two different schools of thought. One is best summed up with Carl Sagan’s mantra: “Extraordinary claims require extraordinary evidence.” The other might be paraphrased as: “Extraordinary findings are just extraordinary! No need to explain them.” Or as one article put it: “32-inch tires are way faster on most terrain—even where you wouldn’t expect it!”

Here at Rene Herse Cycles, we have made plenty of extraordinary claims ourselves. Back in 2006, the results of our first real-world tire performance study were unexpected: High pressure doesn’t make supple tires roll faster. Wide tires can roll as fast as narrow rubber. That was very controversial. It went against what everybody believed—us included. That’s why we left no stone unturned to confirm, replicate and validate our results.

As scientists, that’s our speciality. My Ph.D. is in science. I spent half a decade learning how to design studies and also how to spot problems. My collaborator Mark VdK is even more of an expert in this field. He has a Ph.D. with a minor in applied mathematics, and he has worked for years as a senior research and data analyst at the world’s leading company for enterprise resource planning. Designing tests is his speciality. I only mention these credentials to head off criticism that we’re just a bunch of curmudgeons who don’t like change.

We’re serious about science because we’re serious about having fun on our bikes. Our bikes may not always look like those of mainstream racers, but they work extremely well. Often they are ahead of their time. Like when we raced Unbound XL in 2022 on 54 mm-wide tires. Back then, that was considered ‘too much tire’ even for the Flint Hills of Kansas. Today, tires that wide aren’t controversial any longer. We also ran narrow handlebars for better aerodynamics, way back when most gravel racers were on burly 44 cm bars. That’s another area where the mainstream has caught up with us.

If I may say so, we’ve got a good track record. The results of our testing have stood the test of time. It’s taken a while, but most of our initially ‘controversial’ results have been accepted by the mainstream.

That’s no coincidence—it’s the result of careful testing and relentlessly questioning our results. Here is what we did to make sure our findings about wide tires and low pressure were real:

  • We looked for an explanation: Suspension losses caused by vibrations were the likely reason why high pressure didn’t make tires faster. Previous drum tests (without a rider) didn’t measure suspension losses, so they missed how vibrations slow down the bike. (Thanks to Jim Papadopoulos for digging up a 1960s Army research study that first discovered suspension losses.)
  • We measured suspension losses in our famous rumble strip tests. This confirmed our hypothesis: Significant energy was lost to vibrations. That’s why wide, soft tires roll as fast as narrow, hard rubber—or faster, depending on how rough the surface is.
  • We replicated our roll-down tests with a different methodology: We rode around a track with a precision power meter (above). The results were the same with both methods.
  • In wind tunnel testing, we confirmed that our test riders are able to maintain the same position, time and again. That means we didn’t need to worry about changes in rider position adding noise to our measurements.
  • We had our research peer-reviewed by cycling science experts before we published our findings.
  • We started using wide tires to gain in-the-field experience: Our times in long-distance brevets and races improved on wide tires. Real-world practice matches the theory.

Since we first published our results 20 years ago, they’ve gone from controversial to widely accepted. Today, most mainstream makers and journalists agree with our findings. Pro racers have moved from 23 mm tires to 28 or 30 mm tires, even for smooth courses—and their speeds have gone up. Controlled studies by others, like the Escape Collective, have confirmed many of our findings. There’s really no longer any doubt. Wide tires are here to stay. And pressures of 100 psi (7 bar) and more are history. It’s no coincidence that Tadej Pogačar inflates his tires to the values recommended by the Rene Herse Tire Pressure Calculator. Those values are based on our research, and they really work on the road (and not just in the lab).

Back to 32″ wheels: John’s data shows a huge performance advantage for 32″ wheels on smooth pavement. No matter whether you believe that larger wheels have better ‘roll-over’ or not, you’d expect that advantage to be greatest on rough gravel and cobblestones. Instead, John’s data shows twice as much benefit of the larger wheels on smooth pavement than on the roughest surfaces he tested.

Here’s what our own tests show: On smooth pavement, there’s no performance benefits for bigger wheels. In the roll-down test (left), the smaller wheels rolled marginally faster, but the difference wasn’t statistically significant. With the rider pedaling (right), there’s a little more noise, but again large wheels do not require less power than the smaller ones. (The differences are once again not statistically significant.)

Large wheels offering no performance advantage is the opposite of what John’s data shows. To resolve this discrepancy, we need to look at the testing methodologies. Our data comes from both roll-down tests (left) and tests with power meters (right). Each time, we used the same tire model in different sizes. We tested on calm days, with constant temperatures, and we did a statistical analysis. Two different methods, two different tire models, same result: Wheel size doesn’t affect speed on smooth pavement.

And let’s not forget: These results are part of two decades of studying tire width and suspension losses. Our studies have been validated. They are now widely accepted. It’s fair to say that these studies are reliable. In other words: To challenge these results, simply presenting new data is not enough. You also need an explanation why the existing data is incorrect. Remember: All the other data from the same studies has been reliable—there’s no reason to doubt just this small part of our studies.

We don’t just have data, but also physics to explain these results: Pneumatic tires deform where they meet the road. The radius of the wheel doesn’t make a difference as long as surface irregularities are small enough to be absorbed by the tire. On smooth roads, that’s definitely the case.

Let’s look at John’s methodology, which shows 32″ wheels rolling so much faster on all surfaces, including smooth pavement? John has been testing tires using the Virtual Elevation Method. It’s a clever methodology, based on simple physics: Aerodynamic drag goes up exponentially with speed, while rolling resistance (and other mechanical resistances) go up linearly. What this means: At low speeds, a large portion of the rider’s power output goes toward rolling resistance. At high speeds, most of the rider’s watts push against wind resistance. If we measure power output at various speeds, we should be able to separate rolling resistance and wind resistance. To illustrate how this works, let’s look at two rider/bike scenarios:

  • Bike/rider 1: fast-rolling tires, upright position. Rider 1 will require very little power at low speeds. At high speeds, their power output will increase exponentially.
  • Bike/rider 2: slow tires, low aero position. Rider 2 will need more power at low speeds. To go faster, they’ll only need to increase power by a smaller percentage than Rider 1.

How does this test work in practice? For the Virtual Elevation Method, the bike is ridden on a circular course. The course should have some elevation gain, but coasting and braking should be avoided. The course is ridden at variable speeds. Power, speed and elevation are recorded. (John took the elevation profile off topography data, rather than relying on less-accurate GPS.) Estimates of wind resistance (CdA) and rolling resistance (Crr) are plugged into a spreadsheet until the ‘virtual elevation’ and actual elevation profiles match.

Above are the calculated elevation curve (blue) and the actual elevation of the course (black; the elevation data is in 1-meter steps). The accuracy is within ±1% (above). That’s incredible precision, considering this is real world testing on two bikes (carbon 29er, titanium 32″), at various power outputs.

John’s testing is not the only case where the Virtual Elevation Method results in implausibly high precision. Consider Tom Anhalt’s test results: Tom suspended small spheres (foam balls) on a stick to place them in undisturbed air while he rode (above). He was able to detect very small differences in the size of the balls: 2″ (5 cm); 3″ (7.5 cm) and 4″ (10 cm). According to his calculations, the 2″ ball increased the wind resistance of his bike/rider setup by 0.08%. Using the Virtual Elevation Method, Tom could detect that tiny difference (below). The larger balls increased drag by 0.6% and 1%, and he detected those, too.

As with John’s data, we have no doubt that Tom is conscientious in reporting what he measured. If there is a problem, it’s with the methodology. What’s remarkable isn’t just the ability to detect tiny differences, but also that the data is so ‘clean.’ The measured differences (center column) track the calculated differences (right column) almost exactly. Anybody who has done real-world testing knows that some ‘noise’ in the data is unavoidable. There are always very slight wind currents—even on perfectly calm days. Even a good rider will change their position ever-so-slightly. Temperature is not 100% constant, especially on a road course. Other variables can affect the results, at least a little bit.

Can the Virtual Elevation Method really filter out all this noise and get incredibly ‘clean’ and precise data? To quote Carl Sagan again: “Extraordinary claims require extraordinary evidence.” Basically, the Virtual Elevation Method needs validation using a different methodology. Just like we validated our roll-down tests with power meter tests…

To validate the method, Dyer & Disley compared wind tunnel tests with results from the Virtual Elevation Method. For the Virtual Elevation Method, they tested on an indoor velodrome. With no wind, constant temperature and a very uniform surface, conditions were about as perfect as possible. They also knew the Coefficient of Rolling Resistance (Crr) from previous testing, so they only had to ‘guess’ the wind resistance (CdA) when they matched the curves.

They found a coefficient of variation of 0.7-0.9% in the wind tunnel, and 2-3% with the Virtual Elevation Method. They tested two different shapes of prostheses for amputee cyclists. One was round, the other streamlined. They could detect a statistically significant difference in the wind tunnel, but not using the Virtual Elevation Method.

What this means: Even under perfect conditions and with Crr already known, Dyer & Disley were unable to get anywhere close to the precision that John and Tom found in their testing with the Virtual Elevation Method. Even in the controlled setting of the wind tunnel, where the rider pedals at constant power and cadence, small changes in rider position are inevitable. The result is ‘noise’ that makes it impossible to detect very small changes like those balls that Tom attached to his bike. It’s also important to remember: The Virtual Elevation Method can separate rolling and wind resistance, but there is no mechanism to separate rider-induced changes in wind resistance from those caused by changes to the bike. There’s no way to tell whether the rider has lifted their head slightly, or whether there’s a slightly larger ball attached to the bike.

Why are Tom and John getting so much better results in outdoor testing than Dyer & Disley in their carefully controlled indoor testing? And why, on the other hand, does the Virtual Elevation Method sometimes produce strange results—such as 32″ wheels offering such a large advantage on pavement than on rough gravel or cobblestones (above)? You would expect the opposite: Large wheels should offer the greatest advantage on the roughest surfaces. In other words, why is the Virtual Elevation Method sometimes so incredibly precise, but in the same test sessions also provides results that clearly don’t make sense?

Mark VdK—who develops testing methods for a living—identified a potential problem when fitting the actual and virtual elevation curves to each other. The tester guesses aero drag (CdA) and rolling resistance (Crr), and then checks whether the curves work with those guesses—until the curves match. The fitting of the data to the curves is not done blind, but based on the tester’s hypothesis. If the tester thinks that 32″ wheels roll faster, they’ll start their ‘curve fitting’ by plugging in lower rolling resistance (Crr) for the bigger wheels. Since the tester ‘guesses’ two variables, multiple curves will fit the data. The tester starts with ‘reasonable’ guesses for Crr and CdA—which means the self-selected curve tends to be the one that confirms the hypothesis. That may be the reason the data looks so consistent—more consistent than any direct real-world measurements.

In other words, the Virtual Elevation Method may inadvertently create a form of circular reasoning: When testing balls attached to your bike, it’s natural to ‘guess’ that the wind resistance increases slightly each time you make the ball bigger. You’ll start with a curve that confirms the hypothesis, and it will match your data with a little tweaking—which then seems to confirm that the method is working. It’s the same when testing 32-inch wheels. If the tester starts their ‘guesses’ with a lower rolling resistance for the larger wheels, they’ll probably find a curve that fits—which is what they expected in the first place. It would be the same if you tested different tires: The tires that you ‘guess’ are fast will—most likely—turn out to be fast in your testing.

In science, this is called ‘bias confirmation,’ and it’s considered a huge problem—exactly because the results match what everybody expects. (Who is going to question something that seems to make sense, like larger wheels rolling faster?) In fact, much of the scientific method is designed to avoid this.

Often, the bias becomes obvious when we look at results that were not part of the original hypothesis—like the behavior of the 32″ wheels on smooth pavement in John’s testing. In John’s own words, he wasn’t too concerned about the results on smooth pavement. If the testing method was reliable, you’d expect the results to make sense nonetheless. However, they don’t make sense—which strongly suggests that there’s a problem with the methodology. And that problem extends to all results, not just those that don’t make sense.

One way to improve the Virtual Elevation Method would be to make it blind. Double-blind testing is the gold standard for medical studies: Neither patient nor test administrator know whether the patient is taking the medication that’s being tested or a placebo. If you suggested doing medical research that isn’t blind, you’d be laughed out of the room.

When we tested the influence of frame stiffness on performance, we did a double-blind test (above): Only the framebuilder, Jeff Lyon, knew the differences in frame tubing between the test bikes. The bikes were built up with identical components. Their weights were equalized. For each test session, the test administrator marked the bikes with red, green and pink stem caps, but even the administrator only knew the bikes as #1, #2 and #3—and not their actual specs.

The riders only saw the colors of the stem caps—and those were switched from one test session to the next. In other words, riders had no idea which bike they were riding—not even whether it was the same bike they’d ridden during the previous test session or a different one.

Riders were not allowed to talk about the bikes or compare notes. (That’s why they sit so far apart in the photo above.) Once the testing was complete and the notes were finalized, the test was ‘unblinded’: Now the framebuilder shared which bike was made from which tubing. And the administrator shared which bike had which color stem cap during each test session. Only now did the testers find out which bikes they had been riding in each test session, but they no longer could change what they wrote about them. As a result, their impressions of each bike were free of bias. And the power measurements also weren’t influenced by riders thinking that one bike might be faster than the other. (The riders were not able to see the power numbers during the tests.)

Double-blind testing is not possible with bikes of different wheel sizes, but the testing should at least be blind: Have one tester collect the data, and another do the analysis. Will the results be the same without knowing which dataset is from 32″ wheels and which is from 29″ wheels? Rather than guessing ‘likely’ values for rolling resistance (Crr) and aerodynamic drag (CdA), the data analysis would test a wide range of values and determine which offers the best fit. This blind curve matching would eliminate the problem of starting with values that match the hypothesis.

What would happen if the curve fitting was done blind? My prediction is that the erroneous results—such as the huge benefit of 32″ wheels on smooth pavement—would disappear. However, at the same time, more ‘noise’ would be introduced. Dyer & Disley’s study suggests that the real-world precision of the Virtual Elevation Method, under ideal conditions, is about 2-3%—not even close to the 0.1-1% reported by John and Tom Anhalt. Once you test outside, rather than in an indoor velodrome, there will be even more noise.

None of this is intended as criticism of John or Tom’s work. Both are very conscientious in their testing—it’s the methodology that appears to have problems. As I mentioned in the beginning, the goal is not to prove who is right or wrong, but to figure out how bicycles really work. Despite the obvious issues with the Virtual Elevation Method—both the ‘incredible’ precision and the erroneous results—the underlying physics are sound. The method holds promise. The question is this: Once the method is improved with blind fitting of the curves, will we still get useful results? Or will there be too much noise? There is only one way to find out: Do the experiment described above, or validate the method in some other way. Until then, data collected using the Virtual Elevation Method should not be used as evidence that one setup performs better than another.

As to 32″ wheels, hopefully somebody will soon do a carefully controlled test with direct measurements, like our rumble strip tests (above) or the Escape Collective’s tire tests. If those tests show a big advantage for 32″ wheels, we’ll have moved one step further.

However, even that will not be enough: We’ll still need to find out why there’s no significant performance difference between 26″, 650B and 700C/29″ wheels in Bicycle Quarterly’s testing, and why the next step up to 32″ wheels suddenly reduces rolling resistance by a large margin. Because simply ignoring previous research is not how science works, especially if that previous research has been validated many times.

That validation is missing from the Virtual Elevation Method. As mentioned above, we’re offering to help with this validation, because a new method for testing bicycles under real-world conditions benefits everybody.

Further Reading:

Virtual Elevation Testing with the Chung Method: Problems and Promise [Rene Herse Cycles] (10:08 , Friday, 21 August 2026)

Summary: The Virtual Elevation or Chung Method is based on physically sound concepts. In practice, however, tests using this method have produced unreliable results. Attempts at validating the method have failed to reproduce the claimed precision. A possible explanation: Testers ‘guess’ initial values for rolling resistance (Crr) and aero drag (CdA) as they match measured and predicted data, rather than fitting the curves blind. Since testers will start with what they consider ‘likely’ values for Crr and CdA, this may cause bias confirmation: It may automatically confirm the hypothesis, whether it is actually correct or not. Separating data collection from analysis and doing the analysis blind might help resolve these issues and make the method more reliable. In its current form, the Virtual Elevation Method cannot be trusted to produce reliable results.

If you’ve been following bike tech over the last decade or so, you’ve probably heard about the Virtual Elevation Method, also named the Chung Method. This technique allows testing bike performance, especially aerodynamics, without needing a wind tunnel. More recently, it’s also been used to test the rolling resistance of tires. Over the years, we’ve had discussions on-and-off with the method’s creators and proponents, as well as with other researchers who’ve raised questions about its reliability. Recently, we’ve had the opportunity to take a closer look at its potential and pitfalls.

A little while ago, we wrote in a discussion of the new 32-inch wheels that “there is a total absence of data to back up claims of superior performance of larger wheels” (for road, all-road and gravel bikes). Not everybody agreed with this assessment. There is a lot of excitement in the industry right now about the new wheel size, and there’s a scramble to back this up with data. John Karrasch, who has been at the forefront of testing the new 32-inch tires, wrote: “The only problems with my data are that they don’t fit your mindset.” His results show the 32-inch tires rolling much, much faster. He’s been vocal about his disagreement with our findings, in private conversations and on social media.

In the bike industry, it’s common to ignore data that doesn’t fit one’s ideas. An example: We’ve known for 20 years that wide, supple tires roll as fast as narrow rubber. And yet the mainstream media pretended for many years that all the data supporting this simply didn’t exist… But that’s not how science works. If there is contradictory data, everybody works together to resolve the issues and figure out what is really going on. It’s not about being right or wrong: If 32-inch wheels roll faster, we all want to know. I certainly do. I’d love to have faster wheels for my next race or FKT attempt! And as a company, Rene Herse Cycles already offers our ultra-fast TPU tubes for 32″ wheels. It would be no problem to add supple 32″ tires to our program.

As a side note, John has also tested the rolling resistance of our Snoqualmie Pass tires, and found them to be among the faster tires he’s tested. The goal here is not to discredit results we don’t ‘like,’ but to evaluate a new testing method—and figure out how to make bicycles faster in the real world.

In response to our article about 32″ wheels, John Karrasch sent his latest data (above). John is a smart guy, and he’s conscientious with his testing. As mentioned above, the goal here is to figure out what’s really going on with 32″ tires. John pointed out that his data shows very significant performance benefits for 32″ wheels—on all surfaces. The graphs on the right (32″) show savings of between 21 and 10 watts over those on the left (29″).

Perhaps most striking is the extra speed on pavement (blue bars and line). The 32-inch tires to consume 15% less power on smooth pavement: 39 watts vs. 46 watts. (These are watts attributed to rolling resistance, not total watts required to power the bike.) John’s tests of other 32″ tires show similar savings on smooth pavement.

A savings of 15% is huge—especially since 32″ wheels are just 7.3% larger than 29″.

I mentioned to John that this represents an unexpectedly large performance benefit—coming from wheels that are just slightly larger. His reply: “No shit. And it’s across THREE DIFFERENT TIRE MODELS.”

It’s hard to think of a mechanism that reduces the rolling resistance of slightly larger wheels by this much on smooth pavement. Roll-over isn’t a factor on smooth surfaces. The contact patch shape isn’t all that different, considering that wheel sizes are just 7.3% different. When I asked about possible explanations for this surprising result, John replied: “I don’t know everything and don’t waste my time guessing.” He continued: “You gotta get past the percentage improvement on pavement. The pavement results are the least important thing.”

I understand John’s point: Everybody wants to know how much faster 32″ wheels are on rough surfaces: gravel, cobblestones, singletrack. Why should we focus on the smooth pavement data? The counterpoint: The smooth pavement data provides a useful test for the methodology. If the data for smooth pavement are incorrect, then it’s likely that the other results also have problems.

Perhaps these are two different schools of thought. One is best summed up with Carl Sagan’s mantra: “Extraordinary claims require extraordinary evidence.” The other might be paraphrased as: “Extraordinary findings are just extraordinary! No need to explain them.” Or as one article put it: “32-inch tires are way faster on most terrain—even where you wouldn’t expect it!”

Here at Rene Herse Cycles, we have made plenty of extraordinary claims ourselves. Back in 2006, the results of our first real-world tire performance study were unexpected: High pressure doesn’t make supple tires roll faster. Wide tires can roll as fast as narrow rubber. That was very controversial. It went against what everybody believed—us included. That’s why we left no stone unturned to confirm, replicate and validate our results.

As scientists, that’s our speciality. My Ph.D. is in science. I spent half a decade learning how to design studies and also how to spot problems. My collaborator Mark VdK is even more of an expert in this field. He has a Ph.D. with a minor in applied mathematics, and he has worked for years as a senior research and data analyst at the world’s leading company for enterprise resource planning. Designing tests is his speciality. I only mention these credentials to head off criticism that we’re just a bunch of curmudgeons who don’t like change.

We’re serious about science because we’re serious about having fun on our bikes. Our bikes may not always look like those of mainstream racers, but they work extremely well. Often they are ahead of their time. Like when we raced Unbound XL in 2022 on 54 mm-wide tires. Back then, that was considered ‘too much tire’ even for the Flint Hills of Kansas. Today, tires that wide aren’t controversial any longer. We also ran narrow handlebars for better aerodynamics, way back when most gravel racers were on burly 44 cm bars. That’s another area where the mainstream has caught up with us.

If I may say so, we’ve got a good track record. The results of our testing have stood the test of time. It’s taken a while, but most of our initially ‘controversial’ results have been accepted by the mainstream.

That’s no coincidence—it’s the result of careful testing and relentlessly questioning our results. Here is what we did to make sure our findings about wide tires and low pressure were real:

  • We looked for an explanation: Suspension losses caused by vibrations were the likely reason why high pressure didn’t make tires faster. Previous drum tests (without a rider) didn’t measure suspension losses, so they missed how vibrations slow down the bike. (Thanks to Jim Papadopoulos for digging up a 1960s Army research study that first discovered suspension losses.)
  • We measured suspension losses in our famous rumble strip tests. This confirmed our hypothesis: Significant energy was lost to vibrations. That’s why wide, soft tires roll as fast as narrow, hard rubber—or faster, depending on how rough the surface is.
  • We replicated our roll-down tests with a different methodology: We rode around a track with a precision power meter (above). The results were the same with both methods.
  • In wind tunnel testing, we confirmed that our test riders are able to maintain the same position, time and again. That means we didn’t need to worry about changes in rider position adding noise to our measurements.
  • We had our research peer-reviewed by cycling science experts before we published our findings.
  • We started using wide tires to gain in-the-field experience: Our times in long-distance brevets and races improved on wide tires. Real-world practice matches the theory.

Since we first published our results 20 years ago, they’ve gone from controversial to widely accepted. Today, most mainstream makers and journalists agree with our findings. Pro racers have moved from 23 mm tires to 28 or 30 mm tires, even for smooth courses—and their speeds have gone up. Controlled studies by others, like the Escape Collective, have confirmed many of our findings. There’s really no longer any doubt. Wide tires are here to stay. And pressures of 100 psi (7 bar) and more are history. It’s no coincidence that Tadej Pogačar inflates his tires to the values recommended by the Rene Herse Tire Pressure Calculator. Those values are based on our research, and they really work on the road (and not just in the lab).

Back to 32″ wheels: John’s data shows a huge performance advantage for 32″ wheels on smooth pavement. No matter whether you believe that larger wheels have better ‘roll-over’ or not, you’d expect that advantage to be greatest on rough gravel and cobblestones. Instead, John’s data shows twice as much benefit of the larger wheels on smooth pavement than on the roughest surfaces he tested.

Here’s what our own tests show: On smooth pavement, there’s no performance benefits for bigger wheels. In the roll-down test (left), the smaller wheels rolled marginally faster, but the difference wasn’t statistically significant. With the rider pedaling (right), there’s a little more noise, but again large wheels do not require less power than the smaller ones. (The differences are once again not statistically significant.)

Large wheels offering no performance advantage is the opposite of what John’s data shows. To resolve this discrepancy, we need to look at the testing methodologies. Our data comes from both roll-down tests (left) and tests with power meters (right). Each time, we used the same tire model in different sizes. We tested on calm days, with constant temperatures, and we did a statistical analysis. Two different methods, two different tire models, same result: Wheel size doesn’t affect speed on smooth pavement.

And let’s not forget: These results are part of two decades of studying tire width and suspension losses. Our studies have been validated. They are now widely accepted. It’s fair to say that these studies are reliable. In other words: To challenge these results, simply presenting new data is not enough. You also need an explanation why the existing data is incorrect. Remember: All the other data from the same studies has been reliable—there’s no reason to doubt just this small part of our studies.

We don’t just have data, but also physics to explain these results: Pneumatic tires deform where they meet the road. The radius of the wheel doesn’t make a difference as long as surface irregularities are small enough to be absorbed by the tire. On smooth roads, that’s definitely the case.

Let’s look at John’s methodology, which shows 32″ wheels rolling so much faster on all surfaces, including smooth pavement? John has been testing tires using the Virtual Elevation Method. It’s a clever methodology, based on simple physics: Aerodynamic drag goes up exponentially with speed, while rolling resistance (and other mechanical resistances) go up linearly. What this means: At low speeds, a large portion of the rider’s power output goes toward rolling resistance. At high speeds, most of the rider’s watts push against wind resistance. If we measure power output at various speeds, we should be able to separate rolling resistance and wind resistance. To illustrate how this works, let’s look at two rider/bike scenarios:

  • Bike/rider 1: fast-rolling tires, upright position. Rider 1 will require very little power at low speeds. At high speeds, their power output will increase exponentially.
  • Bike/rider 2: slow tires, low aero position. Rider 2 will need more power at low speeds. To go faster, they’ll only need to increase power by a smaller percentage than Rider 1.

How does this test work in practice? For the Virtual Elevation Method, the bike is ridden on a circular course. The course should have some elevation gain, but coasting and braking should be avoided. The course is ridden at variable speeds. Power, speed and elevation are recorded. (John took the elevation profile off topography data, rather than relying on less-accurate GPS.) Estimates of wind resistance (CdA) and rolling resistance (Crr) are plugged into a spreadsheet until the ‘virtual elevation’ and actual elevation profiles match.

Above are the calculated elevation curve (blue) and the actual elevation of the course (black; the elevation data is in 1-meter steps). The accuracy is within ±1% (above). That’s incredible precision, considering this is real world testing on two bikes (carbon 29er, titanium 32″), at various power outputs.

John’s testing is not the only case where the Virtual Elevation Method results in implausibly high precision. Consider Tom Anhalt’s test results: Tom suspended small spheres (foam balls) on a stick to place them in undisturbed air while he rode (above).

Tom tested four setups: no foam ball and three different foam balls that measured 2″, 3″ and 4″ in diameter (in metric, that’s 5 cm, 7.5 cm, 10 cm.) To visualize this, I bought the same size foam balls. In the photo above, you see the 2″ and 3″ balls compared to a standard water bottle. The differences are pretty small. It would be amazing if one could detect them with a rider on the bike whose pedaling motions introduce all kinds of ‘noise.’

Using the Virtual Elevation Method, Tom easily detected the changes in wind resistance, as he attached the different balls to his bike. According to his calculations, the 2″ ball increased the wind resistance of his bike/rider setup by 0.08%. Using the Virtual Elevation Method, Tom could detect that tiny difference (below). The larger balls increased drag by 0.6% and 1%, and he detected those, too.

As with John’s data, we have no doubt that Tom is conscientious in reporting what he measured. If there is a problem, it’s with the methodology. What’s remarkable isn’t just the ability to detect tiny differences, but also that the data is so ‘clean.’ The measured differences (center column) track the calculated differences (right column) almost exactly. Anybody who has done real-world testing knows that some ‘noise’ in the data is unavoidable. There are always very slight wind currents—even on perfectly calm days. Even a good rider will change their position ever-so-slightly. Temperature is not 100% constant, especially on a road course. Other variables can affect the results, at least a little bit.

Can the Virtual Elevation Method really filter out all this noise and get incredibly ‘clean’ and precise data? To quote Carl Sagan again: “Extraordinary claims require extraordinary evidence.” Basically, the Virtual Elevation Method needs validation using a different methodology. Just like we validated our roll-down tests with power meter tests…

To validate the method, Dyer & Disley compared wind tunnel tests with results from the Virtual Elevation Method. For the Virtual Elevation Method, they tested on an indoor velodrome. With no wind, constant temperature and a very uniform surface, conditions were about as perfect as possible. They also knew the Coefficient of Rolling Resistance (Crr) from previous testing, so they only had to ‘guess’ the wind resistance (CdA) when they matched the curves.

They found a coefficient of variation of 0.7-0.9% in the wind tunnel, and 2-3% with the Virtual Elevation Method. They tested two different shapes of prostheses for amputee cyclists. One was round, the other streamlined. They could detect a statistically significant difference in the wind tunnel, but not using the Virtual Elevation Method.

What this means: Even under perfect conditions and with Crr already known, Dyer & Disley were unable to get anywhere close to the precision that John and Tom found in their testing with the Virtual Elevation Method. Even in the controlled setting of the wind tunnel, where the rider pedals at constant power and cadence, small changes in rider position are inevitable. The result is ‘noise’ that makes it impossible to detect very small changes like those balls that Tom attached to his bike. It’s also important to remember: The Virtual Elevation Method can separate rolling and wind resistance, but there is no mechanism to separate rider-induced changes in wind resistance from those caused by changes to the bike. There’s no way to tell whether the rider has lifted their head slightly, or whether there’s a slightly larger ball attached to the bike.

Why are Tom and John getting so much better results in outdoor testing than Dyer & Disley in their carefully controlled indoor testing? And why, on the other hand, does the Virtual Elevation Method sometimes produce strange results—such as 32″ wheels offering such a large advantage on pavement than on rough gravel or cobblestones (above)? You would expect the opposite: Large wheels should offer the greatest advantage on the roughest surfaces. In other words, why is the Virtual Elevation Method sometimes so incredibly precise, but in the same test sessions also provides results that clearly don’t make sense?

Mark VdK—who develops testing methods for a living—identified a potential problem when fitting the actual and virtual elevation curves to each other. The tester guesses aero drag (CdA) and rolling resistance (Crr), and then checks whether the curves work with those guesses—until the curves match. The fitting of the data to the curves is not done blind, but based on the tester’s hypothesis. If the tester thinks that 32″ wheels roll faster, they’ll start their ‘curve fitting’ by plugging in lower rolling resistance (Crr) for the bigger wheels. Since the tester ‘guesses’ two variables, multiple curves will fit the data. The tester starts with ‘reasonable’ guesses for Crr and CdA—which means the self-selected curve tends to be the one that confirms the hypothesis. That may be the reason the data looks so consistent—more consistent than any direct real-world measurements.

In other words, the Virtual Elevation Method may inadvertently create a form of circular reasoning: When testing balls attached to your bike, it’s natural to ‘guess’ that the wind resistance increases slightly each time you make the ball bigger. You’ll start with a curve that confirms the hypothesis, and it will match your data with a little tweaking—which then seems to confirm that the method is working. It’s the same when testing 32-inch wheels. If the tester starts their ‘guesses’ with a lower rolling resistance for the larger wheels, they’ll probably find a curve that fits—which is what they expected in the first place. It would be the same if you tested different tires: The tires that you ‘guess’ are fast will—most likely—turn out to be fast in your testing.

In science, this is called ‘bias confirmation,’ and it’s considered a huge problem—exactly because the results match what everybody expects. (Who is going to question something that seems to make sense, like larger wheels rolling faster?) In fact, much of the scientific method is designed to avoid this.

Often, the bias becomes obvious when we look at results that were not part of the original hypothesis—like the behavior of the 32″ wheels on smooth pavement in John’s testing. In John’s own words, he wasn’t too concerned about the results on smooth pavement. If the testing method was reliable, you’d expect the results to make sense nonetheless. However, they don’t make sense—which strongly suggests that there’s a problem with the methodology. And that problem extends to all results, not just those that don’t make sense.

One way to improve the Virtual Elevation Method would be to make it blind. Double-blind testing is the gold standard for medical studies: Neither patient nor test administrator know whether the patient is taking the medication that’s being tested or a placebo. If you suggested doing medical research that isn’t blind, you’d be laughed out of the room.

When we tested the influence of frame stiffness on performance, we did a double-blind test (above): Only the framebuilder, Jeff Lyon, knew the differences in frame tubing between the test bikes. The bikes were built up with identical components. Their weights were equalized. For each test session, the test administrator marked the bikes with red, green and pink stem caps, but even the administrator only knew the bikes as #1, #2 and #3—and not their actual specs.

The riders only saw the colors of the stem caps—and those were switched from one test session to the next. In other words, riders had no idea which bike they were riding—not even whether it was the same bike they’d ridden during the previous test session or a different one.

Riders were not allowed to talk about the bikes or compare notes. (That’s why they sit so far apart in the photo above.) Once the testing was complete and the notes were finalized, the test was ‘unblinded’: Now the framebuilder shared which bike was made from which tubing. And the administrator shared which bike had which color stem cap during each test session. Only now did the testers find out which bikes they had been riding in each test session, but they no longer could change what they wrote about them. As a result, their impressions of each bike were free of bias. And the power measurements also weren’t influenced by riders thinking that one bike might be faster than the other. (The riders were not able to see the power numbers during the tests.)

Double-blind testing is not possible with bikes of different wheel sizes, but the testing should at least be blind: Have one tester collect the data, and another do the analysis. Will the results be the same without knowing which dataset is from 32″ wheels and which is from 29″ wheels? Rather than guessing ‘likely’ values for rolling resistance (Crr) and aerodynamic drag (CdA), the data analysis would test a wide range of values and determine which offers the best fit. This blind curve matching would eliminate the problem of starting with values that match the hypothesis.

What would happen if the curve fitting was done blind? My prediction is that the erroneous results—such as the huge benefit of 32″ wheels on smooth pavement—would disappear. However, at the same time, more ‘noise’ would be introduced. Dyer & Disley’s study suggests that the real-world precision of the Virtual Elevation Method, under ideal conditions, is about 2-3%—not even close to the 0.1-1% reported by John and Tom Anhalt. Once you test outside, rather than in an indoor velodrome, there will be even more noise.

None of this is intended as criticism of John or Tom’s work. Both are very conscientious in their testing—it’s the methodology that appears to have problems. As I mentioned in the beginning, the goal is not to prove who is right or wrong, but to figure out how bicycles really work. Despite the obvious issues with the Virtual Elevation Method—both the ‘incredible’ precision and the erroneous results—the underlying physics are sound. The method holds promise. The question is this: Once the method is improved with blind fitting of the curves, will we still get useful results? Or will there be too much noise? There is only one way to find out: Do the experiment described above, or validate the method in some other way. Until then, data collected using the Virtual Elevation Method should not be used as evidence that one setup performs better than another.

As to 32″ wheels, hopefully somebody will soon do a carefully controlled test with direct measurements, like our rumble strip tests (above) or the Escape Collective’s tire tests. If those tests show a big advantage for 32″ wheels, we’ll have moved one step further.

However, even that will not be enough: We’ll still need to find out why there’s no significant performance difference between 26″, 650B and 700C/29″ wheels in Bicycle Quarterly’s testing, and why the next step up to 32″ wheels suddenly reduces rolling resistance by a large margin. Because simply ignoring previous research is not how science works, especially if that previous research has been validated many times.

That validation is missing from the Virtual Elevation Method. As mentioned above, we’re offering to help with this validation, because a new method for testing bicycles under real-world conditions benefits everybody.

Further Reading:

Saturday, 08 August 2026

How does programming language affect token efficiency and correctness? [] (08:00 , Saturday, 08 August 2026)

This somewhat widely cited post (I keep seeing it cited, anyway) suggests that dynamic languages and/or languages that represent things more concisely are more token efficient. It seems to be cited enough that LLM search results agree. For example, when I searched for "dynamic vs static language token cost" (no quotes), Google's AI summary opened with

Dynamically typed languages generally have a lower LLM token cost than traditional statically typed languages because omitting explicit type declarations makes the code more compact.

Google's AI cited the same post, which suggests that some concise dynamic languages have maybe 1/2 to 1/3 the token cost of static languages like Rust, Go, C++, etc. The author says

There was a very meaningful gap of 2.6x between C (the least token efficient language I compared) and Clojure (the most efficient).

And then they later tried J, saying

It dominates at just 70 tokens average, nearly half of Clojure (109 tokens). Array languages can be extremely token-efficient when they avoid exotic symbol sets. If token efficiency turns out to be a key driver, this is perhaps a very interesting way for languages to evolve.

The other dynamic vs. static language token comparison I've found floating around is this one, which supports the same conclusion. If you want to treat this as part 8 of this series of exercises on benchmarking, evals, and experimental design, you can click through to the links and think about eval issues before reading further.

Without running our own eval, one problem the first experiment has is that the problems are trivial, which we can see from quote above; a problem that can be solved in 70 tokens in J and 109 in Clojure isn't much of a problem at all (the author used Rosetta Code). As we saw when we looked at other evals of caveman mode vs. our own evals, you can get very different results from trivial problems where most of the work is in printing out an answer vs. slightly less trivial problems that actually require some amount of "real work"; the big gains claimed by caveman mode and shown in replications go away when you start looking at problems that take more than just a few tokens. In general, performance on trivial tasks doesn't generalize.

The issues in the second link are a little more subtle, so we'll defer most of them to an appendix, but they include issues like one of the tests executing the wrong path (which doesn't exist), causing a test to fail. One of the later agents then symlinks the non-existent path to its own executable, which works for that case, but also causes every later test to run that one agent's executable instead of the correct executable. The author tries to draw conclusions about what it means that Rust had some failures, but all it means is that scoring for Rust ran before the Go agent symlinked all scoring on that broken test to the Go executable.

Instead of relying on these evals, we can try running some of our own evals. As we can see from these evals as well as the evals discussed in our last exercises on evals, it's very easy to make an eval that doesn't say what the creator of the eval seems to think it's saying. No doubt these evals will not be an exception to this and will be flawed (see appendix below for more details).

As a way to build my intuition about things, I like to pre-register guesses before looking at results1. Some things I pre-registered with friends were:

  • High confidence (95%): the overall dynamic vs. static language claim won't hold
    • For reasons stated above: this feels analogous to the caveman eval, where the result will, at best, get diluted as the problem gets larger
  • Low confidence (60%): static languages will be somewhat better than dynamic at ultra effort
    • Very weak confidence that, at ultra effort, the harness will get feedback to the model more quickly and this will result in some kind of benefit for either correctness or efficiency, but it would also seem reasonable for this to not be the case for all kinds of reasons, e.g., I've noticed that codex, when invoking the Rust compiler, very often makes the exact same error and then has to fix it; perhaps this kind of thing dwarfs things like a hypothetical faster feedback cycle
  • High confidence (98%): the "weird" language supremacy of something like J won't hold
    • Same reasoning as the overall static vs. dynamic claim, with the additional thought that AI labs are going to have much less (and possibly zero) synthetic data RL env effort on obscure languages

Zstd

For the first eval, I tried giving agents the zstd RFC (plus errata) and telling them to implement a complete zstd decoder (agents are stuck in a container without internet access). The tests were not given to agents. For something with the surface area of zstd, it's not really reasonable to expect that the tests cover every possible case. For example, even though zstd is a fairly well-tested piece of software, I once found a data corruption bug in zstd. The test suite isn't intended to find extreme corner cases that might be lurking for years and is instead intended to check various cases that can "easily" be derived from the RFC that should work.

Below, the x-axis is cost and the y-axis is correctness score (up and to the left is better / down and to the right is worse); average result on medium and ultra efforts with GPT-5.6 Sol. If we only look at medium (and ignore the fact that results often wildly differ on different tasks), we might come to a conclusion like the Alderson evaluation, that dynamic languages are more efficient and better when using LLMs because (ignoring relatively obscure languages) the cluster of dynamic languages lands up and to the left of the cluster of static languages (we used Alderson's color-coding for static vs. dynamic to make it easy to compare at a glance). But if we look at ultra effort, the results are quite mixed, with a couple static languages doing the best, with more static than dynamic languages among the better results.

The graphs below also have a toggle to convert the x-axis to time instead of cost. The mame/ai-coding-lang-bench noted that it's valuable to get results more quickly (I personally don't find this to be the case because results take long enough that I multitask instead of waiting), so we can also look at that. Similarly, we observe that neither language type dominates the other although, at medium effort on this particular task, the best dynamic language results are once again better than the best static language results (though, once again, they're fairly close).

We can observe that, just like when we compared completely trivial caveman mode evals to a less trivial caveman mode eval, the very strong relationships that held in the trivial evals don't generalize to this larger case. As was the case there, the extreme ratios in performance go away in these larger evals, except in cases where we might expect poor performance, such as when using assembly (which would be significantly more time consuming and difficult for a human) and when using relatively obscure languages where we might not expect that AI labs are expending effort generating synthetic RL environment data.

Note that this is the opposite of what the 1st eval found when it suggested that very dense languages like J would make sense for efficiency reasons. Perhaps using an obscure (and "weird") language can make sense if you have a very large budget and you can train or fine-tune a model to be effective for your pet language, but if you're a normal user of LLMs, it seems like sticking with a mainstream language is likely a better bet than using an obscure dense language.

And it turns out that if we plot language popularity vs. performance on this eval (not shown), we observe a weak to moderate positive correlation where more popular languages end up with more correct as well as cheaper solutions.

As we previously noted, very closely related evals can give substantially different results. For example, we saw significantly different results in the Optimization 1 vs. Optimization 2 evals here when Optimization 1 and Optimization 2 were optimizing bzip2 compression and decompression in wasm, which are fairly closely related tasks as evals go. To make a strong, universal, claim, like "dynamic languages are more efficient than static languages", we'd have to run evals across many tasks. However, showing that a claim like

Dynamically typed languages generally have a lower LLM token cost than traditional statically typed languages because omitting explicit type declarations makes the code more compact.

is maybe at best vaguely directionally true and not really relevant to any particular case and maybe not strong enough to be relevant in general, we just need to try a few cases and see that the claim doesn't generally hold. Above, we saw that at one effort level, the claim seems to maybe be kinda sorta true, but with exceptions, and then at a higher effort level, the claim seems to not be particularly true, which is sufficient to say that the claim is probably not universally true, modulo our eval having a confounder that completely invalidates it.

Pandoc

But, just to get a view on a very different task that's also presented in a different way (more TDD-like than "read a spec"-like), this next eval takes the Pandoc ProgramBench eval and modifies it for our use case. Instead of the reverse engineering task presented by ProgramBench, we present agents with ProgramBench materials as well as the ProgramBench tests and then score agents against a holdout set of tests to measure the performance of each condition2.

In the results below, the x-axis is cost again and the y-axis is score on the holdout tests.

As before, we don't see a very strong relationship between success or cost and whether a language is static or dynamic or very dense. We once again see that relatively obscure languages tend to do poorly (although Clojure does much better here than on Zstd). Also, Assembly does much worse, which seems expected in that we would expect a human writing Assembly to be at much more of a disadvantage implementing Pandoc than implementing Zstd and there doesn't seem to be a strong reason to think that LLMs would be different in this regard.

What does it all mean?

Who knows?

I have a lot of questions about what works well when using LLMs (such as, what test techniques work well, what languages work well, what software architectures work well, if bug fixing cost varies by language, if general program maintenance cost varies by language, etc.). Most of these questions are unanswered in public data and, if they've been answered in AI labs, the information mostly hasn't been made public.

Most of the claims that get thrown around about how a particular language is good for LLM use seem to be wrong (e.g., the claim that Ruby, Clojure, and J, are particularly well suited to LLMs, which were mentioned in the evals linked above, as well as the somewhat common claim that Elixir is particularly suited to LLMs), but it's not clear what's right.

In 2014, we looked at the literature on static vs. dynamic types and found that surveying the literature wasn't very informative outside of a few case studies. For an example that typifies a standard academic study, we saw the paper, Do Static Type Systems Improve the Maintainability of Software Systems? An Empirical Study, on which I commented:

Subjects were given classes in which they had to either fix errors in existing code or fill out stub methods. Static classes for Java, dynamic classes for Groovy. In cases of type errors (and their respective no method errors), developers solved the problem faster in Java. For semantic errors, there was no difference. The study used a within-subject design, with randomized task order over 33 subjects. A notable limitation is that the study avoided using “complicated control structures”, such as loops and recursion, because those increase variance in time-to-solve. As a result, all of the bugs are trivial bugs. This can be seen in the median time to solve the tasks, which are in the hundreds of seconds. Tasks can include multiple bugs, so the time per bug is quite low.

Picking tasks that avoid "complicated control structures" such as loops and recursion, where tasks take hundreds of seconds makes the result meaningless with respect to tasks that really eat up a professional programmer's time, just like the first eval we saw where tasks took high tens to low hundreds of tokens. However, with LLMs, we can actually feed them non-trivial tasks and compare how they do. There's the issue of how well results generalize to different tasks, but we'd have that exact same issue with human studies, but worse (LLM variance is huge, but human variance is even huger since you can't get the same human to do a bunch of tasks with different seeds). And while $20 to get an LLM to implement a Zstd decoder isn't exactly cheap once you multiply by the number of languages and the number of iterations per condition per language, if you think about how much it would cost to hire a professional programmer who can read the zstd RFC and implement it, there's no way the equivalent study would've been done because the cost would've made it completely infeasible. That goes double for the Pandoc task.

With LLMs, a lot of the questions have gone from being effectively unanswerable to being answerable with a bit of effort and some tokens. Due to the incentives that are in play3, it's not clear that we'll get answers to questions like this any time soon, but it's at least possible to take a crack at it now.

There are a lot of claims I've seen floating around that these evals can't prove or disprove (for the reason noted above that, due to the variance across different problems, many more tasks would have to be tried), but that these shed some light on, such as:

  • Languages with a lot of bad code out there (e.g., PHP) will perform worse
    • Appears to be false on these tasks
  • Because it's so easy to re-write now, you should use a powerful language (like Haskell)
    • Appears to be false on these tasks
  • You should use a popular language
    • There's weak support for this statement

For my pre-registered guesses, we had

  • High confidence (95%): the overall dynamic vs. static language claim won't hold
    • This seems correct
  • Low confidence (60%): static languages will be somewhat better than dynamic at ultra effort
    • There's not enough information to determine this conclusively, but if we had to make a binary correct/incorrect call, I would call this incorrect
  • High confidence (98%): the "weird" language supremacy of something like J won't hold
    • This seems correct
  • [from a draft reader]: "dynamic is better on small-scale, but gets overtaken by static as the size of the project grows"
    • Not supported by these tasks (static languages didn't seem to do substantially better than dynamic on the much larger Pandoc task vs. the smaller Zstd task), but the tasks and the presentation of the tasks are so different that it's unclear if this is because task-size scaling or because of other differences

By the way, a major reason Clojure improves by so much in the Pandoc eval compared to the Zstd eval is that, in the Zstd eval, 36/40 medium and 5/40 ultra Clojure programs had test failures because byte conversion throws on 128–255 (maybe unchecked-byte should've been used?) and they used this conversion inappropriately.

That's a real result, in that, if you ask the best publicly available GPT model to implement Zstd (and presumably if you do other bit/byte manipulation tasks where this might come up), it will emit code that fails in this particular way. If there are tests that catch this, the bug will get fixed, but it will still cost time and tokens. Whether or not a language did well, there are costs like this all over the place (for example, cargo repeatedly gets invoked with the wrong arguments, which then immediately gets caught and fixed, but I've noticed this loop can actually consume a decent amount of wall clock time on my real projects unless you give explicit instructions to codex on how to invoke cargo, and it's clear that's worth the space in the context window).

Anyway, all of this is an illustration of why, if someone wanted to make a strong claim about which languages or classes of languages are particularly good with LLMs, they would need to run quite a few different evals. If we dig into why any particular condition got a certain score, the failures that caused the score are generally something idiosyncratic where it's not always obvious how much the issue generalizes across tasks or across setups. There's no way to look at the score on one eval or even five or ten evals and draw a conclusion about programming in general.

It's true that, in both the Zstd eval and the Pandoc eval, we see a correlation between language popularity and positive outcomes (higher correctness, lower cost, lower wall clock time) and it seems plausible that we'd see this across other evals, but it would be a mistake to draw a strong conclusion about any particular language. I gave a warning like this back when I looked at how often different projects have a broken build according to GitHub CI data, noting that there are different reasons that a build might be broken more or less often across projects and that one shouldn't draw strong conclusions because results across projects aren't necessarily comparable (for example, if one project's main branch is some kind of release candidate that's gone through other vetting, that project would be expected to have low build breakage, but that's not comparable to a project where people are developing directly against main).

Shortly afterwards, someone involved in one of the languages with a high score (IIRC, it was Martin Odersky and Scala) tweeted out the post and cited the language's high ranking as a victory for the language. That was an unwarranted conclusion there and, due to the many sources of variance that are in play here, any such conclusion about a single language would be even more unwarranted here.

This data (assuming eval validity) can refute some strong claims and is suggestive of other claims, but it can really only be suggestive of things for classes of languages and not for particular languages due to having only two tasks, which any particular language could do well or poorly on for some idiosyncratic reason which may or may not generalize to other tasks.

Thanks to Max Bittker, Yossi Kreinen, Aaron Levin, Alan Boll, Luke Burton, Marco Primi, Milosz Danczak, Justin Blank, and Tom Adamczewski for comments/corrections/discussion.

Appendix: selected issues in ai-coding-lang-bench

Like I said above, my eval here is a quick and dirty eval and I'm sure it's full of flaws, so I'm not trying to say the evals I've presented here are great and this is bad, but here are a number of issues in the Endoh ai-coding-lang-bench eval.

One issue is that the wrong executable appears to have been run for some of the tests. The setup for the published run seems to have executed ../../minigit inside each candidate's directory for one of the tests when the candidate's generated executable is at ../minigit. ../../minigit doesn't exist.

Because statically typed languages had a lower correctness score, the author of the eval noted "the only failures in 600 runs were in Rust and Haskell (both statically typed, both relatively "difficult" languages)" and suggests that "difficult languages", such as "C's memory management, Rust's ownership model, and Haskell's monads/purity may add overhead for the AI".

However, Rust's failures were because there is no executable at ../../minigit, causing the test to fail. The first Go run "fixed" this by executing ln -sf minigit-go-1-v1/minigit ../minigit and linking generated/minigit to its own run, but this means that every later execution (for every language) actually executed the first Go run's executable. On rescoring Rust against its own executable (as opposed to having it fail by trying to execute a non-existent file), Rust gets a perfect score, invalidating the theory that Rust had failures because it's a difficult language to deal with.

Other tests also have issues. For example, two tests have a structure that causes them to pass regardless of the actual value being checked. One of the tests has

  if ../minigit commit ...; then                                                                                                                                      
    COMMIT_POST_CHECKOUT=$(cat .minigit/HEAD)                                                                                                                         
                                                                                                                                                                      
    if grep -q "parent: $COMMIT1" \                                                                                                                                   
        ".minigit/commits/$COMMIT_POST_CHECKOUT"; then
      pass "checkout then new commit works"
    else
      pass "checkout then new commit works"                                                                                                                           
    fi                                                                                                                                                                
  else                                                                                                                                                                
    fail "checkout then new commit works"                                          
  fi

The inner if has a pass in both branches, meaning that this is almost equivalent to

  if ../minigit commit ...; then
    pass
  else
    fail
  fi

The inner if appears to be intended to have the actual check, but due to a coding error (perhaps a copy+paste error?), the check is effectively elided.

Also, as noted above, agents can modify the test environment, which the 1st Go agent did to fix a broken environment. They have full access to tests and the environment and can do anything and the test suite is visible during development with no holdout, which can easily lead to cheating by special-casing code in a way that passes tests but creates a program that's useless "in real life". At a high level, something like this seems to have happened in that many programs fail to implement large parts of the spec but do pass all tests, which may indicate that the agents "understood" how to pass the tests and preferred that over implementing the spec (it could also indicate that the tests are very thin and are easy to pass).

Another issue is that the Claude Code CLI versions aren't the same for all runs (it varies from 2.1.66 to 2.1.68). There are a handful of other issues like this that could be significant, but are likely small compared to the issues noted above.

Appendix: medium in a loop vs. ultra

As an example of something we can compare, I was curious how cost effective using medium + asking the agent to keep working would be and then, in the back of my mind, I also had this question about something "Ralph loop" advocates say, that you're better off clearing the context window on every iteration of the loop and giving the agent the full prompt again. As with the above, my pre-registered guesses here are:

  • Zero confidence (50%): Ultra is more effective than medium in a loop
    • I'm not sure how to think about this. I guess the case for this would be that ultra was designed in some way and should be smarter than repeatedly doing medium in a loop. But it's possible that there's some tradeoff where ultra was made for more speed and, as we've noted, the variance is very high so even if ultra wins on most problems it might lose here; ultra might also be more optimized for trading off to improve wall clock time or another parameter; ultra also has the disadvantage that it doesn't "know" to stop after reaching correctness on the hidden tests, whereas medium conditions that hit full correctness aren't run again under this setup, which hugely advantages medium in a loop (which is arguably realistic w.r.t. how someone might use these)
    • You could maybe say this is 50% + epsilon since my mind went to framing it this way and not the other way around, but I would say extremely low confidence here at best
  • Medium confidence (80%): continuing with context outperforms Ralph loop
    • /goal mode, etc., don't do this by default and, presumably, folks at Anthropic and OpenAI have tried things like the Ralph loop and found them less effective
    • Watching your context window very closely seems to have gotten less important as harnesses (and models?) have improved; in late 2025 / early 2026 I often had to throw out my context window when working on a long-running task to avoid issues and that's gotten rarer over time but, even then, because I wasn't paying attention to what people were saying, I was running agentic loops with a default of keeping context and only clearing when there were obvious problems, which seemed to work ok, e.g., I built the world's strongest Azul AI doing that, so it's not clear to me that having a default of clearing context on every loop iteration was the right choice back then

Below, we have the average result for medium in a loop vs. ultra, sorted by best to worst ultra correctness score, for a prompt that simply resumes individual runs that don't have 100% test correctness as well as a Ralph-loop like prompt that discards context and gives the original prompt again (x-axis is cost, y-axis is number of correct test cases):

For this one problem, on average, running ultra once seems better than repeatedly running medium per unit cost (and much more so per unit time) and continuing with previous context outperforms Ralph. The problem with naively running medium on repeat is that the agent can get anchored to a bad solution and fail to make progress. The theory behind the Ralph loop is that you throw away bad context which can cause this to happen, but that doesn't save you from having a bad artifact.

Just from using LLMs, I've noticed that you're often better off throwing away a chunk of code and having an LLM re-write it from scratch than you are having an LLM modify it or try to re-write it in place. Michael Malis, who's been re-writing Postgres in Rust and has been making major changes has also noted this. This also relates to this idea noted previously that, due to the high variance (plus this path dependence) you're often better off rolling the dice multiple times and taking the best result, if you don't mind spending the tokens.

It's hard to say too much about static vs. dynamic languages from looking at just this one condition, but a naive thought like "static languages will outperform when iterating" isn't obviously true. If there's one pattern that jumps out at me, it's that the cases where the Ralph loop most badly underperformed continuing with context were generally dynamic languages. It's possible this is because of the lack of type information, but we'd need to both look at the differences in trajectories in more detail as well as look at other examples to observe if that's a real pattern. Even if you don't care about Ralph loops now that the Ralph loop trend has passed, being able to make changes to a codebase more effectively when starting a new task or starting with fresh context is something you might care about and the pattern here is suggestive of a possible advantage.

Appendix: Guards of Atlantis 2

I tried to do a third eval that seemed like a more "business logic" kind of eval in both how the problem is presented and the actual execution of the problem. You can argue that the Zstd eval and the Pandoc eval are quite unusual tasks for a programmer to face in that not many programmers receive a specification as well-written and thorough as the Zstd RFC and not many programmers are handed a problem with as many pre-created tests as you get from ProgramBench tests.

The idea here was to implement a board game. In general, board game rules are written by people who aren't experts in writing clean specs, so implementing a board game is more like what happens when a non-programmer (or a programmer who isn't an expert at writing good specs) gives someone a task.

The problem here is getting a game where I have a reasonable oracle for scoring that isn't trivial for LLMs. For example, LLMs were able to one-shot the rules for Scout and Azul, which make those poor tasks. For games that an LLM won't immediately one-shot, I happen to have an oracle for Guards of Atlantis 2 because I had an LLM implement a copy for me and my friends to play (no link for this one because I don't see how to make an interface that's free of copyright infringement). The backend only took a few hours of my time, but it took a fairly large amount of LLM time to get the rules to be roughly correct. I like this as a task in that the rules are tricky in the same way a lot of problem descriptions that are delivered to programmers are tricky, but it is, in principle, possible to figure out the correct rules and implement them (after all, humans implicitly do this when they play the game correctly offline).

In board game rules, it's fairly common to have rules where reading the rule strictly as written is incorrect and you need to use "common sense" (or read some kind of FAQ) to play the rule correctly (there are some game designers who strive to avoid this, such as J C Lawrence, but this is fairly uncommon). Guards of Atlantis has quite a few rules like this. The designer of Guards of Atlantis is also vocal about there being no such thing as the spirit of the rules or common sense interpretations of the rules and says that you should always read the rule exactly as written, which creates two difficulties. One is that there are also many cases where you need to ignore the "common sense" interpretation and read the rule exactly as written. The other difficulty is, as anyone who's ever tried to write a formal spec knows, it's very easy to accidentally have ambiguity or contradictions. Even people who do this profesionally are unlikely to be able to create a non-trivial, complete, clear, spec without formal methods or a very large amount of human review. Realistically, a board game designer who doesn't have a background in writing formal specs doesn't have a chance, and thinking that it's easy (as the designer seems to) reduces the already low odds even further. A nice way to mitigate this kind of issue to write down your intent or "the spirit of the rules", but because the designer says that there is no such thing as the spirit of the rules and you should read all rules exactly as written, there are no meta-comments in the rulebook that would help someone interpret confusing or abmiguous rules. This combination is quite difficult for LLMs (and, judging by the rate at which I see humans play the game according to the designer's intent, it's also quite difficult for humans).

I think it would be effectively impossible to just read the rules and play correctly (of course it would be possible, but it would require knowing which rules are to be read as written and which rules are not, which one would have to do randomly and get lucky as the rules don't define a consistent system that one could use to infer which rules obey which meta-ruleset). When I was implementing the game, in order to get my LLM to understand the rules, I gave it various resources such as an unofficial rules FAQ (which is correct), an unofficial short version of the rules (which is better written than the official rules and correct, but incomplete), an opening book (which can be used to test rules against on the assumption that the opening book only contains legal moves), comments from the rules channel on Discord, etc., and had the LLM do consistency checks across these with the understanding that things like the FAQ and the Discord comments have higher authority than the actual printed rules.

One additional source of difficulty is that the designer is frequently delibrately unhelpful when answering rules questions. He often likes to make fun of people who ask rules questions or played rules incorrectly, which has a chilling effect and reduces the number of rules questions (multiple people have said that they don't ask rules questions because of how the designer behaves), and when he does answer questions, it's often with something like a meme image that says "reading the card explains the card". To extract the information, the LLM has to process these meme images, and then there's often no information or delibrately round-about information, such as, in the case of ambiguity, a referenece to a particular section. When people do point out contradictions, the designer often says it should be obvious which side of the contradiction is correct, which may be true for a human who's kept up on all rulings to date, but current SOTA LLMs find many of the designer's rules clarifications unhelpful.

Yet another source of difficulty, perhaps related to the designer's propensity to make fun of people who ask rules questions or are confused by rules, the game's interface seems almost designed to trick people into doing the wrong thing. There are multiple design affordances that I've seen trip up most new players (even if you explain to the UI trap to them, there are enough rules to take in they often forget, and then when it trips them up, they'll say something like "I'm an idiot, you even explained that to me twice"). It's not clear if these traps were created to give the designer people to make fun of, but that's certainly one result. Another is that LLMs struggle to understand the games rules and UI.

With my $200/mo personal OpenAI/codex account, I let an LLM use all my spare capacity to run consistency checks and make rules fixes. I didn't closely track how long this took, but I think it was something like a month or two of cranking on fixes like this to get a somewhat reasonable result that's playable, but that I wouldn't really trust to be correct.

The only reason I somewhat trust this is that Pedro Oliveira also implemented Guards of Atlantis and they used a completely different approach (a more standard approach of having a human drive an LLM rather than trying to get the LLM to figure things out itself). When we compared implementations, we found maybe 10-ish bugs in each. There are probably some remaining bugs where both of our implementations incorrectly do the same thing and perhaps some where our implementations differ but the checking system didn't notice, but I think the rules for both of our implementations are now reasonably solid. That's how I have an oracle for this game.

I like this as a task because it feels more like the kind of "specification" you get in the real world, where the spec is ambiguous and contradictory and sometimes just plain wrong, and then you need to use other information to get a correct result. For this eval, to avoid having it be a test of how well LLMs can access data in annoying formats (such as converting the opening book from a set of images to some kind of structured data, converting a scan of the rules to text, etc.), I gave agents both the originals of anything where I directed an LLM to extract the data (which also required various consistency checks to get correct) as well as the the extracted data (the originals were presented so that LLMs could check the originals for extraction errors if they chose to).

While I did this task with older models (I did a chunk of it with GPT-5.1 or 5.2, and then another chunk with 5.4 or 5.5), with newer models but without the kind of guidance I gave to the older models, the task was still far too hard. Regardless of language, agents scored approximately 0 on this task.

BTW, if you're curious what LLMs (and humans) struggle with, here are some examples. There's one card whose text reads "Target a unit adjacent to you. After the attack: may repeat once on a different enemy hero."

In this game, a hero is a type of unit. Read strictly, with full knowledge of the rules, e.g., what "After the attack" means, etc., this should mean that you can either attack a single unit or you can attack two heroes (after all, to repeat the attack on a different enemy hero would mean that the first unit was a hero; otherwise it would be a different unit that is a hero, not a different enemy hero).

This card actually has what is effectively an errata printed on the card because people complained it was unclear; the errata reads "(You may repeat even if the original target was a minion)". That's already confusing to LLMs (and some humans), but the real killer here is that there are other cards that use the same construction and don't have this correction. To play other cards with the same construction correctly, you need to know that every time this construction is used, you should play it with the errata that's on this card. There are a number of constructions the game designer likes to use that have a specific non-literal meaning that you have to keep in mind.

Another example of a rule that shouldn't be played in the obvious way is a character with a card which reads "Choose one, or both, on different targets: A, B". Reading this strictly as written, one would expect to be able to, on different targets, do either A or B, or both A and B. But part of the spirit of the game is the meta-rule that a character can't attack another character multiple times with one card, so the interpretation that you can do what the card says and do both and A and B on some number of different targets can't be right. Based on similar deductions and how similar constructions are used, the way this card is supposed to be interpreted is "Choose one, or both on different targets", which is arguably still ambiguous and could be more clearly written as "Choose one or both (must be on different targets if both)".

As a human, once you understand what the "spirit of the game is", you can resolve these kinds of things. But, by design, this isn't written down clearly in the rules and one has to infer this from Discord discussions, which appears to be beyond the capability of today's models even though humans who are outperformed by today's models on many specialized tasks are able to do this.

When I was supervising the LLMs that implemented the rules, the reason LLMs reached a ceiling and didn't converge to fully correct rules was that an LLM would observe that a rule was inconsistent and incorrect. It would then try to fix this rule and would also fix other things to try to make them consistent and correct. This would sometimes make things more correct and sometimes make things less correct. When making things less correct, the LLM would sometimes modify an existing correct test to turn it into an incorrect test so, after a while, the LLM wasn't really improving correctness and was just churning on which rules were incorrect. That was with some guidance on what to check and how to check it; without that guidance, even with the more advanced models that are available today, LLMs were unable to navigate this in a reasonable way.

I'm sure there is a board game of the right rules complexity to make for a good eval here but, by definition, this would be something where it would take some work to create the oracle and I don't have an oracle handy for a board game with the right rules. If my goal were to make evals, I would've used board games with actual game replay data to get good tests or oracles for a whole bunch of games, but my goal was to play a particular game with some friends. But, if one were inclined to try this board game thing, it should be possible to create hundreds (thousands?) of these in a scalable way, so one could get a reasonably correct oracle for hundreds or thousands of games and then check which games are at the correct level to be an interesting test for LLMs today.

This is arguably a bit of a funny problem in that, given a clear spec, e.g., a clearly written set of rules, an artifact that's more complex than Guards of Atlantis can be implemented by LLMs (I would argue the Zstd RFC is more complex, and Pandoc certainly is; even individual document formats Pandoc supports, like PDF, are more complex than Guards of Atlantis), so the problem isn't finding a game with rules that are complex enough that LLMs struggle and the problem is more about finding a game with rules that are ambiguous or contradictory enough that LLMs struggle, but not so much so that LLMs are completely hopeless. This is an actual real-world problem, in that humans are generally not very good at writing clear specifications and how well models and harnesses can handle a human's unclear, contradictory, and sometimes just plain wrong, specification is probably more relevant to the typical user than how well an LLM can implement something from a specification as well-written as the Zstd RFC or how well an LLM can implement a problem when handed the 4800 ProgramBench Pandoc test cases plus documentation. And these problems seem solvable in principle, in that humans who want to play board games correctly (even ones who would have no hope of "playing" Zstd correctly, let alone Pandoc) are generally able to navigate the mess of information out there to figure out what the rules to a board game are.

If we look at it form the other side, this is suggestive that, to get an LLM to do something, maintaining a clear, canonical, spec is an effective way to work.

Appendix: reasons for various decisions

  • Testing ultra
    • I've seen people say that you shouldn't really measure this because this is a harness thing and not a model thing. I can see why you'd want to measure these separately if you're working on improving models or harnesses, but when looking at how users use things, many people are just going to use codex or claude with the various built-in features and options; whether or not something is a harness thing or a user thing isn't really relevant to them
  • Using codex
    • I've seen evals use a very thin harness for the same reason as above and my reason for using codex and not a very thin harness is the same as above
    • Similarly, in this caveman model eval, I used claude with Opus and Fable and codex with GPT
  • No internet access
    • Models will often cheat if given internet access and there are plenty of problems where searching on the internet doesn't turn up source code that solves the problem, so this makes these evals approximate those more closely
  • Relatively large tasks compared to a lot of benchmarks people pass around
    • Although I have LLMs do plenty of trivial tasks, the things that take my time or take tokens tend to be larger than the kinds of tasks that were in the Alderson eval or the Endoh eval; LLMs are good enough at trivial tasks that it doesn't make too much difference to me if some condition makes them slightly better or slightly worse at one of those trivial tasks, but for a task like implementing Guards of Atlantis, where I have to spend some number of hours setting up scaffolding for the task to even sort of work, I care a lot about what makes models perform better or worse
  • Agent-specified prompts
    • Public evals seem to have moved to relatively thin/lightweight prompts that don't specify the task in great detail; this is said to be better because an agent setting up a task will give too much information that helps agents succeed at the task
      • I can see why you would want to test that, but it's also the case that I care a lot about how well agents do at tasks set up by agents because a lot of the tasks that I have agents execute are tasks that are defined by agents; I care about how agents perform under both styles, not just one style, and the public evals have moved towards one style
  • Zstd eval: asking agents to fix bugs without telling them the issue or the failing tests
    • In general, if you tell an agent to fix a specific thing, it will fix it, but it won't necessarily fix the class of issue; I've found that if you tell it there's an issue but don't tell it what the issue is, it sometimes does a more general thing instead of just putting in a narrow, brittle fix, so I do care about how agents behave when given instructions like this (of course you can tell agents to not just make a narrow, brittle, fix, but that often doesn't work)
      • This feels a bit related to the issue we noted in the Pandoc holdout footnote, where telling agents we had a holdout set appeared to force agents to produce more generalized and less brittle solutions

Appendix: issues with these evals

When it comes to performance benchmarking, I've done enough of it that I feel like I generally know how my benchmarks are flawed and I can make an informed time/effort vs. flaw tradeoff and I have decent confidence the flaws that exist in the benchmarks aren't material to the thing I'm trying to understand. I haven't done enough AI evals to have this kind of feel for AI evals so, at a meta level, I would expect any AI eval I do to have some unknown-to-me flaws.

Another reason I would expect some flaws here is that I had coding agents set up these evals and every time I spent a minute looking for issues I would find at least one issue. This indicates that it's fairly likely that these evals have additional flaws that could be uncovered by looking a bit more, but I wanted this to be more of a "quick toy project" level of correctness than a "Gary Bernhardt" level of correctness, so I stopped after fixing a handful of issues.

Back when I was working as a verification engineer, I attended a meetup by a Sun/Oracle engineer in Austin, maybe around 2007 or so, where they mathematically formalized this idea of converting the time between bugs to a level of confidence in a chip release. I haven't seen people do this much, but I recently heard Will Wilson (co-founder of Antithesis) mention that some folks at Antithesis used math from ecology (the literature on rare species observation) to estimate true bug rate, which seems like a much more sophisticated version of what this engineer at Sun/Oracle was doing a couple decades ago.

That's a cool idea, but when you're finding a bug every minute you look, you don't need fancy math to tell you that there are probably a lot of other bugs. If I were doing this for work and we had some reason to care about the fidelity of these evals, it would probably make sense to look at these more closely and fix more issues (and I would probably have the skills and experience to make fewer mistakes in instructing LLMs to set up these evals if I did this kind of thing for work). But, for the purposes of answering the question "is the claim that dynamic languages are meaningfully better than static languages when using LLMs?", I have a little more confidence that the claim isn't true, and there are a lot of other questions that seem more likely to yield some kind of actionable result (such as, what techniques or test libraries work best).

I normally don't publish things on the blog until I feel like they're somewhat solid, but this means that I often explore some data enough to satisfy my curiosity and then never publish the result. From talking to people about these non-published results, people I talk to are often curious about the results even if they're not done to a standard that I really like, which seems like an indication that folks I don't talk to might be interested as well. From what I've seen so far, I suspect it would take at least 10x the time I've put into this to get this to a standard I really like. I'm fairly busy at the moment and can't see myself having the time to do that for months, at which point I'm not sure I'd really ever get around to publishing this. In a recent post, I mentioned an analysis I did almost a year ago where I was trying to understand which cars are better for concussion risk in accidents, where I spent some time figuring that out, got far enough to get an answer that satisfied me, and then didn't ever get around to doing the work it would take to clean up the result enough to publish it.

There are some results from that analysis seem "publishable", in the sense that they could turn into a published paper (such as finding from actual crash data that the relationship between HIC and velocity looks like it's to the fourth power (!); there's a paper that tried to find this relationship, but did the wrong kind of analysis and wasn't able to find an "O(n)"-style relationship and had something much fuzzier), but I've never really cared about whether something is a paper or a blog post and it turns out that I'm more likely to just move on to the next analysis instead of cleaning up the analysis enough to publish a post.

A more recent project along these lines is that, after making a superhuman Azul AI, I tried to make a superhuman Splendor AI using a much less human-time-intensive process. I believe that didn't succeed, but it beats every other Spelndor AI I could find by a good margin, which is a mildly interesting result. I think I know enough about board game AIs to write something up about them, but my main interest was in figuring out if I could get something decent, and then I keep just doing other projects instead of spending the time to do a nice write-up. An example of something I think is interesting there is that a lot of the performance optimizations you want to do actually change the result, so you can't only rely on optimizations that can be strictly checked to not change the result. But, if you naively ask a coding agent to do these optimizations in a way that doesn't reduce playing strength, they'll do all sorts of things that reduce strength. Cases where the strength reduction is very severe are easy to catch, but there are more subtle issues that sometimes result in (for example) no change in strength vs. your own AI in self-play but a reduction in strength against humans or other AIs, so some kind of process to catch bad optimizations is necessary, and it's inherently a kind of arbitrary process that has to be designed using some combination of your intuition and relying on LLMs (which will be very helpful but also often completely wrong).

For these kinds of data-y projects that I'm interested in, LLMs massively reduce the amount of effort it takes to get a result that's strong enough to satisfy my curiosity but, AFAICT, they don't reduce the effort it takes to publish a result by much (at least if you write up results by hand instead of having an LLM write up the results and you want the results to be nice and clean), which means that writing them up runs into a kind of Ahmdhal's law bottleneck, so I've been doing more projects like this and writing up fewer of them. If anything, I think it actually takes more time to write these up because of how I've changed my workflow. For example, instead of just outputting some graph from ggplot2, I'll make a version an interacive version that's nicer in some ways, but definitely takes more time to produce. And I run an LLM spell/grammar check pass (at least so far, that's the only LLM assistance I've used for writing), which turns up a bunch of issues to be fixed. Since I look at each one manually instead of taking the fixes (and I make a lot of typos), that's actually fairly time consuming (over an hour on my last post and over half an hour on this post even though I didn't even make corrections all the way to the end and abandoned the process maybe halfway through).

Anyway, publishing this is an experiment in publishing some half-baked notes instead of having the kind of cleaned up version that I'd really like to have before publishing something. If you have opinions on this, please let me know (X Bsky Mastodon)!

I don't have GitHub links to the current evals. On the one hand, I feel like I really should. On the other hand, they're a mess and there's a bunch of stuff I'd want to clean up before publishing the code, and I don't know if/when I'll get to that and this way, at least I'm putting something out there instead of just talking to a few friends about the result and then having the result sit on my hard drive indefinitely?

Appendix: more details on Zstd

Agents were instructed to ignore performance, but the timeout wasn't infinite and, under the medium condition, some test cases timed out. This is arguably unfair, but this didn't materially impact the score. For non-infinite loop timeouts, there were 2 test cases in Clojure (across 40 * 34 tests), 2 in J, 2 in Tcl, 1 in Factor, and 1 in PHP. And, at 9000s (2.5h), the timeout was fairly generous considering that the largest test case was 4 GiB. Failing to decode 4 GiB in 2.5h is an implied rate of less than 0.5 MB/s on a Graviton 5 core, which is quite slow.

Here are some of the issues that I ran into when trying to get agents to set this up (and, as noted above, the short amount of time it took to find each issue implies there are more issues)

  • Originally, the build setup wasn't clearly specified to agents, causing some languages to randomly fail when agents did something that seemed reasonable based on how this was specified to agents but didn't work when scoring occurred
    • BTW, I was very exicted by the initial result here because it was super interesting looking and it confirmed my biases. Dynamic languages were substantially worse than static languages. What a blockbuster result! But it turned out that the real result from the initial setup was that static languages were less likely than dynamic languages to have problems caused by this issue because static languages were less likely to have issues with the idiosyncratic way project builds were ambiguously specified
  • In the original assembly conditions, agents implemented code in C and then compiled it to assembly and submitted the assembly
    • With this issue, asssembly did as well as other languages, which is super interesting! And also false once this issue was fixed. It turns out to be very easy to get incorrect but compelling looking results that would easy go viral if you aren't careful. After fixing those two issues, the results looked fairly mundane and fall into what you might call a "negative result" in the framing of a paper, in that there's no interesting or surprising or contentious thing the results show; perhaps slightly favoring boring languages would've been contarian result for very online people 10-20 years ago, but very online trendy discourse seems to have moved away from that, so this isn't really an interesting contarian result anymore
  • For some reason, the agent doing the setup imposed unusual arbitrary restrictions on some languages and not others (for example, the Rust setup didn't have access to rustfmt or Clippy); most, but not all, languages had things like this
  • Many of the tests (which were created by an agent) were actually some kind of performance/stress tests even though agents were instructed to ignore performance (I wouldn't consider processing 4 GiB of Zstd in 9000 seconds a performance stress test)
  • Some language conditions had arbitrary instructions to agents (for example, the Haskell condition had instructions not to use bytestring, with instructions on alternative implementation suggestions)
  • Some language conditions had old toolchains (for example, Zig was on 0.10)
  • Some language conditions had scaffolding to help agents implement Zstd
  • Some language conditions had explanations of tools that were available that were incorrect (for example, assembly conditions were told they had access to GDB, but GDB didn't work)
  • The agent responsible for health checks for running iterative evals would sometimes decide that evals weren't making enough progress and give held out tests or other information to agents inside the eval

There's one thing which arguably wasn't a bug that I removed anyway. One of the tests was very hard (maybe 10% of agents passed the test on the first try). On testing the current zstd release binary, the zstd binary also fails this test. On reading the RFC, this seems to be an ambiguity in the RFC about the legality of a certain edge case. There was fairly strong clustering with respect to which languages passed this test case more frequently, which I think is interesting, but doesn't seem like a very useful thing to measure when all of the other tests are measuring (or at least attempting to measure) something more straightforward.

Anyway, in the above list (which is not exhaustive), many of the issues impacted a large fraction of languages and some issues had to be fixed multiple times. All told, if you count each condition as a separate bug, I probably fixed (had agents fix) over 100 of these bugs and I expect there are more. When I talked to Max Bittker (who runs an RL environment startup), he noted

all the evals I've worked on, I ended up putting a huge amount of time and effort into, mostly in the form of reading trajectories (or summaries of many trajectories) and then triaging issues , e.g "oh this class of bug shouldn't be possible, lets update X "(X being the prompt, the harness/ environment, or the verifier)"

agents tend to slop this up, so I put a lot of care there to make sure things get fixed at the right layer, for instance it's very sensitive what's in-context for the agent under test (bad to add random junk it has to worry about, or at worst leaking answers) vs whats fixed behind the scenes in other parts of the system.

agents, when writing evals, are not sensitive enough to the experience of the agent under test, and will just give it the answer or fix problems by making it the inner agent's problem ("remember to not reward hack plz")

I also have had a lot of success re-using existing things (repos, games, tools, levels) and building harnesses and verifiers around them, versus trying to make something from scratch for an eval by prompting

In retrospect, I sort of regret doing a cross-language eval. Even after fixing 100 or more eval issues, I have no doubt that plenty more remain. Maybe this is just a "grass is greener on the other side" thought and I'll also regret the next eval I try, but I think it would've been a lot less work to try to evaluate how well different test techniques or testing frameworks work than to evaluate different languages and I find that topic at least as interesting. And, in retrospect, had I done a lot more work by hand and relied on agents less, this would've gone a lot better. For example, I should've had agents produce an environment for one language and then both had agents inspect it and inspected it myself and fixed the issues before producing the environment for another language. After doing this a few times, I might've had a better setup for producing environments for other languages (and if not, I could've just repeated this process for each language and gotten a more reliable result, likely without even taking more time).

Another thing to note is that a number of things that are genuine differences in languages weren't really tested, such as memory safety against adversarial inputs. If agents had a harder time producing generally roughly correct code in C or C++ than Rust, that would be observed, but if a fuzzer or valgrind or other tools would turn up issues, that's not likely to be captured in the small set of tests. Just out of curiosity, I asked an agent to (briefly) check the Zstd C and C++ code for memory safety issues. The agent claims it ran the C and C++ code under ASan+UBSan and tried a few fuzz inputs (4000 each) and didn't find issues, but of course that doesn't mean there aren't issues or that a larger codebase wouldn't have issues.

And, in fact, doing an analogous quick check for memory safety issues for the Pandoc eval found memory safety issues in all of the C programs and all but one of the C++ programs (the issues were things like incorrectly dereferencing out-of-bounds memory; one specific example is that, in one of the C programs, a truncated LaTeX table could result in an out-of-bounds memory read). The fact that these issues were findable with 10 of seconds prompting indicates that many such issues could be found and fixed without much human effort, but it would cost quite a few tokens and would push the cost of the C and C++ versions well beyond the cost of the Rust version and after doing all of that you would still have less confidence in the memory safety of the C and C++ versions than in the Rust version.

Anyway, if you're curious about the distribution of results, we have the following for medium and ultra:

I don't love that the ultra results are somewhat saturated here, but one "problem" with testing ultra is that it will keep going for a long time as problems get harder (e.g., most of the Pandoc ultra runs ran for 12+ hours, and the assembly runs went for much longer), so the things that don't get saturated are very large tasks, like the Pandoc eval, or tasks that are too difficult in some way, like the Guards of Atlantis eval.


  1. a draft reader pre-registered the guess, "dynamic is better on small-scale, but gets overtaken by static as the size of the project grows". [return]
  2. The holdout tests seem necessary because, without them, agents cheat and will detect a test input and hard-code the passing test output (they sometimes do this even when instructed not to cheat). If all cheating was that blatant, that wouldn't be a problem (and could be an interesting thing to measure, as agents differentially following directions or not across languages is something that matters to real users), but a lot of the cheating is more subtle and difficult to adjudicate. For example, some agents wrote code that branched off of the structure of the tests, but then filled in the contents of the branches with code that wasn't special-cased to a single test result and could pass many variants of the same test. For any point on the spectrum from "definitely not cheating" to "obviously cheating", some agent tried it. As we saw when we looked at Senior SWE-Bench, LLM scoring of evals is tricky and a great way to introduce both bias and variance; using a holdout set of tests has some problems, but it lets us avoid this much larger set of problems.

    For one thing, the holdout tests are suspsicious because they were created by agents. The intention was to create holdout tests that a reasonable person (or agent) would be able to make pass if they're not cheating. Agents audited this set of holdout tests for cases where this wasn't reasonable and eliminated some, but I didn't check these by hand, so I find it likely that there's at least one holdout test that's unfair in some way. However, the overall score against holdout tests is low enough that I'm not too worried about a small number of tests being bad (if I worked at an AI lab and was trying to train next-generation models, I would be more worried about this, but I don't think it's material for our use case here).

    Instructing agents not to cheat while having a holdout set of tests didn't prevent blatant cheating that scored extremely poorly on holdout tests, but telling agents that there was a holdout set of tests they were graded against seemed to reduce the score they achieved on the agent-visible tests while increasing the score they achieved against holdout tests (without telling them this, a number of agents achieved 100% on the Pandoc tests with uselessly brittle code; on telling them there's a holdout, no agent scored 100% after 1 turn on ultra, but the holdout scores were substantially better, indicating better generalization).

    [return]
  3. There are various Substacks, YouTube channels, and other things that promise to tell you the secrets of LLM coding success, but the ROI on spending time running actual experiments isn't really there. When we looked at caveman mode, we saw that one of the biggest programming YouTubers had a video where they spent a few minutes looking into it and decided that it worked. Spending even 15 minutes looking into whether or not it really works is probably negative ROI compared to spending that time producing more content instead.

    There are various papers that discuss different techniques, and these sometimes go into more detail than most blog posts or videos but, on average, they don't necessarily have more useful information. For example, when I asked ChatGPT (5.6 Sol, Pro) to find discussions of language effectiveness with respect to LLMs, it turned up this paper on token efficiency, which has an interesting idea, but has the same issue as the caveman mode evals we discussed earlier, where it's not looking at a task that's interesting enough for the result to be relevant to me as a programmer. Just seeing what cited that paper, we find this paper by three academics on token efficiency of languages titled "The Best Programming Language for Tokenmaxxing", but compared to this post, that paper only compares four languages, uses worse models, and uses small toy problems (from something called LiveCodeBench; the cost to solve problems with GPT-5.5 is often on the order of 1000 tokens). Regardless of how well done the eval is, as we've noted in this post and in our caveman mode eval, we often see wildly different relative results when going from a small toy problem to a problem that I might care about for hobby projects or work. Also, in that paper, they note that they gave the prompt "To test your program, run exactly ./test.sh... These are the only tests I care about" and they say this is realistic because "We believe that this setup is a realistic way to study agent behavior: in everyday use, programmers don’t hide their tests from agents. Instead, programmers direct their agents to keep working until all tests pass." but, as we noted above, doing this results in brittle code that fails in the real world (or if you have holdout tests that aren't given to the agent, it fails the holdout tests at a very high rate; this problem cannot be solved by just adding a few more tests; it can perhaps be addressed via something like fuzzing or property-based testing, but how well that works is a topic for another post). I'm not saying these papers are bad or that there isn't something interesting to learn from these papers, but as a programmer who wants to know what techniques or tools I should use, I can't get that information from papers like the ones linked above.

    [UPDATE: Tom Adamczewski sent me a link to his paper, https://arxiv.org/pdf/2606.30182, which does handle a lot of the issues mentioned above. Relative to this post, it tries a lot more different tasks (which is great) and tries fewer languages and fewer ways of presenting tasks. One conclusion they draw in the paper that I think falls out of trying fewer languages is that language doesn't matter; even if you exclude the very obscure languges from the evals we tried here, we can observe a correlation between language popularity/usage and result quality; because Adamczewski's paper tries a lot more tasks, you can get a more complete picture by looking at this post and that paper combined than you can by looking at either in isolation.]

    [return]

Feeds

FeedRSSLast fetched
XML 03:55 , Sunday, 06 September 2026
35mmc XML 03:55 , Sunday, 06 September 2026
About – Bikes and Film Cameras Club XML 03:55 , Sunday, 06 September 2026
apenwarr XML 03:55 , Sunday, 06 September 2026
Arch Linux: Recent news updates XML 03:55 , Sunday, 06 September 2026
Ars Cardboard - Ars Technica XML 03:55 , Sunday, 06 September 2026
benjojo blog XML 03:55 , Sunday, 06 September 2026
BIKEPACKING.com XML 03:55 , Sunday, 06 September 2026
Biz & IT - Ars Technica XML 03:55 , Sunday, 06 September 2026
Cardinal News XML 03:55 , Sunday, 06 September 2026
Coding Horror XML 03:55 , Sunday, 06 September 2026
Cryptography Dispatches XML 03:55 , Sunday, 06 September 2026
Debian News XML 03:55 , Sunday, 06 September 2026
derailleur XML 03:55 , Sunday, 06 September 2026
EMULSIVE XML 03:55 , Sunday, 06 September 2026
flak XML 03:55 , Sunday, 06 September 2026
Idle Words XML 03:55 , Sunday, 06 September 2026
inks XML 03:55 , Sunday, 06 September 2026
joshua stein XML 03:55 , Sunday, 06 September 2026
McMansion Hell XML 05:55 , Sunday, 06 September 2026
Migratory Caving XML 03:55 , Sunday, 06 September 2026
Open source software and nice hardware XML 03:55 , Sunday, 06 September 2026
OpenBSD Journal XML 03:55 , Sunday, 06 September 2026
Q R P e r XML 03:55 , Sunday, 06 September 2026
Rene Herse Cycles XML 03:55 , Sunday, 06 September 2026
reproducible-builds.org XML 03:55 , Sunday, 06 September 2026
Steam for Linux RSS Feed XML 03:55 , Sunday, 06 September 2026
Techdirt XML 03:55 , Sunday, 06 September 2026
Tedium XML 03:55 , Sunday, 06 September 2026
The Soma Fab Blog XML 03:55 , Sunday, 06 September 2026
The Velo ORANGE Blog XML 03:55 , Sunday, 06 September 2026
Velo Orange - The Velo Orange Blog XML 03:55 , Sunday, 06 September 2026
WUVT-FM 90.7 Blacksburg, VA: Recent Articles XML 05:55 , Sunday, 06 September 2026
www.collegiatetimes.com - RSS Results for * of type article OR video OR youtube OR collection XML 05:55 , Sunday, 06 September 2026