Lese-Ansicht

China's Moonshot AI reportedly used Nvidia Blackwell chips for training Kimi K3 — company circumvented both U.S. export and Chinese import controls to acquire compute

Keeping the upper hand in the AI arms race has become a vital goal for both the U.S. and China, and Nvidia's Blackwell AI chips are one of many flashpoints in that fight. The US government bars their sale to Chinese firms, while Chinese policies block their import as the country tries to spin up an advanced AI chip industry of its own.

But as we've discussed multiple times and then some more, Chinese AI firms are quite creative with workarounds for these restrictive policies. That's the case of Moonshot AI, which has reportedly made good use of Blackwell for training the recently released Kimi K3 frontier-level model, and is seemingly looking to obtain additional access in preparation for Kimi K4.

The Information says "people with knowledge of the matter" told it that Moonshot employed two Chinese firms that have Blackwell chips in their respective datacenters despite the bilateral restrictions we mentioned. Given that those chips are scarce enough right now even when obtained legitimately, it's unsurprising that neither firm had enough of them on hand to let Moonshot train K3. This reportedly forced Moonshot to figure out how to join multiple eight-chip Blackwell servers together and across datacenters in order to harness the necessary computing power.

The report also mentions "a researcher at a major Chinese tech firm who works on model training" as stating that Kimi K3 has "started a new round of arms race" in the country's AI industry. They further added that training frontier models is difficult or impossible with the promising but slowly developed homegrown chips. By that source's account, Chinese AI accelerators remain a generation or two behind Nvidia's current offerings and are reportedly several months in backorder.

For inference work, Moonshot reportedly relies on Nvidia's China-market HGX H20, a last-gen chip that isn't blocked by trade laws on either side of the Pacific. The firm recomends setups with at least 64 H20 GPUs for running Kimi K3. Those requirements, combined with that frontier model's desirability, meant that Moonshot quickly ran out of computing capacity to run K3 and currently has subscriptions on a waiting list. Given it's an open-weight model, and that its weights were released this week, many other inference providers are serving it, perhaps alleviating that bottleneck.

Meanwhile, White House Director Michael Kratsios claimed last week in a tweet that that Moonshot AI both "acquired GB300-equipped servers and has accessed GB300s in Thailand." While buying Blackwell chips is illegal, renting them is apparently fair game, at least until the proposed Remote Access Security Act takes effect. That law is designed to prevent the rental loophole by treating remote access as an export event. There's no telling exactly how the U.S. would enforce this law across other jurisdictions, though.

At any rate, the Department of Commerce is formally investigating if Chinese firms are accessing advanced U.S. chips like Blackwell GPUs, and that's likely to be an ongoing point of contention as the war for frontier model supremacy continues.

In China, it's an open secret that many of the country's high-level own or have access to Blackwell and other advanced chips, but despite all the trade restrictions and pushing the usage of local-made chips, the CCP has seemingly yet to crack down on said AI players. Some have theorized that the turning of this blind eye is intentional so Chinese firms like Moonshot can catch up to the likes of Anthropic and OpenAI.

  •  

Mystery reviewer 'finds' Nvidia RTX Spark prototype laptop and puts it through its paces — Microsoft Surface Laptop Ultra with Nvidia N1X chip shows promise, though prototype warts are still quite visible

Everyone loves previews of pre-release hardware that just happened to fall off the back of a truck — particularly when it's a piece of kit that's been lighting up news headlines. Fouquin, a TechPowerUp reader, "accidentally" stumbled upon a Microsoft Surface Laptop Ultra prototype that ensconces an Nvidia RTX Spark N1X SoC.

As a quick recap, the N1X is supposed to herald a generation of "AI laptops", meaning it's an alternative to the RTX Spark desktop, Mac Studio, Ryzen AI Max machines, and the Macbook Pro. This is due to the fact that they're all designs with large pools of unified memory and competent GPUs, making them suitable for non-datacenter AI work. The N1X packs 10 MediaTek-designed ARM CPU cores in a 5p5e configuration with multi-threading for 20 threads total, plus a GPU purportedly equivalent to a desktop RTX 4070 or a mobile RTX 5070. It can be wired up to 128 GB of LPDDR5X onboard in a unified pool, though the tested machine only has 24 GB.

Fouquin notes the limitations of their "review" and notes they're not an AI person, though they tried running the Phoronix AI test suite nonetheless. This yielded a mixed bag of results, as very few tests would reliably work, seemingly due to pre-production drivers. Some Vulkan tests did run properly, and the reported performance appears to be indeed in the ballpark of an RTX 4070, though there are two very important caveats.

First, this system and its drivers appear to be very buggy and unoptimized, particularly around power management. Fouquin says the system came with ancient 591.33 drivers from November 2025, eventually upgraded to 616.00 preview drivers with CUDA and Vulkan support. Apart from that, the tester found the machine's TDP spends and power plans behaving erratically, and none of the Phoronix CUDA tests produced results.

Second, head-to-head benchmarks are hard to come by, as RTX-series desktop cards have limited pools of VRAM compared to the 24 GB on the test machine. All told, there's a fair chance that final hardware might perform better, especially considering this pre-production machine displayed many outstanding issues with the LCD display, idle current draw, and that plugging it in or running it from battery didn't make any difference. Interestingly, despite the 616.00 drivers adding features, performance actually took a significant turn for the worse.

The Cinebench 2024 multi-threaded CPU test produced a result of 1386 points, trailing the 12-core Macbook M4 Pro, while Cinebench 2026 yielded 5771, a fair bit behind the Apple M4 Max. Fouquin did run some GPU tests, and the best results in 3DMark were with the older drivers. Port Royal showed 26-31 FPS on older drivers, while Nomad got 20-21 FPS. Meanwhile, Unigine Superposition actually improved with the 616.00 software, netting a maximum of 56-58 FPS.

Compared to desktop cards, those results would roughly match RTX 3060 or RTX 3060 Ti, once again underscoring the pre-release nature of the Nvidia drivers and related software. As if more illustration was needed, the Balanced power plan almost always produced the best results.

The story isn't much different in gaming, as Fouquin says that while most of their library of games ran fine, the performance was very sluggish, with multiple-second stutters and with GPU clock speed "bouncing all over the place." Helldivers 2 would outright crash the system, as did the GPGPU integer tests. Fouquoin believes these are all issues specific to this particular implementation of the N1X rather than a problem with the platform altogether. It's worth adding that presumably the games ran under the ARM emulation layer, which can introduce its own set of issues.

N1X chip and fixable warts aside, Fouquin actually liked the machine, remarking on the high build quality, the keyboard feel, touchpad input, and display quality. Given there were previous claims the machine would be serviceable, Fouquin put those to the test and opened it, finding that while that claim is technically true, "many snap-on thin aluminum sheet panels covering all the primary components (even encasing the SSD) means actually attempting to service the laptop will lead to much bending and/or breaking." Here's to hoping that this too is fixed before final release.

  •  

Daring coder gets Doom running with regular expressions at 180 seconds per frame, like playing 'correspondence chess with a shotgun' — nearly 14 million substitutions to render a frame at 80,000 substitutions per second

Running the 1992 Doom on the most random piece of hardware around has become probably the most common unofficial programming challenge. We've seen the game running on anything from toasters to an Anker charger, and even a pregnancy test. Enterprising coders also get it running on the weirdest software possible, and just recently Artem Lytkin got it running in regular expressions.

Developers in the audience are probably recoiling in horror, as that sentence is definitely cursed. You see, regular expressions (regexes) are a utility language used in programs for finding and replacing text. They're incredibly powerful, but the syntax is often said to be write-only, as it looks just like gibberish. For example, /.*(\d{4}).*/g would find the "2026" in "Tom's Hardware 2026 articles." They can be exceedingly complicated, as they include conditional statements, elaborate character-jumping, and substitution rules. However, this also means they fulfill all the technical requirements to be a programming language.

Leveraging those capabilities, Lytkin created a 96 MB plain-text string that contains sections for the virtual CPU's registers, some RAM, a video output (framebuffer), the game's WAD data, plus I/O and other bits and bobs. Once it's all started, the regex will start text-matching and substituting characters in the string to pretend they're the numbers in each processor's register, then accessing and writing to the "memory," so on and so forth.

As you can imagine, this is spectacularly slow, and Lytkin says that it takes around 180 seconds to produce a single frame of game output. Each of those needs nearly 14 million substitutions, though (a) it actually works and (b) Lytkin claims it's byte-identical to the actual Doom output running. You can even control the game with the keys, but as the daring coder poignantly illustrates, playing it "is closer to correspondence chess with a shotgun than to a twitch shooter."

Particularly nerdy devs will be happy to know how Lytkin wrote the memory access part: essentially a binary tree, by jumping from "branch" to branch using standard regex character-jump instructions. This avoids having to scan the entire 96 MB of text repeatedly just to find the "#M" marker bookending it. Lytkin notes the challenge was not about whether it could be done, but whether the game would run "before the heat death of the universe," as the engine fires up 80,000 text substitutions per second.

The doom-regex repository is here, and you can download a demo to run it on your own computer. The project's website demonstrates how the regex machine runs in both visual and text format, and it's mesmerizing to watch. It's quite reminiscent of the time we spent watching defragging utilities do their thing when we were young'uns.

  •  

Moonshot AI releases weights for Kimi-K3, firing a shot across the bow of OpenAI and Anthropic — open-weight model performs almost as well as frontier models while being 2-3x easier to run

Well, the artificially intelligent cat is out of the bag. After publishing a blog post and API documentation for the minty-fresh Kimi K3, Chinese outfit Moonshot AI delivered on its promise to release the model's weights for free, meaning that most anyone with a contemporary rack of AI GPUs can run it and charge for it, with few restrictions.

This is quite the shot across the bow of the big AI players, namely but not only Anthropic and OpenAI. Those companies' latest models are Claude Fable and GPT-5.6 Sol, respectively, and it happens that Kimi K3's capabilities outright beat previous generations of Claude and GPT in Moonshot's benchmarks, and closely trail Fable and Sol— all while seemingly being around 2-3x cheaper to run, up to 10x if a particular query lands in the cache. Moonshot's technical write-up seemingly backs up the benchmarks published last week, as the company reveals which exact software was used for testing.

For its inference cost comparisons, Moonshot says that its costs "are measured internally" versus the publicly available token pricing for other companies, but the figures are quite impressive. For input, Moonshot charges $3 per million tokens for Kimi K3. Meanwhile, Fable costs $10/1M, while Sol goes for $5/1M. That figure is standard non-cached input and is already pretty good-looking, but Kimi K3's caching structure seemingly has a 90% hit ratio for coding tasks, turning those $3 into $0.30/1M if your use case hits the cache a lot. The story is pretty similar for output tokens.

One of the likely reasons why Kimi K3 is so efficient is that it uses a mix of MXFP4 for weights and MXFP8 for input activation, both data types with relatively low precision and thus amenable to running on far less VRAM. Out of Kimi's 2.8 trillion parameters, only 104.2 billion are activated at a time, too.

Interestingly, Moonshot's write-up only mentions Nvidia's H20 being used for running Kimi for some coding tests, a fairly low-end chip by today's standards. That GPU doesn't have native support for MX floating-point types, unlike the export-controlled Blackwell B-series chips.

In turn, this can mean that Kimi K3's optimizations make it particularly amenable to run on lower-end hardware, but it's an equally reasonable guess that running it on something like Nvidia Blackwell or other MXFP-native silicon could make it even more cost-effective than in the presented benchmarks. We'll have to wait for more official figures to confirm this speculation.

Additionally, Kimi K3 doesn't use a conventional ever-expanding key-value (KV) store, instead relying on a fixed-size state handler called Kimi Delta Attention, again theoretically saving both on VRAM and execution time. Its mixture-of-experts (MoE) is particularly sparse with only 16 activated at each time out of 896, further contributing to inference cost reductions. Broadly speaking, Moonshot went for optimization at every layer of inference to avoid unnecessary overhead and bring inference cost down.

This is could be bad news for OpenAI and Anthropic, given that most anyone with decent AI GPUs can now become their direct competitor, and the fact that Kimi K3 is open-weight also gives off the impression that "free" software is nearly as good, and far cheaper to run, than its proprietary competitors. It's worth noting that open-weight does not mean open-source; the training process and dataset are still Moonshot's special secret sauce.

  •  

Game Compressor can save you hundreds of GB across your game library — as storage prices remain high, utility leverages Windows' built-in LZX compression for substantial space savings

With NAND prices flying into the stratosphere, buying a large SSD has become a cold-sweat-inducing event, with even the cheapest 2 TB models now approaching $250 — and desirable performant models going for well beyond that. With no end to either memory or storage shortages in sight, the next best option for gamers is to save on used disk space — and the recently-upgraded Game Compressor utility is a fine tool for that end.

Depending on the game in question, savings can be massive, with ARK Survival Evolved reportedly going down from 169 GB to 91 GB (54% of the original size) and Crimson Desert slimming down by 32 GB. Not every game bears high compressibility, but if even just one of your titles goes down by tens of gigabytes, that's enough to install another. This is handy on most every PC, and even more so for handhelds and machines where disk space is at a premium.

How do you log in to your PC?

The application, which costs $5.99 / 5.99€ on Steam, leverages Windows' built-in LZX folder and file compression to transparently compress on-disk games — sometimes with massive savings. I've used NTFS compression for over two decades with no ill effects, and I occasionally enjoy improved game load times. That might sound counterintuitive, but with a compressed game, there's less I/O data to shuffle around. The only cost is a little bit of CPU overhead, which is almost guaranteed to be available as almost no games can fully utilize all the system's cores.

While it's easy to say that Game Compressor is just a UI for Windows' built-in functionality, the amount of convenience is impressive. You first set up a number of watched folders, of which presumably the ones for Steam games will be filled in. After Compressor scans which games are available, it will actually preview what kind of space savings you can expect, to save you the hassle of going through a long-running task for no benefit.

Once you select your games to compress, the tasks go into a queue, and can be paused or cancelled at any point — a convenience the Windows command line doesn't offer. You can also selectively decompress games later, and there's a log of operations just in case something goes awry. As you might expect, Compressor tries to detect which files (such as videos) aren't worth squeezing further.

One of Game Compressor's neat tricks is that it can monitor games that have received patches to reapply compression. Normally if, say, a 50GB game file gets a patch, it could be decompressed post-patch — this way, the space savings are kept in place.

Just about the only "con" to this notion is that it's not advisable to use compression on DirectStorage-enabled games, because that API relies on IO-to-VRAM operations and adding a CPU decompression step could introduce stuttering. Those titles are few and far between, though.

Some of you may be scoffing at this article and pointing out that ever since Windows 10, the "compress" command supports LZX compression. But paying just $5.99/5.99€ for Game Compressor will save your time, add visibility into compression status, and automatically recompress patches. If you'd rather strike a middle ground, there's also the open-source CompressGUI — which is also pretty nifty, though not quite as feature-filled.

  •  

OpenAI's HuggingFace breach heralds an unprecedented age of AI cyber warfare — contemporary LLMs have caused massive upheaval in cybersecurity, and it's only going to get worse

This week, OpenAI revealed that during a purported capability test with no safeguards, a set of bots, including its upcoming GPT-5.6 Sol, hacked their way out of their locked-down network and into Hugging Face's production infrastructure. Only months ago, Anthropic made a splash in the news when its CEO, Dario Amodei, said its new Mythos model had cyberwarfare capabilities, which prompted a strong reaction in the AI space and among government entities, most notably the U.S. Bureau of Industry and Security, which issued an export-control order for the model, which it has since slightly loosened.

Despite the bluster that AI CEOs like Dario Amodei and Sam Altman make over the capabilities of new models, frontier-level LLMs are now proven to be stalwarts in cybersecurity.

It's a fact that LLMs adept at coding are equally suited to spotting security vulnerabilities in source code. Exploits fall almost universally into a handful of categories, and LLMs are literally designed for pattern recognition. So much so that the Zero Day Clock (ZDC) project currently registers a zero-day exploit's time-until-exploit at negative 8 hours, meaning that malfeasants using AI bots are now routinely finding vulnerabilities before actual security researchers or vendors.

Driving that point home further, 81% of disclosed vulnerabilities are zero-day, and only a tiny portion even go one week before being exploited. All of this only counts security exploits with public disclosure. Predictably, among many advisories, the ZDC recommends preemptively using AI in every step of the development process. The industry-standard 90-day disclosure window, still used by most vendors' bug bounty programs, appears effectively dead, leaving looming implications for the rest of us.

AISI report on frontier models

(Image credit: UK AISI)

Back in March, the UK's AI Security Institute published a paper where it tested contemporary AI models in security exploitation scenarios, and the results were sobering. Most bots went through four out of nine exploitation milestones. A more recent comparison, which included Claude Mythos 5 and GPT-5.6 Sol, showed that every single milestone up to and including full network takeover was reached, at least in one of the many attempts.

Aikido also published its latest cybersecurity benchmark results on July 16. In this case, the test was having the bots recall (find again) multiple known exploits in a varied set of software. The results were sobering, with the GPT-5.6 variants in the lead at an 88.5% recall rate. Perhaps most importantly still, the price per exploitation was incredibly cheap — even GPT-5.6 Terra came in at only ~$750 per full run.

This study also revealed that even with less-powerful, cheaper models, you can reach the same number of total exploits if you run them enough times. Considering these aggregate results, GPT-5.6 Terra at $247/run was just as good as GPT-5.6 Sol Max at $870/run.

Aikido frontier model benchmark pricing

(Image credit: Aikido.dev)

Aikido also redid its testing after the debut of Moonshot Kimi K3, to staggering results. Kimi K3's results were similar to OpenAI's GPT 5.6 Terra, while being 15% cheaper. Compared to OpenAI's leading model, GPT-5.6-Sol, the difference is even starker, with Kimi K3 being four times cheaper when discovering cybersecurity vulnerabilities.

The fact that an open-weight model is often trading blows with even the über-expensive offerings from OpenAI and Anthropic is rattling Western closed-source companies. Why pay Big AI for pricey models when you can just rent servers and run Kimi K3 instead?

Aikido Kimi K3 benchmarks

(Image credit: Aikido.dev)

Furthermore, Moonshot is not the only Chinese AI company developing frontier models, as Z.ai's GLM 5.2 (also an open-weight model) and 360 Security's Tulongfeng are reportedly adept at security workloads.

So, what are companies expected to do? The answer, perhaps unfortunately, is deploying AI agents of their own. According to Hugging Face, the recent intrusion by OpenAI's bots was stopped with its own fleet of AI agents. Given the speed of the attacks and the fact that HuggingFace's defenses were mostly made up of other AI agents, it's quickly becoming clear that it is infeasible for humans to keep up.

Google AI Threat Defense, MindGard, and HiddenLayer are but a few of the many names popping up in the AI cyberdefense arena. Besides the UK AISI, the European Systemic Risk Board and the Australian Cyber Security Center have both issued concerning advisories on the situation.

Using AI for defense raises yet another question: When both attack and defense are swarms of non-deterministic algorithms, there will be a point where we won't even know what the AI models are doing on either side, or at least not until it's too late. These scenarios were originally envisioned by classic Sci-Fi authors — now it's a reality that, for better or worse, the cybersecurity industry must face.

  •  

Oregon will start charging service providers for undersea cables using its sea floor — millions accrued over years will go toward state schools

After decades of charging only a one-time fee for undersea cables along its coast, Oregon's lawmakers have recently been discussing a missed opportunity for tax revenue. The state is set to charge cable operators a one-time fee for cable installations within three miles of its coastline, and will be funneling that money into its Common School Fund.

The currently existing fee structure dates back to 2001 (25 years ago), and is simply a $5,000 application fee. The new pricing is still undergoing discussion, but the simplified version is $3 per linear foot of cable, plus $7 per each section of bore pipe — the protective structure that covers the portion closest to shore. Oregon's Department of State Lands (DSL) estimates that on a 20-year lease, this would bring in millions per undersea cable over that time span. The DSL did some rough math with Amazon's Bitfrost cable, and came up with a grand total of $1,450,380 for two decades, paid in one lump sum.

It's worth noting that the aforementioned rates were revised from an earlier, more expensive proposal that would take bore cross-section into account (a calculation that was dismissed as too complicated and, presumably, too onerous). The state wants to continue attracting investment to its tech sector, which includes around 140 datacenters and 16 undersea cables of various types. Oregon also doesn't have sales tax, and its enterprise zoning incentives has made it an attractive location for players like Amazon, Microsoft, and Google.

The DSL calls the revised rate "competitive-to-low," reinforcing the notion that Oregon is trying to stay competitive with its southern neighbor, California, which charges an estimated $5 per cable foot on an annual basis.

It's worth noting that spending a million or two over two decades is peanuts compared to the initial outlay required to put an undersea cable down in the first place, especially one coming from across the world. These reportedly ring in at around $250 million for a transatlantic connector — and a whopping $400 million to cross the Pacific.

Although it's not set in stone, it seems that already-approved cables might be almost entirely exempt from the new fees. Some recent contracts, such as the one for Amazon Bifrost, did include a clause that said that any fees applied by new state laws must be paid for. However, the contract also included a specific $300,000 "out" for this clause — which Amazon paid at the time.

  •  

OpenAI's GPT-5.6 Sol and unreleased AI models break out of testing environment in 'unprecedented cybersecurity incident' — rogue agents hacked HuggingFace's production servers with 'thousands of individual actions across a swarm of short-lived sandboxes'

Not too long ago, Anthropic CEO Dario Amodei described Claude Mythos as capable of cyber-warfare, spawning all sorts of mythology that became popular reading at investors' desks, and even at the U.S. government table, which issued an export control order for it. Now, it's OpenAI's turn to do the same, with a blog post describing how a bot cyber-gang including its upcoming GPT-5.6 Sol and an "even more capable pre-release model" all broke out of their virtual cage and eventually into Hugging Face's production infrastructure, during an attack capability test. OpenAI described the snafu as an "unprecedented cyber incident."

Some may decry these revelations as marketing stunts meant to assuage increasingly suspicious investors, yet there's no denying the proven effectiveness of AI models' pattern recognition when it comes to finding vulnerabilities in software. Many professionals have even decried the industry-standard 90-day security vulnerability disclosure window as obsolete.

The usual modus operandi, however, involves granting a bot access to source code repositories so it can scan them for mistakes easy for our poor human eyes to miss. In the test that OpenAI described, though, Sol and its managed to find multiple vulnerabilities by analyzing its containment network and then Hugging Face's production machines, all seemingly without any initial access to source code.

OpenAI had set up an isolated environment in which Sol and its posse supposedly couldn't reach the internet. Importantly, the bots were operating "without production classifiers used to prevent models from pursuing high-risk cyber activity" — a lot of words for "no guardrails," essentially. The only network access they had was a software package installer proxy. In theory, this likely also prevented them from using any Retrieval-Augmented Generation (RAG), meaning they couldn't search for solutions online.

Given that bots have limitless patience and no need for naps, they eventually dug their own virtual tunnel to the internet at large by way of finding and exploiting a zero-day vulnerability in the package proxy software, an impressive feat if performed as described. OpenAI says it's disclosed the technical details to the vendor in question.

Once out on the internet, much like cheating students, the bots reportedly reasoned that instead of performing the ExploitGym challenges (a set of security capture-the-flag exercises for bots), they could simply find the solutions, and became "hyper-focused" on that exact goal. They deduced the data existed at Hugging Face's servers and proceeded to break into the premises using "stolen" credentials and additional unspecified zero-day vulnerabilities to gain remote code execution privileges.

Interestingly, Hugging Face's own writeup says the attack comprised "thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services" — a live illustration of a scenario in many a science-fiction book and flick, proving once again that life imitates art. Hugging Face says that none of its customer-facing services are compromised, and that it stopped the attack using AI capabilities of its own — another familiar cliché.

For its part, OpenAI says it's going to add controls to the bot "at the cost of research velocity," and that the incident "points to the need to further strengthen our model’s alignment, cyber protections during evaluation time, and monitoring during internal testing" — a statement that might be an obvious fact, or intended to oversell the model's capabilities. Whichever it may be, the fact is that when it comes to digital security, bots have proven quite capable. After all, even if they're not virtual Bruce Schneiers, they just need to be marginally more effective than average humans.

  •  

The enduring paradox of the AI economy — models get better and more efficient, yet costs can still easily spiral out of control

Much has been written regarding the questionable economics of the AI space, but most of the discussion revolves around high-level concepts like market shares, datacenter investments, and power expenditure. The general expectation of technology is that it gets cheaper as it improves, but the AI space has a rather peculiar problem: Usage costs are actually soaring even as models get better, as detailed in a write-up at VentureBeat.

With all the advancements in models over the last two years, having a bot that answers questions of simple-to-moderate difficulty is now old news, as they all do that with reasonable accuracy. Agentic workloads are where the real potential is at, letting bots loose on multi-stage, repeatable tasks that would previously take a human days, if not weeks, to perform. Grant a bot access to your billing system, some Excel spreadsheets, and your CRM, and ask it something like "who are my profitable customers by category and what are their trends" becomes child's play.

While it's trivial for you to ask that question and get a fairly accurate answer back in a couple minutes, behind the scenes there is a lot of processing going on, and far more than you'd expect. AI computing time is measured in tokens — a short question and answer might take somewhere between 200 to 2,000 tokens, and one that requires the models to do some internet research might be around 1,000 to 4,000. An agentic task, though, can easily spend millions of tokens on a seemingly innocuous request. How? Token amplification, a recently coined term.

In a simplified manner, because a model has no memory or cognition, every time you ask it another question in a conversation, it will re-load and process the entire exchange — everything you wrote, everything the bot replied with, and every file you uploaded. That means that additional questions in a long conversation progressively get more costly. Each question in a conversation might only need 500 tokens by itself, but reprocessing all the previous information adds up, so the second one might spend 600, and so on and so forth. The conversation as a whole uses up the cumulative number of all individual interactions, and as an added penalty: Response speed also tends to get slower as chats drag on.

The aforementioned task of generating a report will have to be run in stages, say three for looking up Excel sheets, four for the CRM, perhaps a half-dozen web searches for contextual information about products, and a good dozen intermediary processing and calculation steps. Each step tacks on potentially several thousands of tokens, and by the end of it, you may be looking at millions of tokens cumulatively spent for all the steps combined.

Costs, then, can spiral out of control very quickly — and that's assuming a best-case scenario where all of the data is easy to interpret, the model won't need any retries, and you're using a moderately powerful model rather than something like Claude Opus.

In financial terms, this means that one innocent question might spend 280,000 tokens and cost $1, with ballpark estimates at current prices. That may not sound like much, but it was just one task. If you have this scheduled to run every 15 minutes for a dashboard, then suddenly this costs $96 per day, or $2,880 in a month. And that task example was fairly simple; a tricky multi-stage report might need millions of tokens, turning that $1 into, say, $10, or a grand total of $28,800. For the one report, for a limited number of users, in one department of a corporation.

That's the key reason why Anthropic, Microsoft, OpenAI, and almost every other player moved their major offerings to usage-based billing back in April, and severely limited token spends in fixed-rate plans. This generated many sticker shock events, as developers around the globe found out how much their vibe coding actually cost and were left in a daze.

In a large company using many AI agents, such costs can become prohibitive and oftentimes costlier than standard-issue human employees. Firms like Uber, Microsoft, Amazon, and Walmart, among many others, have responded by curbing AI spend. Token expenditure is suddenly an issue for both financial and engineering departments, as cost control becomes paramount. As VentureBeat succinctly puts it, for agent-heavy companies, "a prompt redesign is a margin event," and more illustratively, "a poorly bound agent loop is an outage with a credit card attached."

Although most AI outfits still offer fixed monthly pricing plans, they often come with harsh token limits. As is often the case with most fixed-pricing services, however, the most capable bot-wranglers will also inevitably be the ones burning through the most tokens, not just due to their intensive usage per se, but because the type of tasks they will perform will be precisely the ones hardest hit by token amplification.

This throws a wrench in the usual financial model of the casual users easily offsetting the cost of the small percentage of professionals — the pros' usage can take a serious bite off monthly profits, or chew through them entirely.

The AI companies aren't sitting still, and getting the per-token cost down is likely to be the primary task for most of their engineering teams, at this point. Prompt caching, model routing, batch processing, semantic caching, and context window management are among many technical measures that can massively cut down on token spend, each of those netting a two-digit percentage drop.

And yet, costs keep soaring for the simple reason that the smarter the models get, and the more adept people become at using them, the more complex and long-running the agentic tasks will become, adding multiple orders of magnitude to token counts.

For the time being, it looks like a losing race, despite recent advancements like Deepseek V4 Pro and V4 Flash deeply undercutting comparable Western models at up to a reported 17x for the former and 25x for the latter. Out of the major Western players, only Anthropic has predicted its first profitable quarter, though much like anyone else, it's still in the hole for many a billion for cumulative spend.

  •  

Security engineer ports password cracker hashcat to Gameboy Advance — 16.8 MHz chip can perform a meager 727 hashes a second, 30 million times slower than a modern rig

In the modern age, password cracking is an activity that generally begets quite powerful hardware, usually high-powered GPUs. Hashing algorithms have gotten trickier, and brute-forcing them has become computationally expensive. It's only natural, then, that a new software project has come up to port the common hashcat cracking utility to the mighty... Gameboy Advance (GBA).

I "ported" @hashcat to the GBA, the ultimate password cracking rig ever (my dumbest project yet).The monstrously powerful 16.78 MHz ARM7TDMI chip does an astronomical 727 SHA256 hashes per second!! which is about thirty million times slower than a modern cracking rig. 350 days… pic.twitter.com/kWDD2B27KMJuly 17, 2026

The gba-hashcat software is the brainchild of security engineer solstICE (Ice), to whom the question of "why" is seemingly answered with "why not." Running on an original GBA, the app can power through SHA256 hashes at 727 h/s. At that rate, the author estimates that one year's worth of gba-hashcat would be equivalent to one second of a modern GPU-accelerated rig.

The GBA's main processor is an ARM7TDMI chip, a 32-bit RISC affair clocked at all of 16.8 MHz. It connected to most of its data buses and memory on a 16-bit bus, though, and the machine only carried 288 KB of RAM, plus 98 KB of VRAM. Most password-cracking software makes use of precomputed tables to help speed the repeated attempts along, but given that standard GBA cartridges are limited to 32 MB in size, all that Ice could include was the ignis-1M (one million) word list that rings in at 8 MB.

The software itself is quite small, given that Ice made good use of the Butano engine, a library that allows developers to easily write GBA games using plain C++ code. The gba-hashcat user interface is quite sparse, showing only an intro screen, progress statistics, and the current password attempt. It's not like you need much more for this application, save perhaps for a "please be patient" message.

In the GitHub commit, Ice calls gba-hashcat "[their] dumbest app of all time." That's quite the statement, but X commenters were quick to request networking support so they could make a cluster of password-cracking Gameboys, an idea that may have merit considering the price of graphics cards these days. At any rate, joke all you want, but for the one really important password retrieval I was once tasked with — recovering access to an Excel payroll sheet — even the original Gameboy would have sufficed. The password was "123".

  •  

Linus Torvalds rebukes anti-AI stances in the Linux kernel code review process, says 'Linux is not one of those anti-AI projects' — creator embraces AI as just a tool and 'clearly a useful one'

AI-generated slop code has been a plague for some open-source projects, namely but not only Gentoo Linux, Curl, and Ghostty, limiting or outright banning LLM-created contributions. And yet, just like both the models themselves get better and the people using them become more considerate, the landscape may be changing. Linus Torvalds, Linux's creator and kernel manager, has seemingly taken an accepting stance of AI-assisted tooling.

In a long comment on the Linux kernel mailing list, Torvalds spelled it out fairly clearly: "I realize that some people really dislike AI, but this is an area where I'm willing to absolutely put my foot down [...] Linux is not one of those anti-AI projects, and if somebody has issues with that, they can do the open-source thing and fork it. Or just walk away."

This discussion came regarding the usage of Sashiko, an opt-in (per-mailing-list) and apparently quite effective multi-stage code review tool that analyzes kernel patches. The project page says the tool can find 53.6% of bugs on proposed patches, and argues that that metric already puts it above human level, as the patches in question already supposedly went through human review. The false positive rate "is harder to measure," being pinned "within 20%." Crucially, Sashiko only comments on patches, and does not take action by itself.

Developer Laurent Pinchart suggested that Sashiko's output be triaged before comments were sent out to patch authors, basing the notion on the Software Freedom Conservancy's guidelines on AI-generated code. Google's Roman Gushchin, one of Sashiko's creators, pointed out that doing that would undermine the utility of the tool, and that Pinchart's position was quite anti-LLM — a sentiment echoed by Linus Torvalds in his reply.

Torvalds put it clearly: "AI is a tool, just like other tools we use. And it's clearly a useful one. It may not have been that 'clearly' even just a year ago, but it's no longer in question today," a statement that reflects his changing stance on the matter since he initially dismissed AI tools as overhyped back in 2024. He also noted that Linux is not a "social warrior" project, and that it's always been about improving technology.

To drive the point home, he remarked that the tool "keeps finding embarrassing bugs," adding that he "will very loudly ignore people who try to argue against other people from using it," while highlighting the software's rapid evolution.

Perhaps quite poignantly, Torvalds remarked that resistant developers could use some self-awareness, as "it's not like natural intelligence is always all that great either," underscoring the fact that while AI tools may not be perfect, they generally only need to be good enough for their respective use cases. The fact that Sashiko seemingly finds errors in code that underwent human review is quite illustrative.

  •  

Scientists synchronize 105,000 nano-oscillators in just 45 nanoseconds — paving the way for a highly efficient and fast alternative to transistors

"Oscillator-based computing" is a term that doesn't make many headlines, but this area of computation is evolving and showing promise. The latest impressive development comes from an experiment in which boffins managed to synchronize 105,000 nano-oscillators in just 45 nanoseconds, reportedly using very little energy.

In layman's terms, the entire grid of tiny magnets, once perturbed, synchronized itself entirely within 45 ns, all just using the magnets' inherent spin — think of ripples on a water surface. Each oscillator measures 10-20 nm across, and the 105,000-count result is nearly a 1000x upgrade over the previous demonstration with 64 oscillators, proving that the technology can be scaled. In this new experiment, synchronization time barely increased with additional oscillators: it was 10 ns with 100 oscillators and rose only to 45 ns at 105,000.

What this means for computing is that grids can solve certain classes of problems that lend themselves to representation via propagating waves, directly or indirectly. Broadly speaking, most anything involving waves, statistics, approximation, and pattern recognition is eligible. The article mentions Ising machines and reservoir computing as being implementable by oscillator grids. At some point, the grids could become programmable by manipulating the oscillators' frequencies, phases, and coupling strengths. The result is then read by measuring how the grid settles into a synchronized state.

Going from there, practical applications include high-speed communication networks, financial and scientific modeling, real-time data analytics, and even AI acceleration. The research paper specifically notes that the grids could operate at tens of GHz and spend comparatively little energy doing so. The 45-nanosecond figure for the oscillator grid to stabilize would be roughly analogous to the time it would take a regular CPU to perform one calculation across an entire matrix.

Unlike quantum computing, which requires extensive and difficult error correction to maintain coherence, the oscillator array produces an exceedingly clear signal once it settles. The quality factor of the oscillator experiment was over one million, meaning the resulting wave frequency was well-defined and easy to read — think of the exact pitch carried by a tuning fork. To get the full details, be sure to read the research paper here.

  •  

US gov't allows Chinese telecom giant ZTE to purchase Nvidia H200 AI chips — firm joins Alibaba, Tencent, and ByteDance in access to Hopper tech

The Sino-American chip wars have resulted in many back-and-forth salvos and negotiations as the countries try and strike a balance between technology access and trade. Currently, both sides have set respective import and export controls, letting specific companies on a case-by-case basis. Today, Reuters reports that Chinese telecoms giant ZTE and server firm Maginfra have received U.S. approval to buy Nvidia's last-gen H200 "Hopper" chips.

ZTE joins a club that counts Alibaba, Tencent, ByteDance, and JD.com among the roughly 10-strong group of Chinese companies with U.S. clearance for those purchases. Additionally, an apparent subsidiary of Kingsoft Cloud got approval to buy AMD accelerators equivalent to Nvidia's H200, presumably Instinct MI300X-class chips.

Over on the Chinese side of the table, Reuters remarks that there's no word on whether the respective authorities will give ZTE the go-ahead for import, as the country has taken on a protectionist stance as it tries to grow its own chip industry. The country has discouraged firms from purchasing foreign tech and has instead pushed companies to acquire homegrown accelerators. Huawei in particular has made great strides both technologically and financially.

But even with those domestic production initiatives, the Chinese hunger for AI silicon is so deep that six months ago, Reuters said the nation's tech firms had more than two million H200 chips on order, far more than what Nvidia had on hand at the time. We'd venture that hunger has barely subsided.

ZTE might not be a familiar name Stateside, but the corporation is one of China's largest telecommunication conglomerates, and among many other ventures, it sells all sorts of carrier network gear that's installed worldwide, along with corresponding client-facing equipment, including phones and IoT equipment. Like most any sizable technological venture, ZTE has joined in on the cloud computing and AI push, so it needs accelerators to make those ambitions reality.

The current status of the AI chip trade situation is roughly that the U.S. allows Chinese firms to buy AI chips up to and including the Hopper family (meaning no Blackwell chips), with a 25% export tariff, though final decisions are made on a case-by-case basis. Over on Chinese shores, Beijing's authorities play their cards close to their chest and dole out approvals as they see fit, with no clear rules seemingly set. But China is, of course, a global power with trade connections to most everyone, so interested firms were able to get their hands on Blackwell chips through various creative (and potentially illicit) means.

Whether this change will actually clear the way for any great volumes of H200 accelerators to make their way into ZTE's data centers remains to be seen. CNBC cites a U.S. trade official who today stated that "very few shipments against licenses for H200s and equivalents have taken place. It’s a very small quantity of chips" during a congressional hearing. If H200 shipments become material to Nvidia's bottom line, we'll almost certainly hear about it in future comments or earnings reports.

  •  

Modder successfully runs Counter-Strike clone at 60 FPS on the original Sony PSP — created his own Rust-based 3D engine to power 480 x 272 gameplay, also works on PS Vita

Barely a day goes by without some retro system developer accomplishing a technical feat that beggars belief. OpenStrike is the latest mad project, creating a proof-of-concept reimplementation of the seminal Counter-Strike for the Sony PlayStation Portable (PSP), a 22-year-old system that arguably has no business running the game.

The port was made by Yifeng Wang (aka doodlestrike), who apparently created his own Rust-based 3D engine called Pocket3D for the affair, along with the PocketJS JavaScript engine for the game rules and UI. Despite the proof-of-concept status, Wang says that elimination matches with bots are already playable, though the game's famous buying stages aren't yet functional. The full set of eight original Counter-Strike maps have already been tested, and interested modders should easily be able to roll their own.

Yes, that's Counter Strike on a PSP!Yes, that's DevTools inspecting game UI!Yes, it's PocketJS! Clean room implemented FPS engine, fully open JavaScript mod API, ~12MB RAM footprint, 60fps, fully open source. ⚡️ pic.twitter.com/OprLN3E71JJuly 10, 2026

The engine manages to run the game with bots at a steady 60 FPS at the PSP's native 480x272 resolution, thanks to pre-processing graphical assets to directly bake lightmaps into vertex colors (as opposed to real-time or on-startup calculations). Interestingly, Wang kept the old-school BSP (Binary Space Partitioning) rendering method to cull non-visible areas rather than use a fancier modern algorithm, lending credence to the "don't fix what isn't broken" motto.

The game also runs on the PS Vita, a system with 4x the effective resolution of the original at 960x544, with native graphics, without upscaling of either 3D rendering or 2D assets. You should also be able to run the project as a standard desktop game, and naturally, it's playable on the popular PPSSPP emulator. To get the project going, you'll have to provide your own copy of the Counter-Strike asset data, similar to using Doom's original WADs in one of the many ports.

The project's entire technical architecture is impressive, as Wang went way beyond the scope of "make a 3D engine that load maps." Both the Pocket3D and PocketJS 2D engines are generic, and OpenStrike is merely the first "product" that uses them. The code runs on a server/client/event architecture, and the event handling (ex: shooting) is decoupled from the rendering core, preventing nasty FPS dips. The game is testable offline, and most everything is trivially moddable. You should check out the OpenStrike repo if you're interested.

  •  

Colibrì proof-of-concept gains frontier-level 1.5-TB AI model — novel approach runs on only 25GB of RAM and shows promise for local AI setups

Running LLMs and agents in home lab setups is steadily gaining popularity due to the rising cost of AI bot subscriptions and concerns about data privacy. Unfortunately, an Nvidia NVL72 rack is ever so slightly out of the financial reach of most people, so enthusiasts have to make do with models that can run in limited amounts of memory. Italian engineer Vincenzo (aka JustVugg) seemingly wanted to have his cake and eat it, so he created ColibrÌ to run the 744-billion-parameter 1.5-TB GLM-5.2 model on a modest CPU, a mere 25 GB of RAM, and a 1 GB/s virtual NVMe drive.

Let's get the elephant out of the way: Colibrì's speed on Vincenzo's setup is only about 0.05 to 0.1 tokens per second on average, a measure that's unusable for practical conversation — imagine just one question taking hours to answer. Higher-end setups provide far better figures, but for now, they still don't meet the 20-30 tokens per second required for real-time use.

Having said that, GLM-5.2 is a Mixture-of-Experts (MoE) model with frontier-level capability, at least somewhere in viewing distance of the finest offerings from Anthropic, OpenAI, et al. This means that the quality of the answers ought to be excellent, and Vincenzo himself says his limited testing produced some impressive results. The way Colibrì works is simple enough to describe, and yet hard to do right: loading the model in slices to RAM. We're going to oversimplify for clarity's sake.

An MoE model like GLM-5.2 includes hundreds of expert sub-models to answer different topics, and these are chosen per token, not per query — meaning that when you ask a question, your words get split into tokens (chunks). For each token, the bot activates the best experts for it. The experts might always be the same for the entire question, but more often than not, a query might reel in tens of experts, possibly going into triple digits.

Whereas normally large chunks of the model, or the entire model, are loaded onto interconnected datacenter GPUs, Colibrì takes advantage of the MOE architecture and repeatedly loads/unloads the experts required per token, allowing even a cheap machine to use a large model at a steep performance penalty. For speed and simplicity's sake, Colibrì's expert-selection code is a single C file with very few dependencies. Additionally, the GLM-5.2 model is quantized down (simplified with lossy encoding) to take up less space to begin with.

If you're thinking that loading and unloading data for every piece of a question's words is going to be a hard hit on storage I/O and memory bandwidth, you're exactly on the right track. In this type of setup, NVMe storage speed is the first major bottleneck, but the proverbial funnel varies across configurations. Give it enough storage bandwidth, then you're up against RAM limitations. Fix that, then you need more CPU cores, and so on.

Colibrì is currently a proof-of-concept and doesn't yet run on GPUs, though it's worth noting that even then, shuffling data to/from the card will almost certainly be the biggest constraint. Even still, the project has barely been released, and it's already proving quite popular. Vincenzo is collecting benchmark data and running fixes as we speak, so be sure to visit the repository to contribute if you can. Maybe at some point it'll be feasible to run a really clever model on high-end consumer hardware at a decent enough clip.

  •  

Chat Control 1.0 sneaks through the EU Parliament, letting companies scan user data without warrants — legal tactic used to force a majority-required re-vote on eve of Parliament break

The Chat Control 1.0 law that enables warrantless mass scanning of digital communications has been voted against multiple times by the EU Parliament. And yet, just like a movie zombie, it keeps getting resurrected by various legal sleight-of-hand moves. Yesterday, one of those tricks worked, as Chat Control 1.0 passed (or rather, was not rejected) in a forced re-vote that required an absolute majority (50% + 1) for active refusal. This brings back the law until 2028, and sets a different stage for September's upcoming discussion on Chat Control 2.0.

After the impending publication in the EU Official Journal, online direct-communication platforms will be allowed to mass-scan their users' data without the need for a warrant, under the guise of looking for child sexual abuse material (CSAM).

The scanning is not mandatory, but big tech firms will have a legal mechanism to rifle through user data. EU firms have historically refrained from doing so, presenting privacy and data sovereignty as selling points, but the legal door is nevertheless now officially open.

The obvious platforms where monitoring can now take place will be e-mail and chat services. Immediate examples include Gmail, iCloud, Hotmail, Discord, Instagram, Slack, Teams, Snapchat, Xbox, and Google Chat.

Although the law's scope is for "interpersonal communications services," the legal mechanism might hypothetically extend to some gray areas like Google Drive, where sending someone a link to a cloud file could be within the scope of the law.

It's worth noting that "direct communication" isn't restricted to one-to-one chats, as it includes group chats; just not public or undirected communications. Additionally, EU law enforcement is still beholden to the same warrant requirement as before — Chat Control 1.0 does not grant a blank pass to authorities to mass-scan user data, or request companies to do so without a targeted warrant.

Thanks to two amendments in yesterday's vote, end-to-end-encrypted (E2EE) communications means (ex: WhatsApp) stay exempt. That means that for now, Chat Control 1.0 isn't a commandment to break encryption, something that has been regularly suggested by lawmakers around the world.

It's as good a time as any to remind people that Instagram messages are no longer E2EE as of May, and that although WhatsApp's messages are encrypted, the service leaks out every single bit of metadata about them — sender, recipient, time, size, etc. As always, Signal is recommended as a privacy-focused communications app.

This latest development in the EU parliament is eliciting widespread public outcry due to the nature of the law itself, but also due to the manner in which it happened. Critics and opponents of the rule are suggesting this move is unprecedented.

Chat Control 1.0 has already been shot down repeatedly, most recently in March. However, European Parliament President Roberta Metsola forced a second reading of the law, and invoked Rule 163's "urgent procedure" mechanism. This had many effects, including bringing up a law that was voted against for discussion yet again; turning the decision into a denial vote (vote-to-deny, not vote-to-pass); exploiting the second-reading requirement that demands an absolute majority vote (50% + 1); and letting the President herself set the schedule. Metsola scheduled the second reading to the very last day before the European Parliament summer recess.

The result was that out of 720 representatives, only 607 actually cast a vote. Of those, 315 (over half) voted against Chat Control 1.0. That figure did not meet the supermajority threshold of 361, which was calculated against a full chamber.

Opponents to Chat Control have posted resources at the Fight Chat Control website, including a breakdown of member-state and individual representative voting positions and contact information.

  •  

Minecraft shown running on Game Boy Color and Game Boy in 3D with textures — developer coaxed 3D look out of barely-there hardware

Tobias Friedly, also known as Game of Tobi on YouTube, is quite the wizard when it comes to making the seemingly impossible happen on early Nintendo hardware. His latest venture is getting a limited version of Minecraft running on the Game Boy Color (GBC)... but in 3D. As an added bonus, he even got the game working on the original Game Boy in limited fashion, thanks to both machines' interoperability.

Friedly demonstrated his work in a short YouTube video, where it's plainly visible that he managed to coax the GBC's 8x8 sprites into something visually resembling Minecraft in a three-dimensional projection. Although there are no enemies, inventory, or game logic, the feat is exceedingly impressive given that the hardware was never meant to have 3D games.

This is a significant departure from existing Minecraft demakes for old Nintendo gear, seemingly all of which tried implementing the original game's mechanics in a flat two-dimensional space. Friedly's project appears to be an offshoot of his existing Minecraft 3D for the Gameboy Advance.

Even with the limited GBC hardware, he even went as far as adding a map generator that can create flat or bumpy maps. Blocks of various types (including portals) can be placed and removed, and he even added the Nether area of the game. Game saving and loading is included, and there's an option to enable block textures. Predictably the results are a bit iffy given the limited resolution of the display and the inherent difficulty in faking texture mapping with 8x8 tiles.

Despite running at different speeds and color levels, the Game Boy Color and original Game Boy are for the most part compatible, and Friendly even showed that his clone does run on the original black-and-green Game Boy. It's hard to distinguish blocks on that screen — but it works. If you're interested in trying it out yourself, download the cart files that Friedly published.

  •  

New hack exploits AI hallucinations to trick agents into running malicious code — 'HalluSquatting' attack exploits a fundamental weakness in every available model

Ever since the advent of agentic AI, security researchers have been yelling from the top of their lungs about how it's a bad idea to grant user-level permissions to an LLM — for all purposes, a program with non-deterministic outputs and inconsistent handling of inputs. A research paper on HalluSquatting from researchers at Tel Aviv University, Technion, and Intuit, shows how easily one can fool modern AI bots and harness them into a massive army of AI agents, with the research showing that agents can hallucinate potentially malicious code repositories up to 85% of the time.

The mechanism for HalluSquatting (aka "adversarial hallucination squatting") is surprisingly simple, and takes advantage of the fact that when met with unfamiliar terms, bots will not know they're incorrect and hallucinate a "correct" answer. Adding to that, the methods the bots use to come up with said answer are predictable, for example, owner/repository or toolname/toolname GitHub URLs. This is different than just standard typo-squatting, as it exploits the hallucination mechanism itself.

An attacker first identifies an application, code repository, programming library, or bot skill that's gained popularity only in recent months or years — let's say, a new GitHub repo with the URL OriginalOwner/WindowsTelemetryOff. As the bots' training data is not recent enough to contain information about it, GitHub URLs owner/repo combinations SuperHacker/WindowsTelemetryOff , and WindowsTelemetryOff/WindowsTelemetryOff look just as peachy. Likewise, WindowsTelemetryOf and WindowTelemetryOff (note the typos) will be valid candidates.

The attacker then creates a malicious repository using those generated names. When Claude or another code agent is asked to "run the windowstelemetryoff scripts" or a similar instruction, chances are they'll hallucinate the repo name (sometimes even having run a web search), run into the malicious version that looks like the original, and happily run whatever's in there.

From that point, all bets are off now that the attacker's code is running on the user's machine. The most obvious outcome could be creating a reverse shell (the user's machine opens a command line that's controlled remotely). Now having access to the user's account, the attacker can siphon off their data and passwords, install software, run crypto miners, or harness their AI agent for further malfeasance, all with the power of entire data centers at their disposal.

And here's the kicker: just the one HalluSquatted piece of software has the potential to bait and reel in tens of thousands of bots, if not more, in a proverbial blink of an eye. A crafty attacker would be kind enough to include all the original code in their poisoned version, adding yet another layer of unawareness to the mix.

The research team found that an LLM will hallucinate the location of a recent code repository up to 85% of the time, a figure that can reach 100% for trending agentic skills. Every single model is widely affected, up to and including Anthropic's mighty Claude Opus 4.5. At the application level, the figures are better, but still pretty bad.

The scientists are working on common LLM-backed programming applications, including Cursor, Windsurf, and OpenClaw, among others. In this scenario, the bots stand a better chance given they're working with more context information, but even still, the success rates for hacking ranged from 20%-35% for Cursor, Gemini CLI, and Copilot, and increased massively to close to 80-100% on OpenClaw and its variants. The exploit mechanism doesn't even need to be crafted specifically for any bot; the researchers' results show it's universal and transferable, too.

The mean hallucination rate for names of sample GitHub repositories published in 2025 is 92.4%, while predictably, bots get the URLs wrong 0.9% for those from 2019 or earlier, though that's arguably still a concerning figure. The most effective mitigation is adjusting workflow: instructing bots to always run web searches before installing software, and providing them with additional context. Unfortunately, that's not the default way most people appear to use them.

Cybersecurity professionals have long advocated for not blindly trusting a bot's actions and severely restricting the access level granted to AI agents. And yet it's not uncommon to see bots with wide-ranging permissions over users' machines, API keys, access keys, and service accounts, to name a few — all in a bid to make it "easier" for the bot to vibe-code their pointy-haired-boss' latest brilliant idea.

  •  

Nvidia touts Vera CPU's single-threaded performance as its agentic AI advantage, reveals next-gen 'Rigel' Arm CPU cores — frames chip as a 'max single-threaded CPU at scale,' not a parallel monster

Only a little while back, Phoronix got the chance to test-drive one of Nvidia's upcoming Arm-based Vera CPUs. In certain approved workloads, the chip put up an impressive showing, nipping at the heels of its Xeon and Epyc x86 competitors. In specific single-threaded scenarios, Vera "absolutely dusted the competition" (our words). But AMD had some things to say about the Phoronix test, firing back with its own metrics of a 3.3x performance gain over Vera for the projected output of a 100 kW rack of its hardware.

And Nvidia is already thinking about this future. It revealed that its next-gen Rigel Arm v9.2 CPU core, shipping as part of its Rosa CPU, will deliver even higher per-core performance than Vera's Olympus core within the same silicon footprint via "better instruction delivery," more L2 cache, and better memory handling.

Now, Nvidia is reasserting Vera's advantage for AI work by describing it with a new product category: a "max single-threaded CPU at scale" rather than a parallel-processing beast. Instead of simply maximizing the core count per socket, Nvidia says Vera's monolithic 88-core design is meant to provide strong performance per core under load, enough memory bandwidth per core to keep active cores supplied with data, and predictable latency.

Nvidia describes AI inference workloads as being bound by single-thread speed. For example, a reasoning AI will run the model for one step, and will run the model again as many times as it takes until the answer is generated. Since each step needs the output from the previous one, no amount of parallelism will help — the speed at which one thread can run is most important. The situation is similar in agentic workloads, as agent B can't get its work started without knowing what happened with agent A.

Nvidia Vera performance profile

(Image credit: Nvidia)

Vera's design, then, appears to be one aimed at both having and eating the proverbial cake: high single-thread speed with a large number of available threads. Vera is an 88-core design with SMT support for 176 total threads. And to supply each of those cores with adequate bandwidth, Nvidia says Vera talks to LPDDR5X RAM at 1.2 TB/s, and that its monolithic compute die keeps cores well fed and avoids bottlenecks thanks to 3.4 TB/s of core-to-core bandwidth. The company says the latter figure is 3x that of "any other data center CPU."

There are many ways to measure inter-core bandwidth, so direct comparisons are tricky at best, but given the bespoke design of Vera for AI inference tasks, the claim is at least plausible.

The company's latest blog post about the new silicon reiterates this point, claiming its new silicon delivers 1.8x higher performance versus its x86 competition in "loaded CPU workloads that represent agentic execution," 1.5x higher perf in coding workflows, and 3x faster work in database analytics.

The numbers Nvidia touts purportedly come from real-world scenarios, starting with those from Perplexity, whose usage of Vera in coding agent work delivered a claimed 1.5x performance increase over x86, and a 1.9x speedup running concurrent sandboxes.

The claimed speed increases are wider still in database workloads, with Starburst (federated database firm) clocking a 3x uplift in large-scale SQL analytics, while Redpanda's real-time analytics saw a claimed 6x latency drop. According to Nvidia, all this purported performance is delivered by Vera's particular architecture, one that aims to deliver maximal single-thread performance with high thread counts.

We should note that vendor-approved benchmarks should always be taken with a bucket of salt, particularly those for hardware in a field that can shuffle trillions of dollars in a single day. The company doesn't say which precise x86 chips it tested Vera against, but it's a fair guess that they're mid- to high-end Intel Xeon and AMD Epyc models.

Nevertheless, in the blog post, Nvidia describes a conundrum that's familiar to most any server administrator: big-iron server chips can pack obscene amounts of cores, making them ideal for processing many tasks at once. However, the more cores you add, the slower they need to be to keep thermal performance and power draw in check. But that scale is an obstacle for tasks that need to be done now, parallelization be darned.

And the architectural decisions involved in using chiplets to scale to high core counts aren't free, either. Nvidia calls this "chiplet tax", and it says that scaling using chiplets creates memory access and performance inconsistencies that Vera's monolithic design is specifically meant to avoid.

We've long emphasized the importance of high single-threaded performance for a fast and responsive experience for client PCs, and it seems like AI agents are going to end up placing similar demands on hardware as they do their thing. If that's how the agentic AI future plays out, Nvidia's particular design optimizations for Vera make greater sense than prioritizing core count above all, as it might be for a general-purpose server chip meant to satisfy different economic and customer demands.

We'll have to see if Intel and AMD respond with "max single-threaded CPUs at scale" of their own.

  •  

Arrest and extradition of Scattered Spider hacker shines light on how Windows telemetry GDIDs can identify and track users — Microsoft device identifier is just one digital fingerprint in a software world rife with them

The Internet is buzzing over news that 19-year-old Estonian "hacker" Peter Stokes got nabbed by the authorities and extradited to the U.S. on digital crime charges, mostly thanks to Microsoft Windows' built-in telemetry. The FBI seemingly subpoenaed Microsoft, which coughed up telemetry logs that contained both Stokes' GDID (Global Device Identifier) and websites he visited using his main Windows machine.

The existence of GDID isn't new by itself, as Windows telemetry's data collection has been extensively analyzed and reported on. It's also been known, and publicly explained by Microsoft, that the extended telemetry modes (Full/Optional instead of Required/Basic) can upload lists of URLs analyzed by SmartScreen and Defender, together with the GDID. In fact, using the Edge browser in this setup can even send every visited URL. The court documents do not reveal which exact mechanism triggered the telemetry upload, though.

This data collection has long been the source of heated debate and general public disgust. Even though the data is genuinely useful and necessary for debugging (by Microsoft or systems administrators in enterprise environments), the fact that it comes enabled by default in Windows Home and Professional editions is questionable. The fact that those versions don't have a simple, user-facing "Off" switch to fully disable telemetry also adds insult to injury.

The Peter Stokes arrest appears to be the first public case where these Windows GDIDs were both used as a tracking identifier and contained telemetry data including some of the URLs the defendant visited. The case also prompted a renewed analysis of the GDID by a security researcher that you might want to look into. From what we can ascertain, it's likely Stokes had his Windows telemetry set to Optional/Full, as Required/Basic doesn't appear to transmit URLs by default.

Using the telemetry GDID, the FBI easily connected the dashing rogue to his ngrokaccount, because he used that tool in the same session in which he accessed his Facebook and Snapchat accounts. The agents also established a link between travel records, a New York IP address, and a rental at the Empire Hotel, likely facilitated by the photos Stokes posted of his hotel room. The criminal mastermind was equally sneaky (read: not) in his visit to Thailand.

As many hackers do, he enjoyed some time off playing an obscure game, in this case Ubisoft's Growtopia, shortly before accessing his Apple logins, as well as the aforementioned Facebook and Snapchat logins over the following weeks. Besides Microsoft, Google and Apple also collaborated on the hunting effort, with Google linking Stokes' phishing phone number to the same exact IP address and date where he created the ngrok account. Ever the stealthy craftsman, Stokes had created the ngrok account using the same GMail address connected to a second phone number where he made phishing calls from.

While it's easy and arguably quite necessary to hoist pitchforks at Microsoft for collecting detailed information about billions of computers by default, security professionals will be quick to remind users that Windows' telemetry is merely one of the many ways to track a user. Even if not by malice, a lot of software simply requires GDID-like identifiers for things like tracking usage, subscription and licensing limitations, activation requests, and hardware detection. And every company behind such software can be subpoenaed by authorities, as exemplified in Stokes' case by Microsoft, Google, Apple, ngrok, and others. Even privacy-oriented services like Proton are careful enough to describe what they can and cannot reveal to authorities under a court order.

If you're wondering the steps Stokes took to cover his tracks, though, you'd be looking at a small list. He did route his connections through a VPN hosted at servers from Tzulo along with the developer-oriented ngrok tunneling service and teleport.sh. Unfortunately, the modern digital world allows for many forms of identification, and hiding one's source IP address is merely one of them.

Using a VPN is recommended for digital anonymity, but it's merely the first of many necessary steps and can even backfire when not set up carefully. If misconfigured, a VPN may allow certain applications and operating system features to talk to the outside world using the original IP instead of the hidden one. Plus, the VPN will not stop the operating system or any application from sending out identifying information to begin with.

Perhaps more worryingly still, modern-day device and user fingerprinting is far more insidious and hard to counter. For example, plain web browsers are notorious leakers of personal information, as data-harvesting companies can weaponize features like TLS levels, HTML5 Canvas functionality, the fonts list, and even Widevine DRM in a combination that uniquely identifies a visitor. Stokes now has plenty of time to read up on the EFF's surveillance self-defense guides and get acquainted with the scripts at the Privacy Is Sexy website.

  •