阅读视图

发现新文章,点击刷新页面。
✇Tomshardware

Sanctioned Chinese supercomputer maker stripped of IO500 benchmark crown, Intel-powered Aurora retakes the lead — record-breaking ParaStor F9000 storage system doesn't meet reproducibility requirements

The IO500 Committee has removed storage subsystems based on Sugon's ParaStor F9000 all-flash storage systems from its Production IO500 list, as the system does not meet reproducibility requirements, which include sufficient architecture details and general availability, as noticed by Glenn K. Lockwood. The machines powered by ParaStor F9000 storage systems have been moved to the Research IO500 list and are still among the world's highest-performing storage devices; meanwhile, Intel's Aurora has retaken the top spot on the Production list.

"After further review, the Sugon ISC26 submission has been transferred from the Production List to the Research List, as it did not satisfy the criteria for the highest level of Reproducibility due to the lack of widely available architectural details and limited general availability of the file system," a statement by the IO500 Committee reads. "Accordingly, the previous #1 position on the Production and Production 10-Client lists (Argonne’s DAOS system) has been restored."

IO500 is essentially the storage counterpart to TOP500, but rather than ranking supercomputers by computational performance, it ranks HPC storage systems by their I/O performance in terms of overall bandwidth and I/O.

At ISC 2026, SCNet submitted two systems based on ParaStor storage software and ParaStor file system and F9000 all-flash storage systems. The larger SCNet AICS-A submission ran the IO500 benchmark from 500 client nodes with 64,000 client processors and achieved an IO500 score of 79,110.05, with 26,888.39 GiB/s of bandwidth and 232,754.76 kIOPS of metadata performance. The smaller AICS-B was a 10-client submission with 2,560 client processors. It scored 7,839.30, with 2,551.40 GiB/s and 24,086.69 kIOPS. Both submissions identify Sugon as the storage vendor and ParaStor as the file system.

The results substantially exceeded Argonne National Laboratory's Aurora running a custom storage subsystem featuring Intel's Optane Persistent Memory modules, SSDs, and DAOS file system. A comparable Aurora Production result scored 32,165.90, with 10,066.09 GiB/s of bandwidth and 102,785.41 kIOPS, which means AICS-A's overall score was about 2.46X higher. In the 10-client category, AICS-B's 7,839.30 was about 2.72X faster than Aurora (which scored 2,885.57). Thus, when initially accepted into the Production lists, the two SCNet submissions displaced Aurora from the top positions in both the main Production and 10-Client Production rankings.

Just like the Top 500 list, which Top 20 largely includes one-off supercomputers, the IO500 accepts completely bespoke storage subsystems based on exotic hardware and custom parallel file systems. However, the IO500 requires its Production-list submissions to meet its highest reproducibility standard, which means their architecture must be sufficiently documented and the underlying file system generally available so that the results can be independently understood and reproduced.

So, while it is hard to expect someone trying to reproduce Aurora’s 230PB storage subsystem in their garage or data center, its architecture and software can be independently examined and reproduced on a smaller scale because DAOS is open source, downloadable, extensively documented, and has publicly available architecture, implementation details, and hardware/software requirements.

By contrast, Sugon's ParaStor is proprietary and far less transparent: IO500 said the F9000 submission lacked widely available architectural details and that the file system had limited general availability, which prevents independent examination and reproduction sufficient to meet the Production list's reproducibility standard.

The same applies to Huawei's OceanFS and SuperFS architectures as well as other proprietary architectures developed in China, which the Research list includes. For example, the Research IO500 list is led by Pengcheng Laboratory's CloudBrain system with Huawei OceanStor A800 storage and the OceanFS file system, which achieved an IO500 score of 603,334.56 with 8,291.11 GiB/s of sequential throughput and 43,903,983.64 KIOPS random performance.

What is perhaps a bit odd is that while Pengcheng Laboratory's CloudBrain and CloudBrain-II submissions are clearly marked as 'proprietary' in the reproducibility column of the Research IO500 list, the SCNet-A submission carries a 'fully reproducible' badge.

✇Tomshardware

The supercomputer race no longer means what it used to, as rankings lose relevance in the AI era — as privately held compute clusters are built, running HPL becomes a distraction

China's LineShine is the world's fastest supercomputer – though having said that, it does depend on who exactly is counting. The machine, based in the National Supercomputing Center in Shenzhen, debuted at number one in the June 2026 TOP500 list published in mid-July because of its eye-popping results, with nearly 2.2 exaflops, or more than two quintillion floating-point calculations per second. The system also topped the competing HPCG ranking. But the picture was more mixed taken in the round: LineShine finished fourth on the mixed-precision HPL-MxP test and was well off the pace on the Green500 energy-efficiency table.

The gap for the supercomputer sits depending on who's ranking it, showing how what used to be a simple race for supremacy has actually become several overlapping competitions. It also highlights how "fastest" has different definitions depending on what is measured, how it is measured, and which operators publish their results.

We previously reported how LineShine uses 13.79 million cores built around China's domestic LingKun platform, 304-core LX2 processors, the proprietary LingQi interconnect, and Kylin operating system. It's also the first CPU-only machine on the TOP500 to sustain more than two exaflops, reaching around 80% of its theoretical peak while drawing 42.2 megawatts. That makes it impressive enough – but it’s also a big win on a geopolitical stage.

China rising

El Capitan

(Image credit: AMD)

LineShine is the first Chinese system to lead the TOP500 since 2017, ending a period in which China stopped submitting its most advanced machines to the public ranking. Jack Dongarra, emeritus professor of computer science at the University of Tennessee and one of the TOP500's founders, has previously said that two unreported Chinese systems may already have been faster than the then-leading US machine, Frontier.

The TOP500 has used the High Performance Linpack benchmark, or HPL, since the ranking began in 1993. In the test, computers have to solve a dense web system of linear equations using double-precision arithmetic. Like many tests, the regularity of it can be heavily tuned, producing a clean output that invites long-run comparisons across architectures and decades.

It’s tempting to look at rankings like this as the be-all and end-all, but even TOP500 says its result doesn’t reflect a machine's overall performance, because no single number can. But HPL remains valuable because it tests whether millions of processor cores and the links between them can be corralled into one tightly coordinated calculation. "It shows you can go 400 kilometres per hour with your Porsche," said Julian Kunkel, professor of high-performance computing at the University of Göttingen and deputy head of high-performance computing at GWDG, in an interview with Tom’s Hardware Premium. "This is the maximum speed. TOP500 does not tell us anything other than we have an instrument that could go that far."

That speed matters because supercomputers are as much scientific instruments as anything else, used to simulate climate systems, aircraft engines, nuclear accidents and prospective fusion reactors. Scientists want answers relatively quickly. "Imagine I do the weather prediction for tomorrow, and you need 25 hours to calculate it," says Kunkel. By the time the answer arrives, so has tomorrow, rendering it useless.

A mixed picture

The HPCG test and its sparse calculations make less efficient use of processors while putting greater strain on memory bandwidth, latency, and communication between nodes – the sorts of things that scientific applications really need. LineShine's first-place score of 22 petaflops on this measure suggests that its memory and interconnect systems are as powerful as its CPUs. "Those results describe different capabilities, not a contradiction," said Dongarra, responding to Tom’s Hardware Premium’s questions by email.

But another ranking looks at performance in a different way. The way HPL-MxP is designed favors GPUs and dedicated tensor or matrix engines. On this one, LineShine nabbed fourth place globally with 7.92 exaflops, a 3.6-fold increase over its HPL result. But the published HPL and HPL-MxP lists show the accelerator-heavy El Capitan achieved a 9.2-fold improvement, Aurora 11.5-fold and Frontier 8.4-fold.

Dongarra said the different results on different checks show a trade-off in specialization, rather than a universal weakness. LineShine's CPU-centered design has less dedicated low-precision throughput, so other machines do better on mixed-precision workloads – but he said it would be churlish to say LineShine is less capable overall. And the IO500 – which Kunkel co-founded — tests storage bandwidth and metadata performance, while Green500 divides HPL performance by power consumption.

The different metrics by which supercomputers can be measured mean that those procuring them often overlook the rankings in favor of things like time-to-solution and reliability or energy use.

What about AI?

The race for public supercomputing has been disrupted by the rise of private AI infrastructure. Although Microsoft's Eagle system appears at number seven in the latest TOP500, most frontier AI clusters publish only GPU counts, theoretical FP8 performance, or tokens-per-second claims – and some don’t really disclose much at all publicly, worried that their competitors might pick up on it.

There are some AI-based testing regimes out there, including MLPerf, which offers standardized AI training and inference tests, including new DeepSeek-V3 and GPT-OSS 20B workloads in its 2026 training suite. However, submissions to the scoring system remain voluntary, and AI labs might not want to submit them in case they score badly, torpedoing their public perception — meaning we know very little about some of the world's most consequential computing systems.

It’s also a distraction, said the experts. "There is no sense for any private AI system to run HPL," said Kunkel. Doing so would take thousands of expensive accelerators away from commercial work and require days of tuning a result their customers don’t use, and probably don’t care about.

Dongarra reckons LineShine probably isn’t best at everything, even though among those systems that are submitting public results, it showed the strongest measured double-precision and HPCG performance. That, he believes, makes it good enough to say this is a very good supercomputer. Any credible assessment, he said, "requires a portfolio of measurements" — covering time-to-solution, sustained performance under mixed workloads, reliability, ease of programming, cost and power alongside the headline benchmarks.

So while LineShine's HPL victory is significant and technically impressive, it also highlights how there is no longer one supercomputer race. Systems whose results are never disclosed may be faster for particular AI workloads, "but without comparable public evidence,” said Dongarra, “that remains a claim rather than a ranking".

✇Tomshardware

China tops the list of fastest supercomputers with a CPU-only behemoth, ending US champion El Capitan's reign — 2.198 exaflops of performance without a single GPU

China's LineShine supercomputer has taken the top spot on the 67th-edition TOP500 list, posting 2.198 exaflops on the High Performance Linpack benchmark and pushing the AMD-powered El Capitan into second place by more than 20%. The system, installed at the National Supercomputing Centre in Shenzhen (NSCS) and built by the Shenzhen Cloud Computing Center, used no GPUs or accelerators of any kind, and reached the figure with 13,789,440 cores of domestically designed silicon, the first machine on the list to clear two exaflops of double-precision performance on CPUs alone. It’s also the first China-based system to lead the TOP500 since Sunway TaihuLight in 2017.

The fact that a sanctioned country has managed to build an exascale flagship without a single Western accelerator is one thing, but what’s more telling is that China has decided to put it on the list. For years, its fastest machines have stayed off the rankings entirely, and the decision to submit a chart-topper now is a deliberate change of posture.

A domestic stack from core to OS

LineShine is built on what NSCS calls the LingKun platform. Each of its 20,480 compute nodes carries two LX2 processors, Armv9-based parts with 304 cores running at 1.55 GHz, organized as eight clusters of 38 cores. Every core includes Arm's Scalable Vector Extension and Scalable Matrix Extension units covering FP64, FP32, BF16, FP16, and INT8.

Each of those LX2s pairs 32 GB of on-package HBM rated at up to 4 TB/s with as much as 256 GB of off-package DDR5, an arrangement that’s closer to Fujitsu's A64FX in Japan's Fugaku than to a conventional server CPU. Nodes are tied together by the proprietary LingQi interconnect, and the machine runs the homegrown Kylin OS.

It’s not known who designs the LX2 — NSCS names no vendor — but Jon Peddie Research has attributed the chip to Huawei, and the project's pilot phase reportedly ran on Huawei Kunpeng servers. The fabrication node and foundry are likewise unconfirmed. SMIC's 7nm-class process is the obvious domestic candidate by elimination, given that EUV tooling and TSMC capacity are both off the table, but nobody has documented the part to date.

Not an AI crown

LineShine also took first on HPCG, the test that rewards memory- and communication-bound workloads closer to real scientific code, at 22.00 petaflops. But on HPL-MxP, the mixed-precision benchmark that approximates AI training math, it came in only fourth at 7.92 exaflops, a 3.6 times uplift over its FP64 score.

In other words, the accelerator-based machines it beat on Linpack pull far ahead the moment precision drops. Per the TOP500 announcement, El Capitan posts 16.7 exaflops on HPL-MxP, a 9.2 times jump over its standard result, with Aurora and Frontier showing similar multipliers. Reduced-precision throughput is exactly where GPUs and APUs separate from CPUs, and LineShine has nowhere to hide it.

We can see similar issues cropping up in terms of power. LineShine draws 42,220 kW and returns 52.07 gigaflops per watt on its Linpack run. That beats Intel’s Aurora comfortably but trails El Capitan's 60.94 gigaflops per watt, so LineShine produces more total FP64 output than the Livermore system while burning roughly 42% more power to do it.

It’s worth holding onto this distinction because the TOP500 ranking is decided on FP64 Linpack, the one regime where a wide, HBM-fed CPU can still go toe-to-toe with accelerators. LineShine is a genuine double-precision champion, but it’s not a world-leading AI training machine, and its fourth-place HPL-MxP result says so.

So, why did China submit it?

China stopped submitting its fastest systems to the TOP500 around 2021, after a run of entity-list additions hit Sunway's Wuxi center and Sugon. The community has long believed that the country operated exascale hardware well before this entry: the Sunway successor OceanLight and the NUDT-built Tianhe-3 both appeared via Gordon Bell Prize science papers without ever appearing on the list. TOP500 co-founder Jack Dongarra has said for years that Chinese researchers told him they weren’t permitted to submit, and that omissions were about avoiding U.S. attention rather than any lack of capability.

Last June's list, which AMD topped while Chinese HPC remained absent, was especially conspicuous, but putting LineShine forward now reverses that. It has been reported that the system was developed without public funding, which lowers the political exposure of disclosing it, and the all-domestic design means there’s no dependency on Western parts for Washington to choke off after the fact.

Addison Snell, chief executive of HPC analyst firm Intersect360 Research, told Reuters he wasn’t surprised by the performance but by the disclosure itself, noting the surprise was that China submitted the result and wanted recognition for it. Ultimately, submitting a number-one system that runs entirely on indigenous parts is a statement that the sanctions regime hasn’t closed the gap China cares about.

AMD still dominates

The top of the list might have changed hands, but the bulk of it hasn’t. The U.S. still dominates with three of the top five in El Capitan (1.809 exaflops), Frontier (1.353 exaflops), and Aurora (1.012 exaflops), and Germany's JUPITER Booster remains the first and only European exascale system at an even 1.000 exaflops.

AMD’s silicon underpins most of the accelerated field with the company, per its own blog, now powering 191 systems on the list, up 11% year over year, and 41% of this edition's new entries. It holds three top-10 slots — El Capitan, Frontier, and the newly deployed HPC7 at Italian energy firm Eni — and contributes more than 40% of combined top-10 Linpack performance. On efficiency, it powers 56% of the top 50 Green500 systems, and its first Instinct MI355X deployments, two Cambridge Zenith systems in the UK, entered at positions 67 and 68.

None of that is dented by LineShine, not least because the two aren’t competing for the same workload. AMD’s MI300A and MI355X parts are built for mixed-precision AI arithmetic, where LineShine places fourth, and the rest of the Western labs are optimizing for that, not FP64 leaderboard positions.

El Capitan, Frontier, and Aurora all post HPL-MxP scores several times their Linpack results, enabled by hardware that LineShine doesn’t have. So, while it’s true the TOP500 crown moved to Shenzen, it did so on a benchmark that Western labs are no longer chasing with their fastest machines.

✇Tomshardware

China's LineShine supercomputer dethrones US' El Capitan, secures first place in Top 500 list — first machine in the rankings to sustain more than 2 ExaFLOPS of double-precision performance using only CPUs

China's LineShine supercomputer has dethroned El Capitan as the world's number one supercomputer, going straight to the top of the charts after the National Supercomputer Center in Shenzhen (NSCS) submitted its results.

LineShine hit 2.198 FP64 ExaFLOPS in the Linpack benchmark and became the industry's first machine in the Top 500 list to sustain more than 2 ExaFLOPS of double-precision performance using only CPUs. The system is deployed at the National Supercomputing Centre in Shenzhen and was built by the Shenzhen Cloud Computing Center using semi-custom 304-core LX2 processors based on the Armv9 instruction set architecture and running at 1.55 GHz. The machine employs 13.79 million cores in total, uses proprietary LingQi interconnect, and consumes 42.2 MW of power.

From a performance-per-watt point of view, the LineShine machine delivers 52.07 GFLOPS/W, which is below El Capitan's 60.94 GFLOPS/W. However, LineShine by far outperforms Fugaku — another CPU-only supercomputer that used to be the No.1 HPC system several years ago — that can only deliver 14.78 – 16.84 GFLOPS/W depending on whether its efficiency is optimized or not.

LineShine also moved to the top of the HPCG ranking with 22.00 HPCG-PFLOPS. However, the supercomputer achieved 7.92 mixed-precision EFLOPS in HPL-MxP, which puts it behind El Capitan, Frontier, and Aurora. This limits LineShine's usability for AI training and inference, but this can be justified with its exceptional performance for traditional supercomputer tasks.

Each LX2 CPU relies on two compute chiplets and has a total of 304 CPU cores organized into eight CPU clusters containing 38 cores each. Every core includes Arm SVE (Scalable Vector Extension) and SME (Scalable Matrix Extension) units that accelerate vector and matrix operations used in AI training and scientific computing that support FP64, FP32, BF16, FP16, and INT8 data formats. The chip features a rather unusual memory architecture that pairs 32 GB of on-package HBM, offering up to 4 TB/s of bandwidth with as much as 256 GB of external DDR5 memory to maximize both bandwidth and capacity.

Despite this, the processor only gains 3.6X performance when moving from FP64 to mixed-precision data, which is lower compared to systems that integrate low-precision accelerators, such as AMD's Instinct MI300A or Intel's Ponte Vecchio. While an Armv9 CPU with SVE/SME can accelerate FP16/BF16/INT8 workloads, its mixed-precision uplift remains limited compared to systems with accelerators due to many reasons, including memory bandwidth, software maturity, and interconnect efficiency. That said, it may be too early to make final conclusions about the LX2 and its usability for mixed-precision workloads.

In any case, the very fact that a Chinese supercomputer has achieved extraordinary FP64 performance is remarkable. Furthermore, the fact that NSCS has actually submitted results to Top 500 indicates that the organization is confident that the LineShine supercomputer relies exclusively on domestic technologies and the U.S. government cannot affect the production of these technologies.

✇Tomshardware

Elon Musk restarts Dojo3 'space' supercomputer project as AI5 chip design gets in 'good shape' — will be first Tesla-built supercomputer to feature all-in-house hardware, with no help from Nvidia

Elon Musk confirms the restart of the Dojo3 supercomputer project. Success in AI5 chip design has enabled Tesla to start porting resources over to Musk's next-gen supercomputer aimed at "space AI".

✇Tomshardware

Nvidia and partners to build seven AI supercomputers for the U.S. gov't with over 100,000 Blackwell GPUs —combined performance of 2,200 ExaFLOPS of compute

Nvidia, Oracle, and the U.S. Department of Energy will build seven ExaFLOPS-class AI supercomputers for Argonne National Laboratory — including the Oracle-built Equinox and Solstice systems with over 100,000 Blackwell GPUs delivering up to 2,200 FP4 ExaFLOPS — to power next-generation AI and scientific research.

✇Tomshardware

Nvidia GPUs and Fujitsu Arm CPUs will power Japan's next $750M zetta-scale supercomputer — FugakuNEXT aims to revolutionize AI-driven science and global research

Japan is investing over $750 million in FugakuNEXT, a zetta-scale supercomputer built by RIKEN and Fujitsu. Powered by FUJITSU-MONAKA3 CPUs and advanced accelerators, the system will integrate AI into scientific research, aiming to achieve performance 1,000× beyond current models.

✇Tomshardware

AMD supercomputers take gold and silver in latest Top500 as Chinese HPC remains shrouded in secrecy

The Top500 project's 65th list of performance results reveals U.S. leadership in supercomputing. The AMD-based El Capitan, Frontier, and Intel-powered Aurora take the top three spots. AI-focused systems like Microsoft's Eagle and Germany’s GH200-powered Jupiter Booster make early appearances. China submitted no new entries for this Top500 edition.

© AMD

❌