One of the primary functions of NRENs in supporting science is to serve as the circulatory system for scientific workflows, experiments, and data analysis campaigns and this requires both the transportation and storage of vast quantities of data. Indeed, most data workflows involve storage systems at one or both ends of the workflow: the data is read from storage at the beginning, written to storage after acquisition/analysis, or both. Thus, while the network is the key component that interconnects the different pieces of the science workflow, the capabilities of the storage systems used by the workflow define the performance envelope for most scientific workflows today (even streaming workflows usually have storage at either the data source or sink).
This instalment of the Global Science Network Forum, was held in Helsinki as a side meeting to TNC26. Over 70 people participated in person, and the forum explored how user’s related storage challenges impact on NRENs from several perspectives:
- Performance: what are the capabilities of the storage systems actually deployed in scientific and research environments?
- Workload: how are storage systems actually used by science collaborations, and how does that affect traffic to or from storage systems?
- Capacity: how much storage do science collaborations have to work with, and do the cost and policy aspects of storage acquisition by science collaborations have an impact on NREN traffic?
- Design: there are many ways to design storage systems, with performance, scale, and reliability implications for the users of the system. How do the design aspects of storage systems influence the traffic?
- Cost: how does the TCO of the storage play a part in planning and design? What compromises are needed? What mitigations can be put in place?
The forum had a lineup of experts in the field and explored these different aspects, providing multiple and diverse points of view into the many aspects of storage-related issues for scientific production
- Emanuele Vitali (LUMI) – LUMI – An overview on storage systems – LUMI_storage.pdf
-
- This presentation discussed the three-tier storage architecture in the LUMI supercomputer and why storage is critical. The strengths and weaknesses of Lustre and object storage were compared, highlighting performance, scalability, reliability, and usability trade-offs. Finally,the requirements for next-generation HPC storage systems that are simpler for users and better suited for emerging AI workloads were outlined.
- Eli Dart (ESnet) – PoV: streaming vs data copy, and you’re the NREN – Streaming Data.pdf
- Modern science increasingly depends on moving data efficiently across research infrastructures, clouds, and HPC systems and while file transfer remains well understood and reliable, storage systems are becoming a bottleneck for large-scale data analysis.
This presentation provided an NREN perspective on scientific data movement, comparing traditional file-copy workflows with streaming-based approaches.
Several ESnet examples demonstrated petabyte-scale transfers, cloud-to-HPC workflows, and real-time analysis use cases. The conclusion is that future scientific workflows will increasingly rely on high-performance networks rather than on pre-positioned data copies.
- Modern science increasingly depends on moving data efficiently across research infrastructures, clouds, and HPC systems and while file transfer remains well understood and reliable, storage systems are becoming a bottleneck for large-scale data analysis.
- Ian Collier (UKRI-STFC) – Data management plans for the SKA Regional Centres – SRCNet Global and Local Network Provisioning.pdf
- This presentation provided an update on the networking requirements for the SKA Regional Centre (SRC) infrastructure. It reviewed the expected SKAO data volumes, storage growth, and associated network bandwidth requirements across the global network of SRC sites.
Updated forecasts indicate lower data rates than previously anticipated, reducing pressure on research and education networks. The presentation also reported on initial SRCNet test campaigns that validated large-file and high-file-count data transfers using tools such as Rucio and FTS. The key findings included the continued need for high-capacity connections, despite relatively low average utilisation, and also the crucial ongoing coordination between NRENs, SRC sites, and existing large-scale science networking initiatives.
- This presentation provided an update on the networking requirements for the SKA Regional Centre (SRC) infrastructure. It reviewed the expected SKAO data volumes, storage growth, and associated network bandwidth requirements across the global network of SRC sites.
- Kalle Happonen (CSC)/Bo Nygaard Bai (SIKT) – A Pan-European Object Storage infrastructure – Bo-Kalle.pdf
- Growing research data volumes, compliance requirements, user expectations, and data sovereignty are the key drivers for the proposal of a pan-European research storage infrastructure based on federated object storage.
This presentation provided a vision of a sustainable, sovereign, interoperable storage platform built from compatible provider-operated object storage services.
Design principles emphasise reuse of existing technologies, interoperability, and enabling communities to build services on top of shared infrastructure. The outcome would be a trusted, resilient European storage ecosystem supporting research collaborations at scale.
- Growing research data volumes, compliance requirements, user expectations, and data sovereignty are the key drivers for the proposal of a pan-European research storage infrastructure based on federated object storage.
- Jakob Tendel (DFC) – The data movement problem(s) – Data Movement – Problem Statement.pdf
- What are the challenges faced by the so-called “mid-tail” researcher groups when moving datasets between institutions?
Although datasets have grown from gigabytes to petabytes, many researchers still rely on manual tools such as SCP, SFTP, FTP, and rsync. The main barriers are usability, reliability, automation, interoperability, and security rather than raw network bandwidth.
The authors propose creating a simplified “data movement layer” that automates transfers, selects appropriate protocols, and guides users through compliant workflows. Objectives include making transfers in the 10 GB–10 TB range effortless and consistent across research domains. The presentation seeks community validation of both the problem statement and potential NREN-led solutions.
- What are the challenges faced by the so-called “mid-tail” researcher groups when moving datasets between institutions?
A lively Q&A followed at the end of the presentation slots, enriching the conversation well beyond the already comprehensive scope laid out by the speakers. At the end of the session, two proposals have emerged as next steps:
1 – Build upon the shared experience from the NRENs to further contribute to building a solid shared knowledge of the challenge at hand, providing user stories that document how NRENs have supported their users with solutions related to data storage/movement challenges. These case studies will help to build up the content for the next workshop that will be centred on the NRENs’ experiences.
2 – The development of a position paper to capture the challenges identified during the workshop, in particular the ones raised by Jakob about the data movement problem, and indicate possible way forward in the development of new NREN services to tackle those.
A heartfelt Thank You to all the speakers but also to all the participants, for making this workshop a big success.
If you would like to contribute to either of these proposals, please get in touch with Enzo and Eli – vincenzo.capone@geant.org, dart@es.net.
Listen on Spotify to Enzo Capone talking about the Global Science Network Forum.








