As governments and enterprises turn to sovereign AI to strengthen global competitiveness and digital independence, CIOs and their teams are accelerating investments in sovereign infrastructure to support these initiatives. To date, the EU, Japan, and other nations, in partnership with local GPU-as-a-service (GPUaaS) providers and national telecommunications companies, have committed billions to secure their digital borders. While each sovereign AI initiative now faces the challenge of converting policy ambition and capital investment into value, the broader challenge of turning AI investments into operational outcomes is shared by enterprises and neocloud providers alike.

The SK Telecom (SKT) Haein cluster is one of the first major sovereign AI deployments aligned with South Korea’s $50 billion AI initiative to build national AI infrastructure supporting 52 million people. Built in collaboration between SKT and Penguin Solutions in Gasan, Seoul, this 1,000+ GPU, 10 PB cluster went live in August 2025 and serves as a validated blueprint for rapid sovereign AI infrastructure deployment.

Read Full Deployment Story

Importantly, the cluster has completed a full year in production. This milestone provides real-world operational insight from a national-scale sovereign AI deployment, demonstrating best practices and paths to success for any organization making AI infrastructure investments.

For organizations planning similar initiatives, the lesson is clear: designing and building owned AI infrastructure is only part of the challenge; deploying it within the bounds of available energy and other critical resources and operating it within consistent control and governance is what defines long-term success.

Four Essential Insights for Sovereign Deployments, Enterprises, and Neoclouds

Deploying Haein—a 1,000+ GPU cluster—in just 60 days is a monumental achievement. By launching so rapidly, the Penguin and SKT teams proved that time-to-value (TTV) for a national-scale AI cluster can be measured in months, not years. This accelerated timeline is possible when design, deployment, and execution are meticulously planned and coordinated ahead of implementation.

The hard test began on Day 61.

Moving immediately into operation, the joint team had established performance SLAs with customers within the Republic of Korea’s Sovereign AI Foundation Model initiative to develop and tune native large language and multimodal models, support production workloads, and build user confidence in citizen-focused AI. Based on a year of sustained operations, Haein’s track record offers practical guidance for organizations planning or scaling their own AI infrastructure.

1. Success is Defined by Seamless Operations Following Deployment

As impressive as the 60-day cluster build-out was, it only marked the starting point. Organizations that treat go-live as the finish line often discover that the real challenges begin in operations—meeting customer SLAs, establishing operational procedures, monitoring and optimizing performance, training local support teams, and scaling GPU nodes and AI workloads.

Haein’s first goal after stand-up was to establish and validate a stable operational framework for this real-world AI service environment. Over the first 12 months, Haein demonstrated that sovereign AI infrastructure can sustain the strict consistency and governance disciplines that sovereign LLMs and production workloads demand.

This sustained operational record proves that AI infrastructure is a dependable production environment—often the deciding factor for organizations considering whether to build LLMs and AI workloads on a sovereign AI Factory platform. Penguin Solutions full-stack AI Factory Platform reduces deployment and operational risk for these organizations, accelerating their TTV.

The first year of Haein operations reinforced that the real measure of success is what happens after the cluster goes live: sustained node availability, consistent performance across a multi-tenant environment, rapid issue identification, and the ability to evolve the platform as workload demands change. Today, the platform successfully supports a diverse mix of government, enterprise, and research use cases across South Korea's national AI ecosystem, leveraging infrastructure from SKT, which maintains the nation's longest-standing record for customer satisfaction and operational trust.

For technology leaders, the operational phase is where AI infrastructure either generates strategic returns or erodes them. The success of AI infrastructure is determined not only by deployment, but by seamless operation.

2. The Operating Model is as Important as the Infrastructure Design

While the ability to move to operations is critical, it must deliver sustained, long-term value. Haein’s ability to enter continuous operation proved the importance of architecting an operating model with the same design rigor as the hardware procurement and deployment.

Operating model decisions—who owns what responsibilities, how incidents are triaged and escalated, how tenants are onboarded and governed, and how capacity is allocated and reported—shape day-to-day cluster performance as directly as GPU configuration or network topology. In a multi-tenant sovereign AI environment with multiple ministries, universities, and enterprises sharing a common infrastructure, operational governance is a core technical and strategic requirement.

To address this, the Penguin and SKT team established clear delineation of responsibilities before the cluster entered production:

  • Penguin Solutions Managed Services manages the underlying infrastructure, delivering 24/7 support, continuous health monitoring, rapid issue identification, automated remediation, and established standardized AI infrastructure operating procedures (SOPs) to maintain over 99% cluster availability for jobs and achieve 100% compliance with SLA commitments.
  • SKT maintains platform-level control, managing tenants, user access, and sovereign data governance through its Petasus AI Cloud and AI Cloud Manager platforms.

This deliberate design removed ambiguity under pressure, enabling both parties to move faster when requirements change.

A prime example of this model’s success was supporting SKT in obtaining South Korea’s strict Cloud Security Assurance Program (CSAP) certification for its neocloud business initiative. As the first Blackwell GPU-based AI cloud infrastructure in South Korea to achieve CSAP certification, the platform's reliability was validated under real-world conditions—serving as the active training infrastructure for the highly competitive national Sovereign AI Foundation Model Project, where SKT stands as a leading contender.

For organizations planning sovereign AI deployments, the takeaway is clear: treat operating model design as a first-class workstream and build operational readiness into your program from Day One. Retrofitting it after deployment is considerably harder.

3. Customer Requirements Evolve Rapidly, Requiring Adaptability

Haein’s first year illustrated how quickly customer requirements evolve beyond initial expectations. During the deployment phase, the primary objective was to provide a stable, high-performance AI infrastructure. Once the platform entered day-to-day operations, new requirements continually emerged, expanding beyond raw compute priorities to include complex operational processes, security controls, governance frameworks, and architectural changes. Navigating those changes required close, continuous collaboration across infrastructure, operations, governance, and support teams.

Throughout the first year, SKT continued to enhance Haein to improve its GPUaaS capabilities and support evolving customer needs. As more Korean developers, enterprises, and institutions began leveraging the platform's GPUaaS model, the scope and sophistication of their demands expanded accordingly.

This experience reinforced a lesson that many infrastructure programs discover only after launch: deployment creates capability, but continuous adaptation creates long-term value. Adaptability is especially relevant for sovereign AI infrastructure, where clusters that attract a broad range of production workloads inevitably face pressure to support larger models, longer context windows, higher concurrency, and more complex multi-tenant configurations than originally anticipated.

To prevent costly, disruptive overhauls, organizations must design for adaptability from day one. Infrastructure architectures that support non-disruptive expansion, modular pod design, and flexible software orchestration are better positioned to accommodate evolving requirements without extended downtime or full-stack replacements. Penguin Solutions' ClusterWareAI™ platform, which underpins Haein's operational management, was specifically built into the design to enable this kind of ongoing evolution—handling firmware updates, capacity expansion, and configuration changes without disrupting active workloads.

4. Operational Excellence is the Greatest Competitive Advantage

Haein’s rapid 60-day deployment was a breakthrough for a cluster of this size, but its greatest strategic value was that it allowed the team to start learning immediately and enabled an immediate transition into a highly optimized production environment. The sooner an AI cluster goes live under experienced management, the sooner an organization shifts from infrastructure setup to realizing actual compute value.

There remains a wide gap between theoretical AI cluster design and the practical realities of running large-scale environments under daily pressure. Bridging the gap requires deep institutional knowledge. Penguin Solutions’ expertise—built on accumulated troubleshooting experience, documented failure modes, and refined operational playbooks—provides this established maturity. For organizations like SKT, relying on this foundation delivers a durable competitive advantage, ensuring their platform remains resilient and expertly managed as workloads scale.

At the cluster level, these capabilities translate directly into performance. Every anomaly resolved, every silent degradation identified before it impacts a workload, and every escalation pathway refined makes the next operational challenge faster and less disruptive to address.

For organizations evaluating sovereign AI partnerships, the implication is straightforward: operational track record matters as much as deployment capability. A partner that has navigated the full lifecycle of a national-scale AI cluster brings the critical institutional knowledge required to accelerate an organization's own path to operational maturity.

This is why Penguin Solutions continues to serve as SKT’s trusted AI Factory Platform partner and is trusted by other major enterprise companies and neoclouds. Backed by field experience with nearly 100,000 GPUs deployed and managed, and more than 4 billion hours of GPU runtime, Penguin provides a foundation of accumulated operational intelligence that directly benefits new sovereign AI deployments from Day One.

Putting These Lessons to Work

Haein's relevance extends well beyond South Korea. In a global market where most sovereign AI deployments are still navigating the planning and build phases of early adoption, Haein stands as one of the few national-scale clusters with a complete year of operational history in production. That track record provides reference data that remains exceptionally rare at this scale.

Haein's multi-tenant GPUaaS model offers a validated reference architecture. It proves that organizations can deliver shared compute to multiple stakeholders with total flexibility and cost efficiency, without ever compromising the strict governance that defines sovereign infrastructure.

Ultimately, the Haein cluster demonstrates that sovereign AI at national scale can quickly be achieved—but moreover, that the organizations most likely to succeed with AI are those that plan for operational complexity with the same seriousness they bring to infrastructure design.

These insights apply broadly: to enterprises deploying private AI infrastructure, to neoclouds building GPUaaS platforms, and to nation states establishing the AI foundations for long-term economic competitiveness and digital independence.

Penguin Solutions brings the design, build, deployment, and managed services expertise to help organizations translate these insights into operational reality. To discuss how these experiences apply to your AI infrastructure program, contact our team today.

Author Image

Related Articles

Server aisle

Talk to the Experts at
Penguin Solutions

At Penguin, our team designs, builds, deploys, and manages high-performance, high-availability HPC & AI enterprise solutions, empowering customers to achieve their breakthrough innovations.

Reach out today and let's discuss your infrastructure solution project needs.

Let's Talk