logo
منزل القضايا

DDN and Nebul Validate KV Cache Acceleration for NVIDIA-Based AI Factories

شهادة
الصين Beijing Qianxing Jietong Technology Co., Ltd. الشهادات
الصين Beijing Qianxing Jietong Technology Co., Ltd. الشهادات
زبون مراجعة
موظفو المبيعات في Beijing Qianxing Jietong Technology Co. ، Ltd محترفون وصبورون للغاية. يمكنهم تقديم الاقتباسات بسرعة. كما أن جودة المنتجات وتعبئتها جيدة جدًا. تعاوننا سلس للغاية.

—— 《Festfing DV LLC

عندما كنت أبحث عن وحدة المعالجة المركزية Intel CPU و Toshiba SSD بشكل عاجل ، أعطتني Sandy من Beijing Qianxing Jietong Technology Co.، Ltd الكثير من المساعدة وحصلت على المنتجات التي أحتاجها بسرعة. أنا حقا أقدرها.

—— كيتي ين

ساندي من بكين Qianxing Jietong Technology Co. ، Ltd هو بائع دقيق للغاية ، يمكنه تذكيرني بأخطاء التكوين في الوقت المناسب عندما أشتري خادمًا. المهندسون محترفون للغاية ويمكنهم إكمال عملية الاختبار بسرعة.

—— ستريلكين ميخائيل فلاديميروفيتش

نحن سعداء جدًا بتجربتنا في العمل مع شركة بكين تشيانشينغ جيتونغ. جودة المنتج ممتازة، والتسليم دائمًا في الموعد المحدد. فريق المبيعات لديهم محترف، صبور، ومفيد جدًا في الإجابة على جميع أسئلتنا. نحن نقدر حقًا دعمهم ونتطلع إلى شراكة طويلة الأمد. موصى به بشدة!

—— أحمد نافيد

الجودة: تجربة رائعة مع موردي. كانت ميكروتيك RB3011 مستخدمة بالفعل، لكنها كانت في حالة جيدة جدا وكل شيء يعمل بشكل مثالي. التواصل كان سريعا وسلاسة،وكل مخاوفي تمت معالجتها بسرعةمُزود موثوق به جداً

—— جيران كوليسيو

ابن دردش الآن

DDN and Nebul Validate KV Cache Acceleration for NVIDIA-Based AI Factories

July 17, 2026
At Paris’s RAISE Summit, DDN showcased its ongoing collaboration with Nebul, a European sovereign hybrid cloud provider, focused on boosting efficiency for large-scale AI inference deployments. Unveiled last week, the joint initiative unites Nebul’s inference platform, DDN’s Infinia data intelligence architecture, and NVIDIA accelerated computing to resolve a key production AI bottleneck: data movement costs and performance limitations during inference workloads.

أحدث حالة شركة حول DDN and Nebul Validate KV Cache Acceleration for NVIDIA-Based AI Factories  0

DDN frames the collaboration around critical production metrics: GPU utilization, token throughput, cost per token, and latency. While model training builds AI asset value, inference defines its operational and commercial returns. The rising adoption of agentic AI, retrieval-augmented generation (RAG), and high-concurrency inference means storage and data infrastructure directly impact accelerator efficiency and AI response speeds.

This active proof-of-concept project has yielded promising early results. The partners have recorded measurable improvements in time-to-first-token with KV cache enabled and completed validation for RoCE-based infrastructure. Ongoing benchmarking covers longer inference sequence lengths, unlocking further optimization potential for the Infinia platform. The collaboration also expands to joint NVIDIA efforts on benchmarking frameworks, scalability verification, and upcoming technical publications.

The integrated platform leverages distributed KV cache services, GPU-native data movement, intelligent data orchestration, and high-performance storage architecture. KV cache acceleration delivers notable inference gains by preserving and rapidly retrieving pre-computed attention states, cutting redundant calculations and eliminating data delivery delays that cause GPU idling.

أحدث حالة شركة حول DDN and Nebul Validate KV Cache Acceleration for NVIDIA-Based AI Factories  1

Leaders from DDN, Nebul, and NVIDIA highlighted a major industry shift: AI infrastructure priorities are moving from raw GPU deployment to operational efficiency, maximizing returns from existing accelerator hardware. DDN CEO Alex Bouzari and Nebul CEO Arnold Juffer noted that past focus on larger model scales has given way to optimizing inference economics to make production AI commercially viable via lower per-token costs. NVIDIA Cloud Infrastructure VP Rod Evans added that large-scale agentic workloads now measure infrastructure success by GPU utilization and latency, rather than sheer compute power.

DDN emphasizes that AI infrastructure must evolve beyond basic storage functions to actively support AI execution workflows. The firm’s infrastructure platforms currently power over one million GPUs worldwide, serving hyperscalers, cloud providers, enterprises, governments, and research institutions.

Modern AI infrastructure teams now prioritize these core production metrics:
GPU utilization: Measures effective accelerator activity during inference, maximizing value of high-end GPU hardware.
Cost per token: Links infrastructure performance directly to AI model output operational costs.
Tokens per watt: Evaluates energy efficiency of AI inference output.
Time to first token: Determines interactive AI application responsiveness and user experience.
Time to production: Quantifies operational effort to migrate AI services from testing to scalable commercial deployment.

As inference becomes the dominant AI workload, delivering cached context and enterprise data to GPUs with low, stable latency will be critical to sustaining high GPU utilization and controlling long-term operational costs.

Beijing Qianxing Jietong Technology Co., Ltd.
Sandy Yang/Global Strategy Director
WhatsApp / WeChat: +86 13426366826
Email: yangyd@qianxingdata.com
Website: www.qianxingdata.com/www.storagesserver.com
Business Focus:
ICT Product Distribution/System Integration & Services/Infrastructure Solutions
With 20+ years of IT distribution experience, we partner with leading global brands to deliver reliable products and professional services.
“Using Technology to Build an Intelligent World”Your Trusted ICT Product Service Provider!

تفاصيل الاتصال
Beijing Qianxing Jietong Technology Co., Ltd.

اتصل شخص: Ms. Sandy Yang

الهاتف :: 13426366826

إرسال استفسارك مباشرة لنا (0 / 3000)