48.9% Gone: How WEKA Uncovered Hidden S3 Spend

48.9% Gone: How WEKA Uncovered Hidden S3 Spend
Johan de Keulenaer
Partnerships & Channel Growth,
Johan de Keulenaer
Partnerships & Channel Growth,

WEKA was seeing significant Amazon S3 costs that didn’t add up. Its storage configuration looked healthy, the obvious explanations didn’t account for the spend, and the scale of the environment made a conventional full scan impractical. CloudZone was brought in to uncover the source without disrupting production.

About WEKA

WEKA is an AI-native data platform built for high-performance, data-intensive workloads. Its platform helps organizations accelerate demanding AI and data workloads across cloud and hybrid environments.

Industry: Data Infrastructure / Cloud Storage

WEKA was seeing significant Amazon S3 costs that didn’t add up. Its storage configuration looked healthy, the obvious explanations didn’t account for the spend, and the scale of the environment made a conventional full scan impractical. CloudZone was brought in to uncover the source without disrupting production.

The Challenge

WEKA was already using S3 Intelligent-Tiering, and its visible object configuration appeared healthy. Yet a significant portion of its storage spend remained unexplained.

The initial hypothesis was that small objects below the 128 KB tiering threshold were responsible. But standard object-level analysis couldn’t confirm it.

The bigger challenge was scale. A full scan of WEKA’s environment wasn’t practical, and comparable vendor workflows had failed to complete. Enabling additional analytics services would have added more time and complexity without guaranteeing an answer.

CloudZone needed a way to identify what was driving the additional spend without scanning the entire environment or putting production workloads at risk.

The Solution

Instead of attempting another full-scale scan, CloudZone designed a statistical investigation that could isolate the source of the spend using a representative sample of the environment.

Using Wiv.ai, CloudZone combined custom statistical sampling with S3 metadata analysis to test the initial hypothesis and investigate storage that wasn’t apparent through the standard object-level view.

The analysis revealed that small objects were not the main issue.

Approximately 95% of the unexplained spend was coming from incomplete multipart uploads, which had left behind billable storage fragments that continued to accumulate over time.

Once the source was identified, CloudZone developed an automated remediation process to safely remove the accumulated data without affecting WEKA’s production environment.

CloudZone also implemented an S3 lifecycle policy to automatically clear incomplete multipart uploads going forward, preventing the same waste from building up again.

The Results

The engagement significantly reduced WEKA’s Amazon S3 spend with zero impact on production.

  • 48.9% reduction in total Amazon S3 expenditure
  • ~95% of previously unexplained S3 spend identified and resolved
  • Zero impact on WEKA’s production environment
  • A preventive lifecycle policy to stop future accumulation of incomplete multipart uploads

The investigation also became the foundation for a reusable CloudZone automation approach. The same methodology was later extended to identify obsolete object versions and uncover additional S3 savings opportunities across other customer environments.


Meir Tapiro:

Cloud-Eng Team Lead at Weka.io

Maximum you, with us!

Read more case studies

FinOps
Scaling Cloud Efficiency Through FinOps Excellence
FinOps
48.9% Gone: How WEKA Uncovered Hidden S3 Spend
AI
Legal Document Automation with GenAI: Hahn Law Case Study

Let’s push your cloud to the max

Thanks for reaching out

We’ve received your request, and one of our experts will be in touch shortly.
Form submission failed!