ON-PREMISES DATA & AI PLATFORM

Tarantula Lakehouse

An open lakehouse combining cost-efficient object storage, distributed SQL analytics and AI-ready data management.

8B75
Home/Tarantula Lakehouse

Open analytics
for BI and AI.

Tarantula Lakehouse brings structured, semi-structured and unstructured data into an open architecture. Storage and compute scale independently, while sensitive data can remain on premises.

Tarantula Lakehouse integrates structured, semi-structured and unstructured data using open table formats. Separating storage and compute supports SQL analytics and AI/ML workloads together. Keep sensitive data in-house while scaling with data volumes and analytical demand.

OpenOpen table formats
ScaleSeparate storage and compute
BI+AIOne shared data foundation
On-premPredictable costs

Combine the strengths of a data lake and warehouse in one operating model

01

Data integration

Bring business databases, logs, files, images and external data into one storage layer to reduce departmental silos.

  • Store structured and unstructured data
02

High-performance SQL

Distributed SQL and columnar data structures support large-scale aggregations, joins and interactive analytics.

  • Trino MPP SQL Engine
03

Storage efficiency

Object storage and compression reduce long-term retention costs and extend the useful life of data assets.

  • ToS (Tarantula Object Storage)
04

Openness

Standard formats and interfaces establish a data foundation independent of a particular analytics tool or cloud.

  • Iceberg open table format
05

AI Ready

Connect BI, data science, model training and RAG pipelines to the same data.

  • 100% S3 API compatibility
06

Hybrid architecture

An on-premises foundation provides an expansion path to cloud analytics services when needed.

  • Integration experience with AWS and Snowflake

From ingestion to analytics and AI
one connected data flow

Ingest→Meta Data→Store→Object Storage→Trino→Analyzie→Usage
Data ingestion scope

Ingestion requires development tailored to each data source and is delivered with UNNET’s specialist partners.

Aggregation and join performance validated on large transaction datasets

A PoC measured aggregations and joins on structured datasets ranging from tens of millions to billions of records, demonstrating strong analytical performance against the comparison systems.

Query workload Comparison A Comparison B Tarantula Lakehouse
Daily aggregation (10 million records) 8 s 0.7 s 0.4 s
Monthly aggregation (hundreds of millions of records) 13 s 1 s 0.8 s
Annual aggregation (billions of records) 26 s 2 s 4 s
Daily join across two tables 12 s 1.2 s 1 s
Monthly join across two tables 29 s 3.3 s 2.5 s
Annual join across two tables 52 s 15 s 19 s

Turn rapidly growing transaction data into an asset

120GB+

Data generated daily

40TB+

Annual data growth

150TB

Deployed storage capacity

6~8

Production servers

Transaction data grew four to five times faster than initially projected. A PoC on hundreds of millions to billions of records validated long-term retention and analytics. After assessing performance, cost and usability, deployment began in November 2025 and the operational foundation was completed in approximately four months.

Start small and scale with your data.

Standard Edition

Minimum configuration

Designed for project research and development at hospitals, laboratories, local governments and small or medium-sized enterprises. Suitable for 500GB–10TB, it provides high-speed structured-data analytics and a repository for straightforward AI projects.

  1. Single compute configuration
  2. Community Version object storage
  3. 8×5 technical support / emergency response within one day

ENTERPRISE

Recommended configuration

An enterprise lakehouse for finance, manufacturing and retail. Suitable for 50TB and above, it provides high-speed warehouse-class structured-data analytics and a highly available platform for structured and unstructured AI workloads.

  1. Highly available compute with active-active expansion
  2. Enterprise Version object storage
  3. 24×7 technical support / emergency response within four hours
License and capacity planning

Licensing and server specifications are determined by retained data volume, daily ingestion, concurrent queries, retention period and required availability.

01

Predictable cost

Control infrastructure and processing costs with an on-premises-first model.

02

Open architecture

Use open table formats, catalog services and standard interfaces without lock-in.

03

AI ready

Serve the same governed data to BI, data science, model training and RAG.

Separate storage
from compute.

BI · AI/ML · Data AppsDistributed SQL EngineCatalog · Open Table FormatTarantula Object Storage (TOS)

Design the right data platform
for your environment.

Contact us