A global streaming and media leader unifying its data science foundation
As Hulu and Disney+ content and audiences converged, the company’s Central Data & Intelligence team needed a production data foundation that could keep its subscriber data science running reliably across a complex, multi-platform environment. Dozens of models — driving subscriber lifecycle, retention, and upgrade-propensity decisions — depended on a steady supply of clean, governed features.
The challenge was twofold: keep the feature store fed as data sources shifted across platforms, and make the underlying data delivery repeatable, governed, and safe enough to run in production at the scale of a global streaming service.
Platform convergence strained a Hulu-only data foundation
As Hulu users increasingly consumed Disney content, the team’s feature store faced a structural risk: its legacy upstream sources relied strictly on Hulu-only data, while the models needed comprehensive engagement signals across both platforms. At the same time, the delivery process had its own friction:
-
The feature store risked a data deficit as audiences moved across platforms, threatening models already in production
-
Models needed features delivered into whichever environment a downstream consumer used — Snowflake, Hive, or S3 — without bespoke plumbing each time
-
Manual, UI-driven deployments created toil and inconsistency across environments
-
Production runtime and security needed to be uniform and auditable, not configured by hand per job
The team needed a feature store that was agnostic to where data lived, and a delivery platform that made every deployment governed, repeatable, and safe.
Merlin: an agnostic feature store on Databricks, with governed delivery
In partnership with Qubika, the team runs its production feature store — Merlin — and its data delivery on Databricks.
A Spark/Delta feature store powering 30+ models
-
Merlin’s data processing runs on Spark via Delta Lake in Databricks, powering 30+ Hulu data science models for subscriber lifecycle and retention.
-
It supports 100+ profile-compliant virtual features spanning both active and cancelled subscribers — including the features driving upgrade-propensity models (Bundle, Trio, Live TV).
-
The feature store is structurally agnostic: it reads from Snowflake, Hive, and S3, and lands model-serving views precisely where each downstream partner needs them. Four Gold production tables anchor the serving layer.
-
(Context: the broader Hulu-on-Disney+ data unification — “Hulk R2” — makes the combined engagement data available in Snowflake, which Merlin consumes as one of its agnostic sources.)
Governed, repeatable delivery on Databricks
-
Databricks (Unity workspaces) is the core compute and orchestration platform for the team’s ETL, notebooks, and production jobs.
-
Assets are packaged as Databricks Asset Bundles (DABs) — workflows, notebooks, cluster specs, init scripts, Unity Catalog SQL, and versioned Python packages — and deployed through centralized, Terraform-backed CI/CD, eliminating manual UI configuration.
-
Unity Catalog centrally declares data-access and warehouse paths to enforce governed access; cluster policies and instance profiles enforce uniform runtime and security; production runs rely strictly on service principals.
-
Scheduled workflows automatically clean up ephemeral PR and non-prod bundles, and local deployments are isolated with per-developer suffixes for fast, collision-free iteration.

Operational impact
A single foundation for subscriber data science
Merlin powers 30+ production data science models from one feature store, with 100+ profile-compliant features covering active and cancelled subscribers.
Resilient to platform convergence
By reading agnostically across Snowflake, Hive, and S3, the feature store absorbs shifting upstream sources without breaking the models that depend on it.
Repeatable, low-toil deployments
Standardized Asset Bundles and managed clusters replace manual UI setup, enabling rapid, repeatable deployments through shared CI/CD.
Uniform, auditable production
Unity Catalog governance, cluster policies, and service-principal-only production runs enforce consistent, secure, and auditable environments across every job.
Faster, safer developer iteration
Documented tooling and isolated local deployments let engineers validate and ship changes quickly without colliding with production-serving data.
A governed Databricks foundation for streaming-scale data science
By running Merlin on Spark and Delta Lake in Databricks and standardizing delivery on Asset Bundles, Unity Catalog, and Terraform-backed CI/CD, the team turned a Hulu-only, manually deployed data foundation into a governed, multi-platform feature store that keeps dozens of production models running through platform convergence — and gives its engineers a repeatable, auditable way to ship.



