Data Product interoperability: Native Integration for Databricks, Google, SAP and Snowflake
AI has raised the stakes on data products. Every major platform now ships its own: SAP, Snowflake, Databricks and Google. But a data product is only as good as the context around it, and that context fragments the moment products are scattered across platforms. It's already a barrier: according to Gartner, 90% of enterprises struggled with data fragmentation impeding agentic AI deployment in 2025. As AI and agents consume more of this data, ungoverned products become a real liability. Collibra makes data products interoperable and trustworthy, wherever they're built.
What’s new
Collibra now interoperates with the data product constructs our partner platforms create. It ingests Snowflake Listings, Databricks listings, SAP datasphere marketplace data products and Google Cloud Data Products from Knowledge Catalog as first-class Collibra Data Product and Data Product Output Port assets. The point isn't to replace where a product is built; each partner platform remains its source of truth. It's to add a unifying and consistent layer of context: policies, stewardship, business terms, quality, publication and contract metadata, and cross-platform lineage that connects every product to the technical assets beneath it and to everything else in your catalog. Ingested products use the same asset types as data products built directly in Collibra, so they're governed identically the moment they land. Build a data product anywhere in your ecosystem; govern it consistently in Collibra.
How data product interoperability helps
The real challenge is that every data product sits behind its own platform boundary, subject to its own governance and visible only to the teams within. Connectors could surface the underlying tables, but everything wrapped around them, ownership, terms, refresh cadence and publication state, remained stranded on the source platform. So consumers were forced to switch between tools, stewards recreated assets manually, and AI initiatives drew on data of unverifiable provenance. Interoperability dissolves these boundaries while leaving each platform free to do what it does best.
Problems it solves:
- Cross-platform visibility: Data products from every partner platform appear alongside Collibra-native ones in a single enterprise marketplace, so stewards govern them all without recreating anything.
- Broken lineage: Each ingested product links to the technical assets that implement its ports, so business users trace from a product they consume back to its physical source, and up to its declared inputs.
- Added context, not duplicated effort: Owner, description, refresh frequency, and labels flow automatically from the partner platform, and Collibra layers on policies, glossary terms, quality, and ownership workflows the source doesn't carry. Collibra’s Data Product Lifecycle Manager then resolves any gaps or inconsistencies within Collibra, so what the platforms provide is inherited, not rebuilt, and everything ends up fully governed.
- Contract and compliance visibility: Refresh cadence, access terms, publication state, and audience surface in Collibra, so compliance teams verify commitments inside the governance workflow.
- Supported, partnership-aligned integrations: Customers who relied on manual documentation or community-built bridges move to officially supported, maintained integrations aligned with both Collibra's and our partners' roadmaps.
How the data product integrations work
The integration builds on Collibra's existing connectors rather than adding new infrastructure to manage. Where the platform previously cataloged technical assets (databases, schemas, tables, and columns), it now also reads the data product construct native to each partner platform and maps it to Collibra's standard Data Product and Data Product Output Port asset types. Those are the same out-of-the-box asset types used for data products built directly in Collibra, so an ingested partner product and a Collibra-native product sit side by side, governed identically: searchable, policy-attachable, and routable through stewardship. Once it's in Collibra, you can enrich it with context the source platform doesn't carry: business glossary links, data quality rules, classifications, and ownership workflows.
For Google Cloud, the integration reads Data Product entries and their associated output ports from the @dataplex entry group in Knowledge Catalog (a distinct entry type that wasn't previously ingested) and establishes business lineage to the BigQuery and CloudSQL assets already cataloged from the @bigquery and @cloudsql entry groups. Owner, refresh frequency, description, and GCP labels populate automatically, and declared input ports link to the upstream systems they draw from. Ingestion runs inside the existing Dataplex capability behind a configuration toggle, with no separate Edge capability or connection to deploy. Domain include/exclude mappings apply to Data Products exactly as they do to tables and schemas.
For Snowflake, the integration captures Listings, Snowflake Horizon's vehicle for packaging data with business context (title, description, usage terms, publication state, and audience). Collibra previously cataloged only the underlying technical assets via Edge; now Listings map to Data Product assets and Snowflake Shares link to Data Product Output Ports. This complements Snowflake Horizon by making its Listings discoverable and governable in Collibra alongside data products from every other platform, and replaces earlier community-built approaches with an officially supported integration.
For Databricks and SAP, the same pattern applies: each partner's data product construct is ingested and mapped to Collibra's Data Product and Data Product Output Port asset types, with lineage to the underlying physical assets and product metadata carried over from the source, then enriched with Collibra context.
Why you should be excited
- Data Analysts & AI Agents: Both reach every governed data product across the enterprise from one place—analysts through a single enterprise data marketplace, agents through an MCP server. Instead of navigating siloed platforms or ungoverned schemas, they consume curated, context-rich products with business terms, quality rules, and lineage attached—improving trust and analyst productivity while boosting agent accuracy and reducing token consumption.
- Data Product Owners: Declared input ports surface in Collibra linked to their upstream systems, so stakeholders trace every dependency back to its source without leaving the catalog.
- Data Governance Managers: Data products from every platform appear automatically after each sync, so you assign stewards, set policies, and track status across your whole ecosystem, partner-built and Collibra-built alike, without manually creating assets.
- Data Stewards: Source metadata arrives pre-populated, and you layer on the context the source doesn't carry: glossary terms, quality rules, classifications, and ownership. You enrich rather than re-enter.
Use cases
- Trustworthy discovery for AI agents: A data consumer (or an AI assistant answering on their behalf) searches Collibra for a "customer 360" product and surfaces the right one regardless of whether it was built in Snowflake, Databricks, Google, or Collibra, complete with usage terms, owner, and full lineage. The productized, enriched layer gives both humans and AI the business context that raw technical assets can't.
- AI-ready data products with governed context: Rather than pointing AI at ungoverned schemas, teams select governed, context-rich products—each landing in Collibra with its business terms, quality rules, classifications, and lineage attached. These curated data products give AI agents cleaner inputs, which improves accuracy and lowers token consumption.
Key takeaways
Data products from across your ecosystem, governed under one set of asset types, policies and workflows. Instead of a separate governance story for each cloud, or a decision about where data products are even allowed to live, stewards, analysts, and data product teams work from one consistent, automatically synchronized marketplace. Each data product is enriched with Collibra context, connected by lineage to the data beneath it, and carried through a single, end-to-end data product lifecycle that closes the gaps no individual platform covers on its own. Our partnerships with SAP, Snowflake, Databricks, and Google run deep, and this is interoperability that makes every platform in your ecosystem more valuable.
Where to learn more about the data product integrations
Sign up for the private preview to get early access to data product integrations, explore the latest capabilities firsthand, and share your feedback to help shape what comes next.
Keep up with the latest from Collibra
I would like to get updates about the latest Collibra content, events and more.
Thanks for signing up
You'll begin receiving educational materials and invitations to network with our community soon.