For an enterprise that runs on SAP, AI should run where the business meaning (semantics) lives because it understands the data better and produces faster, more accurate business insights. Replicating SAP data into a general-purpose external platform like Databricks, Snowflake, or Microsoft Fabric forces teams to rebuild relationships and logic downstream. This leads to additional cost, delay, and trust erosion.
Key takeaways
- Even as models improve, their effectiveness is still governed by the quality and readiness of underlying data context. A June 2026 study found that only 7% of enterprises say their data is AI-ready.
- Moving SAP data to an external platform strips business semantics, the relationships and logic that tell a model what an “order” or a “delivery” actually means.
- A delay-prediction build delivered by HGS kept modeling native to SAP, preserved business logic, and scored at-risk orders with an explainable model and minimal data movement.
- The rule is to preserve business context by default and move data out only when the workload needs something the native stack can’t provide.
Where should AI run for an enterprise that runs on SAP?
Close to your SAP data. That sounds obvious until you look at what several SAP enterprises do. They buy Databricks, Snowflake, or Microsoft Fabric, and SAP Business Data Cloud and SAP Datasphere are relegated to being the extraction layer that feeds them. Those platforms are good. SAP partners with all three. But treating data movement as the default is where the trouble starts.
SAP data carries business context
Here is the thing about SAP data. It is not self-explanatory. Every field value carries meaning that comes from the relationships, rules, and business logic built around it. An order is not just a number. It is tied to a customer, a delivery schedule, a set of rules about how it moves through order-to-cash, and a web of relationships that took years to configure. Within the SAP ecosystem, all of that is preserved automatically. The semantic layer in S/4HANA and SAP Datasphere maintains those relationships and business rules as a function of how the system is built. A model running on that layer understands your business the way your business actually works.
What happens when business semantics gets left behind
When you move that data to an external platform without carrying the context with it, the model at the other end gets the values but not the meaning. Someone has to rebuild the business logic downstream. That rebuild is slow, it introduces errors, and it makes the reasoning behind an AI decision hard to trace. Given that a data layer strongly influences what a model can actually deliver, it deserves as much strategic thought as the model selection itself.
Agentic AI’s credibility depends on the semantic layer
The stakes to get this architecture right are only going up. Agentic AI makes decisions with less human review at each step. When an agent acts on a prediction, you need to be able to show what that prediction was based on. A preserved semantic layer gives you that lineage by default. A rebuilt downstream copy means reconstructing it from the agent’s perspective. As AI agents take on more autonomous actions, they cannot operate reliably across disconnected platforms where business context has to be reconstructed at every boundary. They need a single governed layer where data, rules, and oversight travel together.
A delivered build: Predicting order delays without moving the data
We put this principle to the test on a real problem. The goal was to predict which open sales orders are likely to be delayed, so teams can step in before the customer feels it. The whole thing ran natively on SAP data and analytics infrastructure.
The architecture stayed close to the business context end to end. S/4HANA was the data foundation for orders, deliveries, and logistics. SAP Datasphere handled integration, modeling, and graphical views, with data products published into the governed store. The SAP Databricks machine-learning layer, embedded in BDC, was reached through zero-copy Delta Sharing rather than bulk extraction. The model learned delay patterns from historical completed orders, then scored every open order with a delay probability and a high, medium, or low risk classification. It also surfaced the key influencing factors, so the output explained itself.
The report moved from descriptive to predictive to prescriptive without ever leaving the governed environment. Teams saw which orders were at risk, why, and what to do about it.
Why the approach worked
The end outcome rested heavily on business context. Instead of extracting and re-engineering SAP data externally, we leveraged native modelling capabilities within with minimal data movement. We modeled order-to-cash quickly, preserved the business logic with no downstream redefinition, and ran feature engineering on trusted data. That’s the Realized AI pattern: applied AI against a measured outcome, with governance and explainability intact, on a 90-day proof-of-value timeline. The architecture choice made the result defensible.
When moving data still makes sense
In only those situations, when a workload genuinely needs something the native stack can’t give you, moving data out is reasonable. For heavy exploratory data science across diverse non-SAP sources, for large unstructured workloads, or when you’re standardizing analytics across many non-SAP systems, Databricks, Snowflake, or Microsoft Fabric is often the stronger platform. Barring these exceptions, forcing everything into one place would be detrimental.
The default that costs you
In most SAP environments, data moves out by default rather than by design. The vendor’s architecture assumes it. The integration was built that way. Nobody explicitly decided otherwise. What makes that habit particularly costly in an SAP context is that the infrastructure to avoid it already exists. SAP Datasphere, the semantic layer in S/4HANA, and the AI capabilities embedded in the Business Technology Platform were built precisely to run AI close to governed, context-rich data. The pass-through habit is therefore not just architecturally risky, it is also commercially irrational. Organizations are bypassing infrastructure they have already invested in to rebuild it elsewhere, at additional cost, with less context and weaker governance.
| Capability |
SAP BDC |
External platforms |
| Business context |
Native |
Requires re-modeling |
| Data movement |
Minimal |
Often replication-heavy |
| Data products |
Built-in |
Custom implementation |
| Lineage |
Business + technical |
Mostly technical |
| Time-to-value |
Faster |
Longer for SAP data |
Additionally, zero-copy approaches like Delta Sharing have shifted what is technically necessary. External platforms can now query SAP data in place for many workloads, without it leaving the governed environment. That gives organizations a real choice where previously there was not one. A partner with deep SAP ecosystem knowledge and experience would easily be able to leverage SAP’s own setup. Therefore, the question before any data movement decision should be straightforward: does this use case actually require moving data out, or can the same outcome be achieved while keeping the business context where it lives?
Frequently asked questions
What is SAP Business Data Cloud?
SAP Business Data Cloud (BDC) is a managed SaaS that unifies and governs SAP and third-party data through a business data fabric, with data products, a catalog, and an embedded machine-learning layer. SAP positions it as the trusted data foundation for analytics and AI.
For an enterprise that runs on SAP, where should AI and machine learning run?
As close to the SAP business context as possible. Keeping models on top of the SAP semantic layer preserves the relationships and logic that make outputs trustworthy, and it reaches a working result faster than rebuilding that context on an external platform.
What do you lose when you move SAP data to an external platform?
You lose business semantics. The relationships, hierarchies, and rules that define what your data means have to be rebuilt downstream, which adds cost and delay and weakens the lineage behind any AI decision.
How does zero-copy sharing let you keep AI close to SAP data?
Zero-copy Delta Sharing gives a platform like Databricks, Snowflake, or Microsoft Fabric live access to SAP data products without bulk extraction or duplicate copies. The governed, context-rich version stays the source of truth, so you can reach other tools without stripping the semantics.
When does it still make sense to move SAP data to an external platform?
When a workload needs capabilities the native stack cannot provide, such as heavy data science on diverse non-SAP data or large unstructured processing on Databricks, Snowflake, or Microsoft Fabric. Use zero-copy Delta Sharing so business context survives the move.