For all the progress in model capability, some of the hardest problems in AI are starting to look surprisingly old-fashioned. There’s not enough context, records are incomplete, information gets trapped in separate systems and while some data exists, it cannot be used. That matters most in the sectors where AI is expected to make consequential decisions.
Credit models need more than repayment histories to understand people with thin financial files. Fraud systems need signals that cross banks, merchants and payment networks. Healthcare models become more useful when they can learn from information spread across hospitals, laboratories, insurers and primary-care systems. But none of those organizations sees the full picture.
A recent World Economic Forum article highlighted a South African experiment that makes the problem unusually concrete. Banks combined financial information with grocery-shopping behavior using privacy-preserving technology. Eight million people with little or no conventional credit history could be assessed; 3.2 million qualified for affordable credit. Models using grocery data showed a 41% Gini lift (World Economic Forum, 2026).
Much of AI development has focused on accumulation: more parameters, more compute, larger training corpora. Yet the next improvement in a lending model may not come from ingesting another mountain of public data. It may come from a small, highly relevant dataset sitting behind somebody else’s firewall.
Healthcare has the same problem. So does insurance. Cybersecurity too. Useful signals are often scattered across institutions precisely because modern economies are organized around separate companies, databases and legal responsibilities.
Bringing them together is not simply an engineering task. Financial records reveal debt, income and vulnerability. Medical records contain some of the most sensitive information a person generates. Location, purchasing and telecom data can reconstruct everyday behavior with uncomfortable precision. GDPR purpose limitation and data minimization rules exist precisely to constrain how personal information can be collected, combined and reused (European Commission).
AI faces a paradox: richer context can improve decisions, but acquiring that context can make the system harder to justify.
Synthetic data has become one response. It is valuable for software testing, simulations and situations where exposing real records would be unnecessary. But it cannot reliably fill every hole in the information landscape. Patterns missing from the original data remain difficult to recreate faithfully. Minority populations can remain underrepresented. New fraud behaviours do not appear because someone generated another million artificial rows. NIST has similarly warned that synthetic data can reduce accuracy for sub-populations, introduce systematic bias and propagate errors into downstream uses (NIST, 2025).
And high-stakes AI eventually has to leave the simulation anyway. Someone applies for the loan. A patient arrives at the hospital. An unusual transaction hits the network.
Privacy-enhancing technologies, or PETs, offer a different route. Rather than moving entire datasets between organizations, techniques such as secure computation, federated learning, trusted execution environments and controlled data environments can allow useful analysis while limiting exposure of the underlying records. The OECD has identified PETs including secure multi-party computation, trusted execution environments, differential privacy and homomorphic encryption as tools for using and sharing AI models while protecting sensitive information (OECD, 2025).
Their importance goes beyond privacy engineering. PETs point toward a broader change in what constitutes AI infrastructure. Competitive advantage may increasingly come from the ability to connect fragmented, high-quality information safely, not simply from owning the largest proprietary dataset.
Who has the most data matters less than who can combine the right data, at the right time, without creating a privacy problem worse than the one the AI was meant to solve. As the South African example suggests, some of the information AI needs most already exists.
What is missing is permission and an architecture capable of using it.
.png?width=1816&height=566&name=brandmark-design%20(83).png)