Safe Data Technologies: 3 Key Lessons for Data Privacy and Innovation

Lisa Chang
8 Min Read



Article on Data Privacy

Earlier this month, a new Commerce Department order shifted the conversation on how our government protects the data it collects from all of us. The directive, focused on tools like data suppression rather than newer methods like “noise infusion,” highlights a persistent and painful tension. On one side, there’s the public’s right to privacy. On the other, the critical need for researchers and policymakers to access accurate, usable data to inform everything from economic policy to public health. It’s a high-stakes balancing act, and it’s not just a problem for the Census Bureau. Every federal agency that collects statistics faces the same dilemma: how to share useful evidence without compromising individual confidentiality.

This is where the promise of Privacy-Enhancing Technologies, or PETs, comes in. These are the technical tools designed to square that circle. But moving a PET from a compelling research paper to a trusted, operational system is notoriously difficult. A collaborative effort called the Safe Data Technologies (SDT) project, involving the IRS’s Statistics of Income Division, the Urban Institute, and other partners, offers a rare and valuable case study in how to do just that. While its initial focus was on creating synthetic versions of confidential tax data, the lessons learned are a blueprint for any agency wrestling with these challenges. Having followed the evolution of data privacy tools for years, I see the SDT project not just as a technical success, but as a masterclass in the human and organizational elements required for real innovation. Here are three crucial takeaways that extend far beyond tax forms.

The first lesson is that breakthrough innovation rarely happens in a vacuum or with a single check. The SDT project’s journey from concept to reality was powered by a mosaic of partnerships, each contributing at a different stage. Early, high-risk exploration was funded by philanthropic organizations like Arnold Ventures and the Alfred P. Sloan Foundation. This kind of support is vital; it creates the runway for pure experimentation without the immediate pressure of delivering a final product. As the prototypes showed promise, research grants from places like the National Science Foundation helped refine the technology and rigorously evaluate its performance. Finally, direct investment from the IRS’s SOI division brought the operational perspective, focusing on implementation, sustainability, and integration into real agency workflows.

This layered funding model is a blueprint others should study. Philanthropy de-risks the initial leap. Academic grants deepen the science. Agency investment ensures it actually works in the real world. The key insight is that no single entity possessed all the resources or expertise needed. Success depended on the collaborative alchemy of official statisticians, academic researchers, technology builders, and patient funders all aligned around a shared, long-term vision. For any agency looking to innovate, the message is clear: cultivate an ecosystem of partners, not just a single vendor or grant.

A second, more pragmatic lesson emerged as the SDT team looked beyond their original use case. In theory, a PET that works for tax data should be adaptable for, say, labor statistics or census records. The shared challenges—protecting confidentiality, enabling secure access—are virtually identical across agencies. But in practice, the path to adoption is littered with unique local obstacles. During my conversations with technologists working in this space, a common theme arises: one agency’s solution is another agency’s problem. Differences in legacy IT infrastructure, security protocols, governance approvals, and even staff technical expertise can be monumental.

The SDT team realized that building a tool wasn’t enough; they had to build for portability. They are now actively engaging with state and local partners to map out these variations. Which components of their system can be standardized into an “off-the-shelf” package? Which parts must remain flexible to meet specific legal or operational constraints? This proactive step—engaging a diverse set of future users early in the design process—is what separates a functional prototype from a widely adoptable solution. The goal is to avoid reinventing the wheel for every new dataset while acknowledging that no two statistical agencies have the exact same IT wheelbase.

Perhaps the most profound takeaway from the SDT experience is that the hardest problems aren’t technical; they’re human. New technology introduces new questions for the people responsible for it. How do you audit a synthetic dataset for privacy risks? What level of statistical accuracy is “good enough” for a given policy decision? Who has the authority to approve its use? The SDT project invested heavily in creating the documentation, governance frameworks, and training programs needed to answer these questions. They didn’t just deliver a system; they worked alongside SOI staff to build internal competency through co-design sessions and hands-on technical assistance.

This focus on capacity building is the unsung hero of sustainable innovation. A flashy new PET is useless if the agency lacks the internal expertise to manage, evaluate, and trust it over the long term, especially after the original development team has moved on. The real impact comes from empowering agency staff with the knowledge and confidence to own the technology. As one developer involved told me, “We’re not just coding a server; we’re helping codify a new practice.”

The implications of these lessons are already spreading. Projects like the NSF’s SEDSyn-23, which is creating synthetic data for the Survey of Earned Doctorates, are applying this same multifaceted model of collaboration, practical design, and capacity investment. The Commerce Department’s recent order makes the exploration of such PETs more urgent than ever. The path forward isn’t about finding a single magic technological bullet. It’s about building the resilient partnerships, adaptable systems, and, most importantly, the human expertise within our institutions to navigate the delicate balance between privacy and progress. The future of public data depends not just on what we build, but on how wisely we prepare to use it.

  • Public’s right to privacy
  • Need for accurate, usable data
  • Sustainable innovation through capacity building
  • Engaging diverse partners early
  • Standardization vs flexibility
  • Real-world implementation and trust
Challenge Solution
Data Privacy Implement Privacy-Enhancing Technologies
Funding Initial Projects Philanthropic Organizations
Implementing Tools Collaborative Partnerships
Adopting Solutions Engaging Future Users
Building Trust Training and Documentation
Long-term Sustainability Internal Competency Development


Share This Article
Follow:
Lisa is a tech journalist based in San Francisco. A graduate of Stanford with a degree in Computer Science, Lisa began her career at a Silicon Valley startup before moving into journalism. She focuses on emerging technologies like AI, blockchain, and AR/VR, making them accessible to a broad audience.
Leave a Comment