Statements
%20(1).png)
Shelley Labs' submission on Singapore's draft Digital Infrastructure Bill, with technical and research support from DataFund, DINA Lab at Hongik University, and Neouly Inc.: a statutory raw/derived distinction, a class exemption for data trusts, and a mandatory tiered framework for machine unlearning.
An insurer in Singapore wants to price crop failure risk in the Mekong Delta. A regional agritech wants a yield model that works on smallholder plots rather than industrial farms. Both need the operating records of thousands of small producers across Cambodia and Vietnam.
Under the standard architecture, that means the records move — collected by a cooperative, exported to a cloud region, pooled, trained on. The contributor loses custody at the first hop. Increasingly it also means breaking the law, as data localisation requirements across the region tighten.
So the data stays where it is and the models never get built. The insurer prices on national averages. The smallholder stays uninsurable.
This is not a failure of willingness. Contributors are generally willing to have models trained on their data. What they are unwilling to do, reasonably, is surrender the data itself, permanently, to an institution in another country, in exchange for nothing they can point to.
Shelley Labs exists to separate those two things: use of data, and custody of data.
What we're building
A Singapore foundation acting as trustee of a cross-border data trust, starting with roughly 3,000 agricultural contributors in Cambodia and Vietnam.
Raw data never moves. Traversal learning — a refinement of federated learning developed with our research partners — trains across contributor-controlled nodes so that only intermediate activations and parameter updates are ever exchanged. Never records. Accuracy matches centralised training, and the method yields per-instance attribution: a quantified measure of each contributor's marginal effect on the model.
Computation is cryptographically verifiable. Processing on shared hardware runs inside hardware-attested Trusted Execution Environments executing reproducibly built, published code. The guarantee is cryptographic impossibility, not a contract clause nobody will audit.
Withdrawal reaches the model. A withdrawing contributor's node is excluded from subsequent training and, where required, the affected model is retrained without that contribution.
Singapore hosts the trustee, the orchestration software, and one small attested query gateway. That is the entire footprint. A Singapore insurtech submits a query and gets back a non-identifying result computed against data sitting in Cambodia. This is infrastructure for complying with data-sovereignty law, not for circumventing it.
On 21 July we filed a submission on the draft Digital Infrastructure Bill, with technical and research support from DataFund (Slovenia), DINA Lab at Hongik University (Korea), and Neouly Inc. (Korea). Here is what we asked for, and why.
The Bill wasn't written for architectures like this
The Bill licenses data centres at 3 MW or above and major cloud services above roughly S$100 million in Singapore revenue. For hyperscale facilities and systemic platforms, sound. The friction is in the definitions.
A cloud computing service is defined as an elastic pool of shareable computing resources, including where distributed across several locations. Read literally, that is an orchestrated network of on-premise learning nodes. Separately, the Bill doesn't say whether geographically separate nodes aggregate toward the 3 MW threshold, or whether an orchestrator setting attestation policy for equipment it neither owns nor occupies thereby "operates" a facility.
Nobody believes IMDA intends to license a cooperative learning network. But the barrier is ambiguity, not obligation. Our status today rests on inference — and inference doesn't close a financing round, satisfy a partner's counsel, or stop precautionary legal spend a foundation can't absorb.
Ask one: define derived data artefacts in statute
We asked for a statutory definition of derived data artefact — parameter, parameter update, gradient, activation, aggregate statistic, attested query result — and for distributed-learning orchestration to be excluded from the regulated category where raw data stays with its originator and only such artefacts cross the network.
Critically, the definition is keyed to risk, not type. An artefact qualifies only if it cannot reasonably be used to reconstruct the underlying data or identify an individual. The literature is unambiguous that raw gradients can permit reconstruction in some settings; a rule treating "gradient" as inherently safe would be wrong and would deserve to be ignored.
The drafting is technology-descriptive but operator-neutral. It turns on where raw data sits and what crosses the network — not on who is running it. And it would give Singapore, for the first time, a working statutory formulation of the raw/derived distinction that PDPC guidance can cross-reference without further amendment.
Ask two: make machine unlearning mandatory, and tier it
Consent withdrawal is meaningless if it stops at the database. Today the record is deleted and the model trained on it retains that influence indefinitely. The right is honoured in the storage layer and voided in the layer that makes decisions about people.
Regulators elsewhere are forcing this from both directions — the FTC has ordered model destruction repeatedly, from Cambridge Analytica through Rite Aid, and the EDPB's Opinion 28/2024 holds that trained models cannot be presumed anonymous. Singapore's PDPA has no model-level removal duty at all. That gap is a chance to define a workable standard rather than inherit an unworkable absolute one through litigation.
A naive mandate would fail, and the failure modes dictate the design. Full retraining per request is prohibitive. Efficient approximate-unlearning methods are shallow and reversible, with recovery rates above 88% reported against state-of-the-art techniques. And unlearning cannot be verified from the final model at all — identical weights can arise from runs that did and did not include a given datum.
The only auditable object is the procedure. So:
- Regulate the procedure, not the model state. A Verified Unlearning Procedure: per-source data lineage, reproducible pipeline, attested removal runs, independent process audit.
- Layer 1, immediate and unconditional. On withdrawal: cease collection, exclude from every subsequent training cycle, suppress at output within 30 days. No thresholds, no exceptions, every prescribed provider.
- Layer 2, batched weight-level removal, triggered by the earliest of a significant bloc of accumulated withdrawals, the next retraining checkpoint subject to a maximum latency, or regulator direction where data was unlawfully obtained. The bloc threshold gates only timing — never whether an individual's data keeps being used.
- Tiers. Unlearning-native architectures comply now and are rewarded with certification and procurement weighting; mid-scale providers get 36 months; frontier providers get 60.
The collective trigger has a second virtue: it gives cooperatives and data unions a right their members can exercise together, turning the data strike from an adversarial threat into an orderly statutory mechanism.
The mandate belongs in the PDPA or future AI legislation, not an infrastructure statute, and we didn't pretend otherwise. What this Bill's process can deliver now is the incentive layer and the evidence. Our pilot runs unlearning-native from day one.
The rest, briefly
Non-aggregation. Confirm distributed nodes don't sum toward 3 MW, orchestrators aren't operators, colocation tenants aren't operators.
A class exemption for qualifying data trusts, on objective criteria — fiduciary duties to contributors, raw data at origin, TEE-attested processing, under 1 MW per site, under S$10M Singapore revenue — operating by notification, not approval. Not a named carve-out; we said so.
Attestation as evidence in the codes of practice, so codes written around perimeter fences don't disadvantage architectures with stronger but differently evidenced guarantees.
Joint IMDA–PDPC guidance on cross-border flows: outbound queries engage PDPA s 26 only where they embed personal data; inbound non-identifying artefacts aren't personal data; node-to-node exchange abroad never touches Singapore.
A study on data contributions as compensable inputs. We didn't ask the Bill to decide whether a farmer whose records improve a crop-finance model holds something the law can see. We asked it to secure the evidence — audited per-contributor attribution across 3,000 contributors, a first for the region.
Why this is Singapore's win
Data trusts let Singapore host the governance layer of regional data collaboration — trustee, orchestrator, attestation verifier, compensation ledger — precisely because the raw data need not be hosted here. That layer is high-value and still unclaimed.
Every cooperative routing queries through a Singapore trustee is net new demand for Singapore-licensed infrastructure, not a substitute for it. If ambiguity prices small trusts out at formation, the trusteeship layer forms in whichever jurisdiction clarified first.
And the asks cost the licensing regime nothing. No facility above 3 MW and no service above the revenue threshold is excluded from anything.
The consultation closed on 22 July 2026. We've offered IMDA and the PDPC a technical briefing with our partners, access to the learning-engine repository and preprint, and the Cambodia–Vietnam pilot as an inaugural PET Sandbox case.
If you're building on privacy-preserving architecture in the region, or working on the same regulatory questions, get in touch.
Submitted by Liyana Mahirah, Founder, Shelley Labs Private Limited. With technical and research support from Gregor Žavcer & Tadej Fius (DataFund, Slovenia), Young Yoon, Ph.D. (DINA Lab, Hongik University, Korea), and Neouly Inc. (Korea). Contact: liyana@helloshelley.com