Urban multi-mode digital twin data asset safety management method
By employing a unified access gateway and multi-agent reinforcement learning, the dynamic control of data privacy and model accuracy in the urban digital twin framework was addressed, enabling secure management and collaborative analysis of multimodal data, and enhancing the operational resilience and public trust of smart cities.
Patent Information
- Application Number
- CN202511280546.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-11-11
Smart Images

Figure CN120934890A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of urban digital twin and data security technology, specifically to a method for secure management of urban multimodal digital twin data assets. Background Technology
[0002] As smart cities are rapidly being implemented, multimodal data sources, such as traffic cameras, vehicle-to-everything (V2X) devices, public security systems, environmental monitoring systems, medical images, and government archives, are converging on city digital twin platforms on an unprecedented scale for real-time simulation, precise law enforcement, and public service optimization. At the same time, city-level digital twin systems deployed in various regions have begun to explore technologies such as blockchain, differential privacy, and federated learning in order to break through data silos and improve credibility.
[0003] Existing practices have shown that differential privacy provides a statistical publishing channel for open data such as traffic flow, public health, and electricity load, but it has been criticized for its difficulty in balancing accuracy and privacy in the long run due to static noise budget configuration. Federated learning allows multiple departments to complete model training locally, but in practical applications, it is still limited by node gradient leakage, contribution difficulty, and asynchronous training rounds. Blockchain is used to record data access and model updates to provide audit traceability and cross-entity trust, but its throughput bottlenecks, smart contract conflicts, and cross-chain anchoring failures still lack unified solutions. At the same time, frequent privacy leakage incidents at home and abroad—such as fitness apps exposing the whereabouts of important personnel and the hijacking of urban IoT sensor networks—highlight the huge risks of multi-source, high-frequency data without dynamic security governance. Existing urban digital twin frameworks mostly rely on the idea of "fixed indicator thresholds + offline security hardening," and have not yet formed a dynamic defense link throughout the entire lifecycle, from data access, hierarchical protection, on-chain auditing, collaborative modeling to top-level strategic control. Furthermore, they lack an evolvable privacy-performance coordination mechanism to cope with continuous changes in regulations and business.
[0004] Against this backdrop, when morning rush hour coincides with extreme weather or large-scale events, the surge in risk indices forces the system to temporarily increase noise and tighten permissions, causing a sharp drop in the accuracy of the federated model. Conversely, when the risk subsides, the high noise from the static configuration remains uncorrected, leading to wasted budget and model degradation, creating a "safety-performance oscillation." Without an automated governance framework capable of understanding regulatory strength, budget margins, and business trends, and enabling cross-cycle learning, city platforms either sacrifice predictive accuracy due to overprotection, affecting the timeliness of traffic scheduling, energy allocation, and emergency decision-making, or suffer large-scale privacy breaches due to insufficient protection, resulting in damaged public trust and hefty compliance fines.
[0005] The aforementioned imbalance scenarios not only occur during peak load periods, but are also amplified when regulations are updated, links fail, or nodes go offline. Once the governance layer is unable to reallocate the budget in real time and simultaneously coordinate the strength of differential privacy, permission granularity, and federated training rounds, the data security system established in the early stages will quickly fail, causing model drift, expansion of the attack surface, and a surge in the difficulty of auditing and accountability, ultimately weakening the overall operational resilience of smart cities and public trust. Summary of the Invention
[0006] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a method for secure management of urban multimodal digital twin data assets. It involves constructing a unified access gateway to aggregate multi-departmental, multi-type data and generate sensitivity labels and privacy-level metadata. Subsequently, the data provider implements differential privacy or desensitization based on the labels and dynamically calibrates parameters. Access, authorization, and training are recorded on a permissioned consortium blockchain, and fine-grained access control is implemented. With on-chain authorization, each node only uploads perturbed and encrypted model updates, resulting in a global model after secure aggregation. The governance layer continuously monitors logs, budgets, and model metrics, triggering closed-loop policy adjustments. Furthermore, it introduces multi-agent reinforcement learning with privacy and performance agents, forming a unified multi-objective optimizer through distillation, and refreshing privacy, federation, and permission policies online to achieve collaborative analysis and compliant shared operation without data leaving the domain. This solves the technical problems described in the background section.
[0007] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: a method for secure management of urban multimodal digital twin data assets, comprising: establishing a unified access gateway to aggregate multimodal data from multiple departments, automatically classifying and grading the data, generating sensitivity labels for each piece of data and binding privacy-level metadata; Locally at the data provider, differential privacy or desensitization is implemented according to the metadata at different levels, and noise parameters are dynamically adjusted through policy mapping to balance availability and risk. A permissioned consortium blockchain is built, incorporating departmental nodes. Smart contracts are used to record access, authorization, and model training on the blockchain, and fine-grained access control and tamper-proof auditing are implemented according to the strategy. With the on-chain authorization, each node trains locally and only uploads model updates processed by differential privacy and encryption. Secure multi-party computation or homomorphic encryption is used to aggregate and generate a global model. Deploy a centralized policy governance layer to continuously collect on-chain logs, privacy budgets, and model metrics; when thresholds are reached, automatically issue updates to differential privacy, permissions, and federated training configurations to form a cross-step closed-loop control. A multi-agent reinforcement learning network with privacy agent and performance agent is deployed at the governance layer. Based on the real-time index, it self-plays and distills into a multi-objective optimizer, and refreshes the differential privacy parameters, federated training hyperparameters and permission policies online.
[0008] Furthermore, a cross-departmental protocol adaptation and mode conversion module is configured in the unified access gateway to complete the mode mapping and time alignment of structured, semi-structured and unstructured data; Generate a globally unique identifier and source fingerprint for each piece of data, and bind and version the identifier and source fingerprint together with the privacy level metadata.
[0009] Furthermore, sensitivity scores are calculated for the accessed data based on fields and context, combined with entropy and distribution difference measures and summarized by modality weights. Then, regularization updates are performed on the feature map constructed based on attribute relationships to smooth out mislabeling. Sensitivity labels and label version hashes are output and written into the privacy level metadata.
[0010] Furthermore, a differential privacy processing pipeline is established locally at the data provider. First, norm pruning is performed on the data or gradient that needs to be protected. Based on the sensitivity label, the privacy budget pointer is queried and the noise scaling parameter is mapped. The noise generator is called to generate perturbations and generate noise fingerprints. At the same time, the deduction record is recorded in the budget ledger.
[0011] Furthermore, the noise scaling parameters are dynamically fine-tuned based on real-time risk and availability indicators using a rule engine or reinforcement learning strategy. A safety barrier is formed by logarithmic domain updates and rate of change constraints. All parameter adjustments generate control transactions, which are then released in a canary manner by the governance layer and support version rollback and full-link audit traceability.
[0012] Furthermore, for each access or training operation, an original transaction is constructed locally. The transaction body includes a resource identifier, invocation intent, permission credentials, budget credentials, and data or model digest hash. After being jointly multi-signed, it is submitted to the permissioned consortium blockchain, written to the ledger through consensus, and an index mapping with the policy table is established.
[0013] Furthermore, the permission contract employs an attribute-based access control strategy, evaluating the subject attribute, purpose attribute, environment attribute, and sensitivity tag to generate an access token with expiration and scope; it supports token revocation and strategy change event callbacks, and records the changes and trigger reasons on the blockchain.
[0014] Furthermore, each node performs training locally, calculates the local gradient, and forms a protected update after norm pruning and differential privacy perturbation. The protected update is then encrypted using homomorphic encryption or an equivalent encryption scheme to generate ciphertext, which is submitted to the aggregation end along with the gradient hash, algorithm version, and budget increment.
[0015] Furthermore, the aggregation end sums the elements of the ciphertext in the ciphertext field to obtain the aggregated ciphertext. After threshold decryption or secure multi-party computation, the protected update sum is recovered. The global model update is calculated according to a predetermined weight and a model version number is generated. The model summary and aggregation process metadata are registered and sent back to each node.
[0016] Furthermore, the governance layer continuously collects on-chain logs, budget ledgers, and model evaluation results to construct a state vector with a fixed field order and performs dimension-by-dimensional standardization according to robustness centers and robustness scales to obtain a unified state; the unified state is then matched with trend features and a rule base to generate a draft strategy recommendation.
[0017] Furthermore, the governance layer compiles the proposed strategy into update instructions for differential privacy parameters, permission policies, and federated training hyperparameters. After applying rate limits and boundary constraints, these instructions form control transactions that are then sent to the corresponding modules. All update actions are associated with the version hash and canary releases are supported.
[0018] Furthermore, a policy network for privacy agents and performance agents is constructed, which uniformly receives the unified state as input. The action space covers noise scaling, permission thresholds, and training concurrency rounds. Each agent is trained using its own reward definition, and the output candidate actions are processed by action masking and projection to meet budget and compliance constraints before entering arbitration.
[0019] Furthermore, the teacher distribution is obtained by normalizing the action scores generated by the arbitration or teacher side through temperature, and the divergence of the distribution output by the distillation strategy network is calculated to form the distillation loss. The parameters are updated in the shadow domain according to the set period and canary release to the production domain. At the same time, the version information of the distillation parameters and rollback points are recorded.
[0020] (III) Beneficial Effects This invention provides a method for secure management of urban multimodal digital twin data assets, which has the following beneficial effects:
[0021] When data enters the first gateway, protocol isomorphism, semantic alignment, spatiotemporal synchronization, and sensitivity labeling are completed. Subsequently, a multi-granularity protection shell is built at the source end through budget pool subdivision and differential privacy noise scaling to ensure that massive multimodal data always participates in subsequent analysis in a "usable but invisible" form.
[0022] On-chain, identity registration contracts, three-way sparse policy trees, chained hash auditing, and zero-knowledge proof verification together constitute a distributed root of trust. All authorization, access, and model training operations are written to an immutable ledger in real time, and multi-signature and cross-chain anchoring mechanisms provide time consistency for operational behavior.
[0023] The federated learning phase integrates differential privacy noise injection, gradient pruning, homomorphic encryption aggregation, and threshold key splitting to ensure that aggregation nodes cannot see the gradients of individual nodes. At the same time, it quantifies the model contributions of each party through a contribution ledger, thus protecting data sovereignty while fully exploring the value of cross-domain collaboration.
[0024] The strategy governance layer introduces real-time monitoring, risk quantification, multi-objective optimization, and closed-loop control links. Once it detects a decline in model accuracy, excessive budget consumption, or regulatory changes, it can automatically generate and issue instructions to increase or decrease the budget, adjust noise, or modify permissions. Sliding window evaluation and shadow chain write-back ensure that governance actions effectively reduce risks without causing excessive interference to business, realizing a modern governance model of "less intervention, stronger supervision, and faster response" for managers.
[0025] The multi-agent reinforcement learning game layer allows privacy agents and performance agents to compete and cooperate in the same state space for a long time. The game results are injected into the multi-objective optimizer through temperature distillation and parameter mapping, so that the system no longer depends on static weights, but continues to evolve with business peaks, budget cuts or regulatory upgrades, maintaining a dynamic balance between privacy and model accuracy. Attached Figure Description
[0026] Figure 1 This is a schematic diagram of the process for the urban multimodal digital twin data asset security management method of the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] Please see Figure 1 This invention provides a method for secure management of urban multimodal digital twin data assets, including: In the multimodal digital twin system of a city, only by achieving refined management of the entire chain of data source-security label-privacy budget can a trusted data foundation be laid for subsequent differential privacy protection, blockchain auditing and federated collaboration.
[0029] The mission of Step One is to break down the heterogeneous data silos from transportation, public security, medical care, power and other sectors, and to seamlessly aggregate them into a unified gateway. Then, with the help of automated sensitivity labeling technology, privacy-level metadata is bound to each piece of data. This ensures that subsequent processes can make quantitative judgments on the value and risk of data, and can also accurately track the same data entity in each processing stage through uniquely mapped metadata, avoiding naming drift and policy deviation.
[0030] Step 1: For urban multimodal data assets, build a fully automated link from heterogeneous collection to sensitive measurement, so that each piece of data can obtain unique and traceable privacy-level metadata in the first second after entering the system, providing a reliable reference for all subsequent security and sharing strategies.
[0031] Step 101: Construction of Multimodal Data Access Gateway and Alignment with Global Input City business systems originate from different periods and suppliers, resulting in a wide variety of underlying protocols and data formats. Without a unified first-level gateway for orchestration, any subsequent security algorithms cannot function effectively. By using coherent processing logic to transform raw data into standard input tensors, a structured, time-aligned, and traceable input medium is provided for sensitivity labeling, ultimately ensuring accurate privacy labeling.
[0032] To completely eliminate interface fragmentation, four virtual channels are defined on the gateway side based on modal category priority: sensor stream channel, video frame channel, text log channel, and structured table channel. Each channel declares a unique set of catalog fields, and protocol isomorphism is achieved using JSON-Schema and Avro-Schema double-layer validation. Therefore, when the RTSP stream from the traffic camera and the HL7 message from the medical HIS arrive simultaneously, the scheduling engine can quickly schedule them using a modality-protocol-channel ternary index, and then uniformly convert them into an internal binary block format, achieving physical layer decoupling.
[0033] Raw data fields often exhibit representational differences, such as names / patient names and license plate numbers. To unify the semantic space, an embedding matrix is used in the gateway phase. Perform online entity mapping to obtain semantic vectors : ; In the formula: Inbound field tags; For field labels The one-hot vector representation has a dimension equal to the gateway vocabulary size; The semantic embedding matrix has the shape of vocabulary size × embedding dimension. Each row (or column, depending on the implementation convention; in this scheme, it is defined as a row vector in the gateway) stores a continuous representation of a field label in a unified semantic space. When in use, synonymous fields are identified and merged using cosine similarity, and a mapping table between the original field and the standard field is established simultaneously, so that subsequent sensitivity models can directly reference the unified field name.
[0034] Perform sliding window alignment between high-frequency sensor data and low-frequency service data. Set a unified time base for the city. For any mode Use variable window length Perform resampling: ; In the formula: Modality In the The average result of features within a time window; The original signal; the original stream is processed by a distributed NTP polarization corrector. Corrected to obtain; For discrete-time indexing; Modal time window length, value range Dynamically adjust according to the clock deviation of the data source When used, it can ensure that different modalities have comparable sampling step sizes at the same timestamp, providing temporal consistency for subsequent cross-modal risk fusion.
[0035] For subsequent blockchain traceability, each input record is appended with a source fingerprint. : ; Without exposing the real organization number and equipment serial number, one-way hashing can be used to verify the source.
[0036] Step 101 establishes a continuous chain of format isomorphism, semantic alignment, spatiotemporal synchronization, and source solidification, providing structured, traceable, and deambiguous high-quality input for sensitivity assessment. It seamlessly integrates heterogeneous data, completing semantic deambiguation, time synchronization, and reliable source traceability instantly at the gateway. This achieves a truly evolvable data access system; as the city expands, adding new data sources will not trigger system reconstruction, significantly reducing later maintenance costs and improving data quality and reliability.
[0037] Step 102: Automated classification and grading, and solidification of privacy-level metadata.
[0038] Even with a unified input, data risk still needs to be quantified for accurate protection. This is addressed by using multimodal feature extraction and adversarial risk propagation models to calculate data sensitivity scores. This is then mapped to a privacy level. Finally, the metadata is written to the on-chain registry, providing precise coordinates for dynamic noise injection in the differential privacy layer, where: Combining modal weights With domain risk factors Define sensitivity score : ; In the formula: , No. The windowed feature sequence or its statistical representation of the modality in the current time window; Modal weights, normalized to ; The total number of modalities is determined by the currently accessed multimodal channels; Domain risk factor, value It is set according to the specific circumstances; Information entropy function: measures the information content and distinguishability of the modality within the current window. It is implemented using parameterless nearest neighbor entropy or binning entropy and normalized to between zero and one in the range given by the governance layer. : Deviation from magnification factor, value This is used to amplify the "deviation impact from the relative public form" during regulatory windows or in specific business scenarios; Jensen–Shannon distance measures the structural deviation of the current distribution from the public baseline, normalized to between zero and one; The distinguishability coefficient is assigned different weights for direct identifiers and quasi-identifiers; : No. The public baseline distribution of the modality is obtained from public data or long-term urban statistics that have been processed in compliance with regulations; Based on sensitivity score The landing area will determine the privacy level of the data. Divided into three levels: ; In the formula: Risk thresholds are periodically issued by the policy governance layer; continuous scores are mapped to discrete levels to form the privacy budget pointer for the subsequent differential privacy layer. ; Using graph convolutional networks Correcting risk propagation between cross-modal entities and obtaining calibrated sensitivity. Identifying hidden associations (such as the same license plate + medical ID) increases the risk of aggregation, including: ; In the formula: Softplus activation; Scene attribute vectors centrally carry governance factors that are not directly reflected in the distribution; The calibration operator is used to assess sensitivity scores. With scene attribute vector The combination is an intermediate fraction; it can be implemented using functions with monotonic constraints (such as piecewise linear calibration, equidistant grids, and lightweight networks with monotonicity), and satisfies monotonicity constraints: Source fingerprint Privacy level Privacy Budget Guidelines Sensitivity after calibration Encapsulated as a meta data packet Synchronously write to the on-chain data dictionary contract: ; This ensures that any subsequent access request can index sensitivity and privacy budget with a single click, improving the efficiency of differential privacy layer parameter scheduling.
[0039] Through a closed loop of risk function-propagation calibration-level mapping-on-chain solidification, step 102 quantifies abstract risk into computable strategies and stores them in the form of on-chain contracts for the long term, thereby achieving a strong binding between data assets and privacy budgets.
[0040] This approach transforms dynamic and multidimensional privacy risks into machine-executable security policies. Compared to existing hierarchical methods that rely solely on static rules, this scheme utilizes graph neural networks to capture cross-modal hidden risk relationships and combines adaptive thresholds to maintain consistency between the hierarchical structure and the city's operational status, achieving a "real-time checkup" of data sensitivity. Furthermore, metadata is bound to on-chain contracts as an immutable reference standard, enabling the differential privacy layer to retrieve budgets in milliseconds. This significantly reduces policy computation latency in large-scale query scenarios and provides a stable and reliable source of privacy parameters for subsequent federated learning.
[0041] In summary, through a dual-drive approach of a multimodal unified gateway and automated risk quantification, the entire process of transforming raw heterogeneous data into secure label metadata is completed: Step 101 first performs protocol isomorphism, semantic alignment, spatiotemporal synchronization, and source fingerprint encapsulation on the input to ensure that the data has complete traceability characteristics when it enters the system; then Step 102 introduces a sensitivity function combining information entropy and Jensen-Shannon distance, superimposed with a graph convolutional network to combat the risk of hidden associations, and maps the score to a privacy level and solidifies it on the chain along with the differential privacy budget pointer.
[0042] In urban multimodal digital twin scenarios, step two plays a crucial role in transitioning from identifying risks to ensuring protection: it must implement the privacy-level labels generated in step one. and privacy budget pointers The noise is parsed into specific noise vectors and then precisely injected into local data or model updates. This not only suppresses attackers' inferences about individual pieces of information but also ensures that the data retains its analyzable value after being disturbed. Source nodes from different departments possess significant differences in computing power, bandwidth, and data structures. Without a unified and adaptive privacy control framework, the solidification of noise parameters will inevitably lead to the leakage of highly sensitive data or the excessive contamination of low-sensitivity data, thereby weakening the accuracy of subsequent blockchain audits and the convergence speed of federated models.
[0043] Step 2: For graded data assets, construct a dynamic differential privacy protection link to synchronize local noise intensity with real-time risks and model utility, ensuring that data undergoes verifiable and traceable privacy enhancement before leaving the source.
[0044] Step 201: Differential Privacy Budget Mapping and Intelligent Noise Injection Step one has already assigned a privacy level to each piece of data. With privacy budget pointers However, the risks and availability of different data units vary across different scenarios. Therefore, a continuous chain is constructed sequentially: budget pool - noise allocation - granular noise injection - hotspot masking. This ensures that noise energy is precisely targeted, avoiding information waste or leakage gaps caused by a one-size-fits-all approach.
[0045] First, a privacy budget pool mapping table is established on the node side, and the privacy budget pointers are... The budget is divided into three categories: query budget, deployment budget, and training budget, and allocated accordingly. This ensures that different use cases have their own independent budgets, preventing high-frequency queries from exhausting the global budget and leaving no budget available during the training phase. ; in: : Check budget; : Release the budget; Training budget; : Allocation coefficient, satisfying It is issued by the strategy governance layer; By segmenting the budget pool, each business scenario can independently calculate privacy costs, thereby enabling more precise control over the total amount of noise injection.
[0046] To improve resistance to combined attacks, the following measures were introduced: Differential privacy perspective, constructing noise scaling factor and Step Divergence correlation: ; in, For sensitivity, , : Privacy strength; Noise standard deviation; : Order, specified in the strategy layer ; Current remaining budget; use Privacy can provide tighter budget accumulation bounds during multi-round synthesis, making total budget consumption more controllable. When in use, the noise generator works in conjunction with the budget pool to scale the noise scaling factor in real time for the remaining budget, extending the data lifecycle.
[0047] The node divides the raw data into three granularities: record-level, field-level, and fragment-level, using different noise sampling strategies for each: record-level uses Laplace perturbation to maintain global distribution, field-level uses Gaussian perturbation to protect sensitive fields, and fragment-level uses random permutation to hide the temporal sequence. The process executes pipelined noise injection in the local buffer, following the order of field-first, record completion, and fragment rearrangement to ensure that the data meets the required privacy level before leaving the node, and records the specific noise type and parameters in the metadata. Multi-granularity noise injection directs noise towards the parts that truly need protection, maximizing data utility.
[0048] For highly sensitive fields, a dynamic mask is generated based on the real-time risk heatmap. When the risk increases sharply, the mask immediately triggers a surge coverage mode, replacing the corresponding field with a placeholder and marking it as retrievable. Once the risk subsides, the data is retrained using incremental data to restore it. The dynamic mask provides a more flexible protection method than strong noise, preventing the entire data from being wiped out due to extreme risks.
[0049] The budget pool makes the budget a thread-level access token, preventing budget overdraft from the root of the process; noise double buffering and traceable marking provide a complete "traceable, reproducible, and backable" path for noise injection, facilitating subsequent auditing and fault backtracking; the hotspot mask fusion risk predictor enables proactive protection against privacy spikes, rather than reactive remediation, significantly improving the real-time performance and resilience of local protection.
[0050] Step 202: Adaptive Noise Parameter Control and Data Output Consistency Verification After the initial noise injection, it is still necessary to continuously monitor data utility and signs of privacy leakage, and dynamically adjust noise parameters. By forming a closed loop of monitoring-decision-execution-verification, the noise intensity is kept in sync with environmental changes, ensuring that the privacy budget is not wasted and the model training does not degrade.
[0051] Construct state variables Among them, the risk of leakage For real-time leakage risk estimates, accuracy loss For model accuracy loss, the remaining budget ratio The remaining budget ratio. The agent, based on... The function selects actions to add noise, reduce noise, or maintain the action, and uses rewards as the basis for these actions. The assessment, in which, For the weighting factor; This approach prioritizes protection when risks increase and appropriately relaxes restrictions when accuracy decreases. The reinforcement learning agent enables the parameter tuning process to evolve autonomously, reducing the cost of manual intervention. When the agent decides to increase or decrease noise, it calls the recalibration function to recalculate the noise scaling factor and obtain a new noise scaling factor. : ; In the formula: Step size coefficient, with a positive value, set according to scenario level; The action value represents the tendency to increase or decrease noise intensity under the current risk-utility-budget trade-off: increase noise = +1, decrease noise = -1. Among them, exponential recalibration avoids data distribution instability caused by large jumps, and smooth scale updates provide stable input for model training and avoid drastic fluctuations.
[0052] Iterate continuously until the dual objective thresholds are met: and If the objective is not achieved before the budget is exhausted, an emergency strategy is triggered to publish a minimum available subset of the data, outputting only downsampled summary data. The dual-objective convergence ensures that the balance between privacy and utility is clearly prioritized.
[0053] Before the final output data, the node performs a distribution consistency check and sensitivity level backtracking on the noise-injected data. If the check passes, the final noise scaling factor, proxy action trajectory, and budget consumption are recorded in the metadata and written back to the blockchain smart contract; otherwise, the data enters the rollback queue for parameter retuning. This consistency verification closed loop ensures that any data leaving the node has a traceable protection history, providing a complete chain for the permission audit in step three.
[0054] The dual-network structure and protective lock mechanism solve the problems of concurrency conflicts and oscillations; the compressed sensing-assisted consistency check breaks through the shallow verification of traditional methods that only focus on mean and variance, providing a deeper level of protection for the integrity of the data structure. The design of putting action trajectories and budget consumption on the blockchain transforms the entire parameter tuning process into an auditable event stream, pioneering the realization of "decision transparency" and "compliance as a service" in differential privacy dynamic parameter tuning scenarios.
[0055] Step two involves segmenting the budget pool. Noise generation, multi-granularity noise injection, and dynamic masking construct a localized differential privacy protection shell. Subsequently, a reinforcement learning agent is used to adaptively adjust the noise parameters, and the protection history is solidified through output consistency verification in a closed loop. This process elevates the privacy level of step one. With privacy budget pointers This translates into real-time noise injection decisions, while simultaneously writing the remaining budget and agent trajector traces back to the blockchain. This provides a detailed log for the access audit in step three and a stable, usable data stream for the federated training in step four.
[0056] If a city's multimodal digital twin is to truly break free from the constraints of untrustworthy data and difficulties in implementing collaboration, it must implant a distributed, tamper-proof, and traceable pillar of trust into the cross-departmental call chain.
[0057] Step 3: Map node identity, data permissions, and operational behaviors to on-chain smart contract events, and form a verifiable and traceable continuous audit trail to ensure compliance throughout the cross-departmental data call and model training process.
[0058] Step 301: On-chain identity authentication and permission policy mapping In a consortium blockchain scenario, without a unified identity and permission system, each node is prone to becoming an isolated island, unable to definitively verify the legitimacy of access behavior.
[0059] When the blockchain is launched, the policy governance layer uses a two-layer certificate system to issue identities to participating departments, first issuing a level-one certificate with the city root private key. Then, a secondary certificate is issued using the department-level private key. Identity hash fingerprint calculation formula: ; In the formula: On-chain node identity; These are Level 1 and Level 2 certificates, respectively. A hash function that outputs a cryptographic hash of fixed length; When in use, it ensures that node identities are publicly verifiable and hierarchically managed. The identity is written into the identity registration contract, and any operation or transaction must reference the on-chain node identity. Furthermore, the tiered certificate system solves the problem of difficulty in refining the granularity of permissions when multiple institutions participate at the city level.
[0060] The data fingerprint-privacy level-privacy budget pointer triple generated in step one is written into the data directory contract, and a policy tree is constructed based on a three-dimensional index of data type, department level, and risk level. Policy leaf nodes record the set of access actions and contract functions. By using depth-first traversal to verify whether the caller's identity matches the policy leaf node attributes, the policy tree evolves the permission rules from a static list into a programmable structure, supporting subsequent dynamic adjustments.
[0061] When a node initiates a data read or model training request, it first constructs the original transaction locally. Each department node encapsulates an access / training / authorization action into a standardized message that is not on-chain, unsigned, or pending signature; Then, the department-level private key is used for signing, and together with the co-signature of the superior regulatory node, a multi-signature transaction is formed. Smart contracts Once all signatures are verified, the transaction status is set to authorized. Multi-signature verification prevents any single node from exceeding its authority, constituting a chain-level two-factor authentication mechanism.
[0062] To ensure consistency between the time series of the consortium blockchain and external government blockchains, nodes call a verifiable random function to generate timestamps when writing blocks. and anchored contracts using cross-chain Submitting the latest block hash to an external chain enables bidirectional referencing between the main chain and side chains. This cross-chain anchoring avoids the risk of orphan chains and provides a higher level of credible temporal evidence for subsequent regulatory audits.
[0063] Tiered certificates integrate city-level policy versions with node identities, achieving a closed loop of policy iteration and identity re-issuance; a three-way sparse policy tree and lazy pruning mechanism enable permission updates and retrospection to be achieved simultaneously, eliminating the problem of "historical gaps"; multi-signature combined with budget prediction avoids authorization without budget and protects privacy budgets from being wasted; cross-chain anchoring provides main chain-level time notarization through two-way verification, enhancing the legal compliance of the consortium ledger.
[0064] Step 302: Full-process access audit and violation tracing After solidifying identity and permissions, it is still necessary to maintain immutable records for every actual operation and provide efficient means of violation detection and accountability. Data reading, model updates, budget consumption, and noise scaling factor changes are uniformly abstracted into audit action logs. : ; Smart Contracts Record audit actions Write blocks and perform chained hashing on adjacent transactions to form an immutable log. Any tampering will destroy the hash chain, where: ; Where: Audit Summary : No. Stroke operation hash, the first A rolling summary up to the last audit action; For the first The summary obtained after each action is processed; Fixed-length cryptographic hash function When used, chained hashing allows auditors to detect changes without traversing the entire chain.
[0065] To avoid revealing sensitive fields during audits, design Proof: The node submits a proof that does not contain plaintext parameters. Validate the contract This confirms that the noise scaling factor and budget consumption meet the differential privacy threshold without exposing the internal noise sequence. Zero-knowledge proofs achieve a balance between privacy remaining on the blockchain and compliance verifiability.
[0066] To mitigate block size expansion, a hierarchical hashing approach is introduced: transactions within the same time window are first generated locally on the node. Leaf hash, upload root hash On-chain. Subsequent auditors can recalculate the root hash by extracting a subset of leaf nodes using a verifiable random function, achieving lightweight auditing. This approach combines hierarchical hashing with random sampling to compress storage requirements while maintaining audit feasibility.
[0067] On-chain event streams enter the sidechain analytics engine and utilize streaming rules. Matching abnormal patterns, such as unusually rapid budget consumption or missing multi-signature entries. Once a rule is triggered, the contract automatically records the violating transaction. Write the information into the responsibility locking table and freeze the permissions of the relevant nodes; administrative regulatory departments can use the locking table to trace the identity of nodes on the chain. fingerprints from the source Determining responsibility is crucial. Real-time violation detection allows governance to mitigate risks during the incident, rather than resorting to post-incident remediation.
[0068] Among them, streaming rules This refers to an independently enabled, versionable real-time anomaly matching rule in the event stream processing layer. It is used to identify suspicious behaviors or policy violations online on continuously arriving audit events and business events (such as AccessEvent, ModelLedger, BudgetPool, ControlTx, VerifyMultiSig result streams), and hand over the identification results to the governance and audit module to perform subsequent actions (alarms, freezes, rollbacks, forced threshold escalation, etc.). Step three, within the consortium blockchain framework, solidifies identity, permissions, and time into verifiable anchors through four stages: hierarchical certificates, policy trees, multi-signatures, and cross-chain anchoring. Subsequently, a closed loop of operation-verification-accountability is constructed using a transaction model, chained hashing, zero-knowledge proofs, layered hashing, and real-time violation detection. Metadata, noise history, and budget consumption from steps one and two are all embedded into an immutable log. Thus, in the subsequent federated learning step four, data can be accessed simply by referencing authorized transaction hashes. Auditors only need to verify zk-SNARK proofs without decrypting noise to confirm compliance with differential privacy thresholds. The policy governance layer can also monitor on-chain events related to budget consumption in real time, dynamically adjusting permissions and synchronously writing back the policy tree if anomalies are detected.
[0069] Transaction semantic normalization and ring signature marking ensure the semantic continuity of the same business process across multiple on-chain operations, helping auditors quickly locate abnormal links; the budget window accumulation gate incorporates "total compliance" into a single verification through zero-knowledge proofs, achieving efficient blocking of "ant-moving" type violations; layered hash prefix cohesion combined with random sampling allows on-chain logs to achieve a new balance between compressed storage and verifiability; executable audit scripts extend the automation of accountability to the sandbox reenactment level, providing a highly efficient and low-interference means of obtaining evidence of violations for large-scale nodes at the city level.
[0070] After completing the data source labeling, local differential privacy noise injection, and on-chain access auditing loop in the aforementioned three steps, the city's multimodal digital twin still requires cross-departmental model collaboration to unlock global insights. However, model collaboration is subject to two rigid constraints: first, no node may disclose raw data or noise history; second, the global model must maintain verifiable consistency between privacy budget consumption and the authorization chain. Step four bridges this dual constraint by allowing the algorithm to move while the data remains static: on the one hand, it utilizes a federated learning framework to break down the training process into nodes, using privacy budget indicators... Random perturbations are injected to the limit of the budget pool balance; on the other hand, secure multi-party computation and homomorphic encryption are introduced to keep the model parameter aggregation process in ciphertext state for the aggregator, and cooperate with on-chain smart contracts to complete the aggregation digest notarization and legality proof.
[0071] Step 4: Within the consortium blockchain authorization framework, the local gradient with differential privacy noise is encrypted and aggregated to generate a global model, and the model contribution and budget consumption are locked by on-chain notarization, so as to realize data not leaving the domain and reliable model evolution.
[0072] Step 401 Federal Training Task Scheduling and Node Admission Federated learning is only feasible when the four dimensions of participating nodes, data batches, privacy budget, and authorization hash are synchronized and consistent. By serializing task release, budget alignment, gradient normalization, and on-chain admission, it ensures that the input vectors in the subsequent cryptographic aggregation stage are completely consistent within the semantic and security boundaries.
[0073] The governance layer generates training task descriptors. The authorization hash is the multi-signature transaction that has been completed in step three. The block hash is used to bind the permission chain.
[0074] The descriptor is written to the on-chain TaskRegistry contract, triggering an event broadcast. It's important to note that the descriptor binds training metadata to the authorization hash, preventing unauthorized nodes from forging tasks.
[0075] After receiving the task, the node first calls the budget mapper to read the remaining training budget in the local budget pool. If the budget is insufficient to support the set number of rounds, the node refuses to join and writes a DeclineEvent on the chain, which the governance layer uses to adjust the budget or reduce the number of rounds. The budget mapper strongly binds training resource consumption to the budget pool, preventing nodes with empty budgets from joining.
[0076] After each round of local training iterations, the node first calculates the local gradient. implement Clipping, obtaining the clipped gradient To prevent extreme gradients from amplifying aggregation errors and the risk of privacy leaks, including: ; In the formula: The pruning threshold is given by the task descriptor; Local gradient It is at the department node The parameter gradient vector calculated when the current global model is trained using local data for one or more iterations after authorization on the chain is approved. Node calculation after pruning gradient hash The JoinRequest transaction is submitted along with the budget serial number, and the admission contract verifies the node's identity. Budget serial number and post-pruning gradient hash After signing, the node is added to the aggregation whitelist for this round. The admission contract implements a chain-level whitelist, prohibiting any unpruned or over-budget gradients from entering the aggregation pool.
[0077] By employing a combined mechanism of "descriptor version locking—budget accumulator—layered gradient thresholding—random sampling admission," the traditional "handshake" process in the early stages of federated training is thoroughly secured and refined. Descriptors explicitly label model sensitivities to prevent unauthorized sensitive inference; the budget accumulator and shortage alerts provide proactive budget scheduling during training; the multi-layered threshold design of the gradient pruner reduces the dual risks of vanishing and exploding gradients in deep networks; and random sampling admission breaks the common blind trust model of "upload equals aggregation," enabling source-end authenticity verification.
[0078] Step 402 Encrypted Aggregation and Model Backhaul Verification After the nodes have collected the safety gradients, it is necessary to ensure that the aggregator cannot solve the gradient of a single node, and at the same time, ensure that the aggregation result matches the budget consumption and authorized links.
[0079] The node uses the remaining training budget from step two. Call the noise generator to produce a Laplacian noise vector And add noise to the clipping gradient : ; Where: Noise injection gradient At the node After cropping the gradient, noise is added to obtain the protected gradient for uploading / aggregation. Noise vector, at the department node The above represents the local differential privacy noise injected into the protected object; Based on privacy budget pointers noise scaling factor obtained from mapping Call the noise generator to produce a Laplacian noise vector. And add it to the cropped object to form a protected result, which is then audited along with the noise fingerprint and budget flow chain; When in use, the information of a single node can be further obscured, adding a privacy layer to the dense aggregation.
[0080] All nodes use homomorphic encryption public keys. Encrypt and inject noise gradients and upload them: ; In the formula: : Ciphertext vector, node Encrypted upload payload; Encryption operators An operator that maps plaintext vectors to ciphertext vectors, satisfying additive homomorphism (adding plaintext corresponds to adding ciphertext). Implementations can be one of Paillier, BFV, CKKS, etc. The aggregation node performs addition on the ciphertext field to obtain the aggregated ciphertext. And the global gradient is obtained after decryption. ,in: ; in, The number of participating nodes, the number of valid nodes participating in the upload in this round, is derived from the task descriptor and the whitelist; Converged ciphertext Without decryption, the ciphertext is obtained by adding the ciphertext of each node element by element; due to the homomorphic property of addition, subsequent threshold decryption can directly obtain the sum of the plaintext gradients of each node; When in use, encrypted aggregation prevents the aggregator from viewing the gradient of a single node, ensuring that data remains within its domain. The aggregation node calculates the contribution weight of each node: ; In the formula: All are clipping gradients; For node indexing, For summation index; The node's identity, weight, and budget consumption triple are then written into the ModelLedger contract to form a contribution ledger. This contribution ledger provides a quantitative basis for subsequent incentives or budget reallocation.
[0081] After the global model parameters are updated, the governance layer performs zero-knowledge proofs on the model to confirm that the parameter changes are consistent with the budget history, and then writes the model hash and version number into the ModelRegistry contract. When a node pulls a new model, it first verifies that the model hash is consistent with the authorization hash, and then replaces the local model to enter the next round. It should be noted that the model is verified before use to prevent malicious models or version drift.
[0082] This solution comprehensively addresses the pain points of traditional federated learning aggregation processes, such as "single-point key issues," "unremembered contributions," and "model tampering." Adaptive noise ensures balanced privacy protection across all gradient channels; threshold key design eliminates the risk of single-point leakage during decryption; contribution / budget dual-dimensional snapshots provide a quantitative reference for cross-departmental incentives; and model governance tickets establish an on-chain identity for new models, allowing for direct version identification during subsequent audits and rollbacks.
[0083] Step four involves binding the authorization hash to the task descriptor, aligning the training budget with the budget mapper, and using a gradient pruner and admission contract for double-insurance filtering of out-of-bounds inputs. Subsequently, differential privacy noise and homomorphic encryption are used to transform the gradient into a dense vector, perform invisible addition operations at the aggregation node, and write it into the ledger according to contribution.
[0084] After the model output's budget history is confirmed by zero-knowledge proof, the on-chain version is solidified using the model hash, achieving a complete closed loop where the input has a budget, the process is invisible, and the output is traceable. The governance layer can monitor the weight and budget changes in the ModelLedger in real time. If abnormal weight concentration or budget exhaustion is detected, step five will adjust the policy tree or add budget.
[0085] The first four steps of urban multimodal digital twins have completed a five-in-one operational loop of data access, risk labeling, privacy annotation, on-chain auditing, and federated aggregation. However, in the real urban operating environment, any static configuration will eventually face dynamic challenges such as regulatory revisions, emergencies, budget fluctuations, and model drift. If these disturbances cannot be captured in a timely manner at a global scale and the balance between privacy, security, and performance cannot be adjusted in real time, it will quickly fall into the dual risks of compliance mismatch or model failure.
[0086] Step 5: Based on real-time monitoring and multi-dimensional evaluation, dynamically adjust the privacy budget, permission rules and federated training configuration, and write the update results back to the blockchain and the source end to achieve an adaptive closed loop of data security and model performance.
[0087] Step 501: Comprehensive Monitoring and Risk Quantification Engine A continuous monitoring system that spans multiple levels and modalities is needed to integrate blockchain operation logs, budget pool balances, federated model indicators, and external regulatory signals into a decision-making risk profile.
[0088] The event collector captures data from four channels: on-chain AccessEvent, on-chain ModelLedger, off-chain BudgetPool snapshot, and off-chain ModelMetric stream. The collector appends a generic identifier—event type, node identity, and timestamp—to each event and writes it to the time-series database. It also maintains a sliding window view for downstream second-level querying.
[0089] The original state vector output by the data collector Perform piecewise orthonormalization to obtain the normalized index vector. Smoothing out the differences in magnitude makes the risk synthesis function compatible with multimodal indicators, including: ; Where: Normalized index vector The state after removing units and dimensions from each dimension index is used for threshold comparison, rule matching, and optimizer / agent input. robust central vector On the sliding history window The robust center calculated dimension by dimension (usually the quantile median) represents the "normal level" of each indicator; : Scale coefficient vector, derived from historical quantiles, applied within the same historical window. The scaling factor for dimension-wise estimation; inverse of a diagonal matrix ,Will Inverting the vector after placing it on the diagonal to form a diagonal matrix is equivalent to performing a "division by scale" operation on each dimension of the vector. Original state vector In time step The collected multidimensional indicator vectors, assembled in a fixed order according to the "state dictionary," have not yet been standardized. Typical components are derived from: window statistics of the final sensitivity score, privacy budget consumption rate, the trajectory of the noise scaling factor, the utility and convergence signal of the federated training model, blockchain audit alarm density, permission policy tightening level, distribution drift intensity, and data quality labeling ratio, etc. By fusing fuzzy entropy weights with multi-head attention, an instantaneous comprehensive risk index is output. : ; In the formula: This is a normalized index vector; Subhead weights, normalized to , No. The non-negative weights assigned to each channel reflect the importance of each objective in the current scenario; The target number of channels to be included in the summary; : No. The attention subfunction, the th ... The mapping function of each "target channel" maps the overall state. Or, a selected subvector can be mapped to a scalar score; This allows it to simultaneously capture multi-dimensional risk signals such as a sudden drop in budget, a sharp decline in model accuracy, and abnormal call frequency.
[0090] When comprehensive risk Exceeding the threshold Or fall into the low-risk range At that time, the rule compiler starts; based on risk labels, node identities, and sensitivity levels, it generates a set of candidate actions. The actions include: increasing or decreasing budgets, raising noise levels, tightening or loosening permissions, and adjusting federal rounds. Each action is tied to a logical premise and a projected impact window. This translates risk scores into machine-executable regulatory actions, providing parameterized input for the decision-making process.
[0091] The logical clock correction and dual-buffering system of the collector significantly improve the verifiability of event arrival order, ensuring that even network fluctuations at edge nodes do not affect the main chain's real-time perception of risk trends; outlier guards prevent single-point extreme values from misleading risk scoring, maintaining the monitoring system's immunity to spoofing attacks; the combination of multi-head attention and regulatory strength coefficients enables the risk synthesis function to be dynamically adjustable between business priorities and regulatory rigidity, providing a more flexible governance tool than single-weight or linear weighting; the hot-load rule compiler allows the governance layer to update rules without restarting the system when laws and regulations are temporarily adjusted, achieving true online governance.
[0092] Step 502: Closed-loop decision issuance and effect evaluation Once the candidate action set is available, a trade-off must be struck between budget constraints and model performance, and the optimal action must be efficiently and safely distributed to each module.
[0093] Set the optimization objective: Minimize the overall risk index. Maximizing model utility Constraining total budget consumption Not exceeding the upper limit. Using the NSGA-III algorithm from... Searching for Pareto optimal actions : ; Where: action vector An ordered set of joint adjustment instructions (such as noise scaling fine-tuning, federated training concurrency and rounds, permission thresholds, budget reallocation ratios, etc.) from the governance layer at the current time step for system parameters. Decision-making timeline index for governance and monitoring; : Budget cap, the upper limit of the single-step or sliding window budget allowed under the current regulatory window and scenario; Thus, a balance is found in the three-dimensional space of risk, utility, and budget, avoiding extreme bias in a single indicator.
[0094] Once the decision-maker selects the Pareto optimal action... Immediately generate a ControlTx transaction on-chain, writing the action summary, target node, budget flow, and effective time into the GovernanceContract; simultaneously, push precise parameters to the target node's control daemon process via gRPC stream, ensuring atomic consistency between on-chain Tx and off-chain execution. This dual-write mechanism between the chain and the endpoint prevents situations where changes are made on-chain but not on the node, or where changes are made on the node but not on the chain, resulting in a loss of connectivity.
[0095] Upon receiving the instruction, the node updates its local budget pool, noise parameters, or federated training configuration, and generates an AckEvent write chain. The governance layer defines a sliding window. Assess changes in the ternary indicators and generate window differences. : ; Where: Observation index sequence : at time step Collected indicator vectors ; Window width The number of intervals used in the trend calculation is a positive integer; The average first difference within the window, a positive value indicates that the indicator is generally upward, a negative value indicates that it is downward, and the larger the absolute value, the faster the change; If the rate of change of risk The rate of change in task utility decreased while the rate of change in budget expenditure remained stable. If the instruction is within the budget track, it is considered successful; otherwise, it enters the re-optimization loop.
[0096] ,by Based on, take the nearest The average first difference of each interval; model utility (e.g., dimensionless scores for model performance and service quality) are based on the average first-order difference; Measured by budget consumption (which can be defined as the amount of privacy budget consumed per unit time / unit round, or its standardized component) is the average first-order difference; To mitigate the impact of long-term environmental changes, the governance layer trains the meta-policy network in parallel, autonomously learning the risk-action-result triple sequence and dynamically adjusting the NSGA-III weights. The network input is... The output is a new weight vector. To reduce decision convergence time and adapt to regulatory changes, meta-policy learning injects experience into the optimizer, making the system smarter with use.
[0097] Step 501 normalizes the multi-source event flow from both on-chain and off-chain sources and synthesizes a comprehensive risk index. This provides a set of candidate actions for step 502; step 502 selects the optimal action through multi-objective optimization. The weights are then iteratively upgraded using sliding window difference and meta-policy learning, with the parallel chain and end-to-end dual-channel distribution.
[0098] The decision-maker introduces a dual mechanism of Monte Carlo simulation and meta-policy weighting, elevating governance decisions from static thresholds to adaptive intelligence based on probability prediction and experience fusion. Chain-end synchronization employs transaction replay and timeout alarms to ensure that each control command has a verifiable execution trajectory, eliminating governance black holes caused by chain-end inconsistencies. Shadow chain diversion and red-green signal visualization greatly reduce the pressure of governance metrics on the main chain performance while providing decision transparency. The meta-policy network, through LSTM long memory and regulatory cold start distillation, enables the system to operate stably during long-term evolution and regulatory changes.
[0099] Step five establishes a closed-loop governance system of monitoring, quantification, decision-making, execution, feedback, and learning: a multi-source event pipeline collector gathers blockchain and off-chain indicators, outputs comprehensive risk through indicator normalization and attention fusion; the rule compiler translates risk threshold triggers into action sets, and the decision-maker uses multi-objective optimization to find the best action in the three-dimensional Pareto plane of risk, utility, and budget; the chain-end synchronization mechanism ensures that governance instructions are implemented atomically, the sliding window evaluation ensures that governance actions truly bring about the expected improvement, and the meta-policy network feeds governance experience back into the optimizer, making the next round of decision-making more accurate and efficient.
[0100] Within the first five-step framework, the city's multimodal digital twin system has established a five-level closed loop encompassing data access, local privacy, on-chain auditing, security aggregation, and policy regulation, enabling cyclical security self-adaptation. However, urban situations and regulatory environments often exhibit a mix of sudden and gradual changes, raising concerns about privacy budgets. The optimal solution, in relation to model utility and accuracy, drifts over time rather than being a static constant. While a single multi-objective optimizer can instantaneously integrate the risk index... Maximizing model utility Constraining total budget consumption However, it struggles to handle long-term game-theoretic contradictions: privacy and security tend to increase noise at the expense of accuracy, while performance requirements tend to reduce noise and improve prediction; the two are mutually constraining and subject to external constraints. Therefore, adding a self-game learning layer to the governance layer is imperative.
[0101] Step 6: Introduce a dual-agent game network of privacy agent and performance agent at the governance layer, continuously play strategy games on risk index, model utility and budget consumption, and distill the optimal policy into a multi-objective optimizer to achieve cross-cycle privacy-performance balance and adaptation.
[0102] Step 601: Game Environment Modeling and Multi-Agent Training Loop Traditional single-agent or static optimization methods cannot capture privacy-performance adversarial structures; by constructing a game environment that includes dynamic regulations and budget constraints, privacy agents are trained in parallel. With performance agent This allows two parties to compete and cooperate in the same state space to obtain the Pareto boundary.
[0103] Environment in time step Broadcast state vector to two agents : ; in, Comprehensive risk index; Model utility score; Remaining budget rate; The current highest sensitivity level; Regulatory strength coefficient; Strength of external business trends; Privacy Proxy Action Set Includes noise amplification, restricted permissions, and reduced federation rounds; performance proxy action set. Includes noise reduction, budget expansion, and increased federal rounds. Each time both sides make a move simultaneously, the environment reference mapping table... Analyze and handle conflicts: When actions are mutually exclusive, execute the rule of high regulatory priority and low budget conservatism to avoid decision deadlock.
[0104] Building Privacy Proxy Rewards : ; In the formula: Risk penalty weight, value ; : Regulatory violation indication; if the current step triggers a regulatory or compliance red line, the value is 1, otherwise it is 0; determined by streaming rules. In accordance with compliance module criteria (certificate revocation still in use, unauthorized access, failure to meet multi-signature threshold, etc.); Performance Agent Rewards : ; In the formula: : Budget consumption penalty weight, non-negative calibration parameter; Utility reward weight, a non-negative calibration parameter; : Constraints on penalties for violations, non-negative calibration parameters; Budget red line indicator function: 1 when the budget balance / sliding window quota touches or exceeds the preset red line, otherwise 0; determined jointly by the budget contract and the monitor (including scenarios such as low water level, over-quota, and abnormal accelerated consumption). Privacy agents emphasize reducing In contrast to budget consumption, performance agents emphasize improving... and lower This creates competition.
[0105] Using distributed PPO, the strategies of both sides are updated in parallel on a multi-core GPU server. To avoid learning from getting stuck in infinite adversarial learning, an experience alignment cache is set up, and high-value action fragments of the opponent in the previous round are injected into the replay buffer of the player. This enhances the diversity of strategies and promotes the game toward Nash equilibrium. This ensures that adversarial learning is stable and does not cause oscillations due to an overly strong unilateral strategy.
[0106] The stable penalty term of the dual reward ensures that the system will not experience "sawtooth" privacy-performance fluctuations due to proxy adversarial behavior, while the innovation incentive keeps the strategy evolution diverse and prevents early convergence to suboptimal. The dual-tower Actor-Critic and cold-backup weights provide steady-state guarantees for large-scale parallel training.
[0107] Step 602: Online Refresh of Strategy Distillation and Treatment Parameters The strategy obtained from game convergence needs to be transformed into parameters that can be executed by the governance layer; otherwise, it can only remain at the simulation level.
[0108] Extract from each agent's policy network , with temperature Distillation is performed to obtain the soft target distribution, and then a lightweight unified policy is trained. Approximate the weighted sum of both sides, using a merged adversarial strategy as a single model to reduce deployment overhead: ; In the formula: For distillation loss; teacher's numerical vector The action score (logits) vector generated by the "teacher side" has the same dimension as the action space; The temperature of softmax controls the "softness" or "hardness" of the teacher distribution, and its value is greater than 0.
[0109] ,right The summation of the component exponents forms the denominator of the softmax function; Kullback–Leibler divergence; The parameter mapper will implement a lightweight, unified strategy. The output action is mapped to three types of governance parameters: differential privacy noise scaling factor. Budget redistribution vector Permission compaction level The mapping rules are statically constrained by the risk-sensitivity level table and can be hot-updated to prevent the policy from generating out-of-bounds parameters.
[0110] The governance layer packages the updated parameters into the DistillCtrlTx transaction and writes them to all nodes on both the blockchain and the client side; at the same time, it updates the multi-objective optimizer weight vector. This ensures consistency with the distillation results. DistillCtrlTx carries a version stamp and hash, so that any subsequent node can only reference the latest version of the parameters, preventing the remnants of old version strategies.
[0111] Shadow Chain Recycling Performance Changes at the Start of the Next Monitoring Period Risk changes If the measured deviation exceeds the distillation confidence interval, a fast distillation correction is triggered, updating the lightweight unified strategy with new logits. The chain is rewritten to implement a monitoring-distillation-correction microcycle, ensuring that the strategy performs in the real world in accordance with the simulation prediction.
[0112] Each agent policy network refers to a parameterized decision model constructed for the privacy agent and the performance agent respectively, using a unified robust standardized state vector. As input, the output is a constrained distribution of control actions: Privacy-side strategy A performance-side strategy that tends to increase noise scaling and permission thresholds to reduce risk. The approach prioritizes optimizing federated training concurrency and rounds to improve efficiency; each network uses its own immediate rewards. and (Depend on , , The candidate actions are updated based on the red line indicator and other parameters. First, "action mask / projection" satisfies budget and compliance constraints, and then the arbitration module synthesizes it into a unified control command. Distribute to noise mapping, permission contracts, and federated scheduling; policy parameters The version hash of the weighting factor is recorded with each transaction.
[0113] Step 601 uses multi-source state vectors to drive adversarial learning between the privacy agent and the performance agent within the PPO framework, which will have a long-term effect. - Precision game mapping is used to optimize rewards; step 602 adopts a temperature distillation compression adversarial strategy and transforms the game output into actual governance parameters through a parameter mapper. Then, the differential privacy coefficient, budget allocation and permission level are refreshed through chain-end instructions, and online correction is ensured by performance write-back, which complements the multi-objective optimizer in step five.
[0114] By employing adaptive temperature distillation and policy pruning, the system maintains initial policy exploration space while reducing long-term deployment costs. The hybrid mechanism of Piece-wise mapping and soft circuit breakers eliminates uncertain jumps at the end of policy mapping, stabilizing system parameters. CRDT concurrent tracing fields and semantic version numbers allow multiple batches of distillation to be linearly merged and clearly traced in a distributed environment. Trust interval fast distillation correction combined with canary releases provides a "try before you go" safety buffer for real-time governance policies.
[0115] Step Six: Add a privacy-performance self-playing intelligent agent pair to the top level of governance, and use a competition-cooperation mechanism to mine... The long-term equilibrium point of the indicator is achieved by injecting game theory knowledge into the existing multi-objective optimizer through temperature distillation and parameter mapping, realizing a dynamic win-win situation for privacy and utility; chain-end instructions and version control ensure that distillation parameters are implemented in real time and are fully traceable, and performance write-back ensures that the strategy is synchronized with reality.
[0116] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0117] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0118] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0119] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0120] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for secure management of urban multimodal digital twin data assets, characterized by: include, Establish a unified access gateway to aggregate multi-department, multi-modal data, automatically classify and grade it, generate sensitivity labels for each piece of data and bind privacy-level metadata; Locally at the data provider, differential privacy or desensitization is implemented according to the metadata at different levels, and noise parameters are dynamically adjusted through policy mapping to balance availability and risk. A permissioned consortium blockchain is built, incorporating departmental nodes. Smart contracts are used to record access, authorization, and model training on the blockchain, and fine-grained access control and tamper-proof auditing are implemented according to the strategy. With the on-chain authorization, each node trains locally and only uploads model updates processed by differential privacy and encryption. Secure multi-party computation or homomorphic encryption is used to aggregate and generate a global model. Deploy a centralized policy governance layer to continuously collect on-chain logs, privacy budgets, and model metrics; when thresholds are reached, automatically issue updates to differential privacy, permissions, and federated training configurations to form a cross-step closed-loop control. A multi-agent reinforcement learning network with privacy agent and performance agent is deployed at the governance layer. Based on the real-time index, it self-plays and distills into a multi-objective optimizer, and refreshes the differential privacy parameters, federated training hyperparameters and permission policies online.
2. The asset security management method according to claim 1, characterized in that: Configure a cross-departmental protocol adaptation and mode conversion module in the unified access gateway to complete the mode mapping and time alignment of structured, semi-structured and unstructured data; Generate a globally unique identifier and source fingerprint for each piece of data, and bind and version the identifier and source fingerprint together with the privacy level metadata.
3. The asset security management method according to claim 2, characterized in that: Sensitivity scores are calculated for the accessed data based on fields and context. Entropy and distribution difference measures are combined and aggregated according to modality weights. Then, regularization updates are performed on the feature map constructed based on attribute relationships to smooth out mislabeling. Output the sensitivity label and label version hash and write it into the privacy level metadata.
4. The asset security management method according to claim 3, characterized in that: A differential privacy processing pipeline is established locally at the data provider. First, norm pruning is performed on the data or gradient that needs to be protected. Based on the sensitivity label, the privacy budget pointer is queried and the noise scaling parameter is mapped. The noise generator is called to generate perturbations and generate noise fingerprints. At the same time, the deduction record is recorded in the budget ledger.
5. The asset security management method according to claim 4, characterized in that: The noise scaling parameters are dynamically fine-tuned based on real-time risk and availability metrics using a rule engine or reinforcement learning strategy. A safety barrier is formed by logarithmic domain updates and rate of change constraints. All parameter adjustments generate control transactions, which are released in a canary manner by the governance layer and support version rollback and full-chain audit traceability.
6. The asset security management method according to claim 5, characterized in that: For each access or training operation, an original transaction is constructed locally. The transaction body includes a resource identifier, invocation intent, permission credentials, budget credentials, and data or model digest hash. After being jointly multi-signed, it is submitted to the permissioned consortium blockchain, written to the ledger through consensus, and an index mapping with the policy table is established.
7. The asset security management method according to claim 6, characterized in that: The permission contract employs an attribute-based access control policy, which evaluates the subject attribute, purpose attribute, environment attribute, and sensitivity tag to generate an access token with expiration and scope. It supports callbacks for token revocation and policy change events, and records the changes and trigger reasons on the blockchain.
8. The asset security management method according to claim 7, characterized in that: Each node performs training locally, calculates the local gradient, and forms a protected update after norm pruning and differential privacy perturbation. The protected update is then encrypted using homomorphic encryption or an equivalent encryption scheme to generate ciphertext, which is submitted to the aggregation end along with the gradient hash, algorithm version, and budget increment.
9. The asset security management method according to claim 8, characterized in that: The aggregation end sums the elements of the ciphertext in the ciphertext field to obtain the aggregated ciphertext. After threshold decryption or secure multi-party computation, the protected update sum is recovered. The global model update is calculated according to the predetermined weight and a model version number is generated. The model summary and aggregation process metadata are registered and sent back to each node.
10. The asset security management method according to claim 9, characterized in that: The governance layer continuously collects on-chain logs, budget ledgers, and model evaluation results to construct a state vector with a fixed field order and performs dimension-by-dimensional standardization according to robustness center and robustness scale to obtain a unified state. The unified state is then matched with trend features and rule base to generate a draft strategy recommendation.
11. The asset security management method according to claim 10, characterized in that: The governance layer compiles the proposed strategy into update instructions for differential privacy parameters, permission policies, and federated training hyperparameters. After applying rate limits and boundary constraints, these instructions form control transactions that are then sent to the corresponding modules. All update actions are associated with the version hash and canary releases are supported.
12. The asset security management method according to claim 11, characterized in that: A policy network for privacy and performance proxies is constructed, which uniformly receives the unified state as input. The action space covers noise scaling, permission thresholds, and training concurrency rounds. Each action is trained using its own reward definition, and the output candidate actions are processed by action masking and projection to meet budget and compliance constraints before entering arbitration.
13. The asset security management method according to claim 12, characterized in that: The teacher distribution is obtained by normalizing the action scores generated by arbitration or the teacher side through temperature. The divergence of the distribution output by the distillation strategy network is calculated to form the distillation loss. The parameters are updated in the shadow domain at a set period and then released to the production domain in grayscale. At the same time, the version information of the distillation parameters and rollback points is recorded.
Citation Information
Patent Citations
Federal learning-based privacy protection type large-scale model training and deployment method
CN118734360A
Data privacy protection method
CN119848936A
Federal learning method and system based on bidirectional feedback knowledge distillation and differential privacy
CN119990373A
Model encryption and privacy protection method oriented to artificial intelligence algorithm
CN120068123A
Financial privacy security alignment method and system based on federated learning and adversarial training
CN120470624A
Cited By
Rapid soil pollution detection method
CN121186336A
Engine system based on integrated privacy federal computing power
CN121327042A
High overload resistant data analysis method and system based on multi-dimensional buffer protection
CN122086681A
Multi-source heterogeneous asset management system information interaction method and device
CN122222305A
Multi-agent medical care teaching virtual tutor system, teaching method and medium
CN122347892A