Methods and systems for cross-border VPN data anonymization and transmission that are compatible with multiple regions and comply with regulations.
By parsing cross-border data packet headers to generate context metadata, dynamically constructing compliance and utility objective functions, and optimizing the de-identification execution plan, the problem of compliance adaptation and data utility balance in cross-border data transmission of VPN technology is solved, and the maximum preservation of data packet compliance and utility is achieved.
Patent Information
- Application Number
- CN202511952709.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2045-12-23
AI Technical Summary
Existing VPN technologies lack the ability to deeply perceive and dynamically adapt data content during cross-border data transmission, resulting in data distortion and a lack of cross-regional compliance adaptation, making it impossible to achieve an effective balance between compliance constraints and data utility.
By parsing the header of cross-border data packets to generate data flow context metadata, dynamically loading compliance constraint sets and task utility objective functions, performing semantic scanning to generate pre-analysis feature sets, constructing decision context information, and optimizing the optimal de-identification execution plan in real time, and combining difference set operations and compliance decision encapsulation data packets, dynamic compliance and utility maximization are achieved.
It maximizes the utility of retained data while meeting compliance constraints, avoids data distortion and unintended damage to non-sensitive data caused by traditional VPNs, and supports dynamic adaptation of compliance and business value for cross-border data transmission.
Smart Images

Figure CN121396664B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cross-border data security transmission, and in particular to a cross-border VPN data anonymization transmission method and system adapted to multi-regional compliance. Background Technology
[0002] In global cross-border data exchange, ensuring the secure and compliant flow of data across different jurisdictions has become a critical issue in the field of network communications. However, existing Virtual Private Network (VPN) technologies are mostly limited to encryption of the transmission channel, lacking deep awareness and dynamic adaptation capabilities regarding data content. How to address the problems of traditional VPN encryption's lack of knowledge about content, data distortion caused by static anonymization strategies, and the lack of cross-regional compliance adaptation amidst a complex and ever-changing international legal environment and diverse business needs has become a core challenge that urgently needs to be solved in the intelligent upgrading of cross-border data transmission technologies.
[0003] Chinese patent application CN120378135A discloses a method for ensuring cross-border data security. This method includes: acquiring the original data stream; marking the original data stream with sensitive information to obtain sensitive information annotations; parsing the contextual relationships of the sensitive information annotations to obtain a contextual relationship structure; performing semantic abstraction and replacement based on the sensitive information annotations and the contextual relationship structure to obtain preliminary anonymized data; generating a de-identified feature map based on the preliminary anonymized data to obtain a de-identified feature map; formulating a source-end strategy based on the de-identified feature map to obtain an initial source-end strategy; negotiating the strategy based on the initial source-end strategy and dynamically adjusting the strategy to obtain a negotiated and adjusted strategy.
[0004] However, current technologies still face numerous challenges. In the common scenario of multinational corporations conducting global collaborative training of deep learning models based on multi-source heterogeneous data, existing cross-border data transmission solutions generally face a zero-sum game between compliance constraints and data utility. Traditional VPN channels only provide blind encryption at the transport layer, lacking semantic awareness of the data payload content. This one-size-fits-all approach destroys the non-linear correlations within the data, leading to the loss of key semantic information. Once downstream training models receive distorted data that has lost its statistical regularity, they will be unable to capture effective gradient information, resulting in the loss function failing to converge or the trained model having poor generalization ability. This forces companies to choose between two risks when facing cross-border business: if they retain data utility, they face compliance triggers and export restrictions; if they adopt excessive anonymization, the training quality of core AI business will be difficult to maintain, resulting in economic losses. Summary of the Invention
[0005] To achieve the above objectives, this invention provides a cross-border VPN data anonymization transmission method that is compatible with compliance in multiple regions. The specific technical solution is as follows:
[0006] The header of the original cross-border data packet is parsed to generate data flow context metadata. The source compliance region, destination compliance region and business type identifier in the data flow context metadata are used as indexes to dynamically load the compliance constraint set and the task utility objective function. Semantic scanning is performed on the data payload of the original cross-border data packet in parallel to generate a pre-analysis feature set. The data flow context metadata, compliance constraint set, task utility objective function and pre-analysis feature set are combined to construct decision context information.
[0007] Extract the task utility objective function, pre-analysis feature set, and compliance constraint set from the decision context information. Combine them with the semantic tag set to instantiate the objective function and transformation strategy space for the undetermined transformation strategy. Execute the constraint optimization algorithm to solve for the optimal transformation strategy. Using the data position in the pre-analysis feature set as the spatial index, compile the optimal transformation strategy into an optimal de-identification execution plan composed of atomic execution instructions.
[0008] Each atomic execution instruction in the optimal desensitization execution plan is executed to generate a new data feature set. This set is then combined with the non-sensitive data extracted through difference operations to reconstruct a new data payload. Compliance decision and audit labels are then assembled, and the new data payload is injected to output the second cross-border data packet.
[0009] The second cross-border data packet is encapsulated into an encrypted third cross-border data packet at the VPN communication layer. Based on the decision context information, the optimal transformation strategy, and the second cross-border data packet, the log generation function is called to generate a decision audit log. The log is decrypted and quantified to compare the compliance decision and audit label with the target end's local compliance policy. If the comparison results are inconsistent, a policy update trigger signal is generated and sent back via the reverse channel, driving the source end policy update in a closed loop.
[0010] Furthermore, the method for constructing the decision context information includes:
[0011] Extract the five-tuple information from the header of the original cross-border data packet, use the five-tuple information as an index to query the pre-built federated mapping function, and output the data flow context metadata; the five-tuple information includes the source IP address, destination IP address, source port number, destination port number, and transport protocol; the data flow context metadata includes the data flow unique identifier, the source origin compliance region, the destination compliance region, and the service type identifier.
[0012] Based on the origin and destination compliance regions in the data flow context metadata, the corresponding compliance constraint sets are dynamically queried and loaded from the multi-region compliance policy database. These compliance constraint sets contain... The compliance constraints consist of data field types, risk quantification functions, and statutory risk thresholds;
[0013] Based on the business type identifier in the data flow context metadata, the task utility objective function is dynamically loaded from the pre-configured task utility configuration library. The task utility objective function quantifies the business requirements into a weighted sum of utility values based on utility characteristics through a utility quantification function and business importance weight.
[0014] Semantic scanning is performed on the data payload of the original cross-border data packets to extract the original data blocks. A pre-analysis feature set is constructed using a pre-analysis model and aggregated into decision context information by combining data flow context metadata, compliance constraint set, and task utility objective function.
[0015] Furthermore, the joint mapping function is a pre-built mapping mechanism, constructed based on a lookup table or rule matching engine, and the specific mapping dimensions executed include the geographic-legal mapping dimension and the network-business mapping dimension;
[0016] The geographic-legal mapping dimension is used to extract the source IP address or target IP address from the five-tuple information, match it in the pre-stored IP address segment index, and map it to the corresponding geographic-legal region identifier. The geographic-legal region identifier is used to index the legal jurisdiction to which the IP address segment belongs.
[0017] The network-service mapping dimension is used to extract the target port number or transmission protocol from the five-tuple information, perform association matching in the pre-stored service feature library, and map it to the corresponding service type identifier.
[0018] Furthermore, the dynamic loading step of the task utility objective function includes:
[0019] Using the business type identifier in the data flow context metadata as the index key, a matching search is performed in the pre-configured task utility configuration library, and the corresponding task utility objective function is instantiated.
[0020] The task utility objective function aims to maximize... The goal is the sum of the weighted utility values of each utility feature;
[0021] The method for calculating the weighted sum of utility values includes: traversing the associations of the business type identifiers. For each utility feature, the corresponding data variable is input as a formal parameter into a predefined utility quantification function to calculate a utility score; the utility score is then multiplied by the pre-configured business importance weight of the utility feature to obtain a weighted utility value; for all... The weighted utility values of each utility feature are summed to generate a total weighted utility value.
[0022] Furthermore, the compilation method for the optimal de-identification execution plan composed of atomic execution instructions includes:
[0023] Extract the task utility objective function and pre-analysis feature set from the decision context information. Use the semantic tag set to dynamically bind the business importance weight and utility quantification function in the task utility objective function with the original data blocks in the pre-analysis feature set to construct an objective function with the undetermined transformation strategy as the only independent variable.
[0024] Extract compliance constraint sets and pre-analysis feature sets from decision context information, establish the mapping relationship between compliance constraints and original data blocks using semantic tag sets, and instantiate risk quantification functions and statutory risk thresholds into inequality constraints for undetermined transformation strategies, defining them as transformation strategy space;
[0025] The constrained optimization algorithm is used to solve the undetermined transformation policy that maximizes the objective function. If the undetermined transformation policy belongs to the transformation policy space, the objective function is used as a positive reward; otherwise, a negative reward is applied. The process is iteratively converged and the optimal transformation policy is output.
[0026] Using the data locations in the pre-analyzed feature set as spatial indices, the optimal transformation strategy is compiled into an optimal de-identification execution plan, which is then generated by... Each instruction consists of atomic execution instructions that include specific actions, location information, and action parameters.
[0027] Furthermore, the step of constructing the objective function with the undetermined transformation strategy as the sole independent variable includes:
[0028] Using the semantic label set carried by each data feature instance in the pre-analyzed feature set as the association index, the objective function instantiation operation is performed;
[0029] The objective function instantiation operation includes: establishing a binding relationship between the predefined business importance weight and utility quantification function in the task utility objective function and the original data block in the pre-analysis feature set, and constructing an objective function with the undetermined transformation strategy as the only independent variable;
[0030] The calculation logic of the objective function includes: traversing each data feature instance in the pre-analysis feature set in the outer layer, and traversing all utility feature indices contained in the semantic label set of each data feature instance in the inner layer; calling the utility quantization function according to the utility feature index, inputting the original data block processed by the pending transformation strategy into the utility quantization function to obtain the utility score; performing a multiplication operation on the utility score and the corresponding business importance weight to generate a weighted utility value, and performing an accumulation operation on the weighted utility values corresponding to all utility feature indices of all data feature instances, and using the accumulated result as the output value of the objective function.
[0031] Furthermore, the inequality constraint includes: applying the undetermined transformation strategy to the original data block of the data feature instance to generate a transformed intermediate data state; calling the risk quantification function corresponding to the compliance constraint, using the intermediate data state as input parameters, to calculate and obtain a risk score; verifying whether the risk score is less than or equal to the statutory risk threshold corresponding to the compliance constraint; when the undetermined transformation strategy satisfies the verification for all data feature instances, it is determined that the undetermined transformation strategy satisfies the inequality constraint and is determined to fall into the transformation strategy space.
[0032] Furthermore, the method for outputting the second cross-border data packet includes:
[0033] Traverse each atomic execution instruction in the optimal desensitization execution plan, locate the original data block in the data payload based on the position information, call the instruction execution function to perform the transformation operation defined by the specific action and action parameters to generate a new data block, and aggregate all the new data blocks through the union operation to generate a new data feature set;
[0034] Aggregate the location information of all atomic execution instructions in the optimal desensitization execution plan to construct the sensitive data region. Use the difference operation to remove the untouched non-sensitive data from the original data payload. Then call the payload reconstruction function to physically splice the non-sensitive data and the new data feature set to output the new data payload.
[0035] The optimal transformation strategy is applied to a hash function to generate decision criteria. The compliance objectives, final risk score, and final utility score in the decision context information are extracted, assembled into compliance decision and audit labels, and the new data payload is injected by calling the embedding function to output the second cross-border data packet.
[0036] Furthermore, the method for driving the source-side policy update includes:
[0037] The second cross-border data packet is used as the payload input of the VPN communication layer. The VPN standard encryption function is called to perform protocol encapsulation and encryption, generating an encrypted third cross-border data packet.
[0038] After the third cross-border data packet is sent, the log generation function is called to generate a decision audit log based on the decision context information, the optimal transformation strategy, and the second cross-border data packet.
[0039] The third cross-border data packet is decrypted to extract compliance decision and audit tags. The policy verification function is used to quantitatively compare the source decision results carried in the compliance decision and audit tags with the target local compliance policy. If the comparison results are inconsistent, a policy update trigger signal is generated and sent back through the reverse channel to trigger the source policy update.
[0040] A cross-border VPN data anonymization transmission system adapted to multi-regional compliance is provided to implement the aforementioned cross-border VPN data anonymization transmission method adapted to multi-regional compliance. The system includes a decision context construction module, a joint policy generation module, a data reconstruction module, and an encryption and auditing module.
[0041] The decision context construction module is used to parse the header of the original cross-border data packet to generate data flow context metadata. Using the source compliance region, destination compliance region, and business type identifier in the data flow context metadata as indexes, it dynamically loads the compliance constraint set and task utility objective function, performs semantic scanning on the data payload of the original cross-border data packet in parallel to generate a pre-analysis feature set, and combines the data flow context metadata, compliance constraint set, task utility objective function, and pre-analysis feature set to construct decision context information.
[0042] The joint strategy generation module is used to extract the task utility objective function, pre-analysis feature set and compliance constraint set from the decision context information, combine the semantic tag set, instantiate the objective function and transformation strategy space for the undetermined transformation strategy, execute the constraint optimization algorithm to solve the optimal transformation strategy, and compile the optimal transformation strategy into an optimal desensitization execution plan composed of atomic execution instructions using the data position in the pre-analysis feature set as the spatial index.
[0043] The data reconstruction module is used to execute each atomic execution instruction in the optimal desensitization execution plan to generate a new data feature set, combine it with the non-sensitive data extracted by the difference set operation to reconstruct a new data payload, assemble compliance decision and audit labels, inject the new data payload and output the second cross-border data packet;
[0044] The encryption and auditing module is used to encapsulate the second cross-border data packet into an encrypted third cross-border data packet of the VPN communication layer, and based on the decision context information, the optimal transformation strategy and the second cross-border data packet, call the log generation function to generate a decision audit log, decrypt and quantify the comparison between the compliance decision and audit label and the target end local compliance policy. If the comparison results are inconsistent, a policy update trigger signal is generated and sent back through the reverse channel to drive the source end policy update in a closed loop.
[0045] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0046] This invention generates data flow context metadata by parsing the header of the original cross-border data packet. Based on this metadata and the pre-analyzed feature set, decision context information is constructed. Before the data enters the encrypted channel, compliance constraints and task utility objectives are quantified into mathematical optimization inputs. This solves the problems of traditional VPNs being unaware of data content and unable to perform differentiated compliance, as well as the data distortion caused by static desensitization. It maximizes the preservation of data utility under compliance constraints.
[0047] This invention constructs a joint optimization objective function and a compliance constraint space with undetermined transformation strategies as independent variables, and uses a constraint optimization algorithm to solve them in real time. This enables the dynamic generation of desensitization strategies that maximize business utility while strictly meeting regional compliance constraints, thus solving the problem of business value loss caused by static desensitization strategies failing to balance compliance rigidity and utility flexibility.
[0048] This invention performs precise transformation of sensitive data by executing the optimal desensitization execution plan, and reconstructs the payload by combining the stripped non-sensitive data with difference operations. This achieves partitioning processing of the original data packets, avoiding the problems of inadvertently damaging non-sensitive data and zeroing out data utility caused by the one-size-fits-all processing of traditional static desensitization.
[0049] This invention solves the problems of unauditable data content caused by traditional VPN encryption and the inability of static policies to adapt to dynamic changes in the regulations of the target side by encapsulating compliance decisions and audit labels into cross-border data packets and transmitting them in an encrypted manner, and by implementing a reverse feedback closed loop at the target end based on the quantitative comparison of label content with local policies. Attached Figure Description
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0051] Figure 1 This is a flowchart illustrating the principle of the cross-border VPN data anonymization transmission method adapted to multiple regions and compliant in this invention.
[0052] Figure 2 This is a functional block diagram of the cross-border VPN data anonymization transmission system adapted to multiple regions and compliant with the present invention. Detailed Implementation
[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] Example 1:
[0055] Please see Figure 1 As shown, this embodiment provides a method for cross-border VPN data anonymization transmission that is compatible with compliance in multiple regions, including:
[0056] Step S1000: Parse the original cross-border data packet. masthead To generate data stream context metadata Data flow context metadata The source of compliance area compliant areas of the destination and business type identifier For indexing, dynamically load compliance constraint sets. and task utility objective function Parallel processing of the original cross-border data packets Data payload Perform semantic scanning to generate a pre-analyzed feature set. Combined with the data flow context metadata Compliance constraint set Task utility objective function and pre-analysis feature set Constructing decision context information .
[0057] Specifically, this step aims to address the issue of traditional Virtual Private Networks (VPNs) lacking semantic awareness of transmitted content, thus hindering compliance processing. Before data enters the encrypted channel, the original cross-border data packets are processed. We conduct in-depth analysis and proactively model and encapsulate the network, legal, and business attributes of data flows into structured decision context information. This approach unifies legal compliance issues and business objective issues into mathematical optimization inputs at the data parsing layer, providing a technical prerequisite for achieving real-time and dynamic trade-offs downstream.
[0058] Further, step S1000 includes:
[0059] Step S1100: Extract the original cross-border data packet. Interim report header The five-tuple information is used as an index to query a pre-built union mapping function and outputs data stream context metadata. The five-tuple information includes the source IP address, destination IP address, source port number, destination port number, and transport protocol; the data stream context metadata Including unique identifiers for data streams compliant areas of origin compliant areas of the destination and business type identifier .
[0060] Specifically, this step aims to address the shortcomings of traditional VPNs, which lack awareness of data stream identity and thus cannot enforce differentiated compliance, by analyzing the incoming raw cross-border data packets. The process involves parsing anonymous network data packets and assigning them dual contexts of legal and business attributes, outputting structured data flow context metadata. The data stream context metadata This will serve as the index for loading compliance constraints in subsequent steps S1200 and loading task utility targets in step S1300, ensuring that the strategy loading is targeted and adaptable.
[0061] In the specific implementation process, this step involves processing the original cross-border data packets at the VPN data path entry point. masthead Perform deep flow detection. The header... Includes original cross-border data packets The network and transport layer technical information does not involve the data payload content, but only carries the five-tuple information used for network routing, namely the source IP address, destination IP address, source port number, destination port number, and transport protocol.
[0062] Unlike traditional technologies such as firewalls or Deep Packet Inspection (DPI), which are limited to classification or blocking operations at the network security level, the innovation of this step lies in actively constructing a data context using the results of deep stream inspection, rather than passively defending against security threats. The system uses the five-tuple information as a composite query key to access a pre-built network-business-geographic joint mapping function, thereby achieving dynamic binding of network address attributes with legal and business attributes.
[0063] The joint mapping function is a pre-configured, deterministic mapping mechanism that can be implemented through lookup tables or rule matching engines. Its functionality includes two dimensions: First, the geographic-legal mapping dimension. It extracts the source or target IP address from the 5-tuple information, matches it against a pre-stored IP address range index, and maps it to its corresponding geographic-legal region. For example, the IP address range "198.51.100.x" is mapped to the geographic-legal region "DE-GDPR_Zone," indicating that the IP address belongs to the jurisdiction of Germany. Second, the network-service mapping dimension. It extracts the target port number or combines it with the transmission protocol, performs association matching against a pre-stored service feature library, and maps it to a specific service type identifier. For example, when the target port is 8080 and the host domain name is "dl.example.com," it resolves to the service type "DL_Training_Driver_Model," representing the service scenario of driver model training.
[0064] Using the aforementioned joint mapping function, this step performs real-time reverse parsing at the packet level of the network layer and transport layer observables contained in the five-tuple information into data flow context metadata with both legal and business semantics. .in, It represents data flow context metadata and is a structured dataset; It represents a unique identifier for a data stream, which is a string or hash value. It is usually generated by hashing the 5-tuple information. Its value remains unique within the same business session and is used to identify and associate contexts belonging to the same session. The originating region is indicated by the joint mapping function based on the header. The geographic-legal region identifier obtained by reverse lookup of the source IP address corresponds to a specific legal jurisdiction. The destination compliance region is indicated by the joint mapping function based on the header. The geographic-legal region identifier obtained by reverse lookup of the target IP address, its value and the compliance region of the origin. The legal framework for cross-border data transfers will be jointly determined. The business type identifier representing the downstream task is determined by the joint mapping function based on the header. The service classification obtained by reverse parsing the target port number and the transmission protocol corresponds to a specific data utility target.
[0065] Step S1200, based on data flow context metadata The source of compliance area and destination compliance areas Dynamically query and load the corresponding compliance constraint sets from the multi-region compliance policy database. The compliance constraint set Include The data field type Risk quantification function and statutory risk threshold Compliance constraints constituted .
[0066] Specifically, this step aims to address the problem of data utility loss caused by a one-size-fits-all approach of strong anonymization in existing technologies such as static desensitization, by utilizing the data stream context metadata output in step S1100. Especially the compliant areas of origin. and destination compliance areas Instead of using static, directive sets of desensitization rules, it proactively quantifies and constructs a set of computable mathematical constraints from ambiguous legal provisions in different judicial regions, and encapsulates them into a structured set of compliance constraints. This provides a legal basis with binding boundaries for subsequent steps.
[0067] In the specific implementation process, this step is based on the data flow context metadata. The source of compliance area and destination compliance areas The compliance constraint loader dynamically queries and loads the corresponding, structured sets of compliance constraints from the pre-configured Multi-Region Compliance Policy Database (MR-CPD). The set of compliance constraints It is a package Compliance constraints The set of is expressed by the following formula:
[0068] ;
[0069] in, Represents the set of compliance constraints The A compliance constraint is a triple, containing the first... All the definition information required for a constraint, i.e., the data field type. Risk quantification function and statutory risk threshold ; Indicates compliance constraints The index is a compliance constraint. The sequence number, with a value range from 1 to... Positive integers used in the compliance constraint set Each structured compliance constraint is uniquely identified in the code. ; Indicates compliance constraints The total number of them, whose value is equal to the set of compliance constraints. Includes all compliance constraints The total number, i.e., the metadata for the current data stream context activated during this load. The sum of constraint rule entries.
[0070] The Indicates the first Compliance constraints The applicable data field type is an identifier whose value is a predefined string used to apply the compliance constraint. It is then associated and matched with the corresponding fields in subsequent steps.
[0071] The Indicates the first Compliance constraints The corresponding risk quantification function is a predefined mathematical function or assessment model that takes anonymized data as input and outputs the compliance risk score remaining in the data, thereby achieving quantitative calculation of ambiguous legal requirements.
[0072] The Indicates the first Compliance constraints The corresponding statutory risk threshold is a fixed numerical upper limit defined by regulations or preset by strategies, used to limit the risk quantification function. The valid range of output values provides rigid constraint boundaries for subsequent steps.
[0073] Step S1300, based on data flow context metadata Business type identifier in Dynamically load the task utility objective function from the pre-configured task utility configuration library. The task utility objective function By using a utility quantification function and business importance weights, business needs are quantified into a weighted sum of utility values based on utility characteristics.
[0074] Specifically, this step is executed in parallel with step S1200, aiming to provide a two-way balanced loading mechanism for compliance constraints and task utility objectives, based on the data flow context metadata output in step S1100. Especially the business type identifier Instead of loading simple, static weight lists or sets of empirical parameters, it proactively quantifies vague business requirements into a set of computable mathematical optimization objectives, encapsulating them into a task utility objective function. This provides guidance for maximizing utility in subsequent steps.
[0075] In the specific implementation process, this step is based on the data flow context metadata. Business type identifier in Using this as the query key, the corresponding task utility objective function is retrieved from the pre-configured task utility configuration library and instantiated. The innovation of this step lies in the loading of the task utility objective function. With the risk quantification function defined in step S1200 Mathematically, they are dual: the former aims to maximize utility, while the latter uses risk satisfaction constraints as its boundary. Together, they form the dual-objective structure for desensitization optimization in subsequent steps. This step transforms qualitative business requirements, such as preserving visual characteristics, into computationally calculable and optimizable mathematical objectives.
[0076] The task utility objective function This is the final output of this step, which is itself an optimization objective, aiming to make... The sum of the weighted utility values of each utility feature is maximized to define the mathematical expression that needs to be maximized in subsequent steps when making desensitization decisions. The calculation method for the sum of the weighted utility values is defined as follows: by maximizing the sum of the weighted utility values of the first utility feature... Each utility characteristic's data variable is input into its corresponding utility quantification function to calculate a utility score; then, this utility score is multiplied by the business importance weight corresponding to that utility characteristic to obtain the 1st utility characteristic. The weighted utility values of each utility characteristic; ultimately for all The weighted utility values of each utility feature are summed.
[0077] in, This represents the total number of utility characteristics, and its value is equal to the current business type identifier. The total number of utility feature entries on which it depends; the data variables of the utility features serve as input independent variables or formal parameters, i.e., placeholders, for use in subsequent steps to evaluate the utility quantification function from the original cross-border data packet. Data payload Extracted, corresponding to the The specific data blocks of each utility feature are linked and used as actual data in the subsequent optimization solution to be substituted into the utility quantification function for calculation; the utility quantification function is a predefined mathematical function or evaluation model, whose input is "data processed by a certain desensitization operation", and whose output is the residual utility contribution of the data to the downstream task, that is, a normalized scalar. An index representing a utility characteristic, with values ranging from 1 to... A positive integer, used in the task utility objective function. The summation calculation iterates through all relevant utility features; the business importance weight is a pre-configured non-negative scalar, used to define the first [value] in the weighted summation. The relative importance of each utility feature to downstream tasks.
[0078] Step S1400, process the original cross-border data packet. Data payload Perform a semantic scan to extract the raw data blocks. Construct a pre-analysis feature set using a pre-analysis model. and combined with data flow context metadata Compliance constraint set Task utility objective function Aggregated into decision context information .
[0079] Specifically, this step aims to process the input raw cross-border data packets. Data payload The data is parsed to identify the data blocks to be decided, and then bound to the abstract functions loaded in steps S1200 and S1300 using semantic tags, outputting a pre-analysis feature set. This completes the decision context information in step S1000. The construction.
[0080] In the specific implementation process, this step runs in parallel with steps S1200 and S1300, and processes the original cross-border data packets. Data payload Perform a semantic scan. The data payload... It is the original cross-border data packet. The part that carries business data, such as images and text. The semantic scanning is performed through a pre-analysis model, such as a set of or a multimodal classification and detection model, like the YOLO-Fastest model or the multimodal CLIP model.
[0081] The pre-analysis model analyzes the data payload. Perform semantic scanning to locate and extract the original data blocks. Record the original data block In data net load Data location in and the original data block Perform semantic analysis on the content, generate one or more semantic tags, and encapsulate them into a semantic tag set. .
[0082] After semantic scanning, the pre-analysis model outputs a pre-analysis feature set. This pre-analysis feature set It is a package Item by index Identified data feature instances A collection of each data feature instance It is a block of raw data Semantic tag set and data location The triplet. Among them, This represents the total number of data feature instances, and its value is equal to the pre-analyzed feature set. The elements contained therein, that is, the semantic scan in the data payload of this data scan. Data feature instances identified in sum; The index represents the data feature instance, with a value ranging from 1 to... Positive integers used in the pre-analysis feature set Each identified data feature instance is uniquely identified. .
[0083] The original data block From data payload The actual data extracted is used as the risk quantification function defined in step S1200. The actual input data when the utility quantification function defined in step S1300 is executed in subsequent steps.
[0084] The semantic tag set It is a set of labels used to separate raw data blocks and the risk quantification function defined in step S1200 It is bound to the utility quantification function defined in step S1300. Specifically, this semantic tag set... The included tags are used to link to two abstract definitions simultaneously: first, the set of compliance constraints encapsulated in step S1200. The Middle Compliance constraints Applicable data field types Secondly, it matches the task utility objective function encapsulated in step S1300. The Middle Each utility characteristic corresponds to a data variable.
[0085] The data location It is the original data block In data net load The offset and length information are used for positioning and replacement operations after desensitization decisions are made in subsequent steps.
[0086] Thus, the four sub-steps of step S1000 are completed collaboratively, namely steps S1100, S1200, S1300, and S1400. The pre-analysis feature set output by this step... The data stream context metadata output in step S1100 The set of compliance constraints encapsulated in step S1200 The task utility objective function encapsulated in step S1300 The data is aggregated together to output decision context information. .
[0087] Step S2000: Extract decision context information Medium task utility objective function Pre-analysis feature set and compliance constraint set Combined with semantic tag set Instantiate the undetermined transformation strategy objective function and transformation strategy space The optimal transformation strategy is solved by executing a constrained optimization algorithm. and with pre-analysis feature set Data location in For spatial indexing, the optimal transformation strategy will be used. Compiled into instructions that can be executed atomically. The optimal de-identification execution plan .
[0088] Specifically, this step no longer loads a fixed de-identification template, but instead receives a constraint optimization problem instance for the current data stream compiled in step S1000, i.e., decision context information. Regarding the context information of this decision The algorithm performs real-time solution and dynamically generates an optimal transformation strategy that simultaneously satisfies the rigid compliance boundary in step S1200 and maximizes the business utility in step S1300. And compile it into an optimal de-identification execution plan that can be executed in subsequent steps. .
[0089] Further, step S2000 includes:
[0090] Step S2100, from decision context information Extract the task utility objective function and pre-analysis feature set Using semantic tag sets The objective function of task utility Business importance weights and utility quantification functions and pre-analysis feature sets raw data blocks in Perform dynamic binding and construct a transformation strategy to be determined. Objective function with only one independent variable .
[0091] Specifically, this step aims to receive the decision context information output by step S1000. And the task utility objective function encapsulated therein from step S1300. and the pre-analysis feature set from the output of step S1400 The two are linked by semantic tags as a bridge, and an instantiation is generated with an undetermined transformation strategy. An objective function that can be maximized with a single independent variable. .
[0092] In the specific implementation process, this step directly uses the decision context information. Extract the task utility objective function and pre-analysis feature set By pre-analyzing the feature set The semantic tag set provided in As a linking bridge, the task utility objective function The abstract business objectives defined in the text, namely the business importance weights and utility quantification functions, are related to the pre-analysis feature set. Each raw data block Perform the binding. Through this binding, this step instantiates a variable with a pending transformation strategy. Objective function with only one independent variable The undetermined transformation strategy It is the objective function The independent variable represents the definition of one or a set of abstract desensitization operations, such as blurring, noise overlay, feature extraction, or character discarding.
[0093] Instantiated target function The computational logic is as follows: traverse the pre-analyzed feature set through an outer summation layer. Each data feature instance in ; and for each data feature instance Perform an inner summation, traversing the utility feature index set. Index of each utility feature in The utility feature index set is mentioned above. It is a data feature instance semantic tag set Index of the utility features corresponding to all labels that identify utility features. A set used to determine the current data feature instance. Which utility quantification functions should be used to evaluate the data block?
[0094] In the inner summation process, from the task utility objective function Index of Utility Features The corresponding business importance weights and utility quantification functions are then used, and these utility quantification functions are applied to data feature instances. raw data blocks Strategy determined by transformation The resulting utility score is calculated and then multiplied by the business importance weight.
[0095] Step S2200, from decision context information Extract compliance constraint set and pre-analysis feature set Using semantic tag sets Establish compliance constraints and raw data blocks The mapping relationship will quantify the risk function. and statutory risk thresholds Instantiated for undetermined transformation strategy The inequality constraints are defined as the transformation policy space. .
[0096] Specifically, this step aims to receive the decision context information output by step S1000. and the set of compliance constraints encapsulated therein from step S1200 and the pre-analysis feature set from step S1400 By using semantic tags as a bridge to link the two, a rigid, mathematical transformation strategy space is instantiated and generated. .
[0097] In the specific implementation process, this step is executed in parallel with step S2100. Instead of loading or executing static, descriptive rules, such as "must be anonymized", it directly inherits the compliance functionization in step S1200.
[0098] This step involves decision context information. Extract compliance constraint set and pre-analysis feature set By pre-analyzing the feature set The semantic tag set provided in As a linking bridge, it will integrate compliance constraints. The abstract compliance constraints defined in the document, i.e., the risk quantification function Less than or equal to the statutory risk threshold Instantiated into the pre-analyzed feature set Each related raw data block Together, they define a transformation policy space. .
[0099] The transformation strategy space It is a set of undetermined transformation strategies that satisfy the universal quantization condition. The set. The universal quantization condition is a rigid constraint for the pre-analyzed feature set. Each data feature instance and through semantic tag sets Link to this data feature instance Every compliance constraint All must satisfy inequality constraints. These inequality constraints will determine the transformation strategy to be determined. Applied to pre-analysis feature sets Each data feature instance raw data blocks The results obtained are then fed into compliance constraints. The corresponding risk quantification function The calculated risk score must be less than or equal to the compliance constraint. The corresponding statutory risk threshold This allows for the definition of undetermined transformation strategies for subsequent steps. A rigid mathematical boundary that must be satisfied.
[0100] Step S2300: Solve the objective function using a constrained optimization algorithm. Maximizing the undetermined transformation strategy If the undetermined transformation strategy Belongs to the transformation policy space Then the objective function Positive rewards are applied, while negative rewards are applied otherwise. The process iteratively converges and outputs the optimal transformation policy. .
[0101] Specifically, this step aims to receive the objective function instantiated in step S2100. and the transformation strategy space instantiated in step S2200 By using a constrained optimization algorithm, compliance and utility are weighed in real time and dynamically, and an optimal solution is found and output that satisfies the objective function. Maximize and simultaneously satisfy the transformation policy space Optimal transformation strategy for constraints .
[0102] In practice, the core of this step lies in moving away from matching static rules and instead solving a dynamically constructed optimization problem in real time to find the optimal transformation strategy. The optimal transformation strategy It is obtained through a constraint maximization operation, and the solution must be found in the transformed policy space. Undetermined transformation strategy within This allows the objective function to be maximized and instantiated. .
[0103] To achieve this constraint maximization solution, this step employs a constraint optimization algorithm. Taking reinforcement learning (RL) as an example of a non-restrictive implementation of the constraint optimization algorithm, its dynamic trade-off process is as follows: This reinforcement learning algorithm transforms the constraint optimization problem instantiated in steps S2100 and S2200 into a policy learning task. Here, the policy is the undetermined transformation policy to be solved. Its action space corresponds to all alternative desensitization operations, such as different levels of blurring, feature extraction, noise superposition, etc.
[0104] The reward function of the reinforcement learning algorithm is defined as a combination of reward and penalty for the objective and constraints. Specifically, if the reinforcement learning algorithm selects a determinate transformation strategy... Belongs to the transformation policy space That is, if compliance constraints are met, then the objective function will be... As a positive reward. If the reinforcement learning algorithm selects a determinate transformation strategy... Not part of the transformation policy space In other words, failing to meet compliance constraints results in a significant negative reward, such as a preset minimum penalty value, which teaches the algorithm to avoid any behavior that violates rigid compliance. This reward and punishment mechanism reinforces the training objective of the learning algorithm—maximizing cumulative rewards—aligning it with the aforementioned objective of maximizing the solution to constraints.
[0105] Step S2400, using the pre-analyzed feature set Data location in For spatial indexing, the optimal transformation strategy will be used. Compile into the optimal de-identification execution plan The optimal de-identification execution plan Depend on The item contains specific actions Location information and motion parameters Atomic execution instructions constitute.
[0106] Specifically, this step aims to transform the optimal transformation strategy output in step S2300. This is compiled and visualized into an atomic sequence of operation instructions that can be executed immediately in subsequent steps, i.e., the optimal de-identification execution plan. This step utilizes the pre-analysis feature set output in step S1400. Data locations included This involves completing the crucial transformation from abstract decision-making to concrete execution, namely, from optimal transformation strategy. To the optimal desensitization execution plan The conversion.
[0107] In the specific implementation process, the optimal transformation strategy This is an abstract, mathematical solution, such as a policy network or high-level instructions. This step translates this abstract optimal transformation policy into a compilation function. Compile and visualize this into a sequence of operational instructions that can be executed immediately in subsequent steps, i.e., the optimal de-identification execution plan. The compilation function is a deterministic transformation function that receives an abstract optimal transformation strategy. As input, and utilizing the pre-analyzed feature set Data location in To obtain the optimal transformation strategy to execute. The specific data locations required will be used to output the optimal de-identification execution plan. .
[0108] The optimal de-identification execution plan It is a collection Atomic execution instructions The set. Among them, This indicates the optimal de-identification execution plan. The Atomic execution instructions, including the first All the execution information required for a given operation, i.e., the specific actions. Location information and motion parameters ; Indicates atomic execution instructions The index, whose value ranges from 1 to... A positive integer used to identify the currently executed atomic instruction. In the optimal desensitization execution plan The order in; Indicates atomic execution instructions The total number is the optimal desensitization execution plan. Atomic execution instructions included The sum; Indicates the first Atomic execution instructions The specific action is an opcode or identifier, and its value is a predefined string. Indicates the first Atomic execution instructions The location information is a specific data address or range. Indicates the first Atomic execution instructions The action parameters are determined by the optimal transformation strategy. Compiled from, used for specific actions Provide specific operational details.
[0109] Step S3000: Execute the optimal de-identification plan. Each atom executes instructions To generate new data feature sets Combined with non-sensitive data extracted through difference operations Reorganized into a new data payload And assemble compliance decision and audit labels. Inject new data payload Output the second cross-border data packet .
[0110] Specifically, this step aims to implement the optimal de-identification plan output by step S2400. Execution to the original cross-border data packet Data payload This step aims to achieve the ultimate goal of jointly optimizing compliance and effectiveness. Instead of using a fixed de-identification template, this step involves processing the original cross-border data packets... The data is reconstructed, and compliance decision tags are embedded into the reconstructed data packets, ultimately outputting a second cross-border data packet. This ensures that the VPN achieves intelligent compliance and effectiveness before it even enters the encryption process.
[0111] Further, step S3000 includes:
[0112] Step S3100: Traverse the optimal de-identification execution plan Each atom executes instructions Based on location information In data net load Locating the raw data block Call the instruction to execute the function Execution is determined by specific actions and motion parameters The defined transformation operations generate new data blocks, and all new data blocks are aggregated through a union operation to generate a new data feature set. .
[0113] Specifically, this step aims to receive the optimal desensitization execution plan for atoms, output from step S2400. and physically execute it onto the original cross-border data packet. Data payload Above, new data feature sets that meet multi-regional compliance requirements are generated physically. This enables the final decision-making process of dynamic trade-offs in step S2000, avoiding the loss of data utility caused by a one-size-fits-all approach to desensitization.
[0114] In the specific implementation process, this step receives the optimal de-identification execution plan. It then iterates through each atom and executes the instructions. Execute instructions for each atom. This step calls the instruction execution function. According to atomic execution instructions Location information included In the original cross-border data packet Data payload The upper precisely locates the raw data block identified in step S1400. and execute atomic execution instructions. Specific actions defined in and motion parameters Return to execute the atomic instruction. A new data block. Finally, all instructions are executed via a union operation. The returned new data blocks are merged to form a new data feature set. The specific process formula is as follows:
[0115] ;
[0116] in, The new data feature set is a function that contains all instruction execution functions. The generated set of new data blocks, whose values are determined by the optimal de-identification execution plan. middle Atomic execution instructions Perform a union operation on the execution structure Generated; Union operation; This represents the instruction execution function, whose purpose is to receive an atomic execution instruction. and data payload and execute instructions according to atomic principles. Specific actions included Location information and motion parameters In data net load The atomic execution instruction is executed on the above. Returns the transformed new data block.
[0117] Step S3200: Aggregate the optimal de-identification execution plan All atoms in the instruction execution Location information To construct sensitive data regions, the difference operation is used to extract data from the original data payload. Stripping away untouched non-sensitive data and call the payload reconstruction function to remove non-sensitive data. and new data feature sets Perform physical splicing to output new data payload. .
[0118] Specifically, this step aims to extract data from the original cross-border data packets. Data payload Extract untouched non-sensitive data Combined with the new data feature set generated in step S3100 By intelligently recombining the two, a new, transformed data payload that is both compliant and effective is output. This physically avoids data distortion caused by a one-size-fits-all approach to desensitization.
[0119] In the specific implementation process, this step starts from the original cross-border data packet. Data payload The process identifies all instances where the optimal de-identification execution plan was not implemented. Any location information contained therein Non-sensitive data touched This identification process is the inverse selection of the execution area in step S3100, and the specific formula is as follows:
[0120] ;
[0121] The logic of the above formula is that it uses a union operation. Optimal Desensitization Execution Plan All Atomic execution instructions Location information Merge them to form a complete sensitive data area; then in the data payload Perform difference operation on That is, from the net data load Subtract the sensitive data area from the output, and finally output the non-sensitive data. .in, This represents the difference operation; The union operation is used to combine all... Atomic execution instructions Location information The covered areas are merged into a single sensitive data area.
[0122] Based on non-sensitive data and new data feature sets The payload reconstruction function is called to physically merge the two data payloads and output the new data payload. .
[0123] Step S3300, for the optimal transformation strategy Applying hash functions to generate decision-making criteria Extract decision context information compliance objectives The final risk score and final utility score are then assembled into compliance decision and audit labels. And call the embedded function to inject the new data payload. Output the second cross-border data packet .
[0124] Specifically, this step aims to leverage the decision context information output in step S1000. The optimal transformation strategy output in step S2300 Generate a compliance decision and audit label By labeling this compliance decision and audit The new data payload is embedded into the output of step S3200. In the process, a second cross-border data packet with self-certification compliance capability is output. This makes it auditable before it enters the VPN encrypted channel.
[0125] In the specific implementation process, addressing the shortcomings of traditional VPNs that cannot detect the compliance of encrypted content and only provide transport layer encryption, this step involves writing verifiable audit evidence, namely compliance decisions and audit labels, before data enters the VPN encryption channel. This enables the upper-level system to trace the de-identification strategies, implementation rationale, and compliant areas of cross-border data.
[0126] The compliance decision and audit label It is a metadata structure, and its generation process is as follows:
[0127] First, the optimal transformation strategy Apply hash functions to generate unforgeable decision-making criteria. The decision is based on a fixed-length hash value, which is used for unforgeable decision tracing and auditing in subsequent steps.
[0128] At the same time, from the context of decision-making information Retrieving data stream context metadata Destination compliant areas compliance objectives The compliance target value is a region identifier used to identify the regulatory region of the compliance target that the data packet follows, such as a region code like "CN" or "DE".
[0129] Furthermore, a decision result is obtained from the solution result of step S2300. The decision result is the result of the constrained optimization algorithm in step S2300 solving for the optimal transformation strategy. During the process, to verify that it satisfies the transformation policy space Rigid constraints, i.e. compliance constraints The corresponding risk quantification function Less than or equal to compliance constraints The corresponding statutory risk threshold The calculated final risk score, and the final utility score calculated to determine that it is the optimal solution, serve as quantitative proof of compliance.
[0130] Based on the data generated and retrieved above, i.e., the decision-making basis Compliance Objectives The final risk score and the final utility score are combined to form compliance decision and audit labels. Then, the embedded function is called to add compliance decisions and audit labels. Embedded into the new data payload Output the second cross-border data packet .
[0131] Step S4000, the second cross-border data packet Encrypted third-party cross-border data packets encapsulated as VPN communication layers And based on decision context information Optimal transformation strategy and the second cross-border data packet Call the log generation function to generate decision audit logs. Deciphering and quantifying the comparison of compliance decisions and audit labels and target-side local compliance strategy If the comparison results are inconsistent, a policy update trigger signal will be generated. It is then transmitted back via the reverse channel, driving the source-end policy update in a closed loop.
[0132] Specifically, this step aims to pass the second cross-border data packet output in step S3300. Perform standard VPN encrypted transmission, while coordinating the decision context information output in step S1000. And the optimal transformation strategy output in step S2300 Generate verifiable decision audit logs. This addresses the shortcomings of VPNs, such as ignorance and lack of traceability, and utilizes generated third-party cross-border data packets. The tags in the code enable cross-domain policy validation, forming a dynamic adaptive closed loop.
[0133] Further, step S4000 includes:
[0134] Step S4100, the second cross-border data packet As the payload input to the VPN communication layer, it calls the VPN standard encryption function to perform protocol encapsulation and encryption, generating an encrypted third-party cross-border data packet. .
[0135] Specifically, this step aims to implement an encrypted transmission process with audit evidence, avoiding indiscriminate encryption of cross-border data and instead receiving the encapsulated compliance decisions and audit labels output in step S3300. Second cross-border data packet It then treats the entire data packet as an encrypted object, generating an encrypted third-party cross-border data packet. This makes it a carrier of compliance decisions and audit labels. A safe carrier.
[0136] In the specific implementation process, this step receives the new data payload. Compliance decision-making and audit labels Second cross-border data packet The second cross-border data packet This serves as the payload input to the VPN communication layer. The VPN communication layer can employ standard encryption and encapsulation operations using either Internet Protocol Security (IPsec) or Secure Sockets Layer / Transport Layer Security (SSL / TLS).
[0137] The encryption and encapsulation process is performed by a VPN standard encryption function, which receives the payload to be encrypted, i.e., the tagged second cross-border data packet. The application uses a selected encryption protocol to encrypt and encapsulate the data, ultimately generating an encrypted third-party cross-border data packet. The third cross-border data packet It is an opaque encrypted data packet, suitable for secure transmission in public network environments.
[0138] Step S4200, in the third cross-border data packet After sending, based on decision context information Optimal transformation strategy and the second cross-border data packet Call the log generation function to generate verifiable decision audit logs. .
[0139] Specifically, this step aims to achieve traceability and mathematical verifiability of compliance and utility decisions, addressing the auditing challenge of traditional VPNs' inability to explain cross-border anonymization activities under content invisibility conditions. Unlike traditional logs, this step is based on the decision context information output from step S1000. The optimal transformation strategy output in step S2300 And the compliance decision and audit label output by step S3300 Record and generate decision audit logs It continuously stores all core decision-making elements in a structured manner, supporting post-verification and cross-regional compliance traceability.
[0140] In the specific implementation process, the third cross-border data packet is transferred in step S4100. After sending, this step calls the log generation function in parallel at the source end to generate verifiable decision audit logs. .
[0141] The decision audit log It is a persistent decision audit log, whose values are serialized data structures organized according to a predetermined structure. It can be stored in an immutable ledger or secure database to provide regulatory agencies with a complete chain of evidence of cross-border data processing.
[0142] The log generation function is used to receive the decision context information corresponding to the current timestamp. Optimal transformation strategy and the second cross-border data packet The three data points are then serialized to generate a decision audit log. Among them, the decision context information Includes data flow context metadata Compliance constraint set Task utility objective function And a pre-analysis feature set, used for decision audit logs Provide complete problem input; the optimal transformation strategy Used to describe decision audit logs The provided execution decision, i.e., what decision to take; this second cross-border data packet Compliance decision and audit labels encapsulated in the middle Includes decision-making basis Physical credentials, used to transfer decision audit logs The third cross-border data packet transmitted in step S4100 Perform association verification.
[0143] Step S4300: Decrypt the third cross-border data packet. To extract compliance decision and audit tags Using policy validation functions to link compliance decisions and audit labels The source-side decision-making results and the target-side local compliance strategies carried within. A quantitative comparison is performed; if the comparison results are inconsistent, a policy update trigger signal is generated. And it is sent back via the reverse channel, triggering the source policy update.
[0144] Specifically, this step aims to construct a cross-domain adaptive closed loop to address the deficiency of static desensitization strategies in not being able to update in real time as the target changes. At the target end, the encrypted third-party cross-border data packets transmitted in step S4100 are verified. The compliance decision-making and audit labels carried in Verify whether the source decisions remain valid under the current jurisdiction. If compliance decisions and audit labels are valid... and target-side local compliance strategy If there is a mismatch, a policy update trigger signal will be generated. And then send the updated strategy back to the source via the reverse channel.
[0145] The target-side local compliance strategy It is a policy library that is maintained independently on the target side, configured and updated locally in real time by the target side's compliance management system, and represents the sovereign normative standards of that jurisdiction.
[0146] In the specific implementation process, the target end receives third-party cross-border data packets. The data packet is then decrypted using a standard VPN decryption function to recover the second cross-border data packet from step S3300. Subsequently, the second cross-border data packet Analyze the data to extract compliance decision and audit tags. .
[0147] Then, the policy verification function is called to verify the extracted compliance decisions and audit tags. Perform policy verification, comparing the source-side decisions it carries with the target-side local compliance policies configured locally in the compliance management system. Perform a matching verification.
[0148] The matching verification not only compares region identifiers but also performs quantitative verification by executing a policy verification function. Specifically, this policy verification function considers compliance decisions and audit labels. Extract the decision-making results from the source, such as the final risk score, and extract them from the target's local compliance strategy. Search for currently valid local compliance requirements, such as the target's local compliance policy. Thresholds corresponding to regulatory updates. Compliance decisions and audit labels are determined mathematically by comparing the aforementioned quantitative values. That is, source-end decision-making and target-end local compliance strategies. This refers to local requirements, specifically whether the structures are consistent. If the aforementioned quantization comparison structures are inconsistent, the matching verification operation outputs a strategy update trigger signal. It is sent back to the source end through the reverse channel of the communication link to drive the source end to update the strategy.
[0149] Example 2:
[0150] This embodiment, based on Embodiment 1, provides a cross-border VPN data anonymization and transmission system that is compatible with compliance in multiple regions, such as... Figure 2 As shown, it includes a decision context construction module, a joint policy generation module, a data reconstruction module, and an encryption and auditing module;
[0151] The decision context construction module is used to parse the original cross-border data packets. masthead To generate data stream context metadata Data flow context metadata The source of compliance area compliant areas of the destination and business type identifier For indexing, dynamically load compliance constraint sets. and task utility objective function Parallel processing of the original cross-border data packets Data payload Perform semantic scanning to generate a pre-analyzed feature set. Combined with the data flow context metadata Compliance constraint set Task utility objective function and pre-analysis feature set Constructing decision context information .
[0152] The joint strategy generation module is used to extract decision context information. Medium task utility objective function Pre-analysis feature set and compliance constraint set Combined with semantic tag set Instantiate the undetermined transformation strategy objective function and transformation strategy space The optimal transformation strategy is solved by executing a constrained optimization algorithm. and with pre-analysis feature set Data location in For spatial indexing, the optimal transformation strategy will be used. Compiled into instructions that can be executed atomically. The optimal de-identification execution plan .
[0153] The data reconstruction module is used to execute the optimal de-identification plan. Each atom executes instructions To generate new data feature sets Combined with non-sensitive data extracted through difference operations Reorganized into a new data payload And assemble compliance decision and audit labels. Inject new data payload Output the second cross-border data packet .
[0154] The encryption and auditing module is used to transmit the second cross-border data packet. Encrypted third-party cross-border data packets encapsulated as VPN communication layers And based on decision context information Optimal transformation strategy and the second cross-border data packet Call the log generation function to generate decision audit logs. Deciphering and quantifying the comparison of compliance decisions and audit labels and target-side local compliance strategy If the comparison results are inconsistent, a policy update trigger signal will be generated. It is then transmitted back via the reverse channel, driving the source-end policy update in a closed loop.
[0155] The parts of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of corresponding technical solutions in the prior art have not been described in detail to avoid excessive elaboration.
[0156] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A cross-border VPN data desensitization transmission method suitable for multi-region compliance, characterized in that, The method comprises the following steps: parsing the header of the original cross-border data packet to generate data stream context metadata, indexing the source compliance region, the destination compliance region and the business type identifier in the data stream context metadata, dynamically loading the compliance constraint set and the task utility target function, performing semantic scanning on the data payload of the original cross-border data packet in parallel to generate a pre-analysis feature set, and constructing decision context information based on the data stream context metadata, the compliance constraint set, the task utility target function and the pre-analysis feature set; extracting the task utility target function, the pre-analysis feature set and the compliance constraint set from the decision context information, combining the semantic label set, instantiating the target function and the transformation strategy space about the pending transformation strategy, executing a constraint optimization algorithm to solve the optimal transformation strategy, and compiling the optimal transformation strategy into an optimal desensitization execution plan composed of atomic execution instructions based on the data location in the pre-analysis feature set as a spatial index; executing each atomic execution instruction in the optimal desensitization execution plan to generate a new data feature set, recombining the non-sensitive data extracted by set difference operation into a new data payload, assembling compliance decisions and audit labels, and injecting the new data payload to output a second cross-border data packet; encapsulating the second cross-border data packet into an encrypted third cross-border data packet of the VPN communication layer, calling a log generation function to generate a decision audit log based on the decision context information, the optimal transformation strategy and the second cross-border data packet, decrypting and quantifying the compliance decisions and audit labels and the target local compliance strategy, and generating a policy update trigger signal and returning it via a reverse channel if the comparison result is inconsistent, thereby driving source policy update in a closed loop.
2. The method for adapting multi-region compliance cross-border VPN data desensitization transmission according to claim 1, characterized in that, The construction method of the decision context information comprises: extracting the quintuple information of the header in the original cross-border data packet, querying a pre-constructed joint mapping function based on the quintuple information, and outputting data stream context metadata; the quintuple information includes source IP address, target IP address, source port number, target port number and transmission protocol; the data stream context metadata includes data stream unique identifier, source compliance region, destination compliance region and business type identifier; According to the source and destination compliance regions in the data flow context metadata, the corresponding compliance constraint set is dynamically queried and loaded from the multi-region compliance policy database, and the compliance constraint set contains A compliance constraint composed of a data field type, a risk quantification function, and a legal risk threshold. dynamically loading a task utility target function from a pre-configured task utility configuration library according to the business type identifier in the data stream context metadata; the task utility target function quantifies business requirements into a weighted utility value sum based on utility features through an utility quantification function and a business importance weight; performing semantic scanning on the data payload of the original cross-border data packet to extract original data blocks, constructing a pre-analysis feature set using a pre-analysis model, and aggregating the data stream context metadata, the compliance constraint set and the task utility target function into decision context information.
3. The method of claim 2, wherein the method further comprises: The joint mapping function is a pre-constructed mapping mechanism based on a lookup table or a rule matching engine, and the specific mapping dimensions include geographical-legal mapping dimension and network-business mapping dimension; The geographic-legal mapping dimension is used to extract a source IP address or a target IP address in the five-tuple information, match in a pre-stored IP address segment index, and be mapped into a corresponding geographic-legal region identifier, which is used to index a legal jurisdiction range to which the IP address segment belongs. The network-service mapping dimension is used to extract a target port number or a transmission protocol in the five-tuple information, perform associated matching in a pre-stored service feature library, and be mapped into a corresponding service type identifier.
4. The method of claim 2, wherein the method further comprises: The dynamic loading step of the task utility target function comprises: Taking the service type identifier in the data flow context metadata as an index key, the corresponding task utility target function is searched and instantiated in the pre-configured task utility configuration library. The task utility objective function aims to maximize a weighted utility value sum of individual utility features. The calculation method of the weighted utility value sum comprises: traversing utility characteristics associated with the service type identifier , inputting a data variable corresponding to each utility characteristic as a formal parameter into a predefined utility quantification function, and calculating to obtain a utility score; multiplying the utility score by a service importance weight of the utility characteristic preconfigured to obtain a weighted utility value; and performing summation operation on the weighted utility values of all utility characteristics to generate the weighted utility value sum.
5. The method for adapting multi-region compliance cross-border VPN data desensitization transmission according to claim 1, characterized in that, The compiling method of the optimal desensitization execution plan composed of atomic execution instructions comprises: The task utility target function and the pre-analysis feature set are extracted from the decision context information, the business importance weight and the utility quantification function in the task utility target function and the original data block in the pre-analysis feature set are dynamically bound by using the semantic label set, and a target function with the pending transformation strategy as the only independent variable is constructed; The compliance constraint set and the pre-analysis feature set are extracted from the decision context information, the mapping relationship between the compliance constraint and the original data block is established by using the semantic label set, the risk quantification function and the legal risk threshold are instantiated as inequality constraints for the pending transformation strategy, and the transformation strategy space is defined; The constraint optimization algorithm is used to solve the pending transformation strategy that maximizes the target function, if the pending transformation strategy belongs to the transformation strategy space, the target function is used as a positive reward, otherwise a negative reward is applied, and the optimal transformation strategy is output after iterative convergence; Compiling the optimal transformation strategy into an optimal de-identification execution plan, the optimal de-identification execution plan being based on the spatial indexing of the data locations in the pre-analysis feature set and the optimal transformation strategy The atoms execution instructions comprise specific actions, location information and action parameters.
6. The method for adapting multi-region compliance cross-border VPN data desensitization transmission according to claim 5, characterized in that, The step of constructing the target function with the pending transformation strategy as the only independent variable comprises: The semantic label set carried by each data feature instance in the pre-analysis feature set is used as an associated index to perform a target function instantiation operation; The target function instantiation operation comprises: the pre-defined business importance weight and the utility quantification function in the task utility target function are bound to the original data block in the pre-analysis feature set, and a target function with the pending transformation strategy as the only independent variable is constructed; The calculation logic of the target function comprises: the outer layer traverses each data feature instance in the pre-analysis feature set, and the inner layer traverses all utility feature indexes contained in the semantic label set in each data feature instance; the utility quantification function is called according to the utility feature index, the original data block processed by the pending transformation strategy is input into the utility quantification function, and an utility score is obtained; the utility score is multiplied by the corresponding business importance weight to generate a weighted utility value, and the weighted utility values corresponding to all utility feature indexes of all data feature instances are accumulated to obtain an accumulated result as an output value of the target function.
7. The method of claim 5, wherein the method further comprises: The inequality constraint comprises: applying a pending transformation strategy to a raw data block of a data feature instance to generate a transformed data intermediate state; calling a risk quantification function corresponding to a compliance constraint to calculate a risk score with the data intermediate state as an input parameter; checking whether the risk score is less than or equal to a legal risk threshold corresponding to the compliance constraint; and when the pending transformation strategy satisfies the check for all data feature instances, determining that the pending transformation strategy satisfies the inequality constraint and is determined to fall into the transformation strategy space.
8. The method for adapting multi-region compliance cross-border VPN data desensitization transmission according to claim 1, characterized in that, The output method of the second cross-border data packet comprises: traversing each atomic execution instruction in the optimal desensitization execution plan, locating a raw data block in a data payload according to position information, calling an instruction execution function to execute a transformation operation defined by a specific action and action parameters to generate a new data block, and aggregating all new data blocks through a union operation to generate a new data feature set; aggregating position information of all atomic execution instructions in the optimal desensitization execution plan to construct a sensitive data region, stripping non-sensitive data that is not touched from the original data payload through a difference set operation, and calling a payload reconstruction function to physically splice the non-sensitive data and the new data feature set to output a new data payload; applying a hash function to the optimal transformation strategy to generate a decision basis, extracting a compliance target and final risk score and final utility score in the decision context information, assembling a compliance decision and an audit label, and calling an embedding function to inject the new data payload to output the second cross-border data packet.
9. The method for adapting multi-region compliance cross-border VPN data desensitization transmission according to claim 1, characterized in that, The driving method of the source end policy update comprises: inputting the second cross-border data packet as a payload of a VPN communication layer, calling a VPN standard encryption function to perform protocol encapsulation and encryption to generate an encrypted third cross-border data packet; after sending the third cross-border data packet, based on the decision context information, the optimal transformation strategy and the second cross-border data packet, calling a log generation function to generate a decision audit log; decrypting the third cross-border data packet to extract the compliance decision and the audit label, and using a policy checking function to quantitatively compare the decision result of the source end and the target end local compliance policy carried in the compliance decision and the audit label, and if the comparison result is inconsistent, generating a policy update trigger signal and returning through a reverse channel to trigger the source end policy update.
10. A cross-border VPN data de-identification transmission system adapted for multi-region compliance, for implementing the method of cross-border VPN data de-identification transmission adapted for multi-region compliance according to any one of claims 1-9, characterized in that, The system comprises a decision context construction module, a joint policy generation module, a data reconstruction module, and an encryption and audit module; The decision context construction module is configured to parse a header of an original cross-border data packet to generate data stream context metadata, use a source and destination compliance region, a destination compliance region and a business type identifier in the data stream context metadata as indexes to dynamically load a compliance constraint set and a task utility target function, perform semantic scanning on a data payload of the original cross-border data packet in parallel to generate a pre-analysis feature set, and combine the data stream context metadata, the compliance constraint set, the task utility target function and the pre-analysis feature set to construct decision context information. The joint strategy generation module is configured to extract a task utility target function, a pre-analysis feature set and a compliance constraint set in the decision context information, combine a semantic label set, instantiate the target function and a transformation strategy space about a pending transformation strategy, execute a constraint optimization algorithm to solve an optimal transformation strategy, and compile the optimal transformation strategy into an optimal desensitization execution plan composed of atomic execution instructions, with data positions in the pre-analysis feature set as spatial indexes. The data reconstruction module is configured to execute each atomic execution instruction in the optimal desensitization execution plan to generate a new data feature set, recombine non-sensitive data extracted through a difference set operation into a new data payload, assemble a compliance decision and an audit label, and inject the new data payload to output a second cross-border data packet. The encryption and audit module is configured to encapsulate the second cross-border data packet into an encrypted third cross-border data packet of a VPN communication layer, call a log generation function to generate a decision audit log based on the decision context information, the optimal transformation strategy and the second cross-border data packet, decrypt and quantify a comparison between the compliance decision and the audit label and a target end local compliance strategy, and if the comparison result is inconsistent, generate a strategy update trigger signal and return it via a reverse channel, to drive a source end strategy update in a closed loop.
Citation Information
Patent Citations
Data cross-border compliance management and control method and device, computer equipment and storage medium
CN114760149A
Data cross-border security guarantee method
CN120378135A