Data flow conversion management method and system of multi-port data platform and storage medium
Patent Information
- Application Number
- CN202610687351.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-19
- Publication Date
- 2026-09-11
AI Technical Summary
[0003]现有技术中,普遍采用单点式管控的技术思路,针对数据流转的单一环节设置独立的加密方案,不同环节使用相互独立的加密系统,导致全链路的加解密策略不统一,权限管控分散,无法实现数据流转的全流程追溯;同时,现有技术采用静态线性加权的管控方式,无法捕捉多源数据聚合后的敏感属性跃迁效应,也无法量化风险沿流转路径的传导放大效应,跨场景流转时需要执行解密后重加密操作,存在严重的明文泄露风险;此外,现有技术无法平衡安全管控与业务效率的核心矛盾,要么为了安全牺牲业务效率,要么为了效率降低安全标准,且管控闭环缺失,执行结果无法反向优化管控体系,无法满足卫生健康等领域敏感政务数据的高合规、高安全管控要求
本申请以单条数据的数字孪生体为唯一核心载体,构建了非线性敏感演化、有向风险传导、多目标链路编排、时序闭环校验的管控体系,通过非线性演化算法精准捕捉多源数据聚合后的敏感属性跃迁效应,量化风险沿流转路径的放大效应,实现了安全管控与业务效率的动态平衡,通过国密密文透传协议实现了跨场景全链路无明文中间环节,彻底消除了解密重加密带来的泄露风险,同时通过时序滑动窗口校验算法实现了全流程的实时合规管控,结合数字孪生体实现了数据流转全生命周期的可追溯、可审计、可优化,大幅提升了多端口政务数据平台敏感数据流转的安全性、合规性与业务适配性,能够很好地满足敏感政务数据的高等级管控要求,具备极强的实用价值与推广前景。
Smart Images

Figure CN122741104A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data security, and in particular to data flow management methods, systems and storage media for multi-port data platforms. Background Technology
[0002] With the continuous improvement of the government data sharing system, especially the government data platform in the health sector, it needs to connect with hundreds of medical institutions, health commissions at all levels, and various business application systems. Data flow covers multiple links such as data collection, development, operation and maintenance, business use, and cross-institutional sharing. The operating entities, data usage needs, and security control requirements of different links vary greatly.
[0003] Existing technologies generally adopt a single-point control approach, setting up independent encryption schemes for a single link in the data flow. Different links use independent encryption systems, resulting in inconsistent encryption and decryption strategies across the entire chain, fragmented access control, and an inability to achieve full-process traceability of data flow. At the same time, existing technologies use a static linear weighted control method, which cannot capture the sensitive attribute transition effect after multi-source data aggregation, nor can it quantify the transmission and amplification effect of risks along the flow path. When flowing across scenarios, decryption and re-encryption operations are required, posing a serious risk of plaintext leakage. In addition, existing technologies cannot balance the core contradiction between security control and business efficiency. They either sacrifice business efficiency for security or lower security standards for efficiency, and the control loop is missing. The execution results cannot be used to optimize the control system, and they cannot meet the high compliance and high security control requirements of sensitive government data in fields such as health. Summary of the Invention
[0004] To address the aforementioned technical issues, this application provides a data flow management method, system, and storage medium for a multi-port data platform.
[0005] Firstly, this application provides a data flow management method for a multi-port data platform, employing the following technical solution: The data flow management method for a multi-port data platform includes the following steps: S1. In advance, configure the encryption and decryption policy library, national unified key system, permission rules for each operation role, and basic constraint rules of the directed acyclic graph of the data flow path in the data security sharing and management platform for each data flow scenario. S2. When the target data in the multi-port data platform is detected to trigger a transfer operation, a unique and tamper-proof digital identity identifier (DID) is generated for the target data. The digital twin with the time-series three-dimensional architecture is constructed using the digital identity identifier as an index. The target data is dynamically perceived in all dimensions through the nonlinear evolution algorithm of the dynamic sensitivity level of the twin. The constructed digital twin is then synchronized to the data security sharing and management platform. S3. Based on the digital twin of the target data, the full-link risk quantification is completed by the directed acyclic graph full-link risk transmission quantification algorithm, and then the encryption and decryption strategy matching and end-to-end ciphertext pass-through component link arrangement is completed by the multi-target adaptive strategy matching and link arrangement algorithm, and the arrangement result is updated to the digital twin of the target data. S4. Based on the updated digital twin, the orchestrated encryption and decryption component links are invoked to complete the end-to-end encrypted transmission execution. At the same time, the full-process compliance verification is completed through the time-series sliding window full-link consistency closed-loop verification algorithm, and the execution result is updated back to the digital twin to realize the self-iterative optimization of the control system. S5. Log the entire process of this transfer operation and synchronize the logs to the data security sharing and management platform and the digital twin of the target data.
[0006] Optionally, in step S2, the digital twin includes a static base layer, a dynamic evolution layer, and a risk transmission layer; The static base layer is used to store fixed attributes of the target data, which do not change over time; The dynamic evolution layer is used to store the real-time flow status of the target data and is dynamically updated over time. The risk transmission layer is used to store the directed acyclic graph of the target data flow path and the risk transmission coefficient, which is dynamically updated along the flow path.
[0007] Optionally, the formula for the nonlinear evolution algorithm of the twin dynamic sensitivity level in step S2 is: ; in, For target data with digital identity (DID) in time series The real-time sensitivity level, with a value range of [0, 10]; The initial sensitivity level of the target data is set, with a value range of [0, 10], and is pre-assigned according to the government data classification and grading specifications. is the Sigmoid non-linear activation function, with an output range of [0,1]. For time sequence The number of aggregation dimensions of the target data at that time, with a positive integer value; For time sequence The total number of scene nodes where the target data has been successfully transferred, and the value is a positive integer; The sensitivity level increment of the target data after the i-th scene node is processed, with a value range of [0,1]. For time sequence The historical cumulative risk value of the target data at that time, with a value range of [0,1]. , , The weighting coefficients for aggregation dimension, flow-sensitive increment, and historical risk value are respectively, satisfying the following conditions. .
[0008] Optionally, the directed acyclic graph end-to-end risk transmission quantification algorithm in step S3 includes: ; ; in, This is a risk matrix for all nodes in the entire link, with the following dimensions: , Let be the risk value of the i-th scene node, with a value range of [0, 10]. This is a node sensitivity level matrix with dimension 1. The element represents the real-time sensitivity level of the target data at the i-th node; The node-inherent risk weight diagonal matrix has dimensions of . diagonal elements Let be the inherent risk weight of the i-th scene node, with a value range of [0,1]. This is a risk propagation adjacency matrix with dimension 1. ,element Let be the risk transmission coefficient from the i-th node to the j-th node; This represents the overall risk value for the entire circulation path, with a value range of [0,10]. This represents the total number of scene nodes in the flow path. For node time-series weights, satisfying .
[0009] Optionally, the multi-objective adaptive policy matching and link orchestration algorithm in step S3 is a constrained multi-objective optimization algorithm, including: ; The constraints are: ; in, These are the decision variables, namely the encryption / decryption strategies and component link combinations to be matched; For the plan The security objective function has a value range of [0,1]. For the plan The efficiency objective function has a range of [0,1]. Let be the objective function for link reliability of scheme s, with a value range of [0,1]. This represents the maximum allowed processing time for the business. For the plan The computation time for the i-th node.
[0010] Optionally, the formula for the timing sliding window end-to-end consistency closed-loop verification algorithm in step S4 is: ; ; The execution rules are as follows: when If the verification passes, proceed to the next node. when If this occurs, execution will be immediately interrupted, the flow of target data will be frozen, and a security alert will be triggered. in, For time sequence The average execution deviation within the sliding window, with a value range of [0,1]. The width of the sliding window, which takes the value of a positive integer; This represents the expected execution result for the k-th node. The actual execution result of the kth node, after normalization, takes a value in the range [0,1]. For time sequence The deviation threshold at that time, with a value range of [0,1]; The baseline deviation threshold.
[0011] Optionally, the flow scenarios include data development scenarios, asynchronous encryption scenarios for data collection pre-library, data operation and maintenance scenarios, data management business usage scenarios, and cross-organizational data return scenarios; the encryption and decryption component links include one or more of the following: UDF encryption and decryption functions, AOE-Client encryption client, data operation and maintenance gateway integrating UDF encryption and decryption functions, AOE-Plugin application encryption plugin, and AOE-Gateway encryption proxy gateway, and adjacent components communicate through a ciphertext pass-through protocol compatible with national cryptographic algorithms.
[0012] Optionally, after step S5, the method further includes: tracing and auditing the entire data flow process based on the full-link operation logs and execution deviation data stored in the digital twin; and simultaneously inputting the execution deviation data back into the nonlinear evolution algorithm for the dynamic sensitivity level of the digital twin in step S2 to iteratively update the algorithm weight parameters.
[0013] Secondly, this application provides a data flow management system for a multi-port data platform, which adopts the following technical solution: A data flow management system for a multi-port data platform includes a data security sharing and control platform, and multiple data source ports, multiple business usage ports, and multiple management ports that are communicatively connected to the data security sharing and control platform. The data security sharing and control platform includes: The strategy configuration module is used to pre-configure the encryption and decryption strategy library, the national unified key system, the permission rules for each operation role, and the basic constraint rules of the directed acyclic graph of the data flow path for each scenario. The twin construction module is used to generate a unique and tamper-proof digital identity for the target data that triggers the flow, construct a digital twin with a time-series three-dimensional architecture, and complete the full-dimensional dynamic perception of the target data through a nonlinear evolution algorithm; The strategy orchestration module is used to complete the full-link risk transmission quantification based on digital twins, and to complete the link orchestration of encryption and decryption strategy matching and end-to-end ciphertext pass-through components through multi-objective optimization algorithms; The encryption / decryption execution module is used to call the encryption / decryption component link to complete end-to-end encrypted pass-through execution based on the orchestration results in the digital twin, and at the same time complete the full-process closed-loop verification through the time-series sliding window algorithm. The log management module is used to record logs of the entire process of the workflow and store them synchronously in the data security sharing and control platform and the digital twin of the target data.
[0014] Thirdly, this application provides a storage medium, which adopts the following technical solution: The storage medium stores the program of the data flow management method for the multi-port data platform described in any one of the above.
[0015] In summary, this application includes at least the following beneficial technical effects: This application uses a digital twin of a single data entry as the sole core carrier to construct a control system encompassing nonlinear sensitive evolution, directed risk transmission, multi-objective link orchestration, and time-series closed-loop verification. Through a nonlinear evolution algorithm, it accurately captures the sensitive attribute transition effects after multi-source data aggregation, quantifies the amplification effect of risks along the flow path, and achieves a dynamic balance between security control and business efficiency. By employing a national cryptographic encryption protocol, it achieves a cross-scenario, end-to-end plaintext-free intermediate link, completely eliminating the leakage risk caused by decryption and re-encryption. Simultaneously, a time-series sliding window verification algorithm enables real-time compliance control throughout the entire process. Combined with the digital twin, it achieves traceability, auditability, and optimizability throughout the entire data flow lifecycle, significantly improving the security, compliance, and business adaptability of sensitive data flow on multi-port government data platforms. It effectively meets the high-level control requirements of sensitive government data and possesses strong practical value and promising prospects for widespread application. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the data flow management method in this application. Detailed Implementation
[0017] The embodiments of this application are described in detail below, and examples of the embodiments are shown in the accompanying drawings.
[0018] In the description of this specification, the references to "certain embodiments," "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples" refer to specific features, structures, materials, or characteristics described in connection with the described embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0019] The multi-port data platform in this embodiment is specifically a health and wellness government affairs data platform that connects to multiple medical institutions, health and wellness management departments of various districts, and various business application systems. It includes multiple data source ports (business systems of various medical institutions), multiple business usage ports (one-screen viewing application, health profiling application, data quality control and analysis system, etc.), and multiple management ports (management systems of health and wellness commissions of various districts). All ports are connected to a unified data security sharing and management platform through the government extranet. Data flow covers the entire process of collection, development, operation and maintenance, business use, and cross-institutional return. All steps in this embodiment are implemented based on this platform.
[0020] This application discloses a data flow management method for a multi-port data platform. Figure 1 This includes the following steps: S1. In advance, configure the encryption and decryption policy library, national unified key system, permission rules for each operation role, and basic constraint rules of the directed acyclic graph of the data flow path in the data security sharing and management platform for each data flow scenario. S2. When the target data in the multi-port data platform is detected to trigger a transfer operation, a unique and tamper-proof digital identity identifier (DID) is generated for the target data. The digital twin with the time-series three-dimensional architecture is constructed using the digital identity identifier as an index. The target data is dynamically perceived in all dimensions through the nonlinear evolution algorithm of the dynamic sensitivity level of the twin. The constructed digital twin is then synchronized to the data security sharing and management platform. S3. Based on the digital twin of the target data, the full-link risk quantification is completed by the directed acyclic graph full-link risk transmission quantification algorithm, and then the encryption and decryption strategy matching and end-to-end ciphertext pass-through component link arrangement is completed by the multi-target adaptive strategy matching and link arrangement algorithm, and the arrangement result is updated to the digital twin of the target data. S4. Based on the updated digital twin, the orchestrated encryption and decryption component links are invoked to complete the end-to-end encrypted transmission execution. At the same time, the full-process compliance verification is completed through the time-series sliding window full-link consistency closed-loop verification algorithm, and the execution result is updated back to the digital twin to realize the self-iterative optimization of the control system. S5. Log the entire process of this transfer operation and synchronize the logs to the data security sharing and management platform and the digital twin of the target data.
[0021] Step S1 involves completing the basic configuration work in the data security sharing and management platform beforehand, specifically including: Configure encryption / decryption strategy library: For five core flow scenarios, namely data development scenario, asynchronous encryption scenario of data collection front-end library, data operation and maintenance scenario, data management business use scenario, and cross-organization data return scenario, configure corresponding encryption / decryption strategies, including encryption / decryption field rules, national cryptographic SM4 / SM2 encryption algorithm, de-identification rules, key binding rules, etc., to form a standardized encryption / decryption strategy library; Configure a national unified key system: Establish a symmetric and asymmetric key management system that is one-to-one bound to encryption and decryption strategies. The entire process of key generation, distribution, updating, and destruction is uniformly managed through a data security sharing and control platform to ensure the consistency and compliance of keys throughout the entire chain. Configure operation role permission rules: For the five core operation roles of developers, operations and maintenance personnel, business personnel, medical personnel and security management personnel, configure corresponding operation permissions, encryption and decryption permissions and data access scope respectively, so as to achieve one-to-one binding of permissions and roles; Configure the basic constraint rules for the directed acyclic graph of the data transfer path: clarify the legal path for the transfer of government data, prohibit loop transfers and cross-authority transfers, and ensure that the transfer path complies with the compliance requirements of government data.
[0022] In step S2, when a target data in the multi-port data platform triggers a transfer operation, a unique and immutable digital identity identifier (DID) is generated for the target data. Using this DID as an index, a time-series three-dimensional digital twin is constructed. A nonlinear evolution algorithm for the dynamic sensitivity level of the digital twin is then used to achieve full-dimensional dynamic perception of the target data. Specifically, this includes: S21. Construct a digital twin with a temporal three-dimensional architecture, the three-layer architecture being as follows: Static base layer: Used to store fixed attributes of target data, including initial sensitivity level, data source information, data classification and grading labels, and unique data identifier. The data in this layer does not change over time and is only written during twin initialization. Dynamic Evolution Layer: Used to store the real-time flow status of target data, including real-time sensitivity level, number of aggregation dimensions, flowed scenario nodes, historical operation records, and cumulative risk value. The data in this layer is dynamically updated over time t. Risk transmission layer: A directed acyclic graph used to store the flow path of target data. The data at this layer includes risk transmission coefficients between nodes, expected flow paths, and potential risk nodes. This data is dynamically updated as the flow paths change. S22. Collect the static basic information of the target data, write it into the static basic layer of the twin, complete the initialization of the twin, and assign a globally unique digital identity identifier (DID) to the target data as the unique index of the twin. Target data without a corresponding DID and twin will be directly intercepted and prohibited from circulation. S23. Calculate the real-time sensitivity level of the target data at time t using the twin dynamic sensitivity level nonlinear evolution algorithm, and write it into the twin's dynamic evolution layer. The algorithm formula is as follows: ; in, For target data with digital identity (DID) in time series The real-time sensitivity level, with a value range of [0, 10], where the higher the value, the higher the sensitivity level; The initial sensitivity level of the target data is set, with a value range of [0,10]. It can be pre-assigned according to the "Government Data Classification and Grading Specification" and the requirements for health data management. For example, the initial sensitivity level of patient identity information and diagnosis and treatment information is set to 10, and the initial sensitivity level of publicly available statistical data is set to 2. It is a sigmoid nonlinear activation function with an output range of [0,1]. Its core function is to simulate the sensitive attribute transition effect after multi-source data aggregation. The output range is [0,1]. When the input z exceeds the threshold, the output quickly approaches 1, realizing the nonlinear transition of the sensitivity level and solving the core defect that the linear weighted formula cannot capture sensitive mutations. For time sequence The number of aggregation dimensions of the target data, i.e., how many different sources and categories of sub-data were aggregated to generate this data. The value is a positive integer, for example, after aggregating ID card numbers, medical record information, and medical insurance information... =3; For time sequence The total number of scene nodes where the target data has been successfully transferred, and the value is a positive integer; This is the sensitivity level increment of the target data after it flows through the i-th scene node. The value range is [0,1], which is determined by the operation type and data processing method of the node. For example, the increment for data processing operation is assigned a value of 0.3, and the increment for read operation only is assigned a value of 0. For time sequence The historical cumulative risk value of the target data is in the range of [0,1]. It is calculated by normalizing the number of historical abnormal operations and the number of unauthorized operations. When there are no abnormal operations, the value is 0. , , The weighting coefficients for aggregation dimension, flow-sensitive increment, and historical risk value are respectively, satisfying the following conditions. For health-sensitive data, the recommended configuration is as follows: , , Prioritize the identification of sensitive transitions in aggregated data; S24. Based on the historical flow records of the target data, construct a directed acyclic graph of the flow path. Where V is the set of scene nodes and E is the set of flow edges between nodes, calculate the risk transmission coefficient matrix between nodes and write it into the risk transmission layer of the twin; S25. Synchronize the completed digital twin to the data security sharing and management platform in real time, which will serve as the sole legal input for subsequent steps S3 and S4, ensuring that all operations throughout the process revolve around the digital twin and achieving unified management.
[0023] Step S3 is implemented based on the target data digital twin output from S2, completing end-to-end risk quantification, encryption / decryption strategy matching, and component link orchestration, specifically including: S31. Extracting the directed acyclic graph of flow paths from twins. Using the real-time sensitivity level matrix, and through a directed acyclic graph full-link risk transmission quantification algorithm, the full-link node risk matrix and comprehensive risk value are calculated. The algorithm formula is as follows: ; ; in, This is a risk matrix for all nodes in the entire link, with the following dimensions: , Let be the risk value of the i-th scene node, with a value range of [0, 10]. The node sensitivity level matrix is the sole input source for the directed acyclic graph (DAG) end-to-end risk transmission quantification algorithm in step S3. It corresponds to each scenario node in the DAG of the flow path and is a structured representation of the real-time sensitivity level of the target data at each node, with dimensions of [missing information]. The element represents the real-time sensitivity level of the target data at the i-th node. ; The node-inherent risk weight diagonal matrix has dimensions of . diagonal elements This represents the inherent risk weight of the i-th scenario node, with a value range of [0,1]. It is pre-assigned based on the openness of the scenario, such as a cross-institutional return scenario. =0.9, internal operation and maintenance scenario =0.4; This is a risk propagation adjacency matrix with dimension 1. ,element Let be the risk transmission coefficient from node i to node j. The value is determined by the following rule: if there is a directed edge from i to j in the flow path, then... ,in The basic transmission coefficient, with a value range of (1, 1.2], simulates the amplification effect of risk along the flow path; the longer the path, the stronger the amplification effect. If there is no directed edge, then... =0, ensuring that risks are only transmitted forward along the flow path, which is in line with the business logic of data flow; This represents the overall risk value for the entire circulation path, with a value range of [0,10]. This represents the total number of scene nodes in the flow path. For node time-series weights, satisfying The closer to the final destination node, The larger the value, the closer the simulation is to the data outflow stage, and the higher the priority of risk control in business characteristics.
[0024] S32. Based on the calculated node risk matrix and comprehensive risk value This algorithm, employing multi-objective adaptive strategy matching and link orchestration algorithms, matches the optimal encryption / decryption strategy for each scenario node in the flow path, along with a unified national cryptographic key system bound to the entire link. This constrained multi-objective optimization algorithm solves for the Pareto optimal solution, as shown in the formula: ; The constraints are: ; in, These are the decision variables, namely the encryption / decryption strategies and component link combinations to be matched; For the plan The security objective function has a value range of [0,1] and is determined by the strength of the encryption and decryption algorithm, the key management level, and the granularity of access control. The higher the value, the stronger the security. For the plan The efficiency objective function, with a value range of [0,1], is determined by the total time consumed in the end-to-end encryption and decryption calculations; the higher the value, the higher the execution efficiency. Let be the link reliability objective function of scheme s, with a value range of [0,1]. It is determined by the probability of fault-free operation of the component link and the probability of ciphertext pass-through consistency. The higher the value, the stronger the reliability. Constraint 1: The security level must cover the comprehensive risk value of the entire link. The higher the risk, the higher the minimum security requirement, so as to achieve adaptive risk matching. Constraint 2: The maximum allowed processing time for the business. The computation time of scheme s at the i-th node is given, ensuring that the total computation time does not exceed the business limit; Constraint 3: The link reliability shall not be lower than 99%, meeting the availability requirements of the government data platform; S33. Based on the matching end-to-end encryption and decryption strategy, orchestrate end-to-end encryption and decryption component links for the expected flow path of the target data. The component links include UDF encryption and decryption functions, AOE-Client encryption client, data operation and maintenance gateway integrating UDF encryption and decryption functions, AOE-Plugin application encryption plugin, and AOE-Gateway encryption proxy gateway. Adjacent components communicate with each other through a ciphertext pass-through protocol compatible with the national cryptographic SM4 algorithm, ensuring that cross-scenario flow does not require intermediate decryption links. S34. Update the entire completed end-to-end encryption / decryption strategy, component links, unified key system, and expected execution matrix into the digital twin of the target data, which will serve as the sole execution and verification benchmark for step S4.
[0025] Step S4 is implemented based on the target data digital twin updated in step S3, completing end-to-end encrypted pass-through execution, full-process closed-loop verification, and iterative updates of the digital twin. Specifically, it includes: S41. Extract the component links, full-link encryption and decryption strategies, unified key system and expected execution matrix arranged in step S3 from the digital twin of the target data. Pre-allocate temporary keys and operation permissions to the encryption and decryption components of each scenario node. Components without twin authorization cannot call the key, ensuring the compliant use of the full-link key. S42. Trigger the execution of the encryption / decryption component chain. After the target data completes initial encryption at the first scenario node, it is directly processed or transferred in all subsequent scenario nodes through the national cryptographic ciphertext pass-through protocol of the adjacent components. Only at the final target transfer port, decryption or desensitization display at the corresponding granularity is completed according to the permissions of the operating subject. There is no plaintext intermediate link throughout the entire process. The specific execution method is as follows for the 5 core transfer scenarios: Data development scenario: Match UDF encryption and decryption functions, extract the encrypted data from the returned library back to the MRS-Hive data lake cluster to create a test encrypted library, create a permanent UDF encryption and decryption function in the data governance platform DGC, and directly perform calculations on the encrypted data through the UDF encryption and decryption function without decryption; Asynchronous encryption scenario for data collection front-end library: Match the AOE-Client encryption client, deploy the client on the data source front-end library server, the client obtains the encryption policy and key from the management platform, reads plaintext data, completes encryption, and writes it back to the database, supporting periodic incremental encryption; Data operation and maintenance scenario: Matching the data operation and maintenance gateway that integrates UDF encryption and decryption functions, when operation and maintenance personnel trigger operations through the database client, the gateway calls the UDF encryption and decryption functions to encrypt and decrypt the data, and verifies the operation and maintenance personnel's permissions throughout the process; Data management business use case: Match the AOE-Plugin application encryption plugin, integrate the plugin into the business application server, and when users access data, the plugin decrypts and displays the data or performs anonymization processing according to the policy, without modifying the core code of the business application; Cross-organizational data return scenario: Matching the AOE-Gateway encrypted proxy gateway, the data extracted by ETL is encrypted by the gateway and stored in the return pre-return library. The receiver completes local decryption through an integrated plugin, realizing end-to-end encrypted data transfer; S43, throughout the entire process of component link execution, the consistency of the execution result of each node with the expected execution matrix in the twin is verified in real time through a time-series sliding window full-link consistency closed-loop verification algorithm. The algorithm formula is: ; ; The execution rules are as follows: when If the verification passes, proceed to the next node. when If the event occurs, execution will be immediately interrupted, the flow of target data will be frozen, a security alarm will be triggered, and the abnormal event will be updated to the digital twin in real time. in, For time sequence The average execution deviation within the sliding window, ranging from [0,1]. The higher the value, the greater the deviation between the execution and the expectation. The value is the width of the sliding window, which can be a positive integer. A value of 3 is recommended, which means that the execution consistency between the current node and the previous two nodes is checked each time, covering the encrypted pass-through link of the adjacent nodes. The expected execution result of the k-th node is derived from the expected execution matrix arranged in step S3 of the twin, which includes the ciphertext hash value, the execution status of the encryption / decryption strategy, and the permission verification result. The actual execution result of the k-th node, and The dimensions are completely consistent, and the normalized values range from [0,1]. For time sequence The deviation threshold at that time, with a value range of [0,1]; The recommended value for the basic deviation threshold is 0.05, meaning that under normal circumstances, a deviation exceeding 5% is considered abnormal.
[0026] The overall risk value of the entire link is output by the algorithm in step S3. The higher the risk, the stricter the deviation threshold, so as to realize the adaptive verification and control of risk. S44. After the component link is completed, the entire link execution record, sensitivity level change, verification deviation value, and transfer result are updated to the digital twin of the target data. At the same time, the verification deviation value is input in reverse into the nonlinear evolution algorithm of the twin dynamic sensitivity level in step S2 to iteratively update the weight parameters of the algorithm, complete the closed-loop control of this transfer, and provide optimized perception basis for the subsequent transfer of the data, so as to realize the self-iterative optimization of the control system.
[0027] In the specific implementation of step S5, the entire process of this transfer operation is logged. The log content includes the operation subject, operation time, transfer scenario, encryption and decryption strategy used, data range of operation, execution result, verification deviation value, abnormal events, etc. The log is synchronously stored in the data security sharing and management platform and the digital twin of the target data for subsequent compliance audit and traceability.
[0028] This application also discloses a data flow management system for a multi-port data platform, including a data security sharing and control platform, and multiple data source ports, multiple business use ports, and multiple management ports connected to the data security sharing and control platform via the government extranet. The data source ports are the business system ports of various medical institutions, the business use ports are the ports of business systems such as the One-Screen View application and the Health Profile application, and the management ports are the management system ports of the health commissions of various districts. The data security sharing and control platform is deployed on the Shanghai Municipal Government Cloud and includes a policy configuration module, a twin construction module, a policy orchestration module, an encryption / decryption execution module, and a log management module. The modules communicate with each other to collaboratively complete the full-process control of data flow.
[0029] This embodiment also provides a storage medium, which may be a read-only memory, a random access memory, a disk, an optical disk, or other storage media. The storage medium stores a computer program, and when the computer program is executed by the processor, it implements the data flow management method of the multi-port data platform described in this embodiment.
[0030] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A data flow management method for a multi-port data platform, characterized in that, Includes the following steps: S1. In advance, configure the encryption and decryption policy library, national unified key system, permission rules for each operation role, and basic constraint rules of the directed acyclic graph of the data flow path in the data security sharing and management platform for each data flow scenario. S2. When the target data in the multi-port data platform is detected to trigger a transfer operation, a unique and tamper-proof digital identity identifier (DID) is generated for the target data. The digital twin with the time-series three-dimensional architecture is constructed using the digital identity identifier as an index. The target data is dynamically perceived in all dimensions through the nonlinear evolution algorithm of the dynamic sensitivity level of the twin. The constructed digital twin is then synchronized to the data security sharing and management platform. S3. Based on the digital twin of the target data, the full-link risk quantification is completed by the directed acyclic graph full-link risk transmission quantification algorithm, and then the encryption and decryption strategy matching and end-to-end ciphertext pass-through component link arrangement is completed by the multi-target adaptive strategy matching and link arrangement algorithm, and the arrangement result is updated to the digital twin of the target data. S4. Based on the updated digital twin, the orchestrated encryption and decryption component links are invoked to complete the end-to-end encrypted transmission execution. At the same time, the full-process compliance verification is completed through the time-series sliding window full-link consistency closed-loop verification algorithm, and the execution result is updated back to the digital twin to realize the self-iterative optimization of the control system. S5. Log the entire process of this transfer operation and synchronize the logs to the data security sharing and management platform and the digital twin of the target data.
2. The data flow management method for a multi-port data platform according to claim 1, characterized in that, In step S2, the digital twin includes a static base layer, a dynamic evolution layer, and a risk transmission layer; The static base layer is used to store fixed attributes of the target data, which do not change over time; The dynamic evolution layer is used to store the real-time flow status of the target data and is dynamically updated over time. The risk transmission layer is used to store the directed acyclic graph of the target data flow path and the risk transmission coefficient, which is dynamically updated along the flow path.
3. The data flow management method for a multi-port data platform according to claim 1, characterized in that, The formula for the nonlinear evolution algorithm of the twin dynamic sensitivity level in step S2 is as follows: ; in, For target data with digital identity (DID) in time series The real-time sensitivity level, with a value range of [0, 10]; The initial sensitivity level of the target data is set, with a value range of [0, 10], and is pre-assigned according to the government data classification and grading specifications. is the Sigmoid non-linear activation function, with an output range of [0,1]. For time sequence The number of aggregation dimensions of the target data at that time, with a positive integer value; For time sequence The total number of scene nodes where the target data has been successfully transferred, and the value is a positive integer; The sensitivity level increment of the target data after the i-th scene node is processed, with a value range of [0,1]. For time sequence The historical cumulative risk value of the target data at that time, with a value range of [0,1]. , , The weighting coefficients for aggregation dimension, flow-sensitive increment, and historical risk value are respectively, satisfying the following conditions. .
4. The data flow management method for a multi-port data platform according to claim 1, characterized in that, The directed acyclic graph end-to-end risk transmission quantification algorithm in step S3 includes: ; ; in, This is a risk matrix for all nodes in the entire link, with the following dimensions: , Let be the risk value of the i-th scene node, with a value range of [0, 10]. This is a node sensitivity level matrix with dimension 1. The element represents the real-time sensitivity level of the target data at the i-th node; The node-inherent risk weight diagonal matrix has dimensions of . diagonal elements Let be the inherent risk weight of the i-th scene node, with a value range of [0,1]. This is a risk propagation adjacency matrix with dimension 1. ,element Let be the risk transmission coefficient from the i-th node to the j-th node; This represents the overall risk value for the entire circulation path, with a value range of [0,10]. This represents the total number of scene nodes in the flow path. For node time-series weights, satisfying .
5. The data flow management method for a multi-port data platform according to claim 4, characterized in that, The multi-objective adaptive policy matching and link orchestration algorithm in step S3 is a constrained multi-objective optimization algorithm, including: ; The constraints are: ; in, These are the decision variables, namely the encryption / decryption strategies and component link combinations to be matched; For the plan The security objective function has a value range of [0,1]. For the plan The efficiency objective function has a range of [0,1]. Let be the objective function for link reliability of scheme s, with a value range of [0,1]. This represents the maximum allowed processing time for the business. For the plan The computation time for the i-th node.
6. The data flow management method for a multi-port data platform according to claim 4, characterized in that, The formula for the time-series sliding window end-to-end consistency closed-loop verification algorithm in step S4 is as follows: ; ; The execution rules are as follows: when If the verification passes, proceed to the next node. when If this occurs, execution will be immediately interrupted, the flow of target data will be frozen, and a security alert will be triggered. in, For time sequence The average execution deviation within the sliding window, with a value range of [0,1]. The width of the sliding window, which takes the value of a positive integer; This represents the expected execution result for the k-th node. The actual execution result of the kth node, after normalization, takes a value in the range [0,1]. For time sequence The deviation threshold at that time, with a value range of [0,1]; The baseline deviation threshold.
7. The data flow management method for a multi-port data platform according to claim 1, characterized in that, The data transfer scenarios include data development scenarios, asynchronous encryption scenarios for data collection front-end libraries, data operation and maintenance scenarios, data management business usage scenarios, and cross-organizational data return scenarios. The encryption and decryption component links include one or more of the following: UDF encryption and decryption functions, AOE-Client encryption client, data operation and maintenance gateway integrating UDF encryption and decryption functions, AOE-Plugin application encryption plugin, and AOE-Gateway encryption proxy gateway. Adjacent components communicate with each other through a ciphertext pass-through protocol compatible with national cryptographic algorithms.
8. The data flow management method for a multi-port data platform according to claim 1, characterized in that, The process following step S5 further includes: tracing and auditing the entire data flow based on the full-link operation logs and execution deviation data stored in the digital twin; simultaneously, inputting the execution deviation data back into the nonlinear evolution algorithm for the dynamic sensitivity level of the digital twin in step S2 to iteratively update the algorithm weight parameters.
9. A data flow management system for a multi-port data platform, characterized in that, The data flow management method of the multi-port data platform as described in any one of claims 1-8 includes a data security sharing and management platform, and multiple data source ports, multiple service usage ports, and multiple management ports that are communicatively connected to the data security sharing and management platform. The data security sharing and management platform includes: The strategy configuration module is used to pre-configure the encryption and decryption strategy library, the national unified key system, the permission rules for each operation role, and the basic constraint rules of the directed acyclic graph of the data flow path for each scenario. The twin construction module is used to generate a unique and tamper-proof digital identity for the target data that triggers the flow, construct a digital twin with a time-series three-dimensional architecture, and complete the full-dimensional dynamic perception of the target data through a nonlinear evolution algorithm; The strategy orchestration module is used to complete the full-link risk transmission quantification based on digital twins, and to complete the link orchestration of encryption and decryption strategy matching and end-to-end ciphertext pass-through components through multi-objective optimization algorithms; The encryption / decryption execution module is used to call the encryption / decryption component link to complete end-to-end encrypted pass-through execution based on the orchestration results in the digital twin, and at the same time complete the full-process closed-loop verification through the time-series sliding window algorithm. The log management module is used to record logs of the entire process of the workflow and store them synchronously in the data security sharing and control platform and the digital twin of the target data.
10. A storage medium, characterized in that, The system contains a computer program that, when executed by a processor, implements the data flow management method of the multi-port data platform according to any one of claims 1-8.