Cross-platform data control flow circulation method, device and medium based on data space
By generating and iteratively updating flow control information through a third-party data space platform, combined with intelligent processing trace detection, the problems of chaotic paths and difficulty in monitoring data processing traces in cross-platform data circulation have been solved. This has enabled full-link controllability and reliable traceability of data flow, and improved the security and compliance of data sharing.
Patent Information
- Application Number
- CN202511461366.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-10-14
AI Technical Summary
Existing technologies have problems with data security and compliance in cross-platform data circulation, especially in multi-level circulation scenarios where paths are chaotic and responsibilities are unclear, and traditional electronic signatures cannot identify data processing traces.
By generating and iteratively updating flow control information through a third-party data space platform, a chain-like flow path is formed. Combined with an intelligent processing trace detection mechanism, the entire chain of data packets can be controlled and reliably traced.
It achieves end-to-end controllability and reliable traceability of data flow, provides intelligent proactive detection capabilities for data processing traces, enhances the security and compliance of cross-platform data sharing, and reduces regulatory costs and the complexity of manual auditing.
Smart Images

Figure CN120956526B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data security and data circulation, and particularly relates to a cross-platform data control circulation method based on data space, equipment and medium. BACKGROUND
[0002] In today's digital economy, data as a core production factor, its security, compliance, controllable circulation is essential to release the value of data. Cross-organizational, cross-platform data sharing and exchange needs are growing, but at the same time, it also brings serious challenges: how to ensure the safety of data in the process of circulation, how to verify the identity of each party, how to trace the data flow, and how to prevent unauthorized tampering or processing of data.
[0003] Traditional data circulation solutions rely on centralized data trading platforms or simple point-to-point transmission. These solutions have inherent defects: first, the centralized platform itself becomes a single point of failure for data security and privacy, and it is difficult to achieve equal cooperation among multiple parties; second, point-to-point transmission lacks a global perspective and cannot effectively monitor and control the subsequent circulation of data.
[0004] Major manufacturers in the industry have recognized these problems and proposed various solutions. Patent document 1 (CN113992553B) proposes a data sharing method based on blockchain, which records data transactions through a distributed ledger, although it improves transparency, but the performance bottleneck and storage cost of blockchain limit its application in large-scale data circulation scenarios, and this method still cannot effectively control the use of data after leaving the blockchain network. Patent document 2 (CN114640453A) discloses a data security circulation scheme, which ensures the security of data processing through trusted execution environment (TEE) technology. This scheme focuses on protecting data privacy during computing, but cannot solve the problem of secondary dissemination control after authorized use of data, and the special requirements of TEE technology for hardware environment also limit its application range. Patent document 3 (CN114064997A) proposes a data use control method, which traces the source of data leakage through digital watermarking technology. This scheme has certain effect in post-tracing, but belongs to passive protection and cannot achieve active control of data circulation process, and digital watermarking technology may affect data quality, with the risk of being removed or destroyed.
[0005] In the prior art, electronic signature or digital signature technology is often used to verify the signature of a data packet to ensure the authenticity and integrity of the data source. However, in a multi-level circulation mode, a new electronic signature is added at each link, resulting in a linear increase in the number of signatures and forming a complex signature chain. This not only increases the verification calculation and storage overhead, but more importantly, the mere stacking of signatures makes the entire circulation path unclear and difficult to trace. Detection of data processing traces is another technical difficulty. Data providers usually require that the original data they provide remain unchanged after circulation, and any processing behavior may violate the data provision agreement. However, the existing electronic signature mechanism can only guarantee that the data has not been tampered with since the last signature, but cannot identify whether the data content is obtained from the original data through compliant or non-compliant processing.
[0006] Therefore, there is an urgent need in the art for an innovative technical solution that can implement full-process management and control of data circulation under a decentralized architecture, while having efficient traceability and intelligent processing trace detection functions, thereby ensuring the safety and compliance of the entire life cycle of data circulation. SUMMARY
[0007] To solve the above technical problems, the technical solution adopted by the present application is as follows:
[0008] According to the first aspect of the present application, a cross-platform data control circulation method based on data space is provided, comprising the following steps:
[0009] S100, receiving an original data packet from a data provider, the original data packet containing original data and a corresponding first electronic signature.
[0010] S200, verifying the validity of the first electronic signature before distributing the original data packet to a data user;
[0011] S300, generating and attaching circulation control information to the original data packet to form a controlled data packet.
[0012] S400, when the current data user performs data re-circulation to a target data user, the controlled data packet needs to be submitted to a third-party data space platform.
[0013] S500, after the third-party data space platform verifies the identity of the current data user and the integrity and authenticity of the historical circulation control information in the controlled data packet, it generates new circulation control information containing the current user identifier, the target user identifier and the platform electronic signature, iteratively updates the controlled data packet, and distributes the updated controlled data packet to the target user, thereby forming a traceable chain circulation path.
[0014] S600, for the controlled data packet in circulation, perform processing trace detection operation.
[0015] According to a second aspect of the present application, an electronic device is provided, comprising a processor and a memory; the processor is configured to execute the steps of the method according to the first aspect of the present application by invoking programs or instructions stored in the memory.
[0016] According to a third aspect of the present application, a computer readable storage medium is provided, which stores programs or instructions for causing a computer to execute the steps of the method according to the first aspect of the present application.
[0017] The data space-based cross-platform data control circulation method provided by the present application can produce the following significant technical effects compared to the prior art:
[0018] (1) Realize the full-link controllable and credible traceability of data circulation: by generating and iteratively updating circulation control information containing electronic signatures by a third-party data space platform, an unalterable chain circulation path is established for the data package. This technical effect directly solves the pain points of path confusion and unclear responsibility in multi-level circulation scenarios, enabling any party to quickly and accurately trace the exact source of the data and the entire division life cycle, providing a solid technical foundation for data auditing and compliance supervision.
[0019] (2) Provide intelligent data processing trace active detection capability: innovatively propose two complementary processing trace detection mechanisms (direct comparison and feature analysis). This not only enables accurate comparison when there is a source data copy, but also enables intelligent inference of whether data has been processed by analyzing the internal statistical characteristics (such as entropy value, data distribution) of the data when the source data cannot be obtained. This technical effect overcomes the limitations of traditional electronic signatures that can only prevent tampering but cannot identify processing, enabling effective supervision of whether the data user has violated the "no processing" agreement, greatly enhancing the confidence of data providers in sharing original data.
[0020] (3) Enhance the security and compliance of cross-platform data sharing: electronic signature identity verification is conducted throughout each circulation link, and combined with blockchain technology to record circulation records, a decentralized trust environment is built. This technical effect ensures that the identity of the participants is trusted, the data is complete and unaltered, and the operation records are not falsifiable, providing strong technical support for the safe, compliant, and efficient circulation of data elements, directly helping enterprises meet the compliance requirements of laws and regulations such as the Data Security Law.
[0021] (4) The level of automatic operation and maintenance of the data flow ecological environment is improved: the method changes the compliance audit with human participation into a platformized automatic detection process. Once the system finds a non-compliant processing behavior, it can automatically generate and send a warning message. This technical effect greatly reduces the regulatory cost and the complexity of manual audit of data flow, improves the efficiency, makes large-scale and high-frequency data flow possible, and promotes the prosperity of the data ecosystem.
[0022] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0024] Figure 1 The flow chart of the cross-platform data control flow method based on data space provided by the embodiments of the present application. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings of the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used in the specification of the present application are only for the purpose of describing the specific embodiments and are not intended to limit the present application. The term "and / or" used herein includes any and all combinations of one or more related listed items.
[0027] It is to be understood that some of the example embodiments are described in terms of a process or method depicted as a flowchart. Although each step in a flowchart can be identified with a reference number and can describe an operation, which can be implemented, directly or indirectly, in computer program instructions, it also be appreciated that such phrases are meant to connote a computer-related action that is implemented by programmed computers. The described implementations herein can provide a data space-based cross-platform data control circulation method to solve the problem of uncontrolled multi-level data circulation path and difficult monitoring of processing trace, and ensure the whole process of data circulation to be traceable, auditable and controllable.
[0028] For the purpose of describing the embodiments of the present application, the key terms are defined as follows:
[0029] Electronic signature: refers to a digital signature technology implementation based on public key infrastructure (PKI) technology, generated by the private key of the signatory, which can be used to verify the identity of the signatory and ensure the integrity of the data.
[0030] Data provider: refers to an entity that generates or owns data ownership and entrusts the platform to control data distribution.
[0031] First data user: refers to an entity that first obtains data reception rights from the data provider.
[0032] Subsequent data user / target user: refers to an entity that obtains data reception rights from the previous data user in the data circulation chain.
[0033] Unless otherwise specified, 'data user' described below refers to all data receiving entities.
[0034] As shown in Figure 1 The data space-based cross-platform data control circulation method provided by the embodiments of the present application can include the following steps:
[0035] S100, receiving an original data packet from a data provider, the original data packet containing original data and a corresponding first electronic signature;
[0036] S200, verifying the validity of the first electronic signature before distributing the original data packet to a data user;
[0037] S300, generating and attaching circulation control information to the original data packet to form a controlled data packet;
[0038] S400, when the current data user performs data re-circulation to the target data user, the controlled data packet must be submitted to the third-party data space platform;
[0039] S500, the third-party data space platform verifies the identity of the current data user and the integrity and authenticity of the historical flow control information in the controlled data package, generates new flow control information containing the current user identifier, the target user identifier and the platform electronic seal, iteratively updates the controlled data package, and distributes the updated controlled data package to the target user, thereby forming a traceable chain flow path.
[0040] S600, for the controlled data package in circulation, a processing trace detection operation is performed.
[0041] The method provided by the embodiment of the application realizes full-process traceability of data distribution by constructing a chain flow path, and solves the problem of out-of-control flow. By introducing an intelligent data feature analysis mechanism, the limitation of being able to verify integrity but not being able to identify processing is broken through, and active detection of data processing behavior is realized. Through platform-based automatic control, the compliance risk and supervision cost of data circulation are greatly reduced, and the safe and efficient circulation of data elements is promoted.
[0042] Further, the cross-platform data control circulation method based on a data space provided by the embodiment of the application is applied to a third-party data space platform, and the third-party data space platform is in communication connection with terminals of at least two data providers and two data users respectively.
[0043] In the embodiment of the application, the data provider refers to the owner or authorized party of the original data in the data circulation ecosystem, and is the starting point of the data value chain. Its core feature is to have data sovereignty. The role is a “seller” or “authorized party” of data. The data provider must provide original and unprocessed data, be responsible for classifying and grading data, and set data use strategies and compliance requirements (for example: “only for model training”, “not for resale”, “prohibited processing”). The goal of the data provider is to realize the realization of data value or business cooperation through data sharing on the premise of ensuring data security and compliance. For example, a bank (data provider) provides its desensitized user transaction records to a compliant financial technology company (user) for credit model analysis.
[0044] In the embodiments of the present application, the data user refers to a party that obtains data and consumes and applies the data in the data circulation ecology, and is the realization end of data value. The core feature is to enjoy the right to use data. The role is the "buyer" or "authorized party" of data. The data user must strictly use data in accordance with the agreement and data use strategy signed with the data provider. Before obtaining data, the identity authentication of the third-party platform is required to prove that it is a legal and authorized receiver. The goal of the data provider is to solve specific business problems, train AI models, conduct scientific research or generate new insights by obtaining external data. For example, a research institution (user) obtains medical image data from multiple hospitals (provider) to train an AI model for auxiliary diagnosis.
[0045] In the data space constructed by the present application, the data provider and the data user are not directly connected point to point, but interact through the third-party data space platform. As a trusted intermediary, the platform ensures that the control right of the provider and the use right of the user are realized under the premise of compliance, and restricts the behavior of the user through technical means (such as electronic signature, flow control, processing detection) to protect the interests of the provider.
[0046] In the embodiments of the present application, the third-party data space platform is the core executive subject and data circulation hub of the present application. It is not a simple data storage warehouse or transmission pipeline, but a standardized architecture-based, decentralized, and trusted data circulation infrastructure and governance environment. The "third-party" attribute emphasizes its core features of technical neutrality and business independence, that is, it does not produce data, does not consume data, and does not favor any participant. Its only role is to act as a "mediator" trusted by all participants, to execute pre-recognized rules through technical means to ensure the order of data circulation. The specific roles include:
[0047] Trusted neutral coordinator: As a third party trusted by all data providers and users, the third-party data space platform is responsible for executing rules such as identity authentication, authorization management, and flow strategy execution, and solving the trust problem between multiple parties.
[0048] Data circulation hub: all data outflow and inflow must pass through the platform. It does not centrally store business data, but centrally manages metadata, flow path information, and access strategies of data, and is the "traffic hub" that controls data flow.
[0049] Compliance technical executor: the platform converts legal terms in the data sharing agreement (such as "no processing" and "limited to specific use") into automatically executable technical rules (such as processing trace detection and access control), and ensures compliance through code rather than manpower.
[0050] To achieve the above-mentioned roles, the platform usually contains the following key functional modules or technical components:
[0051] Identity and access management: responsible for the identity registration, authentication and authorization of all participants. Usually based on the public key infrastructure (PKI) system, digital certificates are issued to each participant to ensure the authenticity of each interaction identity.
[0052] Policy enforcement point: the core rule engine of the platform. It parses and executes the data use policies set by the data provider (e.g. "number of forwarding times", "use period", "whether to allow processing"), controls whether the data can be sent out and whether it can be forwarded.
[0053] Data traceability and evidence module: responsible for generating, recording and maintaining the flow control information as described in the claims. This module forms an unalterable record of the key information of each data flow (who, when, to whom), and builds a complete chain path. To achieve unalterability, this module is often combined with blockchain technology, and the hash value of the record is stored on the chain.
[0054] Data processing trace detection engine: one of the core innovations of the invention. This engine has built-in detection algorithms as described in the claims, which can actively or on request analyze the data packets in circulation, determine whether they are compliant, and generate evidence of non-compliance.
[0055] Connector: standardized interface for the platform to connect with external data provider / use system. It is responsible for data standardization packaging, secure transmission, electronic signature application and verification, etc. to ensure that different platforms and different technology systems can seamlessly access the data space.
[0056] In the context of the invention, the term "terminal" does not refer to the commonly understood personal mobile phone or computer, but to a software system or client program that represents an organization for automated data exchange. The connection between the third-party data space platform and the terminal is not a simple web access, but a secure and programmable communication between machines. It can be connected through API-based secure communication connection and standardized interface and data format, i.e. the connection between the third-party platform and the terminal is a machine-to-machine communication based on standard API, through high-strength encrypted channel (HTTPS), and two-way digital certificate authentication.
[0057] Next, the steps are described in detail.
[0058] S100, receiving the original data packet from the data provider, the original data packet containing original data and a corresponding first electronic signature.
[0059] In the embodiments of the present application, step S100 is the starting point and the basis of trust for data to enter the controlled circulation space.
[0060] In the embodiments of the present application, the original data packet is not a simple data set, but a data object containing multiple information after structured processing. It is mainly composed of original data and an electronic seal.
[0061] Among them, the original data refers to the initial data directly generated or collected by the data provider without any derivative processing. Its format can be structured (such as database table, CSV file), semi-structured (such as JSON, XML log) or unstructured (such as picture, audio and video file). One of the core goals of the present application is to ensure that the original data or its traceable copy is always circulated after this link.
[0062] The first electronic seal is a digital signature attached to the original data, which is generated by the data provider using its unique private key. It is not a simple electronic seal, but a cryptographic certificate created based on asymmetric encryption algorithm (such as RSA, SM2) and hash algorithm (such as SHA-256, SM3). The core role of the first electronic seal is to verify the identity of the data provider and ensure the integrity of the original data, and its technical principle is as follows:
[0063] Identity verification: the first electronic seal is the result of encrypting the hash value of the original data using the data provider's private key. After receiving the data packet, the third-party data space platform can use the pre-stored or obtained from the authoritative certificate authority (CA) public key of the data provider to decrypt the seal. Successful decryption can prove that the data packet indeed comes from the claimed data provider, achieving identity authentication of the data source.
[0064] Integrity assurance: Before generating the seal, the data provider will first calculate a fixed-length hash value (digital fingerprint) of the original data. Any minor modification of the original data will cause a huge change in the calculated hash value. When verifying, the platform will use the same hash algorithm to recalculate the hash value of the received original data and compare it with the original hash value decrypted from the seal. If they are exactly the same, it proves that the data has not been tampered with during transmission; if they are not the same, it immediately determines that the data is invalid and refuses to receive, thereby ensuring the integrity of the data.
[0065] In the embodiment of the present application, the data provider transmits the generated original data packet to the third-party data space platform through a secure communication protocol (such as HTTPS, MQTT with TLS). After receiving, the security gateway of the platform will immediately trigger an automated verification process to verify the validity of the first electronic seal. This process is the premise of all subsequent circulation control. Only the data packet that passes the verification will be accepted by the platform and enter the next circulation distribution link. This eliminates the risk of fake identity and data destruction in the middle of transmission from the source, laying a solid foundation for the trust system of the entire data space.
[0066] S200, before distributing the original data packet to the data user, verifying the validity of the first electronic seal.
[0067] This step is the first key security check point of building the entire data space trust chain. Its purpose is to ensure the authenticity and integrity of the data from the source before data distribution, preventing fake data or tampered data from entering the circulation link. The verification process is automatically performed by the third-party data space platform, which includes the following levels of verification:
[0068] 1. Verification content and principle
[0069] The verification process is mainly based on the public key infrastructure (PKI) cryptography principle, including two core verification items:
[0070] (1) Identity authenticity verification:
[0071] Purpose: To confirm that the received original data packet indeed comes from the claimed data provider, not an impostor.
[0072] Technical process: The platform uses the public key publicly disclosed by the data provider to decrypt the first electronic seal attached in the original data packet. If it can be successfully decrypted, it proves that the seal is generated by the private key paired with the public key and only owned by the data provider, thereby verifying the identity of the data provider.
[0073] (2) Data integrity verification:
[0074] Purpose: To confirm that the original data has not been tampered with, added, deleted or destroyed by any third party since it was signed by the data provider.
[0075] Technical process: The platform uses the same hash algorithm (such as SHA-256) as the data provider to recalculate the original data received and generate a new hash value (also known as a digital fingerprint). At the same time, from the successfully decrypted signature, the original hash value generated by the data provider when signing is extracted. The platform compares the two hash values. If they are exactly the same, it proves that the data is intact; if there is even a byte of change in the data, the two hash values will be completely different, and the verification will fail.
[0076] 2. Verification process and logical judgment
[0077] The verification logic of the platform is a strict automated process:
[0078] (1) Trigger: Before receiving the data packet and preparing to distribute it to the target user, the platform automatically calls the signature verification service.
[0079] (2) Execution:
[0080] a. The platform queries the digital certificate of the data provider from the secure certificate library and checks whether the certificate itself is within the valid period and whether it is signed by a trusted certificate authority (CA) (certificate validity verification).
[0081] b. Use the public key in the certificate to try to decrypt the first electronic signature.
[0082] c. If decryption fails, the process is immediately terminated and determined as "signature verification failed" (identity not authentic).
[0083] d. If decryption is successful, the hash value comparison described above is performed.
[0084] e. If the hash values do not match, it is determined that "data integrity check failed".
[0085] (3) Result processing:
[0086] Verification success: The platform logs the successful verification and the data packet is allowed to proceed to the next step (S300, additional flow control information is added).
[0087] Verification failure: The platform immediately suspends any subsequent processing of the data packet and sends a security alert to the system administrator and the data provider, indicating "data source authentication failure" or "data may be tampered with". The data packet will be placed in the quarantine area for audit analysis and will never be distributed to the data user.
[0088] This step is not only a necessary technical check, but also carries important legal and compliance significance:
[0089] (1) Establishing initial trust anchor: It is the basis of the entire chain of trust traceability. All subsequent electronic signature verification in the circulation link depends on the confidence of the initial data packet.
[0090] (2) Key evidence of responsibility division: The record of successful verification shows that the platform has fulfilled the due diligence obligation. If there is a problem in the subsequent link, it can be proved that the problem does not come from the transmission process between the data provider and the platform.
[0091] (3) Compliance requirements: Meet the requirements of relevant laws and regulations such as the Electronic Signature Law on reliable electronic signature, which provides technical support for the legality of data circulation.
[0092] S300, generating and attaching flow control information to the original data packet to form a controlled data packet.
[0093] This step is the core innovation point of the present application for realizing controllable data circulation and full-link traceability. After verifying the validity of the original data packet (S200), the third-party data space platform is no longer a simple data forwarding, but as a trusted intermediary, it "packs" and "marks" the data packet, and marks it with platform-level control, thereby forming a controlled data packet.
[0094] In the embodiment of the present application, the flow control information is a set of standardized metadata generated by the platform, which at least contains the following key elements:
[0095] Data provider identification: uniquely identifies the original data source (for example, a unique identity ID or digital certificate hash is assigned).
[0096] Current receiver identification: uniquely identifies the designated data user in this circulation.
[0097] Third-party electronic signature: the certificate generated by the platform using its own private key to digitally sign the key information (such as provider ID, receiver ID, timestamp, etc.) of this circulation.
[0098] Timestamp: records the exact time when the platform generates the control information, which is used for subsequent audit and timeliness judgment.
[0099] Circulation serial number: records the sequence number of this circulation in the entire chain path (for example, the first circulation is #1, and the next time is #2).
[0100] The third-party data space platform first extracts key information from the current transaction (such as the provider ID from S100, the target user ID), and generates a unique timestamp and serial number. Then, the platform uses its own private key to digitally sign the combination of these information, generating a third-party electronic seal. Next, the platform serializes these information together with the generated seal according to the predetermined data format (such as JSON, XML), forming a structured control information block. Finally, the control information block is attached to the verified original data packet in one of the following two ways:
[0101] Method A (as metadata header): The control information block is attached as an additional data header (Header) in front of the original data packet. This is a lightweight and easy-to-parse way.
[0102] Method B (digital envelope mode): The original data packet and the control information block are packaged together and encrypted using the target user's public key to form a new secure data packet. This way is more secure, ensuring that only the target user can decrypt and view.
[0103] The final controlled data packet is equal to the original data packet plus the transfer control information.
[0104] S300 realizes the qualitative change from "data" to "controlled data", and its technical effect is:
[0105] Establish platform control rights: By attaching the platform's own electronic seal, the third-party platform formally intervenes in the supervision of data circulation. Any subsequent transfer must be returned to the platform for verification and unpacking, ensuring the platform's continuous control over data flow.
[0106] Create the "first ring" of the traceability chain: The control information generated this time becomes the starting point of the traceable chain in the entire data transfer life cycle. It is like a "birth certificate" that records the first authorized distribution of data.
[0107] Ensure the non-repudiation of the transfer event: Since the control information contains the platform's digital signature, it becomes a strong evidence at the level of cryptography, proving that "at a certain time point, the platform authorized the provision of A's data to B", which cannot be denied by any party.
[0108] Lay the foundation for subsequent iterative transfer: Each subsequent transfer (S400, S500) will be based on the current controlled data packet, with a new control information block attached. All these blocks are linked together through cryptography to form a complete and verifiable trust chain.
[0109] S300 step by generating and attaching standardized flow control information, a static raw data packet is converted into a dynamic, traceable, and continuously monitored data entity (controlled data packet). This is the primary and key technical means to achieve the "controllable flow" goal described in the present application.
[0110] In step S300, after the platform generates and attaches flow control information, optionally, the data processing trace detection engine can be called to perform initial feature analysis on the raw data packet, extract its baseline features (such as data structure schema, statistical features, information entropy, etc.), and store the baseline features associated with the data packet unique identifier in the pre-defined data processing rule library. This baseline feature will serve as the original reference for comparison with the features of the data packet in the flow when performing subsequent processing trace detection.
[0111] S400, when the current data usage direction target data usage party performs data re-flow, the controlled data packet must be submitted to the third-party data space platform.
[0112] In the embodiment of the present application, this step is the core mandatory rule to ensure that the data is continuously controlled in the multi-level flow process. Its core purpose is to completely eliminate private flow between data usage parties, ensure that every data distribution event is under the monitoring, authorization and recording of the platform, and thus build a complete and tamper-proof flow traceability chain.
[0113] The "must submit" is a combination of technical enforcement and protocol constraints, and its specific implementation and expansion includes the following aspects:
[0114] (1) Technical enforcement implementation mechanism
[0115] Unauthorized forwarding design: the controlled data packet obtained by the current data usage party (hereinafter referred to as "forwarding party") from the platform has a decryption key or access permission only for its own use. The platform does not grant any data usage party a forwarding token or key in design. Therefore, the forwarding party cannot technically generate a new controlled data packet that can be recognized by the target data usage party (hereinafter referred to as "recipient") and the platform.
[0116] Client integrated SDK / connector: in the data processing environment of the forwarding party, integrate the special client software (SDK) or connector provided by the platform. The client is pre-configured to: any attempt to send a controlled data packet outside will be intercepted and redirected to the target API interface of the platform. This step enforces behavior compliance from the code level.
[0117] Policy Enforcement Point (PEP) integration: The data access policy issued by the platform explicitly contains forwarding constraint clauses (such as <No_Redistribution>). The policy is enforced by the client, and any attempt to violate the policy will be prevented by the policy enforcement point.
[0118] (2) Submission process and data preparation
[0119] When the forwarding party decides to forward to the recipient, the following standardized process is triggered:
[0120] a. Request initialization: The forwarding party initiates a "data forwarding request" through its integrated platform client. The request contains at least the following metadata:
[0121] requestor_id: Unique identity of the forwarding party (such as digital certificate).
[0122] target_recipient_id: Unique identification registered by the platform.
[0123] data_package_id: Unique identification of the controlled data package to be forwarded.
[0124] (partially) purpose: Declaration of the purpose of this forwarding.
[0125] b. Data package preparation: The client encapsulates the entire controlled data package (including original data, all historical flow control information, and electronic signature) without any modification.
[0126] c. Secure transmission: The client submits the forwarding request and the complete controlled data package as payload to the platform's designated, highly available data flow API endpoint (such as POST https: / / <platform-domain> / api / v1 / data-redistribution).
[0127] The technical effect of S400 is:
[0128] (1) Ensure the full record of flow events: Forced back to the platform ensures that every authorized transfer from the data provider to the end user is recorded by the platform, avoiding the information black hole caused by private flow, laying the foundation for complete traceability.
[0129] (2) Dynamic policy check: The platform can check in real time whether the forwarding request violates the initial policy set by the data provider (for example, "maximum forwarding 2 times", "prohibit forwarding to specific types of entities") when receiving the forwarding request. Policy checking is dynamic, allowing providers to update policies at any time and take effect immediately.
[0130] (3) Maintain the continuation of the trust chain: Provide the necessary premise for the platform to generate new, verifiable flow control information (including platform signatures) in the next step (S500). Only with the authorization of the platform can the new flow relationship be accepted by other participants in the trust chain.
[0131] S400 step combines "technical enforcement" and "agreement constraints" to centralize all data reflow behaviors to the third-party data space platform for authorization and execution. This is not only a necessary condition for realizing chain path recording, but also a core hub for converting data providers' business rules and security policies into executable technical rules, fundamentally eliminating the risk of out-of-control data in the flow process, and ensuring the controllability, traceability and auditability of the whole data circulation process.
[0132] S500, after the third-party data space platform verifies the identity of the current data user and the integrity and authenticity of the historical flow control information in the controlled data package, it generates new flow control information containing the current user identifier, target user identifier and platform electronic signature, iteratively updates the controlled data package, and distributes the updated controlled data package to the target user, forming a chain flow path that can be traced back.
[0133] In the embodiment of the application, this step is the core processing link of building a trusted data flow traceability chain. When the platform receives the reflow request submitted by the current data user (forwarding party), it does not simply act as a transmission pipeline, but as a trusted intermediary to perform a rigorous "verification-generation-update-distribution" process, aiming to extend the existing trust chain and ensure that every data transfer event is recorded and cannot be denied.
[0134] The specific extension implementation includes the following sub-steps:
[0135] (1) Multi-factor authentication and audit
[0136] The platform first initiates an automated multi-factor verification process to comprehensively check the received request and data:
[0137] a. Identity authenticity verification: The platform uses the digital certificate registered by the forwarding party to verify the validity of the identity credentials (such as digital signatures) carried in its request. This ensures that the request indeed comes from a legitimate authorized user, preventing impersonation forwarding.
[0138] b. Historical flow chain integrity verification: This is the core of the verification. The platform parses all existing flow control information blocks in the controlled data package and performs step-by-step cryptographic verification:
[0139] The platform starts with the latest control information block, uses the platform's own public key to verify the validity of the electronic signature attached. If the verification is successful, it confirms that this flow record (e.g., from A to B) is authentic.
[0140] Then, extract the "previous block hash" or similar pointer contained in this information block to locate the previous control information block.
[0141] Repeat this verification process to verify each electronic signature on the link in turn until the first electronic signature generated by the data provider is traced back (which has been verified in S200). This process ensures that the entire historical path has not been tampered with since its inception, and any attempt to modify the historical record will result in the failure of all subsequent signature verification.
[0142] c. Policy compliance check: The platform parses the verified historical flow control information to check whether this forwarding violates the data provider's set policy. For example, check if the number of transfers has exceeded the "maximum number of transfers" limit, or if the target recipient is within the "allowed sharing" whitelist.
[0143] (2) Generate new flow control information
[0144] After all verifications are passed, the platform will create an authoritative record for this new flow event:
[0145] a. Information assembly: The platform generates a new structured flow control information block, which at least includes:
[0146] From: The unique identity of the forwarding party.
[0147] To: The unique identity of the target data user.
[0148] Timestamp: The exact timestamp of this flow occurrence.
[0149] PreviousBlockHash: A key field. This is the hash value of the latest control information block in the previous (i.e. the one held by the forwarding party) controlled data package. This field links the new and old blocks cryptographically, like a "chain".
[0150] b. Platform signature: The platform uses its own private key to digitally sign the entire new control information block, generating a platform electronic signature, and attaches it to the information block. This signature proves that "the platform authorized this transfer from B to C at a certain time".
[0151] (3) Iterative update of controlled data package
[0152] The platform attaches the newly generated, signed control information block to the original controlled data package. At this time, the structure of the data package becomes:
[0153] [original data] + [control information block 1 (A->B)] + [control information block 2 (B->C)] +...
[0154] Important note: This process is "append" rather than "replace", all historical control information is retained in its entirety, forming a growing, ordered data structure over time.
[0155] (4) Secure distribution to target user
[0156] The platform distributes the updated controlled data package to the target data user through a secure channel (such as mTLS). After receiving it, the target user can independently verify the validity of all platform electronic signatures on the entire link, thus confirming the complete transfer history of the data.
[0157] (5) Build a chain of transfer path
[0158] Through the above process, each successful transfer adds a new link certified by the platform authority to the data package. Each new link contains the cryptographic hash of the previous link, making the entire transfer path:
[0159] Traceable: Anyone can trace back from the latest link to the original data provider along the "PreviousBlockHash" pointer.
[0160] Verifiable: Each link has a platform digital signature that can be independently audited.
[0161] Tamper-proof: Any modification to the historical link will change its hash value, causing the "PreviousBlockHash" verification of all subsequent links to fail, making it easy to detect.
[0162] S500 step through the iteration mechanism of "verification-signature-link", like putting a notarization stamp that cannot be forged for each movement of data, and linking all these stamps in order, eventually forming a solid, traceable chain of trust. This not only realizes the full visual monitoring of the data life cycle, but also provides strong technical evidence for data compliance audit and dispute traceability, which is the technical cornerstone of the invention to realize the goal of controllable data circulation.
[0163] S600, for the controlled data package in circulation, performing processing trace detection operation.
[0164] This step is the core innovation of the invention to realize active compliance supervision. Its purpose is to break through the limitation of traditional electronic seal that can only verify whether the data has been tampered with, and further actively detect whether the data has been "processed in violation of regulations", so as to ensure that the data user strictly complies with the agreement terms signed with the provider.
[0165] In the embodiment of the invention, the processing trace detection operation can be triggered by one or more of the following ways: temporary detection request initiated by the data provider or the supervisor; automatically triggered by the third-party data space platform according to the preset strategy (such as periodic polling, before the next circulation of the data package occurs). Among them:
[0166] Periodic polling detection: the platform automatically checks the key data copies in circulation according to the predetermined plan (such as once a day).
[0167] Event-driven detection: triggered when a specific event occurs, such as when the data is about to be circulated for a new round (before S500) or when the data user applies the data for a new purpose.
[0168] Request-based detection: the data provider or supervisor can initiate a detection request for a certain data to the platform at any time.
[0169] Further, the execution of the processing trace detection operation includes:
[0170] Step one, compare the data in the current data package with the original data traced back to the data provider; and / or
[0171] Step two, based on the pre-defined data processing rule library, analyze the characteristics of the current data package to determine whether the current data package is derived from the original data;
[0172] Step three, if it is detected that there is a processing behavior that violates the agreement of the data provider, generate and send a breach prompt information to the data provider and / or supervisor, specifically, if there is a high suspicion of violating the agreement of the data provider based on the detection result, generate and send a breach prompt information to the data provider and / or supervisor for manual final decision.
[0173] Among them, step one is a direct comparison method, this method is suitable for scenarios that can accurately trace and obtain the original data copy, it is an accurate "digital fingerprint" comparison. Specifically, by analyzing the electronic seal sequence in the chain flow path, it is verified and located to the original data provider step by step, and the original data copy is obtained for comparison. The specific process can include:
[0174] Tracing and positioning: the platform analyzes all electronic seals in the chain flow path attached to the data packet (current data packet) to be detected. The platform starts from the latest seal, verifies its validity step by step, and finally locates the original data provider and its corresponding unique data identifier. This process benefits from the complete trust chain constructed by S300-S500.
[0175] Data acquisition: the platform retrieves the original data copy from the data provider's storage (or the platform's own cache) according to the located information.
[0176] Accurate comparison: the platform compares the original data copy and the data in the current data packet. This is not a simple byte comparison, but a more intelligent comparison:
[0177] Consistency comparison: calculate and compare the hash values of the two. If they are not consistent, it is immediately determined that the data has been modified.
[0178] Content difference analysis: if the protocol allows some processing (such as anonymization), but prohibits others (such as aggregation), the platform will perform more detailed content analysis, such as checking whether certain sensitive fields are retained, whether the numerical precision is changed, etc.
[0179] Technical effect: this method can provide irrefutable certainty evidence, directly proving whether the current data is completely consistent with the original data.
[0180] Further, step two is a feature analysis method, which is suitable for scenarios where original data cannot be directly obtained (such as due to privacy regulations, data has been deleted) or needs to be quickly and preliminarily screened. It is an intelligent inference based on probability and statistics. Specifically, it includes:
[0181] By analyzing the data structure, statistical characteristics, entropy value or whether there are traces of known data processing algorithms of the current data packet, the derivative relationship between the current data packet and the possible source data set is inferred.
[0182] The present application analyzes the multi-dimensional characteristics of the current data packet, and intelligently compares them with the source data benchmark characteristics or general data characteristics stored in the pre-defined data processing rule library, thereby inferring the derivative relationship between the current data packet and the possible source data set.
[0183] In the embodiment of the present application, the data processing rule library is a knowledge base that stores the characteristic fingerprints left by various data processing operations. The rules in the rule library can be defined by experts or generated through machine learning training.
[0184] Among them, data structure analysis is used to analyze the changes in the data mode level, which is the most direct and fastest detection method, including: comparing the field list of the current data packet with the source data field list recorded in the data processing rule library, and comparing the data type and precision of the fields of the current data packet.
[0185] Among them, if some source data fields (such as mobile phone number and detailed address are deleted) are missing in the current data packet, it is determined that data desensitization or field filtering may have been performed, and if new fields (such as age segmentation) are added, it is strongly suggested that data derivation calculation has been performed.
[0186] For example, a birth date field of DateTime type in the source data becomes an age field of Integer type in the current data packet, indicating that it has undergone derivation calculation. A longitude and latitude field of decimal(10,6) becomes decimal(6,3), indicating that it has undergone precision reduction processing, which is a common fuzzification processing.
[0187] Among them, statistical feature analysis deeply analyzes the changes in the data content itself through quantitative indicators, which is suitable for scenarios where the field structure does not change but the content changes, including:
[0188] The number of data records in the current data packet is counted. If the number of rows is significantly less than the number of source data rows recorded in the data processing rule library, it is possible that data sampling or data filtering has been performed. If the number of rows increases, it is very likely that data synthesis or connection operations have been performed.
[0189] The statistical values (mean, variance, median, extreme value) of numerical fields (such as income, transaction amount) are calculated and compared with the baseline values in the data processing rule library. Significant changes in the mean value may be due to filtering or bias; significant reduction in variance may indicate that the data has been smoothed or the range has been limited (such as removing extremely high or low values). The data processing rule library predefines difference thresholds for different data fields. When the change rate of the statistical characteristic value exceeds its corresponding threshold, it is determined to be significant.
[0190] The number of unique values of a categorical field (such as the city where it is located, the product type) is counted. A decrease in the number of unique values is a strong signal of data aggregation (such as aggregating cities into provinces) or filtering.
[0191] Information entropy is an index for measuring the degree of disorder or uncertainty of data, and is particularly suitable for detecting the loss of data diversity. The principle of information entropy analysis is that any processing operation that reduces data diversity (such as aggregation, sampling, and screening) will result in a decrease in information entropy. The specific steps of information entropy analysis include:
[0192] If the information entropy value of a certain classification field (such as occupation) in the current data packet decreases by more than a predetermined threshold (for example, 40%) compared to the field source data benchmark entropy value recorded in the data processing rule library, it is determined that the data may have undergone aggregation or sampling processing; for example, aggregating specific occupation types into "blue-collar" and "white-collar" categories will directly result in a decrease in the entropy value of the field.
[0193] Known data processing algorithm trace detection is used to detect statistical "fingerprints" left after using specific privacy computing or data synthesis algorithms, which can include differential privacy noise detection, synthetic data detection, and K-anonymity verification. Among them, differential privacy technology adds noise of a certain mathematical distribution (such as Laplace distribution, Gaussian distribution) to the data. The platform can analyze the noise in the data and determine whether it follows the above distribution through hypothesis testing (such as K-S test). If it matches, it can be determined that the data has been processed using differential privacy technology. Synthetic data detection is based on synthetic data generated by generative models (such as GANs), which often have micro-statistical characteristics (such as certain correlation anomalies, excessively smooth distribution) that do not exist in real data. A pre-trained synthetic data detection model can be integrated into the rule library to identify these characteristics and determine whether the data is synthetic derivative data. K-anonymity verification checks whether the data satisfies K-anonymity. That is, it checks whether all quasi-identifiers (such as zip code, gender, age) combinations appear at least K times. If the number of occurrences of all combinations is exactly an integer multiple of K, it is an obvious trace of K-anonymity.
[0194] Further, the pre-defined data processing rule library analyzes the characteristics of the current data packet to determine whether the current data packet is derived from original data, specifically including:
[0195] S601, one or more of data structure features, statistical features, information entropy features, and algorithm trace features are extracted from the current data packet.
[0196] The data structure feature is the difference in structure between the current data packet and the source data mode recorded in the rule library, including the increase or decrease of the number of fields, the change of the field name, the conversion of the field data type (for example, from string to integer), and the adjustment of the field precision (for example, the reduction of the decimal places of the floating point number). Extraction is automatically completed by parsing the metadata of the data packet. Specifically, the DDL statement of the database table, the Schema of the Parquet file, and the structure definition of the JSON can be parsed to obtain relevant information and compare it with the source data mode.
[0197] The statistical feature is a quantitative index calculated from specific numerical values in the data packet, including the total number of data records, the mean, variance, standard deviation, maximum value, minimum value, and median of numerical fields, the number of unique values of categorical fields, and the number of occurrences of the highest frequency value. Extraction can be performed by executing SQL query statements (such as COUNT(), AVG(), VARIANCE()) or using built-in functions of big data analysis frameworks (such as Spark, Pandas) to calculate the numerical values in the data packet to obtain these indicators.
[0198] The information entropy feature is the information entropy calculated for a categorical field, which is used to measure the uncertainty or diversity of the field. The higher the entropy value, the more chaotic and diverse the data, and the lower the entropy value, the more ordered and single the data. The calculation follows the formula H(X) = -∑(p(x i ) × log2(p(x i ))) where p(x i ) is the probability of occurrence of each unique value in the field x i , and i takes values from 1 to n, where n is the number of fields. The platform automatically calculates the entropy value of a specific field in the current data packet. Extraction automatically calculates the entropy value of a specific categorical field in the current data packet by calculating the probability of occurrence of each unique value in the field and substituting it into the information entropy formula.
[0199] The algorithm trace feature is a statistical "fingerprint" left by known data processing algorithms in the data, such as numerical distribution conforming to Laplace distribution or Gaussian distribution (reflecting the trace of adding noise for differential privacy), or having the characteristics of synthetic data. Extraction is performed by statistical hypothesis testing (such as using K-S test to verify distribution hypothesis) or calling a special pre-trained machine learning model for reasoning to detect whether these algorithm traces exist in the data.
[0200] S602, match the extracted features with the pre-defined rules in the data processing rule library, where each rule defines a mapping relationship between a specific data processing operation and a feature change.
[0201] Each rule is an "IF-THEN" logic statement, encapsulating domain knowledge. The data processing rule base is an extensible collection of multiple rules. For example:
[0202] Rule 1: IF
Variance of age field significantly reduced, e.g., reduced by more than 60%
Processing type: numerical range limited; Confidence contribution: 0.7
[0203] Rule 2: IF
Information entropy of city field significantly attenuated, e.g., reduced by more than 40%
Processing type: aggregation or sampling; Confidence contribution: 0.8
[0204] Rule 3: IF
Precision of latitude / longitude field reduced, e.g., from 6 decimal places to 2 decimal places
Processing type: precision reduction / fuzzification; Confidence contribution: 0.9
[0205] Rule 4: IF
Noise in data conforms to Laplace distribution through K-S test
Processing type: differential privacy processing; Confidence contribution: 0.95
[0206] The platform compares the extracted feature set with each rule in the rule base. If the current data state meets the conditions of a rule (IF part), the rule is triggered or matched successfully.
[0207] S603, according to the rule matching result, determine whether the current data packet is derived from the original data processing and its processing type.
[0208] S603 specifically includes:
[0209] S6031, according to the confidence contribution value of the triggered rule in the rule matching result, a final data processing confidence is calculated using a predefined synthesis algorithm.
[0210] The platform uses an algorithm to synthesize the confidence contribution values provided by all triggered rules into a final confidence. Various algorithms can be used:
[0211] Weighted average method: different weights are assigned to different types of rules, and then the weighted average is calculated. For example, algorithm trace rules have higher weights, and statistical feature rules have lower weights.
[0212] Probability model method: use Bayesian networks and other models to treat each rule as a piece of evidence, updating the posterior probability of the existence of processing behavior.
[0213] Maximum value method: among multiple rules, take the highest value of confidence contribution as the final value (suitable for scenarios where features are very obvious).
[0214] A data processing confidence between 0 and 1 (e.g. 0.85) is finally generated, indicating the likelihood that the current data packet is derived from the original data processing.
[0215] S6032, based on the data processing confidence and the triggered rules, a processing trace detection report is generated, including at least the possible data processing type, the confidence, and the judgment evidence.
[0216] The platform automatically generates a structured detection report as the final output. The report includes at least the possible data processing type, the final data processing confidence, and detailed judgment evidence; the judgment evidence includes the list of triggered rules and the corresponding feature change details, which can include:
[0217] Possible data processing type: list all processing operations (e.g. "sampling", "aggregation", "differential privacy processing") pointed by the triggered rules.
[0218] Final data processing confidence: a clear numerical value (e.g. 0.92).
[0219] Judgment evidence: detailed list of each triggered rule and its details, for example:
[0220] "Rule #2 triggered: detected 'city' field information entropy from 2.1 to 1.2, decay rate 42.8% (exceeding 40% threshold)".
[0221] "Rule #4 triggered: K-S test shows that the added noise conforms to Laplace distribution (p-value < 0.01)".
[0222] Conclusion and suggestion: according to the confidence threshold (e.g. greater than 0.8), give a preliminary conclusion (e.g. "highly suspected illegal processing") and suggest the data provider for manual review.
[0223] In the embodiments of the present application, step three is the key link to realize the post-supervision and responsibility implementation of data flow. Its core lies in establishing a human-machine collaborative judgment and response mechanism: the system performs preliminary detection and risk rating based on algorithms, and generates a structured evidence package; human experts make the final decision based on the evidence provided by the machine, so as to balance the supervision efficiency and the fairness of the decision.
[0224] The specific implementation includes the following sub-steps:
[0225] (1) Quantitative judgment of high confidence suspicion
[0226] Instead of simply giving a "yes / no" conclusion of violation, the system calculates a quantitative violation risk score (e.g. 0-100) based on the processing trace detection results. The score takes into account multiple factors:
[0227] Feature matching degree: the degree of matching between the current data features and the pre-defined processing feature patterns in the rule library.
[0228] Confidence synthesis: synthesizing an overall confidence by weighted average or probability model according to the confidence contribution values of multiple triggered detection rules.
[0229] Behavior abnormality degree: comparing with the historical behavior patterns of the data user, judging whether the operation deviates from the normal significantly.
[0230] Policy violation severity: the severity level of the suspected processing behavior violating the agreement terms (e.g. the severity of "light desensitization" is different from that of "complete reconstruction").
[0231] When the risk score exceeds the pre-set high-risk threshold (e.g. ≥80 points), the system determines that there is a high-confidence suspicion and triggers the subsequent process.
[0232] (2) Automatic generation of breach prompt information
[0233] The system automatically generates a structured breach suspicion report containing machine-readable and human-readable information, which at least includes:
[0234] a. Metadata information:
[0235] Suspicion event ID: unique tracking identifier.
[0236] Suspected breach party: the identity information of the current data user.
[0237] Associated data package ID and traceability path: the data package suspected to be processed in violation of the rules and its complete flow history.
[0238] Detection time and timestamp.
[0239] b. Technical evidence details:
[0240] Detection method: clearly indicate whether it is a direct comparison method or a feature analysis method.
[0241] Detailed findings: list the triggered rules one by one, the specific numerical values of feature changes (such as "field A entropy value decreased by 45%"), the calculated risk score and confidence.
[0242] Suspected processing type: the possible processing operation inferred by the system (such as "data aggregation", "synthetic data generation").
[0243] Data sample snapshot: provide anonymized sample fragments of suspected breach data and original data for comparison under strict compliance with privacy protection.
[0244] c. Disposal suggestions:
[0245] Based on a predefined policy library, preliminary handling suggestions are automatically generated (such as "It is recommended to suspend the data access permissions of this account and wait for verification").
[0246] (3) Multi-channel distribution and delivery guarantee
[0247] The generated default notification message is sent simultaneously through multiple reliable channels to ensure delivery:
[0248] API Push: Structured reports are pushed to the data provider's internal management system or regulatory platform in real time via a secure API callback interface.
[0249] Email / Message Notifications: Send alert emails or messages to designated contacts of data providers and regulators, including a report summary and a link to view the full report.
[0250] Blockchain-based evidence storage: The hash value of the report is stored on the blockchain to ensure that its generation time and content cannot be tampered with, providing legal evidence for possible subsequent arbitration.
[0251] All sending operations have status receipts and retry mechanisms to ensure reliable delivery of information.
[0252] (4) Manual final decision and feedback closed loop
[0253] Once the data provider or authorized personnel receive the notification, they can view the full suspect report on their management portal:
[0254] Manual review: Experts combine the evidence provided by the system, their own domain knowledge, and other contextual information to make a final ruling ("confirmed violation", "false alarm", "further investigation required").
[0255] Execution: Select the ruling result on the platform and issue an execution instruction (such as "Violation confirmed, permanently revoke the party's access rights"). The instruction will be executed automatically by the platform.
[0256] Feedback learning: The final ruling will serve as an important sample to feed back into the system's data processing rule base and detection model, which will be used to optimize detection rules, adjust confidence weights and risk thresholds, and enable the system to learn and continuously optimize itself.
[0257] Step three combines precise automated detection with rigorous human adjudication to build an efficient, reliable, and auditable closed loop for violation detection and handling. This not only significantly reduces monitoring costs for data providers and regulators and enables rapid response to violations, but also continuously improves the accuracy of supervision through a feedback learning mechanism, thereby effectively deterring potential violations and maintaining fairness and order in the data circulation market.
[0258] In the embodiments of the present application, "system" refers to the general term of a complete set of technical entities for implementing the aforementioned data control circulation method, which is a complete technical solution composed of software, hardware, network, protocol and rules, and is a highly automated and intelligent technical execution and control platform. The core is the software program running on the server, but it also includes the supporting hardware infrastructure, security protocols, data rule library and interactive interface.
[0259] In the embodiments of the present application, the circulation control information is stored in the form of a block chain, and the record of the chain circulation path is realized by a transaction hash chain on the block chain to achieve non-tamperable evidence and backtracking.
[0260] In the embodiments of the present application, after the third-party data space platform generates the circulation control information each time in steps S300 or S500, it does not only attach it in the data packet, but writes its key content (such as data provider identifier, receiver identifier, timestamp, data packet hash value, etc.) as a transaction into the block chain network. The specific implementation method includes:
[0261] Transaction generation: the platform packages the core metadata of this circulation and signs it using its private key to generate a legal block chain transaction.
[0262] Consensus on-chain: the transaction is broadcast to the nodes in the block chain network, verified and packaged into a new block through a network consensus mechanism (such as PBFT, Raft, etc.). Once the block is added to the chain, the record is permanently fixed and cannot be tampered with or deleted by a single institution.
[0263] Hash association: the "chain circulation path" is realized through a transaction hash chain. Specifically:
[0264] When the circulation control information is generated for the first time (S300), its block chain transaction hash value TxHash1 is calculated.
[0265] When the data is recalled and circulated again (S500), the newly generated circulation control information will explicitly contain the hash value TxHash1 of the last circulation transaction as its input or association field when writing to the block chain.
[0266] This process is iterated continuously, forming a chain of records on the block chain, which are mutually associated and linked, with the hash chain TxHash1->TxHash2->TxHash3->… This hash chain completely corresponds to the external data circulation path.
[0267] Through the above mechanism, the records of all flow events are stored in each node of the blockchain, and no longer rely on the data storage of any single centralized platform. Any attempt to tamper with or deny historical flow records requires tampering with all blocks after the record on the blockchain and controlling more than 51% of the node computing power, which is almost impossible to achieve in practice. Thus, it provides high-strength, publicly verifiable, tamper-proof evidence for the entire data flow process. Any participant or regulatory agency can independently verify and trace the on-chain records, clearly auditing the complete path of data from the origin to the final user, greatly enhancing the credibility and transparency of the entire system.
[0268] (Embodiment)
[0269] Suppose that data provider A company needs to share a set of original user behavior data with user B company, and agrees that B company cannot process the data.
[0270] First, A company uses its private key to generate a digital signature (first electronic seal) for the original data, and sends the data and seal to the third-party data space platform P.
[0271] After receiving it, platform P verifies the validity of the seal using A company's public key, confirming that the data is authentic and has not been tampered with. Then, platform P generates the first flow control information (e.g., From: A, To: B, Timestamp: T1) and signs it with the platform's private key, and attaches this information to the original data package, forming a controlled data package sent to B company.
[0272] Thereafter, if B company wants to share the data with partner C company, it cannot send it directly and must submit the entire controlled data package back to platform P. Platform P will check B company's identity and verify whether the entire flow history chain (including A and P's seals) submitted by B company is valid and complete.
[0273] After verification, platform P generates new flow control information (e.g., From: B, To: C, Timestamp: T2) and signs it, and attaches this new information to the data package, updates it, and then sends it to C company. At this point, the data package contains a complete and verifiable flow path from A -> B -> C.
[0274] To monitor whether B company has violated the agreement by processing the data, platform P can initiate processing trace detection on the data copy held by B company. Since direct access to B company's internal data may be difficult, platform P uses a feature analysis mode:
[0275] Platform P retrieves the baseline features of this data set provided by A company from the rule library, such as the baseline entropy value H0 for the "city" field.
[0276] The platform P requests to analyze the entropy value H1 of the "city" field in the data of company B.
[0277] The change rate of the entropy value is calculated. If it is found that H1 is much smaller than H0 (for example, the attenuation exceeds 50%), it indicates that the data of company B may only contain a partial city user sample and has been screened; or multiple cities have been combined into a large area and have been aggregated.
[0278] The platform P generates a default suspicion report accordingly and automatically sends it to the data provider A company and the regulatory party.
[0279] In the above manner, the present application realizes effective control and intelligent supervision of the whole process of data circulation.
[0280] The embodiment of the present application also provides an electronic device, comprising: at least one processor; and a memory in communication connection with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are arranged to execute the method described in the embodiment of the present application.
[0281] The embodiment of the present application also provides a computer readable storage medium, which stores computer executable instructions, and the computer executable instructions are used to execute the method described in the embodiment of the present application.
[0282] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, the steps described in the present application can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in the present application can be achieved, which are not limited herein.
[0283] The above specific embodiments do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A data space-based cross-platform data control flow circulation method, characterized in that, The method is applied to a third-party data space platform, and the method comprises the following steps: S100, receiving an original data package from a data provider, the original data package containing original data and a corresponding first electronic signature; S200, verifying the validity of the first electronic signature before distributing the original data package to a data user; S300, generating and attaching flow control information to the original data package to form a controlled data package; the flow control information at least contains a data provider identifier, a current receiver identifier, and a third-party electronic signature; S400, when a current data user performs data re-flow to a target data user, the controlled data package needs to be submitted to the third-party data space platform; S500, after the third-party data space platform verifies the identity of the current data user and the integrity and authenticity of the historical flow control information in the controlled data package, the third-party data space platform generates new flow control information containing a current user identifier, a target user identifier, and a platform electronic signature, iteratively updates the controlled data package, and distributes the updated controlled data package to the target user, so as to form a traceable chain flow path; the flow control information is stored in the form of a block chain, and the record of the chain flow path is realized by a transaction hash chain on the block chain to realize the non-tamperable evidence and traceability; S600, performing a processing trace detection operation on the controlled data package in circulation.
2. The method of claim 1, wherein, The first electronic signature is used to verify the identity of the data provider and ensure the integrity of the original data.
3. The method of claim 1, wherein, The processing trace detection operation comprises: comparing the data in the current data package with the original data traced back to the data provider; and / or analyzing the characteristics of the current data package based on a predefined data processing rule library to determine whether the current data package is derived from the original data by processing; if it is detected that there is a processing behavior that violates the agreement of the data provider protocol, generating and sending a breach prompt information to the data provider and / or a supervisory party.
4. The method of claim 3, wherein, The comparison of the data in the current data package with the original data traced back to the data provider specifically comprises: verifying and locating to the original data provider step by step by analyzing the sequence of electronic signatures in the chain flow path, and obtaining the original data copy for comparison.
5. The method of claim 3, wherein, The analysis of the characteristics of the current data package based on the predefined data processing rule library to determine whether the current data package is derived from the original data by processing specifically comprises: extracting one or more of data structure characteristics, statistical characteristics, information entropy characteristics, and algorithm trace characteristics from the current data package; matching the extracted characteristics with the predefined rules in the data processing rule library, wherein each rule defines a mapping relationship between a specific data processing operation and a characteristic change; determining whether the current data package is derived from the original data by processing and the processing type according to the rule matching result.
6. The method of claim 5, wherein, The determination of whether the current data package is derived from the original data by processing and the processing type according to the rule matching result specifically comprises: calculating a final data processing confidence value by using a predefined synthesis algorithm according to the confidence contribution value of the triggered rule in the rule matching result. Based on the data processing confidence and the triggered rule, a processing trace detection report is generated, the report including at least possible data processing type, confidence and judgment evidence.
7. The method of claim 2, wherein, The electronic signature is generated based on a public key infrastructure digital certificate, and each time of circulation verification includes checking validity of a certificate issuer and whether the certificate is in a valid period.
8. An electronic device, comprising: The system comprises a processor and a memory; The processor is configured to execute the steps of the method according to any one of claims 1 to 7 by calling programs or instructions stored in the memory.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium is configured to store programs or instructions, which enable the computer to execute the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
A microservice-based platform-based traffic generation system, method, computer, and storage medium
CN113992553B
Artificial intelligence power dispatching decision-making system based on big data
CN114064997A
Authentication and secret key negotiation method suitable for wireless sensor
CN114640453A
Electronic file full life cycle identification system, method, equipment and medium
CN117390695A
Block chain-based electric energy data management and sharing method and system
CN119848133A