Cross-platform data control circulation method and device based on data space and medium
By generating and iteratively updating flow control information through a third-party data space platform, combined with intelligent detection, the problems of chaotic paths and difficulty in monitoring processing traces in cross-platform data circulation have been solved. This has enabled full-link controllability and reliable traceability of data flow, and improved the security and compliance of data sharing.
Patent Information
- Application Number
- CN202511461366.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-10-14
AI Technical Summary
Existing technologies have problems with data security and compliance in cross-platform data circulation, especially in multi-level circulation scenarios where paths are chaotic and responsibilities are unclear, and traditional electronic signatures cannot identify data processing traces.
By generating and iteratively updating circulation control information containing electronic signatures through a third-party data space platform, an unalterable chain circulation path is formed. Combined with an intelligent processing trace detection mechanism, the entire chain of data packets can be controlled and reliably traced.
It achieves end-to-end controllability and reliable traceability of data flow, provides intelligent proactive detection capabilities for data processing traces, enhances the security and compliance of cross-platform data sharing, and reduces regulatory costs and the complexity of manual auditing.
Smart Images

Figure CN120956526A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of data security and data circulation, and in particular to a cross-platform data control and circulation method, device and medium based on data space. Background Technology
[0002] In today's digital economy, data, as a core production factor, is crucial for unlocking its value through secure, compliant, and controllable circulation. The demand for cross-organizational and cross-platform data sharing and exchange is growing, but this also brings serious challenges: how to ensure data security during circulation, how to verify the identities of all parties, how to trace data flow, and how to prevent unauthorized alteration or processing of data.
[0003] Traditional data circulation solutions often rely on centralized data trading platforms or simple peer-to-peer transmission. These solutions have inherent flaws: First, the centralized platform itself becomes a single point of failure for data security and privacy, and it is difficult to achieve peer-to-peer collaboration among multiple parties; second, peer-to-peer transmission lacks a global perspective and cannot effectively monitor and control the subsequent flow of data.
[0004] Major manufacturers in the industry have recognized these problems and proposed various solutions. Patent document 1 (CN113992553B) proposes a blockchain-based data sharing method that records data transactions through a distributed ledger. While this improves transparency, the performance bottlenecks and storage costs of blockchain limit its application in large-scale data circulation scenarios. Furthermore, this method cannot effectively control the use of data after it leaves the blockchain network. Patent document 2 (CN114640453A) discloses a data security transfer scheme that uses Trusted Execution Environment (TEE) technology to ensure the security of the data processing process. This scheme focuses on protecting data privacy during computation but cannot solve the problem of controlling the secondary dissemination of data after authorized use. The specific hardware requirements of TEE technology also limit its application scope. Patent document 3 (CN114064997A) proposes a data usage control method that uses digital watermarking technology to trace the source of data leakage. This scheme has some effect in post-event traceability but is a passive protection method that cannot achieve proactive control over the data flow process. Moreover, digital watermarking technology may affect data quality and is at risk of being removed or destroyed.
[0005] In existing technologies, electronic signatures or digital signatures are commonly used to verify data packets, ensuring the authenticity and integrity of the data source. However, in multi-level circulation models, each link adds a new electronic signature, leading to a linear increase in the number of signatures and forming a complex signature chain. This not only increases the computational and storage overhead of verification, but more importantly, the simple stacking of signatures makes the entire circulation path obscure and difficult to trace. Detecting data processing traces is another technical challenge. Data providers typically require that the original data they provide remain unchanged after circulation; any processing may violate the data provision agreement. However, existing electronic signature mechanisms can only guarantee that the data has not been tampered with since the last signature, but cannot identify whether the data content was obtained by compliant or non-compliant processing of the original data.
[0006] Therefore, there is an urgent need in this field for an innovative technical solution that can achieve full-process control of data flow under a decentralized architecture, while possessing efficient traceability capabilities and intelligent processing trace detection functions, thereby ensuring the security and compliance of data circulation throughout its entire lifecycle. Summary of the Invention
[0007] To address the aforementioned technical problems, the technical solution adopted by this invention is as follows: According to a first aspect of the present invention, a cross-platform data control and circulation method based on data space is provided, comprising the following steps: S100, receive a raw data packet from the data provider, the raw data packet containing raw data and the corresponding first electronic signature.
[0008] S200, before distributing the original data packet to the data user, verify the validity of the first electronic signature; S300, generate and attach flow control information to the original data packet to form a controlled data packet.
[0009] S400, when the current data user is transferring data to the target data user, the controlled data packet must be submitted to a third-party data space platform.
[0010] S500: After the third-party data space platform verifies the identity of the current data user and the integrity and authenticity of the historical flow control information in the controlled data packet, it generates new flow control information containing the current user identifier, the target user identifier, and the platform's electronic signature, iteratively updates the controlled data packet, and then distributes the updated controlled data packet to the target user, thereby forming a traceable chain flow path.
[0011] S600 performs processing trace detection operations on controlled data packets in circulation.
[0012] According to a second aspect of the present invention, an electronic device is provided, including a processor and a memory; the processor executes the steps of the method described in the first aspect of the present invention by invoking a program or instructions stored in the memory.
[0013] According to a third aspect of the present invention, a computer-readable storage medium is provided that stores a program or instructions that cause a computer to perform the steps of the method described in the first aspect of the present invention.
[0014] The cross-platform data control and circulation method based on data space provided by this invention can produce the following significant technical effects compared with the prior art: (1) It realizes end-to-end controllability and reliable traceability of data flow: By generating and iteratively updating flow control information containing electronic signatures through a third-party data space platform, an immutable chain flow path is established for data packets. This technology directly solves the pain points of chaotic paths and unclear responsibilities in multi-level flow scenarios, enabling any party to quickly and accurately trace the exact source of the data and the entire distribution lifecycle, providing a solid technical foundation for data auditing and compliance supervision.
[0015] (2) It provides intelligent proactive detection capabilities for data processing traces: It innovatively proposes two complementary processing trace detection mechanisms (direct comparison and feature analysis). This not only enables accurate comparison when a copy of the source data is available, but also, when the source data cannot be obtained, it can intelligently infer whether the data has been processed by analyzing the inherent statistical characteristics of the data (such as entropy value and data distribution). This technology overcomes the limitations of traditional electronic signatures, which can only prevent tampering but cannot identify processing, and achieves effective supervision of whether data users violate the "no processing" agreement, greatly enhancing the confidence of data providers in sharing original data.
[0016] (3) Enhanced security and compliance of cross-platform data sharing: Electronic signature authentication is integrated into every step of the data transfer process, and blockchain technology is used to store transfer records, creating a decentralized trust environment. This technology ensures that the identities of the participants are trustworthy, the data is complete and tamper-proof, and the operation records are non-repudiable at every stage from provision and transmission to use. It provides strong technical support for the secure, compliant, and efficient circulation of data elements, and directly helps enterprises meet the compliance requirements of laws and regulations such as the Data Security Law.
[0017] (4) Improved the level of automated operation and maintenance of the data circulation ecosystem: This method transforms manual compliance auditing into a platform-based automated detection process. Once the system detects a violation, it can automatically generate and send a warning message. This technology significantly reduces the regulatory cost and complexity of manual auditing for data circulation, improves efficiency, makes large-scale, high-frequency data circulation possible, and promotes the prosperity of the data ecosystem.
[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A flowchart of a cross-platform data control and circulation method based on data space provided in an embodiment of the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0023] It should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of these steps can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the steps can be rearranged. A process can be terminated when its operation is complete, but it may also have additional steps not included in the figures. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. This invention provides a cross-platform data control and flow method based on data space to solve the problems of uncontrollable multi-level data flow paths and difficulty in monitoring processing traces, ensuring that the entire data flow process is traceable, auditable, and controllable.
[0024] For the purposes of describing the embodiments of the present invention, key terms are defined as follows: Electronic signature: refers to a digital signature technology based on public key infrastructure (PKI) technology, generated by the signer's private key, which can be used to verify the signer's identity and ensure data integrity.
[0025] Data provider: refers to the entity that generates or owns the data and entrusts the platform to control the distribution of the data.
[0026] First data user: refers to the entity that first obtains the right to receive data from the data provider.
[0027] Subsequent data users / target users: refers to entities that obtain the right to receive data from the previous data user in the data flow chain.
[0028] Unless otherwise stated, the term 'data user' as used below refers to all data receiving entities.
[0029] like Figure 1 As shown in the figure, an embodiment of the present invention provides a cross-platform data control and circulation method based on data space, which may include the following steps: S100, Receive raw data packet from data provider, the raw data packet containing raw data and corresponding first electronic signature; S200, before distributing the original data packet to the data user, verify the validity of the first electronic signature; S300, Generate and attach flow control information to the original data packet to form a controlled data packet; S400, when the current data user is transferring data to the target data user, the controlled data packet must be submitted to a third-party data space platform; S500: After the third-party data space platform verifies the identity of the current data user and the integrity and authenticity of the historical flow control information in the controlled data packet, it generates new flow control information containing the current user identifier, the target user identifier and the platform's electronic signature, iteratively updates the controlled data packet, and then distributes the updated controlled data packet to the target user, thereby forming a traceable chain flow path. S600 performs processing trace detection operations on controlled data packets in circulation.
[0030] The method provided in this invention achieves end-to-end traceability of data distribution by constructing a chain-like flow path, thus solving the problem of uncontrolled flow. By introducing an intelligent data feature analysis mechanism, it overcomes the limitation of only verifying integrity but not identifying processing, and achieves proactive detection of data processing behavior. Through platform-based automated management and control, it significantly reduces compliance risks and regulatory costs in data circulation, promoting the safe and efficient flow of data elements.
[0031] Furthermore, the cross-platform data control and circulation method based on data space provided in this embodiment of the invention is applied to a third-party data space platform, which is connected to the terminals of at least two data providers and two data users respectively.
[0032] In this embodiment of the invention, the data provider refers to the owner or authorizer of the original data in the data circulation ecosystem, and is the starting point of the data value chain. Its core characteristic is possessing data sovereignty. Its role is that of a "seller" or "authorizer" of the data. The data provider must provide raw, unprocessed data, be responsible for classifying and grading the data, and set data usage strategies and compliance requirements (e.g., "for model training only," "no resale," "no processing"). The data provider's goal is to realize the monetization of data value or business cooperation through data sharing, while ensuring data security and compliance. For example, a bank (data provider) provides its anonymized user transaction records to a compliant fintech company (user) for credit model analysis.
[0033] In this embodiment of the invention, the data user refers to the party in the data circulation ecosystem that acquires, consumes, and applies data; they are the end point for realizing data value. Their core characteristic is the right to use the data. Their role is that of the "buyer" or "authorized party" of the data. The data user must strictly adhere to the agreement and data usage strategy signed with the data provider. Before acquiring data, they must pass identity verification through a third-party platform to prove they are a legitimate and authorized recipient. The data provider must not perform operations prohibited by the agreement (such as unauthorized processing, resale, or disclosure). The data provider's goal is to solve specific business problems, train AI models, conduct scientific research, or generate new insights by acquiring external data. For example, a research institution (user) acquires medical image data from multiple hospitals (providers) to train an AI model for assisted diagnosis.
[0034] In the data space constructed by this invention, data providers and data users do not connect directly point-to-point, but interact through a third-party data space platform. The platform acts as a trusted intermediary, ensuring that the provider's control and the user's usage rights are realized within compliance limits. It also uses technical means (such as electronic signatures, flow control, and processing detection) to constrain the user's behavior and protect the provider's rights.
[0035] In this embodiment of the invention, the third-party data space platform is the core execution entity and data circulation hub. It is not a simple data storage warehouse or transmission pipeline, but a decentralized, trusted data circulation infrastructure and governance environment built on a standardized architecture. Its "third-party" attribute emphasizes its core characteristics of technological neutrality and business agnosticness; that is, it neither produces nor consumes data, nor does it favor any participant. Its sole role is as a "mediator" trusted by all participants, enforcing pre-agreed rules through technical means to ensure the order of data circulation. Specific roles include: A trustworthy and neutral coordinator: As a third party trusted by all data providers and users, the third-party data space platform is responsible for enforcing rules such as identity verification, authorization management, and flow strategy execution, thus resolving trust issues among multiple parties.
[0036] The hub of data flow: All data inflows and outflows must pass through this platform. It does not centrally store business data, but centrally manages data metadata, flow path information, and access policies, serving as a "transportation hub" for controlling data flow.
[0037] Compliance technology enforcers: The platform transforms legal clauses in data sharing agreements (such as "no processing" and "for specific purposes only") into automatically enforceable technical rules (such as processing trace detection and access control), ensuring compliance through code rather than human intervention.
[0038] To fulfill these roles, the platform typically includes the following key functional modules or technical components: Identity and Access Management: Responsible for the registration, authentication, and authorization of all participating parties. Typically based on a Public Key Infrastructure (PKI) system, digital certificates are issued to each participant to ensure the authenticity of their identity in each interaction.
[0039] Policy enforcement point: This is the core rule engine of the platform. It parses and executes the data usage policies set by the data provider (e.g., "number of forwardings", "usage period", "whether processing is allowed"), controlling whether data can be sent and forwarded.
[0040] Data traceability and evidence storage module: Responsible for generating, recording, and maintaining the flow control information as described in the claims. This module creates an immutable record of key information for each data flow (who, when, and to whom), constructing a complete chain path. To achieve immutability, this module is often combined with blockchain technology, storing the hash value of the record on the chain.
[0041] Data processing trace detection engine: This is one of the core innovations of this invention. This engine incorporates the detection algorithm described in the claims, enabling it to proactively or upon request analyze data packets in circulation, determine their compliance, and generate evidence of breach of contract.
[0042] Connector: A standardized interface for the platform to connect with external data provider / user systems. It is responsible for standardized data encapsulation, secure transmission, application and verification of electronic signatures, ensuring seamless access to the data space for applications from different platforms and technology systems.
[0043] In the context of this invention, the term "terminal" does not refer to a personal mobile phone or computer as commonly understood, but rather to a software system or client program that represents an organization in automating data exchange. The connection between the third-party data space platform and the terminal is not a simple web page access, but a secure, programmable machine-to-machine communication. This connection is primarily achieved through secure API-based communication connections and standardized interfaces and data formats. In other words, the connection between the third-party platform and the terminal is a machine-to-machine communication based on standard APIs, using a high-strength encrypted channel (HTTPS), and employing two-way digital certificate authentication.
[0044] The steps are described in detail below.
[0045] S100, receive a raw data packet from the data provider, the raw data packet containing raw data and the corresponding first electronic signature.
[0046] In this embodiment of the invention, step S100 is the starting point and foundation of trust for data to formally enter the controlled circulation space.
[0047] In this embodiment of the invention, the original data packet is not a simple collection of data, but a structured data object containing multiple pieces of information. It mainly consists of original data and electronic signatures.
[0048] Here, raw data refers to the initial data directly generated or collected by the data provider, without any derivative processing. Its format can be structured (e.g., database tables, CSV files), semi-structured (e.g., JSON, XML logs), or unstructured (e.g., images, audio / video files). One of the core objectives of this invention is to ensure that what circulates after this stage is always this raw data or its traceable copy.
[0049] The first electronic signature is a digital signature attached to the original data, generated by the data provider using their unique private key. It is not a simple electronic stamp, but a cryptographic credential created based on asymmetric encryption algorithms (such as RSA and SM2) and hash algorithms (such as SHA-256 and SM3). The core function of the first electronic signature is to verify the identity of the data provider and ensure the integrity of the original data. Its technical principle is as follows: Authentication: The first electronic signature is the result of encrypting the hash value of the original data using the data provider's private key. After receiving the data packet, the third-party data space platform can decrypt the signature using a pre-stored public key of the data provider or one obtained from a certified certificate authority (CA). Successful decryption proves that the data packet indeed comes from the claimed data provider, thus achieving authentication of the data source.
[0050] Ensuring Integrity: Before generating a signature, the data provider calculates a fixed-length hash value (digital fingerprint) on the original data. Any minor modification to the original data will cause a significant change in the calculated hash value. During verification, the platform recalculates the hash value of the received original data using the same hash algorithm and compares it with the original hash value decrypted from the signature. If they match perfectly, it proves that the data has not been tampered with during transmission; if they do not match, the data is immediately deemed invalid and rejected, thus ensuring data integrity.
[0051] In this embodiment of the invention, the data provider transmits the generated raw data packets to a third-party data space platform via a secure communication protocol (such as HTTPS or MQTT with TLS). Upon receiving the data, the platform's security gateway immediately triggers an automated verification process to validate the validity of the first electronic signature. This process is a prerequisite for all subsequent circulation controls; only data packets that pass verification are accepted by the platform and proceed to the next stage of circulation and allocation. This eliminates the risk of identity forgery and data corruption during transmission from the outset, laying a solid foundation for the trust system of the entire data space.
[0052] S200, before distributing the original data packet to the data user, verify the validity of the first electronic signature.
[0053] This step is the first critical security checkpoint in building the entire data space trust chain. Its purpose is to ensure the authenticity and integrity of data from the source before distribution, preventing forged or tampered data from entering the circulation process. This verification process is automatically executed by a third-party data space platform and specifically includes the following levels of verification: 1. Verification Content and Principles The verification process is primarily based on Public Key Infrastructure (PKI) cryptographic principles and includes two core verification items: (1) Identity verification: Objective: To confirm that the received raw data packets do indeed come from the claimed data provider, and not from an imposter.
[0054] Technical Process: The platform uses the public key publicly disclosed by the data provider to decrypt the first electronic signature attached to the original data packet. If decryption is successful, it proves that the signature was generated using a private key paired with the public key and held exclusively by the data provider, thus verifying the data provider's identity.
[0055] (2) Data integrity verification: Objective: To confirm that the original data has not been tampered with, added to, deleted from, or destroyed by any third party during its transmission to the platform since it was signed by the data provider.
[0056] Technical Process: The platform uses the same hash algorithm (such as SHA-256) as the data provider to recalculate the received original data, generating a new hash value (also known as a digital fingerprint). Simultaneously, the platform extracts the original hash value generated by the data provider when signing from the successfully decrypted signature. The platform compares these two hash values. If they match perfectly, the data is intact; if even a single byte of the data is modified, the two hash values will be completely different, and the verification will fail.
[0057] 2. Verification process and logical judgment The platform's verification logic is a strictly automated process: (1) Trigger: Before receiving the data packet and preparing to distribute it to its target user, the platform automatically calls the signature verification service.
[0058] (2) Execution: a. The platform queries the data provider's digital certificate from a secure certificate store and verifies whether the certificate itself is valid and whether it was issued by a trusted Certificate Authority (CA) (certificate validity verification).
[0059] b. Using the public key in the certificate, attempt to decrypt the first digital signature.
[0060] c. If decryption fails, the process is terminated immediately and judged as "signature verification failed" (identity is not genuine).
[0061] d. If decryption is successful, perform the hash value comparison as described above.
[0062] e. If the hash values do not match, the result is "data integrity verification failed".
[0063] (3) Result processing: Verification successful: The platform log records that the verification has passed and the data packet is allowed to enter the next stage (S300, additional flow control information).
[0064] Verification failed: The platform immediately halts any further processing of the data packet and sends a security alert to the system administrator and data provider, indicating "Data source authentication failed" or "Data may have been tampered with." The data packet is placed in an isolated area for audit analysis and will never be distributed to data users.
[0065] This step is not only a necessary technical check, but also carries important legal and compliance implications: (1) Establishing the initial trust anchor: It is the foundation of the entire chain of trust tracing. All subsequent electronic signature verification in all circulation links depends on the confidence in the initial data packet.
[0066] (2) Key evidence for the division of responsibility: Successful verification records indicate that the platform has fulfilled its due diligence obligations. If problems arise in subsequent stages, this can be used to prove that the problem did not originate from the transmission process between the data provider and the platform.
[0067] (3) Compliance requirements: It meets the requirements of the Electronic Signature Law and other relevant laws and regulations regarding reliable electronic signatures, providing technical protection for the legality of data circulation.
[0068] S300, generate and attach flow control information to the original data packet to form a controlled data packet.
[0069] This step is the core innovation of this invention in achieving controllable data flow and end-to-end traceability. After verifying the validity of the original data packet (S200), the third-party data space platform no longer simply forwards the data, but acts as a trusted intermediary to "encapsulate" and "mark" the data packet, imprinting it with platform-level control, thereby forming a controlled data packet.
[0070] In this embodiment of the invention, the flow control information is a set of standardized metadata generated by the platform, which includes at least the following key elements: Data Provider Identifier: A unique identifier of the original data source (e.g., a uniquely assigned identity ID or digital certificate hash).
[0071] Current recipient identifier: Uniquely identifies the designated data user in this data transfer.
[0072] Third-party electronic signature: A certificate generated by the platform using its own private key to digitally sign key information in this transaction (such as provider ID, recipient ID, timestamp, etc.).
[0073] Timestamp: Records the precise time when the platform generated this control information, for subsequent auditing and timeliness assessment.
[0074] Flow sequence number: Records the sequence number of this flow in the entire chain path (for example, the first flow is #1, and the next flow is #2).
[0075] The third-party data space platform first extracts key information (such as the provider ID from S100 and the target user ID) from this transaction and generates a unique timestamp and sequence number. Then, the platform uses its own private key to digitally sign this information combination, generating a third-party electronic signature. Next, the platform serializes this information along with the generated signature according to a predetermined data format (such as JSON or XML) to form a structured control information block. Finally, this control information block is appended to the verified original data packet in one of two ways: Method A (as metadata header): The control information block is appended as an additional data header to the beginning of the original data packet. This is a lightweight and easy-to-parse method.
[0076] Method B (Digital Envelope Mode): The original data packet and control information block are packaged together and encrypted using the target user's public key to form a completely new secure data packet. This method offers higher security, ensuring that only the target user can decrypt and view it.
[0077] The final controlled data packet is equal to the original data packet plus the flow control information.
[0078] The S300 achieves a qualitative leap from "data" to "controlled data," and its technological advantages are as follows: Establishing platform control: By attaching its own electronic signature, the third-party platform formally intervenes in the supervision of data flow. Any subsequent transfers must be transmitted back to the platform for verification and unpacking, thereby ensuring the platform's continuous control over the flow of data.
[0079] The first link in creating the traceability chain: The control information generated this time becomes the starting point of the traceable chain path in the entire data flow lifecycle. It's like a "birth certificate," recording the first authorized distribution of the data.
[0080] Ensuring the non-repudiation of the transfer event: Because the control information contains the platform's digital signature, it becomes a strong cryptographic-level evidence proving that "at a certain point in time, the platform authorized the provision of A's data to B," which neither party can deny.
[0081] This lays the foundation for subsequent iterations: each subsequent iteration (S400, S500) adds a new block of control information to the current controlled data packet. All these blocks are cryptographically linked together to form a complete, verifiable chain of trust.
[0082] Step S300 transforms a static raw data packet into a dynamic, traceable, and continuously monitored data entity (controlled data packet) by generating and attaching standardized flow control information. This is the primary and crucial technical means to achieve the "controlled flow" objective described in this invention.
[0083] In step S300, after the platform generates and attaches flow control information, optionally, it can call the data processing trace detection engine to perform initial feature analysis on the original data packet, extract its baseline features (such as data structure schema, statistical features, information entropy value, etc.), and associate these baseline features with the unique identifier of the data packet and store them in a predefined data processing rule base. This baseline feature will serve as the original reference for comparison with the features of data packets in circulation when subsequent processing trace detection is performed.
[0084] S400 When the current data user is transferring data to the target data user, the controlled data packet must be submitted to a third-party data space platform.
[0085] In this embodiment of the invention, this step is a core mandatory rule to ensure that data is continuously controlled during multi-level flow. Its core purpose is to completely eliminate private transfers between data users and ensure that every data distribution event is monitored, authorized, and recorded by the platform, thereby building a complete and tamper-proof traceability chain.
[0086] The "mandatory submission" design is a combination of technical enforcement and protocol constraints, and its specific implementation and extension include the following aspects: (1) Technical mandatory implementation mechanism Unauthorized forwarding design: The decryption key or access permissions of the controlled data packets obtained by the current data user (hereinafter referred to as the "forwarder") from the platform are limited to its own use. The platform is designed not to grant any data user a forwardable token or key. Therefore, the forwarder is technically unable to generate a new controlled data packet that can be mutually recognized by both the target data user (hereinafter referred to as the "receiver") and the platform.
[0087] Client Integration SDK / Connector: In the forwarding party's data processing environment, a dedicated client software (SDK) or connector provided by the platform is integrated. This client is pre-configured so that any attempt to send controlled data packets outwards will be intercepted and redirected to the platform's target API interface. This enforces behavioral compliance at the code level.
[0088] Policy Enforcement Point (PEP) Integration: The data access policy issued by the platform explicitly includes forwarding constraints (such as...).<No_Redistribution> This policy is enforced on the client side; any forwarding attempt that violates the policy will be blocked by the policy enforcement point.
[0089] (2) Submission process and data preparation When the forwarding party decides to re-transfer the data to the receiving party, the following standardized process is triggered: a. Request Initialization: The forwarding party initiates a "data forwarding request" through its integrated platform client. This request must contain at least the following metadata: requestor_id: The unique identity credential of the forwarder (such as a digital certificate).
[0090] target_recipient_id: A unique identifier registered on the platform.
[0091] data_package_id: A unique identifier for the controlled data packet to be forwarded.
[0092] (Optional) Purpose: Statement of the purpose of this forwarding.
[0093] b. Data packet preparation: The client encapsulates the entire controlled data packet to be forwarded (including the original data, all historical flow control information and electronic signatures) without any modification.
[0094] c. Secure Transmission: The client submits the forwarding request and a complete, controlled data packet as the payload to the platform-specified, highly available data transfer API endpoint (such as POSThttps: / / ) via a secure, mutually authenticated channel (mTLS). <platform-domain> / api / v1 / data-redistribution).
[0095] The technical advantages of the S400 are: (1) Ensure full recording of transfer events: Forced return to the platform ensures that every authorized transfer from the data provider to the end user is recorded by the platform, avoiding information black holes caused by private transfers and laying the foundation for complete traceability.
[0096] (2) Implement dynamic policy checks: When the platform receives a forwarding request, it can verify in real time whether the forwarding violates the initial policy set by the data provider (e.g., "allow a maximum of 2 forwardings" or "prohibit forwarding to specific types of entities"). The policy check is dynamic, allowing the provider to update the policy at any time and make it effective immediately.
[0097] (3) Maintaining the continuity of the trust chain: This provides the necessary prerequisite for the platform to generate new, verifiable flow control information (including platform signature) in the next step (S500). Only with platform authorization can the new flow relationship be accepted by other participants in the trust chain.
[0098] The S400 process combines "technical enforcement" with "protocol constraints" to centralize the authorization and execution of all data transfer activities on a third-party data space platform. This is not only a necessary condition for achieving chain-like path recording, but also a core hub that transforms the data provider's business rules and security policies into executable technical rules. This fundamentally eliminates the risk of data getting out of control during the transfer process, ensuring that the entire data circulation process is controllable, traceable, and auditable.
[0099] S500: After the third-party data space platform verifies the identity of the current data user and the integrity and authenticity of the historical flow control information in the controlled data packet, it generates new flow control information containing the current user identifier, the target user identifier, and the platform's electronic signature, iteratively updates the controlled data packet, and then distributes the updated controlled data packet to the target user, thereby forming a traceable chain flow path.
[0100] In this embodiment of the invention, this step is the core processing link in building a trusted data flow traceability chain. When the platform receives a retransfer request submitted by the current data user (forwarder), it does not simply act as a transmission channel, but rather acts as a trusted intermediary to execute a rigorous "verification-generation-update-distribution" process, aiming to extend the existing trust chain and ensure that every data transfer event is authoritatively recorded and cannot be denied.
[0101] Its specific extended implementation includes the following sub-steps: (1) Multi-factor verification and auditing The platform first initiates an automated multi-factor authentication process to comprehensively verify the received requests and data: a. Identity Verification: The platform uses the digital certificate registered by the forwarding party to verify the validity of the identity credentials (such as digital signatures) carried in the request. This ensures that the request truly comes from a legitimate, authorized user, preventing impersonation.
[0102] b. Historical Flow Chain Integrity Verification: This is the core of the verification process. The platform parses all existing flow control information blocks in the controlled data packet and performs step-by-step cryptographic verification: Starting with the latest control block, the platform uses its own public key to verify the validity of the accompanying electronic signature. If the verification is successful, the transaction record (e.g., from A to B) is confirmed to be authentic.
[0103] Then, extract the "previous block hash" or similar pointer contained in the information block to locate the previous control information block.
[0104] This verification process is repeated, verifying each electronic signature along the chain sequentially, until the first electronic signature generated by the data provider (which has already been verified in S200) is traced back to. This process ensures that the entire historical path has not been tampered with since its inception, and any attempt to modify the historical record will cause all subsequent signature verifications to fail.
[0105] c. Policy Compliance Check: The platform parses verified historical transfer control information to check whether the current forwarding violates the policy set by the data provider. For example, it checks whether the number of transfers has exceeded the "maximum number of forwards" limit, or whether the target recipient is on the "allow sharing" whitelist.
[0106] (2) Generate new flow control information Once all verifications are successful, the platform will create an authoritative record for this new transaction event: a. Information Assembly: The platform generates a new structured flow control information block, the content of which includes at least: From: The unique identifier of the forwarding party.
[0107] To: The unique identifier of the target data user.
[0108] Timestamp: The precise timestamp of this transfer.
[0109] PreviousBlockHash: A key field. This is the hash value of the latest control information block in the previous (i.e., the one held by the forwarder) controlled data packet. This field acts like a "chain," cryptographically linking the old and new blocks.
[0110] b. Platform Signature: The platform uses its own private key to digitally sign the entire new control information block, generating a platform electronic signature, and attaches it to the information block. This signature proves that "the platform authorized this transfer from B to C at a certain moment."
[0111] (3) Iteratively update the controlled data packets The platform appends the newly generated, signed control information block to the existing controlled data packet. At this point, the data packet structure becomes: [Original Data] + [Control Block 1 (A->B)] + [Control Block 2 (B->C)] + ... Important note: This process is an "addition" rather than a "replacement". All historical control information is fully preserved, forming an ordered data structure that grows over time.
[0112] (4) Securely distribute to the target user The platform distributes the updated, controlled data packets to the target data user via a secure channel (such as mTLS). Upon receiving the packets, the target user can independently verify the validity of all platform e-signatures throughout the entire chain, thereby confirming the complete flow history of the data.
[0113] (5) Constructing a chain-like flow path Through the above process, each successful transfer adds a new, platform-certified step to the data packet. Each new step contains the cryptographic hash of the previous step, making the entire transfer path: Traceable: Anyone can trace back from the latest point to the original data provider by following the "PreviousBlockHash" pointer.
[0114] Verifiable: Each step has a digital signature from the platform, which can be independently audited.
[0115] Immutable: Any modification to a historical step will cause its hash value to change, thus causing the "PreviousBlockHash" verification of all subsequent steps to fail, making it easy to detect.
[0116] The S500 process, through an iterative "verification-signature-linking" mechanism, acts like stamping each movement of data with an unforgeable "notary stamp," and sequentially links all these stamps together to form a robust and traceable chain of trust. This not only enables comprehensive monitoring of the data lifecycle but also provides strong technical evidence for data compliance auditing and dispute tracing, forming the technological cornerstone of this invention's goal of achieving controllable data circulation.
[0117] S600 performs processing trace detection operations on controlled data packets in circulation.
[0118] This step is the core innovation of this invention in achieving proactive compliance supervision. Its purpose is to break through the limitation of traditional electronic signatures, which can only verify whether "data has been tampered with," and to proactively detect whether data has been "illegally processed," thereby ensuring that data users strictly abide by the terms of the agreement signed with the provider.
[0119] In this embodiment of the invention, the processing trace detection operation can be triggered by one or more of the following methods: a temporary detection request initiated by the data provider or regulator; or automatic triggering by a third-party data space platform according to a preset strategy (such as periodic polling or before the next data packet transfer). Wherein: Regular polling and inspection: The platform automatically conducts random checks on key data copies in circulation according to a predetermined plan (such as once a day).
[0120] Event-driven detection: Triggered when a specific event occurs, such as when a new round of data transfer is about to begin (before S500) or when a data user requests to use the data for a new purpose.
[0121] On-demand testing: Data providers or regulators can request testing of a particular piece of data from the platform at any time.
[0122] Furthermore, the operation of performing machining trace detection includes: Step 1: Compare the data in the current data packet with the original data traced back to the data provider; and / or Step 2: Based on a predefined data processing rule base, analyze the characteristics of the current data packet to determine whether the current data packet is derived from the original data processing. Step 3: If any processing behavior that violates the data provision agreement is detected, a breach of contract warning message will be generated and sent to the data provider and / or the regulator. Specifically, if the detection results indicate a high degree of suspicion that there is a processing behavior that violates the data provision agreement, a breach of contract warning message will be generated and sent to the data provider and / or the regulator for final manual adjudication.
[0123] Step one is a direct comparison method, suitable for scenarios where the original data can be accurately traced and a copy obtained. It's a precise "digital fingerprint" comparison. Specifically, it analyzes the electronic signature sequence in the chain-like flow path, verifies and locates the original data provider level by level, and obtains the original data copy for comparison. The specific process may include: Source tracing and location: The platform parses all electronic signatures in the chain-like flow path attached to the data packet to be tested (the current data packet). Starting from the latest signature, the platform verifies its validity level by level, and finally locates the original data provider and its corresponding unique data identifier. This process benefits from the complete trust chain built by S300-S500.
[0124] Data Acquisition: Based on the located information, the platform retrieves the original copy of the raw data from the data provider's repository (or the platform's own cache).
[0125] Precise comparison: The platform performs a comparison between the original data copy and the data in the current data packet. This is not a simple byte comparison, but a more intelligent comparison: Consistency check: Calculate and compare the hash values of the two hashes. If they do not match, it is immediately determined that the data has been modified.
[0126] Content Differential Analysis: If the protocol allows certain processing (such as anonymization) but prohibits others (such as aggregation), the platform will perform more granular content analysis, such as checking whether specific sensitive fields are retained or whether numerical precision has changed.
[0127] Technical effect: This method can provide irrefutable, definitive evidence that directly proves whether the current data is completely consistent with the original data.
[0128] Furthermore, step two involves feature analysis, a method suitable for scenarios where raw data cannot be directly obtained (e.g., due to privacy regulations or data deletion) or where rapid, preliminary screening is required. It is an intelligent inference based on probability and statistics. Specifically, it includes: By analyzing the data structure, statistical characteristics, entropy value, or the presence of known data processing algorithm traces in the current data packet, we can infer the derivation relationship between the current data packet and the possible source dataset.
[0129] This invention analyzes the multi-dimensional features of the current data packet and intelligently compares them with the source data benchmark features or general data features stored in the predefined data processing rule base, thereby inferring its derivation relationship with possible source datasets.
[0130] In this embodiment of the invention, the data processing rule base is a knowledge base that stores the feature fingerprints left by various data processing operations. The rules in the rule base can be defined by experts or generated through machine learning training.
[0131] Among them, data structure analysis is used to analyze changes at the data pattern level. This is the most direct and fastest detection method, which includes: comparing the field list of the current data packet with the source data field list recorded in the data processing rule base, as well as comparing the data type and precision of the fields in the current data packet.
[0132] If certain source data fields are missing from the current data packet (such as the deletion of mobile phone number or detailed address), it is determined that data anonymization or field filtering may have been performed. If new fields are added (such as age segmentation), it strongly suggests that data derivation calculations have been performed.
[0133] For example, a birth date field that is a DateTime type in the source data becomes an Integer type age field in the current data packet, indicating that it has undergone a derivation calculation. A latitude and longitude field that is decimal(10,6) becomes decimal(6,3), indicating that it has undergone a precision reduction process, which is a common fuzzification process.
[0134] Among them, statistical feature analysis uses quantitative indicators to deeply analyze changes in the data content itself. It is suitable for scenarios where the field structure remains unchanged but the content has changed, including: Count the number of data records in the current data packet. If the number of rows is significantly less than the number of source data rows recorded in the data processing rule base, data sampling or filtering may have occurred. If the number of rows increases, data synthesis or join operations are very likely to have been performed.
[0135] Calculate the statistical values (mean, variance, median, extreme values) of numerical fields (such as income, transaction volume) and compare them with benchmark values in the data processing rule base. Significant changes in the mean may stem from filtering or bias; a significant decrease in variance may indicate that the data has been smoothed or its range limited (e.g., extreme high or low values have been removed). The data processing rule base predefines difference thresholds for different data fields. A statistical characteristic value is considered significant when its rate of change exceeds its corresponding threshold.
[0136] The number of unique values in statistical categorization fields (such as city or product type). A decrease in the number of unique values is a strong signal of data aggregation (such as summarizing cities into provinces) or filtering.
[0137] Information entropy is a metric that measures the disorder or uncertainty of data, and is particularly suitable for detecting the loss of data diversity. The principle of information entropy analysis is that any processing operation that reduces data diversity (such as aggregation, sampling, and filtering) will lead to a decrease in information entropy. The specific steps of information entropy analysis include: If the entropy value of a certain category field (such as occupation) in the current data packet decreases by more than a predetermined threshold (e.g., 40%) compared to the baseline entropy value of the source data of that field recorded in the data processing rule base, it is determined that the data may have undergone aggregation or sampling processing; for example, aggregating specific occupation types into categories such as "blue-collar" and "white-collar" will directly lead to a decrease in the entropy value of that field.
[0138] Known data processing algorithm trace detection is used to detect statistical "fingerprints" left after using specific privacy-preserving computation or data synthesis algorithms. This can include: differential privacy noise detection, synthetic data detection, and K-anonymity verification. Differential privacy techniques introduce noise with a specific mathematical distribution (such as Laplace or Gaussian distribution) into the data. The platform can analyze the noise in the data and determine whether it conforms to the above distribution through hypothesis testing (such as the KS test). If a match is found, it can be determined that the data has been processed using differential privacy techniques. Synthetic data detection, based on synthetic data generated by generative models (such as GANs), often contains micro-statistical features not present in real data (such as abnormal correlations or overly smooth distributions). Pre-trained synthetic data detection models can be integrated into the rule base to identify these features, thereby determining whether the data is synthetically derived. K-anonymity verification checks whether the data satisfies K-anonymity. That is, it checks whether all combinations of quasi-identifiers (such as zip code, gender, and age) appear at least K times. If the frequency of all combinations is exactly a multiple of K, this is a clear indication of K-anonymity.
[0139] Furthermore, the step of analyzing the characteristics of the current data packet based on a predefined data processing rule base to determine whether the current data packet is derived from the original data processing specifically includes: S601, extract one or more of the following from the current data packet: data structure features, statistical features, information entropy features, and algorithm trace features.
[0140] The data structure characteristics refer to the structural differences between the current data packet and the source data pattern recorded in the rule base. These differences include the increase or decrease in the number of fields, changes in field names, conversions in field data types (e.g., from string to integer), and adjustments in field precision (e.g., reduction in decimal places for floating-point numbers). Extraction is automatically completed by parsing the metadata of the data packet. Specifically, this involves parsing the DDL statements of database tables, the schema of Parquet files, and the structure definition of JSON, extracting relevant information, and comparing it with the source data pattern.
[0141] Statistical features are quantitative indicators calculated from specific numerical values in a data packet. These include the total number of data records, the mean, variance, standard deviation, maximum value, minimum value, and median of numerical fields, and the number of unique values and the frequency of occurrence of the highest-frequency value for categorical fields. These indicators can be extracted by executing SQL queries (e.g., COUNT(), AVG(), VARIANCE()) or by using built-in functions of big data analytics frameworks (e.g., Spark, Pandas) to calculate the numerical values in the data packet.
[0142] Information entropy is a feature calculated for categorical fields. It measures the uncertainty or diversity of a field. A higher entropy value indicates more chaotic and diverse data, while a lower entropy value indicates more ordered and homogeneous data. The calculation follows the formula H(X)=-Σ(p(x)). i )×log2(p(x i ))), where p(x i ) is the field x i The probability of each unique value appearing in the data is calculated, where i ranges from 1 to n, and n is the number of fields. The platform automatically calculates the entropy value of a specific field in the current data packet. The entropy value of the specific category field in the current data packet, calculated automatically by the platform, is extracted by statistically analyzing the probability of each unique value appearing in the field and substituting it into the information entropy formula for further calculation.
[0143] Algorithmic traces are statistical "fingerprints" left by known data processing algorithms that may exist in the data. Examples include numerical distributions conforming to a Laplace or Gaussian distribution (reflecting traces of noise added by differential privacy), or characteristics of synthetic data. Extraction is performed by conducting statistical hypothesis tests (e.g., using the KS test to verify distribution hypotheses) or by invoking a dedicated pre-trained machine learning model for inference, to detect the presence of these algorithmic traces in the data.
[0144] S602, the extracted features are matched with predefined rules in the data processing rule base, wherein each rule defines the mapping relationship between a specific data processing operation and feature changes.
[0145] Each rule is an "IF-THEN" logical statement that encapsulates domain knowledge. The data processing rule base is an extensible collection containing multiple rules. For example: Rule 1: IF [The variance of the age field has decreased significantly, for example, by more than 60%] THEN [Process type: Numerical range limited; Confidence contribution: 0.7] Rule 2: IF [The information entropy of the city field significantly decreases, for example, by more than 40%] THEN [Processing type: aggregation or sampling; confidence contribution: 0.8] Rule 3: IF [The precision of the latitude and longitude field is reduced, for example, from 6 decimal places to 2 decimal places] THEN [Processing type: precision reduction / fuzzification; confidence contribution: 0.9] Rule 4: IF [The noise in the data passes the KS test and conforms to a Laplace distribution] THEN [Processing type: Differential privacy processing; Confidence contribution: 0.95] The platform compares the extracted feature set with each rule in the rule base. If the current data state meets the condition (IF part) of a rule, then that rule is triggered or a match is successful.
[0146] S603, based on the rule matching result, determine whether the current data packet is derived from the original data and its processing type.
[0147] S603 specifically includes: S6031, based on the confidence contribution value of the triggered rule in the rule matching result, a final data processing confidence is calculated using a predefined synthesis algorithm.
[0148] The platform uses an algorithm to combine the confidence contribution values provided by all triggered rules into a final confidence score. Several algorithms can be used: Weighted average method: Different weights are assigned to different types of rules, and then a weighted average is calculated. For example, algorithmic trace rules have higher weights, while statistical feature rules have slightly lower weights.
[0149] Probabilistic model method: Using models such as Bayesian networks, each rule is treated as a piece of evidence, and the posterior probability of the existence of the processing behavior is updated.
[0150] Maximum value method: Among multiple rules, take the value with the highest confidence contribution as the final value (suitable for scenarios with very obvious features).
[0151] Finally, a data processing confidence score between 0 and 1 (e.g., 0.85) is generated, indicating the probability that the current data packet is derived from the original data processing.
[0152] S6032, Based on the data processing confidence level and the triggered rules, generate a processing trace detection report, the report including at least the possible data processing type, confidence level and judgment evidence.
[0153] The platform automatically generates a structured detection report as the final output. The report includes at least the possible data processing types, the final data processing confidence level, and detailed judgment evidence. This judgment evidence includes a list of triggered rules and corresponding feature change details, specifically including: Possible data processing types: List all processing operations (such as "sampling", "aggregation", "differential privacy processing") that are triggered by the rules.
[0154] Final data processing confidence level: a specific numerical value (e.g., 0.92).
[0155] Evidence for judgment: List each triggered rule in detail, including its specifics, for example: "Rule #2 has been triggered: The entropy of the 'City' field was detected to have dropped from 2.1 to 1.2, with a decay rate of 42.8% (exceeding the 40% threshold)." "Rule #4 was triggered: the KS test showed that the added noise followed a Laplace distribution (p-value < 0.01)." Conclusions and recommendations: Based on the confidence threshold (e.g., greater than 0.8), a preliminary conclusion is given (e.g., "highly suspected of being illegally processed"), and it is recommended that the data provider conduct a manual review.
[0156] In this embodiment of the invention, step three is a crucial step in realizing post-event supervision and accountability for data circulation. Its core lies in establishing a human-machine collaborative judgment and response mechanism: the system performs preliminary detection and risk assessment based on algorithms, generating a structured evidence package; human experts make the final ruling based on the evidence provided by the machine, thus balancing regulatory efficiency with the fairness of the decision.
[0157] Its specific extended implementation includes the following sub-steps: (1) Quantitative determination of high confidence level suspicion The system does not simply provide a "yes / no" conclusion regarding violation. Instead, based on the results of processing trace detection, it calculates a quantified violation risk score (e.g., 0-100 points). This score integrates multiple factors: Feature matching degree: The degree to which the current data features match the predefined processing feature patterns in the rule base.
[0158] Confidence synthesis: Based on the confidence contribution values of multiple triggered detection rules, an overall confidence score is synthesized through weighted averaging or a probability model.
[0159] Abnormality of behavior: Compare with the historical behavior patterns of the data user to determine whether this operation significantly deviates from the norm.
[0160] Strategy violation severity: The severity level of the alleged processing behavior that violates the terms of the agreement (e.g., the severity of "mild desensitization" is different from that of "complete reconstruction").
[0161] When the risk score exceeds the preset high-risk threshold (e.g., ≥80 points), the system determines that there is "high confidence level suspicion" and triggers subsequent processes.
[0162] (2) Automated generation of default notice information The system automatically generates a structured report on suspected breach of contract, containing both machine-readable and human-readable information. This report must include at least: a. Metadata information: Suspect ID: Unique tracking identifier.
[0163] Suspected violator: The identity information of the current data user.
[0164] Associated data packet ID and traceability path: Data packets suspected of being illegally processed and their complete flow history.
[0165] Detection time and timestamp.
[0166] b. Details of technical evidence: Detection method: Clearly indicate whether it is direct comparison or feature analysis.
[0167] Detailed findings: List the triggering rules, the specific numerical values of the feature changes (e.g., "Field A entropy value decreased by 45%), and the calculated risk score and confidence level.
[0168] Suspected processing type: Possible processing operations inferred by the system (such as "data aggregation" or "synthetic data generation").
[0169] Data Sample Snapshot: Under the premise of strictly adhering to privacy protection, we provide anonymized sample fragments of suspected illegal data for comparison with the original data.
[0170] c. Recommendations for handling the situation: Based on a predefined policy library, preliminary handling suggestions are automatically generated (such as "It is recommended to suspend the data access permissions of this account and wait for verification").
[0171] (3) Multi-channel distribution and delivery guarantee The generated default notification message is sent simultaneously through multiple reliable channels to ensure delivery: API Push: Structured reports are pushed to the data provider's internal management system or regulatory platform in real time via a secure API callback interface.
[0172] Email / Message Notifications: Send alert emails or messages to designated contacts of data providers and regulators, including a report summary and a link to view the full report.
[0173] Blockchain-based evidence storage: The hash value of the report is stored on the blockchain to ensure that its generation time and content cannot be tampered with, providing legal evidence for possible subsequent arbitration.
[0174] All sending operations have status receipts and retry mechanisms to ensure reliable delivery of information.
[0175] (4) Manual final decision and feedback closed loop Once the data provider or authorized personnel receive the notification, they can view the full suspect report on their management portal: Manual review: Experts combine the evidence provided by the system, their own domain knowledge, and other contextual information to make a final ruling ("confirmed violation", "false alarm", "further investigation required").
[0176] Execution: Select the ruling result on the platform and issue an execution instruction (such as "Violation confirmed, permanently revoke the party's access rights"). The instruction will be executed automatically by the platform.
[0177] Feedback learning: The final ruling will serve as an important sample to feed back into the system's data processing rule base and detection model, which will be used to optimize detection rules, adjust confidence weights and risk thresholds, and enable the system to learn and continuously optimize itself.
[0178] Step three combines precise automated detection with rigorous human adjudication to build an efficient, reliable, and auditable closed loop for violation detection and handling. This not only significantly reduces monitoring costs for data providers and regulators and enables rapid response to violations, but also continuously improves the accuracy of supervision through a feedback learning mechanism, thereby effectively deterring potential violations and maintaining fairness and order in the data circulation market.
[0179] In this embodiment of the invention, "system" refers to the collective term for all technical entities that execute the aforementioned data control and flow method. It is a complete technical solution composed of software, hardware, network, protocols, and rules, and is a highly automated and intelligent technology execution and management platform. Its core is the software program running on the server, but it also includes the supporting hardware infrastructure, security protocols, data rule base, and interaction interfaces.
[0180] In this embodiment of the invention, the flow control information is stored in the form of a blockchain, and the record of the chain flow path is made immutable and traceable through the transaction hash chain on the blockchain.
[0181] In this embodiment of the invention, after the third-party data space platform generates flow control information in step S300 or S500 each time, it does not simply append it to the data packet, but instead writes its key content (such as data provider identifier, receiver identifier, timestamp, data packet hash value, etc.) as a transaction into the blockchain network. The specific implementation includes: Transaction generation: The platform packages the core metadata of this transaction and signs it using its private key to generate a legitimate blockchain transaction.
[0182] Consensus on-chain: The transaction is broadcast to nodes in the blockchain network, verified by the network consensus mechanism (such as PBFT, Raft, etc.), and packaged into a new block. Once the block is added to the chain, the record is permanently fixed and cannot be tampered with or deleted by a single institution.
[0183] Hash Association: The "chained flow path" is implemented through a transaction hash chain. Specifically: When the flow control information (S300) is generated for the first time, its blockchain transaction hash value TxHash1 is calculated.
[0184] When the data is retrieved and processed for the next transaction (S500), the newly generated transaction control information will be written to the blockchain and will explicitly include the hash value TxHash1 of the previous transaction as its input or associated field.
[0185] This process iterates continuously, forming a chain of interconnected records on the blockchain, consisting of transaction hashes TxHash1->TxHash2->TxHash3->... This hash chain corresponds perfectly to the external data flow path.
[0186] Through this mechanism, records of all data transfer events are distributed and stored across various nodes of the blockchain, no longer relying on data storage for any single centralized platform. Any attempt to tamper with or deny historical transfer records would require simultaneously altering all blocks following that record on the blockchain and controlling more than 51% of the node computing power, which is virtually impossible in practice. This provides a high-strength, publicly verifiable, and immutable record of evidence for the entire data flow process. Any participant or regulatory body can independently verify and trace the on-chain records, clearly auditing the complete path of data from its origin to the end user, greatly enhancing the credibility and transparency of the entire system.
[0187] (Example) Suppose that data provider Company A needs to share a raw user behavior dataset with user Company B, and agrees that Company B shall not process the data.
[0188] First, Company A uses its private key to generate a digital signature (first electronic signature) on the original data, and then packages the data and signature together and sends them to a third-party data space platform P.
[0189] Upon receiving the data, platform P uses company A's public key to verify the validity of the signature, confirming that the data source is authentic and has not been tampered with. Subsequently, platform P generates the first flow control information (e.g., From:A, To:B, Timestamp:T1) and signs it with its private key, attaching this information to the original data packet to form a controlled data packet sent to company B.
[0190] Subsequently, if Company B wants to share data with its partner Company C, it cannot send it directly; it must submit the entire controlled data packet it received back to platform P. Platform P will check Company B's identity and verify the validity and completeness of the entire transaction history chain it submitted (including the signatures of A and P).
[0191] After successful verification, platform P generates new flow control information (e.g., From:B, To:C, Timestamp:T2) and signs it. This new information is then appended to the data packet, updated, and sent to company C. At this point, the data packet contains a complete and verifiable flow path from A->B->C.
[0192] To monitor whether Company B has processed data in violation of regulations, platform P can proactively initiate processing trace detection on the data copies held by Company B. Since directly obtaining Company B's internal data may present obstacles, platform P employs a feature analysis approach: Platform P retrieves the baseline features of the dataset provided by Company A from the rule base. For example, the baseline entropy value of the "city" field is H0.
[0193] Platform P requests analysis of the entropy value H1 of the "City" field in Company B's data.
[0194] Calculate the rate of change of entropy. If H1 is found to be much smaller than H0 (e.g., the decay exceeds 50%), it indicates that Company B's data may only contain user samples from some cities and has been filtered and processed; or multiple cities may have been merged into a larger region and aggregated.
[0195] Based on this, platform P generates a suspected breach of contract report and automatically sends it to data provider A and the regulator.
[0196] Through the above methods, this invention achieves effective control and intelligent supervision of the entire data circulation process.
[0197] This invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being configured to perform the method described in this invention.
[0198] This invention also provides a computer-readable storage medium storing computer-executable instructions for performing the methods described in this invention.
[0199] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.
[0200] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A cross-platform data control and circulation method based on data space, characterized in that, Includes the following steps: S100, Receive raw data packet from data provider, the raw data packet containing raw data and corresponding first electronic signature; S200, before distributing the original data packet to the data user, verify the validity of the first electronic signature; S300, Generate and attach flow control information to the original data packet to form a controlled data packet; S400, when the current data user goes to the target data user for data transfer, the controlled data packet must be submitted to a third-party data space platform; S500: After the third-party data space platform verifies the identity of the current data user and the integrity and authenticity of the historical flow control information in the controlled data packet, it generates new flow control information containing the current user identifier, the target user identifier and the platform's electronic signature, iteratively updates the controlled data packet, and then distributes the updated controlled data packet to the target user, thereby forming a traceable chain flow path. S600 performs processing trace detection operations on controlled data packets in circulation.
2. The method according to claim 1, characterized in that, The first electronic signature is used to verify the identity of the data provider and ensure the integrity of the original data; the flow control information includes at least the data provider identifier, the current recipient identifier, and the third-party electronic signature.
3. The method according to claim 1, characterized in that, The operation of performing machining trace detection includes: Compare the data in the current data packet with the original data traced back to the data provider; and / or Based on a predefined data processing rule base, the characteristics of the current data packet are analyzed to determine whether the current data packet is derived from the original data processing. If any processing behavior that violates the data provision agreement is detected, a breach notification message will be generated and sent to the data provider and / or the regulator.
4. The method according to claim 3, characterized in that, The step of comparing the data in the current data packet with the original data traced back to the data provider specifically includes: By analyzing the electronic signature sequence in the chain-like circulation path, the system verifies and locates the original data provider at each level, and obtains the original copy of the data for comparison.
5. The method according to claim 3, characterized in that, The method, based on a predefined data processing rule base, analyzes the characteristics of the current data packet to determine whether the current data packet is derived from the original data processing, specifically including: Extract one or more of the following features from the current data packet: data structure features, statistical features, information entropy features, and algorithm trace features; The extracted features are matched with predefined rules in the data processing rule base, where each rule defines the mapping relationship between a specific data processing operation and feature changes; Based on the rule matching results, determine whether the current data packet is derived from the original data and the type of processing.
6. The method according to claim 5, characterized in that, The step of determining whether the current data packet is derived from the original data processing and its processing type based on the rule matching results specifically includes: Based on the confidence contribution value of the triggered rule in the rule matching result, a final data processing confidence is calculated using a predefined synthesis algorithm; Based on the data processing confidence level and the triggered rules, a processing trace detection report is generated. The report includes at least the possible data processing types, confidence levels, and judgment evidence.
7. The method according to claim 2, characterized in that, The circulation control information is stored in the form of a blockchain, and the records of the chain circulation path are made immutable and traceable through the transaction hash chain on the blockchain.
8. The method according to claim 2, characterized in that, The electronic signature is generated using a digital certificate based on public key infrastructure. Verification of each transfer includes checking the validity of the certificate issuer and whether the certificate is valid.
9. An electronic device, characterized in that, Including processor and memory; The processor executes the steps of the method as described in any one of claims 1 to 8 by invoking programs or instructions stored in the memory.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a program or instructions that cause a computer to perform the steps of the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
A microservice-based platform-based traffic generation system, method, computer, and storage medium
CN113992553B
Artificial intelligence power dispatching decision-making system based on big data
CN114064997A
Authentication and secret key negotiation method suitable for wireless sensor
CN114640453A
Method, system and equipment for checking authenticity of electronic archive file based on original handwriting signature and medium
CN115952560A
Electronic file full life cycle identification system, method, equipment and medium
CN117390695A