Data lifecycle control method, system, device, medium and program

By dynamically generating data governance and usage policies, the problem of permission management in dynamic scenarios in existing data access control methods is solved, and compliance and security are achieved throughout the data lifecycle.

CN121211481BActive Publication Date: 2026-05-01INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
Filing Date
2025-11-25
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing data access control methods struggle to achieve dynamic authorization management in scenarios such as data transactions, large-scale model training, and secondary data circulation. They are unable to generate associated resources on the provider side according to the specific requirements of data contracts and lack extended control capabilities after data flow, leading to uncontrolled permissions and security risks.

Method used

By invoking the state extraction algorithm, data management and usage policies are generated, permissions are dynamically adjusted, data usage is evaluated in real time, resources associated with the provider and user sides are generated, and preprocessing and evaluation are performed to form a closed loop of the entire lifecycle process.

Benefits of technology

It enables dynamic and precise control of data permissions, ensuring compliance and security in data use, avoiding control failures caused by rigid permissions, and preventing data leakage and abuse.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121211481B_ABST
    Figure CN121211481B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, and discloses a data full life cycle use control method, system, device, medium and program, which solves the problem that the existing access control cannot dynamically and differentially adjust the permission and evaluate the absence according to the data contract and its performance state, data use, data secondary circulation and other factors. The method comprises the following steps: a data use direction provider initiates a request and signs a data contract; a state extraction algorithm is called to extract a state set, generate a data management and use strategy; according to the strategy and the contract, a providing side associated resource, a transmission resource and a using side associated resource are generated, and further, a usable resource is generated; the use of the usable resource is evaluated, and when it is not in compliance, the strategy is fed back and updated. The system comprises a signing system, an authorization system and a data processing system. The present application is suitable for the use control of the full life cycle of data multiple copies, realizes the dynamic differential authorization and full-process use control evaluation of the full life cycle of data multiple copies, and guarantees the compliant use of data.
Need to check novelty before this filing date? Find Prior Art

Description

Data lifecycle control methods, systems, devices, media, and procedures. Technical Field

[0001] This invention relates to the field of data processing technology, specifically to data security and access control technology, and more particularly to a method, system, device, medium, and program for controlling the use of data throughout its entire lifecycle. Background Technology

[0002] With the rapid development of the digital economy, data has become a factor of production and a component of transactions. In scenarios such as data trading, large-scale model training, and secondary data circulation, the need to manage the entire data lifecycle—from request initiation, data contract signing, data transmission to usage evaluation—is becoming increasingly urgent. Among these, data access control and usage compliance assessment are key links in ensuring the secure and orderly flow of data.

[0003] Current mainstream data access control methods mainly include role-based access control (RBAC) and attribute-based access control (API). These methods have significant limitations. Firstly, permission decisions rely on the subject's role, subject attributes, and object attributes. In current new application scenarios, such as data transactions, large-scale model training, and secondary data circulation, dynamic authorization management is not feasible. In data transactions, permissions need to dynamically change depending on the contract type and execution stage; in model training, changes in model type and training purpose necessitate dynamic permission adjustments; changes in data usage, data environment (including physical and network environments), scope of data use, data protection algorithms and parameters, and data usage period also require dynamic permission changes; during secondary data circulation and sharing, different operations are performed on the data, or it is forwarded between different subjects, resulting in multiple copies of the data. Permissions also dynamically change depending on the subject, the data copy, and the stage of circulation. Existing methods struggle to dynamically adjust data access and usage permissions in real time based on these various changes in these scenarios. Secondly, there is a lack of extended control capabilities after data transfer. After data is transmitted from the provider to the user, it is impossible to effectively assess the behavior at each stage of data use, such as exceeding the time limit, exceeding the scope of use, or failing to destroy data as required after use. It is also difficult to trigger the permission adjustment mechanism in a timely manner for discovered non-compliant behaviors. Thirdly, a closed-loop operation for the entire process of resource generation, transmission, use, and evaluation has not been formed. It is impossible to generate relevant resources on the provider side as needed according to the specific requirements of the data contract. The data preprocessing stage, including privacy protection and format adaptation, also lacks precise matching with the actual needs of the data user, making it difficult to ensure data compliance while also considering data usability.

[0004] The aforementioned issues can lead to situations such as loss of control over permissions and gaps in evaluation during the entire data lifecycle, which not only restricts the effective release of the value of data elements, but may also trigger security risks such as data leakage and data abuse. Summary of the Invention

[0005] The purpose of this invention is to provide a method for controlling the use of data throughout its entire lifecycle, comprising the following steps: a data user initiates a usage request to a data provider, and the data user and the data provider sign a data contract; a state extraction algorithm is invoked to extract a state set, and a data management strategy and a data usage strategy are generated based on the state set; provider-side associated resources are generated from data provider-side resources according to the data management strategy and / or the data contract; transmission resources are generated according to the provider-side associated resources and the data management strategy; a data usage authorization result is generated according to the data usage strategy; user-side associated resources are generated according to the data usage authorization result and the transmission resources; the user-side associated resources are preprocessed based on the data usage authorization result to generate usable resources; or the user-side associated resources are used as usable resources based on the data usage authorization result; the usage of the usable resources is evaluated and / or constrained based on the usage authorization result; and the usable resources are used based on the usage authorization result.

[0006] According to one embodiment of the present invention, the step of generating provider-side associated resources from data provider-side resources specifically includes: the data provider-side resources include data provider-side main resources and existing data provider-side associated resources; searching for data required by the data contract from the data provider-side main resources and / or existing data provider-side associated resources, and using it as the data provider-side associated resources; confirming that the data required by the data contract does not exist, generating the data required by the data contract from the data provider-side main resources and / or existing data provider-side associated resources, and using the data required by the data contract as the data provider-side associated resources.

[0007] According to one embodiment of the present invention, the data management strategy includes any combination of one or more of the following: data subject, data type, data content, data transmission protection algorithm and its parameters, data transmission method, data transmission path, data transmission subject, data transmission time, and data receiving subject.

[0008] According to one embodiment of the present invention, the data usage strategy includes any combination of one or more of the following: data receiving entity, data protection capability of the data receiving entity, data user entity, data protection capability of the data user, data storage protection algorithm and its parameters, data purpose, data usage environment, data usage time limit, data usage scope, data destruction method after use, training model type and training algorithm, and post-training purpose.

[0009] According to one embodiment of the present invention, the data use authorization result includes one or more of the following: allowed data transmission protection algorithms and their parameters, data storage protection algorithms and their parameters, data use environment, data use time limit, data user, data protection capability of the data user, data purpose, scope of data use, type of training model, and any combination thereof.

[0010] According to one embodiment of the present invention, the preprocessing of the user-side associated resources to generate usable resources specifically includes: performing any combination of one or more of the following operations on the user-side associated resources: model training, privacy protection, statistical analysis, denoising, sorting, completion, and fusion with other user-side associated resources, to generate data that meets the needs of the data user as the usable resources.

[0011] According to one embodiment of the present invention, the method further includes: confirming that the use of the available resources is non-compliant and providing feedback; generating and issuing a new data usage policy based on the feedback of non-compliance and an existing data usage policy, or solely based on the feedback of non-compliance; and cyclically executing the method from the point where a data usage authorization result is generated based on the data usage policy, according to the new data usage policy.

[0012] According to one embodiment of the present invention, the setting forms of the data contract terms reference, the selection of the call status extraction algorithm, the parsing method of the data management strategy, the parsing method of the data usage strategy, and the parsing method of the data usage authorization result include: based on rules, configuration files, buttons, circle, checkmark, mark, key, scroll wheel, menu, voice, video, eye contact, gesture, text, bioelectric signals, and virtual reality, at least one of which is also provided by the present invention.

[0013] This invention also provides a data lifecycle usage control system, comprising: a contract signing system, a provider-side associated resource generation module, a transmission resource generation module, a user-side associated resource generation module, an authorization system, and a data processing system; the contract signing system is used for data users to initiate usage requests to data providers, and the data users and data providers sign a data contract; the provider-side associated resource generation module is used to generate provider-side associated resources from data provider-side resources according to data management strategies and / or the data contract; the transmission resource generation module is used to generate transmission resources according to the provider-side associated resources and the data management strategies; the user-side associated resource generation module is used to generate user-side associated resources according to data usage authorization results and the transmission resources; the authorization system comprises: a status acquisition module, used to call a status extraction algorithm to extract a status set; a data control strategy generation module, used to generate a data management strategy based on the status set; a data usage strategy generation module, used to generate a data usage strategy based on the status set; and a remote control module, used by the data provider to remotely verify data usage. The system includes: a data management module for generating data usage authorization results based on the data usage strategy; a data preprocessing module for preprocessing the associated resources based on the data usage authorization results to generate usable resources; or for designating the associated resources as usable resources based on the data usage authorization results; a data usage evaluation module for evaluating and / or constraining the usage of the usable resources based on the usage authorization results; and a data calculation module for the data user to use the usable resources based on the usage authorization results, and for designating the associated resources as usable resources based on the data usage authorization results. The data management module also sends the data usage authorization results to the data preprocessing module, the data usage evaluation module, and the data calculation module. The data usage evaluation module also provides feedback on non-compliant data usage to the authorization system and the data provider.

[0014] According to one embodiment of the present invention, the remote control module is used to receive information from the data usage evaluation module and feed it back to the data control policy generation module and the data usage policy generation module; the data control policy generation module is further used to generate a new data management policy based on the non-compliance feedback from the data usage evaluation module and existing data usage policies, or only based on the non-compliance feedback from the data usage evaluation module; the data usage policy generation module is further used to generate a new data usage policy based on the non-compliance feedback from the data usage evaluation module and existing data usage policies, or only based on the non-compliance feedback from the data usage evaluation module; the instruction issuing module is used to issue the new data usage policy to the data... The system comprises a management and control module, a data usage evaluation module, a data calculation module, and a remote control module. The instruction issuing module is further configured to issue the new data management and control strategy to the provider-side associated resource generation module, the transmission resource generation module, and the user-side associated resource generation module. The data management and control module generates a new data usage authorization result based on the new data usage authorization result. The data processing system begins a new round of data processing and calculation based on the new data usage authorization result. The new round of data processing includes processing data resources based on the data usage authorization result received by the data preprocessing module. The new round of data calculation includes updating the data resources used in the calculation task based on the data usage authorization result received by the data calculation module, and performing task calculations.

[0015] According to one embodiment of the present invention, the non-compliance includes one or more of the following: data user, data purpose, data usage environment, data usage time period, data usage scope, type of training model and training algorithm, data transmission protection algorithm and its parameters, and data storage protection algorithm and its parameters.

[0016] According to one embodiment of the present invention, the working mode and calculation result usage mode of each module of the data lifecycle use control system and the authorization system include: based on rules, configuration files, buttons, circle, check, mark, key, pull wheel, menu, voice, video, eye contact, gesture, text, bioelectric signals, and virtual reality, at least one of which is also provided by the present invention.

[0017] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the data lifecycle use control method described above.

[0018] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements all or part of the steps of the data lifecycle use control method of the above embodiments.

[0019] The present invention also provides a computer program product, the computer program product including computer-executable instructions, which, when executed, are used to implement all or part of the steps of the data lifecycle use control method of the above embodiments.

[0020] This invention dynamically extracts data contract type, contract execution status, data usage purpose, data protection measures, data usage environment, and different data copies, and dynamically generates management and usage strategies based on these different statuses to achieve dynamic authorization and usage control. Simultaneously, this invention dynamically evaluates data usage, promptly detects abnormal behavior during use, and adjusts management and usage strategies in a timely manner based on abnormal behavior, achieving extended control after data flow. Through dynamic usage control evaluation throughout the entire lifecycle, this invention effectively improves the accuracy and security of data usage control.

[0021] This invention uses a state extraction algorithm to obtain a state set encompassing data transaction information, usage environment, data purpose, scope of use, and data format. Based on this state set, it generates data management and usage strategies. When the data contract type, contract execution stage, training model, data purpose, scope of use, or usage environment changes, or when different data is generated during data circulation, the strategies can be updated and adjusted in real time. This enables dynamic and precise control of data permissions, effectively solving the problem of management adaptation in dynamic scenarios. Compared to traditional access control methods that rely on roles or attributes, this invention can accurately match data permission requirements in dynamic scenarios, avoiding management failures caused by rigid permissions and ensuring that data permissions always align with the actual usage scenario.

[0022] This invention constructs a data lifecycle usage control and evaluation system, filling the gap in extended control after data flow. From the moment the data user initiates a request to the provider and signs a data contract, to the generation of provider-related resources, transmission resources, and user-related resources based on data management strategies and contracts, and finally to the evaluation of the usage of available resources, a complete closed-loop process is formed. The data usage evaluation module can monitor key dimensions in real time, such as the data user, data purpose, data usage period, data usage scope, training model type, and training algorithm. Once non-compliance is detected, it can promptly report to the authorization system and data provider, thereby triggering the generation and issuance of new data management and data usage strategies. This achieves an integrated iterative management process of evaluation, feedback, strategy adjustment, and execution, effectively avoiding problems such as data exceeding its scope, exceeding its time limit, or failure to destroy data as required after use, ensuring compliance throughout the entire data flow process.

[0023] This invention achieves on-demand adaptation and compliance assurance of data resources, ensuring both data security and usability. It can search for required data from the data provider's main resources and existing related resources based on data contracts to generate related resources on the provider side. If no corresponding data exists, it generates data based on existing resources, ensuring a precise match between resources and contract requirements. The data preprocessing module can also perform model training, privacy protection, statistical analysis, noise reduction, sorting, completion, or fusion with other related resources on the user-side related resources, generating usable resources that meet the actual needs of the data user. Simultaneously, through the setting of data transmission protection algorithms and parameters, data storage protection algorithms and parameters, and constraints on data usage permissions, it effectively prevents security risks such as data leakage and misuse while adapting to data usage needs, achieving a balance between data security and usability. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced one by one below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 is a flowchart of the data lifecycle usage control method provided by the present invention.

[0026] Figure 2 is a flowchart of the process of generating associated resources from the data provider side.

[0027] Figure 3 is a schematic diagram of the contents of the data governance strategy.

[0028] Figure 4 is a schematic diagram of the contents of the data usage strategy.

[0029] Figure 5 is a schematic diagram of the content included in the data usage authorization result.

[0030] Figure 6 is a schematic diagram of the content included in the preprocessing of associated resources on the user side.

[0031] Figure 7 is a flowchart of the follow-up process of Figure 1. It is a flowchart of generating and issuing a new data usage policy when using available resources based on the usage authorization result but not in compliance.

[0032] Figure 8 is a block diagram of the data lifecycle use control system provided by the present invention.

[0033] Figure 9 is a schematic diagram of the structure of the electronic device provided by the present invention.

[0034] Explanation of reference numerals in the attached figures:

[0035] 100. Data lifecycle usage control system; 110. Contract signing system; 120. Provider-side associated resource generation module; 130. Transmission resource generation module; 140. User-side associated resource generation module; 150. Authorization system; 151. Status acquisition module; 152. Data control strategy generation module; 153. Data usage strategy generation module; 154. Remote control module; 155. Command issuance module; 160. Data processing system; 161. Data management and control module; 162. Data preprocessing module; 163. Data usage evaluation module; 164. Data calculation module; 910: Processor; 920. Communication interface; 930. Memory; 940. Communication bus. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention. To more clearly explain how this invention solves the above problems, the specific execution flow, key steps, and application logic of the data lifecycle usage control method of this invention will be described in detail below with reference to Figure 1.

[0037] This invention provides a data lifecycle usage control method, applicable to every entity in the data lifecycle, including data providers, data users, contracting systems, authorization systems, and data processing systems. Data providers and data users are the fundamental participants in data flow, responsible for initiating or responding to data usage requests and signing data contracts; the contracting system specifically supports the signing of data contracts; the authorization system is the entity responsible for strategy generation and dynamic management, undertaking functions such as status extraction, management and usage strategy generation, and strategy distribution; the data processing system is the entity responsible for data processing and usage evaluation, responsible for generating authorization results, data preprocessing, usage evaluation, and data usage support. Figure 1 is a flowchart of the data lifecycle usage control method provided by this invention. As shown in Figure 1, the method includes:

[0038] In step S110, the data user initiates a usage request to the data provider, and the data user and the data provider sign a data contract.

[0039] Before initiating a data usage request, the data user must first clarify their data needs and scenarios. These scenarios can include data transactions, large-scale model training, and secondary data circulation resulting in multiple copies. The request must include key information that allows the data provider to accurately understand the needs, including the type, scope, intended use, and usage range of the required data, as well as preliminary requirements such as the duration and environmental conditions of data usage. This information will lay the foundation for subsequent contract negotiations between the two parties.

[0040] Upon receiving a request, the data provider will conduct an assessment based on the requirements outlined in the request to determine whether it possesses the data resources to meet the needs, or whether it can generate the required data based on existing resources. If the assessment is successful, the data provider will create a draft data contract. The contract must cover several key elements, including the contract execution and usage phases, contract terms, sharing agreement, the specific scope, type, and price of data use, and may also include binding provisions such as data destruction methods after use and the scope of use of the trained model, thus clarifying the rights and obligations of both parties. Data contracts can be set up in multiples, depending on actual needs. For example, separate contracts can be signed for different types of datasets, or segmented contracts can be signed for different usage phases of the same dataset.

[0041] The contract terms and their format are diverse. Standard terms can be preset based on rules, or custom terms can be defined through configuration files. Specific content can be confirmed via button checkboxes, menu selections, text input, and voice confirmation. It even supports confirmation of contract elements through interactive methods such as virtual reality and bioelectric signals, adapting to the operating habits and technological environments of different stakeholders. The data user reviews the draft contract, confirming that the data type, quantity, price, scope of use, and usage constraints meet their needs, before jointly signing the data contract with the data provider. Once signed and effective, the contract not only serves as the legal basis for data transfer between the two parties but also provides a basis for the subsequent authorization system to extract state sets and generate data management and usage strategies, ensuring a clear execution benchmark for the entire data lifecycle usage control process.

[0042] Data duplication generated in scenarios such as data trading, large model training, and secondary data circulation can be defined as follows: Assuming data set A and data set B, if there exists ①A=B, ②A B If A and B are copies of each other, then A and B are said to be copies of each other; or if B is generated from A, then B is said to be a copy of A.

[0043] Step S120: Invoke the state extraction algorithm to extract the state set, and generate data management strategy and data usage strategy based on the state set.

[0044] After the data user and data provider complete the data contract signing, the authorization system's status acquisition module starts working. Following preset logic, it invokes a status extraction algorithm to acquire information related to the current data lifecycle management from multiple dimensions, and then integrates this information to form a status set. These information sources revolve around the data contract, data provider resources, data user needs, and the data's own attributes, ensuring that the status set comprehensively reflects the preconditions and constraints of data flow and use.

[0045] The state extraction algorithm specifically extracts the following information: First, information related to the data contract, such as the contract type, the scope of data use stipulated in the contract, the contract execution stage, validity period, data delivery nodes, and other details of the terms. Second, resource information on the data provider side, covering the type, quantity, and storage location of the data provider's main resources, as well as the usage records and compliance status of existing related resources on the data provider side. Third, relevant information on the data user side, including the data user entity, the data user's data protection capabilities, the planned usage environment, the scope of data use, the method of data use, the method of data destruction after use, and the expected training model type or data purpose. Fourth, the basic attributes of the data itself, such as data type, data content format, and whether it has undergone preliminary processing. Fifth, data transmission method, data pre-transmission path, data transmission entity, data transmission time, and data receiving entity. Through the screening, integration, and verification of this information, the state extraction algorithm ultimately forms a state set, providing a precise basis for subsequent strategy generation.

[0046] Based on the aforementioned state set, the authorization system generates data management and control policies and data usage policies respectively. The data management and control policy must cover the constraint rules for each stage of data flow, specifically including one or more combinations of data subject, data type, data content, data transmission protection algorithm and its parameters, data transmission method, data transmission path, data transmission subject, data transmission time, and data receiving subject. For example, based on information such as transmission method and data transmission path in the state set, the data management and control policy explicitly specifies the corresponding data transmission protection algorithm and its parameters, while also limiting a specific transmission time window.

[0047] Data usage strategies focus on the standardized requirements of data usage. These strategies may include one or more combinations of the following: data receiving entity, the data recipient's data protection capabilities, the data user entity, the data user's data protection capabilities, data storage protection algorithms and their parameters, data purpose, data usage environment, data usage timeframe, data usage scope, data destruction methods after use, training model type and training algorithm, and post-training uses. For example, combining the information "data is used for training a model in a specific domain" from the state set, the data usage strategy could explicitly define the type of training model and the training algorithm, while also specifying the data usage scope, the specific destruction methods after use, and the timeframe.

[0048] The two generated strategies need to be matched with the information in the state set. If the state set changes due to changes in data contracts or adjustments in usage requirements, the authorization system can call the state extraction algorithm again to update the state set and adjust the data management strategy and data usage strategy accordingly. This ensures that the strategy is always consistent with the actual situation throughout the data lifecycle and provides a clear basis for the subsequent generation, transmission and use evaluation of data-related resources.

[0049] Step S130: Generate provider-related resources from data provider-side resources according to data governance policies and / or data contracts.

[0050] First, it needs to be clarified that the data provider-side resources involved in this step include two specific types of resources: one type is the data provider-side entity resources, which refers to the raw data resources directly held by the data provider and used as the basic source of data, such as unprocessed data accumulated in the provider's daily business and archived historical data; the other type is the data provider-side existing related resources, which refers to data resources that the data provider has previously generated or that have been associated with and retained in the process of interacting with other entities, and that have potential adaptability to the current data needs, such as preliminary processed data prepared for similar scenarios in the past and data fragments retained after being shared with partners. These two types of resources together constitute the basic source for generating the provider-side related resources.

[0051] The process of generating related resources on the provider side must be guided by the usage requirements clearly defined in the data contract, and simultaneously constrained by the data management strategy generated in step S120. Specifically, it involves two execution scenarios: The first scenario is search matching, where targeted screening is conducted between the main resources on the data provider side and existing related resources on the data provider side, based on the requirements stipulated in the data contract, including the required data type, content scope, and format standards. During this process, the data management strategy further limits the screening scope. For example, the "data subject scope" and "data type priority" explicitly stated in the management strategy ensure that the screened data not only meets the usage requirements stipulated in the contract but also satisfies data security and compliance management requirements. If sufficient and fully suitable data can be obtained through the search, this data is directly identified as related resources on the provider side.

[0052] The second scenario involves supplementary generation. After the aforementioned search, it is confirmed that there is no data among the data provider's resources and existing related resources that meets the contract requirements, or that the quantity and format of the existing data cannot meet the actual usage requirements. In this case, the data required by the contract needs to be generated based on the data contract requirements and the attribute characteristics of the data provider's resources and existing related resources. During the generation process, strict adherence to data governance strategies is required, such as the "data generation format standards" and "data content compliance boundaries" stipulated in the governance strategy, to ensure that the generated data not only meets the contract's usage requirements but also complies with relevant data security governance requirements.

[0053] Regardless of whether the provider-side associated resources are generated through search matching or supplementary generation, they must simultaneously meet the usage requirements of the data contract and the compliance constraints of the data control strategy. Their role is to provide compliant and adaptable basic data support for the subsequent step S140 "generating transmission resources based on provider-side associated resources and data control strategy", ensuring that the subsequent data transmission links are compliant and applicable from the source.

[0054] Step S140: Generate transmission resources based on the associated resources and data management policies of the providing side.

[0055] This step uses the provider-side associated resources generated in step S130 as the basic data source, and combines the data management strategy determined in step S120 with the rules related to data transmission. Through systematic processing, it forms transmission resources that can be securely transferred between the data provider and user sides. The goal is to ensure that the data transmission process complies with compliance requirements, while ensuring the integrity and security of the data.

[0056] First, it is necessary to clarify the key elements guiding the generation of transmission resources in the data management strategy. These elements are generated based on a state set in the early stages and specifically cover data transmission protection algorithms and their parameters, data transmission methods, data transmission paths, data transmission subjects, data transmission times, and data receiving subjects. These elements constitute the hard constraints in the transmission resource generation process, and every processing action must be based on these rules to ensure that it does not deviate from the management requirements.

[0057] The specific process for generating transmission resources includes any combination of one or more of the following three key stages: The first stage is data protection processing, which involves processing the associated resources on the providing side according to the data transmission protection algorithm and its parameters specified in the data management strategy. If the management strategy specifies a particular encryption algorithm or desensitization rule, the original data of the associated resources on the providing side is processed according to the corresponding algorithm and parameters to ensure that sensitive information is not leaked during transmission; if the strategy has specific requirements for the data format, the data format also needs to be adjusted synchronously to meet the transmission standards.

[0058] The second step is configuring transmission parameters, which involves determining the specific method and path for data transmission based on the data management policy. The choice of transmission method must strictly match the policy agreement. For example, if the policy specifies a particular transmission protocol or channel type, the corresponding method will be used to initiate transmission. The configuration of the transmission path must comply with the requirements of the "data transmission path" in the policy, ensuring that data is forwarded only through pre-defined authorized nodes to avoid data passing through unauthorized network links. Simultaneously, in conjunction with the "data transmission time" clause in the policy, a reasonable transmission window period is set to ensure that data is transmitted only within the permitted time frame, reducing security risks during unauthorized periods.

[0059] The third step is the embedding of control identifiers, which involves attaching identification information corresponding to the data control policy to the protected data packets. This identification information includes the unique identifier of the data transmission subject, the designated information of the data receiving subject, the unique number of the transmission resource, and the identifier of the transmission protection algorithm used. This facilitates the subsequent verification of the data source and compliance by the receiving party and provides a basis for data flow tracking. The embedding of identification information must adopt the same security processing method as the data packets to prevent tampering or theft.

[0060] After the transmission resources are generated, they need to be verified for compliance by the verification module in the data processing system to confirm whether one or more of the data protection processing, transmission parameter configuration, and control identifier embedding meet the requirements of the data control policy. If any non-compliance is found during verification, the process is returned for reprocessing until all steps meet the control rules. Transmission resources that pass verification will be temporarily stored in a dedicated cache module on the data provider side, waiting for the window period that meets the "data transmission time" clause to initiate transmission, providing secure and compliant data source support for the subsequent step S160 "generating user-side associated resources based on data usage authorization results and transmission resources".

[0061] Step S150: Generate data usage authorization results based on the data usage policy.

[0062] This step is specifically executed by the data management module in the data processing system. The key action is to transform the data usage strategy generated in step S120 into specific and actionable authorization information, providing a clear operational basis for the subsequent generation of related resources on the user side, data preprocessing, and data usage evaluation, and ensuring that data usage behavior strictly complies with data contract agreements and compliance management requirements.

[0063] Before generating data usage authorization results, the data management module will first comprehensively analyze the data usage policy, extracting the constraints directly related to data usage. According to the definition of the data usage policy based on this invention, its scope includes, but is not limited to: the identity restrictions of the data receiving entity, the data protection capability standards that the data receiving party must meet, the identity restrictions of the data using entity, the data protection capability standards that the data using party must meet, the protection algorithms and their parameters to be used during data storage, the specific scope of data usage, the environmental requirements for data usage (including physical environment, network environment, computing environment, etc.), the time limit for data usage, the entities allowed to manipulate the data, the scope and restrictions of data usage, the type of training model and the corresponding training algorithm, and the method of data destruction after use. During the analysis process, these policy clauses will be processed to transform vague descriptions into clear parameters or standards, avoiding ambiguity in subsequent execution stages.

[0064] After parsing, the data management module will adapt and adjust the parsed policy content based on the actual situation at the current stage of the data's entire lifecycle to ensure that the authorization results are both compliant and practical. For example, if the data usage policy stipulates that "data is only used for training specific types of models," the data management module will clearly indicate the specific type of model and the corresponding training algorithm in the authorization results, while also supplementing the computing environment requirements adapted to the model training. If the data usage policy specifies that "data storage must use a specified encryption algorithm," the authorization results will list the parameter configuration of the encryption algorithm in detail, making it convenient for subsequent data users to perform storage protection operations as required. If there are multiple overlapping constraints in the data usage policy, such as "within a specific time period, a designated entity may use data in a limited quantity," the data management module will integrate and correlate these constraints, clarifying the priority and relationship of each condition to ensure that the logic of each clause in the authorization results is consistent and there are no execution conflicts.

[0065] The final data use authorization result strictly corresponds to the constraints of the data use policy, specifically including key information such as permitted data transmission protection algorithms and their parameters, data storage protection algorithms and their parameters, data use environment, data use period, data user, data purpose, data use scope, training model type, and training algorithm. This authorization result is synchronously stored in a designated module of the data processing system and simultaneously sent by the data management module to the data preprocessing module, data use evaluation module, and data computation module: This clarifies the compliance boundaries of preprocessing operations for the data preprocessing module, such as the algorithm standards required for privacy protection; provides the data use evaluation module with the basis for subsequent evaluation, such as whether there is any overdue or out-of-scope data use; and limits the scope and method of data use for the data computation module, such as allowing data to be used only for training calculations of specified models, ensuring that all subsequent data operations are executed in an orderly manner within the authorization framework.

[0066] Step S160: Generate usage-side associated resources based on the data usage authorization result and transmission resources.

[0067] This step is executed by the user-side associated resource generation module in the system. The key is to perform compliance adaptation processing on the transmission resources generated in step S140 and the data usage authorization results output in step S150, forming user-side associated resources that adapt to the subsequent operation requirements of the data user side, realizing the compliant connection of data from the transmission stage to the usage stage, and providing a data source that meets the authorization requirements for the preprocessing or direct use in step S170.

[0068] First, the resource generation module on the user side needs to complete the reception and basic compliance verification of the transmission resources. After receiving the transmission resources from the data provider, the module first parses the control identification information or control policies embedded in the transmission resources. This information includes the data transmission subject identifier, data receiving subject identifier, data user subject identifier, transmission protection algorithm identifier, and the data user's data protection capabilities. During the verification phase, the module compares these identifiers one by one with the constraints in the data usage authorization results: confirming that the "data receiving subject identifier" of the transmission resource matches the "allowed data receiving subject" in the authorization results, the "transmission protection algorithm identifier" matches the "allowed data transmission protection algorithm" in the authorization results, the "data user subject identifier" matches the "allowed data user" in the authorization results, and the "data user's data protection capabilities" matches the "required data user's data protection capabilities" in the authorization results. If a mismatch occurs, the module immediately terminates the subsequent processing flow and reports the anomaly to the data usage assessment module, ensuring that only compliant transmission resources can proceed to the next processing stage.

[0069] After verification, the module will perform targeted processing on the transmission resources based on the constraints in the data usage authorization result. On one hand, regarding the limitations such as "scope of data use" and "data type" in the authorization result, the module will obtain data content from the transmission resources that meets the requirements of the data contract. For example, if the authorization result stipulates that the range of business data of a specific type is [0, 1000] records, the module will receive 1000 records of business data that meet the requirements from the transmission resources, and the range of usable data is from record 0 to record 1000. On the other hand, regarding the requirements of "data transmission protection algorithm and its parameters" in the authorization result, the module will call the appropriate decryption or deprotection tool to process the transmission resources according to the parameters specified in the authorization result, transforming the encrypted transmission resources into data that can be further operated. During the processing, data security requirements will be strictly followed to avoid data leakage or tampering.

[0070] Furthermore, if the data usage authorization result includes format requirements adapted to subsequent usage scenarios, such as adapting to the input format of a specific training model or the field structure of statistical analysis, the user-side associated resource generation module will also perform basic format adjustments on the filtered transmission data. For example, it may retain key data fields according to the authorization requirements, delete redundant fields unrelated to the authorized purpose, or adjust the data encoding format and field naming rules to ensure that the generated user-side associated resources can be connected to the operation in step S170 without additional adjustments.

[0071] The final generated user-side associated resources must meet two requirements: first, the data content must fully comply with the constraints of the data usage authorization results, and the data type, quantity, protection level, etc., must not exceed the authorized scope; second, the data format must be adapted to the subsequent operational needs of the data user side, and can directly support the preprocessing in step S170 or be used directly. After the user-side associated resources are generated, the module will report the generation results to the data management module so that the system can synchronously record the data flow status and provide a basis for subsequent evaluation and traceability.

[0072] Step S170: Based on the data usage authorization result, preprocess the associated resources on the user side to generate usable resources; or based on the data usage authorization result, use the associated resources on the user side as usable resources.

[0073] This step is performed by the data preprocessing module in the data processing system. The goal is to ensure that the associated resources on the user side meet the actual operational needs of the data user, while ensuring that the entire process complies with the constraints of the data usage authorization results, so as to provide a compliant and usable data source for subsequent data usage and evaluation stages.

[0074] The data preprocessing module first determines whether the associated resources on the user side need to undergo preprocessing based on the data usage authorization result: if the authorization result explicitly requires that the data must be processed before it can be used, or if the current state of the associated resources on the user side, such as data format, protection level, or data integrity, does not meet the usage standards agreed upon in the authorization result, then the preprocessing process is initiated; if the associated resources on the user side fully meet the usage requirements in the authorization result and can be used directly for subsequent operations without additional adjustments, then they are directly identified as usable resources.

[0075] When the preprocessing process is initiated, the data preprocessing module selects an appropriate processing method from multiple operation categories based on the constraints of the data usage authorization results. Specific operation types include model training, privacy protection, statistical analysis, denoising, sorting, completion, and fusion with other related resources on the user side. All operations must strictly match the requirements of the authorization results: for example, if the authorization results stipulate that the data is used for training a specific type of model, the module will perform model training adaptation processing on the related resources on the user side, adjusting the feature dimensions and label formats of the data to conform to the input standards of that model; if the authorization results require the data to meet a specific privacy protection level, the module will use the privacy protection algorithm specified in the authorization results to process the data, ensuring that sensitive information in the data is not leaked; if the data contains missing or outlier values, the module will perform integrity repair according to the completion rules and denoising algorithms allowed by the authorization results to avoid affecting subsequent usage; if the authorization results support data fusion, the module will perform field matching and content integration between the related resources on the user side and other compliant related resources already available to the data user, forming a more comprehensive dataset.

[0076] If the associated resources on the user side do not require preprocessing and can be used directly, the data preprocessing module still needs to perform basic verification: confirm that the data format of the associated resources on the user side is consistent with the usage format agreed upon in the authorization result, that the data protection level meets the authorization requirements, and that the data content does not exceed the scope of authorization, so as to ensure that there will be no compliance risks or operational abnormalities during direct use.

[0077] Regardless of whether usable resources are generated through preprocessing or directly used as usable resources from the user side, the data preprocessing module will perform a final compliance verification on the results after processing. The verification includes ensuring that the data type, quantity, scope of use, protection status, and format of the usable resources fully comply with all constraints of the data usage authorization results. For example, whether the scope of data usage conforms to the authorization results and whether privacy protection processing meets the authorization standards. After successful verification, the data preprocessing module will report the processing results to the data management module and the data usage evaluation module, synchronously recording information such as the generation method and compliance status of the usable resources, providing clear traceability for the evaluation in step S180 and the use in step S190.

[0078] Step S180: Evaluate and / or constrain the use of available resources based on the authorization results.

[0079] This step is performed by the data usage assessment module in the data processing system. Its goal is to track the usage of available resources in real time, ensure that all operations strictly adhere to the constraints of data usage authorization results, promptly detect and handle violations, safeguard the compliance and security of data during use, and provide a basis for possible subsequent policy adjustments.

[0080] First, the data usage assessment module clearly defines the assessment basis, namely the various constraints in the data usage authorization results. These constraints specifically include permitted data users, limited data uses, specified data usage time limits, approved data usage scope, designated training model types and algorithms, required data transmission protection algorithms and their parameters, and explicit data storage protection algorithms and their parameters. The module transforms these constraints into quantifiable assessment metrics. For example, it converts "data usage time limits" into specific time interval thresholds and "data usage scope" into statistically measurable operation counts or data access volume limits, ensuring that the assessment criteria are clear and unambiguous.

[0081] Subsequently, the module conducts evaluation operations in two ways: First, it evaluates data usage behavior in real time. The module connects to the data computation module and collects usage logs of available resources in real time. These logs include information such as the identity of the operator, operation time, operation type (e.g., data reading, model training, result output), the amount of data accessed, and the model and algorithm used. Each piece of collected information is compared in real time with preset evaluation indicators. For example, it verifies whether the operator is a permitted data user in the authorization results, whether the operation time is within the specified time limit, whether the amount of data accessed does not exceed the approved limit, and whether the model and algorithm used are consistent with the specified type. If any information does not match the evaluation indicators, the module immediately triggers an early warning mechanism, suspending the current data usage operation to prevent the violation from continuing. Second, it periodically verifies the data usage status. In addition to real-time monitoring, the module reviews and verifies the overall usage status of available resources at preset intervals. Verification includes whether the cumulative data usage exceeds the authorized total, whether data storage still meets the specified protection algorithm requirements, and whether the data usage has deviated from its intended purpose, such as data originally used for model training being used for other analysis scenarios. Regular verification can compensate for any momentary omissions that may occur in real-time assessments, ensuring that the assessment covers all time periods and scenarios in which the data is used. The data assessment module can also collect information other than logs.

[0082] When the data usage assessment module detects non-compliance, it will execute processing actions according to the set process: First, it records detailed information about the non-compliance, including the time of the violation, the entity involved, the specific content of the violation (such as exceeding the time limit, exceeding the scope of access, using an unauthorized model, etc.), and the specific scope of the non-compliant data, forming a complete violation file; Second, it synchronously reports the non-compliance to the authorization system and the data provider. Reporting to the authorization system is to trigger subsequent policy adjustment processes, and reporting to the data provider is to ensure their right to know about data usage; Finally, if the violation is serious, the module can temporarily freeze the access rights to the available resources according to the preset rules in the authorization results until the violation is resolved and compliance verification is passed.

[0083] Throughout the assessment process, the data usage assessment module continuously records all assessment data and verification results, storing them in the system's assessment log database. These logs can be used to trace violations and provide real-world data support for the authorization system to subsequently optimize data control and data usage strategies, promoting the dynamic improvement of data lifecycle management. The information modalities continuously recorded by the data usage assessment module can also include other information besides logs.

[0084] Step S190: Use available resources based on the usage authorization result.

[0085] This step is performed by the data user using the data computing module in the data processing system. The goal is to complete the compliant use of available resources under the full constraints of the data usage authorization results, and to realize the transformation of data from resource preparation to value realization. All operations must follow the data usage rules defined in the usage strategy to ensure that compliance boundaries are not exceeded.

[0086] First, the data computation module will perform pre-use adaptation preparation. The module needs to receive two key pieces of information: first, usable resources that have passed the verification in step S170; and second, the data usage authorization result generated in step S150 and synchronized to the module. Subsequently, the module will initiate basic adaptation verification to confirm that the data format and protection status of the usable resources are consistent with the requirements of "data usage environment" and "data storage protection algorithm and its parameters" in the authorization result. For example, if the authorization result requires data to be used in a specific encrypted computing environment, the module will first verify whether the encryption configuration of the current computing environment meets the requirements; if the authorization result restricts data input to a specific format, the module will confirm that the format of the usable resources meets the standard to avoid usage anomalies due to environment or format mismatch.

[0087] After the adaptation preparation is complete, data users can conduct actual operations on the available resources through the data calculation module. The operation type must strictly match the constraints of "data purpose" and "training model type and training algorithm" in the authorization result. Specific scenarios include the following two forms: First, model training scenarios. If the authorization result stipulates that the data is used for training a specific type of model, the data calculation module will call the authorized training algorithm and input the available resources as training data into the model. During use, the amount and scope of data calls must be monitored in real time to ensure that the single and cumulative amount and scope of calls do not exceed the "scope of data use" limit in the authorization result; at the same time, key parameters during the training process, such as the number of iterations and model accuracy, are recorded to ensure that the training behavior fully conforms to the authorized purpose and does not generate unauthorized model training tasks. Second, statistical analysis scenarios. If the authorization result stipulates that the data is used for statistical analysis on a specific topic, the data calculation module will calculate the available resources and generate analysis results according to the authorized analysis dimensions, such as time dimension and user dimension. During the process, it is necessary to avoid mining dimensions outside the authorization, and the output format of the analysis results must meet the authorization requirements and not contain unauthorized sensitive data fields to ensure that the analysis behavior does not deviate from the authorized scope.

[0088] During use, the data computation module generates usage logs in real time and synchronizes them to the data usage evaluation module. The logs must fully cover usage details, including the operation time (ensuring it falls within the authorized "data usage period"), the operating entity (which must match the authorized "data user entity"), the amount of available resources used, specific usage behaviors (e.g., algorithm selection for model training, dimensions of statistical analysis), and the storage location of output results. These logs serve as the basis for the data usage evaluation module's real-time monitoring and as evidence for subsequent compliance traceability, ensuring the entire usage process is traceable and verifiable. The data computation module can also generate other information besides usage logs.

[0089] After the resources are used, the data calculation module must process the temporary data generated during the process, such as cached data and intermediate calculation results, according to the "data destruction method after use" requirements in the authorization result. If the authorization result requires "immediate destruction after use", the module will start an automatic destruction program to completely clear the temporary data; if the authorization result allows short-term retention, the module will set an automatic cleanup task according to the agreed retention period and perform the destruction operation after the expiration to avoid the risk of unauthorized data retention.

[0090] Throughout the entire usage process, the data usage authorization result must always be the constraint. All operations must be carried out within the control scope of the data processing system. This ensures that the data user can normally obtain the value of the data. Furthermore, the pre-use adaptation, in-use control, and post-use cleanup processes work in synergy with the evaluation process in step S180 to jointly ensure the security and compliance of the data usage throughout its entire lifecycle, fully complying with the compliance requirements for the data usage process in the document.

[0091] According to an embodiment of the present invention, generating provider-side associated resources from data provider-side resources specifically includes: data provider-side resources including data provider-side main resources and existing data provider-side associated resources; searching for data required by the data contract from the data provider-side main resources and / or existing data provider-side associated resources according to the data contract, and using it as data provider-side associated resources; confirming that the data required by the data contract does not exist, generating the data required by the data contract based on the data provider-side main resources and / or existing data provider-side associated resources, and using the data required by the data contract as data provider-side associated resources.

[0092] Specifically, as shown in Figure 2, in the process of generating provider-related resources from data provider-side resources, the specific composition of the data provider-side resources needs to be clarified in step S131. These resources include two categories: primary data provider-side resources and existing related resources on the data provider side. Primary data provider-side resources are the original basic data resources directly held by the data provider. Taking a hospital as the data provider as an example, these resources may include raw gene sequencing data of tumor patients within a specific time period stored in the hospital's laboratory information system, basic patient diagnosis and treatment information associated with the hospital's information system, and gene annotation data of tumor tissue samples archived in the hospital's pathology department. Before using these resources, a preliminary compliance verification needs to be completed to confirm that the data subject has signed a data usage authorization and that the data storage complies with the requirements of relevant laws and regulations regarding the encryption of sensitive personal information storage, thus avoiding compliance risks associated with the source resources. Existing related resources on the data provider's side refer to data resources that the data provider previously generated for other scenarios or retained after interactions with other entities, and whose attributes are compatible with the current data contract requirements. Examples include a subset of tumor patient gene data previously provided by a hospital to a university research team, gene sequencing quality control report template data retained after cooperation with a third-party testing institution, or intermediate result sets of gene data of patients with specific diseases that were previously processed and temporarily stored. For these resources, an initial compatibility assessment must be conducted first. By comparing the matching degree of the resource's field structure, data type, compliance status with the current data contract requirements, resources with reuse potential are selected. At the same time, the authorization for secondary use of the resources is verified to ensure that the reuse process is legal and compliant.

[0093] After clarifying the resource composition of the data provider, step S132 prioritizes searching and matching based on data contract requirements. Taking a biopharmaceutical company's contract requiring raw gene sequencing data of lung cancer patients within a specific timeframe as an example, the search dimensions are first broken down into data type, data usage scope, data timeframe, data format, fields, data volume, and compliance standards. Search priorities are then set, typically starting with searching the main resources of the data provider, followed by searching related resources already provided by the data provider. During the main resource search, the data provider's data resource management platform retrieval interface may be invoked, filtering step-by-step according to the set dimensions. This can be completed by following these steps: first, filtering out data that meets the timeframe and data type requirements; then, removing data that does not meet format requirements or has missing fields; finally, verifying the data's authorization status and encryption through the compliance management module to obtain data that initially meets the requirements. If the amount of data searched for the main resource does not meet the contract requirements, the system will further search for existing related resources on the data provider side, selecting suitable data to supplement it. Finally, data that fully complies with the data contract requirements and compliance criteria will be identified as the provider-side related resources. A resource list containing unique data identifiers, storage paths, compliance verification results, and field integrity reports will be generated and uploaded to the resource management module of the authorization system. Other processes can also be used to generate provider-side related resources.

[0094] When a search and matching process confirms that no data fully meets the data contract requirements exists in the data provider's main resources and existing related resources, supplementary generation will be performed based on the two existing resource types in step S133. First, a supplementary generation plan is formulated by combining the data contract requirements and existing resource attributes, clarifying the content to be supplemented and the implementation method. For example, if the contract requires data to include specific new fields and specify disease staging, but existing resources lack relevant detailed data and fields, then the data will be generated based on the data provider's main resources through data annotation and field supplementation, while referencing existing related resources on the data provider's side to ensure the adaptability of the generated data. In the specific generation operation, the following steps can be used: first, data that meets the basic conditions is selected from the main resources; then, compliant tools or templates are used to supplement and annotate the data with new fields; finally, the generated data is cleaned and formatted uniformly, removing substandard data and ensuring the format meets requirements. Other steps can also be used. After the generation process is complete, a second compliance and compatibility verification will be conducted. The compliance verification confirms that the data's authorization status, sensitive information anonymization, and encrypted storage status all meet the requirements. The compatibility verification will be jointly conducted by the data user's technical personnel and the data provider's data team to confirm that the data fields and formats meet the data user's actual needs and that the data scope meets the contractual requirements. Finally, data that meets the contractual scope requirements is selected from the verified data and designated as supplementary provider-side associated resources. The resource list is then updated and uploaded to the resource management module of the authorization system, completing the entire provider-side associated resource generation process.

[0095] According to an embodiment of the present invention, the data management strategy includes any combination of one or more of the following: data subject, data type, data content, data transmission protection algorithm and its parameters, data transmission method, data transmission path, data transmission subject, data transmission time, and data receiving subject.

[0096] Figure 3 is a schematic diagram of the contents of the data governance strategy. Specifically, as shown in Figure 3, the data governance strategy constructs a compliance management framework covering the data flow process through a flexible combination of multi-dimensional constraint clauses. The components are interconnected and work together to ensure that the data transmission process from the provider to the user is secure, controllable, and traceable.

[0097] Data subjects, as the fundamental constraint dimension of control strategies, clarify the scope of stakeholders involved in data flow, define the ownership and responsibility boundaries of data, and ensure that all data operations revolve around legitimate stakeholders, avoiding compliance risks arising from unauthorized entities participating in data flow. Data type and data content further refine the control objects. By differentiating data categories and specific information scopes, they provide a basis for differentiated control. For example, for professional data in different fields or data containing sensitive information, stricter transmission and protection requirements can be set in the strategy to ensure that control measures are adapted to the characteristics of the data itself.

[0098] Data transmission protection algorithms and their parameters are means to ensure data transmission security. By clearly defining suitable data protection technologies and corresponding configuration standards, data is ensured to remain secure during transmission, effectively resisting the risks of information leakage or tampering. In different scenarios, appropriate protection algorithms and parameter configurations can be selected based on the sensitivity of the data and transmission requirements, ensuring a precise match between the security level of data transmission and actual needs.

[0099] Data transmission method, data transmission path, data transmission subject, and data transmission time together constitute the constraints of the entire transmission process. The selection of the transmission method must meet data security requirements and the actual application environment to ensure the reliability and security of the data transmission channel; the planning of the transmission path must avoid unauthorized nodes and only allow data to flow through preset compliant paths to reduce external risks faced by data during transmission; the restriction of the transmission subject clearly defines the participants with transmission permissions to prevent unqualified entities from participating in data transmission operations; and the control of transmission time, by setting compliant transmission windows, avoids data transmission during unauthorized periods, further enhancing the security of the transmission process.

[0100] The data receiving entity, as the terminal constraint in the transmission process, clarifies the legitimate recipient of the data, ensuring that the data is only transmitted to the designated entity that meets the requirements, avoiding data flow to unknown or unauthorized recipients, and controlling the compliance of data flow from the transmission terminal.

[0101] Figure 3 presents the positioning and relationship of each control dimension in the overall strategy through a modular presentation. The clauses of each dimension can be arbitrarily combined according to the needs of actual application scenarios to form personalized control strategies that are adapted to different data flow scenarios. This ensures both the comprehensiveness and rigor of control, as well as sufficient flexibility and adaptability, providing a solid compliance guarantee for the transmission of data throughout its entire lifecycle.

[0102] According to an embodiment of the present invention, the data usage strategy includes any combination of one or more of the following: data receiving subject, data protection capability of the data receiving subject, data user subject, data protection capability of the user, data storage protection algorithm and its parameters, data purpose, data usage environment, data usage time limit, data usage scope, data destruction method after use, training model type and training algorithm, and post-training purpose.

[0103] Figure 4 is a schematic diagram of the contents of the data usage strategy. Specifically, as shown in Figure 4, the data usage strategy constructs a compliance management system covering the entire process of data usage through the systematic integration of multi-dimensional constraint clauses. The components are progressive and mutually supportive, ensuring that the entire process from data receipt to completion of use not only complies with the contractual agreement but also resists various security risks.

[0104] The data receiving entity, as the initial constraint in the usage phase, clarifies the legitimate recipient of the data, defines the initial ownership of data usage rights, and ensures that data only flows to specific entities that comply with the contract, preventing unauthorized access or misuse of data from the source. The data receiving entity's data protection capabilities, based on its security level, ensure that it possesses the necessary security management capabilities after receiving the data by limiting its security protection conditions. This includes aspects such as physical environment protection standards and the deployment of data protection algorithms, laying a secure foundation for subsequent data use.

[0105] The data user entity serves as a specific constraint in the data usage process, clarifying the legitimate users of the data, defining the specific ownership of data usage rights, ensuring that data is used only by specific entities that comply with the data contract, specifying the specific objects authorized to operate the data, and limiting data usage operations to specific personnel or departments. This prevents unauthorized personnel from accessing or using the data, reduces the scope of data usage risks, and prevents the misuse of data by unauthorized parties from the source. The data protection capabilities of the data user entity are based on the security level of the data user itself. By limiting its security protection conditions, it ensures that the user entity possesses the necessary security management and protection capabilities to undertake and use the data, such as physical environment protection standards, network security protection measures, and the deployment of data protection algorithms, laying a secure foundation for the entire data lifecycle.

[0106] Data storage protection algorithms and their parameters are crucial for ensuring data storage security. By clearly defining suitable data storage encryption or protection technologies and corresponding configuration standards, the security of data during storage can be guaranteed. Different types and levels of data sensitivity can be matched with different storage protection algorithms and parameter configurations, ensuring a precise match between storage security strength and actual data security needs, effectively preventing data leakage or tampering risks during storage.

[0107] The data usage section clearly defines the legal scenarios and scope of data use, limiting it to the specific purposes agreed upon in the contract, such as model training, statistical analysis, and scientific research. This prevents data from being used in unauthorized scenarios and ensures that data use aligns with the original intent of the contract. The data usage environment section sets constraints across multiple dimensions, including physical, network, and computing environments, limiting data use to environments that meet security standards, such as compliant cloud servers and encrypted local computing environments, thus eliminating risks posed by insecure environments.

[0108] The data usage time limit sets a clear time range for data use, stipulating that data can only be used within the agreed time period. After the time limit is exceeded, the data will no longer be accessible or will need to be processed according to the rules, thus avoiding long-term data occupation or overdue use and ensuring the timeliness and compliance of data use.

[0109] The scope of data use constrains data usage by defining both the specific content and quantity. This limits the range of data used, as well as the amount and scope of data used in a single instance or cumulatively, preventing resource waste or security risks caused by excessive data access and ensuring that data use remains within a reasonable and controllable range. The method of data destruction after use standardizes the processing flow after data is used, specifying that data must be destroyed through specific methods, such as multiple overwrites or physical deletion, to prevent data from being retained or illegally reused, thus forming a closed-loop management system for data use.

[0110] The training model types and algorithms are clearly defined for scenarios where data is used for model training. This ensures that the training process complies with contractual agreements and technical specifications, preventing data misuse or illegal training results due to the use of unauthorized models or algorithms. The post-training usage further extends the constraints, clarifying the application scenarios and scope after model training to prevent the training results from being used for unauthorized commercial activities or other illegal scenarios, forming a complete chain of constraints from data use to result application.

[0111] Figure 4 presents the functional positioning and logical relationship of each control dimension in the overall strategy through a modular approach. The clauses of each dimension can be flexibly combined according to the actual needs of different application scenarios such as data transactions, large model training, and multiple copies generated by secondary data circulation to form personalized usage strategy solutions. This ensures compliance and security throughout the entire data usage process while also possessing sufficient adaptability and flexibility, providing comprehensive and robust compliance protection for all stages of data usage throughout its lifecycle.

[0112] According to an embodiment of the present invention, the data usage authorization result includes one or more of the following: allowed data transmission protection algorithms and their parameters, data storage protection algorithms and their parameters, data usage environment, data usage time limit, data user, data user's data protection capability, data purpose, scope of data usage, type of training model, and any combination thereof.

[0113] Figure 5 is a schematic diagram of the contents of the data use authorization result. Specifically, as shown in Figure 5, the data use authorization result serves as a concrete manifestation of the data use strategy, transforming abstract policy constraints into directly executable and quantifiable operational rules. The various components work together to comprehensively cover the key aspects of data use, providing clear guidance for the compliant operation of data users and providing clear judgment benchmarks for the evaluation process.

[0114] The permitted data transmission protection algorithms and their parameters comply with the requirements of data management and data usage strategies, clearly defining the transmission security standards that must be maintained during data reception and subsequent transfer at the user end. The explicit definition of these algorithms and parameters ensures uninterrupted data security protection from transmission to usage. For example, specifying particular encryption algorithms and key lengths allows data to resist leakage risks during internal transfers at the user end, maintaining consistency and continuity in data security protection.

[0115] Data storage protection algorithms and their parameters define the security boundaries for data storage on the user side, clarifying the encryption, de-identification, and other protective measures and specific configuration standards to be adopted during data storage. Different levels of data sensitivity correspond to different storage protection requirements. By clearly defining algorithms and parameters, the security of data during storage on the user side is ensured, preventing data leakage due to inadequate storage protection, and providing a clear basis for compliance verification of data storage.

[0116] The data usage environment clarifies the scenarios and conditions for the legal use of data from multiple dimensions, including physical environment, network environment, and computing environment, and restricts data operations to only be carried out in environments that meet security standards. For example, it specifies that data must be used in a data center with physical isolation protection or a compliant and certified cloud environment, and prohibits operation on public networks or personal devices without security protection, thereby mitigating external risks during data use from an environmental perspective.

[0117] The data usage period defines a clear timeframe for data use, explicitly defining the start and end dates during which data can be accessed and processed. After this period, the data user will no longer be able to obtain or use the relevant data, and the system will automatically trigger a permission freeze or data cleanup mechanism to prevent data from being occupied for extended periods or used beyond its expiration date. This ensures the timeliness and compliance of data usage and also supports the data provider's subsequent data flow planning.

[0118] The data user entity clearly defines the specific individuals or entities with legitimate data access rights, limiting data usage to only specific personnel, departments, or systems. By precisely defining the user entity, the scope of data access is narrowed, preventing unauthorized personnel or third parties from illegally obtaining or manipulating data, reducing the risk of data misuse, and facilitating accurate tracking of the responsible party for data operations during the evaluation process.

[0119] The data usage guidelines further clarify the legal scenarios and purposes for data use, limiting data to specific matters agreed upon in the contract, such as statistical analysis on a specific topic or training of a specified type of model. This constraint ensures that data use remains consistent with the original intent of the contract, prevents data misuse, protects the legitimate rights and interests of data providers, and also guides data users to focus on their primary needs when conducting data operations, thereby improving data utilization efficiency.

[0120] The scope of data usage allows for precise control over its scale, clearly defining the amount and scope of data that can be used in a single instance or cumulatively. Quantitative constraints prevent excessive data access or reuse, thus avoiding data waste and reducing potential security risks from large-scale data usage. This ensures data usage remains within a reasonable and controllable range and provides a clear reference for resource planning by data users.

[0121] The types of training models and training algorithms specified in the regulations clearly define the permitted models and training technology standards for the scenarios in which the data is used for model training. This constraint ensures that the model training process complies with contractual agreements and technical specifications, avoiding issues such as data misuse and illegal training results caused by the use of unauthorized models or algorithms. It also guarantees the professionalism and compliance of the training process, ensuring that the training results are consistent with the data provider's intended use.

[0122] Figure 5 illustrates the functional positioning and logical relationships of each component in the authorization result through a modular layout. These components can be flexibly combined and adapted to the actual needs of different application scenarios, such as data transactions, large model training, and multiple copies generated from secondary data circulation, to form personalized authorization result solutions. This ensures compliance and security throughout the entire data usage process while providing sufficient practicality and flexibility, offering clear guidance for data users' specific operations, and providing solid support for the evaluation phase of data use throughout its entire lifecycle.

[0123] According to an embodiment of the present invention, preprocessing the user-side associated resources to generate usable resources specifically includes: performing any combination of one or more of the following operations on the user-side associated resources: model training, privacy protection, statistical analysis, denoising, sorting, completion, and fusion with other user-side associated resources, to generate data that meets the needs of the data user as usable resources.

[0124] Figure 6 is a schematic diagram of the preprocessing of user-side associated resources. Specifically, as shown in Figure 6, the preprocessing of user-side associated resources is the link between data reception and actual use. Through a flexible combination of various operations, it ensures that the data meets compliance requirements and accurately adapts to the actual business needs of the data user, transforming the associated resources originally adapted for the transmission process into a data form with direct usability.

[0125] Model training-related preprocessing operations focus on the adaptability of data to the training scenario. They adjust the feature dimensions, label format, and data structure of the data according to the subsequent model training requirements, so that the data can be directly input into the specified type of training model, reducing the cost of format conversion and data adjustment in the subsequent training process. At the same time, they ensure that the feature distribution of the data matches the training objective, laying the foundation for improving the model training effect.

[0126] Privacy protection is a key compliance requirement in the preprocessing stage. Based on the privacy protection standards specified in the data usage authorization results, sensitive information in the associated resources on the user side is processed to ensure that the data retains its value without disclosing personal privacy or trade secrets. This ensures that the data usage process complies with the compliance requirements of relevant laws, regulations, and contractual agreements, balancing data usability and security.

[0127] Statistical analysis focuses on extracting valuable information dimensions from original related resources. Through feature statistics, trend analysis, and pattern mining, it transforms scattered raw data into intermediate data with summaries and regularities, providing more targeted data support for data users' decision analysis or model training, and improving the efficiency and accuracy of data use.

[0128] The noise reduction operation mainly targets outliers, redundant information, and interference data in the data. By using an appropriate data cleaning algorithm, it removes content that does not meet data quality standards, corrects deviations and errors in the data, and ensures that the pre-processed data has high accuracy and consistency, thus avoiding the negative impact of abnormal data on subsequent use.

[0129] The sorting operation organizes the data in an orderly manner according to the data user's business logic and usage habits, based on preset dimensions. This allows the data to be presented in an orderly manner according to specified rules, making it easier for the data user to quickly retrieve and call the target data in subsequent operations, thus improving the convenience and smoothness of data use.

[0130] The completion operation supplements and improves missing values ​​and incomplete fields in the data. Based on the characteristics and patterns of the data and common business sense, it uses reasonable completion rules to fill in the missing information, ensuring the integrity and continuity of the data and avoiding deviations in subsequent analysis or training results due to missing data.

[0131] The integration with other related resources on the user side focuses on enriching and expanding the data dimensions. It matches the fields, integrates the content, and logically associates the current related resources on the user side with other data resources that the data user already has and that meet the compliance requirements, forming a dataset with more comprehensive dimensions and more complete information, so as to meet the data user's needs for collaborative analysis or joint training of multi-source data.

[0132] Figure 6 presents the functional positioning and logical relationships of various preprocessing operations in a modular manner. These operations are not executed in isolation, but can be arbitrarily combined according to data usage scenarios, contract requirements, and business needs. For example, in large model training scenarios, denoising, imputation, model training adaptation, and privacy protection operations can be performed simultaneously; in statistical analysis scenarios, statistical analysis, sorting, and other resource integration operations can be combined. Through diverse combination methods, usable resources that both comply with compliance constraints and accurately match the actual needs of data users are generated, providing efficient, secure, and usable data support for subsequent use stages throughout the entire data lifecycle.

[0133] According to an embodiment of the present invention, the method further includes: in step S200, confirming that the use of available resources is non-compliant and providing feedback; in step S210, generating and issuing a new data usage policy based on the feedback of non-compliance and existing data usage policies, or solely based on the feedback of non-compliance; and in step S220, cyclically executing the method starting from generating a data usage authorization result based on the new data usage policy.

[0134] Figure 7 is a flowchart following Figure 1, illustrating the generation and issuance of a new data usage policy when using available resources based on usage authorization results results in non-compliance. Specifically, as shown in Figure 7, this process is a key component of the dynamic adjustment mechanism in the data lifecycle usage control method. It connects the "assessing available resources based on usage authorization results" and "using available resources" steps in Figure 1. Through precise responses to non-compliant usage behaviors, policy iteration, and process loops, it ensures that data usage always aligns with compliance requirements, forming a complete control loop.

[0135] The data usage assessment module, in fulfilling its responsibilities of real-time monitoring and periodic verification, continuously compares the actual usage of available resources with the constraints of the data usage authorization. When an operation deviates from the authorized scope, a non-compliance confirmation process is initiated. The confirmation phase requires verification from multiple dimensions: checking whether the operating entity matches the permitted data user entity in the authorization results, investigating whether the data usage conforms to the agreed scenario, verifying whether the scope of use exceeds the authorization results, and confirming whether the usage environment meets the prescribed standards. This multi-dimensional cross-verification ensures the accuracy and rigor of the non-compliance determination. Upon confirmation of non-compliance, the data usage assessment module initiates a feedback mechanism, submitting a detailed report to the authorization system. The report includes the time of non-compliance, the scope of data involved, the specific dimensions of the violation, the operating entity information, and the risk level. Simultaneously, a compliance warning is sent to the data provider, informing them of the abnormal data usage.

[0136] After receiving non-compliance feedback, the authorization system's data usage policy generation module is responsible for formulating a new data usage policy. The formulation process will select appropriate logic based on the actual scenario, including but not limited to the following: If the non-compliance is due to blind spots in the existing data usage policy, such as the lack of a clear limit on the number of secondary data calls, then targeted constraint clauses will be added based on the existing data usage policy and the reported non-compliance, refining the rules for controlling the scope of data usage or scenario restrictions. If the non-compliance is due to sudden changes in the scenario, such as the data usage environment suddenly failing to meet security standards, and the existing policy lacks corresponding rules, then entirely new policy content will be constructed based solely on the reported non-compliance, such as adding requirements for the frequency of environmental security checks or temporary permission freeze clauses. The new policy must accurately target the root cause of the non-compliance. For cases of data usage exceeding the scope, the new policy can specify threshold monitoring for real-time data call volume and content scope, immediately suspending usage permissions if the threshold is exceeded. For cases of using unauthorized training models, the new policy can refine rules to only allow specified model types and corresponding algorithms, automatically blocking calls to other models.

[0137] Once the new data usage policy is generated, the authorization system's instruction module will synchronously distribute it to the main modules of the data processing system, including the data management module, data usage evaluation module, and data calculation module. It will also be pushed to the remote control module to ensure that all parties involved in data processing and usage receive the latest policy requirements in a timely manner. A clear effective date will be included during the distribution process to avoid a management vacuum during the transition between the old and new policies, ensuring that the new policy officially replaces the old policy from the specified time.

[0138] After the new policy takes effect, the entire control process will cycle through the steps starting from "Generate data usage authorization results based on the data usage policy" (i.e., step S150) in Figure 1. The data management module regenerates usage authorization results containing updated constraints based on the new usage policy, such as adjusting the scope of data usage and supplementing environmental security verification standards. The data preprocessing module re-adapts the associated resources on the user side based on the new authorization results; if the new policy requires a higher level of privacy protection, it supplements the corresponding de-identification operations. The data usage evaluation module updates the evaluation indicators according to the new policy, strengthening the monitoring of previously non-compliant dimensions, such as increasing the real-time sampling frequency of data call volume. The data calculation module adjusts the data usage logic according to the new authorization results to ensure that subsequent data calls, model training, and other operations fully comply with the constraints of the new policy.

[0139] Figure 7 illustrates the dynamic adaptation logic of non-compliance handling and strategy optimization through a progressive process. This process is not a one-off adjustment, but a continuously iterative control mechanism. Each non-compliance feedback becomes a basis for optimizing data usage strategies, ensuring that the strategy content continuously aligns with the compliance requirements of actual use scenarios. This effectively avoids control failures caused by rigid strategies, further solidifying the compliance and security of data usage throughout its entire lifecycle, and ensuring that data usage is always carried out in an orderly manner within the pre-set control framework.

[0140] According to an embodiment of the present invention, the setting forms for the data contract terms reference, the selection of the call status extraction algorithm, the parsing method of the data management strategy, the parsing method of the data usage strategy, and the parsing method of the data usage authorization result include at least one of the following: based on rules, configuration files, buttons, circle, checkmark, marker, key, scroll wheel, menu, voice, video, eye contact, gesture, text, bioelectrical signals, and virtual reality.

[0141] Specifically, these diverse settings cover multiple dimensions such as system configuration, human interaction, and intelligent perception. They can be flexibly adapted to the security level, ease of operation requirements, and operating habits of participating entities in the data flow scenario, ensuring that the operation of key links such as data contract signing, algorithm selection, and strategy and authorization result parsing is both compliant and controllable, and has sufficient flexibility and adaptability.

[0142] Rule-based configuration is suitable for scenarios with high standardization and low reliance on frequent adjustments. For example, references to common terms in data contracts can be based on pre-defined rules, allowing the system to automatically match corresponding standard terms according to data type and transaction size. Data governance strategy parsing can be based on pre-defined rules to clarify field decomposition logic, ensuring consistency in strategy parsing across different scenarios. Configuration files, on the other hand, are primarily used for scenarios requiring batch configuration or backend adjustments. For instance, when calling a status extraction algorithm, different extraction algorithms for different data types can be specified through the configuration file, allowing algorithm switching without modifying the main code. The parsing method for data usage authorization results can also be defined through configuration files, defining field mapping relationships to adapt to the system interface requirements of different data users.

[0143] Buttons, checkboxes, and menus prioritize ease of human interaction and are suitable for stages requiring participants to make their own choices or confirmations. For example, during data contract signing, data users can select the desired data type, usage period, and other terms via a checkbox menu; the data usage strategy parsing method can be switched between different parsing modes via buttons to adapt to strategies of varying complexity. Circles and markers, on the other hand, are suitable for scenarios requiring precise targeting, such as circling specific clauses in a data contract document as references, or marking key constraints in a data management strategy to facilitate accurate subsequent system parsing.

[0144] Button and scroll wheel options are adapted to physical device operation scenarios. For example, physical buttons can trigger confirmation commands for data contract terms, or scroll wheels can adjust threshold parameters for data access volume in data usage authorization results, meeting the needs of physical device operation in certain scenarios. Text-based options are the most basic setting method, suitable for scenarios requiring detailed descriptions, such as entering custom terms in data contracts or providing textual explanations of special parsing rules in data management strategies, ensuring the accurate transmission of complex requirements.

[0145] Voice and video formats are geared towards remote interaction or contactless operation scenarios. For example, when signing data contracts, participants can confirm the reference of terms through voice commands or complete the verification and confirmation of terms through video conferencing. The parsing method of data usage strategies can be switched through voice commands, improving operational convenience. Eye contact and gestures are suitable for intelligent interaction scenarios. For example, in high-security operating environments, after verifying identity through eye contact, gestures can be used to select the state extraction algorithm; or in virtual reality scenarios, gestures can be used to select data contract terms and adjust strategy parsing parameters, enhancing the intuitiveness of the interaction.

[0146] Bioelectric signals and virtual reality are suitable for high-security or complex scenarios. For example, in the signing of contracts involving sensitive data, the identity of the participants can be verified through bioelectric signals, and then the terms can be visualized and confirmed in a virtual reality environment. At the same time, the data management strategy parsing method can be selected through a virtual interactive interface, which not only ensures operational security, but also improves the efficiency of understanding and setting complex rules.

[0147] These settings are not isolated and can be combined according to actual scenarios. For example, when signing a data contract, standard terms can be selected via a menu, custom content can be added via text, and confirmation can be made via voice. When selecting a status extraction algorithm, a default algorithm can be preset via a configuration file, and then the algorithm type can be manually switched for specific scenarios via a button. This diverse range of settings ensures that participants with different technical skills, operating environments, and security requirements can efficiently complete the setup operations for key stages, supporting the smooth progress of the data lifecycle usage control process.

[0148] Figure 8 is a block diagram of the data lifecycle usage control system provided by the present invention. As shown in Figure 8, the present invention also provides a data lifecycle usage control system 100, including: a contract signing system 110, a provider-side associated resource generation module 120, a transmission resource generation module 130, a user-side associated resource generation module 140, an authorization system 150, and a data processing system 160; the contract signing system 110 is used for the data user to initiate a usage request to the data provider, and the data user and the data provider sign a data contract; the provider-side associated resource generation module 120 is used to generate provider-side associated resources from data provider-side resources according to data management strategies and / or data contracts; the transmission resource generation module 130 is used to generate transmission resources according to provider-side associated resources and data management strategies; the user-side associated resource generation module 140 is used to generate user-side associated resources according to data usage authorization results and transmission resources; the authorization system 150 includes: a status acquisition module 151, used to call a status extraction algorithm to extract a status set; a data control strategy generation module 152, used to generate a data management strategy based on the status set; a data usage strategy generation module 153, used to generate a data usage strategy based on the status set; and a remote... The control module 154 is used by the data provider to remotely verify the data user's use of associated resources according to the data contract; the instruction issuing module 155 is used to issue the data usage authorization result; the processing system includes: a data management module 161, used to generate data usage authorization results according to the data usage strategy; a data preprocessing module 162, used to preprocess the associated resources on the user side based on the data usage authorization result to generate usable resources; or to use the associated resources on the user side as usable resources based on the data usage authorization result; a data usage evaluation module 163, used to evaluate and / or constrain the use of usable resources based on the usage authorization result; a data calculation module 164, used by the data user to use usable resources based on the usage authorization result; and also used to use the associated resources on the user side as usable resources based on the data usage authorization result; the data management module 161 is also used to send the data usage authorization result to the data preprocessing module 162, the data usage evaluation module 163, and the data calculation module 164; the data usage evaluation module 163 is also used to provide feedback on non-compliant data usage to the authorization system 150 and the data provider.

[0149] According to an embodiment of the present invention, the remote control module 154 is used to receive information from the data usage evaluation module 163 and feed it back to the data control policy generation module 152 and the data usage policy generation module 153; the data control policy generation module 152 is further used to generate a new data control policy based on the non-compliance feedback from the data usage evaluation module 163 and the existing data usage policy, or only based on the non-compliance feedback from the data usage evaluation module 163; the data usage policy generation module 153 is further used to generate a new data usage policy based on the non-compliance feedback from the data usage evaluation module 163 and the existing data usage policy, or only based on the non-compliance feedback from the data usage evaluation module 163; the instruction issuing module 155 is used to issue the new data usage policy to the data control module 152 and the data usage policy generation module 153; The system comprises a data management module 161, a data usage evaluation module 163, a data calculation module 164, and a remote control module 154. An instruction issuance module 155 is also used to issue new data management policies to the provider-side associated resource generation module 120, the transmission resource generation module 130, and the user-side associated resource generation module 140. The data management module 161 generates new data usage authorization results based on the new data usage authorization results. The data processing system 160 begins a new round of data processing and calculation based on the new data usage authorization results. This new round of data processing includes processing data resources based on the data usage authorization results received by the data preprocessing module 162. The new round of data calculation includes updating the data resources used in the calculation task based on the data usage authorization results received by the data calculation module 164, and performing task calculations.

[0150] According to an embodiment of the present invention, non-compliance includes one or more of the following: data user, data purpose, data usage environment, data usage time limit, data usage scope, training model type and training algorithm, data transmission protection algorithm and its parameters, and data storage protection algorithm and its parameters.

[0151] According to an embodiment of the present invention, the working mode and calculation result usage mode of each module of the data lifecycle use control system 100 and authorization system 150 include at least one of the following: based on rules, configuration files, buttons, circle, checkmark, mark, key, pull wheel, menu, voice, video, eye contact, gesture, text, bioelectrical signals, and virtual reality.

[0152] The functions of each module of the data lifecycle use control system 100, namely the contract signing system 110, the provider-side associated resource generation module 120, the transmission resource generation module 130, the user-side associated resource generation module 140, the authorization system 150 and its subordinate status acquisition module 151, the data control strategy generation module 152, the data usage strategy generation module 153, the remote control module 154, the instruction issuance module 155, and the data processing system 160 and its subordinate data management module 161, the data preprocessing module 162, the data usage evaluation module 163, and the data calculation module 164, correspond one-to-one with the steps of the data lifecycle use control method recorded in the document. The two can be referred to each other, and the relevant content will not be repeated here.

[0153] Figure 9 is a schematic diagram of the structure of the electronic device provided by the present invention. As shown in Figure 9, the electronic device may include: a processor 910, a communication interface 920, a memory 930 and a communication bus 940, wherein the processor 910, the communication interface 920 and the memory 930 communicate with each other through the communication bus 940. The processor 910 can invoke logic instructions in the memory 930 to execute a data lifecycle usage control method. This method includes: a data user initiating a usage request to a data provider, the data user and the data provider signing a data contract; invoking a state extraction algorithm to extract a state set, and generating a data management strategy and a data usage strategy based on the state set; generating provider-side associated resources from data provider-side resources according to the data management strategy and / or the data contract; generating transmission resources according to the provider-side associated resources and the data management strategy; generating a data usage authorization result according to the data usage strategy; generating user-side associated resources according to the data usage authorization result and the transmission resources; preprocessing the user-side associated resources based on the data usage authorization result to generate usable resources; or using the user-side associated resources as usable resources based on the data usage authorization result; evaluating and / or constraining the usage of the usable resources based on the usage authorization result; and using the usable resources based on the usage authorization result.

[0154] This invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements a data lifecycle usage control method provided by the methods described above. This method includes: a data user initiating a usage request to a data provider; the data user and the data provider signing a data contract; invoking a state extraction algorithm to extract a state set, and generating a data management strategy and a data usage strategy based on the state set; generating provider-side associated resources from data provider-side resources according to the data management strategy and / or the data contract; generating transmission resources according to the provider-side associated resources and the data management strategy; generating a data usage authorization result according to the data usage strategy; generating user-side associated resources according to the data usage authorization result and the transmission resources; preprocessing the user-side associated resources based on the data usage authorization result to generate usable resources; or using the user-side associated resources as usable resources based on the data usage authorization result; evaluating and / or constraining the usage of the usable resources based on the usage authorization result; and using the usable resources based on the usage authorization result.

[0155] This invention also provides a computer program product, comprising a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the data lifecycle usage control method provided by the above-described methods. This method includes: a data user initiating a usage request to a data provider; the data user and the data provider signing a data contract; invoking a state extraction algorithm to extract a state set, and generating a data management strategy and a data usage strategy based on the state set; generating provider-side associated resources from data provider-side resources according to the data management strategy and / or the data contract; generating transmission resources according to the provider-side associated resources and the data management strategy; generating a data usage authorization result according to the data usage strategy; generating user-side associated resources according to the data usage authorization result and the transmission resources; preprocessing the user-side associated resources based on the data usage authorization result to generate usable resources; or using the user-side associated resources as usable resources based on the data usage authorization result; evaluating and / or constraining the usage of the usable resources based on the usage authorization result; and using the usable resources based on the usage authorization result.

[0156] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for controlling the use of data throughout its entire lifecycle, characterized in that, The process includes the following steps: the data user initiates a usage request to the data provider, and the data user and the data provider sign a data contract; a state extraction algorithm is invoked to extract a state set, and a data management strategy and a data usage strategy are generated based on the state set; The process of generating provider-side associated resources from data provider-side resources according to the data management strategy and / or the data contract specifically includes: the data provider-side resources including data provider-side main resources and existing associated resources; searching for data required by the data contract from the data provider-side main resources and / or existing associated resources, and using this data as the data provider-side associated resource; confirming that the data required by the data contract does not exist, generating the data required by the data contract based on the data provider-side main resources and / or existing associated resources, and using this data as the data provider-side associated resource; generating transmission resources based on the provider-side associated resources and the data management strategy; generating data usage authorization results based on the data usage strategy; generating usage-side associated resources based on the data usage authorization results and the transmission resources; preprocessing the usage-side associated resources based on the data usage authorization results to generate usable resources; or using the usage-side associated resources as usable resources based on the data usage authorization results; evaluating and / or constraining the usage of the usable resources based on the usage authorization results; and using the usable resources based on the usage authorization results.

2. The method according to claim 1, characterized in that, The data management strategy includes any combination of one or more of the following: data subject, data type, data content, data transmission protection algorithm and its parameters, data transmission method, data transmission path, data transmission subject, data transmission time, and data receiving subject.

3. The method according to claim 1, characterized in that, The data usage strategy includes any combination of one or more of the following: data receiving entity, data protection capability of the data receiving entity, data user entity, data protection capability of the data user, data storage protection algorithm and its parameters, data purpose, data usage environment, data usage time limit, data usage scope, data destruction method after use, training model type and training algorithm, and post-training purpose.

4. The method according to claim 1, characterized in that, The data usage authorization results include one or more of the following: allowed data transmission protection algorithms and their parameters, data storage protection algorithms and their parameters, data usage environment, data usage time limit, data user, data user's data protection capabilities, data purpose, data usage scope, training model type, and training algorithm.

5. The method according to claim 1, characterized in that, The preprocessing of the user-side associated resources to generate usable resources specifically includes: performing any combination of one or more of the following operations on the user-side associated resources: model training, privacy protection, statistical analysis, denoising, sorting, completion, and fusion with other user-side associated resources, to generate data that meets the needs of the data user as the usable resources.

6. The method according to claim 1, characterized in that, The method further includes: confirming that the use of the available resources is non-compliant and providing feedback; generating and issuing a new data usage policy based on the feedback of non-compliance and existing data usage policies, or solely based on the feedback of non-compliance; and cyclically executing the process starting from the generation of data usage authorization results based on the new data usage policy.

7. The method according to claim 1, characterized in that, The setting formats for the data contract terms reference, the selection of the call status extraction algorithm, the parsing method of the data management strategy, the parsing method of the data usage strategy, and the parsing method of the data usage authorization result include at least one of the following: rule-based, configuration file, button, circle, checkmark, mark, key, scroll wheel, menu, voice, video, eye contact, gesture, text, bioelectrical signal, and virtual reality.

8. A data lifecycle usage control system, characterized in that, include: The system includes a contract signing system, a provider-side associated resource generation module, a transmission resource generation module, a user-side associated resource generation module, an authorization system, and a data processing system. The signing system is used for data users to initiate a usage request to data providers, and the data users and data providers sign a data contract; the provider-side associated resource generation module is used to generate provider-side associated resources from data provider-side resources according to data management policies and / or the data contract; the transmission resource generation module is used to generate transmission resources according to the provider-side associated resources and the data management policies; The user-side associated resource generation module is used to generate user-side associated resources based on the data usage authorization result and the transmission resources; The authorization system includes: a state acquisition module for extracting a state set by calling a state extraction algorithm; a data control strategy generation module for generating a data management strategy based on the state set; a data usage strategy generation module for generating a data usage strategy based on the state set; a remote control module for the data provider to remotely verify the data user's use of the associated resources on the user side in accordance with the data contract; and an instruction issuance module for issuing the data usage authorization result. The processing system includes: a data management module for generating a data usage authorization result based on the data usage strategy; and a data preprocessing module for preprocessing the associated resources on the user side based on the data usage authorization result to generate... The system includes: a data usage assessment module for evaluating and / or constraining the use of the available resources based on the data usage authorization results; a data calculation module for the data user to use the available resources based on the data usage authorization results; and a data management module for sending the data usage authorization results to the data preprocessing module, the data usage assessment module, and the data calculation module; and a data usage assessment module for providing feedback on non-compliant data usage to the authorization system and the data provider.

9. The system according to claim 8, characterized in that, The remote control module is used to receive information from the data usage evaluation module and feed it back to the data control policy generation module and the data usage policy generation module; the data control policy generation module is further used to generate a new data management policy based on the non-compliance feedback from the data usage evaluation module and the existing data usage policy, or only based on the non-compliance feedback from the data usage evaluation module; the data usage policy generation module is further used to generate a new data usage policy based on the non-compliance feedback from the data usage evaluation module and the existing data usage policy, or only based on the non-compliance feedback from the data usage evaluation module; the instruction issuing module is used to issue the new data usage policy to the data management module and the data usage policy generation module. The system comprises a data usage evaluation module, a data calculation module, and a remote control module; the instruction issuing module is further configured to issue the new data management strategy to the provider-side associated resource generation module, the transmission resource generation module, and the user-side associated resource generation module; the data management module generates a new data usage authorization result based on the new data usage strategy; the data processing system begins a new round of data processing and calculation based on the new data usage authorization result; the new round of data processing includes processing data resources based on the data usage authorization result received by the data preprocessing module; the new round of data calculation includes updating the data resources used in the calculation task based on the data usage authorization result received by the data calculation module, and performing task calculation.

10. The system according to claim 8, characterized in that, The non-compliance includes one or more of the following: data user, data purpose, data usage environment, data usage time period, data usage scope, type of training model and training algorithm, data transmission protection algorithm and its parameters, and data storage protection algorithm and its parameters.

11. The system according to claim 8, characterized in that, The data lifecycle usage control system and the authorization system's modules' working methods and calculation results usage methods include at least one of the following: rules, configuration files, buttons, circle, checkmark, mark, key, scroll wheel, menu, voice, video, eye contact, gesture, text, bioelectrical signals, and virtual reality.

12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the data lifecycle use control method as described in any one of claims 1 to 7.

13. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements all or part of the steps of the data lifecycle use control method as described in any one of claims 1 to 7.

14. A computer program product, said computer program product comprising computer-executable instructions, characterized in that, When executed, the instructions are used to implement all or part of the steps of the data lifecycle use control method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data space infrastructure-oriented data use control negotiation method

    CN119004427A

  • File full life cycle management system and method based on cloud computing

    CN120045520A