Data processing method and device, electronic equipment, storage medium and product

By generating watermarks using a resistance network model and embedding them into scientific research data, combined with a trusted execution environment and smart contracts, the problems of watermark imperceptibility and insufficient robustness in traditional encryption algorithms are solved, and the security and legitimacy of scientific research data are guaranteed when shared.

CN120654247APending Publication Date: 2025-09-16PIPECHINA SOUTH CHINA CO +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510734904.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In the existing technology, the encrypted watermarks generated by traditional encryption algorithms have poor imperceptibility and robustness, and cannot resist model reverse engineering attacks, resulting in poor security of scientific research data.

Method used

A generative resistance network model is used to generate watermark information and embed it into scientific research data. By building a trusted execution environment and smart contracts, the security and legitimacy of data sharing are ensured.

Benefits of technology

The imperceptibility and robustness of the watermark are improved, which can resist model reverse engineering attacks and ensure the security and legitimacy of scientific research data when it is shared.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654247A_ABST
    Figure CN120654247A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method and device, electronic equipment, a storage medium and a product. The method comprises the following steps: inputting attribute information corresponding to scientific research data into a pre-trained watermark generation model to obtain watermark information corresponding to the scientific research data; wherein the watermark generation model comprises generation of a resistance network model; embedding the watermark information into the scientific research data to obtain encrypted data; and performing data processing on the encrypted data based on a data sharing request sent by a data sharing party. According to the technical scheme of the embodiment of the invention, the imperceptibility and robustness of the watermark can be improved, the reverse engineering attack of the model can be resisted, and the data security of the scientific research data during sharing is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of data processing technology, and in particular to a data processing method, device, electronic device, storage medium, and product. Background Art

[0002] Oil and gas storage and transportation companies have accumulated a large amount of high-value data in scientific research scenarios such as pipeline corrosion analysis, tank leakage prediction, and pipeline blasting experiments. How to safely protect the accumulated scientific research data has become a very important issue at present.

[0003] In the prior art, traditional encryption algorithms and the timestamp of scientific research data acquisition are typically used to generate encrypted watermarks, which are then used to encrypt scientific research data. However, during the implementation of the present invention, it was discovered that the prior art suffers from at least the following technical issues: encrypted watermarks generated using traditional encryption algorithms are imperceptible and poorly robust, making them vulnerable to model reverse engineering attacks and resulting in poor scientific research data security. Summary of the Invention

[0004] The embodiments of the present invention provide a data processing method, device, electronic device, storage medium and product to improve the imperceptibility and robustness of watermarks, resist model reverse engineering attacks, and achieve the purpose of ensuring the data security of scientific research data when shared.

[0005] According to one aspect of the present invention, there is provided a data processing method, comprising:

[0006] Inputting attribute information corresponding to the scientific research data into a pre-trained watermark generation model to obtain watermark information corresponding to the scientific research data; wherein the watermark generation model includes a generative resistance network model;

[0007] Embedding the watermark information into the scientific research data to obtain encrypted data;

[0008] Based on the data sharing request sent by the data sharing party, data processing is performed on the encrypted data.

[0009] According to another aspect of the present invention, there is provided a data processing apparatus, comprising:

[0010] An information input module is used to input attribute information corresponding to scientific research data into a pre-trained watermark generation model to obtain watermark information corresponding to the scientific research data; wherein the watermark generation model includes a generative resistance network model;

[0011] An information embedding module, used for embedding the watermark information into the scientific research data to obtain encrypted data;

[0012] The data processing module is used to process the encrypted data based on the data sharing request sent by the data sharing party.

[0013] According to another aspect of the present invention, an electronic device is provided, comprising:

[0014] at least one processor; and

[0015] a memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the data processing method described in any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the data processing method according to any embodiment of the present invention when executed.

[0018] According to another aspect of the present invention, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the data processing method according to any embodiment of the present invention.

[0019] The technical solution of the embodiment of the present invention obtains watermark information corresponding to the scientific research data by inputting attribute information corresponding to the scientific research data into a pre-trained watermark generation model; wherein the watermark generation model includes generating a resistance network model; therefore, obtaining watermark information by generating a resistance network model is conducive to improving the imperceptibility and robustness of the watermark, and can resist model reverse engineering attacks; then, the watermark information is embedded in the scientific research data to obtain encrypted data; finally, the encrypted data is processed based on a data sharing request sent by the data sharing party, thereby ensuring the data security of the scientific research data when it is shared.

[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 is a flow chart of a data processing method provided according to an embodiment of the present invention;

[0023] Figure 2 is a flowchart of another data processing method provided according to an embodiment of the present invention;

[0024] Figure 3 is a flow chart of a data authentication process provided according to an embodiment of the present invention;

[0025] Figure 4 This is an architectural diagram of a scientific research data asset management and control system provided according to an embodiment of the present invention;

[0026] Figure 5 is a structural diagram of a data processing device provided according to an embodiment of the present invention;

[0027] Figure 6 It is a structural diagram of an electronic device for implementing the data processing method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "etc." and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0030] It should be noted that the collection, collection, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solutions disclosed herein comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken with respect to user personal information to prevent unauthorized access to user personal information data and maintain the security of user personal information and network security.

[0031] Figure 1 1 is a flow chart of a data processing method according to an embodiment of the present invention. This embodiment is applicable to processing scientific research data. The method can be executed by a data processing device, which can be implemented in the form of hardware and / or software.

[0032] like Figure 1 As shown, the method of this embodiment may specifically include:

[0033] S110. Input the attribute information corresponding to the scientific research data into a pre-trained watermark generation model to obtain the watermark information corresponding to the scientific research data; wherein the watermark generation model includes a generative resistance network model.

[0034] The scientific research data includes oil and gas storage and transportation research data. Attribute information reflects information such as the source, generation time, and sharing permissions of the scientific research data. Sharing permissions reflect the type of permissions that can be used when sharing the scientific research data. For example, permission types include at least one of read type, derivative type, and training model type. Furthermore, attribute information may include information such as the scientific research project number and version number.

[0035] In this embodiment, for scientific research data that needs to be encrypted, the attribute information of the scientific research data can be determined, and the attribute information can be used as input information and input into a pre-trained watermark generation model. Among them, the watermark generation model is a generative adversarial network (GAN) model. The generative adversarial network model includes a generative network and a discriminant network. The generative network can adopt the Transformer model of the attention mechanism and the U-Net structure to generate candidate watermark information; the discriminant network can adopt the VGG19 structure to evaluate the imperceptibility and robustness of the candidate watermark information. Through multiple rounds of adversarial training, the parameters of the generative network are optimized, so that the watermark information obtained by the generative adversarial network model has good imperceptibility and robustness, thereby enhancing its ability to resist model reverse engineering attacks.

[0036] Optionally, the attribute information corresponding to the scientific research data is input into a pre-trained watermark generation model to obtain the watermark information corresponding to the scientific research data, including: taking at least one of the sovereignty identification, permission label and generation timestamp of the scientific research data as the attribute information of the scientific research data; inputting the attribute information into the pre-trained watermark generation model, and determining the watermark information corresponding to the scientific research data based on the output result of the watermark generation model.

[0037] It should be noted that the sovereignty identifier is the identifier corresponding to the creator of the scientific research data, and the permission label is used to reflect the sharing permissions that can be provided when the scientific research data is shared. Among them, the permission label includes at least one of the read right label, the training right label and the derivative right label; the read label indicates that the scientific research data is only for reading, for example, to view the statistical summary of the data. The training right label is used to identify that the scientific research data can be used as model training data for training the model to determine the model weights. The derivative right label is used to indicate that the scientific research data can be further analyzed and processed to obtain other data. As for the income from the data newly derived from the scientific research data, it is necessary to share the income with the provider of the scientific research data.

[0038] By inputting the sovereignty identifier, permission label and generation timestamp as attribute information into the watermark generation model, the obtained watermark information can reflect the sovereignty identifier, permission label and generation timestamp, preventing data from being tampered with; and, generating watermarks based on permission labels facilitates dynamic adjustment of watermarks to ensure the legality and security of data access.

[0039] Through adversarial training, when the discriminant network cannot distinguish between watermarked and original scientific research data, the watermark is considered to be imperceptible. Furthermore, when the scientific research data is an image, to verify the robustness of the watermark, the watermarked image can be subjected to processing such as adding Gaussian noise, image compression, rotation, and cropping. If the watermark can still be correctly extracted from the processed image, the watermark is considered to be robust.

[0040] S120. Embed the watermark information into the scientific research data to obtain encrypted data.

[0041] In this embodiment, watermark information can be embedded in the scientific research data in different ways based on the data type of the scientific research data, and the scientific research data with embedded watermark information can be used as encrypted data. Optionally, specific implementations of embedding watermark information in the scientific research data to obtain encrypted data include: if the scientific research data is structured data, embedding the watermark information in at least one of the least significant bits, singular value decomposition features, and hash values ​​of the scientific research data; or, if the scientific research data is unstructured data, embedding the watermark information in the frequency domain features of the unstructured data.

[0042] Specifically, scientific research data can be divided into structured data and unstructured data. For example, structured data may include tables; unstructured data may include text, audio, and other data.

[0043] If the scientific research data is determined to be structured data, such as a pipeline corrosion rate table, the watermark information can be embedded into the least significant bit, singular value decomposition feature, or hash value of the table. Specific embedding technologies can include digital watermarking and zero watermarking.

[0044] When it is determined that the scientific research data is unstructured data, such as pipeline ultrasonic images, the watermark information can be embedded into the frequency domain features of the image; for example, the frequency domain features include discrete cosine transform coefficients, wavelet transform coefficients, etc.

[0045] This embodiment sets different watermark information embedding methods for different data types of scientific research data, thereby ensuring that watermark information can be effectively embedded into scientific research data for different data.

[0046] S130: Process the encrypted data based on the data sharing request sent by the data sharing party.

[0047] Different data sharing parties are used for different application scenarios. For example, if a company's research data is shared with an external partner, the data sharing party is the external partner; or if the company's research data is shared with an internal subsidiary, the data sharing party is the internal subsidiary. A data sharing request reflects the type of sharing operation and the data sharing party. The sharing operation type can include at least one of a read operation, a derivative operation, and a model training operation.

[0048] In a specific implementation, a received data sharing request can be parsed to determine the data sharing party's information and the type of sharing operation. Based on the data sharing party's information, the sharing permissions granted to the data sharing party are determined. If the sharing permissions match the type of sharing operation, the encrypted data can be processed according to the sharing operation requested in the data sharing request. If the sharing permissions do not match the type of sharing operation, the encrypted data is not processed, and an alert is generated and sent to the data sharing party, allowing the data sharing party to adjust the requested sharing operation based on the alert.

[0049] Furthermore, after the watermark information is determined, it can be written into the sovereign chain. When the data sharing party calls it, it is dynamically bound to the scientific research data in an encrypted form.

[0050] The technical solution of the embodiment of the present invention obtains watermark information corresponding to the scientific research data by inputting attribute information corresponding to the scientific research data into a pre-trained watermark generation model; wherein the watermark generation model includes generating a resistance network model; therefore, obtaining watermark information by generating a resistance network model is conducive to improving the imperceptibility and robustness of the watermark, and can resist model reverse engineering attacks; then, the watermark information is embedded in the scientific research data to obtain encrypted data; finally, the encrypted data is processed based on a data sharing request sent by the data sharing party, thereby ensuring the data security of the scientific research data when it is shared.

[0051] Figure 2 This is a flowchart of another data processing method provided according to an embodiment of the present invention. Based on the above embodiment, this embodiment optionally processes the encrypted data based on the data sharing request sent by the data sharing party, including: determining the sharing authority corresponding to the data sharing party; based on the data sharing request and the sharing authority, performing the data processing operation corresponding to the data sharing request on the encrypted data in a pre-built trusted execution environment. The explanations of the terms that are the same as or corresponding to the above embodiments are not repeated here. Figure 2 As shown, the method includes:

[0052] S210. Input the attribute information corresponding to the scientific research data into a pre-trained watermark generation model to obtain the watermark information corresponding to the scientific research data; wherein the watermark generation model includes generating a resistance network model.

[0053] In this embodiment, the scientific research data includes model parameters of a stress prediction model used to determine tank stress change information; before inputting the attribute information corresponding to the scientific research data into a pre-trained watermark generation model, it also includes: receiving training gradients of the stress prediction model sent by at least two data providers; wherein the training gradient is information after adding noise based on differential privacy technology; based on the at least two training gradients, performing a model aggregation operation, and obtaining the model parameters of the stress prediction model after aggregation.

[0054] It should be noted that when multiple subsidiaries collaborate on scientific research, such as jointly training a stress prediction model for analyzing tank stress change data, federated learning can be used to determine the model parameters of the stress prediction model to protect the data sovereignty of all parties. Each subsidiary can serve as a data provider.

[0055] Specifically, each data provider trains a sub-model corresponding to the stress prediction model based on local data. For example, each data provider pre-deploys a secure sandbox, where they train the sub-model and obtain the training gradient corresponding to the sub-model, which serves as the training gradient for the stress prediction model.

[0056] In order to improve the security of the training gradient, noise can be added to the training gradient based on differential privacy technology, and the training gradient can be updated based on the information after the noise is added. The updated training gradient sent by each data provider is received, and the model aggregation operation is performed based on at least two training gradients. After aggregation, the model parameters of the stress prediction model are obtained, and then the federated learning is completed. The obtained model parameters are used as scientific research parameters, and watermark information is determined for the model parameters. Since the model parameters are obtained by multiple data providers, multi-party joint watermark information can be generated based on the identification of each data provider, and the multi-party joint watermark information can be added to the model parameters for subsequent tracking. Furthermore, a smart contract can be set in advance, and the contribution of each data provider to the model parameters can be determined, so that the model benefits can be provided to each data provider in proportion through the smart contract.

[0057] S220: embed the watermark information into the scientific research data to obtain encrypted data.

[0058] S230. Determine the sharing authority corresponding to the data sharing party; based on the data sharing request and the sharing authority, perform a data processing operation corresponding to the data sharing request on the encrypted data in a pre-built trusted execution environment.

[0059] In this embodiment, different data sharing parties have different sharing permissions. For example, data sharing party A has read permissions, while data sharing party B has model training permissions and derivative permissions. To ensure that data processing operations on encrypted data comply with permission requirements and prevent encrypted data from being leaked or maliciously manipulated, encrypted data can be processed based on sharing permissions.

[0060] Specifically, the sharing permissions granted to the data sharing party are determined based on the data sharing party's information in the data sharing request. To ensure security during data processing, a trusted execution environment (i.e., a secure sandbox) can be pre-built to restrict data processing to encrypted memory, preventing malicious theft or tampering. If the sharing permissions match the requested operation in the data sharing request, data processing on the encrypted data is ensured within the trusted execution environment.

[0061] Furthermore, data usage traces, such as the gradient update path of the model trainer, can be recorded to generate an unalterable chain of evidence, ensuring traceability and transparency during data usage. When performing model training and data analysis in a secure sandbox, operation logs or the gradient update path of each round are recorded and hashed to form an operation and training evidence chain.

[0062] This embodiment builds a trusted execution environment (TEE) to support data processing in an encrypted environment and record data usage on an immutable sovereign blockchain, effectively preventing malicious theft or tampering, thereby enhancing data security and immutability. This TEE enables data to be "available but invisible," supporting multi-party federated learning collaboration based on differential privacy, thereby promoting data sharing and the application of scientific research results.

[0063] In this embodiment, after performing the data processing operation corresponding to the data sharing request on the encrypted data, it also includes: based on the data processing operation, updating the pre-built bloodline map and recording the data information of the encrypted data after the data processing operation; based on the data information, determining the contribution of the data sharing party to the encrypted data; based on the updated bloodline map and contribution, driving the smart contract to determine the derivative value attributes generated by the data sharing party for the encrypted data.

[0064] It's important to note that a data lineage graph can be constructed based on a graph database. This database records the data provided by each data provider and data processing information. This data processing information includes the provided data and the processing operations performed on the data. A lineage graph consists of nodes and edges, where nodes represent data and / or models, and edges represent processing relationships. The lineage graph also records metadata for each node, including the corresponding data volume and quality score. For example, if pipeline A's corrosion data is processed by a corrosion prediction model to generate Report A, the lineage graph would be "Pipeline A Corrosion Data → Corrosion Prediction Model → Report."

[0065] In this embodiment, data processing operations may contribute to the scientific research process. Therefore, based on the data processing operations, the lineage graph can be updated. Specifically, nodes in the lineage graph are updated to record the data after the data processing operations; edges in the lineage graph are also updated to update the data processing relationships. Furthermore, data information generated by the data processing operations can also be recorded. This data information includes at least one of the following: the amount of encrypted data after the data processing operations, a quality score, and model performance.

[0066] In a specific implementation, the contribution of the data sharing party to the encrypted data can be determined based on the data information. For example, the contribution is proportional to at least one of the data volume, quality score, and model performance index.

[0067] Exemplarily, based on the updated lineage map and contribution information, the smart contract is driven to determine the derived value attributes generated by the data sharing party for the encrypted data. This includes: Based on the lineage map and contribution information, the smart contract is driven to automatically assign the derived value attributes. For example, when the report is commercialized, 10% of the model's derived value attributes will be attributed to the original data provider. Alternatively, in an Ethereum-based smart contract, the contract code contains the contribution calculation formula and the derived value attribute allocation rules. When new derived data or models are uploaded to the blockchain, the smart contract automatically calculates the derived value attribute value corresponding to the contribution. The smart contract encodes the contribution calculation formula and the derived value attribute allocation rules.

[0068] For example, the derived value attribute allocation rule for Model 1 is: Data Provider A's derived value attribute value = Total Value Attribute Value × Contribution Ratio, where the Contribution Ratio is the ratio of Data Provider A's contribution to the total contribution value, and the total contribution value is the sum of the contributions of all data providers in Model 1. Alternatively, the derived value attribute value = Preset Use Value Attribute Value + α × Contribution + β × Total Derivative Data Value; where α and β are proportional coefficients, determined by the smart contract based on the data lineage map. The total value of derived data is the total value generated by the derived data.

[0069] Furthermore, watermark information can be used to verify whether scientific research data is used in an illegal manner. Figure 3 As shown, when it is detected that the model is used by a third party without authorization, the model weights can be extracted from the model disclosed by the third party, and the model weights can be decoded to obtain watermark information; the sovereignty chain can be queried to obtain the original data information, such as the recorded data processing link, and an evidence chain can be generated through the original data information for use in rights protection.

[0070] In this embodiment, a data lineage map is constructed through a graph database to record the data processing chain and the contribution of each party, and drive the smart contract to automatically distribute derivative benefits according to the contribution. This solves the problem of loss of rights and interests of the original data contributors, and thus resolves the contradiction between data sharing and rights protection. It effectively breaks the data silo state, significantly improves the data collaboration efficiency in the oil and gas industry, and promotes the transformation of scientific research results into industrialization.

[0071] The data processing method is described in detail above. To make the technical solution of this method clearer to those skilled in the art, a specific application scenario is given below. The data processing method provided in this embodiment can be applied to the scientific research data asset management and control system. The architecture diagram of this system is shown in FIG. Figure 4 As shown, it includes the following four environments:

[0072] A. Support environment for the assetization of scientific research data: including scientific research data security access module, data management, intelligent data search and use portal module and user management module. Among them, the A-1 scientific and technological data security access module realizes the access of the external network to the internal network, and opens up the way for internal and external scientific research projects, field test monitoring and other data to access the internal network securely. The A-2 scientific and technological project data management module establishes the scientific research process business management data ledger and scientific research data metadata of various scientific research projects, and provides basic data query for scientific research data. The A-3 data intelligent search and use module builds an intelligent index of scientific research data assets in the scientific research data asset storage environment, and can provide query and retrieval functions based on the data asset map. The A-4 scientific research data user management module provides management and control of scientific research data-related participants, such as user registration, information authentication and income account allocation.

[0073] B. The research data assetization processing environment provides a B-1 research project data object inventory module, a B-2 research data cleaning module, and a B-3 research data product incubation module within the scope of research project authority. This provides a research project data processing space for researchers to process and clean raw data and incubate assetized products.

[0074] C, the research data asset storage environment, provides a research data asset repository C-1 and a local research data processing repository C-2. C-1 is for storing input and output data for data processing in the research data asset processing environment B. C-2 is for sharing and circulating data asset results in the secure sharing and computing environment C.

[0075] D. Research data asset security sharing and computing environment, which specifically includes: research data asset rights confirmation module, research data security computing module and research data asset operation and circulation module.

[0076] Module D-1: Research Data Asset Ownership Verification Module, primarily implements dynamic watermark embedding and sovereignty anchoring. Dynamic watermarks are generated using a generative adversarial network model, embedding invisible identifiers within the data feature space to resist model distillation attacks. Combined with lightweight zero-knowledge proofs, the watermark verification process remains transparent to the original data.

[0077] Module D-2: Scientific Research Data Security Computing Module, which primarily provides a cross-departmental and cross-unit data sovereignty security computing sandbox environment. This can be done in the following ways:

[0078] Memory Encrypted Execution: Data is decrypted only in the GPU's video memory, while remaining encrypted in the CPU's memory, preventing side-channel attacks. Trusted Execution Space: An isolated, dedicated trusted computing space is established, providing a secure computing environment for third-party data. Computations can be performed on encrypted data. Differential Privacy Injection: Controllable noise is injected before data output, ensuring usability while preventing the original data from being recovered.

[0079] Module D-3: The scientific research data asset operation and circulation module provides data sovereignty smart contracts based on the data sovereignty chain. This allows multiple parties to jointly use data assets while guaranteeing data benefits for all parties. Each participant has different rights and interests in the data, and thus different benefits.

[0080] This embodiment provides a scientific research data asset management and control system that integrates technologies such as dynamic watermarks, sovereign chains, and smart contracts to build a "shareable but non-copyable" circulation system for oil and gas data, thereby improving the efficiency of scientific research data collaboration in the oil and gas industry.

[0081] Figure 5 This is a schematic diagram of the structure of a data processing device provided according to an embodiment of the present invention, which is used to execute the data processing method provided in any of the above embodiments. The device and the data processing methods of the above embodiments belong to the same inventive concept. For details not fully described in the embodiments of the data processing device, please refer to the embodiments of the above data processing methods. Figure 5 As shown, the device includes:

[0082] The information input module 10 is used to input the attribute information corresponding to the scientific research data into the pre-trained watermark generation model to obtain the watermark information corresponding to the scientific research data; wherein the watermark generation model includes a generative resistance network model;

[0083] The information embedding module 11 is used to embed watermark information into scientific research data to obtain encrypted data;

[0084] The data processing module 12 is used to process the encrypted data based on the data sharing request sent by the data sharing party.

[0085] Based on any optional technical solution in the embodiment of the present invention, optionally, the information input module 10 includes:

[0086] The watermark information determination unit is used to use at least one of the sovereignty identification, permission label and generation timestamp of the scientific research data as the attribute information of the scientific research data; wherein the permission label includes at least one of the reading right label, the training right label and the derivative right label; the attribute information is input into the pre-trained watermark generation model, and the watermark information corresponding to the scientific research data is determined based on the output result of the watermark generation model.

[0087] Based on any optional technical solution in the embodiment of the present invention, optionally, the information embedding module 11 includes:

[0088] A first embedding unit is configured to embed watermark information into at least one of the least significant bit, singular value decomposition feature, and hash value of the scientific research data when the scientific research data is structured data; or

[0089] The second embedding unit is used to embed the watermark information into the frequency domain characteristics of the unstructured data when the scientific research data is unstructured data.

[0090] Based on any optional technical solution in the embodiment of the present invention, optionally, the data processing module 12 includes:

[0091] The data processing unit is used to determine the sharing permissions corresponding to the data sharing party; based on the data sharing request and sharing permissions, it performs data processing operations corresponding to the data sharing request on the encrypted data in a pre-built trusted execution environment.

[0092] Based on any optional technical solution in the embodiments of the present invention, optionally, the method further includes:

[0093] An information update module is configured to, after performing a data processing operation corresponding to a data sharing request on the encrypted data, update a pre-built lineage graph and record data information of the encrypted data after the data processing operation based on the data processing operation; wherein the lineage graph is composed of nodes and edges, nodes representing data and / or models, and edges representing processing relationships; and the data information includes at least one of the data volume, quality score, and model performance of the encrypted data after the data processing operation;

[0094] A contribution determination module, used to determine the contribution of the data sharing party to the encrypted data based on the data information;

[0095] The value attribute value determination module is used to drive the smart contract to determine the derived value attributes generated by the data sharing party for the encrypted data based on the updated bloodline map and contribution degree.

[0096] Based on any optional technical solution in the embodiments of the present invention, optionally, the method further includes:

[0097] A data receiving module, configured to receive model parameters of a stress prediction model for determining tank stress variation information from scientific research data; and to receive training gradients of the stress prediction model from at least two data providers before inputting attribute information corresponding to the scientific research data into a pre-trained watermark generation model. The training gradients are information after adding noise based on differential privacy technology.

[0098] The aggregation module is used to perform a model aggregation operation based on at least two training gradients, and obtain model parameters of the stress prediction model after aggregation.

[0099] The technical solution of the embodiment of the present invention obtains watermark information corresponding to the scientific research data by inputting attribute information corresponding to the scientific research data into a pre-trained watermark generation model; wherein the watermark generation model includes generating a resistance network model; therefore, obtaining watermark information by generating a resistance network model is conducive to improving the imperceptibility and robustness of the watermark, and can resist model reverse engineering attacks; then, the watermark information is embedded in the scientific research data to obtain encrypted data; finally, the encrypted data is processed based on a data sharing request sent by the data sharing party, thereby ensuring the data security of the scientific research data when it is shared.

[0100] It is worth noting that in the embodiment of the above-mentioned data processing device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.

[0101] Figure 6 Schematic diagram of the structure of an electronic device that implements the data processing method of an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0102] like Figure 6 As shown, the electronic device 20 includes at least one processor 21, and a memory connected to the at least one processor 21, such as a read-only memory (ROM) 22, a random access memory (RAM) 23, etc., wherein the memory stores a computer program that can be executed by the at least one processor, and the processor 21 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 22 or the computer program loaded from the storage unit 28 to the random access memory (RAM) 23. Various programs and data required for the operation of the electronic device 20 can also be stored in the RAM 23. The processor 21, ROM 22 and RAM 23 are connected to each other via a bus 24. An input / output (I / O) interface 25 is also connected to the bus 24.

[0103] Multiple components in the electronic device 20 are connected to the I / O interface 25, including an input unit 26, such as a keyboard, a mouse, etc.; an output unit 27, such as various types of displays, speakers, etc.; a storage unit 28, such as a magnetic disk, an optical disk, etc.; and a communication unit 29, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 29 allows the electronic device 20 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0104] The processor 21 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 21 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 21 executes the various methods and processes described above, such as the data processing method.

[0105] In some embodiments, the data processing method may be implemented as a computer program that is tangibly contained in a computer-readable storage medium, such as the storage unit 28. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 20 via the ROM 22 and / or the communication unit 29. When the computer program is loaded into the RAM 23 and executed by the processor 21, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, the processor 21 may be configured to perform the data processing method in any other appropriate manner (e.g., by means of firmware).

[0106] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips or systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0107] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0108] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0109] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0110] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0111] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0112] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication unit 29, or installed from the storage unit 28, or installed from the ROM 22. When the computer program is executed by the processor 21, the above-mentioned functions defined in the method of the embodiment of the present invention are performed.

[0113] The computer program product may be implemented by writing computer program code for performing the operations of the present invention in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0114] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0115] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A data processing method, characterized in that: include: Inputting attribute information corresponding to the scientific research data into a pre-trained watermark generation model to obtain watermark information corresponding to the scientific research data; wherein the watermark generation model includes a generative resistance network model; Embedding the watermark information into the scientific research data to obtain encrypted data; Based on the data sharing request sent by the data sharing party, data processing is performed on the encrypted data.

2. The method according to claim 1, characterized in that The method of inputting the attribute information corresponding to the scientific research data into a pre-trained watermark generation model to obtain the watermark information corresponding to the scientific research data includes: Using at least one of the sovereignty identifier, permission label, and generation timestamp of the scientific research data as attribute information of the scientific research data; wherein the permission label includes at least one of a read right label, a training right label, and a derivative right label; The attribute information is input into the pre-trained watermark generation model, and the watermark information corresponding to the scientific research data is determined based on the output result of the watermark generation model.

3. The method according to claim 1, characterized in that The step of embedding the watermark information into the scientific research data to obtain encrypted data includes: In the case where the scientific research data is structured data, the watermark information is embedded into at least one of the least significant bit, singular value decomposition feature and hash value of the scientific research data; or, In the case where the scientific research data is unstructured data, the watermark information is embedded into the frequency domain features of the unstructured data.

4. The method according to claim 1, wherein The processing of the encrypted data based on the data sharing request sent by the data sharing party includes: Determine the sharing permissions corresponding to the data sharing party; Based on the data sharing request and the sharing permission, a data processing operation corresponding to the data sharing request is performed on the encrypted data in a pre-built trusted execution environment.

5. The method according to claim 1, characterized in that After performing the data processing operation corresponding to the data sharing request on the encrypted data, the method further includes: Based on the data processing operation, updating a pre-constructed lineage graph and recording data information of the encrypted data after the data processing operation; wherein the lineage graph is composed of nodes and edges, the nodes represent data and / or models, and the edges represent processing relationships; the data information includes at least one of the data volume, quality score, and model performance of the encrypted data after the data processing operation; Determining, based on the data information, the contribution of the data sharing party to the encrypted data; Based on the updated bloodline map and the contribution degree, the smart contract is driven to determine the derived value attributes generated by the data sharing party for the encrypted data.

6. The method according to claim 1, characterized in that The scientific research data includes model parameters of a stress prediction model for determining the stress change information of the storage tank; before inputting the attribute information corresponding to the scientific research data into the pre-trained watermark generation model, the method further includes: Receiving training gradients of the stress prediction model sent by at least two data providers; wherein the training gradients are information after adding noise based on differential privacy technology; Based on at least two of the training gradients, a model aggregation operation is performed to obtain model parameters of the stress prediction model after aggregation.

7. A data processing device, characterized in that: include: An information input module is used to input attribute information corresponding to scientific research data into a pre-trained watermark generation model to obtain watermark information corresponding to the scientific research data; wherein the watermark generation model includes a generative resistance network model; An information embedding module, used for embedding the watermark information into the scientific research data to obtain encrypted data; The data processing module is used to process the encrypted data based on the data sharing request sent by the data sharing party.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the data processing method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the data processing method according to any one of claims 1 to 6 when executed.

10. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the data processing method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Data access method and device, electronic equipment and computer readable medium

    CN121327877A