Data watermark processing method, device, equipment and readable storage medium

CN122508554APending Publication Date: 2026-08-04CHINA MOBILE SHANGHAI ICT CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE SHANGHAI ICT CO LTD
Filing Date
2026-03-20
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

[0004]本申请实施例提供了一种数据水印处理方法、装置、设备及可读存储介质,能够解决相关技术中数字水印处理对数据溯源效果比较差的技术问题

Benefits of technology

[0017] Fifthly, embodiments of this application provide a computer program product, including computer instructions that, when executed by a processor, implement the steps of the data watermarking processing method as described in the first aspect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122508554A_ABST
    Figure CN122508554A_ABST
Patent Text Reader

Abstract

This application provides a data watermarking processing method, apparatus, device, and readable storage medium. The method includes: acquiring data to be processed, the data to be processed including M field data; performing hierarchical processing on the M field data to obtain N groups of hierarchical data, each group of hierarchical data having a different sensitivity level; employing a watermarking processing strategy corresponding to the sensitivity level of the hierarchical data to watermark each field data in the hierarchical data, obtaining watermarked data for each field data, wherein the watermarking processing intensity of each group of hierarchical data is directly proportional to the sensitivity level corresponding to the hierarchical data; when the sensitivity level corresponding to the hierarchical data is higher than a preset level, the watermarking processing includes desensitization processing, and the desensitization processing intensity of the hierarchical data is directly proportional to the sensitivity level corresponding to the hierarchical data; storing the target data in a blockchain system, the target data including: field data, the sensitivity level of the field data, watermarked data, and the watermarking processing strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data technology, and in particular to a data watermarking processing method, apparatus, device and readable storage medium. Background Technology

[0002] Digital watermarking is a digital signal or pattern that is permanently embedded in the host data and has identifiable characteristics. Digital watermarking can be used to build a traceability and verification system for dynamic monitoring and post-event evaluation.

[0003] Currently, digital watermarking typically involves embedding dynamic watermarks, specifically embedding identifiers within the data flow to enable real-time tracking. However, this method offers relatively coarse-grained verification for source tracing, and the verification result might be "the data flow path has low credibility" or "there is a risk of data disruption"—a qualitative or semi-quantitative assessment. When data tampering is detected, it's difficult to accurately pinpoint which stage or field caused the problem, resulting in poor data source tracing effectiveness. Summary of the Invention

[0004] This application provides a data watermarking processing method, apparatus, device, and readable storage medium, which can solve the technical problem that digital watermarking processing has a poor effect on data traceability in related technologies.

[0005] In a first aspect, embodiments of this application provide a data watermarking method, the method comprising:

[0006] Obtain the data to be processed, which includes M fields, where M is an integer greater than 1;

[0007] The M fields of data are processed in a hierarchical manner to obtain N groups of hierarchical data. Each group of hierarchical data has a different sensitivity level, and N is an integer greater than 1.

[0008] For each group of graded data, a watermarking strategy corresponding to the sensitivity level of the graded data is adopted to watermark each field of the graded data, resulting in watermarked data for each field. The watermarking intensity of each group of graded data is directly proportional to the sensitivity level of the graded data. When the sensitivity level of the graded data is higher than a preset level, the watermarking includes desensitization, and the desensitization intensity of the graded data is directly proportional to the sensitivity level of the graded data.

[0009] The target data is stored in the blockchain system. The target data includes: field data, the sensitivity level of the field data and watermark processing data, and the watermark processing strategy.

[0010] Secondly, embodiments of this application provide a data watermarking processing apparatus, the apparatus comprising:

[0011] The first acquisition module is used to acquire data to be processed, which includes M fields, where M is an integer greater than 1.

[0012] The hierarchical processing module is used to perform hierarchical processing on the M field data to obtain N groups of hierarchical data. Each group of hierarchical data has a different sensitivity level, and N is an integer greater than 1.

[0013] The watermarking module is used to apply a watermarking strategy corresponding to the sensitivity level of the hierarchical data to each group of hierarchical data, and to watermark each field of the hierarchical data to obtain watermarked data for each field. The watermarking intensity of each group of hierarchical data is directly proportional to the sensitivity level of the hierarchical data. When the sensitivity level of the hierarchical data is higher than a preset level, the watermarking process includes desensitization processing, and the desensitization processing intensity of the hierarchical data is directly proportional to the sensitivity level of the hierarchical data.

[0014] The first storage module is used to store target data in the blockchain system. The target data includes: field data, sensitivity level of the field data and watermark processing data, and watermark processing strategy.

[0015] Thirdly, embodiments of this application provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the data watermarking processing method as described in the first aspect.

[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the data watermarking processing method as described in the first aspect.

[0017] Fifthly, embodiments of this application provide a computer program product, including computer instructions that, when executed by a processor, implement the steps of the data watermarking processing method as described in the first aspect.

[0018] In this embodiment, data to be processed is acquired, including M fields, where M is an integer greater than 1. The M fields are then processed hierarchically to obtain N groups of hierarchical data, each group having a different sensitivity level, where N is an integer greater than 1. For each group of hierarchical data, a watermarking strategy corresponding to the sensitivity level of the hierarchical data is applied to each field, resulting in watermarked data for each field. The watermarking intensity of each group is directly proportional to the sensitivity level of the hierarchical data. If the sensitivity level of the hierarchical data is higher than a preset level, the watermarking process includes desensitization processing, and the desensitization intensity is directly proportional to the sensitivity level of the hierarchical data. The target data is then stored in a blockchain system, including: field data, the sensitivity level and watermarked data of the field data, and the watermarking strategy. In this way, by classifying data sensitivity, different watermarking processes can be applied to field data with different sensitivity levels. Furthermore, by storing the field data, the sensitivity level of the field data, the watermarking data, and the watermarking strategy in the blockchain system, fine-grained traceability issues of the watermarking process and specific field data can be accurately detected during data traceability, thereby improving the effectiveness of data traceability. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart of a data watermarking processing method provided in an embodiment of this application;

[0021] Figure 2 This is a structural diagram of a data watermarking processing device provided in an embodiment of this application;

[0022] Figure 3 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] The core of the relevant technology lies in building a traceability and verification system that integrates dynamic monitoring and post-event evaluation. Its technical approach can be summarized as follows:

[0025] 1. Embedded dynamic watermark: Embed the identifier in the data operation flow to achieve real-time tracking.

[0026] 2. Constructing a topology graph and evaluation matrix: Through cross-domain feature fusion, a data evolution topology graph is constructed and a credibility evaluation matrix is ​​generated to graphically represent the data flow path and state changes.

[0027] 3. Risk assessment and instruction generation: Assess the risk level of the traceability chain break and output the verification instruction set through the verification mapping array.

[0028] 4. Report generation and spatial reconstruction: Generate a credibility calibration report and reconstruct the credibility data space accordingly.

[0029] The related technologies may have the following limitations:

[0030] Insufficient protection of the data itself: While adept at monitoring the use of data, it does not detail how to protect the data content itself. For example, it cannot prevent authorized parties from copying, distributing, or tampering with the data after acquisition, lacking direct control over the data content.

[0031] The granularity of traceability verification is relatively coarse: its verification result may be "the data flow path has low credibility" or "there is a risk of data disruption," which is a qualitative or semi-quantitative assessment. When data tampering is discovered, it is difficult to accurately pinpoint which link or field has the problem, and it is also impossible to restore the tampered data to its original state.

[0032] Lack of a data availability balancing mechanism: The solution focuses on security monitoring but does not address how to maximize the value of data utilization while ensuring security. It fails to resolve the core contradiction of "providing data to third parties while preventing the leakage of sensitive information," namely, the balance between data privacy and availability.

[0033] Based on this, the present application provides a new data watermarking method, which aims to solve the following technical problems.

[0034] How to achieve accurate tracing and recovery of data when it is "usable but not visible".

[0035] How to ensure the integrity of data during distribution and use.

[0036] How to achieve fine-grained, clearly defined data authorization and recovery.

[0037] See Figure 1 , Figure 1This is a flowchart of a data watermarking processing method provided in an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:

[0038] Step 101: Obtain the data to be processed, which includes M fields, where M is an integer greater than 1;

[0039] Step 102: Perform hierarchical processing on the M field data to obtain N groups of hierarchical data. Each group of hierarchical data has a different sensitivity level, and N is an integer greater than 1.

[0040] Step 103: For each group of graded data, a watermarking strategy corresponding to the sensitivity level of the graded data is adopted to watermark each field of the graded data, resulting in watermarked data for each field. The watermarking intensity of each group of graded data is directly proportional to the sensitivity level of the graded data. If the sensitivity level of the graded data is higher than a preset level, the watermarking includes desensitization, and the desensitization intensity of the graded data is directly proportional to the sensitivity level of the graded data.

[0041] Step 104: Store the target data in the blockchain system. The target data includes: field data, the sensitivity level of the field data and watermark processing data, and the watermark processing strategy.

[0042] It should be noted that the embodiments of this application involve big data technology and are applied to a data watermarking processing device.

[0043] In step 101, the data to be processed can be data in a trusted data space, which can be a data storage space in a data platform. The data to be processed can be business processing data, such as shopping data, navigation data, etc.

[0044] The data to be processed may include M fields. The data watermarking device may include hierarchical processing nodes. The hierarchical processing unit can extract the field types of the data to be processed through the feature extraction unit to obtain the field data. For example, it can extract field data such as ID card number, which is a direct identifier; extract field data such as transaction amount, which is an indirect identifier; and extract field data such as device ID, which is relatively low-sensitivity data and can be used as sensitive features. For another example, it can extract field data such as data purpose, such as cross-border transactions, which can be used as traceability requirements for full lifecycle traceability; and extract field data such as local consumption, which can be used as traceability requirements for short-term traceability.

[0045] In step 102, the data watermarking processing device may include a hierarchical algorithm unit. The hierarchical algorithm unit may employ a random forest algorithm, using decision data such as 20-50 decision trees, selecting information gain ratio as a feature, and performing hierarchical processing on the field data to generate hierarchical results. For example, the hierarchical results may be S1 level, S2 level, and S3 level.

[0046] In some embodiments, field data can be classified from two dimensions and then combined and graded from these two dimensions. The grading result can be a combined grading of the sensitivity of sensitive features and traceability requirement features. If the sensitivity level of the sensitive features is S1 and the sensitivity level of the traceability requirement features is T1, then the combined grading result can be S1T1. This allows the field data to be classified into S1T1, S1T2, S2T1, S2T2, S3T1, and S3T2, resulting in N groups of tiered data. Each group of tiered data has a different sensitivity level, making the classification more detailed. This approach can adapt to eight core scenarios, including finance (e.g., cross-border transactions / local consumption), healthcare (e.g., diagnosis / scientific research), and government affairs (e.g., long-term archiving / short-term statistics), improving scenario adaptability.

[0047] In some embodiments, the data watermarking processing device may include a classification result verification unit, which can verify the combined classification results. The classification result verification unit can verify the annotation results through 5-fold cross-validation. If the accuracy is ≥95% and the misclassification rate of S1 / S2 level is ≤5%, the classification result is output; otherwise, the random forest parameters are readjusted (e.g., the number of decision trees is increased), and the classification is performed again until the accuracy requirements are met.

[0048] In step 103, hierarchical watermarking can be performed. Different watermarking strategies are applied to the field data in different groups of hierarchical data according to different sensitivity levels. The watermarking strategy can include watermark metadata and watermarking information. The watermarking information can include the watermark embedding method, such as using wavelet transform to embed the watermark into the field data.

[0049] In some embodiments, the watermark processing information may also include a desensitization processing method for field data, such as using a generalized processing method to desensitize the field data, or using an encryption method to desensitize the field data.

[0050] The watermarking intensity for each group of hierarchical data is directly proportional to the sensitivity level of the corresponding data level. In other words, the higher the sensitivity level of the field data (i.e., the more sensitive the field data), the stronger the watermarking intensity, and vice versa. The watermarking intensity can be reflected in the desensitization intensity, the complexity of the watermark metadata, and the watermark embedding method. Furthermore, the desensitization intensity for hierarchical data is also directly proportional to the sensitivity level of the corresponding data level. That is, the higher the sensitivity level of the field data (i.e., the more sensitive the field data), the stronger the desensitization intensity, and vice versa. The desensitization intensity can also be reflected in the desensitization method; for example, using encryption for data desensitization is stronger than using generalized processing methods.

[0051] In this step, watermarking involves de-identifying high-sensitivity fields before embedding the watermark. This allows de-identified data to be used in business applications, maximizing data utilization while ensuring security. This approach resolves the core contradiction of providing data to third parties while preventing the leakage of sensitive information, achieving a balance between data privacy and usability.

[0052] In step 104, the target data can be stored in the blockchain system. The target data includes: field data, the sensitivity level and watermark processing data of the field data, and the watermark processing strategy. The watermark processing strategy can include watermark metadata and watermark processing information. The watermark processing information can include desensitization processing methods. This allows for hierarchical and reversible desensitization of resource data based on the blockchain-based watermark processing strategy. While providing desensitized data for business use, it ensures that the original data can be accurately restored to the required precision under legal authorization. This directly resolves the contradiction between data use and privacy protection, enabling precise traceability and recovery of data that is "usable but not visible."

[0053] Furthermore, by desensitizing the data, the data content itself can be protected, thereby preventing authorized parties from copying, disseminating, or tampering with the data after obtaining it, and enabling direct control over the data content.

[0054] Furthermore, by dividing the data to be processed into field data and watermarking the field data according to different sensitivity levels, when data tampering is detected, the source of the field data can be traced to accurately pinpoint which link and which field has the problem. Through the target data stored on the blockchain, such as the watermarking strategy, any tampering with the desensitized data can be detected quickly and accurately by associating the watermark with the blockchain notarization. This ensures the integrity of the data during distribution and use and enables the restoration of tampered data to its original state.

[0055] In some embodiments, step 103 specifically includes:

[0056] When the sensitivity level of the graded data is the first level, the field data in the graded data is desensitized by encryption to obtain the first desensitized data. Based on the first watermark metadata, the first desensitized data is watermarked to obtain the watermarked data of the field data. The first watermark metadata is generated based on the ownership identifier, sensitivity level and blockchain timestamp of the field data.

[0057] When the sensitivity level of the graded data is the second level, the field data in the graded data is desensitized using a data generalization processing method to obtain the second desensitized data. Based on the second watermark metadata, the second desensitized data is watermarked to obtain the watermarked data of the field data. The first watermark metadata is generated based on the field data's ownership identifier and sensitivity level. The data volume of the second watermark metadata is less than that of the first watermark metadata, and the second level is lower than the first level.

[0058] When the sensitivity level corresponding to the graded data is the third level, watermarking is embedded in the field data in the graded data based on the third watermark metadata to obtain watermarked data of the field data. The third watermark metadata is generated based on the hash value and access address of the field data, and the third level is lower than the second level.

[0059] In some embodiments, the field data can be hierarchically and reversibly desensitized and watermarked simultaneously. Hierarchical and reversible desensitization refers to using different desensitization processing methods according to different sensitivity levels. Reversible means that the desensitization processing is reversible and the field data can be restored to its original state through the processing method corresponding to the desensitization processing method.

[0060] In some embodiments, for field data with a high sensitivity level, i.e., S1 level field data, such as the ID number "110101129001011234", the desensitization processing can be carried out using encryption methods. For example, a Format-Preserving Encryption (FPE) encryption unit can be used, and the AES-256 algorithm (such as ECB mode, 32-byte key) can be used to desensitize the sensitive field data, generating the first desensitized data, such as the first desensitized data "11019901234". At the same time, a mapping table between the field data and the first desensitized data is generated, such as storing the correspondence between "110101129001011234" and "11019901234", so as to achieve subsequent reversible desensitization and restore it to the original data.

[0061] In some embodiments, for S1-level field data, a watermark can be embedded in the first desensitized data based on the first watermark metadata to obtain watermarked data of the field data. The first watermark metadata is generated based on the field data's ownership identifier, sensitivity level, and the corresponding blockchain timestamp. This allows for data traceability, enabling the identification of the data owner and the blockchain timestamp contained in the watermark metadata, thus achieving finer-grained data traceability.

[0062] For example, for S1 level field data, watermark embedding can be performed using wavelet transform watermarking units. The first desensitized data undergoes a 3-level IWT decomposition, selecting the LL3 low-frequency sub-band (16×16 resolution). The first watermark metadata, such as a 128-bit watermark, is embedded according to the rule of "embedding 1 watermark for every 4 coefficients," and the PSNR is controlled to be ≥43dB through least squares error correction (error threshold 0.005). The first watermark metadata can indicate the field data's ownership identifier, such as the data owner ID 0x1234, the blockchain timestamp, such as 1719200000, and the sensitivity level, such as 0x01 (indicating S1 level).

[0063] In some embodiments, for field data with a general sensitivity level, i.e., S2 level field data, such as the transaction amount "12568 yuan", a generalization processing method can be used to desensitize the field data in the hierarchical data to obtain second desensitized data. For example, 12568 yuan can be generalized to 12000-13000 yuan. Then, based on the second watermark metadata, a watermark can be embedded in the second desensitized data to obtain the watermarked data of the field data. The second watermark metadata is generated based on the field data's ownership identifier and sensitivity level, and its data size, such as its length, can be smaller than that of the first watermark metadata. This allows for stronger watermarking of field data with higher sensitivity levels.

[0064] For example, for S2 level field data, the watermark can be embedded using a least squares error correction algorithm to embed a 64-bit second watermark, with an embedding time of ≤9.2ms. The second watermark metadata only indicates the sensitivity level, such as 0x02 (indicating S2 level), and the field data's ownership identifier.

[0065] In some embodiments, for fields with low sensitivity levels, i.e., S3 level data processing, such as the device ID "DEV-2025-001", desensitization is not required; only a third watermark metadata is added to the data header. This supports rapid traceability and verification. The third watermark metadata can indicate the data hash value, such as 0x1A2B3C, and the access address, such as / ipfs / QmXYZ.

[0066] In this embodiment, through hierarchical reversible data anonymization, while providing anonymized data for business use, it ensures that the original data can be accurately restored to the required precision under legal authorization. This directly resolves the contradiction between data use and privacy protection, enabling precise traceability and recovery of data while maintaining its usability without visibility. Furthermore, it can trace the data owner and timestamp contained in the watermark information, making the granularity of traceability verification finer and improving the effectiveness of data traceability.

[0067] In some embodiments, when the sensitivity level corresponding to the graded data is the first level, after embedding a watermark into the first desensitized data based on the first watermark metadata to obtain watermarked data of the field data, the method further includes:

[0068] Generate a mapping table between the field data and the first de-identified data;

[0069] Based on the data space node identifier, blockchain timestamp, and authorizing party fingerprint hash value of the field data, a dynamic recovery key for the field data is generated.

[0070] The dynamic recovery key is bound to the mapping table and stored in the encrypted database.

[0071] When the sensitivity level is S1, a dynamic recovery key (e.g., a 32-byte key) can be generated using the SM4 algorithm by combining the data space node identifier of the field data (e.g., 0xABCDEF), the blockchain timestamp (e.g., 1719200001), and the authorizing party's fingerprint hash (e.g., 0x987654). This dynamic recovery key can be bound to a mapping table and stored in an encrypted database. Subsequently, the data recovery operation itself can be ensured to be secure and auditable through dynamic recovery key and on-chain permission verification, preventing the abuse of recovery permissions. This enables fine-grained and clearly defined data authorization recovery.

[0072] In some embodiments, step 104 specifically includes:

[0073] The first data in the target data is stored in the InterPlanetary File System to obtain the content identifier of the first data, which includes field data and watermark processing data.

[0074] The second data in the target data is stored in the main chain of the blockchain system, and the second data includes the content identifier;

[0075] The third data in the target data is stored in the sidechain of the blockchain system. The third data includes the watermark metadata and watermark processing information in the watermark processing strategy.

[0076] The target data can be stored and verified on the blockchain across the entire chain.

[0077] (1) Off-chain storage: Field data and watermark processing data in the data to be processed can be uploaded to the InterPlanetary File System (IPFS) to obtain the corresponding content identifier (CID), such as field data CID: Qm123 and watermark processing data CID: Qm456;

[0078] (2) Main chain evidence storage: Blockchain evidence storage nodes can write core information, i.e., secondary data, into the main chain through the PBFT consensus mechanism (block generation interval ≤ 0.8 seconds), including: the hash value of the dynamic recovery key, such as 0x7D8E9F, the sensitivity level, such as S1 level, and the IPFS CID;

[0079] (3) Sidechain evidence storage: Detailed information can be written to the sidechain (Polygon), including: watermark metadata such as watermark content and watermark processing information. The watermark processing information can include watermark embedding algorithm, PSNR value and desensitization operation log, such as desensitization processing method, desensitization time: 2025-08-25 10:00:00, operator ID: 0x001, desensitization algorithm version: V2.1;

[0080] (4) Cross-chain synchronization: Data synchronization between the main chain and the side chain can be achieved through Polkadot's XCMP protocol to ensure the consistency of the evidence information and the confirmation delay is ≤0.9 seconds.

[0081] In some embodiments, data traceability can be achieved through target data stored in a blockchain system. Following step 104, the method further includes:

[0082] Obtain the first content identifier of the data to be traced;

[0083] Obtain watermark metadata and watermark processing information corresponding to the first content identifier in the main chain of the blockchain system from the sidechain of the blockchain system, and obtain watermark processing data corresponding to the first content identifier from the InterPlanetary File System.

[0084] Based on the watermark processing information, extract the watermark metadata from the watermark processing data;

[0085] Based on the comparison between the extracted watermark metadata and the watermark metadata stored in the sidechain of the blockchain system, anomaly detection is performed on the data to be traced.

[0086] Scenario 1: Source tracing and verification, such as regulatory agencies checking data integrity.

[0087] The traceability initiator inputs the first content identifier, such as IPFS CID: Qm456, and the traceability verification module sends a query request to the blockchain verification node;

[0088] The blockchain verification node retrieves the corresponding watermark metadata (such as 128-bit watermark content) from the sidechain.

[0089] The source tracing and verification module extracts watermarks from the watermarked data in IPFS using inverse IWT (for S1 level field data) and inverse least squares algorithm (for S2 level field data).

[0090] The extracted watermark is compared with the on-chain stored watermark metadata. If the matching degree is ≥98%, the data is determined to be complete; otherwise, the data is determined to have been tampered with, and an anomaly report is generated.

[0091] In some embodiments, data recovery can be achieved using target data stored in a blockchain system. The second data also includes the hash value and sensitivity level of the dynamic recovery key for the field data. After step 104, the method further includes:

[0092] Obtain the second content identifier and query identity information of the data to be recovered;

[0093] If the identity information verification is successful, the sensitivity level corresponding to the second content identifier in the main chain of the blockchain system is obtained from the side chain of the blockchain system. The successful identity verification includes comparing the hash value of the dynamic recovery key with the hash value of the query key in the identity information. The hash value of the dynamic recovery key is obtained from the side chain of the blockchain system and corresponds to the second content identifier in the main chain of the blockchain system.

[0094] The data to be recovered is processed based on the sensitivity level.

[0095] Scenario 2: Data recovery, such as when an authorized agency needs to obtain the original data for auditing purposes.

[0096] The authorizing party submits a recovery request (including the second content identifier CID of the data to be recovered and the required recovery precision). For example, if the S2 level field data needs to be recovered to "12,500-12,600 yuan", the recovery node can verify the identity of the querying information, such as through biometric comparison and on-chain permission query. In some embodiments, if it is necessary to recover the S1 level field data, the validity of the query key needs to be verified.

[0097] The recovery node can send a key hash query request to the blockchain key verification node to verify the validity of the query key. The hash value of the query key in the query identity information provided by the authorizing party is compared with the hash value of the dynamic recovery key stored on the chain. If they match, it means that the query key is valid.

[0098] Once the verification is successful, the data to be recovered can be processed based on the sensitivity level.

[0099] In some embodiments, for field data at the S1 level, the corresponding mapping table can be retrieved from the encrypted database to completely restore the original data, such as "11019901234" being restored to "110101199001011234".

[0100] For S2 level field data, a gradual partial recovery can be performed, and the recovery can be carried out according to the accuracy requirements. For example, "12000-13000 yuan" can be recovered to "12500-12600 yuan". Dynamic noise (noise intensity ≤0.01) can be added through differential privacy algorithm to ensure that the deviation of data statistical features is ≤5%.

[0101] For S3 level field data, the original data can be directly retrieved, such as the device ID "DEV-2025-001".

[0102] Once the recovery is complete, the recovery operation log (such as recovery time, authorizing party ID, and recovery accuracy) can be written to the blockchain sidechain to support subsequent auditing.

[0103] The following is a detailed description of the data watermarking method provided in the embodiments of this application using a specific example.

[0104] The device is applied to a data watermarking processing device, which may include a data access module, a multi-dimensional hierarchical module, a reversible desensitization and watermarking module, a blockchain evidence storage module, a reversible recovery and traceability module, and an authorization verification module.

[0105] Its function is to receive raw data uploaded by the data owner, perform format verification and integrity checks, and trigger subsequent hierarchical processes after the verification is passed. The raw data can be structured database tables, unstructured image files, etc. Format verification can verify whether the data conforms to the JSON / XML standard, and integrity checks can check whether the data is missing key fields.

[0106] The data access module can interact with the multi-dimensional hierarchical module. Verified raw data can be transmitted to the multi-dimensional hierarchical module with a transmission latency of ≤500ms.

[0107] The components of a multi-dimensional hierarchical module may include:

[0108] Feature extraction unit: can extract sensitive features and traceability requirements features from the original data and output feature vectors (such as [direct identifier, full lifecycle traceability]).

[0109] Hierarchical algorithm unit: The random forest algorithm is used as input, and the combined hierarchical results (such as S1T1) are output.

[0110] Classification result verification unit: The classification accuracy is verified through cross-validation. If the requirements are not met, the results are fed back to the algorithm unit to adjust the parameters.

[0111] The multi-dimensional grading module can receive raw data from the data access module, output grading results to the grading reversible desensitization and watermarking module, and send grading logs (such as grading time and accuracy) to the blockchain evidence storage module.

[0112] The components of a graded reversible desensitization and watermarking module may include:

[0113] FPE encryption unit: Performs AES-256 format encryption on S1 level field data to generate the first desensitized data and a mapping table between the field data and the first desensitized data;

[0114] Wavelet transform watermarking unit: Performs IWT watermark embedding on S1 level field data and lightweight watermark embedding on S2 level field data;

[0115] Dynamic key generation unit: Combines multiple factors to generate dynamic recovery keys, which are then bound to a mapping table.

[0116] The graded reversible desensitization and watermarking module can receive the graded results from the graded module, output watermark processing data, dynamic recovery key, and mapping relationship table to the encrypted database, and send watermark metadata and desensitization operation logs to the blockchain evidence storage module.

[0117] The encrypted database stores mapping relationship tables and dynamic recovery keys in plaintext, supports fast querying by data CID, employs transparent data encryption technology to prevent unauthorized reading of database files, records all query operations, and supports access log auditing.

[0118] The components of a blockchain-based evidence storage module may include:

[0119] On-chain data writing unit: Writes core information (dynamic recovery key hash value, CID) to the main chain and detailed information (watermark metadata, watermark processing information) to the side chain;

[0120] Cross-chain synchronization unit: Realizes data synchronization between the main chain and side chains through the XCMP protocol to ensure consistency.

[0121] The blockchain evidence storage module can receive hierarchical logs from the multi-dimensional hierarchical module and watermark metadata / de-identification logs from the reversible de-identification and watermarking module, and feed back the evidence storage results to the reversible recovery and traceability module.

[0122] The reversible recovery and tracing module may include the following components:

[0123] Watermark extraction unit: Based on the data hierarchy, call the corresponding extraction algorithm (inverse IWT, inverse least squares) to extract the watermark;

[0124] Data recovery unit: Performs tiered recovery by matching the mapping relationship table according to the authorized precision requirements;

[0125] Traceability Report Generation Unit: Generates a traceability report that includes data flow path, operation records, and integrity verification results.

[0126] The reversible recovery and traceability module can receive the identity verification result from the authorization verification module, retrieve the evidence storage information from the blockchain evidence storage module, retrieve the mapping relationship table from the encrypted database, and output the recovery data or traceability report to the authorizing party.

[0127] The authorization verification module uses multi-factor authentication, such as biometrics (e.g., fingerprint / face hash), account password, and on-chain permission query, to verify the legitimacy of the identities of the tracing initiator and the restoration authorization party. Its interaction logic is to send the identity verification result to the reversible restoration and tracing module. If the verification is successful, subsequent operations are allowed; otherwise, the request is rejected.

[0128] The function of IPFS distributed storage is to store raw data and watermarked data, and to uniquely identify files through CID, support distributed access, and reduce the risk of data loss; it can receive raw data / watermarked data from the reversible desensitization and watermarking module, and return the CID to the blockchain evidence storage module.

[0129] In this embodiment, a multi-dimensional hierarchical driven desensitization watermarking collaborative mechanism can be adopted. Based on the dual dimensions of sensitive features and traceability requirements, a combined hierarchical approach is used. Differentiated desensitization (such as FPE encryption, generalization processing, and no desensitization) and watermark embedding (such as 128-bit IWT watermark, 64-bit lightweight watermark, and hash watermark) are performed on field data with different sensitivity levels. This can solve the scenario adaptability problem of unified watermarking processing in related technologies and save computing resources.

[0130] The embodiments of this application have broad commercial value. In the financial field, such as cross-border payments, credit risk control, and audit supervision, data needs to be de-identified and protected, and traced back to its source. Cross-border financial data needs to comply with FATF privacy compliance requirements, and the source of data leakage needs to be traced. The embodiments of this application can meet the spatial requirements of financial data.

[0131] See Figure 2 , Figure 2 This is a structural diagram of a data watermarking processing device provided in an embodiment of this application, as shown below. Figure 2 As shown, the data watermarking processing device 200 includes:

[0132] The first acquisition module 201 is used to acquire data to be processed, the data to be processed includes M fields, where M is an integer greater than 1;

[0133] The hierarchical processing module 202 is used to perform hierarchical processing on the M field data to obtain N groups of hierarchical data, each group of hierarchical data has a different sensitivity level, and N is an integer greater than 1;

[0134] The watermarking module 203 is used to perform watermarking on each field of the hierarchical data for each group of hierarchical data, using a watermarking strategy corresponding to the sensitivity level of the hierarchical data, to obtain watermarked data for each field. The watermarking intensity of each group of hierarchical data is directly proportional to the sensitivity level of the hierarchical data. When the sensitivity level of the hierarchical data is higher than a preset level, the watermarking includes desensitization processing, and the desensitization intensity of the hierarchical data is directly proportional to the sensitivity level of the hierarchical data.

[0135] The first storage module 204 is used to store target data in the blockchain system. The target data includes: field data, sensitivity level of the field data and watermark processing data, and watermark processing strategy.

[0136] Optionally, the watermark processing module 203 is specifically used for:

[0137] When the sensitivity level of the graded data is the first level, the field data in the graded data is desensitized by encryption to obtain the first desensitized data. Based on the first watermark metadata, the first desensitized data is watermarked to obtain the watermarked data of the field data. The first watermark metadata is generated based on the ownership identifier, sensitivity level and blockchain timestamp of the field data.

[0138] When the sensitivity level of the graded data is the second level, the field data in the graded data is desensitized using a data generalization processing method to obtain the second desensitized data. Based on the second watermark metadata, the second desensitized data is watermarked to obtain the watermarked data of the field data. The first watermark metadata is generated based on the field data's ownership identifier and sensitivity level. The data volume of the second watermark metadata is less than that of the first watermark metadata, and the second level is lower than the first level.

[0139] When the sensitivity level corresponding to the graded data is the third level, watermarking is embedded in the field data in the graded data based on the third watermark metadata to obtain watermarked data of the field data. The third watermark metadata is generated based on the hash value and access address of the field data, and the third level is lower than the second level.

[0140] Optionally, if the sensitivity level corresponding to the graded data is the first level, the device further includes:

[0141] The first generation module is used to generate a mapping table between field data and the first de-identified data;

[0142] The second generation module is used to generate a dynamic recovery key for the field data based on the data space node identifier, blockchain timestamp, and authorizing party fingerprint hash value of the field data.

[0143] The second storage module is used to bind and store the dynamic recovery key with the mapping table in the encrypted database.

[0144] Optionally, the first storage module 204 is specifically used for:

[0145] The first data in the target data is stored in the InterPlanetary File System to obtain the content identifier of the first data, which includes field data and watermark processing data.

[0146] The second data in the target data is stored in the main chain of the blockchain system, and the second data includes the content identifier;

[0147] The third data in the target data is stored in the sidechain of the blockchain system. The third data includes the watermark metadata and watermark processing information in the watermark processing strategy.

[0148] Optionally, the device further includes:

[0149] The second acquisition module is used to acquire the first content identifier of the data to be traced.

[0150] The third acquisition module is used to acquire watermark metadata and watermark processing information corresponding to the first content identifier in the main chain of the blockchain system from the side chain of the blockchain system, and to acquire watermark processing data corresponding to the first content identifier from the interplanetary file system.

[0151] The extraction module is used to extract watermark metadata from the watermark processing data based on the watermark processing information.

[0152] The traceability module is used to determine anomalies in the data to be traced based on the comparison between the extracted watermark metadata and the watermark metadata stored in the sidechain of the blockchain system.

[0153] Optionally, the second data further includes the hash value of the dynamic recovery key and the sensitivity level of the field data, and the device further includes:

[0154] The fourth acquisition module is used to acquire the second content identifier and query identity information of the data to be recovered;

[0155] The fifth acquisition module is used to obtain the sensitivity level corresponding to the second content identifier in the main chain of the blockchain system from the side chain of the blockchain system when the identity information is verified. The identity verification is passed by comparing the hash value of the dynamic recovery key with the hash value of the query key in the identity information. The hash value of the dynamic recovery key is obtained from the side chain of the blockchain system and corresponds to the second content identifier in the main chain of the blockchain system.

[0156] The recovery module is used to perform recovery processing on the data to be recovered based on the sensitivity level.

[0157] The data watermarking processing device 200 can implement all the processes implemented in the above-described data watermarking processing method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0158] See Figure 3 The figure shows a structural diagram of an electronic device provided in an embodiment of the present invention. Figure 3 As shown, the electronic device 300 includes: a processor 301, a memory 302, a user interface 303, and a bus interface 304.

[0159] Processor 301 is used to read the program from memory 302 and execute the following procedures:

[0160] Obtain the data to be processed, which includes M fields, where M is an integer greater than 1;

[0161] The M fields of data are processed in a hierarchical manner to obtain N groups of hierarchical data. Each group of hierarchical data has a different sensitivity level, and N is an integer greater than 1.

[0162] For each group of graded data, a watermarking strategy corresponding to the sensitivity level of the graded data is adopted to watermark each field of the graded data, resulting in watermarked data for each field. The watermarking intensity of each group of graded data is directly proportional to the sensitivity level of the graded data. When the sensitivity level of the graded data is higher than a preset level, the watermarking includes desensitization, and the desensitization intensity of the graded data is directly proportional to the sensitivity level of the graded data.

[0163] The target data is stored in the blockchain system. The target data includes: field data, the sensitivity level of the field data and watermark processing data, and the watermark processing strategy.

[0164] exist Figure 3 In this context, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits of one or more processors represented by processor 301 and memory represented by memory 302 together. The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 304 provides an interface. For different user devices, user interface 303 can also be an interface capable of connecting external or internal devices, including but not limited to keypads, displays, speakers, microphones, joysticks, etc.

[0165] The processor 301 is responsible for managing the bus architecture and general processing, while the memory 302 can store the data used by the processor 301 when performing operations.

[0166] In some embodiments, the processor 301 is further configured to:

[0167] When the sensitivity level of the graded data is the first level, the field data in the graded data is desensitized by encryption to obtain the first desensitized data. Based on the first watermark metadata, the first desensitized data is watermarked to obtain the watermarked data of the field data. The first watermark metadata is generated based on the ownership identifier, sensitivity level and blockchain timestamp of the field data.

[0168] When the sensitivity level of the graded data is the second level, the field data in the graded data is desensitized using a data generalization processing method to obtain the second desensitized data. Based on the second watermark metadata, the second desensitized data is watermarked to obtain the watermarked data of the field data. The first watermark metadata is generated based on the field data's ownership identifier and sensitivity level. The data volume of the second watermark metadata is less than that of the first watermark metadata, and the second level is lower than the first level.

[0169] When the sensitivity level corresponding to the graded data is the third level, watermarking is embedded in the field data in the graded data based on the third watermark metadata to obtain watermarked data of the field data. The third watermark metadata is generated based on the hash value and access address of the field data, and the third level is lower than the second level.

[0170] In some embodiments, the processor 301 is further configured to:

[0171] When the sensitivity level corresponding to the graded data is the first level, a mapping table between the field data and the first desensitized data is generated;

[0172] Based on the data space node identifier, blockchain timestamp, and authorizing party fingerprint hash value of the field data, a dynamic recovery key for the field data is generated.

[0173] The dynamic recovery key is bound to the mapping table and stored in the encrypted database.

[0174] In some embodiments, the processor 301 is further configured to:

[0175] The first data in the target data is stored in the InterPlanetary File System to obtain the content identifier of the first data, which includes field data and watermark processing data.

[0176] The second data in the target data is stored in the main chain of the blockchain system, and the second data includes the content identifier;

[0177] The third data in the target data is stored in the sidechain of the blockchain system. The third data includes the watermark metadata and watermark processing information in the watermark processing strategy.

[0178] In some embodiments, the processor 301 is further configured to:

[0179] Obtain the first content identifier of the data to be traced;

[0180] Obtain watermark metadata and watermark processing information corresponding to the first content identifier in the main chain of the blockchain system from the sidechain of the blockchain system, and obtain watermark processing data corresponding to the first content identifier from the InterPlanetary File System.

[0181] Based on the watermark processing information, extract the watermark metadata from the watermark processing data;

[0182] Based on the comparison between the extracted watermark metadata and the watermark metadata stored in the sidechain of the blockchain system, anomaly detection is performed on the data to be traced.

[0183] In some embodiments, the second data further includes the hash value of the dynamic recovery key and the sensitivity level of the field data, and the processor 301 is further configured to:

[0184] Obtain the second content identifier and query identity information of the data to be recovered;

[0185] If the identity information verification is successful, the sensitivity level corresponding to the second content identifier in the main chain of the blockchain system is obtained from the side chain of the blockchain system. The successful identity verification includes comparing the hash value of the dynamic recovery key with the hash value of the query key in the identity information. The hash value of the dynamic recovery key is obtained from the side chain of the blockchain system and corresponds to the second content identifier in the main chain of the blockchain system.

[0186] The data to be recovered is processed based on the sensitivity level.

[0187] Preferably, the present invention also provides an electronic device 300, including a processor 301, a memory 302, and a computer program stored in the memory 302 and executable on the processor 301. When the computer program is executed by the processor 301, it implements the various processes of the above-described data watermarking method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0188] This invention also provides a readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described data watermarking method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here. The readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0189] This application also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the various processes of the above-described data watermarking method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0190] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0191] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0192] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0193] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention, depending on actual needs.

[0194] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0195] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0196] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A data watermarking processing method, characterized in that, The method includes: Obtain the data to be processed, which includes M fields, where M is an integer greater than 1; The M fields of data are processed in a hierarchical manner to obtain N groups of hierarchical data. Each group of hierarchical data has a different sensitivity level, and N is an integer greater than 1. For each group of graded data, a watermarking strategy corresponding to the sensitivity level of the graded data is adopted to watermark each field of the graded data, resulting in watermarked data for each field. The watermarking intensity of each group of graded data is directly proportional to the sensitivity level of the graded data. When the sensitivity level of the graded data is higher than a preset level, the watermarking includes desensitization, and the desensitization intensity of the graded data is directly proportional to the sensitivity level of the graded data. The target data is stored in the blockchain system. The target data includes: field data, the sensitivity level of the field data and watermark processing data, and the watermark processing strategy.

2. The method according to claim 1, characterized in that, The watermarking strategy, employing a sensitivity level corresponding to the hierarchical data, performs watermarking on each field of the hierarchical data to obtain watermarked data for each field, including: When the sensitivity level of the graded data is the first level, the field data in the graded data is desensitized by encryption to obtain the first desensitized data. Based on the first watermark metadata, the first desensitized data is watermarked to obtain the watermarked data of the field data. The first watermark metadata is generated based on the ownership identifier, sensitivity level and blockchain timestamp of the field data. When the sensitivity level of the graded data is the second level, the field data in the graded data is desensitized using a data generalization processing method to obtain the second desensitized data. Based on the second watermark metadata, the second desensitized data is watermarked to obtain the watermarked data of the field data. The first watermark metadata is generated based on the field data's ownership identifier and sensitivity level. The data volume of the second watermark metadata is less than that of the first watermark metadata, and the second level is lower than the first level. When the sensitivity level corresponding to the graded data is the third level, watermarking is embedded in the field data in the graded data based on the third watermark metadata to obtain watermarked data of the field data. The third watermark metadata is generated based on the hash value and access address of the field data, and the third level is lower than the second level.

3. The method according to claim 2, characterized in that, When the sensitivity level corresponding to the graded data is the first level, after embedding a watermark into the first desensitized data based on the first watermark metadata to obtain the watermarked data of the field data, the method further includes: Generate a mapping table between the field data and the first de-identified data; Based on the data space node identifier, blockchain timestamp, and authorizing party fingerprint hash value of the field data, a dynamic recovery key for the field data is generated. The dynamic recovery key is bound to the mapping table and stored in the encrypted database.

4. The method according to claim 1, characterized in that, The process of storing the target data in the blockchain system includes: The first data in the target data is stored in the InterPlanetary File System to obtain the content identifier of the first data, which includes field data and watermark processing data. The second data in the target data is stored in the main chain of the blockchain system, and the second data includes the content identifier; The third data in the target data is stored in the sidechain of the blockchain system. The third data includes the watermark metadata and watermark processing information in the watermark processing strategy.

5. The method according to claim 4, characterized in that, After storing the target data in the blockchain system, the method further includes: Obtain the first content identifier of the data to be traced; Obtain watermark metadata and watermark processing information corresponding to the first content identifier in the main chain of the blockchain system from the sidechain of the blockchain system, and obtain watermark processing data corresponding to the first content identifier from the InterPlanetary File System. Based on the watermark processing information, extract the watermark metadata from the watermark processing data; Based on the comparison between the extracted watermark metadata and the watermark metadata stored in the sidechain of the blockchain system, anomaly detection is performed on the data to be traced.

6. The method according to claim 4, characterized in that, The second data also includes the hash value and sensitivity level of the dynamic recovery key for the field data. After storing the target data in the blockchain system, the method further includes: Obtain the second content identifier and query identity information of the data to be recovered; If the identity information verification is successful, the sensitivity level corresponding to the second content identifier in the main chain of the blockchain system is obtained from the side chain of the blockchain system. The successful identity verification includes comparing the hash value of the dynamic recovery key with the hash value of the query key in the identity information. The hash value of the dynamic recovery key is obtained from the side chain of the blockchain system and corresponds to the second content identifier in the main chain of the blockchain system. The data to be recovered is processed based on the sensitivity level.

7. A data watermarking processing device, characterized in that, The device includes: The first acquisition module is used to acquire data to be processed, which includes M fields, where M is an integer greater than 1. The hierarchical processing module is used to perform hierarchical processing on the M field data to obtain N groups of hierarchical data. Each group of hierarchical data has a different sensitivity level, and N is an integer greater than 1. The watermarking module is used to apply a watermarking strategy corresponding to the sensitivity level of the hierarchical data to each group of hierarchical data, and to watermark each field of the hierarchical data to obtain watermarked data for each field. The watermarking intensity of each group of hierarchical data is directly proportional to the sensitivity level of the hierarchical data. When the sensitivity level of the hierarchical data is higher than a preset level, the watermarking process includes desensitization processing, and the desensitization processing intensity of the hierarchical data is directly proportional to the sensitivity level of the hierarchical data. The first storage module is used to store target data in the blockchain system. The target data includes: field data, sensitivity level of the field data and watermark processing data, and watermark processing strategy.

8. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the data watermarking processing method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the data watermarking processing method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the data watermarking processing method as described in any one of claims 1 to 6.