Data processing method and device, equipment and storage medium

By acquiring and analyzing the lineage propagation information of data, and adjusting the classification and grading information of upstream nodes according to the propagation mode, the problem of inaccurate data classification and grading in the existing technology is solved, and accurate classification and grading in the data propagation process is achieved.

CN115630210BActive Publication Date: 2026-05-05HENAN XINGHUAN ZHONGZHI INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HENAN XINGHUAN ZHONGZHI INFORMATION TECH CO LTD
Filing Date
2022-10-12
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In the process of data classification and security grading, existing technologies make it difficult for scanning and identification programs to fully and accurately identify sensitive data, resulting in inaccurate classification and grading results, and incomplete scanning results after sensitive data is disseminated.

Method used

By acquiring the classification and grading information of data at upstream nodes, as well as the lineage propagation information from upstream nodes to the current node, the classification and grading information of the upstream nodes is adjusted according to the propagation method to obtain the classification and grading information of the current node, including methods such as direct assignment, indirect assignment, aggregation operation, masking, encryption, and reversible and irreversible operation function processing.

Benefits of technology

It improves the accuracy of data classification and grading, and ensures the integrity and accuracy of classification and grading results through kinship transmission analysis and adjustment, especially by dynamically adjusting grading information during data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115630210B_ABST
    Figure CN115630210B_ABST
Patent Text Reader

Abstract

This invention discloses a data processing method, apparatus, device, and storage medium. The method includes: acquiring upstream classification information and upstream hierarchical information of data at an upstream node; acquiring lineage propagation information of data from the upstream node to the current node; determining the data propagation mode based on the lineage propagation information; and adjusting the upstream classification information and the upstream hierarchical information according to the data propagation mode to obtain the current classification information and current hierarchical information of the data at the current node. This technical solution can improve the accuracy of data classification and data hierarchical analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a data processing method, apparatus, device and storage medium. Background Technology

[0002] Data classification and security grading are crucial steps in data processing. A common method is to scan the data to identify its categories and assign them a grade. However, due to data distortion or aggregation during analysis and computation, scanning programs often struggle to comprehensively and accurately identify sensitive data, leading to inaccurate classification and grading results. Furthermore, if sensitive data spreads after scanning, such as to newly created data tables, the previous scan results will become incomplete. Summary of the Invention

[0003] This invention provides a data processing method, apparatus, device, and storage medium that can improve the accuracy of data classification and data grading.

[0004] According to one aspect of the present invention, an embodiment of the present invention provides a data processing method, comprising:

[0005] Obtain upstream classification and hierarchical information of the data at the upstream nodes;

[0006] Obtain the lineage propagation information of data from upstream nodes to the current node;

[0007] The data transmission method is determined based on the aforementioned bloodline transmission information;

[0008] The upstream classification information and the upstream hierarchical information are adjusted according to the data propagation method to obtain the current classification information and the current hierarchical information of the data at the current node.

[0009] Optionally, the data propagation method includes any one of the following: direct assignment propagation, indirect assignment propagation, aggregation operation propagation, masking propagation, encryption propagation, reversible operation function propagation, and irreversible operation function propagation.

[0010] Optionally, the upstream classification information and the upstream hierarchical information may be adjusted according to the data propagation method, including:

[0011] Use the upstream classification information directly as the current classification information of the current node;

[0012] If the propagation method is direct assignment propagation or indirect assignment propagation, then the upstream hierarchical information is directly used as the current hierarchical information of the current node;

[0013] If the propagation method is masked propagation or encrypted propagation, then the upstream hierarchical information is lowered to obtain the current hierarchical information;

[0014] If the propagation method is aggregation operation propagation, then the upstream hierarchical information is adjusted up or down according to the business characteristics of the aggregation operation propagation to obtain the current hierarchical information;

[0015] If the propagation method is irreversible operation function processing, then the upstream hierarchical information is lowered to obtain the current hierarchical information;

[0016] If the propagation method is a reversible operation function, then the upstream hierarchical information is lowered according to the reversibility difficulty to obtain the current hierarchical information.

[0017] Optionally, it may also include: obtaining the upstream hierarchical confidence level of the data in the previous node;

[0018] The adjustment factor is determined based on the propagation method described above;

[0019] The upstream hierarchical confidence level is adjusted according to the adjustment factor to obtain the current hierarchical confidence level of the current node.

[0020] Optionally, the adjustment factor is determined based on the propagation method, including:

[0021] Based on the propagation method, determine the table-level lineage influence factor, the field-level lineage influence factor, and the field propagation influence factor;

[0022] The adjustment factor is obtained by weighted summation of the table-level lineage influence factor, the field-level lineage influence factor, and the field propagation influence factor.

[0023] Optionally, adjusting the upstream hierarchical confidence level according to the adjustment factor to obtain the current hierarchical confidence level of the current node includes:

[0024] The current level confidence score of the current node is obtained by multiplying the upstream level confidence score by the adjustment factor.

[0025] Optional, also includes:

[0026] If the current node's data comes from multiple upstream nodes;

[0027] The data of the current node is then classified and graded by scanning.

[0028] Optionally, obtain upstream classification information and upstream hierarchical information of the data at the upstream node, including:

[0029] Get the scanned data of the current node;

[0030] Data from upstream nodes is obtained based on the scanned data;

[0031] Determine the upstream classification information and upstream hierarchical information of the data at the upstream node.

[0032] According to another aspect of the present invention, embodiments of the present invention also provide a data processing apparatus, comprising:

[0033] The first information acquisition module is used to acquire upstream classification information and upstream hierarchical information of data at upstream nodes;

[0034] The second information acquisition module is used to acquire the lineage propagation information of data from the upstream node to the current node;

[0035] A propagation mode determination module is used to determine the data propagation mode based on the bloodline propagation information;

[0036] The information adjustment module is used to adjust the upstream classification information and the upstream hierarchical information according to the data propagation method, so as to obtain the current classification information and the current hierarchical information of the data at the current node.

[0037] According to another aspect of the present invention, embodiments of the present invention also provide an electronic device, the electronic device comprising:

[0038] At least one processor; and

[0039] A memory communicatively connected to the at least one processor; wherein,

[0040] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data processing method according to any embodiment of the present invention.

[0041] According to another aspect of the present invention, embodiments of the present invention also provide a computer-readable storage medium storing computer instructions for causing a processor to execute and implement the data processing method described in any embodiment of the present invention.

[0042] The technical solution of this invention obtains upstream classification information and upstream hierarchical information of data at upstream nodes; obtains lineage propagation information of data from upstream nodes to the current node; determines the data propagation mode based on the lineage propagation information; and adjusts the upstream classification information and upstream hierarchical information according to the data propagation mode to obtain the current classification information and current hierarchical information of data at the current node. This technical solution can improve the accuracy of data classification and data hierarchical.

[0043] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a flowchart of a data processing method provided according to Embodiment 1 of the present invention;

[0046] Figure 2 This is a general implementation example diagram of a data processing method provided in Embodiment 1 of the present invention;

[0047] Figure 3 This is a flowchart of a data processing method provided according to Embodiment 2 of the present invention;

[0048] Figure 4 This is a schematic diagram of the structure of a data processing device according to Embodiment 3 of the present invention;

[0049] Figure 5 This is a schematic diagram of the structure of an electronic device provided according to Embodiment 4 of the present invention. Detailed Implementation

[0050] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0051] It should be noted that the terms "first," "second," "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0052] Example 1

[0053] Figure 1 This is a flowchart of a data processing method according to Embodiment 1 of the present invention. This embodiment is applicable to situations involving data processing. The method can be executed by a data processing device and specifically includes the following steps:

[0054] Step 110: Obtain the upstream classification information and upstream hierarchical information of the data at the upstream node.

[0055] In this embodiment, during data processing, when the data is at the initial node, it can be understood that the initial node does not have an upstream node. Therefore, the initial node can be manually scanned to obtain its classification and hierarchical information. This embodiment can use data lineage to process the data. Data lineage is a data derivation tracking technology that can discover the propagation path of data throughout its lifecycle. For example, if a new data table is generated from a data table in a database after processing or copying, an upstream and downstream lineage is established between these two data tables; or if a file is processed or copied to generate a new file, an upstream and downstream lineage is established between these two files. In this embodiment, data lineage management collects lineage propagation information to obtain the upstream and downstream data of the current node and the propagation process information between them.

[0056] The data can be a data table, a field within a data table, or a file, or other data information. Data types can include structured data from databases or big data platforms, unstructured data, files, images, and audio, among other forms. An upstream node can be understood as the node from which the current node's data originates. Upstream classification information can be understood as the category information of the data from upstream nodes. Upstream hierarchical information can be understood as the security level information of the data from upstream nodes. In this embodiment, the security level information can be directly defined based on the data's business type or business information. For example, the security level can be from 1 to 10, or other levels, and can be set according to actual needs.

[0057] In this embodiment, upstream classification information and upstream hierarchical information of data at upstream nodes can be obtained.

[0058] In this embodiment, optionally, obtaining upstream classification information and upstream hierarchical information of data at upstream nodes includes: obtaining scanned data of the current node; obtaining data of upstream nodes based on the scanned data; and determining upstream classification information and upstream hierarchical information of data at upstream nodes.

[0059] Here, scanned data can be understood as data that has been discovered and preliminarily classified and graded through a scanning and identification process. Data from upstream nodes can be obtained based on the scanned data. In this embodiment, the scanned data of the current node can be acquired, and then the data of upstream nodes can be obtained based on the scanned data. Furthermore, the upstream classification and grading information of the upstream nodes of the determined data can be calculated.

[0060] For example, the overall implementation example diagram in this embodiment is as follows: Figure 2As shown, scanned data, also known as classified and graded data, refers to data that has been discovered and classified through scanning and identification procedures. Upstream and downstream data are data that have not yet been scanned, identified, and classified: ① Data lineage management collects lineage propagation information to obtain information on the upstream and downstream data of the classified and graded data, as well as the propagation process between them. ② Classification and grading management can obtain lineage propagation information from data lineage management, including pushing lineage propagation information to classification and grading management, or pulling lineage propagation information from data lineage management. ③ Classification and grading management analyzes the lineage propagation information, classifies and grades the upstream and downstream data, and makes necessary adjustments to the classified and graded data. This embodiment, through the collection and analysis of upstream and downstream lineage information, can discover classified and graded data from upstream and downstream, making the classification and grading identification results more complete; and through traversal analysis of the lineage propagation graph data, it can discover problems that were difficult to discover before the introduction of lineage propagation analysis technology, providing users with manual verification and improving the accuracy of data classification and grading. Furthermore, before conducting kinship transmission analysis, this embodiment identifies the initial data, data with significant transmission impact, and data exhibiting transmission loops through graph data analysis. Users are then required to manually review and confirm the classification and grading. This significantly improves the accuracy of the classification and grading results in subsequent transmission analysis and calculations.

[0061] Step 120: Obtain the lineage propagation information of data from the upstream node to the current node.

[0062] The lineage transmission information can be an SQL statement instruction, a file operation instruction, or other instruction information. Understandably, an SQL statement instruction or a file operation instruction can contain instructions to perform certain operations on data; for example, it could be an operation to copy data, an operation to indirectly copy data, or an operation to perform aggregation operations on data, etc.

[0063] In this embodiment, the lineage propagation information of data from the upstream node to the current node can be obtained.

[0064] Step 130: Determine the data transmission method based on the bloodline transmission information.

[0065] In this context, data propagation methods can be understood as the ways in which information is parsed during data propagation. Data propagation methods can be the way source data is propagated to target data, that is, the way upstream data is propagated to downstream nodes. Propagation methods can include the corresponding copying, transmission, and processing of the data. Specifically, data propagation methods can include any of the following: direct assignment propagation, indirect assignment propagation, aggregation operation propagation, masking propagation, encryption propagation, reversible operation function propagation, and irreversible operation function propagation.

[0066] Specifically, live assignment propagation can be a process where source data is copied to target data, essentially replicating data from upstream nodes to downstream nodes. Indirect assignment propagation can be a process where source data is indirectly assigned to target data via intermediate data, similar to how data from upstream nodes is indirectly assigned to downstream nodes. For example, in the SQL statement `[update targetset col1 = "rich man" where id in (select id from src where col1 = "rich man")]`, the value of the `col1` field in the `src` table is indirectly assigned to the `col1` field in the `target` table via the `id` field in both the `target` and `src` tables. Aggregate operation propagation can be the processing of upstream data using aggregate functions, which can be understood as aggregating multiple source data into a single target data through some operation, such as adding multiple source data or using aggregate functions in SQL. Masking propagation can be a method of masking a portion of certain data. For example, if the current binary data is 00111111, the portion "11" can be masked using a mask identifier. Encryption propagation can be a method of encrypting data using an encryption algorithm; for example, the AES algorithm can be used to encrypt the data. Reversible operation function propagation can be understood as using reversible operation functions to process and propagate data. Irreversible operation function propagation can be understood as using irreversible operation functions to process and propagate data. Reversible operation function propagation can include various other reversible operation functions; irreversible operation function propagation can include various other irreversible operation functions. In this embodiment, by analyzing the propagation methods during the lineage propagation process, such as direct assignment and certain aggregation operations, it is convenient to adjust the hierarchical information of the propagated data according to the propagation method.

[0067] In this embodiment, data propagation methods such as direct assignment propagation, indirect assignment propagation, aggregation operation propagation, masking processing propagation, encryption processing propagation, reversible operation function processing propagation, and irreversible operation function processing propagation can be determined based on lineage propagation information.

[0068] Step 140: Adjust the upstream classification information and the upstream hierarchical information according to the data propagation method to obtain the current classification information and the current hierarchical information of the data at the current node.

[0069] Specifically, the current classification information of the current node can be obtained by adjusting the upstream classification information based on the data propagation method. Similarly, the current hierarchical information of the current node can be obtained by adjusting the upstream hierarchical information based on the data propagation method.

[0070] In this embodiment, the upstream classification information is adjusted according to the data propagation method. The upstream classification information can be directly transmitted or remain unchanged. Adjusting the upstream hierarchical information in this embodiment can be understood as changing, raising, or lowering the upstream hierarchical information. This embodiment can achieve hierarchical determination of the propagated data by analyzing the data propagation method during the kinship transmission process, such as direct assignment or certain aggregation operations, such as keeping the security level unchanged, upgrading, or downgrading it.

[0071] In this embodiment, the upstream classification information and upstream hierarchical information can be adjusted according to data propagation methods such as direct assignment propagation, indirect assignment propagation, aggregation operation propagation, masking processing propagation, encryption processing propagation, reversible operation function processing propagation, and irreversible operation function processing propagation, so as to obtain the current classification information and current hierarchical information of the data at the current node.

[0072] In this embodiment, optionally, adjusting the upstream classification information and the upstream hierarchical information according to the data propagation method includes: directly using the upstream classification information as the current classification information of the current node; if the propagation method is direct assignment propagation or indirect assignment propagation, then directly using the upstream hierarchical information as the current hierarchical information of the current node; if the propagation method is masking processing propagation or encryption processing propagation, then lowering the upstream hierarchical information to obtain the current hierarchical information; if the propagation method is aggregation operation propagation, then raising or lowering the upstream hierarchical information according to the business characteristics of the aggregation operation propagation to obtain the current hierarchical information; if the propagation method is irreversible operation function processing propagation, then lowering the upstream hierarchical information to obtain the current hierarchical information; if the propagation method is reversible operation function processing propagation, then lowering the upstream hierarchical information according to the reversibility difficulty to obtain the current hierarchical information.

[0073] In this embodiment, upstream classification information can be directly used as the current classification information of the current node. In this embodiment, during kinship transmission, classification information can be directly transmitted; for example, if the upstream classification information is "personal contact information - mobile phone number", then the classification of the upstream and downstream data is also "personal contact information - mobile phone number", allowing a single piece of data to have multiple classifications.

[0074] In this embodiment, if the upstream node has two separate data sets, such as data A and data B, and the current node needs to combine data A and data B to obtain data C, it can be understood that the current node's data is data obtained by processing data A and data B through certain operations. Therefore, the category of the current data C can be a combination of data A and data B. For example, if the category of data A is "personal contact information" and the category of data B is "age," then the category of data C could be "personal contact information + age."

[0075] In this embodiment, adjusting the upstream hierarchical information by lowering or raising it can be understood as lowering or raising a preset range. The preset range can be set according to time requirements, and this embodiment does not limit this setting. For example, the preset range could be lowering or raising by one level, lowering or raising by two levels, or lowering or raising by three levels, etc.

[0076] In this embodiment, different adjustments can be made depending on the different propagation methods, and the specific adjustment rules can be set according to specific business requirements.

[0077] In this embodiment, upstream classification information can be directly used as the current classification information of the current node. If the propagation method is direct assignment or indirect assignment, upstream hierarchical information can be directly used as the current hierarchical information of the current node. It is understood that in this embodiment, using direct or indirect copying, the current hierarchical information of the current node is the same as the upstream security level information. If the propagation method is masking or encryption, the upstream hierarchical information can be lowered to obtain the current hierarchical information. It is understood that in this embodiment, when using masking or encryption, the current hierarchical information of the current node is obtained by lowering the upstream security level information. If the propagation method is aggregation operation, the upstream hierarchical information can be raised or lowered according to the business characteristics of the aggregation operation to obtain the current hierarchical information. Here, business characteristics can be understood as the characteristics of the actual business. It is understood that in this embodiment, when using aggregation operation propagation, the current hierarchical information of the current node can be obtained by raising or lowering the upstream security level information based on the actual business characteristics of the aggregation operation. If the propagation method uses an irreversible operation function, the upstream security level information can be lowered to obtain the current security level information. Understandably, in this embodiment, when using an irreversible operation function for propagation, the current security level information of the current node can be obtained by lowering the upstream security level information. If the propagation method uses a reversible operation function, the upstream security level information can be lowered based on the reversibility difficulty to obtain the current security level information. The reversibility difficulty can be determined based on the actual situation. Understandably, in this embodiment, when using a reversible operation function for propagation, the current security level information of the current node can be appropriately lowered by adjusting the upstream security level information based on the reversibility difficulty of the reversible function.

[0078] This implementation allows for different tier adjustment methods to be set according to specific business requirements, enabling more flexible adjustment of the current tier information.

[0079] The technical solution of this invention obtains upstream classification information and upstream hierarchical information of data at upstream nodes; obtains lineage propagation information of data from upstream nodes to the current node; determines the data propagation mode based on the lineage propagation information; and adjusts the upstream classification information and upstream hierarchical information according to the data propagation mode to obtain the current classification information and current hierarchical information of data at the current node. This technical solution can improve the accuracy of data classification and data hierarchical.

[0080] Example 2

[0081] Figure 3This is a flowchart of a data processing method according to Embodiment 2 of the present invention. This embodiment is an optimization based on the above embodiment. Specifically, the optimization includes: obtaining the upstream hierarchical confidence level of the data at the previous node; determining an adjustment factor according to the propagation method; adjusting the upstream hierarchical confidence level according to the adjustment factor to obtain the current hierarchical confidence level of the current node. Figure 3 As shown, the method in this embodiment specifically includes the following steps:

[0082] Step 310: Obtain the upstream hierarchical confidence level of the data in the previous node.

[0083] Here, the grading confidence level can be understood as a quantified value of the credibility of the classification and grading results. Since the bloodline transmission analysis in this embodiment is a supplement to the discovery and identification of classification and grading information, it may contain some errors. Therefore, the credibility of the classification and grading results can be quantified by the confidence level value.

[0084] In this embodiment, when initializing the node, the classification and grading information of the initial node data can be obtained through scanning, and an initial grading confidence level can be assigned simultaneously. This embodiment can also obtain the upstream grading confidence level of the data in the previous node.

[0085] Step 320: Determine the adjustment factor based on the propagation method.

[0086] The adjustment factor can be determined based on the propagation method. In this embodiment, the upstream hierarchical confidence level can be adjusted based on the adjustment factor. In this embodiment, the adjustment factor can be determined according to the propagation method; different propagation methods result in different adjustment factors.

[0087] In this embodiment, optionally, determining the adjustment factor according to the propagation method includes: determining the table-level lineage influence factor, the field-level lineage influence factor, and the field propagation influence factor according to the propagation method; and performing a weighted summation of the table-level lineage influence factor, the field-level lineage influence factor, and the field propagation influence factor to obtain the adjustment factor.

[0088] In this embodiment, the influencing factor can be determined by multiple influencing factors. In this embodiment, table-level lineage influencing factor, field-level lineage influencing factor, and field propagation application factor can be determined based on the propagation method; then, the table-level lineage influencing factor, field-level lineage influencing factor, and field propagation influencing factor are weighted and summed to obtain the adjustment factor.

[0089] The adjustment factor can be calculated using the following formula (the sum of all weights in the formula is 1):

[0090] Adjustment factor C = Table-level lineage * weight 1 + Field-level lineage * weight 2 + Field propagation processing * weight 3;

[0091] For example, Table 1 illustrates the calculation of confidence level influence factors using database tables and fields. The parameter values ​​in the table are sample values, and can be specifically set according to business requirements.

[0092]

[0093] Table 1. Explanation of the calculation of the influence factors for confidence level.

[0094] For example, the security level in this embodiment may attenuate during propagation, which can be handled by arithmetic multiplication or other algorithms. For instance, when processing based on arithmetic multiplication, if the confidence level of data A is 80%, and it is propagated to data B according to the indirect assignment in Table 2 above, then referring to Table 2, the confidence level of the classification and grading result of data B is: 80% confidence level of data A * 84% confidence level of indirect assignment = 67.2%.

[0095] This implementation allows for flexible weighting of influencing factors based on different actual business needs, thereby obtaining the confidence level of node classification information and further improving the accuracy of the classification information.

[0096] In addition, in this embodiment, the classification and grading information and confidence level of the data can also be described by labeling. That is, the classification and grading attributes and confidence level of the data can be described as a certain label of the data, or a special description within the label, etc. This embodiment does not limit this.

[0097] Step 330: Adjust the upstream hierarchical confidence level according to the adjustment factor to obtain the current hierarchical confidence level of the current node.

[0098] The current level confidence score of the current node can be obtained by adjusting the upstream level confidence score based on an adjustment factor. Specifically, the adjustment can be achieved by multiplying the upstream level confidence score by the adjustment factor. In this embodiment, the upstream level confidence score can be adjusted based on the adjustment factor to obtain the current level confidence score of the current node.

[0099] In this embodiment, optionally, adjusting the upstream hierarchical confidence level according to the adjustment factor to obtain the current hierarchical confidence level of the current node includes: multiplying the upstream hierarchical confidence level by the adjustment factor to obtain the current hierarchical confidence level of the current node.

[0100] In this embodiment, the current level confidence of the current node can be obtained by multiplying the upstream level confidence by the adjustment factor.

[0101] For example, if the confidence level of an upstream tier is 70% and the adjustment factor is 80%, then the confidence level of the current tier of the current node can be 70% * 80% = 56%.

[0102] In this embodiment, the upstream grading confidence level can be multiplied by an adjustment factor to obtain the current grading confidence level of the current node. Through this setting, the adjustment factor is determined by analyzing the propagation methods during kinship transmission, thereby enabling the calculation of the credibility of the identification results and giving the classification and grading information a quantifiable confidence value.

[0103] Furthermore, in this embodiment, after completing the lineage propagation analysis and the calculation of classification and confidence values, discrepancies in the classification and classification results between some node data can be identified. For example, the classification and classification results of a node's data calculated from multiple propagation paths may be inconsistent or illogical. These discrepancies generally arise from data propagation loops or errors in the classification and classification scanning and identification before the lineage propagation analysis. This embodiment can also identify classification and classification discrepancies between data points through graph operations on the propagation paths. For instance, it may be discovered that upstream high-security-level data propagates downstream through direct assignment, while the upstream data was originally of low priority; this could indicate a classification and classification discrepancy. Propagation path analysis can identify these discrepancies, providing them to users for verification and enhancing the accuracy of the classification and classification results.

[0104] In this embodiment, optionally, it also includes: if the data of the current node comes from multiple upstream nodes, then the data of the current node is classified and graded by scanning.

[0105] In this context, multiple upstream nodes can be understood as two or more upstream nodes.

[0106] In this embodiment, when the current node comes from multiple upstream nodes, the data of the current node can be classified and graded by manual scanning.

[0107] For example, when data at the current node propagates from multiple sources, meaning there are multiple upstream nodes, the data will have multiple classification and grading calculation results. For instance, if data B's upstream data includes data A and data C, then the calculation results of the classification, grading, and confidence scores of data C propagating to data B may be consistent with or inconsistent with the calculation results of data A propagating to data B. This necessitates establishing conflict resolution rules, and these rules must conform to the requirements of the actual business scenario.

[0108] In this embodiment, by setting it up in this way, if the data of the current node comes from multiple upstream nodes, the data of the current node can be classified and graded by scanning, thus avoiding conflicts in the data classification results.

[0109] Example 3

[0110] Figure 4 This is a schematic diagram of a data processing apparatus according to Embodiment 3 of the present invention. This apparatus can execute the data processing method provided in any embodiment of the present invention, and possesses corresponding functional modules and beneficial effects for executing the method. For example... Figure 4 As shown, the device includes:

[0111] The first information acquisition module 410 is used to acquire upstream classification information and upstream hierarchical information of data at upstream nodes;

[0112] The second information acquisition module 420 is used to acquire the lineage propagation information of data from the upstream node to the current node;

[0113] The propagation mode determination module 430 is used to determine the data propagation mode based on the bloodline propagation information;

[0114] The information adjustment module 440 is used to adjust the upstream classification information and the upstream hierarchical information according to the data propagation method, so as to obtain the current classification information and current hierarchical information of the data at the current node.

[0115] Optionally, the data propagation method includes any one of the following: direct assignment propagation, indirect assignment propagation, aggregation operation propagation, masking propagation, encryption propagation, reversible operation function propagation, and irreversible operation function propagation.

[0116] Optional, the information adjustment module 440 is specifically used for:

[0117] Use the upstream classification information directly as the current classification information of the current node;

[0118] If the propagation method is direct assignment propagation or indirect assignment propagation, then the upstream hierarchical information is directly used as the current hierarchical information of the current node;

[0119] If the propagation method is masked propagation or encrypted propagation, then the upstream hierarchical information is lowered to obtain the current hierarchical information;

[0120] If the propagation method is aggregation operation propagation, then the upstream hierarchical information is adjusted up or down according to the business characteristics of the aggregation operation propagation to obtain the current hierarchical information;

[0121] If the propagation method is irreversible operation function processing, then the upstream hierarchical information is lowered to obtain the current hierarchical information;

[0122] If the propagation method is a reversible operation function, then the upstream hierarchical information is lowered according to the reversibility difficulty to obtain the current hierarchical information.

[0123] Optionally, the device further includes:

[0124] The confidence level acquisition module is used to obtain the upstream hierarchical confidence level of the data in the previous node;

[0125] An adjustment factor determination module is used to determine an adjustment factor based on the propagation method;

[0126] The confidence adjustment module is used to adjust the upstream hierarchical confidence level according to the adjustment factor to obtain the current hierarchical confidence level of the current node.

[0127] Optional, the adjustment factor determination module is specifically used for:

[0128] Based on the propagation method, determine the table-level lineage influence factor, the field-level lineage influence factor, and the field propagation influence factor;

[0129] The adjustment factor is obtained by weighted summation of the table-level lineage influence factor, the field-level lineage influence factor, and the field propagation influence factor.

[0130] Optional, confidence adjustment module, specifically used for:

[0131] The current level confidence score of the current node is obtained by multiplying the upstream level confidence score by the adjustment factor.

[0132] Optionally, it also includes: a scanning module, used for:

[0133] If the current node's data comes from multiple upstream nodes;

[0134] The data of the current node is then classified and graded by scanning.

[0135] Optionally, the first information acquisition module 410 is specifically used for:

[0136] Get the scanned data of the current node;

[0137] Data from upstream nodes is obtained based on the scanned data;

[0138] Determine the upstream classification information and upstream hierarchical information of the data at the upstream node.

[0139] The above-described apparatus can execute the methods provided in all the foregoing embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the above methods. Technical details not described in detail in this embodiment can be found in the methods provided in all the foregoing embodiments of the present invention.

[0140] Example 4

[0141] Figure 5This is a schematic diagram of an electronic device according to Embodiment 4 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0142] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0143] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0144] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as data processing methods.

[0145] In some embodiments, the data processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the data processing method by any other suitable means (e.g., by means of firmware).

[0146] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0147] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0148] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0149] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0150] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0151] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0152] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0153] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A data processing method, characterized in that, include: Obtain upstream classification information and upstream hierarchical information of data at upstream nodes; wherein, the upstream classification information is the category information of the data at upstream nodes; and the upstream hierarchical information is the security level information of the data at upstream nodes. Obtain the lineage propagation information of data from upstream nodes to the current node; The data transmission method is determined based on the aforementioned bloodline transmission information; The upstream classification information and the upstream hierarchical information are adjusted according to the data propagation method to obtain the current classification information and the current hierarchical information of the data at the current node; The data propagation method includes any one of the following: direct assignment propagation, indirect assignment propagation, aggregation operation propagation, masking process propagation, encryption process propagation, reversible operation function processing propagation, and irreversible operation function processing propagation; Adjusting the upstream classification information and the upstream hierarchical information according to the data propagation method includes: Use the upstream classification information directly as the current classification information of the current node; If the propagation method is direct assignment propagation or indirect assignment propagation, then the upstream hierarchical information is directly used as the current hierarchical information of the current node; If the propagation method is masked propagation or encrypted propagation, then the upstream hierarchical information is lowered to obtain the current hierarchical information; If the propagation method is aggregation operation propagation, then the upstream hierarchical information is adjusted up or down according to the business characteristics of the aggregation operation propagation to obtain the current hierarchical information; If the propagation method is irreversible operation function processing, then the upstream hierarchical information is lowered to obtain the current hierarchical information; If the propagation method is a reversible operation function, then the upstream hierarchical information is lowered according to the reversibility difficulty to obtain the current hierarchical information.

2. The method according to claim 1, characterized in that, Also includes: Obtain the upstream hierarchical confidence level of the data in the previous node; The adjustment factor is determined based on the propagation method described above; The upstream hierarchical confidence level is adjusted according to the adjustment factor to obtain the current hierarchical confidence level of the current node.

3. The method according to claim 2, characterized in that, Determining the adjustment factor based on the propagation method includes: Based on the propagation method, determine the table-level lineage influence factor, the field-level lineage influence factor, and the field propagation influence factor; The adjustment factor is obtained by weighted summation of the table-level lineage influence factor, the field-level lineage influence factor, and the field propagation influence factor.

4. The method according to claim 2, characterized in that, The upstream hierarchical confidence level is adjusted according to the adjustment factor to obtain the current hierarchical confidence level of the current node, including: The current level confidence score of the current node is obtained by multiplying the upstream level confidence score by the adjustment factor.

5. The method according to claim 1, characterized in that, Also includes: If the current node's data comes from multiple upstream nodes; The data of the current node is then classified and graded by scanning.

6. The method according to claim 1, characterized in that, Obtain upstream classification and hierarchical information of the data at the upstream node, including: Get the scanned data of the current node; Data from upstream nodes is obtained based on the scanned data; Determine the upstream classification information and upstream hierarchical information of the data at the upstream node.

7. A data processing apparatus, characterized in that, include: The first information acquisition module is used to acquire upstream classification information and upstream hierarchical information of data at upstream nodes; wherein, the upstream classification information is the category information of the data at the upstream node; and the upstream hierarchical information is the security level information of the data at the upstream node. The second information acquisition module is used to acquire the lineage propagation information of data from the upstream node to the current node; A propagation mode determination module is used to determine the data propagation mode based on the bloodline propagation information; The information adjustment module is used to adjust the upstream classification information and the upstream hierarchical information according to the data propagation method, so as to obtain the current classification information and current hierarchical information of the data at the current node; The data propagation method includes any one of the following: direct assignment propagation, indirect assignment propagation, aggregation operation propagation, masking process propagation, encryption process propagation, reversible operation function processing propagation, and irreversible operation function processing propagation; The information adjustment module is specifically used for: Use the upstream classification information directly as the current classification information of the current node; If the propagation method is direct assignment propagation or indirect assignment propagation, then the upstream hierarchical information is directly used as the current hierarchical information of the current node; If the propagation method is masked propagation or encrypted propagation, then the upstream hierarchical information is lowered to obtain the current hierarchical information; If the propagation method is aggregation operation propagation, then the upstream hierarchical information is adjusted up or down according to the business characteristics of the aggregation operation propagation to obtain the current hierarchical information; If the propagation method is irreversible operation function processing, then the upstream hierarchical information is lowered to obtain the current hierarchical information; If the propagation method is a reversible operation function, then the upstream hierarchical information is lowered according to the reversibility difficulty to obtain the current hierarchical information.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data processing method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the data processing method according to any one of claims 1-6.