A data classification method for the evolution process of data space information
By obtaining the data operation type and correlation strength evaluation value and dynamically calculating the data level, the problem of repeated grading overhead and level information traceability difficulties in the dynamic evolution scenario of data is solved, the accuracy and consistency of data level is achieved, and efficient management and intelligence of data space are supported.
Patent Information
- Application Number
- CN202510685711.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The existing technology has difficulties in duplicate rating overhead and level information traceability in the dynamic data evolution scenario, and cannot effectively support the efficient management and intelligence of data space.
By obtaining the data operation type and correlation strength evaluation value, dynamically calculate the data level, avoid repeated grading operations, realize data level traceability and consistency, and build a traceability system at the data evolution level.
Reduced redundant calculations, ensured the accuracy and consistency of data level, provided reliable data management and sharing support in multi-node and multi-subject scenarios, and solved the problem of repeated grading overhead.
Smart Images

Figure CN120197083B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data classification, and specifically relates to a data classification method, device, equipment and storage medium for the data space information evolution process. Background Art
[0002] With the advent of the data-driven era, the construction of data spaces has become a core issue in data management and application. By aggregating data from diverse sources and enabling dynamic and complex evolution, data spaces provide crucial support for data analysis and decision-making. In particular, based on Internet of Things (IoD) technologies, data spaces can support the integration and circulation of heterogeneous data, creating the necessary conditions for intelligent and automated data operations.
[0003] Existing technologies have proposed several solutions for data classification, such as Harvard University's Datatags system, which combines knowledge of privacy regulations, data sharing protocols, and security mechanisms to provide an effective approach for scientific data classification. These technologies, centered around static data, design data labeling-based classification strategies, helping researchers classify sensitive data even when they lack specialized knowledge.
[0004] However, these methods are primarily applicable to scenarios where data rarely changes or is static, and their support for complex scenarios involving dynamic data evolution is limited. Furthermore, the Digital Object Protocol in the data space provides basic interaction specifications but does not address the dynamic nature of data classification. This leads to the overhead of repeated classification and the problem of tracing classification information in scenarios where data is dynamically evolving. Summary of the Invention
[0005] The present application aims to provide a data classification method, apparatus, device and storage medium for the data space information evolution process, at least to solve the problem of repeated classification overhead and level information tracing in data classification under the scenario of dynamic data evolution.
[0006] In a first aspect, embodiments of the present application disclose a data classification method for a data space information evolution process, comprising:
[0007] In response to a process of generating second data according to a data operation on first data, obtaining a first data level of the first data and an operation type of the data operation;
[0008] determining, according to the first data and the second data, and the operation type of the data operation, an evaluation value of the association strength between the first data and the second data;
[0009] A second data level of the generated second data is determined according to the first data level, the association strength evaluation value, and the operation type of the data operation.
[0010] In a second aspect, the present application also discloses a data classification device for a data space information evolution process, including:
[0011] an acquisition module, configured to acquire, in response to a process of generating second data according to a data operation on first data, a first data level of the first data and an operation type of the data operation;
[0012] an evaluation module, configured to determine an evaluation value of the strength of association between the first data and the second data according to the first data and the second data, and an operation type of the data operation;
[0013] A grading module is configured to determine a second data level of the generated second data according to the first data level, the association strength evaluation value, and the operation type of the data operation.
[0014] In a third aspect, an embodiment of the present application further discloses an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.
[0015] In a fourth aspect, an embodiment of the present application further discloses a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0016] In summary, in the embodiment of the present application, by obtaining initial information in response to the process of generating data, it is possible to ensure that the data classification process has comprehensive contextual information from the beginning, including the level of the first data and its corresponding data operation type, avoiding the necessity of manual intervention and providing a basis for data classification decisions in subsequent steps; thereby, a calculation method for the association strength evaluation value is introduced to quantify the logical relationship between the original data and the generated data, ensuring that the data level transmission process is more accurate, thereby directly reducing the overhead of repeated classification; and then using the dynamic association strength and operation type to predict and generate the level information of the second data, so as to avoid the separate classification operation for each newly generated data by dynamically generating the level of the second data, reducing redundant calculations from the system architecture level, solving the exponential overhead growth problem caused by the frequent circulation of derivative data, and realizing efficient support for dynamic data evolution scenarios. Thus, based on the method of the embodiment of the present application, the combination of data operation type and level change process is used to achieve the traceability of the association between original data and derivative data, ensuring the accuracy and consistency of data level evolution. Especially in the scenario of multiple nodes and multiple subjects, the traceability mechanism provides reliable support for data management and sharing, overcoming the bottleneck problem of difficult level information traceability in the prior art. It solves the problem of repeated classification overhead in the dynamic evolution of data and successfully builds a traceability system for data evolution levels, providing technical support for efficient management and intelligentization of data space. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In the attached figure:
[0018] Figure 1 This is a flowchart of the steps of a data classification method for the data space information evolution process provided by an embodiment of the present application;
[0019] Figure 2 It is a data-level deduction process under the embodiment of the present application;
[0020] Figure 3 This is a flowchart of another data classification method for the data space information evolution process provided by an embodiment of the present application;
[0021] Figure 4 It is a data level display process under the embodiment of this application;
[0022] Figure 5 This is a block diagram of a data classification device for the data space information evolution process provided by an embodiment of the present application;
[0023] Figure 6 is a block diagram of an electronic device according to an embodiment of the present application;
[0024] Figure 7This is a block diagram of an electronic device according to another embodiment of the present application. DETAILED DESCRIPTION
[0025] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0026] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.
[0027] Consider the following process:
[0028] Assume that the first data set (original data set) is , and the second data set (derivative data set) is , and assume that each data entity The initial level is .
[0029] The set of operation types for data operations is ,
[0030] Data operations may include:
[0031] Downgrade operations (such as desensitization, anonymization, etc.), same-level operations (such as format conversion, text modification, etc.), and upgrade operations (such as multi-source data fusion, etc.).
[0032] Then we can construct a three-dimensional tensor The relationship matrix is established in the form of Represents the original data By operation Generate derived data The strength of association.
[0033] Then the data level can be determined by designing the level transfer function:
[0034] Based on the above process, Figure 1 As shown, a data classification method for the data space information evolution process provided by an embodiment of the present application is shown.
[0035] The method may include the following steps:
[0036] Step 101 : In response to a process of generating second data according to a data operation on first data, a first data level of the first data and an operation type of the data operation are obtained.
[0037] In some embodiments of the present application, in order to ensure that the level information and operation type of the first data can be accurately grasped during the data generation process and to provide the necessary basic data for the subsequent data classification process, the level information of the first data related to the generation of the second data and the data operation type adopted will be recorded in real time by the system during the execution phase of the data operation. The data operation type refers to the specific operation performed during the data processing process, such as data replication and data transformation. The description of the data operation type includes the nature of the relationship between the data and the potential impact on the data classification. In this way, the system can automatically obtain and record contextual information, avoid manual intervention, improve operational efficiency, and provide comprehensive background information for subsequent classification decisions.
[0038] like Figure 2 As shown in the example, when processing data from multiple nodes, the system performs a data fusion operation, combining the multi-source data into a new dataset and automatically calculating the level of the new dataset. The system can then record the original data level information (for example, level 3) and the operation type (for example, data fusion) in real time. This successful recording of the original data level and operation type provides accurate and complete foundational information for predicting the level of the new data.
[0039] Step 102 : Determine an evaluation value of the association strength between the first data and the second data according to the first data and the second data, and the operation type of the data operation.
[0040] In some embodiments of the present application, in order to quantify the logical association between data so that the subsequent data grading process can more accurately reflect the intrinsic relationship between the data, the characteristic values of the first data and the second data, as well as the data operation type, will be used to construct a mathematical model for association strength evaluation and calculate the association strength evaluation value. The association strength evaluation value is a numerical value that measures the degree of logical or semantic connection between data, and is usually determined by analyzing semantic similarity and access permission overlap. The calculation of the association strength evaluation value can directly reflect the association between data, thereby providing technical support for the accuracy of dynamic grading.
[0041] In a specific example, it's necessary to analyze the strength of association between new data generated through data fusion and the original data. At this point, the user has already fused multiple data sources to generate a new derived dataset. The system can then vectorize the text feature values of the original and derived data, calculate their semantic similarity, analyze the overlap of their access permission lists, and then calculate an association strength evaluation based on the set weights. This system successfully generates a quantitative association strength evaluation value to characterize the logical relationship between the first and second data, providing a reliable basis for subsequent data-level prediction and grading.
[0042] Step 103 : determining a second data level of the generated second data according to the first data level, the association strength evaluation value, and the operation type of the data operation.
[0043] In some embodiments of the present application, to achieve the level transfer from original data to derived data and ensure the consistency and accuracy of the data classification and grading process, a mathematical model is used to dynamically calculate the level of the second data based on the level information, association strength evaluation value, and data operation type of the first data. The association strength evaluation value quantifies the logical or semantic relationship between the original data and the derived data, and the data operation type directly affects the result of the level transfer, such as a downgrade operation reduces the level, and an upgrade operation increases the level. This will achieve automatic grading of derived data, avoid the overhead of repeated grading, and ensure the traceability and consistency of the data level evolution process.
[0044] In a specific example, a user performs a data fusion operation on a dataset to generate a new derivative dataset. At this time, the level of the original dataset is 3, and the level of the second data generated by the data fusion operation needs to be redefined. A mathematical model can be used to dynamically predict the level of the second data to be 4 based on the level (3) of the original dataset, the operation type (data fusion), and the calculated association strength evaluation value (for example, 0.85). In this way, the system automatically generates the level information of the new data and records the level evolution process, ensuring the accuracy and efficiency of data classification and grading.
[0045] In summary, in the embodiment of the present application, by obtaining initial information in response to the process of generating data, it is possible to ensure that the data classification process has comprehensive contextual information from the beginning, including the level of the first data and its corresponding data operation type, avoiding the necessity of manual intervention and providing a basis for data classification decisions in subsequent steps; thereby, a calculation method for the association strength evaluation value is introduced to quantify the logical relationship between the original data and the generated data, ensuring that the data level transmission process is more accurate, thereby directly reducing the overhead of repeated classification; and then using the dynamic association strength and operation type to predict and generate the level information of the second data, so as to avoid the separate classification operation for each newly generated data by dynamically generating the level of the second data, reducing redundant calculations from the system architecture level, solving the exponential overhead growth problem caused by the frequent circulation of derivative data, and realizing efficient support for dynamic data evolution scenarios. Thus, based on the method of the embodiment of the present application, the combination of data operation type and level change process is used to achieve the traceability of the association between original data and derivative data, ensuring the accuracy and consistency of data level evolution. Especially in the scenario of multiple nodes and multiple subjects, the traceability mechanism provides reliable support for data management and sharing, overcoming the bottleneck problem of difficult level information traceability in the prior art. It solves the problem of repeated classification overhead in the dynamic evolution of data and successfully builds a traceability system for data evolution levels, providing technical support for efficient management and intelligentization of data space.
[0046] Figure 3 This is another data classification method for the data space information evolution process provided by the embodiment of the present application.
[0047] The method may include the following steps:
[0048] Step 201 : In response to a process of generating second data according to a data operation on first data, a first data level of the first data and an operation type of the data operation are obtained.
[0049] The method shown in this step has been described in step 101 and will not be repeated here.
[0050] Step 202 : Determine an evaluation value of the association strength between the first data and the second data according to the first data and the second data, and the operation type of the data operation.
[0051] The method shown in this step has been described in step 102 and will not be repeated here.
[0052] Optionally, step 202 includes the following sub-steps:
[0053] Sub-step 2021: determining a similarity evaluation value between the first data and the second data based on the contents of the first data and the second data, and determining a permission evaluation value from the first data to the second data based on the operation type of the data operation.
[0054] In some embodiments of the present application, in order to quantify the content relevance and authority matching degree between the first data and the second data, and to provide basic data for the subsequent calculation of the association strength evaluation value, the semantic similarity evaluation value of the first data and the second data can be calculated by extracting the content information of the first data and the second data, and at the same time, the authority information is analyzed according to the operation type of the data operation to determine the authority evaluation value. The semantic similarity evaluation value measures the degree of semantic relevance between the data contents, and usually calculates the similarity between texts through word vector technology. The authority evaluation value measures the overlap of the access rights between the two data, and usually characterizes the overlap by calculating the intersection ratio of the access control list (ACL) through the Jaccard coefficient. This can provide high-quality input data for the subsequent calculation of the association strength evaluation value, thereby improving the accuracy and reliability of the association strength calculation.
[0055] In a specific example, it is necessary to evaluate the strength of association between first and second data. In this case, the first data is a text file containing user behavior data, and the second data is an analysis report generated based on the first data. The text content of the first and second data can be vectorized using word embedding methods, and semantic similarity can be calculated using cosine similarity, resulting in a similarity rating of 0.85. Simultaneously, by analyzing the access control lists (ACLs) of the two, the permission overlap is calculated to be 0.75, and the permission rating is determined to be 0.75. This system successfully generates both semantic similarity and permission ratings, providing accurate input for the weighted calculation of the association strength rating in the next step.
[0056] Sub-step 2022: determining the weighted sum of the similarity evaluation value and the authority evaluation value as the association strength evaluation value.
[0057] In some embodiments of the present application, in order to comprehensively reflect the degree of semantic relevance of data content and the degree of overlap of access rights, thereby providing an accurate quantitative basis for further data level derivation, the similarity evaluation value and the permission evaluation value will be linearly weighted, and their influence weights will be comprehensively considered to form a unified association strength evaluation value. The linear weighted method regulates the proportion of the similarity evaluation value and the permission evaluation value in the final result through weight coefficients, wherein the similarity evaluation value reflects the semantic relevance of the content, and the permission evaluation value reflects the degree of access control overlap. The clear technical effect brought about by executing this step is the generation of a quantitative association strength evaluation value, which serves as the core reference for subsequent level derivation and operation impact assessment, thereby improving the accuracy and consistency of the data grading process.
[0058] In a specific example, it is necessary to evaluate the strength of the association between the first and second data to support predictions at the derived data level. In this case, the first data is a summary text of the user's personal information data, and the second data is a statistical report generated based on the summary. The execution process of the example includes: calculating the similarity evaluation value of the first and second data as 0.85, and the authority evaluation value as 0.70. Then, weighted calculations are performed on the similarity evaluation value and the authority evaluation value according to weights of 0.6 and 0.4, respectively, and the final association strength evaluation value is 0.79. In this way, the quantitative association strength between the first and second data can be successfully determined, providing a clear reference for subsequent level calculations.
[0059] Based on the above scheme, the association strength evaluation value The value of can be calculated as follows:
[0060]
[0061] in:
[0062] Representing the semantic similarity between the original data and the derived data: The textual relevance between the original data and the derived data can be calculated using word vectors. This involves converting the original data and the derived data into vectors, and then calculating the similarity between the vectors. The degree of permission overlap between the original data and the derived data can be represented by calculating the intersection ratio of the access control lists (ACLs) of the original data and the derived data using the Jaccard coefficient. For the corresponding weights, the proportion of semantics and permissions can be adjusted as needed.
[0063] Step 203 : Determine a second data level of the generated second data according to the first data level, the association strength evaluation value, and the operation type of the data operation.
[0064] The method shown in this step has been described in step 103 and will not be repeated here.
[0065] Optionally, step 203 includes the following sub-steps:
[0066] Sub-step 2031: determining a target influence coefficient corresponding to the operation type according to the operation type of the data operation.
[0067] In some embodiments of the present application, in order to quantify and standardize the impact of different data operations on the derived data level, thereby enhancing the accuracy and consistency of data level predictions, a target impact coefficient related to the operation type will be selected according to the characteristics of the operation type, such as downgrade operation, same-level operation and upgrade operation. The target impact coefficient is a numerical parameter used to describe the specific impact of the operation type on the data level transfer process. For example, the impact coefficient corresponding to the downgrade operation ranges from 0.2 to 0.8, the impact coefficient corresponding to the same-level operation is 1.0, and the typical impact coefficient of the upgrade operation is 1.1 to 1.5. The description of the target impact coefficient includes its important role in the dynamic weighting process of the data level. The clear technical effect brought about by executing this step is to achieve a quantitative representation of the impact of data operations, which provides an accurate and dynamic adjustment basis for the subsequent calculation of data levels.
[0068] In a specific example, the user needs to analyze the impact of a multi-source data fusion operation on the derived data level. The user has already generated new derived data through multi-source data fusion and wishes to calculate the level of this data. Based on the data operation type, a target impact coefficient of 1.2 (corresponding to a multi-source data fusion operation) can be set as the weight parameter in the subsequent weighted calculation. This successful setting of the target impact coefficient of 1.2 provides important input for the weighted calculation of association strength and the first data level, resulting in a more accurate prediction of the derived data level.
[0069] Sub-step 2032: Using the product of the target influence coefficient and the association strength evaluation value as a weight, a weighted sum of the first data level is calculated as the second data level.
[0070] In some embodiments of the present application, in order to combine the target impact coefficient and the association strength evaluation value, dynamically adjust the impact of the first data level on the second data level, and thus improve the accuracy of data level prediction, the target impact coefficient is multiplied by the association strength evaluation value to obtain a comprehensive weight value, and then the weight value is used to perform weighted calculation on the first data level to finally obtain the value of the second data level. The target impact coefficient reflects the degree of influence of the data operation, such as the reduction of the data level by the downgrade operation or the improvement of the data level by the upgrade operation, while the association strength evaluation value quantifies the logical and authority relationship between the original data and the derived data. This comprehensive weighted method ensures a reasonable balance of different factors in the derivation of the data level. In this way, the second data level can be accurately derived and the computational overhead caused by repeated classification and grading can be significantly reduced.
[0071] In a specific example, consider calculating the level of derived data generated through a multi-source data fusion operation. The first data levels are 3 and 4, the target impact coefficient for the multi-source data fusion operation is 1.2, and the association strength evaluation value is 0.85. First, the target impact coefficient is multiplied by the association strength evaluation value, resulting in a weight of 1.02. This weight is then used to perform a weighted sum of the first data levels (3 and 4), ultimately yielding a level of 3.84 for the second data (which will be discretized in a subsequent step). This system successfully predicts the continuous value of the second data level, providing the foundational data for discretization and subsequent analysis.
[0072] The level transfer function designed by the above embodiment can be expressed by the following formula:
[0073] ;
[0074] in: is the target influence coefficient. For the second data level, is the first data level, is the evaluation value of the association strength.
[0075] After verification, in some embodiments, the influence coefficient can be set according to the following specific data:
[0076] For downgrade operations (such as desensitization): , the typical value is 0.5; for operations at the same level (such as format conversion): , indicating that it maintains a neutral impact; for upgrade operations (such as data fusion) , typical value 1.2.
[0077] In some embodiments of the present application, in order to determine the specific value of the influence coefficient of each data operation behavior, calculation can be performed using the following method:
[0078] Step 2033: collect first test data and second test data according to the target data to be verified.
[0079] The second test data is test data generated by operating the first test data according to the target data; the first test data has a corresponding first test data level, and the second test data has a corresponding second test data level.
[0080] In some embodiments of the present application, in order to clarify the operational impact relationship between the first test data and the second test data through confirmatory data collection and lay the foundation for the subsequent determination of the target impact coefficient, typical representative data will be selected as the first test data, and the second test data will be generated based on the specified target data operation, while ensuring that the classification and grading information of each data is recorded. Target data operation refers to the act of performing specific processing or processing on data, such as data fusion or data anonymization. In this way, a set of paired test data is successfully generated, providing high-quality data support for analyzing the impact of the target data operation.
[0081] In a specific example, the impact of a data fusion operation on data level needs to be verified. A set of initial data files with a level of 3 are selected from the database as the first test data. The example execution process includes performing a data fusion operation on the first test data, generating a new set of derivative data files as the second test data, and recording the level information of the second test data as 4. The resulting paired test data (first and second test data) and their corresponding level information provide the basic data for calculating the target impact coefficient.
[0082] Step 2034 : Determine a second verification data level of the generated second test data according to the first test data level, the association strength evaluation value between the first test data and the second test data, and the operation type of the data operation.
[0083] In some embodiments of the present application, in order to provide more accurate data level prediction results based on the association and operational characteristics between test data, and to verify the applicability and accuracy of the target impact coefficient, the first test data level will be combined with its association strength evaluation value with the second test data, and the second verification data level of the second test data will be dynamically calculated using the rules or mathematical models defined by the operation type of the data operation. Data operation types include downgrade operations (such as desensitization), same-level operations (such as format conversion), and upgrade operations (such as multi-source data fusion), each of which will have a different impact on the level prediction of the test data. The second verification data level generated in this way can lay the foundation for further evaluating the errors between test data levels, while improving the credibility and accuracy of the verification process.
[0084] In a specific example, test data is needed to verify the suitability of a target impact coefficient for a certain data fusion operation. In this case, the first test data has a level of 3. The second test data is generated through a data fusion operation, has an estimated correlation strength of 0.85 with the first test data, and the data operation type is an upgrade. Based on the weighting rules set for the correlation strength and the data operation type, and combined with the level 3 of the first test data, a second verification data level of 3.8 can be calculated for the second test data. This second verification data level is used to compare the second test data with its actual level to assess the suitability of the target impact coefficient.
[0085] Step 2035: Determine a target influence coefficient corresponding to the target data operation based on the overall error determined by the second test data level and the second verification data level.
[0086] In some embodiments of the present application, in order to provide a quantitative basis for the optimization of the target influence coefficient by analyzing the differences between the test data levels, thereby improving the applicability and accuracy of the data level prediction process, the second test data level and the second verification data level are compared, and the error between the two is calculated. By summarizing the error values of multiple groups of test data, the target influence coefficient is dynamically adjusted using the error minimization criterion. The overall error is calculated by the average or cumulative value of the errors of all test data groups to quantify the scope of influence of data operation behavior on the level prediction results. In this way, by optimizing the target influence coefficient, the accuracy and reliability of the data classification and grading information evolution process are significantly improved.
[0087] In a specific example, it is necessary to verify the rationality of the impact coefficient value of the data fusion operation. In this case, the second test data level is 4, while the system predicts the second verification data level to be 3.8. The absolute error between the two can be calculated as 0.2. After summing up the errors of all test data, the error minimization algorithm is used to optimize the target impact coefficient from the original setting of 1.2 to 1.25. After following the example execution process, the clear result is that the target impact coefficient is successfully adjusted to 1.25, more accurately reflecting the actual impact of the data fusion operation on the level prediction result.
[0088] Optionally, in some embodiments of the present application, the influence coefficient when the overall error is minimized may be determined as the target influence coefficient.
[0089] In some embodiments of the present application, in order to minimize the error in data level prediction by optimizing the value of the target influence coefficient, thereby improving the accuracy of the data grading process, the prediction error of the test data can be statistically analyzed, the overall error value under each candidate influence coefficient can be calculated, and the corresponding influence coefficient can be selected as the target influence coefficient according to the error minimization criterion. The overall error is an indicator that measures the degree of deviation between the predicted level and the actual level, and is usually quantified by the cumulative absolute error or the mean square error. The clear technical effect brought about by executing this step is to automatically determine the target influence coefficient with the best adaptability, provide accurate parameter support for data grading decisions, and ensure the stability of the prediction results.
[0090] In a specific example, it is necessary to determine the target influence coefficient for a data fusion operation. Multiple data sets have been tested using the data fusion operation, and the actual level of each test data set exhibits a certain error between the predicted level and the actual level. The overall error can be calculated for different candidate influence coefficients (for example, 1.1, 1.2, and 1.3). For example, the mean squared error (MSE) is 0.04, 0.02, and 0.05, respectively. A comparison reveals that an influence coefficient of 1.2 minimizes the overall error (MSE = 0.02). This influence coefficient of 1.2 is thus determined as the target influence coefficient for subsequent level prediction.
[0091] The above process can be transformed into the following optimization model:
[0092] ,
[0093] in, is with The corresponding influence coefficient is is the second verification data level, is the second test data level, Function is used to obtain the minimum number of classification errors .
[0094] Step 204: Correct the determined second data level according to a preset rounding mapping function.
[0095] In some embodiments of the present application, in order to map continuous level results to discrete classification levels so that the classification results are more in line with the needs of actual application scenarios, the second data level (continuous value) obtained through dynamic weighted calculation is converted into corresponding discrete level information based on a sensitive level threshold function. The sensitive level threshold function is a mathematical rule used to segmentally map the level calculation value. The accuracy and consistency of the conversion process are ensured by defining a reasonable threshold range (for example, level 1 corresponds to the range from 0 to T1, and level 2 corresponds to the range from T1 to T2). This eliminates the uncertainty of continuous values that may exist in the calculation results and ensures the standardization and uniformity of data level information.
[0096] In a specific example, the level of new data after multi-source data fusion is dynamically calculated, resulting in a continuous value, such as 2.87. This continuous value needs to be normalized to a discrete categorical level for subsequent security policy formulation. Using the sensitivity level threshold function, the continuous value 2.87 can be mapped to the threshold range [T2, T3] and corrected to a discrete level of 3. This successfully corrects the level of the new data to 3, ensuring the applicability and consistency of the level information in the actual system.
[0097] In an embodiment where the maximum data level is 4, the above process can be expressed by the following formula:
[0098] ,
[0099] in: is the second data level obtained by directly calculating the level transfer function, , , is the segmentation threshold, which can be determined according to the actual application scenario.
[0100] Step 205 : Displaying, in a relationship diagram, a process of generating the second data according to the data operation on the first data based on the determined second data level.
[0101] In which, the nodes in the relationship graph are used to represent the first data and / or the second data, the node value of each node is used to represent the first data level of the first data and / or the second data level of the second data corresponding to the node, and the edges in the relationship graph are used to represent the data operation from the first data to the second data.
[0102] like Figure 4As shown, in some embodiments of the present application, in order to display the evolution process of the data level and the association relationship between the data in an intuitive graphical way, so as to facilitate managers to visually track the data classification and grading results, nodes for representing the first data and / or the second data can be created in the relationship graph, and edges for representing data operations can be added between the nodes, and node values are attached to each node to represent the data level. The node value (NodeValue) in the relationship graph represents the classification and grading information of the data corresponding to the node, and the edge (Edge) is used to represent the data operation type and influence relationship between the nodes. The clear technical effect brought about by executing this step is to graphically display the evolution process of the data level and the association operation through the relationship, so that the results of data classification and grading are clearer and convenient for subsequent analysis and management.
[0103] In a specific example, a new dataset generated through data fusion is processed, and it is desirable to record the data level evolution information for easy tracking. In this example scenario, the original dataset has levels L1, L2, and L1, respectively. The new dataset generated through the data fusion operation has levels L2 and L1, respectively. Based on this data, data at levels L3 and L1 are then generated. This allows for accurate and intuitive visualization of the level transition process from original data to derived data, as well as their logical associations, through a relationship diagram, providing effective support for data management and level tracing.
[0104] In summary, in the embodiment of the present application, by obtaining initial information in response to the process of generating data, it is possible to ensure that the data classification process has comprehensive contextual information from the beginning, including the level of the first data and its corresponding data operation type, avoiding the necessity of manual intervention and providing a basis for data classification decisions in subsequent steps; thereby, a calculation method for the association strength evaluation value is introduced to quantify the logical relationship between the original data and the generated data, ensuring that the data level transmission process is more accurate, thereby directly reducing the overhead of repeated classification; and then using the dynamic association strength and operation type to predict and generate the level information of the second data, so as to avoid the separate classification operation for each newly generated data by dynamically generating the level of the second data, reducing redundant calculations from the system architecture level, solving the exponential overhead growth problem caused by the frequent circulation of derivative data, and realizing efficient support for dynamic data evolution scenarios. Thus, based on the method of the embodiment of the present application, the combination of data operation type and level change process is used to achieve the traceability of the association between original data and derivative data, ensuring the accuracy and consistency of data level evolution. Especially in the scenario of multiple nodes and multiple subjects, the traceability mechanism provides reliable support for data management and sharing, overcoming the bottleneck problem of difficult level information traceability in the prior art. It solves the problem of repeated classification overhead in the dynamic evolution of data and successfully builds a traceability system for data evolution levels, providing technical support for efficient management and intelligentization of data space.
[0105] refer to Figure 5 , which shows a data classification device 30 for the data space information evolution process provided by an embodiment of the present application, including:
[0106] An acquisition module 301 is configured to acquire a first data level of the first data and an operation type of the data operation in response to a process of generating second data according to a data operation on the first data;
[0107] Evaluation module 302, configured to determine an evaluation value of the strength of association between the first data and the second data according to the first data and the second data, and the operation type of the data operation;
[0108] The grading module 303 is configured to determine a second data grade of the generated second data according to the first data grade, the association strength evaluation value, and the operation type of the data operation.
[0109] Optionally, when the data types of the first data and the second data are the same, the evaluation module 302 includes:
[0110] a parameter submodule, configured to determine a similarity evaluation value between the first data and the second data based on the contents of the first data and the second data, and to determine a permission evaluation value from the first data to the second data based on the operation type of the data operation;
[0111] The evaluation submodule is used to determine the weighted sum of the similarity evaluation value and the authority evaluation value as the association strength evaluation value.
[0112] Optionally, the grading module 303 includes:
[0113] A coefficient submodule is used to determine a target impact coefficient corresponding to an operation type according to the operation type of the data operation;
[0114] The grading submodule is used to use the product of the target influence coefficient and the association strength evaluation value as a weight, and calculate the weighted sum of the first data level as the second data level.
[0115] Optionally, the data classification device 30 for the data space information evolution process further includes:
[0116] The rounding module is used to correct the determined second data level according to a preset rounding mapping function.
[0117] Optionally, the data classification device 30 for the data space information evolution process further includes:
[0118] A test set module is configured to collect first test data and second test data according to a target data operation to be verified; the second test data is test data generated by the first test data according to the target data operation; the first test data has a corresponding first test data level, and the second test data has a corresponding second test data level;
[0119] a testing module, configured to determine a second verification data level of the generated second test data based on the first test data level, an evaluation value of the strength of association between the first test data and the second test data, and an operation type of the data operation;
[0120] The coefficient determination module is used to determine a target influence coefficient corresponding to the target data operation according to the overall error determined by the second test data level and the second verification data level.
[0121] Optionally, the coefficient determination module includes:
[0122] The coefficient determination submodule is used to determine the influence coefficient when the overall error is minimized as the target influence coefficient.
[0123] Optionally, the data classification device 30 for the data space information evolution process further includes:
[0124] A display module is used to display in a relationship graph a process of generating second data based on a data operation on first data according to a determined second data level; the nodes in the relationship graph are used to represent the first data and / or the second data, the node value of each node is used to represent the first data level of the first data and / or the second data level of the second data corresponding to the node, and the edges in the relationship graph are used to represent the data operation from the first data to the second data.
[0125] In summary, in the embodiment of the present application, by obtaining initial information in response to the process of generating data, it is possible to ensure that the data classification process has comprehensive contextual information from the beginning, including the level of the first data and its corresponding data operation type, avoiding the necessity of manual intervention and providing a basis for data classification decisions in subsequent steps; thereby, a calculation method for the association strength evaluation value is introduced to quantify the logical relationship between the original data and the generated data, ensuring that the data level transmission process is more accurate, thereby directly reducing the overhead of repeated classification; and then using the dynamic association strength and operation type to predict and generate the level information of the second data, so as to avoid the separate classification operation for each newly generated data by dynamically generating the level of the second data, reducing redundant calculations from the system architecture level, solving the exponential overhead growth problem caused by the frequent circulation of derivative data, and realizing efficient support for dynamic data evolution scenarios. Thus, based on the method of the embodiment of the present application, the combination of data operation type and level change process is used to achieve the traceability of the association between original data and derivative data, ensuring the accuracy and consistency of data level evolution. Especially in the scenario of multiple nodes and multiple subjects, the traceability mechanism provides reliable support for data management and sharing, overcoming the bottleneck problem of difficult level information traceability in the prior art. It solves the problem of repeated classification overhead in the dynamic evolution of data and successfully builds a traceability system for data evolution levels, providing technical support for efficient management and intelligentization of data space.
[0126] Reference Figure 6 , electronic device 500 may include one or more of the following components: a processing component 502 , a memory 504 , a power component 506 , a multimedia component 508 , an audio component 510 , an input / output (I / O) interface 512 , a sensor component 514 , and a communication component 516 .
[0127] The processing component 502 generally controls the overall operation of the electronic device 500, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 502 may include one or more processors 520 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 502 may include one or more modules to facilitate interaction between the processing component 502 and other components. For example, the processing component 502 may include a multimedia module to facilitate interaction between the multimedia component 508 and the processing component 502.
[0128] The memory 504 is used to store various types of data to support operations on the electronic device 500. Examples of such data include instructions for any application or method operating on the electronic device 500, contact data, phone book data, messages, pictures, multimedia, etc. The memory 504 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0129] The power supply assembly 506 provides power to the various components of the electronic device 500. The power supply assembly 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 500.
[0130] The multimedia component 508 includes an interface that provides an output interface between the electronic device 500 and the user. In some embodiments, the interface may include a liquid crystal display (LCD) and a touch panel (TP). If the interface includes a touch panel, the interface may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensors can not only sense the demarcation of a touch or slide action, but also detect the duration and pressure associated with the touch or slide action. In some embodiments, the multimedia component 508 includes a front-facing camera and / or a rear-facing camera. When the electronic device 500 is in an operating mode, such as a capture mode or a multimedia mode, the front-facing camera and / or the rear-facing camera can receive external multimedia data. Each front-facing camera and the rear-facing camera can have a fixed optical lens system or have focal length and optical zoom capabilities.
[0131] The audio component 510 is used to output and / or input audio signals. For example, the audio component 510 includes a microphone (MIC) that receives external audio signals when the electronic device 500 is in an operating mode, such as a call mode, a recording mode, or a voice recognition mode. The received audio signals may be further stored in the memory 504 or transmitted via the communication component 516. In some embodiments, the audio component 510 also includes a speaker for outputting audio signals.
[0132] The input / output I / O interface 512 provides an interface between the processing component 502 and peripheral interface modules, such as a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0133] The sensor assembly 514 includes one or more sensors for providing various aspects of status assessment for the electronic device 500. For example, the sensor assembly 514 can detect the open / closed state of the electronic device 500, the relative positioning of components, such as the display and keypad of the electronic device 500. The sensor assembly 514 can also detect changes in the position of the electronic device 500 or a component of the electronic device 500, the presence or absence of user contact with the electronic device 500, the orientation or acceleration / deceleration of the electronic device 500, and temperature changes of the electronic device 500. The sensor assembly 514 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 514 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 514 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0134] The communication component 516 is used to facilitate wired or wireless communication between the electronic device 500 and other devices. The electronic device 500 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 516 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 516 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0135] In an exemplary embodiment, the electronic device 500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to implement the methods provided in the embodiments of the present application.
[0136] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 504 including instructions, which can be executed by the processor 520 of the electronic device 500 to perform the above method. For example, the non-transitory storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0137] Figure 7 FIG2 is a block diagram of an electronic device 600 according to another embodiment of the present invention. For example, the electronic device 600 may be provided as a server.
[0138] Reference Figure 7 The electronic device 600 includes a processing component 622, which further includes one or more processors, and a memory resource represented by a memory 632 for storing instructions executable by the processing component 622, such as an application. The application stored in the memory 632 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 622 is configured to execute the instructions to perform the method provided in the embodiments of the present application.
[0139] The electronic device 600 may further include a power supply component 626 configured to perform power management of the electronic device 600, a wired or wireless network interface 650 configured to connect the electronic device 600 to a network, and an input / output (I / O) interface 658. The electronic device 600 may operate based on an operating system stored in the memory 632, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.
[0140] It should be noted that, for the sake of simplicity, the method embodiments of the present application are described as a series of action combinations. However, those skilled in the art should be aware that the embodiments of the present application are not limited by the order of the actions described, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present application.
[0141] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the application disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0142] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A data classification method for the data space information evolution process, characterized in that: include: In response to a process of generating second data according to a data operation on first data, obtaining a first data level of the first data and an operation type of the data operation; determining, according to the first data and the second data, and the operation type of the data operation, an evaluation value of the association strength between the first data and the second data; determining a second data level of the generated second data according to the first data level, the association strength evaluation value, and the operation type of the data operation; In a case where the data types of the first data and the second data are the same, determining the association strength evaluation value between the first data and the second data according to the first data and the second data, and the operation type of the data operation, includes: Determining a similarity evaluation value between the first data and the second data based on the contents of the first data and the second data, and determining a permission evaluation value from the first data to the second data based on the operation type of the data operation; A weighted sum of the similarity evaluation value and the authority evaluation value is determined as the association strength evaluation value.
2. The data classification method for data space information evolution process according to claim 1, characterized in that: The determining, according to the first data level, the association strength evaluation value, and the operation type of the data operation, of the second data level generated includes: Determining a target impact coefficient corresponding to the operation type according to the operation type of the data operation; The product of the target influence coefficient and the association strength evaluation value is used as a weight, and a weighted sum of the first data levels is calculated as the second data level.
3. The data classification method for data space information evolution process according to claim 2, characterized in that: The data classification method for the data space information evolution process also includes: Collecting first test data and second test data according to target data to be verified; the second test data is test data generated by operating the first test data according to the target data; the first test data has a corresponding first test data level, and the second test data has a corresponding second test data level; determining a second verification data level of the generated second test data according to the first test data level, an evaluation value of the strength of association between the first test data and the second test data, and an operation type of the data operation; A target influence coefficient corresponding to the target data operation is determined according to an overall error determined by the second test data level and the second verification data level.
4. The data classification method for data space information evolution process according to claim 1, characterized in that: The data classification method for the data space information evolution process also includes: The determined second data level is corrected according to a preset rounding mapping function.
5. The data classification method for data space information evolution process according to claim 1, characterized in that: The data classification method for the data space information evolution process also includes: According to the determined second data level, the process of generating the second data based on the data operation on the first data is displayed in a relationship graph; the nodes in the relationship graph are used to represent the first data and / or the second data, and the node value of each node is used to represent the first data level of the first data and / or the second data level of the second data corresponding to the node, and the edges in the relationship graph are used to represent the data operation from the first data to the second data.
6. A data classification device for the data space information evolution process, characterized in that: include: an acquisition module, configured to acquire, in response to a process of generating second data according to a data operation on first data, a first data level of the first data and an operation type of the data operation; an evaluation module, configured to determine an evaluation value of the strength of association between the first data and the second data according to the first data and the second data, and an operation type of the data operation; a grading module, configured to determine a second data level of the generated second data according to the first data level, the association strength evaluation value, and the operation type of the data operation; In the case that the data types of the first data and the second data are the same, the evaluation module includes: a parameter submodule, configured to determine a similarity evaluation value between the first data and the second data based on the contents of the first data and the second data, and to determine a permission evaluation value from the first data to the second data based on the operation type of the data operation; The evaluation submodule is configured to determine a weighted sum of the similarity evaluation value and the authority evaluation value as the association strength evaluation value.
7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the data classification method for the data space information evolution process according to any one of claims 1 to 5 is implemented.
8. An electronic device, characterized in that: The method comprises a processor, a memory and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the data classification method for the data space information evolution process as described in any one of claims 1 to 5.
9. A computer program product, characterized in that The computer program product stores a computer program, and when the computer program is executed by a processor, the steps of the data classification method for the data space information evolution process according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Sensitive data identification protection method and system based on intelligent matching
CN119577815A