Data space information evolution process-oriented data grading method

By responding to data operations in the dynamic evolution scenario of data, obtaining levels and operation types, calculating the correlation intensity evaluation value, and dynamically determining the data level, the repetition overhead and level information traceability problems of data rating are solved, and the precise transmission and traceability of data level are achieved.

CN120197083AActive Publication Date: 2025-06-24BEIJING BIG DATA ADVANCED TECH RES INST
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510685711.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-24
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

In the dynamic data evolution scenario, the prior art is difficult to effectively solve the problem of repeated grading overhead for data grading and the traceability of level information.

Method used

By responding to the data operation process, the level and operation type of the first data are obtained, the correlation strength evaluation value is calculated, and the level of the second data is dynamically determined based on this information, avoiding repeated grading and level information traceability.

Benefits of technology

It realizes accurate data-level transmission, reduces the overhead of repeated grading, ensures data-level traceability and consistency, and supports efficient management and intelligence of data space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197083A_ABST
    Figure CN120197083A_ABST
Patent Text Reader

Abstract

The invention discloses a data spatial information evolution process-oriented data grading method, which comprises the following steps of: in response to a process of generating second data according to a data operation on first data, obtaining a first data level of the first data and an operation type of the data operation; determining an association strength evaluation value between the first data and the second data according to the first data, the second data and an operation type of the data operation; and determining a second data level of the generated second data according to the first data level, the association strength evaluation value and the operation type of the data operation. The problem of repeated grading overhead in the data dynamic evolution process is solved, a traceable system of data evolution levels is successfully constructed, and technical guarantee is provided for efficient management and intelligentization of data space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data classification, and particularly relates to a data classification method, device, equipment, and storage medium for the evolution process of data space information. Background Art

[0002] With the advent of the data-driven era, the construction of data space has become the core issue of data management and application. By aggregating data from different sources and realizing dynamic and complex evolution processes, data space provides important support for data analysis and decision-making. Especially based on the data networking technology, data space can support the integration and circulation of heterogeneous data, creating necessary conditions for intelligent and automated data operations.

[0003] Existing technologies have mainly proposed several solutions for data classification. For example, the DataTags system at Harvard University provides an effective method for scientific data classification by combining privacy regulation knowledge, data sharing protocols, and security mechanisms. These technologies focus on static data and design classification strategies based on data tags to help researchers classify sensitive data in the absence of professional knowledge.

[0004] However, these methods are mainly applicable to scenarios with less data change or static data, and have limited support for the complex situations of data dynamic evolution. In addition, the digital object protocol in the data space provides basic interaction specifications, but does not propose specific solutions for the dynamic changes of data classification, which further causes the problems of repeated classification overhead and traceability of level information in the data dynamic evolution scenario. Summary of the Invention

[0005] This application aims to provide a data classification method, device, equipment, and storage medium for the evolution process of data space information, and at least solve the problems of repeated classification overhead and traceability of level information in the data dynamic evolution scenario.

[0006] In a first aspect, an embodiment of this application discloses a data classification method for the evolution process of data space information, including: In response to the process of generating the second data according to the data operation on the first data, obtaining the first data level of the first data and the operation type of the data operation; Determining the evaluation value of the association strength between the first data and the second data according to the first data and the second data, and the operation type of the data operation; Determining the second data level of the generated second data according to the first data level, the evaluation value of the association strength, and the operation type of the data operation.

[0007] In a second aspect, an embodiment of the present application further discloses a data classification device for the data space information evolution process, including: A collection module, configured to obtain a first data level of the first data and an operation type of the data operation in response to a process of generating second data according to a data operation on the first data; An evaluation module, configured to determine an evaluation value of the association strength between the first data and the second data according to the first data and the second data, and the operation type of the data operation; A classification module, configured to determine a second data level of the generated second data according to the first data level, the evaluation value of the association strength, and the operation type of the data operation.

[0008] In a third aspect, an embodiment of the present application further discloses an electronic device, including a processor and a memory, where the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0009] In a fourth aspect, an embodiment of the present application further discloses a readable storage medium, where a program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0010] In summary, in the embodiment of the present application, by obtaining initial information in response to the process of generating data, it is possible to ensure that the data classification process has comprehensive context information from the beginning, including the level of the first data and its corresponding data operation type, avoiding the need for manual intervention, and providing a basis for data classification decisions in subsequent steps; furthermore, a calculation method for the evaluation value of the association strength is introduced to quantify the logical relationship between the original data and the generated data, ensuring that the process of data level transmission is more accurate, thereby directly reducing the overhead of repeated classification; then, using the dynamic association strength and operation type, the level information of the second data is predicted and generated, so as to avoid the separate classification operation for each newly generated data, reducing redundant calculations at the system architecture level, solving the problem of exponential overhead growth caused by frequent circulation of derivative data, and at the same time realizing efficient support for dynamic data evolution scenarios. Thus, based on the method of the embodiment of the present application, by combining the data operation type with the level change process, the traceability of the association between the original data and the derivative data is realized, ensuring the accuracy and consistency of data level evolution. Especially in the scenario of multiple nodes and multiple subjects, this traceability mechanism provides reliable support for data management and sharing, overcoming the bottleneck problem of difficult traceability of level information in the prior art. It solves the problem of repeated classification overhead in the process of data dynamic evolution, and also successfully constructs a traceable system for data evolution levels, providing technical guarantee for the efficient management and intelligence of the data space. Description of the Drawings

[0011] In the drawings: Figure 1 is a flowchart of steps of a data classification method for the data space information evolution process provided by an embodiment of the present application; Figure 2 is a data level deduction process under an embodiment of the present application; Figure 3 is a flowchart of steps of another data classification method for the data space information evolution process provided by an embodiment of the present application; Figure 4 is a data level display process under an embodiment of the present application; Figure 5 is a block diagram of a data classification device for the data space information evolution process provided by an embodiment of the present application; Figure 6 is a block diagram of an electronic device of an embodiment provided by an embodiment of the present application; Figure 7 is a block diagram of an electronic device of another embodiment provided by an embodiment of the present application. Detailed Description of the Embodiments

[0012] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0013] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0014] Considering the following process: Assume that the set of the first data (original data set) is , and the set of the second data (derived data set) is , and assume that each data entity has an initial level of .

[0015] The set of operation types for data operations is , where data operations may include: Downgrading operations (such as desensitization, anonymization, etc.), peer operations (such as format conversion, text modification in the passage), upgrading operations (such as multi-source data fusion, etc.).

[0016] Then a relationship matrix can be established by constructing a three-dimensional tensor , where the element represents the original data generating the derivative data through the operation of the correlation strength.

[0017] Then the determination of the data level can be carried out by designing a level transfer function: Based on the above process, as Figure 1 shown, it is a data grading method for the data space information evolution process provided by an embodiment of the present application.

[0018] The method may include the following steps: Step 101, in response to the process of generating the second data according to the data operation on the first data, obtain the first data level of the first data and the operation type of the data operation.

[0019] In some embodiments of the present application, in order to ensure that during the data generation process, the level information and operation type of the first data can be accurately grasped, providing the necessary basic data for the subsequent data grading process, during the execution stage of the data operation, the system will record in real time the level information of the first data related to the generation of the second data and the data operation type adopted. The data operation type refers to the specific operation executed during the data processing, such as data replication (datareplication) and data transformation (data transformation). The description of the data operation type includes the nature of the association between data and the potential impact on data grading. In this way, the system can automatically obtain and record the context information, avoiding manual intervention, improving the operation efficiency, and at the same time providing comprehensive background information for the subsequent grading decision.

[0020] As Figure 2As shown, in a specific example, when a user is processing data from multiple nodes and performing a data fusion operation through the system, it is necessary to integrate multi-source data into a new data set and automatically calculate the level of the new data set. Then the system can record the original data level information (for example, the level is 3) and the operation type (for example, data fusion) in real time. In this way, the system successfully records the level information and operation type of the original data, providing accurate and complete basic information for predicting the level of new data in the next step.

[0021] Step 102: Determine the evaluation value of the association strength between the first data and the second data according to the first data, the second data, and the operation type of the data operation.

[0022] In some embodiments of the present application, in order to quantify the logical correlation between data and enable the subsequent data grading process to more accurately reflect the internal relationship between data, a mathematical model for evaluating the association strength will be constructed and the evaluation value of the association strength will be calculated by using the eigenvalue of the first data and the second data, as well as the data operation type. The evaluation value of the association strength is a numerical value that measures the degree of logical or semantic connection between data, and is usually determined by analyzing the semantic similarity and the overlap degree of access permissions. The calculation of the evaluation value of the association strength can directly reflect the correlation between data, thus providing technical support for the accuracy of dynamic grading.

[0023] In a specific example, it is necessary to analyze the association strength between the new data generated through the data fusion operation and the original data. At this time, the user has performed a fusion operation on multiple data sources to generate a new derived data set. Then the system can vectorize the text eigenvalues of the original data and the derived data, calculate their semantic similarity, and analyze the overlap degree of their access permission lists, and then calculate the evaluation value of the association strength according to the set weight. In this way, the system successfully generates a quantitative evaluation value of the association strength to represent the logical association between the first data and the second data, providing a reliable basis for the prediction and grading of the subsequent data level.

[0024] Step 103: Determine the second data level of the generated second data according to the first data level, the evaluation value of the association strength, and the operation type of the data operation.

[0025] In some embodiments of the present application, in order to achieve the level transfer from the original data to the derived data to ensure the coherence and accuracy of the data classification and grading process, a mathematical model will be used to dynamically calculate the level of the second data based on the level information of the first data, the evaluation value of the association strength, and the data operation type. The evaluation value of the association strength quantifies the logical or semantic relationship between the original data and the derived data, and the data operation type directly affects the result of the level transfer. For example, a downgrade operation reduces the level, and an upgrade operation increases the level. In this way, the automatic grading of the derived data can be achieved, avoiding the overhead of repeated grading, and ensuring the traceability and consistency of the data level evolution process.

[0026] In a specific example, a user performs a data fusion operation on a certain data set to generate a new derived data set. At this time, the level of the original data set is 3, and the second data generated through the data fusion operation needs to redefine its level. The mathematical model can be used to dynamically predict that the level of the second data is 4 based on the level (3) of the original data set, the operation type (data fusion), and the calculated evaluation value of the association strength (such as 0.85). In this way, the system automatically generates the level information of the new data and records the level evolution process, ensuring the accuracy and efficiency of data classification and grading.

[0027] In summary, in the embodiments of the present application, by obtaining the initial information during the process of generating the response data, it is possible to ensure that the data grading process has comprehensive context information from the beginning, including the level of the first data and its corresponding data operation type, avoiding the need for manual intervention, and providing a basis for data grading decisions in subsequent steps. Furthermore, the calculation method of the evaluation value of the association strength can be introduced to quantify the logical relationship between the original data and the generated data, ensuring that the data level transfer process is more accurate, thus directly reducing the overhead of repeated grading. Then, the dynamic association strength and operation type are used to predict and generate the level information of the second data. By dynamically generating the level of the second data, the separate grading operation for each newly generated data is avoided, reducing redundant calculations at the system architecture level, solving the problem of exponential overhead growth caused by the frequent circulation of derived data, and at the same time realizing efficient support for dynamic data evolution scenarios. Therefore, based on the method of the embodiments of the present application, by combining the data operation type with the level change process, the traceability of the association between the original data and the derived data is realized, ensuring the accuracy and consistency of the data level evolution. Especially in the scenario of multiple nodes and multiple entities, this traceability mechanism provides reliable support for data management and sharing, overcoming the bottleneck problem of difficult traceability of level information in the prior art. It solves the problem of repeated grading overhead in the data dynamic evolution process, and also successfully constructs a traceable system for data evolution levels, providing technical guarantees for the efficient management and intelligence of the data space.

[0028] Figure 3 This is another data classification method for the data space information evolution process provided by the embodiments of the present application.

[0029] The method may include the following steps: Step 201, in response to the process of generating the second data according to the data operation on the first data, obtain the first data level of the first data and the operation type of the data operation.

[0030] The method shown in this step has been described in step 101 and will not be elaborated here.

[0031] Step 202, determine the evaluation value of the association strength between the first data and the second data according to the first data, the second data, and the operation type of the data operation.

[0032] The method shown in this step has been described in step 102 and will not be elaborated here.

[0033] Optionally, step 202 includes the following sub-steps: Sub-step 2021, determine the similarity evaluation value between the first data and the second data according to the content of the first data and the second data, and determine the permission evaluation value from the first data to the second data according to the operation type of the data operation.

[0034] In some embodiments of the present application, in order to quantify the content relevance and permission matching degree between the first data and the second data, and provide basic data for the calculation of the subsequent association strength evaluation value, the content information of the first data and the second data can be extracted to calculate the semantic similarity evaluation value between the two. At the same time, according to the operation type of the data operation, the permission information is analyzed to determine the permission evaluation value. The semantic similarity evaluation value measures the semantic association degree between data contents, and usually calculates the similarity between texts through the Word Vector technology. The permission evaluation value measures the overlap degree of access permissions between two data, and usually calculates the intersection ratio of the Access Control List (ACL) through the Jaccard coefficient to represent the overlap degree. In this way, high-quality input data can be provided for the calculation of the subsequent association strength evaluation value, thereby improving the accuracy and reliability of the association strength calculation.

[0035] In a specific example, it is necessary to evaluate the association strength between the first data and the second data. At this time, the first data is a text file containing user behavior data, and the second data is an analysis report generated based on the first data. Then, the text contents of the first data and the second data can be vectorized through the word embedding method, and the cosine similarity is used to calculate the semantic similarity, and the similarity evaluation value is obtained as 0.85. At the same time, by analyzing the access control lists (ACLs) of both, the permission overlap degree is calculated as 0.75, and the permission evaluation value is determined as 0.75. In this way, the system successfully generates the semantic similarity evaluation value and the permission evaluation value, providing accurate input for the weighted calculation of the next association strength evaluation value.

[0036] Sub-step 2022, determine the weighted sum of the similarity evaluation value and the permission evaluation value as the association strength evaluation value.

[0037] In some embodiments of the present application, in order to comprehensively reflect the semantic association degree of the data content and the overlap degree of the access permissions, so as to provide an accurate quantitative basis for further data-level derivation, the similarity evaluation value and the permission evaluation value will be linearly weighted and calculated, and their influence weights will be comprehensively considered to form a unified association strength evaluation value. The linear weighting method respectively regulates the proportion of the similarity evaluation value and the permission evaluation value in the final result through the weight coefficients, where the similarity evaluation value reflects the semantic relevance of the content, and the permission evaluation value reflects the access control overlap degree. The clear technical effect brought about after executing this step is to generate a quantitative association strength evaluation value, which serves as the core reference for subsequent level derivation and operation impact assessment, improving the accuracy and consistency of the data classification process.

[0038] In a specific example, it is necessary to evaluate the association strength between the first data and the second data to support the prediction of the derivative data level. At this time, the first data is the abstract text of the user's personal information data, and the second data is the statistical report generated based on this abstract. The execution process of the example includes: calculating the similarity evaluation value of the first data and the second data as 0.85, and the permission evaluation value as 0.70. Subsequently, the similarity evaluation value and the permission evaluation value are weighted and calculated according to the weights of 0.6 and 0.4 respectively, and finally the association strength evaluation value is obtained as 0.79. In this way, the quantitative association strength between the first data and the second data can be successfully determined, providing a clear reference for subsequent level calculation.

[0039] Based on the above scheme, the association strength evaluation value can be calculated in the following way:

[0040] Where: Indicates the semantic similarity between the original data and the derived data: The text correlation between the original data and the derived data can be calculated through word vectors, that is, convert the original data and the derived data into vectors respectively, and then calculate the similarity between the vectors; Indicates the permission overlap degree between the original data and the derived data: The overlap degree can be characterized by calculating the intersection ratio of the access control lists (ACLs) of the original data and the derived data through the Jaccard coefficient; Is the corresponding weight, and the proportion of the roles of semantics and permissions can be adjusted as needed.

[0041] Step 203, determine the second data level of the generated second data according to the first data level, the evaluation value of the association strength, and the operation type of the data operation.

[0042] The method shown in this step has been described in step 103 and will not be elaborated here.

[0043] Optionally, step 203 includes the following sub-steps: Sub-step 2031, determine the target influence coefficient corresponding to the operation type according to the operation type of the data operation.

[0044] In some embodiments of the present application, in order to quantify and standardize the influence of different data operations on the derived data level, thereby enhancing the accuracy and consistency of data level prediction, the target influence coefficient related to the operation type will be selected according to the characteristics of the operation type, such as downgrading operations, same-level operations, and upgrading operations. The target influence coefficient is a numerical parameter used to describe the specific influence of the operation type on the data level transfer process. For example, the influence coefficient range corresponding to the downgrading operation is 0.2 to 0.8, the influence coefficient corresponding to the same-level operation is 1.0, and the typical influence coefficient of the upgrading operation is 1.1 to 1.5. The description of the target influence coefficient includes its important role in the dynamic weighting process of the data level. The clear technical effect brought about after executing this step is to realize the quantitative representation of the data operation influence, providing an accurate and dynamic adjustment basis for the subsequent calculation of the data level.

[0045] In a specific example, it is necessary to analyze the influence of the multi-source data fusion operation on the derived data level. At this time, the user has generated new derived data through multi-source data fusion and hopes to calculate the level of this data. Then the target influence coefficient can be determined as 1.2 according to the data operation type (corresponding to the multi-source data fusion operation), and used as the weight parameter in the subsequent weighted calculation. In this way, the target influence coefficient is successfully set to 1.2, providing an important input for the weighted calculation of the association strength and the first data level, making the prediction of the derived data level more accurate.

[0046] In sub-step 2032, the product of the target impact coefficient and the associated strength evaluation value is used as the weight to calculate the weighted sum of the first data level as the second data level.

[0047] In some embodiments of the present application, in order to combine the target impact coefficient with the associated strength evaluation value and dynamically adjust the impact of the first data level on the second data level, thereby improving the accuracy of data level prediction, the target impact coefficient and the associated strength evaluation value are multiplied to obtain a comprehensive weight value, and then the weight value is used to perform a weighted calculation on the first data level to finally obtain the value of the second data level. The target impact coefficient reflects the degree of impact of data operations. For example, a downgrade operation reduces the data level or an upgrade operation increases the data level, while the associated strength evaluation value quantifies the logical and permission relationships between the original data and the derived data. This comprehensive weighting method ensures a reasonable balance of different factors in the derivation of the data level. In this way, the second data level can be accurately derived, and the computational overhead caused by repeated classification and grading can be significantly reduced.

[0048] In a specific example, it is necessary to calculate the level of the derived data generated by the multi-source data fusion operation. At this time, the levels of the first data are 3 and 4 respectively, the target impact coefficient of the multi-source data fusion operation is 1.2, and the associated strength evaluation value is 0.85. Then, the product of the target impact coefficient and the associated strength evaluation value can be calculated first to obtain a weight value of 1.02, and then the weight value is used to perform a weighted sum on the first data levels (3 and 4) to finally obtain the level of the second data as 3.84 (which will be discretized in subsequent steps). In this way, the system successfully predicts the continuous value of the second data level, providing basic data for the discretization mapping and subsequent analysis of the level.

[0049] The level transfer function designed through the above embodiments can be expressed by the following formula: ; Where: is the target impact coefficient. is the second data level, is the first data level, is the associated strength evaluation value.

[0050] After verification, in some embodiments, the impact coefficient can be set according to the following specific data: For downgrade operations (such as desensitization): , the typical value is 0.5; for same-level operations (such as format conversion): , indicating that it has a neutral impact; for upgrade operations (such as data fusion) , the typical value is 1.2.

[0051] In some embodiments of the present application, to determine the specific value of the influence coefficient for each data operation behavior, the following method can be used for calculation: Step 2033, collect first test data and second test data according to the target data operation to be verified.

[0052] Among them, the second test data is the test data generated from the first test data according to the target data operation; the first test data has a corresponding first test data level, and the second test data has a corresponding second test data level.

[0053] In some embodiments of the present application, in order to clarify the operation influence relationship between the first test data and the second test data through verification data collection, and lay a foundation for the subsequent determination of the target influence coefficient, data with typical representativeness will be selected as the first test data, and the second test data will be generated based on the specified target data operation, while ensuring that the classification and grading information of each data is recorded. The target data operation refers to the behavior of performing specific processing or treatment on data, such as data fusion or data anonymization. In this way, a set of paired test data is successfully generated, providing high-quality data support for analyzing the influence of the target data operation.

[0054] In a specific example, to verify the influence of the data fusion operation on the data level, a set of initial data files is selected from the database as the first test data, and the level of these data is 3. The execution process of the example includes performing a data fusion operation on the first test data to generate a new set of derivative data files as the second test data, and at the same time recording the level information of the second test data as 4. The paired test data (the first test data and the second test data) and their corresponding level information obtained in this way can provide basic data for the calculation of the target influence coefficient.

[0055] Step 2034, determine the second verification data level of the generated second test data according to the first test data level, the evaluation value of the association strength between the first test data and the second test data, and the operation type of the data operation.

[0056] In some embodiments of the present application, in order to provide a more accurate data-level prediction result based on the association and operation characteristics between test data, and to verify the applicability and accuracy of the target impact coefficient, the second verification data level of the second test data will be dynamically calculated by combining the first test data level and based on the evaluation value of the association strength with the second test data, using the rules or mathematical models defined by the operation type of data operation. The data operation types include downgrading operations (such as desensitization), same-level operations (such as format conversion), and upgrading operations (such as multi-source data fusion), where each operation type will have different impacts on the level prediction of test data. The second verification data level generated in this way can lay a foundation for further evaluating the error between test data levels, and at the same time improve the credibility and accuracy of the verification process.

[0057] In a specific example, it is necessary to verify whether the target impact coefficient is applicable to a certain type of data fusion operation through test data. At this time, the first test data has a level of 3, the second test data is generated through a data fusion operation, the evaluation value of its association strength with the first test data is 0.85, and the data operation type is an upgrading operation. Based on the weight rules set by the association strength evaluation value and the data operation type, combined with the level 3 of the first test data, the second verification data level of the second test data can be predicted to be 3.8 through calculation. The second verification data level of the second test data obtained in this way is used to compare with its actual level to evaluate the applicability of the target impact coefficient.

[0058] Step 2035: Determine the target impact coefficient corresponding to the target data operation according to the overall error determined by the second test data level and the second verification data level.

[0059] In some embodiments of the present application, in order to provide a quantitative basis for the optimization of the target impact coefficient by analyzing the differences between test data levels, thereby improving the applicability and accuracy of the data-level prediction process, the second test data level will be compared with the second verification data level, the error between the two will be calculated, and by summarizing the error values of multiple groups of test data, the target impact coefficient will be dynamically adjusted using the error minimization criterion. The overall error is calculated through the average or cumulative value of the errors of all test data groups, so as to quantify the influence range of data operation behavior on the level prediction result. In this way, by optimizing the target impact coefficient, the precision and reliability of the data classification and grading information evolution process are significantly improved.

[0060] In a specific example, it is necessary to verify whether the influence coefficient value of the data fusion operation is reasonable. At this time, the second test data level is 4, and the second verification data level predicted by the system is 3.8. Then the absolute error between the two can be calculated as 0.2. After summarizing the errors of all test data, the error minimization algorithm is used to optimize the target influence coefficient from the original set value of 1.2 to 1.25. The clear result obtained after executing according to the execution process of the example is that the target influence coefficient is successfully adjusted to 1.25, which more accurately reflects the actual influence of the data fusion operation on the level prediction result.

[0061] Optionally, in some embodiments of the present application, the influence coefficient in the case of the minimum overall error can be determined as the target influence coefficient.

[0062] In some embodiments of the present application, in order to minimize the error in data level prediction by optimizing the value of the target influence coefficient, thereby improving the accuracy of the data grading process, the prediction error of the test data can be statistically analyzed, the overall error value under each candidate influence coefficient can be calculated, and the corresponding influence coefficient can be selected as the target influence coefficient according to the error minimization criterion. The overall error is an index to measure the deviation degree between the predicted level and the actual level, and is usually quantified by the cumulative absolute error or the mean square error. The clear technical effect brought by executing this step is to automatically determine the target influence coefficient with the best fitness, provide accurate parameter support for the data grading decision, and ensure the stability of the prediction result.

[0063] In a specific example, it is necessary to determine the target influence coefficient of the data fusion operation. Multiple data sets have been tested through the data fusion operation, and there is a certain error between the actual level and the predicted level of each group of test data. Then the overall error values under different candidate influence coefficients (for example, 1.1, 1.2, 1.3) can be calculated. For example, the mean square errors (MSE) are 0.04, 0.02, and 0.05 respectively. By comparison, it is found that when the influence coefficient is 1.2, the overall error is the smallest (MSE = 0.02). Thus, the influence coefficient 1.2 is determined as the target influence coefficient for use in the subsequent level prediction process.

[0064] The above process can be transformed into the following optimization model: , where, is the influence coefficient corresponding to , is the second verification data level, is the second test data level, The function is used to obtain the that minimizes the number of grading errors.

[0065] Step 204, correct the determined second data level according to a preset rounding mapping function.

[0066] In some embodiments of the present application, in order to map the continuous level result to a discrete classification level, making the classification result more in line with the requirements in the actual application scenario, based on the sensitive level threshold function, the second data level (continuous value) obtained through dynamic weighted calculation is converted into the corresponding discrete level information. The sensitive level threshold function is a mathematical rule for segmentally mapping the level calculation value. By defining a reasonable threshold range (for example: level 1 corresponds to the range from 0 to T1, and level 2 corresponds to the range from T1 to T2), the accuracy and consistency of the conversion process are ensured. In this way, the uncertainty of the continuous value that may exist in the calculation result is eliminated, and the standardization and unity of the data level information are guaranteed.

[0067] In a specific example, the new data level after multi-source data fusion is dynamically calculated, and a continuous value result, such as 2.87, is obtained. It is necessary to standardize this continuous value result into a discrete classification level for subsequent security policy formulation. According to the sensitive level threshold function, the continuous value 2.87 can be mapped to the threshold range [T2, T3] and corrected to the discrete level 3. In this way, the level of the new data is successfully corrected to 3, ensuring the applicability and consistency of the level information in the actual system.

[0068] In an embodiment where the maximum data level is 4 levels, the above process can be expressed by the following formula: , Where: is the second data level directly calculated through the level transfer function, , , is the segmented threshold, which can be determined according to the actual application scenario.

[0069] Step 205, according to the determined second data level, show the process of generating the second data from the first data in the relationship graph.

[0070] Among them, the nodes in the relationship graph are used to represent the first data and / or the second data, the node values of each node are respectively used to represent the first data level of the first data corresponding to the node and / or the second data level of the second data, and the edges in the relationship graph are used to represent the data operation from the first data to the second data.

[0071] Such as Figure 4As shown, in some embodiments of the present application, in order to visually display the evolution process of data levels and the correlation relationships between data in a graphical manner, so as to facilitate the manager to visually track the data classification and grading results, nodes for representing the first data and / or the second data can be created in the relationship graph, edges for representing data operations can be added between the nodes, and node values can be attached to each node to represent the data levels. The node value (NodeValue) in the relationship graph represents the classification and grading information of the data corresponding to the node, and the edge (Edge) is used to represent the data operation type and influence relationship between the nodes. The clear technical effect brought about after performing this step is to graphically display the evolution process of data levels and related operations through the relationship graph, making the results of data classification and grading clearer and more convenient for subsequent analysis and management.

[0072] In a specific example, a new data set generated through a data fusion operation is processed, and it is desired to record the data level evolution information for easy tracking. The specific scenario of the example is that the levels of the original data sets are L1, L2, and L1 respectively, the levels of the new data set generated through the data fusion operation are L2 and L1 respectively, and then data with data levels of L3 and L1 are generated based on this data. In this way, the level transfer process and its logical correlation from the original data to the derived data can be accurately and visually observed through the relationship graph, providing effective support for data management and level traceability.

[0073] In summary, in the embodiments of the present application, by obtaining the initial information through the process of generating data in response, it is possible to ensure that the data classification process has comprehensive context information from the very beginning, including the level of the first data and its corresponding data operation type, avoiding the need for manual intervention and providing a basis for data classification decisions in subsequent steps. Furthermore, a calculation method for the evaluation value of the association strength can be introduced to quantify the logical relationship between the original data and the generated data, ensuring a more accurate data level transfer process, thereby directly reducing the overhead of repeated classification. Then, by using the dynamic association strength and operation type, the level information of the second data is predicted and generated. By dynamically generating the level of the second data, the need for separate classification operations for each newly generated data is avoided, reducing redundant calculations at the system architecture level and solving the problem of exponential overhead growth caused by the frequent circulation of derivative data. At the same time, efficient support for dynamic data evolution scenarios is achieved. Thus, based on the method of the embodiments of the present application, by combining the data operation type with the level change process, the traceability of the association between the original data and the derivative data is realized, ensuring the accuracy and consistency of data level evolution. Especially in scenarios with multiple nodes and multiple entities, this traceability mechanism provides reliable support for data management and sharing, overcoming the bottleneck problem of difficult traceability of level information in the prior art. It solves the problem of repeated classification overhead in the process of data dynamic evolution, and also successfully constructs a traceable system for data evolution levels, providing technical support for the efficient management and intelligence of the data space.

[0074] Reference Figure 5 , which shows a data classification device 30 for the information evolution process of the data space provided by the embodiments of the present application, including: An acquisition module 301, configured to obtain the first data level of the first data and the operation type of the data operation in response to the process of generating the second data according to the data operation on the first data; An evaluation module 302, configured to determine the evaluation value of the association strength between the first data and the second data according to the first data and the second data, and the operation type of the data operation; A classification module 303, configured to determine the second data level of the generated second data according to the first data level, the evaluation value of the association strength, and the operation type of the data operation.

[0075] Optionally, when the data types of the first data and the second data are the same, the evaluation module 302 includes: A parameter sub-module, configured to determine the similarity evaluation value between the first data and the second data according to the content of the first data and the second data, and determine the permission evaluation value from the first data to the second data according to the operation type of the data operation; An evaluation sub-module, configured to determine the weighted sum of the similarity evaluation value and the permission evaluation value as the evaluation value of the association strength.

[0076] Optionally, the grading module 303 includes: A coefficient sub-module, configured to determine a target impact coefficient corresponding to the operation type according to the operation type of the data operation; A grading sub-module, configured to use the product of the target impact coefficient and the correlation strength evaluation value as a weight, and calculate the weighted sum of the first data level as the second data level.

[0077] Optionally, the data grading device 30 for the data space information evolution process further includes: An integer-taking module, configured to correct the determined second data level according to a preset integer-taking mapping function.

[0078] Optionally, the data grading device 30 for the data space information evolution process further includes: A test set module, configured to collect first test data and second test data according to the target data operation to be verified; the second test data is test data generated by the first test data according to the target data operation; the first test data has a corresponding first test data level, and the second test data has a corresponding second test data level; A test module, configured to determine the second verification data level of the generated second test data according to the first test data level, the correlation strength evaluation value between the first test data and the second test data, and the operation type of the data operation; A coefficient determination module, configured to determine a target impact coefficient corresponding to the target data operation according to the overall error determined by the second test data level and the second verification data level.

[0079] Optionally, the coefficient determination module includes: A coefficient determination sub-module, configured to determine the impact coefficient in the case of the minimum overall error as the target impact coefficient.

[0080] Optionally, the data grading device 30 for the data space information evolution process further includes: A display module, configured to display the process of generating the second data according to the data operation on the first data in a relationship graph according to the determined second data level; the nodes in the relationship graph are used to represent the first data and / or the second data, the node value of each node is used to represent the first data level of the first data corresponding to the node and / or the second data level of the second data, and the edges in the relationship graph are used to represent the data operation from the first data to the second data.

[0081] In summary, in the embodiments of the present application, by obtaining initial information through the process of generating response data, it is possible to ensure that the data classification process has comprehensive context information from the very beginning, including the level of the first data and its corresponding data operation type, avoiding the necessity of manual intervention and providing a basis for data classification decisions in subsequent steps; furthermore, a calculation method for the evaluation value of the association strength can be introduced to quantify the logical relationship between the original data and the generated data, ensuring that the process of data level transmission is more accurate, thereby directly reducing the overhead of repeated classification; then, by using the dynamic association strength and operation type, the level information of the second data is predicted and generated, so as to avoid the separate classification operation for each newly generated data, reducing redundant calculations at the system architecture level and solving the problem of exponential overhead growth caused by the frequent circulation of derivative data, while achieving efficient support for dynamic data evolution scenarios. Thus, based on the method of the embodiments of the present application, by combining the data operation type with the process of level change, the traceability of the association between the original data and the derivative data is realized, ensuring the accuracy and consistency of data level evolution. Especially in the scenario of multiple nodes and multiple entities, this traceability mechanism provides reliable support for data management and sharing, overcoming the bottleneck problem of difficult traceability of level information in the prior art. It solves the problem of repeated classification overhead in the process of data dynamic evolution, and also successfully constructs a traceable system for data evolution levels, providing technical guarantee for the efficient management and intelligence of the data space.

[0082] Referring to Figure 6 , the electronic device 500 may include one or more of the following components: a processing component 502, a memory 504, a power supply component 506, a multimedia component 508, an audio component 510, an input / output (I / O) interface 512, a sensor component 514, and a communication component 516.

[0083] The processing component 502 generally controls the overall operation of the electronic device 500, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The processing component 502 may include one or more processors 520 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 502 may include one or more modules to facilitate the interaction between the processing component 502 and other components. For example, the processing component 502 may include a multimedia module to facilitate the interaction between the multimedia component 508 and the processing component 502.

[0084] The memory 504 is used to store various types of data to support the operation of the electronic device 500. Examples of such data include instructions for any application or method operating on the electronic device 500, contact data, phone book data, messages, pictures, multimedia, etc. The memory 504 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0085] The power supply component 506 provides power for various components of the electronic device 500. The power supply component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 500.

[0086] The multimedia component 508 includes an interface that provides an output interface between the electronic device 500 and the user. In some embodiments, the interface may include a liquid crystal display (LCD) and a touch panel (TP). If the interface includes a touch panel, the interface can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 508 includes a front camera and / or a rear camera. When the electronic device 500 is in an operating mode, such as a shooting mode or a multimedia mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0087] The audio component 510 is used to output and / or input audio signals. For example, the audio component 510 includes a microphone (MIC) that is used to receive external audio signals when the electronic device 500 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 504 or transmitted via the communication component 516. In some embodiments, the audio component 510 further includes a speaker for outputting audio signals.

[0088] The input / output (I / O) interface 512 provides an interface between the processing component 502 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a start button, and a lock button.

[0089] The sensor assembly 514 includes one or more sensors for providing a status assessment of various aspects of the electronic device 500. For example, the sensor assembly 514 can detect the on / off state of the electronic device 500, the relative positioning of components, such as the display and keypad of the electronic device 500. The sensor assembly 514 can also detect a change in the position of the electronic device 500 or a component of the electronic device 500, the presence or absence of user contact with the electronic device 500, the orientation or acceleration / deceleration of the electronic device 500, and the temperature change of the electronic device 500. The sensor assembly 514 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 514 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 514 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0090] The communication component 516 is used to facilitate wired or wireless communication between the electronic device 500 and other devices. The electronic device 500 can access a wireless network based on communication standards, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 516 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 516 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0091] In an exemplary embodiment, the electronic device 500 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to implement the methods provided in the embodiments of the present application.

[0092] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 504 including instructions, and the above instructions can be executed by the processor 520 of the electronic device 500 to complete the above method. For example, the non-transitory storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0093] Figure 7is a block diagram of an electronic device 600 according to another embodiment of the present invention. For example, the electronic device 600 may be provided as a server.

[0094] Referring to Figure 7 , the electronic device 600 includes a processing component 622, which further includes one or more processors, and memory resources represented by a memory 632 for storing instructions executable by the processing component 622, such as application programs. The application programs stored in the memory 632 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 622 is configured to execute instructions to perform the methods provided in the embodiments of the present application.

[0095] The electronic device 600 may also include a power component 626 configured to perform power management of the electronic device 600, a wired or wireless network interface 650 configured to connect the electronic device 600 to a network, and an input / output (I / O) interface 658. The electronic device 600 may operate based on an operating system stored in the memory 632, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSD TM or the like.

[0096] It should be noted that, for the method embodiments of the present application, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described action sequence, because according to the embodiments of the present application, certain steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present application.

[0097] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.

[0098] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A data classification method for the data space information evolution process, characterized in that, Including: In response to the process of generating second data based on data operations on first data, obtaining the first data level of the first data and the operation type of the data operation; Determining an association strength evaluation value between the first data and the second data according to the first data, the second data, and the operation type of the data operation; Determining the second data level of the generated second data according to the first data level, the association strength evaluation value, and the operation type of the data operation.

2. The data classification method for the data space information evolution process according to claim 1, wherein When the data types of the first data and the second data are the same, the determining an association strength evaluation value between the first data and the second data according to the first data, the second data, and the operation type of the data operation includes: Determining a similarity evaluation value between the first data and the second data according to the content of the first data and the second data, and determining a permission evaluation value from the first data to the second data according to the operation type of the data operation; Determining the weighted sum of the similarity evaluation value and the permission evaluation value as the association strength evaluation value.

3. The data classification method for the data space information evolution process according to claim 1, wherein The determining the second data level of the generated second data according to the first data level, the association strength evaluation value, and the operation type of the data operation includes: Determining a target influence coefficient corresponding to the operation type according to the operation type of the data operation; Using the product of the target influence coefficient and the association strength evaluation value as a weight, calculating the weighted sum of the first data levels as the second data level.

4. The data classification method for the data space information evolution process according to claim 3, characterized in that The data classification method for the data space information evolution process further includes: Collecting first test data and second test data according to a target data operation to be verified; the second test data is test data generated from the first test data according to the target data operation; the first test data has a corresponding first test data level, and the second test data has a corresponding second test data level; Determining the second verification data level of the generated second test data according to the first test data level, the association strength evaluation value between the first test data and the second test data, and the operation type of the data operation; Determining a target influence coefficient corresponding to the target data operation according to the overall error determined by the second test data level and the second verification data level.

5. The data classification method for the data space information evolution process according to claim 1, characterized in that The data classification method for the data space information evolution process further includes: Correcting the determined second data level according to a preset rounding mapping function.

6. The data grading method for the data space information evolution process according to claim 1, characterized in that, The data classification method for the data space information evolution process further includes: Displaying the process of generating second data based on data operations on first data in a relationship graph according to the determined second data level; the nodes in the relationship graph are used to represent the first data and / or the second data, the node value of each node is used to represent the first data level of the first data corresponding to the node and / or the second data level of the second data, and the edges in the relationship graph are used to represent the data operation from the first data to the second data.

7. A data classification device for the information evolution process of the data space, characterized in that, Comprising: A collection module, configured to obtain a first data level of the first data and an operation type of the data operation in response to a process of generating second data according to a data operation on the first data; An evaluation module, configured to determine an evaluation value of the association strength between the first data and the second data according to the first data and the second data, and the operation type of the data operation; A grading module, configured to determine a second data level of the generated second data according to the first data level, the evaluation value of the association strength, and the operation type of the data operation.

8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the data grading method for the data space information evolution process according to any one of claims 1 to 6.

9. An electronic device, characterized in that, Comprising a processor, a memory, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, it implements the steps of the data grading method for the data space information evolution process according to any one of claims 1 to 6.

10. A computer program product, characterized in that, A computer program is stored on the computer program product, and when the computer program is executed by a processor, it implements the steps of the data grading method for the data space information evolution process according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Intelligent construction method for enterprise data standard hierarchical relationship

    CN117971808A

  • Sensitive data identification protection method and system based on intelligent matching

    CN119577815A

  • Method, electronic apparatus, and storage medium for generating entity information graph

    US20240386055A1