Data classification and grading method and system based on industry knowledge structure and incidence relation

By capturing enterprise business flow data, generating heterogeneous relationship diagrams and multi-layer dynamic flow maps, and combining with regulatory knowledge bases, automatic data classification and grading can be achieved, the problems of manual dependence and static analysis in existing technologies can be solved, and efficient and accurate data classification and grading and continuous management can be achieved.

CN120763261AActive Publication Date: 2025-10-10SHENZHEN ANTECH TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511285013.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-10-10
Estimated Expiration
2045-09-10

AI Technical Summary

Technical Problem

Existing data classification and grading technologies rely on manual intervention and static analysis, which are costly, have high error rates, are difficult to carry out sustainably, and the classification and grading results are inaccurate, which cannot meet the long-term needs of enterprise data management.

Method used

By capturing raw data fragments from corporate business traffic, performing preprocessing and type inference, and generating weighted heterogeneous relationship graphs, combined with multi-layer dynamic flow maps and industry regulatory knowledge bases, data classification and grading are automatically performed to ensure that the results comply with laws and regulations and reflect actual business scenarios.

Benefits of technology

It achieves accuracy and flexibility in data classification and grading, reduces manual intervention, lowers costs, can continuously adapt to changes in enterprise data, and improves data utilization efficiency and management accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763261A_ABST
    Figure CN120763261A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and particularly discloses a data classification and grading method and system based on an industry knowledge structure and an association relationship. The method comprises the following steps: firstly, capturing original data fragments of enterprise business traffic and preprocessing the original data fragments to obtain structured data element information; performing type inference and name semantic analysis on the data items to obtain associated information, and extracting entity relationships to generate a heterogeneous relationship graph; then constructing a multi-layer dynamic circulation map; first classification and grading results are obtained according to the industry regulation knowledge base, and compliance is guaranteed; and finally, a target classification and grading result is obtained by combining the second classification and grading result and the second classification and grading result, and laws and regulations and actual business scenes are comprehensively considered, so that the result is more accurate and reasonable, which is helpful for an enterprise to give full play to data business value while ensuring that data is safe and compliant, and the enterprise experience is improved. The data management strategy is flexibly adjusted, the data utilization efficiency is improved, and the method is efficient and accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a data classification and grading method and system based on industry knowledge structure and association relationships. Background Art

[0002] In the field of data security, classification and grading are key and primary steps in data security governance. Decades of information technology development have amassed vast amounts of data, far exceeding the capacity of human processing. Further processing and unlocking the business value of this data requires clarifying the quantity and distribution of existing data, understanding its data types and sensitivity, and developing governance strategies tailored to the different types and sensitivities of data. Currently, classification and grading techniques employ two approaches: metadata type classification and application scenario classification. In practice, top-down classification is primarily employed, combined with scanning and recognition capabilities to discover data, followed by manual identification and classification.

[0003] However, the current classification and grading technology has obvious shortcomings. Its information technology assistance is lagging behind, and it mainly relies on manpower input and static analysis. The mainstream method is to combine manual labeling, machine scanning and rule discovery to identify data structure correlation relationships, and then use big data models, AI and other technologies to associate and identify more data. However, this method requires a lot of manual intervention and organizational support and coordination, is costly and has a high error rate, and the actual effect is far below expectations. In addition, the classification and grading work takes a long time to complete, requires a lot of manpower and has poor results. In addition, the existing practices are not sustainable and cannot meet the requirements of continuous classification and grading. Summary of the Invention

[0004] In view of this, the embodiments of the present disclosure provide a data classification and grading method and system based on industry knowledge structure and association relationships, which can solve the problems existing in the existing technology such as poor accuracy of data classification and grading results, large manpower investment, high cost, long time consumption, and difficulty in sustainability.

[0005] In a first aspect, embodiments of the present disclosure provide a data classification and grading method based on industry knowledge structure and association relationships, including: Capturing raw data segments from the enterprise's business traffic and preprocessing the raw data segments to obtain structured data metadata; Performing type inference and name semantic parsing on each data item in the structured data metadata to obtain associated information of each data item; Extract entity relationships from all data items based on the association information, identify dynamic association relationships between entities, entities, data, and data, and generate a weighted heterogeneous relationship graph; Based on the heterogeneous relationship graph, a multi-layer dynamic flow transfer map covering entities, data, users, systems, accounts and roles is constructed. A first classification grading result of the data to be analyzed is obtained according to an industry regulation knowledge base. A second classification grading result of the data to be analyzed is obtained according to the multi-layer dynamic flow transfer map. A target classification grading result of the data to be analyzed is obtained based on the first classification grading result and the second classification grading result.

[0006] In a second aspect, the embodiments of the present disclosure further provide a data classification grading system based on industry knowledge structure and association relationship, comprising: A preprocessing module is configured to capture original data segments in business traffic of an enterprise, and preprocess the original data segments to obtain structured data element information. An association information acquisition module is configured to perform type inference and name semantic analysis on each data item in the structured data element information to obtain association information of each data item. A heterogeneous relationship graph acquisition module is configured to perform entity relationship extraction on all data items according to the association information, identify dynamic association relationships between entity-entity, entity-data and data-data, and generate a weighted heterogeneous relationship graph. A multi-layer dynamic flow transfer map construction module is configured to construct a multi-layer dynamic flow transfer map covering entities, data, users, systems, accounts and roles based on the heterogeneous relationship graph. A first classification grading result acquisition module is configured to obtain a first classification grading result of the data to be analyzed according to an industry regulation knowledge base. A second classification grading result acquisition module is configured to obtain a second classification grading result of the data to be analyzed according to the multi-layer dynamic flow transfer map. A target classification grading result acquisition module is configured to obtain a target classification grading result of the data to be analyzed based on the first classification grading result and the second classification grading result.

[0007] In a third aspect, the embodiments of the present disclosure further provide a computer device, which adopts the following technical solution: The computer device comprises: at least one processor; and a memory in communication connection with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the data classification grading method based on industry knowledge structure and association relationship as described above.

[0008] In a fourth aspect, the embodiments of the present disclosure further provide a computer readable storage medium storing computer instructions for causing a computer to execute the data classification and grading method based on industry knowledge structure and association relationship.

[0009] In a fifth aspect, the embodiments of the present disclosure further provide a computer program product comprising computer programs / instructions for implementing the steps of the method according to any one of the preceding aspects when executed by a processor.

[0010] The data classification and grading method based on industry knowledge structure and association relationship disclosed in the present application first extracts original data segments in the business traffic of an enterprise, and pre-processes the original data segments to obtain structured data element information. Secondly, type inference and name semantic analysis are performed on each data item in the structured data element information to obtain association information of each data item. Then, entity relationship extraction is performed on all data items according to the association information to identify dynamic association relationships between entity-entity, entity-data and data-data, and to generate a weighted heterogeneous relationship graph. Next, based on the heterogeneous relationship graph, a multi-layer dynamic flow transfer map covering entities, data, users, systems, accounts and roles is constructed. Finally, a first classification and grading result of the data to be analyzed is obtained according to an industry regulation knowledge base, a second classification and grading result of the data to be analyzed is obtained according to the multi-layer dynamic flow transfer map, and a target classification and grading result of the data to be analyzed is obtained based on the first classification and grading result and the second classification and grading result. The first classification and grading result is obtained according to the industry regulation knowledge base, which ensures that the data classification and grading complies with the requirements of relevant laws and regulations. The second classification and grading result is obtained based on the multi-layer dynamic flow transfer map, which reflects the use and association of the data in actual business. Combining the two can comprehensively consider the regulatory requirements and actual business scenarios, making the data classification and grading result more accurate and reasonable. This comprehensive consideration helps the enterprise to fully utilize the business value of data under the premise of ensuring data security compliance, meets the constraints of regulations, and can flexibly adjust data management strategies according to actual business needs, improving the data utilization efficiency of the enterprise.

[0011] The above description is only a summary of the technical solutions of the present disclosure. In order to more clearly understand the technical means of the present disclosure, the above description can be implemented according to the content of the specification, and in order to make the above and other purposes, features and advantages of the present disclosure more obvious and easy to understand, the following preferred embodiments are described in detail below, and the accompanying drawings are described as follows. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0013] Figure 1 A flowchart of a data classification and grading method based on industry knowledge structure and association relationships provided in an embodiment of the present disclosure.

[0014] Figure 2 A flowchart of a method for obtaining structured data metadata provided by an embodiment of the present disclosure.

[0015] Figure 3 A flowchart of a method for obtaining associated information of each data item provided in an embodiment of the present disclosure.

[0016] Figure 4 A flowchart of a method for constructing a multi-layer dynamic flow map provided in an embodiment of the present disclosure.

[0017] Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0018] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0019] It should be clear that the following embodiments of the present disclosure are described through specific concrete examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that the following embodiments and features in the embodiments can be combined with each other in the absence of conflict. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.

[0020] It should be apparent that the aspects described herein can be embodied in a wide variety of forms and that any specific structure and / or function described herein is merely illustrative. Based on the teachings herein one skilled in the art should appreciate that an aspect described herein can be implemented independently of any other aspects and that an aspect described herein can be implemented both as any claim dependent on another and as an independent claim capable of being implemented in isolation from said dependent claim.

[0021] It should also be noted that the figures provided in the following embodiments are only to schematically illustrate the basic concept of the present disclosure, and only the components related to the present disclosure are shown in the figures, not drawn according to the number, shape and size of the components when actually implemented, and the shape, number and proportion of each component when actually implemented can be a random change, and the layout pattern of the components can also be more complex.

[0022] In addition, in the following description, specific details are provided in order to facilitate a thorough understanding of the examples. However, one skilled in the art will understand that the aspects described can be practiced without these specific details.

[0023] Referring to Figure 1 The present application discloses a data classification and grading method based on industry knowledge structure and association relationship, comprising: S100, grabbing original data segments in business traffic of an enterprise, and preprocessing the original data segments to obtain structured data element information.

[0024] This step grabs original data segments in business traffic of an enterprise and performs preprocessing, converting disordered original data into structured data element information, effectively avoiding the tedious process of manually processing massive data, greatly improving the usability of data, laying a good foundation for subsequent analysis and processing, and making the data processing process more efficient and smooth.

[0025] S200, type inference and name semantic analysis are performed on each data item in the structured data element information to obtain association information of each data item.

[0026] This step performs type inference and name semantic analysis on the data items in the structured data element information to automatically obtain the association information of each data item. This intelligent processing method reduces the time and workload of manual judgment, can quickly and accurately determine the potential relationship between data items, and improves the efficiency of data association analysis.

[0027] S300, extracting entity relationships from all data items based on association information, identifying dynamic association relationships between entity-entity, entity-data, and data-data, and generating a weighted heterogeneous relationship graph.

[0028] This step extracts entity relationships based on associated information, identifies dynamic associations between entities, entities, data, and data, and generates weighted heterogeneous relationship graphs. Compared with traditional static analysis methods, it can more comprehensively and accurately reflect the complex relationships between data, effectively avoid information loss and misjudgment caused by static analysis, and provide a more reliable basis for subsequent data classification and grading.

[0029] S400, based on heterogeneous relationship graphs, builds a multi-layer dynamic flow map covering entities, data, users, systems, accounts and roles.

[0030] This step builds a multi-layer dynamic flow map covering entities, data, users, systems, accounts, and roles based on the heterogeneous relationship diagram. The map comprehensively displays the flow of data at different levels and links, enabling enterprises to clearly understand the flow path and usage of data, helping to classify and grade data more accurately, and improving the accuracy and refinement of data management.

[0031] S500: Obtain a first classification and grading result of the data to be analyzed according to the industry regulations knowledge base.

[0032] This step obtains the first classification and grading results of the data to be analyzed based on the industry regulatory knowledge base, and integrates industry regulatory requirements into the data classification and grading process. This ensures that the company's data management activities comply with relevant laws and regulations, avoids legal risks caused by data classification and grading that do not comply with regulations, and provides compliance protection for the company's data security governance.

[0033] S600: Obtain a second classification and grading result of the data to be analyzed according to the multi-layer dynamic flow map.

[0034] The multi-layer dynamic flow map covers multiple levels such as entities, data, users, systems, accounts and roles, and shows in detail the flow path of data between different links and systems. This enables the classification and grading of data to be closely integrated with actual business usage scenarios. That is, through the multi-layer dynamic flow map, the changes in such data usage scenarios can be accurately captured, thereby classifying and grading data more accurately.

[0035] S700: Obtain target classification and grading results of the data to be analyzed based on the first classification and grading results and the second classification and grading results.

[0036] The first classification and grading result is derived from the industry regulations knowledge base, which ensures that the data classification and grading complies with the requirements of relevant laws and regulations. The second classification and grading result is obtained based on a multi-layer dynamic flow map, which reflects the usage and correlation of data in actual business. Combining the two can comprehensively consider regulatory requirements and actual business scenarios, making the data classification and grading results more accurate and reasonable. This comprehensive consideration approach helps enterprises to fully realize the business value of data while ensuring data security and compliance, which not only meets the constraints of regulations, but also can flexibly adjust data management strategies according to actual business needs, thereby improving the data utilization efficiency of enterprises.

[0037] Traditional data classification and grading work is often a one-time large-investment project that requires enterprises to concentrate a large amount of manpower, material resources and financial resources within a specific time period to complete. This model is not only costly, but as the business of the enterprise develops and the data continues to change, the one-time classification and grading results will soon become outdated and difficult to adapt to new business needs. The method disclosed in this application pioneered the transformation of classification and grading work into a normalized continuous investment work. Through automated data capture, preprocessing and analysis processes, it can continuously process and classify and grade the newly generated data of the enterprise. This means that enterprises no longer need to make large-scale one-time investments, but can gradually and continuously advance classification and grading work according to actual conditions, so that classification and grading work is closely integrated with the daily operations of the enterprise, making it more flexible and adaptable.

[0038] The method disclosed in this application constructs a multi-layer dynamic flow map covering entities, data, users, systems, accounts and roles, which can continuously monitor the data flow process within the organization. This allows enterprises to understand the flow of data between different links and systems in real time, including information such as the source, destination, and frequency of use of the data.

[0039] Furthermore, the application can also help companies continuously improve their data classification and grading based on the results of continuous monitoring. As the business develops and data flow changes, the sensitivity and importance of certain data may change. Through dynamic flow maps, companies can promptly detect these changes and make corresponding adjustments to the classification and grading of data, thereby achieving a near-real-time correspondence between classification and grading results and actual conditions. This real-time adjustment capability ensures that the classification and grading results always accurately reflect the actual status of the data, improving the effectiveness of data management.

[0040] Traditional methods mainly rely on static analysis, which makes it difficult to accurately reflect the complex relationships and dynamic changes between data, resulting in inaccurate classification and grading results. This solution identifies the dynamic association relationships between entities, entities, data, and data, generates weighted heterogeneous relationship graphs, and constructs multi-layer dynamic flow maps. This can comprehensively and accurately reflect the actual situation of the data. At the same time, the target classification and grading results obtained by combining the industry regulatory knowledge base and the dynamic flow map comprehensively consider regulatory requirements and actual business scenarios, further improving the accuracy of classification and grading.

[0041] Existing classification and grading methods are not sustainable and cannot adapt to the continuous growth and changes of enterprise data; the solution of this application can continuously process and classify and grade new data, and can adjust the classification and grading results in real time according to changes in data flow. This continuous improvement mechanism ensures that the classification and grading work can be carried out continuously and effectively, meeting the long-term needs of enterprises for data management.

[0042] Reference Figure 2 The method of S100, "capturing raw data segments from the enterprise's business traffic and preprocessing the raw data segments to obtain structured data metadata," is a method for obtaining structured data metadata, specifically including: S110 captures raw data fragments in the service traffic through bypass monitoring technology.

[0043] Specifically, port mirroring / ERSPAN can be enabled at the core switching node or cloud-native traffic probe to replicate all TCP / UDP traffic.

[0044] In this step, bypass monitoring technology will not interfere with the normal business traffic of the enterprise network, because it monitors by copying the traffic instead of directly intervening in the transmission path of the business data, so it will not affect the performance and stability of the business system; at the same time, it can obtain all business traffic in the enterprise network, including communication between internal systems and interaction with external systems, which enables the enterprise to fully understand the flow of its business data and provides a rich data source for subsequent data processing and analysis.

[0045] S120 , based on a sliding window conversation reorganization algorithm, reorganize all original data fragments into a complete conversation corresponding to the business interaction process.

[0046] Reorganizing scattered data fragments into complete conversations allows businesses to clearly understand the entire business interaction process. This helps businesses analyze whether business processes are running smoothly and whether user behavior is meeting expectations, providing strong support for business optimization. Complete conversation data also facilitates subsequent data analysis and mining. For example, when conducting user behavior analysis, complete conversation data can provide more comprehensive user operation information, helping businesses better understand user needs and preferences, thereby formulating more precise marketing strategies.

[0047] S130 , appending a global unified session ID, a client certificate fingerprint, a server certificate fingerprint, a user identifier, and a nanosecond timestamp to each data fragment in the complete session to form structured data metadata.

[0048] A globally unified session ID allows enterprises to easily manage and query different sessions. Through the session ID, enterprises can quickly locate and retrieve all data fragments for a specific session, improving data retrieval efficiency. At the same time, client and server certificate fingerprints can be used to verify the identities of both communicating parties, ensuring the security of data transmission. User identifiers and timestamps ensure data traceability. Enterprises can use this information to track the source of data, user operation time, and other aspects, facilitating investigations and audits in the event of security issues or business disputes.

[0049] The method for obtaining structured data metadata disclosed in this embodiment ensures the integrity and accuracy of business data through crawling, reorganization and structured processing. Enterprises can obtain complete business interaction information, avoid data loss and errors, and provide a reliable foundation for subsequent data processing and analysis; complete and structured business data enables enterprises to deeply analyze business processes and user behaviors and discover potential problems and opportunities; additional information such as certificate fingerprints, user identifiers and timestamps helps enterprises conduct security monitoring and compliance checks; enterprises can monitor abnormal session activities in real time, promptly discover security threats, and meet the requirements of relevant laws and regulations and industry standards for data traceability and security.

[0050] Reference Figure 3 The method of S200, "performing type inference and name semantic analysis on each data item in the structured data metadata to obtain associated information of each data item," i.e., a method for obtaining associated information of each data item, specifically includes: S210, builds a multi-layer knowledge graph covering industry terms, synonyms, and abbreviations.

[0051] Among them, the multi-layer knowledge graph includes the core layer knowledge graph, the extended layer knowledge graph and the customized layer knowledge graph.

[0052] Furthermore, the core layer knowledge graph contains the most basic and core industry terms in the field. These terms are the cornerstone of the entire knowledge graph and are highly professional and authoritative.

[0053] The extended layer knowledge graph covers synonyms and common abbreviations related to the core layer terms. By expanding the semantic scope of the core layer, the knowledge graph can handle more diverse forms of expression.

[0054] The customized knowledge graph is a collection of terms customized according to specific business needs or specific scenarios, which can be flexibly added or modified to adapt to different application environments.

[0055] Multi-layer knowledge graphs can comprehensively and deeply represent industry knowledge. The core layer provides a basic framework, the extended layer enriches knowledge details, and the customized layer meets specific business needs. Through knowledge graphs, we can better understand the semantics of data items and provide strong support for subsequent entity linking and disambiguation.

[0056] S220, using the BERT-BiLSTM-CRF cascade network to perform sequence labeling on the names of data items in the structured data metadata to obtain entity boundaries.

[0057] BERT can learn powerful semantic representations, BiLSTM can process sequence context information, and CRF can perform accurate sequence labeling; through this cascade network, entity boundaries in data item names can be accurately identified, providing accurate entity information for subsequent entity linking.

[0058] S230, linking entities in the entity boundary with nodes in the multi-layer knowledge graph, and using the PageRank algorithm to disambiguate the link results to obtain the semantic label of each data item.

[0059] Entity linking associates data items with knowledge graphs, giving data items clear semantics; the PageRank algorithm can effectively resolve ambiguity in entity linking, improve the accuracy of semantic labels, and thus better understand the meaning of data items.

[0060] Assume that the entity "iPhone" in the data item "product_name" is obtained through S220. Matching "iPhone" with nodes in the multi-layer knowledge graph reveals that the knowledge graph contains a node for "iPhone." However, there may be multiple pieces of information related to "iPhone," such as different iPhone models. In this case, the PageRank algorithm is used to disambiguate these link results. The PageRank algorithm assigns a weight to each link result based on the importance and link relationships of the nodes in the knowledge graph, and selects the link result with the highest weight as the final semantic label, such as "iPhone 15."

[0061] S240: Extract the data content of each data item in the structured data metadata.

[0062] This step provides actual data samples for subsequent data type analysis. Only after obtaining the data content can its type be judged and analyzed.

[0063] For example, in e-commerce order data, for the data item "price", its corresponding data content, such as "9999", is extracted from the data record; for "product_name", the specific product name is extracted, such as "Apple iPhone 15".

[0064] S250: Use data samples of known types to train the target model, analyze the data content of each data item through the trained model, and obtain a content type label for each data item.

[0065] The associated information includes the semantic label and content type label of the data item.

[0066] Through model training and analysis, the data content type of data items can be automatically and accurately determined, providing important type information for data processing and analysis, which is helpful for subsequent data cleaning, conversion, mining and other operations.

[0067] Specifically, a large number of data samples of known types are collected, such as numeric types (100, 20.5), string types ("apple", "mobile phone"), date types ("2025-08-14"), etc., and these samples are used to train the target model (such as a deep learning model); then, the extracted data content "9999" is input into the trained model, and the model determines that it is a numeric type, thereby obtaining the content type label "numeric type" for the "price" data item.

[0068] By performing type inference and name semantic analysis on data items, combined with multi-layer knowledge graphs and entity link disambiguation, we can more accurately understand the meaning and type of data, providing a more reliable basis for data analysis and decision-making; clear data types and semantic labels help automate data processing processes, reduce manual intervention, and improve the speed and accuracy of data processing; the customization layer and type inference mechanism of the multi-layer knowledge graph can be adjusted and optimized according to different business needs to meet diverse business scenarios; unified data types and semantic labels help data integration and sharing between different data sources, breaking down data silos and increasing the utilization value of data.

[0069] For the method of S300 "entity relationship extraction is performed on all data items according to the association information, dynamic association relationships between entity-entity, entity-data, and data-data are identified, and a weighted heterogeneous relationship graph is generated", the following steps are specifically included: S310: data preprocessing and association information integration. Specifically, 1) the association information of all data items obtained from the S200 step is cleaned to remove noise data such as incorrect semantic labels, incomplete type information, etc. For example, if there is a spelling error in the semantic label of the association information of a certain data item, it is corrected; if the type information is missing, it is supplemented or deleted according to the context or common rules. Process duplicate data items, and for data items with the same association information, only one copy is retained to reduce the complexity of subsequent processing.

[0070] 2) unify the format and expression of the association information. For example, all semantic labels are mapped according to a predefined vocabulary to ensure that semantic labels with the same meaning have a unified representation; standardize the data types, such as unifying different representations of date types into a standard format.

[0071] S320: entity recognition and classification. Specifically, 1) a series of rules are developed to identify entities in data items. For example, according to the semantic labels in the association information, if the semantic labels of a certain data item contain "customer", "product", "order", etc. Key words, it is identified as the corresponding entity.

[0072] 2) Use the type information of the data item to assist entity recognition, such as data items with "object" type are more likely to be entities.

[0073] 3) classify the identified entities according to their business attributes. For example, entities are classified into customer entities, product entities, order entities, etc. At the same time, a set of features is defined for each entity category for subsequent relationship extraction.

[0074] S330: entity relationship extraction. Specifically, 1) machine learning-based relationship extraction can be used, which includes selecting appropriate machine learning algorithms such as support vector machines (SVM), decision trees, etc. to extract relationships between entities-entities, entities-data, and data-data. First, extract features from the association information, such as the similarity of semantic labels, the matching degree of data types, etc. as input to the machine learning model.

[0075] The model is trained using labeled training data containing known entity relationship pairs and their corresponding relationship types. After training, input the association information of all data items into the model, and output the relationship types between entities.

[0076] 2) Rule-based relationship extraction can be used. This involves developing a set of rules to supplement relationships not identified by the machine learning model. For example, if two data items are semantically labeled "order" and "product," and they frequently appear together in a business process, a "contains" relationship can be inferred between them.

[0077] Consider the time sequence and business logic between data items. For example, in a business process, data item A always appears before data item B, and there is a data transfer relationship between them. It can be inferred that there is a "predecessor" relationship between them.

[0078] S340: Relationship Weight Calculation. Specifically, this includes: 1) Optional weight calculation based on association strength. Specifically, for identified entity-entity, entity-data, and data-data relationships, the relationship weight is calculated based on their association strength. Association strength can be measured in various ways, such as semantic similarity and data interaction frequency. For example, if two entities frequently interact in a business process and their semantic labels are highly similar, the relationship weight between them will be higher.

[0079] 2) You can choose to adjust weights based on business importance. Specifically, consider the importance of the relationship to the business and adjust the relationship weights. For example, increase the weight of relationships involved in core business processes and decrease the weight of auxiliary relationships.

[0080] S350: Heterogeneous relationship graph construction. This involves treating the identified entities and data items as graph nodes and the extracted relationships as graph edges. Each node and edge has corresponding attributes. Node attributes include association information and entity category, while edge attributes include relationship type and weight. A graph database (such as Neo4j) is used to construct and store the heterogeneous relationship graph. Node and edge information is inserted into the graph database to form a complete heterogeneous relationship graph. Simultaneously, the graph database is indexed to improve query and analysis efficiency.

[0081] S360: Dynamic Relationship Update. This specifically includes: 1) Continuously monitoring raw data fragments in business traffic and promptly updating structured data metadata and relationship information when new data is generated or existing data changes. 2) Based on the updated relationship information, re-identify entities, extract relationships, and calculate weights, updating the heterogeneous relationship graph to reflect dynamic relationships between entities, entities, data, and data.

[0082] Reference Figure 4 , the method of S400 "building a multi-layer dynamic flow map covering entities, data, users, systems, accounts and roles based on heterogeneous relationship graphs", that is, the method of building a multi-layer dynamic flow map, specifically includes; S410, determine the hierarchy of the initial multi-layer map and the association between each layer.

[0083] The hierarchy includes entity layer, data layer, user layer, system layer, account layer, and role layer.

[0084] The role layer represents different roles in the enterprise, such as administrator, ordinary employee, auditor, etc. Each role has a specific set of responsibilities and permissions, and is an abstract level of permission management in the entire multi-layer map.

[0085] The account layer contains all account information within the enterprise, each account corresponds to a unique identifier, and is associated with a specific role. The account is the specific identity of the user in the system.

[0086] The user layer represents the actual personnel using the enterprise system. A user can have multiple accounts, each account may be associated with different roles, and the user layer embodies the main body of personnel interacting with the system.

[0087] The system layer covers various application systems, database systems, business systems, etc. within the enterprise. These systems are the carriers of data storage and processing, and there may be data interaction and dependency between different systems.

[0088] The entity layer contains various entities in the enterprise business, such as customers, products, orders, etc. The entity is an abstract representation of business data with a clear business meaning.

[0089] The data layer stores the specific data of the enterprise, including structured data and unstructured data. The data is closely related to the entity and is the specific object of business operations.

[0090] The association methods defined include: 1) Role-Account Association: A role can be associated with multiple accounts. This is achieved through a role-account mapping table, which records the role information corresponding to each account and is used for permission allocation and management. 2) Account-User Association: A user can have multiple accounts. Associations are established through a user-account mapping table, which records the correspondence between users and accounts, facilitating user identity management and operation auditing. 3) User-System Association: Users can access different systems through their accounts. This is reflected in a user-system access record table, which records the time, operations, and other information used by users to access the system using their accounts, reflecting the interaction between the user and the system. 4) System-Entity Association: The system is responsible for managing and operating entities. Associations are established through a system-entity mapping table, which records the type and scope of each entity managed by the system, reflecting the business logic relationship between the system and business entities. 5) Entity-Data Association: An entity is an abstraction of data. An entity can correspond to multiple data records. Associations are established through an entity-data index table, which records the correspondence between entities and specific data, facilitating data query and management.

[0091] A clear hierarchical structure helps to organize and manage complex information, making the positioning and classification of each element clear; the clear association method provides a basic framework for subsequent data flow and permission configuration, facilitating system construction and maintenance.

[0092] S420 , mapping entities and data nodes in the heterogeneous relationship graph to corresponding levels of the initial multi-layer map.

[0093] Entity mapping specifically involves traversing the entity nodes in the heterogeneous relationship graph and mapping them to the entity layer of the multi-layer map based on the entity type and business meaning. For example, if the entity is a "customer," it is mapped to the "Customer" category in the entity layer; if it is a "product," it is mapped to the "Product" category. Each entity node mapped to the entity layer is assigned a unique identifier and associated with the corresponding node in the heterogeneous relationship graph for subsequent query and traceability.

[0094] Specifically, data mapping involves mapping data nodes in a heterogeneous relationship graph to the data layers of a multi-layered map based on their entities and data types. For example, data related to the "Customer" entity, such as customer name and contact information, is mapped to the data set corresponding to the "Customer" entity in the data layer. Associations are established between data nodes and entity nodes to ensure that data in the data layer accurately reflects the attributes and status of entities in the entity layer.

[0095] This step realizes the docking of heterogeneous relationship graphs and multi-layer maps, integrates data from different sources and formats, makes the information more unified and orderly, facilitates further processing and analysis of entities and data in multi-layer maps, and avoids data confusion and duplication.

[0096] S430: Building data flow relationships in the initial multi-layer map based on data flow paths and methods between different entities, users, and systems.

[0097] Data flow path analysis specifically involves extracting data flow information between different entities, users, and systems from heterogeneous relationship graphs, including the data's starting, intermediate, and ending nodes, as well as the chronological order and frequency of data flow. It also analyzes the business rules and logic of data flow to determine the triggering conditions and constraints of data flow, such as the approval process required for data to flow between different systems.

[0098] The flow relationship construction specifically involves establishing directed edges between nodes at different levels in a multi-layered map to represent data flow relationships. For example, a directed edge from a user node in the user layer to a system node in the system layer indicates that the user has submitted data to the system; a directed edge from a system node in the system layer to an entity node in the entity layer indicates that the system has updated data on the entity. Attributes are added to each directed edge, including information such as the time of data flow, data volume, and flow method (e.g., real-time transmission, batch transmission), to provide a more detailed description of the data flow process.

[0099] This step intuitively demonstrates the data flow process, helps identify bottlenecks and problems in data flow, optimizes business processes, and provides a basis for subsequent permission control, because different data flow links may require different permissions.

[0100] S440: Determine the operation permissions of different roles on entities and data based on business rules and security policies.

[0101] Business rule analysis specifically involves collecting the company's business rules and process documentation, and analyzing the requirements and restrictions placed on entities and data by different roles in business operations. For example, an administrator role might have create, modify, and delete permissions for all entities and data, while a regular employee role might only have read-only permissions for some entities and data. Based on the company's business development strategy and compliance requirements, the boundaries of operational permissions for different roles in different business scenarios are determined.

[0102] Security policy development specifically involves formulating data access security policies based on the company's security management system and regulatory requirements. For example, for sensitive data, strict access control policies are implemented to ensure access is restricted to authorized roles. Taking into account data confidentiality, integrity, and availability, operational permissions for different roles are refined and graded, such as by setting different access levels and operation types (e.g., read, write, modify, and delete).

[0103] Permission determination involves determining specific operational permissions for each role over entities and data based on business rules and security policies. Each role's operational permissions for each entity and data are recorded using a permissions matrix. The rows of the matrix represent roles, the columns represent entities and data, and the matrix elements represent the role's operational permissions for that entity or data.

[0104] This step ensures the security and integrity of the data and prevents unauthorized access and operations. At the same time, it complies with business rules, and different roles can only perform operations related to their own responsibilities, which improves the standardization and efficiency of the business.

[0105] S450: Configure operation permissions in the initial multi-layer map to obtain a constructed multi-layer dynamic flow map.

[0106] The implementation of permission configuration involves configuring permissions for each node and edge within the multi-layered map's role, account, user, system, entity, and data layers, based on a defined operational permission matrix. For example, at the role layer, each role node is assigned its corresponding permission set; at the account layer, each account node is associated with the permissions of its corresponding role. At the system layer, permission control is configured for the system's access interfaces and functional modules, ensuring that only roles and accounts with appropriate permissions can access and operate.

[0107] The dynamic update mechanism includes establishing a dynamic update mechanism for the multi-layer map. When the enterprise's business rules, security policies, role definitions, account information, etc. change, the permission configuration and data flow relationships in the multi-layer map can be updated in a timely manner. The multi-layer map is regularly audited and evaluated to check the rationality of permission configuration and the compliance of data flow. When problems are found, timely adjustments and optimizations are made, thereby obtaining a multi-layer dynamic flow map that can reflect the enterprise's business and security status in real time.

[0108] This step achieves precise control over the operations of different roles, ensuring the security and stability of the system; the multi-layer dynamic flow map can reflect the flow of data and changes in permissions in real time, providing strong support for corporate decision-making and management.

[0109] The method for constructing a multi-layer dynamic flow map disclosed in this embodiment enables enterprises to better manage and utilize data and reduce data redundancy and errors through a clear hierarchical structure and data flow relationship; role-based permission configuration can effectively prevent data leakage and illegal operations and protect the core information of the enterprise; the intuitive display of the data flow process helps to discover problems in business processes and optimize and improve them; the multi-layer dynamic flow map provides enterprises with a comprehensive information view, helping management make more informed decisions.

[0110] The method of S700 “obtaining a target classification and grading result of the data to be analyzed based on the first classification and grading result and the second classification and grading result” specifically includes: S710, if the first classification and grading results are consistent with the second classification and grading results, the corresponding classification and grading results are used as the target classification and grading results; S720, if the first classification and grading results are inconsistent with the second classification and grading results, obtain the data type of the data to be analyzed; S730, determining the corresponding industry knowledge structure weight and association relationship weight according to the data type; S740, obtaining the classification and grading result corresponding to the largest weight among the industry knowledge structure weight and the association relationship weight, and using it as the target classification and grading result of the data to be analyzed.

[0111] This embodiment can classify and grade data more comprehensively and accurately by comprehensively considering the results of different classification and grading systems as well as the type, industry knowledge and association relationships of the data, thereby reducing the errors that may exist in a single classification and grading system; this solution can handle the situation where the results of different classification and grading systems are inconsistent, has strong adaptability, and is suitable for a variety of different industries and data scenarios; the introduction of industry knowledge structure weights and association relationship weights makes the classification and grading process more scientific and reasonable, avoids subjective arbitrariness, and improves the credibility and authority of the classification and grading results.

[0112] Furthermore, the data classification and grading method based on industry knowledge structure and association relationships disclosed in this application also includes: triggering versioning management of target classification and grading results according to a preset cycle; using a graph difference algorithm based on edit distance and a content difference algorithm based on semantic embedding to automatically identify newly added, deleted, and modified data items and their sensitivity level changes between adjacent versions, and generate a difference report.

[0113] Versioning allows for a clear record of changes to data classification and grading within each preset cycle. This is like taking a photo of each data transformation and arranging them chronologically. To query the classification and grading of a piece of data at a specific point in time, one can directly go back to the corresponding version to clearly identify the data's status at different stages. For example, in medical data management, if errors are subsequently discovered in the classification and grading of a patient's medical record data, versioning allows for rapid identification of the cycle in which the change occurred, as well as the circumstances surrounding the change, facilitating timely error correction and accountability.

[0114] Many industries, such as finance and healthcare, have strict regulations and compliance requirements. Versioning makes the data classification and grading process and results auditable. Regulators or internal auditors can review the classification and grading of different versions of data to ensure that data processing and protection comply with relevant regulations and policies. For example, financial institutions are required to regularly report data security management to regulators. Versioning classification and grading results can serve as strong evidence that the institution's data management is standardized and traceable.

[0115] When business changes, organizational structures adjust, or data processing processes are updated, data classification and grading results may need to be adjusted accordingly. Versioning allows updates without disrupting historical data records, ensuring business continuity. For example, after a business restructuring, a company may need to reassess the importance and sensitivity of data. Versioning can retain the classification and grading results of the previous version as a reference while recording the changes in the new version, ensuring data management stability during the transition period.

[0116] Utilizing graph diffing algorithms based on edit distance and content diffing algorithms based on semantic embedding, we can automatically and quickly identify the addition, deletion, and modification of data items between adjacent versions, as well as changes in sensitivity levels. This significantly improves data monitoring efficiency, eliminating the need for manual comparison of large amounts of data. For example, on large e-commerce platforms, massive amounts of product and user data are updated daily. These algorithms can quickly detect changes in data classification and grading, allowing for timely implementation of appropriate security measures.

[0117] Difference reports clearly identify which data items have changed and how their sensitivity levels have changed, helping companies pinpoint potential risk points. For example, if a data item that was originally classified as low-sensitivity is upgraded to high-sensitivity in a new version, companies can immediately strengthen protection measures for that data to prevent the risk of data leakage.

[0118] Based on the discrepancy report, enterprises can rationally allocate resources for data protection. For newly added data items or those with increased sensitivity, security investments can be increased accordingly; for deleted data items or those with decreased sensitivity, protection resources can be appropriately reduced. For example, in a cloud computing environment, for data with increased sensitivity, storage encryption levels and access control permissions can be increased, while for data with decreased sensitivity, storage costs can be appropriately reduced.

[0119] Difference reports provide clear information for communication between different departments. The data management team can share reports with security teams, business teams, and other teams to ensure that all parties understand data changes. For example, the security team can use the reports to strengthen security measures, while the business team can adjust business strategies based on changes in data sensitivity levels, thereby promoting collaboration between teams and overall business development.

[0120] Furthermore, the data classification and grading method based on industry knowledge structure and association relationships disclosed in this application also includes: submitting the difference report to the user for approval through a multi-role collaborative approval workflow; wherein, the workflow has a built-in explainable presentation module based on role-attribute access control, and the approval process records the complete decision-making trajectory and reason chain; the updated difference content confirmed after approval is merged into the production environment through a zero-downtime grayscale release mechanism, triggering online increments.

[0121] Specifically, it includes: a) adopting a version switching strategy based on a bitemporal database to ensure historical traceability; b) adopting an audit log based on event tracing to record each upgrade action of the classification and grading results; c) adopting a distributed cache invalidation strategy based on consistent hashing to ensure global consistency.

[0122] In this embodiment, regarding "submitting the difference report to the user for approval through a multi-role collaborative approval workflow", assuming that a financial enterprise adopts this data classification and grading method, after the difference report is generated, the system will send the report to different personnel according to the preset roles. For example, it will be sent to the data administrator first. The data administrator will review the difference content of the data classification and grading in the report, evaluate its impact on the existing system and business, and record his or her approval opinions in the system, such as "some differences may affect the risk assessment model and require further review"; the report will then be transferred to the business expert, who will analyze the differences from a business perspective, determine whether they comply with business rules and regulatory requirements, and give approval opinions, such as "some differences comply with new business specifications and can be passed"; finally, the report will be transferred to the security expert, who will evaluate the impact of the differences on data security. If it is confirmed that the security risks are controllable, the entire approval process will be completed.

[0123] The explainable presentation module in the workflow presents the approval basis and decision-making reasons for each role in an intuitive way. For example, the approval opinions, approval times, and relevant reference document links for each role are listed in a table. The decision-making track and reason chain during the approval process are recorded in the system for subsequent audit and traceability.

[0124] Multi-role collaborative approval can evaluate the difference report from different professional perspectives, avoiding the limitations of a single role, and making more comprehensive and accurate decisions. The complete record of decision-making track and reason chain makes the entire approval process transparent and traceable, facilitating subsequent audit and compliance inspection, and also helping to quickly locate responsibilities when problems occur. In many industries, such as finance and medicine, there are strict compliance requirements for data changes. Multi-role collaborative approval and complete records can meet these compliance requirements.

[0125] For the "approved and updated difference content is merged into the production environment through a zero-downtime gray release mechanism, triggering online incremental", the version switching strategy based on dual-time database is adopted to ensure historical traceability. Taking the data classification and grading system of an e-commerce enterprise as an example, a dual-time database is used. When the approved difference content needs to be merged into the production environment, the system creates a new database version. In a dual-time database, each data record has two time dimensions: valid time and transaction time. Valid time represents the valid time period of data in the real world, and transaction time represents the time when data is recorded in the database. When new difference content is merged, it sets new valid time and transaction time for these data. For example, the data classification of a certain commodity changes from "ordinary commodity" to "hot commodity", and the new record will take effect in the new version, while the old record will still be retained. Through valid time and transaction time, the history of data changes can be clearly traced.

[0126] This scheme can facilitate the query of data status at different time points, which is helpful for data analysis, compliance audit and problem troubleshooting; the version switching strategy of dual-time database makes different versions of data clear and manageable, avoiding data loss or confusion; if the newly merged difference content has problems, it can quickly switch back to the old version to ensure system stability.

[0127] For "adopting event-based audit logs to record every upgrade action of classification results", for example, for a data classification system of a logistics company, the system generates an event record every time there is an upgrade action of classification results. For example, when the data classification of a batch of goods is upgraded from "ordinary goods" to "high-value goods", the event record contains detailed information of the event, such as the time of the event, the user or system component that triggered the event, the classification results before and after the upgrade, etc. These event records are stored in the audit log in chronological order, forming a complete event stream.

[0128] The audit log can record every upgrade action of classification results in detail, providing complete evidence for subsequent audits and compliance checks; when data anomalies or business problems occur, the audit log can be used to quickly locate the time and cause of the problem, helping to solve problems in a timely manner; by analyzing the events in the audit log, the trend of data classification can be understood, providing reference for business decision-making.

[0129] For "adopting a consistent hash-based distributed cache invalidation strategy to ensure global consistency", for example, assuming a large Internet company has multiple distributed cache nodes to store data classification results. When new differential content is integrated into the production environment, the data in the cache needs to be updated to ensure global consistency. The consistent hash algorithm is used to map the data key to a hash ring, and each cache node also has a corresponding position on the hash ring. When data changes, only the affected data keys need to be recalculated on the hash ring according to the consistent hash algorithm, and then the corresponding cache nodes are updated. For example, when the classification result of a certain type of data changes, the system will determine which cache nodes store this type of data according to the consistent hash algorithm, and then only update these nodes' cache without affecting other unrelated cache nodes.

[0130] The consistent hash-based distributed cache invalidation strategy can ensure that all cache nodes maintain consistent data in a distributed environment, avoiding business errors caused by inconsistent data; only updating the affected cache nodes avoids full update of all cache nodes, reducing system overhead and performance loss; when adding or reducing cache nodes, the consistent hash algorithm can smoothly adjust the distribution of data, ensuring the scalability of the system.

[0131] Through multi-role collaborative approval and complete record keeping, updates to data classification and grading are ensured to comply with business rules and compliance requirements. The bi-temporal database and event-sourced audit logs further enhance the traceability and transparency of data management, helping to meet various regulatory requirements. The zero-downtime grayscale release mechanism allows the incorporation of new differential content without affecting the normal operation of the system, reducing the impact on the business. The version switching strategy based on the bi-temporal database and the distributed cache invalidation strategy based on consistent hashing ensure data consistency and rollback, improving the stability and reliability of the system. The timely incorporation of approved differential content into the production environment, triggering online increments, can quickly respond to changes in business needs and improve business flexibility and competitiveness. At the same time, the traceability of historical data and event audit logs also facilitate business analysis and decision-making, promoting continuous business optimization.

[0132] In addition, traditional methods require a lot of manual intervention, including manual labeling, identification and classification, which is very costly. The solution of this application reduces dependence on manual labor through automated and intelligent processes. For example, in the data capture and preprocessing stage, the system can automatically complete most of the work, reducing labor costs. At the same time, the continuous investment work model avoids the waste of resources caused by one-time large investments, enabling enterprises to allocate resources more reasonably, further reducing costs. Automated processes and real-time monitoring mechanisms make the sorting of data elements more efficient. The system can quickly process and classify newly generated data, reducing data processing time. Moreover, the comprehensive data association information provided by the multi-layer dynamic flow map helps enterprises identify the type and sensitivity of data more quickly, improves the speed and accuracy of classification and grading, and thus improves the efficiency of data element sorting work as a whole.

[0133] A single classification and grading method may have limitations and be easily affected by factors such as incomplete data and inaccurate analysis methods. The method disclosed in this application can verify and complement each other by combining two different classification and grading results, thereby improving the reliability and credibility of the classification and grading. If the first classification and grading results are consistent with the second classification and grading results, then the accuracy of the classification and grading can be more assured; if there is a difference between the two, the cause of the difference can be further analyzed, and a more in-depth evaluation and judgment can be conducted to obtain a more reliable target classification and grading result. Reliable classification and grading results provide a solid foundation for the company's data security governance. Companies can formulate more effective data protection strategies, access control rules and emergency response plans based on accurate classification and grading results to reduce data security risks and ensure the normal operation of the company.

[0134] As the enterprise's business develops and the external regulatory environment changes, the classification and grading of data also needs to be continuously adjusted and optimized. The target classification and grading results obtained based on the first and second classification and grading results can be used as a dynamic reference. When regulations change, the first classification and grading results can be adjusted according to the new regulatory requirements. When business processes or data flow conditions change, the second classification and grading results can be updated. By continuously updating and optimizing these two classification and grading results and recalculating the target classification and grading results, enterprises can achieve continuous improvement in data governance, ensuring that data classification and grading always adapt to the actual needs and regulatory requirements of the enterprise. This continuous optimization mechanism helps enterprises maintain the advancement and effectiveness of data management, improve the enterprise's data security level and competitiveness. At the same time, it also provides strong support for the digital transformation and innovative development of enterprises, enabling enterprises to better cope with increasingly complex data security challenges.

[0135] In a second aspect, the present application discloses a data classification and grading system based on industry knowledge structure and association relationships, which is used to implement the data classification and grading method based on industry knowledge structure and association relationships disclosed in the first aspect of the present application. The system specifically includes: The preprocessing module is used to capture the original data fragments in the enterprise's business traffic and preprocess the original data fragments to obtain structured data metadata; The associated information acquisition module is used to perform type inference and name semantic analysis on each data item in the structured data metadata to obtain the associated information of each data item; The heterogeneous relationship graph acquisition module is used to extract entity relationships from all data items based on association information, identify dynamic associations between entities, entities, data, and data, and generate weighted heterogeneous relationship graphs. A multi-layer dynamic flow map construction module is used to build a multi-layer dynamic flow map covering entities, data, users, systems, accounts, and roles based on heterogeneous relationship graphs; A first classification and grading result acquisition module is used to obtain the first classification and grading results of the data to be analyzed based on the industry regulations knowledge base; A second classification and grading result acquisition module is used to obtain the second classification and grading results of the data to be analyzed based on the multi-layer dynamic flow map; The target classification and grading result acquisition module is used to obtain the target classification and grading results of the data to be analyzed based on the first classification and grading results and the second classification and grading results.

[0136] A computer device according to an embodiment of the present disclosure includes a memory and a processor. The memory is configured to store non-transitory computer readable instructions. Specifically, the memory can include one or more computer program products that can include various forms of computer readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory, among others. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, among others.

[0137] The processor can be a central processing unit (CPU) or other form of processing unit that has data processing and / or instruction execution capabilities, and can control other components in the computer device to perform desired functions. In one embodiment of the present disclosure, the processor is configured to execute the computer readable instructions stored in the memory, so that the computer device performs all or part of the steps of the data classification and grading method based on industry knowledge structure and association relationship according to the aforementioned embodiments of the present disclosure.

[0138] Those skilled in the art will understand that, in order to solve the technical problem of how to obtain a good user experience effect, the present embodiment can also include well-known structures such as communication buses, interfaces, etc., which should also be included in the protection scope of the present disclosure.

[0139] As Figure 5 A structural schematic diagram of a computer device according to an embodiment of the present disclosure is shown. It shows a structural schematic diagram suitable for implementing the computer device in the embodiments of the present disclosure. Figure 5 The computer device shown is only an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.

[0140] As Figure 5 As shown, the computer device can include a processor (such as a central processor, a graphics processor, etc.) that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) or loaded from a storage device into a random access memory (RAM). In the RAM, various programs and data required for the operation of the computer device are also stored. The processor, the ROM, and the RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.

[0141] Generally, the following devices can be connected to the I / O interface: input devices including, for example, sensors or visual information acquisition devices, etc.; output devices including, for example, display screens, etc.; storage devices including, for example, magnetic tapes, hard disks, etc.; and communication devices. The communication device can allow the computer device to communicate wirelessly or by wire with other devices (such as edge computing devices) to exchange data. Although Figure 5A computer device having various devices is shown, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0142] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the data classification and grading method based on industry knowledge structure and association relationship of the embodiment of the present disclosure are executed.

[0143] For detailed description of this embodiment, please refer to the corresponding description in the aforementioned embodiments, which will not be repeated here.

[0144] According to an embodiment of the present disclosure, a computer-readable storage medium stores non-transitory computer-readable instructions. When the non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the data classification and grading method based on industry knowledge structure and association relationships described in each embodiment of the present disclosure are executed.

[0145] The above-mentioned computer-readable storage media include, but are not limited to, optical storage media (e.g., CD-ROMs and DVDs), magneto-optical storage media (e.g., MOs), magnetic storage media (e.g., magnetic tapes or mobile hard disks), media with built-in rewritable non-volatile memory (e.g., memory cards), and media with built-in ROM (e.g., ROM cartridges).

[0146] For detailed description of this embodiment, please refer to the corresponding description in the aforementioned embodiments, which will not be repeated here.

[0147] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.

[0148] In this disclosure, relational terms such as first and second and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. The block diagram of the devices, apparatus, equipment, systems referred to in this disclosure is merely illustrative and not intended to imply the necessity or arrangement of the connections, arrangement, configuration as shown in the block diagram. As will be appreciated by those skilled in the art, the devices, apparatus, equipment, systems can be connected, arranged, configured in any manner. The words comprising, including, having and the like are to be open ended. As used in this document, the conjunction "or" is to be interpreted in the inclusive sense, i.e. as meaning one or the other, or both. As used in this document, the words "and" and "or" are to be interpreted as having the meaning indicated in the phrase "and / or". As used in this document, the word "such as" is to be interpreted as meaning "such as, but not limited to".

[0149] Also, as used in this document, the word "or" in the cases used to introduce an enumeration of several items, for example, a list of items, is to be interpreted in the inclusive sense, i.e. as meaning one or more, or any combination thereof, of the listed items. Additionally, the phrase "example of" as used in this document is not meant to be limiting in any way. It is not meant to imply that the described example is the only example of the described feature.

[0150] It is also important to note that the systems and methods of the present disclosure can be embodied in a variety of forms including, but not limited to, a data processor, a computer program product, a computer, one or more tangible computer readable storage devices, one or more computer memories, one or more programmable logic devices, one or more application specific devices, one or more computers, one or more processors, one or more microprocessors, one or more microcomputers, one or more microcontrollers, one or more microcontrollers, one or more microprocessors, one or more state machines, one or more integrated circuits, one or more other components, or any combination thereof, and that the systems and methods can comprise, consist of, or consist essentially of such forms.

[0151] Various changes, modifications and alterations in the teachings and techniques described herein can be made without departing from the teachings and techniques defined by the appended claims. Moreover, the scope of the claims of this disclosure is not limited to the specific aspects described above. Rather, the aspects of the disclosure are meant to cover all alternatives, modifications, and equivalents falling within the scope of the claims.

[0152] The foregoing description of the disclosed aspects is intended to be illustrative only and is not intended to be limiting on the aspects of the disclosure. Various modifications to those aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the claims, the principles and the novel features disclosed herein.

[0153] The foregoing description has been presented for the purposes of illustration and description. Furthermore, the description is not intended to limit the embodiments of the disclosure to the forms disclosed herein. Although the various example aspects and embodiments have been described herein with regard to particular aspects and embodiments, those skilled in the art will recognize that certain modifications, changes, substitutions, additions and sub-combinations can be made without departing from the spirit of the disclosure.

Claims

1. A data classification and grading method based on industry knowledge structure and association relationships, characterized by: include: Capturing raw data segments from the enterprise's business traffic and preprocessing the raw data segments to obtain structured data metadata; Performing type inference and name semantic parsing on each data item in the structured data metadata to obtain associated information of each data item; Extracting entity relationships from all the data items based on the association information, identifying dynamic association relationships between entities, entities, data, and data, and generating a heterogeneous relationship graph; Based on the heterogeneous relationship graph, a multi-layer dynamic flow map covering entities, data, users, systems, accounts, and roles is constructed; Obtain the first classification and grading results of the data to be analyzed based on the industry regulations knowledge base; Obtaining a second classification and grading result of the data to be analyzed according to the multi-layer dynamic flow map; A target classification and grading result of the data to be analyzed is obtained based on the first classification and grading result and the second classification and grading result.

2. The data classification and grading method based on industry knowledge structure and association relationship according to claim 1 is characterized in that: The process of capturing raw data segments from the enterprise's business traffic and preprocessing the raw data segments to obtain structured data metadata includes: Capture raw data fragments in business traffic through bypass monitoring technology; Based on a sliding window conversation reorganization algorithm, all the original data fragments are reorganized into a complete conversation corresponding to the business interaction process; A global unified session ID, a client certificate fingerprint, a server certificate fingerprint, a user identifier, and a nanosecond timestamp are appended to each data fragment in the complete session to form structured data metadata.

3. The data classification and grading method based on industry knowledge structure and association relationship according to claim 1 is characterized in that: The performing type inference and name semantic parsing on each data item in the structured data metadata to obtain associated information of each data item includes: Construct a multi-layer knowledge graph covering industry terms, synonyms, and abbreviations; the multi-layer knowledge graph includes a core layer knowledge graph, an extended layer knowledge graph, and a customized layer knowledge graph; Using a BERT-BiLSTM-CRF cascade network to perform sequence labeling on the names of data items in the structured data metadata to obtain entity boundaries; Linking entities in the entity boundary with nodes in the multi-layer knowledge graph, and using the PageRank algorithm to disambiguate the link results to obtain a semantic label for each data item; Extracting data content of each data item in the structured data metadata; Using data samples of known types to train the target model, analyzing the data content of each data item using the trained model to obtain a content type label for each data item; The association information includes a semantic tag and a content type tag of the data item.

4. The data classification and grading method based on industry knowledge structure and association relationship according to claim 1 is characterized in that: The method of constructing a multi-layer dynamic flow map covering entities, data, users, systems, accounts, and roles based on the heterogeneous relationship graph includes: Determine the hierarchical structure of the initial multi-layer map and the relationship between the layers; wherein the hierarchical structure includes an entity layer, a data layer, a user layer, a system layer, an account layer, and a role layer; Mapping entities and data nodes in the heterogeneous relationship graph to corresponding levels of the initial multi-layer map; Based on the data flow paths and methods between different entities, users, and systems, build data flow relationships in the initial multi-layer map; Determine the operational permissions of different roles on entities and data based on business rules and security policies; The operation permissions are configured in the initial multi-layer map to obtain a constructed multi-layer dynamic flow map.

5. The data classification and grading method based on industry knowledge structure and association relationship according to claim 1 is characterized in that: The obtaining of a target classification and grading result of the data to be analyzed based on the first classification and grading result and the second classification and grading result includes: If the first classification and grading results are consistent with the second classification and grading results, the corresponding classification and grading results are used as the target classification and grading results; If the first classification and grading results are inconsistent with the second classification and grading results, obtaining the data type of the data to be analyzed; Determine the corresponding industry knowledge structure weight and association relationship weight according to the data type; The classification and grading result corresponding to the largest weight among the industry knowledge structure weight and the association relationship weight is obtained and used as the target classification and grading result of the data to be analyzed.

6. The data classification and grading method based on industry knowledge structure and association relationship according to claim 1 is characterized in that: Also includes: Triggering versioning management of the target classification and grading results according to a preset period; Using the graph difference algorithm based on edit distance and the content difference algorithm based on semantic embedding, it automatically identifies the newly added, deleted, and modified data items and their sensitivity level changes between adjacent versions and generates a difference report.

7. The data classification and grading method based on industry knowledge structure and association relationship according to claim 6 is characterized in that: Also includes: Submit the difference report to the user for approval through a multi-role collaborative approval workflow. The workflow has a built-in explainable presentation module based on role-attribute access control, and the approval process records the complete decision-making trajectory and reasoning chain. The approved difference content is merged into the production environment through the zero-downtime grayscale release mechanism, triggering an online incremental update.

8. A data classification and grading system based on industry knowledge structure and association relationships, characterized by: include: A preprocessing module is used to capture raw data segments from the enterprise's business traffic and preprocess the raw data segments to obtain structured data metadata; An associated information acquisition module is used to perform type inference and name semantic analysis on each data item in the structured data metadata to obtain associated information of each data item; A heterogeneous relationship graph acquisition module is used to extract entity relationships from all the data items based on the association information, identify dynamic association relationships between entities, entities and data, and data and data, and generate a weighted heterogeneous relationship graph; A multi-layer dynamic flow map construction module is used to construct a multi-layer dynamic flow map covering entities, data, users, systems, accounts and roles based on the heterogeneous relationship graph; A first classification and grading result acquisition module is used to obtain the first classification and grading results of the data to be analyzed based on the industry regulations knowledge base; A second classification and grading result acquisition module is used to obtain a second classification and grading result of the data to be analyzed according to the multi-layer dynamic flow map; The target classification and grading result acquisition module is used to obtain the target classification and grading result of the data to be analyzed based on the first classification and grading result and the second classification and grading result.

9. A computer device, characterized in that: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data classification and grading method based on industry knowledge structure and association relationship described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to enable a computer to execute the data classification and grading method based on industry knowledge structure and association relationship described in any one of claims 1-7.

Citation Information

Patent Citations

  • Maritime accident assisting method and device based on knowledge graph and electronic equipment

    CN116860984A

  • Enterprise-level knowledge base construction method based on large model

    CN119622040A

  • Multi-source heterogeneous data knowledge base system construction method, equipment and medium

    CN120386896A

  • Enterprise intelligent decision-making method and system driven by causal atlas

    CN120542981A