Enterprise data grading and classifying method based on base large model and deep learning, storage medium and equipment

Through the combination of base large model and deep learning, the problem of low accuracy and credibility of enterprise data classification and grading is solved, accurate enterprise data classification is achieved, and the efficiency and accuracy of data security management is improved.

CN120277216APending Publication Date: 2025-07-08JIANGSU HONGXIN SYST INTEGRATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510340680.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing technology has inefficient, strong subjective, and grading disorder in enterprise data classification and classification, and a single data classification and grading automated identification algorithm is difficult to fully capture enterprise data characteristics, resulting in low accuracy and credibility of classification and grading.

Method used

Using a method based on base large model and deep learning, combining rule databases, synonyms and multiple deep learning models, we realize accurate hierarchical classification of enterprise data through rule matching, synonyms expansion and external knowledge base query.

Benefits of technology

It improves the accuracy and credibility of enterprise data classification and grading, reduces misclassification and grading, reduces manual workload, and enhances the guarantee of data security management and compliance use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277216A_ABST
    Figure CN120277216A_ABST
Patent Text Reader

Abstract

The invention discloses an enterprise data grading and classification method based on a base large model and deep learning, a storage medium and equipment, and the method comprises the steps: constructing a rule library according to a rule document, and constructing a synonym library according to a related corpus of an industry; performing synonym expansion on the data item name of each piece of enterprise data in the catalogue through a synonym library; matching the data item names of the expanded enterprise data by using a rule base, and if the data item names are matched, taking the corresponding levels and categories as grading and classification results of the enterprise data; otherwise, predicting a first grading classification result of the enterprise data through various deep learning models, and predicting a second grading classification result of the enterprise data through a base large model; and splicing the data item name of the expanded enterprise data, the first grading classification result and the second grading classification result, inputting XGBoost, and predicting a grading classification result of the enterprise data. According to the invention, the accuracy and credibility of enterprise data classification and grading are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data analysis, and specifically, to a method, a storage medium, and a device for classifying and grading enterprise data based on a base large model and deep learning. Background Art

[0002] With the acceleration of the digital transformation of enterprises and the explosive growth of data volume, the degree of attention to data security has been continuously increasing. Therefore, it is necessary to establish a classification and grading management method for enterprise data to identify important data and core data, so as to establish corresponding data security protection measures and avoid harm to national security, public interests, and the legitimate rights and interests of individuals and organizations.

[0003] At present, when most industries are carrying out the work of classifying and grading enterprise data, they mainly rely on relevant policies and regulations and the enterprise's own business, and rely on manual management of enterprise data to configure and adjust the categories and levels of enterprise data, which has problems of low efficiency, strong subjectivity, chaotic classification, and inconsistent results. In addition, some current enterprise data introduce a single automated recognition algorithm for data classification and grading, but the single automated recognition algorithm for data classification and grading is difficult to comprehensively capture the characteristics of enterprise data, thus reducing the accuracy and credibility of enterprise data classification and grading. Summary of the Invention

[0004] In view of the problems existing in the prior art, the present invention provides a method, a storage medium, and a device for classifying and grading enterprise data based on a base large model, which improve the accuracy and credibility of enterprise data classification and grading by integrating the base large model and deep learning algorithms.

[0005] To achieve the above technical objectives, the present invention adopts the following technical solutions: A method for classifying and grading enterprise data based on a base large model and deep learning specifically includes the following steps:

[0006] Step S1: Obtain the levels and categories of the corresponding enterprise data from the rule document, and form key-value pairs with the text information of the enterprise data and the corresponding levels and categories, and store them in the rule library;

[0007] Step S2: Collect relevant corpora of the industry, perform text data extraction and segmentation, and then map the segmented text data into vectors in a low-dimensional vector space to construct a thesaurus;

[0008] Step S3: Catalog the enterprise data, and expand the synonyms of the data item names of each enterprise data in the catalog through the constructed thesaurus;

[0009] Step S4: Match the data item names of the augmented enterprise data using the key-value pairs in the rule base. If there is a match, use the corresponding level and category as the classification result of the enterprise data; otherwise, execute Step S5;

[0010] Step S5: Extract the text information of the enterprise data, convert it into a feature matrix, and input it into various deep learning models respectively to predict the first classification result of the enterprise data;

[0011] Step S6: Construct the prompt words of the base large model, query the external knowledge base using the RAG method with the extracted text information of the enterprise data, retrieve the relevant knowledge rule content, integrate it with the extracted text data, and input it into the base large model to predict the second classification result of the enterprise data;

[0012] Step S7: Concatenate the data item names, the first classification result, and the second classification result of the augmented enterprise data, and input them into XGBoost to predict the classification result of the enterprise data.

[0013] Furthermore, the text information of the enterprise data includes: data item name, content description, and source information.

[0014] Furthermore, Step S2 includes the following sub-steps:

[0015] Step S2.1: Convert the data in the relevant corpus into a text-based document, perform text data extraction, and after deduplication, filtering, compression, and formatting, extract the key information of the document;

[0016] Step S2.2: Split the key information of the document using the character splitter of langchain, map the split text data into vectors in a low-dimensional vector space to form a thesaurus.

[0017] Furthermore, the process of augmenting the data item names of each enterprise data in the catalog using the constructed thesaurus is as follows: Calculate the semantic similarity between the data item names of the enterprise data and the vectors mapped in the thesaurus, and use the text data corresponding to the vectors that exceed the semantic similarity threshold as the synonyms of the data item names.

[0018] Further, the deep learning models used in step S5 include: LightGBM, Random Forest, BERT, and RoBERTa; before predicting the first classification result of enterprise data, the deep learning models need to be trained. The specific process is as follows: using the text information of enterprise data in the rule base as the training samples of the deep learning models, performing one-hot encoding on the levels and categories corresponding to the training samples as the training labels, with the output of the deep learning models being the first classification result of enterprise data. Calculate the mean squared error loss function between the output first classification result of enterprise data and the corresponding training labels, and optimize the parameters of the deep learning models until the mean squared error loss function converges, thus completing the training of the deep learning models.

[0019] Further, the base large model adopts any one of M3E-base, LLaMA, and GLM.

[0020] Further, the prompts of the base large model include: task description, background knowledge, answer examples, and answer requirements. The task description is to answer the classification result of enterprise data according to the data item name of the enterprise data input by the user, which can refer to the answer examples and comply with the answer requirements; the background knowledge is the integrated relevant knowledge rule content and the extracted text data; the answer examples include: the data item name of the enterprise data is: xxx, the classification result of the enterprise data is: xxx, the classification result of the enterprise data is: xxx, the retrieved relevant knowledge rule content is: xxx, and the source of the retrieved relevant knowledge rule is: xxx; the answer requirement is that you need to answer strictly according to the content of the background knowledge, and directly answer "no relevant answer found" for information you don't know.

[0021] Further, when industry rules or enterprise internal rules change, the rule base and relevant corpora need to be updated in a timely manner to achieve dynamic management of enterprise data classification.

[0022] Further, the present invention also provides a computer-readable storage medium storing a computer program, and the computer program causes a computer to execute the enterprise data classification method based on a base large model and deep learning.

[0023] Further, the present invention also provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the enterprise data classification method based on a base large model and deep learning.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] (1) The enterprise data classification and grading method based on the base large model and deep learning in the present invention preferentially uses the rule base for rule matching during the enterprise data classification and grading process, and takes the level and category matched with the rules as the classification and grading results of enterprise data to achieve precise matching. For the enterprise data that fails to match, the initial classification and grading results are generated with the help of the base large model and the deep learning model, and then the XGBoost is used to predict the classification and grading results of enterprise data, which can integrate the advantages of each model, improve the accuracy of enterprise data classification and grading, provide guarantee for the security management and compliant use of enterprise data, and effectively reduce the workload of manual classification and grading.

[0026] (2) The base large model used in the enterprise data classification and grading method based on the base large model and deep learning in the present invention can more accurately parse complex text data, identify the semantic gap, and query the external knowledge base in the RAG manner, enhancing the understanding ability of the base large model for enterprise data, reducing the situations of misclassification and misgrading, and improving the feasibility of enterprise data classification and grading.

[0027] (3) The enterprise data classification and grading method based on the base large model and deep learning in the present invention constructs a thesaurus to expand the synonyms of the data item names of each enterprise data, so that the semantic information of enterprise data can be enriched during subsequent rule matching, improving the matching efficiency. At the same time, the data item names expanded with synonyms can better understand the potential connections between enterprise data when predicting the classification and grading results by deep learning and the base large model, improving the learning efficiency and prediction accuracy. Description of the Drawings

[0028] Figure 1 is the flowchart of the enterprise data classification and grading method based on the base large model and deep learning in the present invention;

[0029] Figure 2 is the schematic diagram of the prompt word template of the base large model in the present invention. Detailed Embodiment

[0030] The technical solution of the present invention will be further explained below with reference to the drawings.

[0031] As Figure 1 is the flowchart of the enterprise data classification and grading method based on the base large model and deep learning in the present invention, and the enterprise data classification and grading method specifically includes the following steps:

[0032] Step S1: Obtain the level and category of the corresponding enterprise data from the rule document, and form key-value pairs with the text information of the enterprise data and the corresponding level and category, and store them in the rule base. Among them, the text information of the enterprise data includes: data item name, content description, and source information.

[0033] Step S2: After collecting the relevant corpora of the industry, performing text data extraction and segmentation, map the segmented text data into vectors in a low-dimensional vector space, and construct a thesaurus; including the following sub-steps:

[0034] Step S2.1: Collect industry term dictionaries, industry standard specification documents, rule documents, and similar meaning expressions commonly used in the daily operations of enterprise internal business departments. Use this information as the relevant corpora. Since the data in the relevant corpora exists in various formats, such as word, pdf, markdown, etc., it is necessary to convert the data in the relevant corpora into a text version of the document. Use Baidu's PP-Structurev2 to achieve layout analysis, table recognition, etc., complete text data extraction, and after deduplication, filtering, compression, and formatting processing, extract the key information of the document, such as file name, chapter, title, interval, etc.;

[0035] Step S2.2: Considering the token limit of the embedding model and that semantic integrity will affect the overall retrieval effect, use the character splitter of langchain to split the key information of the document. Here, the fixed character length split chunk_size = 128 is used. Map the segmented text data into vectors in a low-dimensional vector space to form a thesaurus, and corresponding indexes can be created for the vectors to enable fast text retrieval functions.

[0036] Step S3: Catalog the enterprise data so that each piece of enterprise data has text information including a data item name, content description, and source information, and remove the enterprise data with identical data item names and content descriptions. The content description provides a detailed textual explanation of aspects such as the scope, definition, specific content, and data format of the enterprise data, helping to more deeply understand the essential characteristics of the enterprise data and serving as auxiliary information for classification and grading. The source information can identify the departments responsible for generating and maintaining the enterprise data, trace the origin of the enterprise data, understand the initial generation background and business logic of the enterprise data, and enhance the judgment of the correlation between enterprise data. Expand the synonyms of the data item names of each enterprise data in the catalog using the constructed thesaurus. Specifically, calculate the semantic similarity between the data item names of the enterprise data and the vectors mapped in the thesaurus, such as cosine similarity calculation, Euclidean distance calculation, Manhattan distance calculation, etc. Use the text data corresponding to the vectors with a semantic similarity exceeding the threshold as the synonyms of the data item names. Expand the synonyms of the data item names of each enterprise data using the constructed thesaurus, enabling the enrichment of the semantic information of the enterprise data during subsequent rule matching and improving the matching efficiency. At the same time, the data item names expanded with synonyms can better understand the potential connections between enterprise data during deep learning and the prediction of classification and grading results by the base large model, enhancing the learning efficiency and prediction accuracy. In addition, the synonyms of the data item names expanded for the enterprise data need to be manually reviewed to ensure data quality.

[0037] Step S4: Match the data item names of the expanded enterprise data using the key-value pairs in the rule library. If there is a match, use the corresponding level and category as the classification and grading result of the enterprise data to achieve precise matching; otherwise, execute Step S5.

[0038] Step S5: Extract the text information of the enterprise data, convert it into a feature matrix, and input it into various deep learning models respectively to predict the first classification and grading result of the enterprise data. The deep learning models used in the present invention include: LightGBM, random forest, BERT, and RoBERTa. The deep learning models improve the accuracy of enterprise data classification and grading by capturing complex features in the feature matrix.

[0039] Before predicting the first classification and grading result of the enterprise data, the deep learning model of the present invention needs to be trained. The specific process is as follows: Use the text information of the enterprise data in the rule library as the training samples of the deep learning model, perform one-hot encoding on the corresponding levels and categories of the training samples as the training labels. The output of the deep learning model is the first classification and grading result of the enterprise data. Calculate the mean squared error loss function between the output first classification and grading result of the enterprise data and the corresponding training labels, and optimize the parameters of the deep learning model until the mean squared error loss function converges to complete the training of the deep learning model.

[0040] Step S6: Construct the prompt words of the base large model, query the external knowledge base for the text information of the extracted enterprise data using the RAG method, retrieve the relevant knowledge rule content, ensure that the enterprise data is updated in real time, and each answer is based on the retrieved evidence. Integrate the retrieved relevant knowledge rule content with the extracted text data, input it into the base large model, predict the second-level classification result of the enterprise data, which can more accurately analyze complex text data, identify the gap between semantics, and query the external knowledge base using the RAG method to enhance the understanding ability of the base large model for enterprise data, reduce the situations of misclassification and misgrading, and improve the feasibility of enterprise data classification and grading.

[0041] In the present invention, the base large model adopts any one of M3E-base, LLaMA, and GLM. For example Figure 2 , the prompt words of the base large model include: task description, background knowledge, answer examples, and answer requirements. Among them, the task description is to answer the classification and grading results of the enterprise data according to the data item name of the enterprise data input by the user, which can refer to the answer examples and comply with the answer requirements; the background knowledge is the integrated relevant knowledge rule content and the extracted text data; the answer examples include: the data item name of the enterprise data is: xxx, the grading result of the enterprise data is: xxx, the classification result of the enterprise data is: xxx, the retrieved relevant knowledge rule content is: xxx, and the source of the retrieved relevant knowledge rule is: xxx; the answer requirement is that you need to answer strictly according to the content of the background knowledge, and directly answer "no relevant answer found" for the information you don't know.

[0042] Step S7: Concatenate the data item names, the first-level classification and grading results, and the second-level classification and grading results of the augmented enterprise data to construct a more comprehensive feature representation, input it into XGBoost, and predict the classification and grading results of the enterprise data. By combining the advantages of deep learning and the base large model through XGBoost, a more comprehensive and stable prediction result can be obtained, further reducing the error, improving the accuracy of enterprise data classification and grading, providing guarantee for the security management and compliance use of enterprise data, and effectively reducing the workload of manual classification and grading.

[0043] When industry rules or internal enterprise rules change, it is necessary to update the rule library and relevant corpora in a timely manner to achieve dynamic management of enterprise data classification and grading. Through data management-related tools, such as data quality management platforms, key indicators of enterprise data are detected in real time or regularly. These indicators reflect the characteristics of enterprise data, such as accuracy, integrity, consistency, timeliness, etc. Once it is found that the relevant indicators of enterprise data exceed the preset thresholds or show abnormal changes, the classification and grading are re-examined in a timely manner. At the same time, a supervision mechanism is established, and personnel with rich experience, familiar with business and regulatory policies are arranged to closely monitor the dynamic development of enterprise business and external standard specifications and policies, evaluate the scope and degree of impact on enterprise data, and readjust the content of the corpus to achieve dynamic management of data category levels.

[0044] In one technical solution of the present invention, there is also provided a computer-readable storage medium storing a computer program, and the computer program causes a computer to execute the enterprise data classification and grading method based on a base large model and deep learning.

[0045] In one technical solution of the present invention, there is also provided an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the enterprise data classification and grading method based on a base large model and deep learning is implemented.

[0046] In the embodiments disclosed in the present application, the computer storage medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the computer storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0047] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0048] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.

Claims

1. An enterprise data grading and classification method based on a base large model and deep learning, characterized in that, Specifically, it includes the following steps: Step S1: Obtain the level and category of the corresponding enterprise data from the rule document, form key-value pairs by combining the text information of the enterprise data with the corresponding level and category, and store them in the rule library; Step S2: Collect relevant corpora of the industry. After text data extraction and segmentation, map the segmented text data into vectors in a low-dimensional vector space, and construct a thesaurus; Step S3: Catalog the enterprise data, and expand the synonyms of the data item names in each enterprise data during cataloging through the constructed thesaurus; Step S4: Match the expanded data item names of the enterprise data with the key-value pairs in the rule library. If there is a match, use the corresponding level and category as the hierarchical classification result of the enterprise data; Otherwise, execute Step S5; Step S5: Extract the text information of the enterprise data, convert it into a feature matrix, and input it into various deep learning models respectively to predict the first hierarchical classification result of the enterprise data; Step S6: Construct prompts for the base large model, query the external knowledge base using the RAG method with the extracted text information of the enterprise data, retrieve relevant knowledge rule content, integrate it with the extracted text data, and input it into the base large model to predict the second hierarchical classification result of the enterprise data; Step S7: Concatenate the expanded data item names, the first hierarchical classification result, and the second hierarchical classification result of the enterprise data, and input them into XGBoost to predict the hierarchical classification result of the enterprise data.

2. The enterprise data grading and classification method based on the base large model and deep learning according to claim 1, characterized in that, The text information of the enterprise data includes: data item name, content description, and source information.

3. The enterprise data grading and classification method based on a base large model and deep learning according to claim 2, wherein Step S2 includes the following sub-steps: Step S2.1: Convert the data in the relevant corpus into a text version of the document, perform text data extraction, and after deduplication, filtering, compression, and formatting processing, extract the key information of the document; Step S2.2: Use the character splitter of langchain to split the key information of the document, map the segmented text data into vectors in a low-dimensional vector space, and form a thesaurus.

4. A method for classifying enterprise data hierarchically based on a base large model and deep learning according to claim 3, characterized in that, The process of expanding the synonyms of the data item names in each enterprise data during cataloging in Step S3 is as follows: Calculate the semantic similarity between the data item names of the enterprise data and the vectors mapped in the thesaurus, and use the text data corresponding to the vectors whose semantic similarity exceeds the threshold as the synonyms of the data item names.

5. A method for classifying enterprise data hierarchically based on a base large model and deep learning according to claim 4, characterized in that, The deep learning models used in Step S5 include: LightGBM, random forest, BERT, and RoBERTa; before predicting the first hierarchical classification result of the enterprise data, the deep learning models need to be trained. The specific process is as follows: Use the text information of the enterprise data in the rule library as the training samples of the deep learning model, perform one-hot encoding on the corresponding levels and categories of the training samples as the training labels, the output of the deep learning model is the first hierarchical classification result of the enterprise data, calculate the mean squared error loss function between the output first hierarchical classification result of the enterprise data and the corresponding training labels, and optimize the parameters of the deep learning model until the mean squared error loss function converges to complete the training of the deep learning model.

6. The enterprise data grading and classification method based on a base large model and deep learning according to claim 4, characterized in that The base large model adopts any one of M3E-base, LLaMA, and GLM.

7. A method for classifying enterprise data based on a base large model and deep learning according to claim 6, characterized in that, The prompts of the base large model include: task description, background knowledge, answer examples, and answer requirements. The task description is to answer the hierarchical classification results of enterprise data based on the data item names of the enterprise data input by the user. You can refer to the answer examples and comply with the answer requirements. The background knowledge is the integrated relevant knowledge rule content and the extracted text data. The answer examples include: the data item name of the enterprise data is: xxx, the hierarchical result of the enterprise data is: xxx, the classification result of the enterprise data is: xxx, the retrieved relevant knowledge rule content is: xxx, and the source of the retrieved relevant knowledge rule is: xxx. The answer requirement is that you need to answer strictly according to the content of the background knowledge. For information you don't know, directly answer that no relevant answer is found.

8. A method for classifying and grading enterprise data based on a base large model and deep learning according to claim 1, characterized in that, When industry rules or enterprise internal rules change, the rule library and relevant corpora need to be updated in a timely manner to achieve dynamic management of enterprise data hierarchical classification.

9. A computer-readable storage medium storing a computer program, characterized in that, The computer program causes the computer to execute the enterprise data hierarchical classification method based on the base large model and deep learning as described in any one of claims 1-8.

10. An electronic device, characterized in that, Including: A memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the enterprise data hierarchical classification method based on the base large model and deep learning as described in any one of claims 1-8 is implemented.