Data classification management method and device and computing equipment

Through a multi-level classification management method, combined with a deep neural network model and multi-level label classification, the problems of insufficient data utilization and slow text classification in existing technologies are solved, and efficient and adaptive data management and fast retrieval are achieved.

CN120687611APending Publication Date: 2025-09-23HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410317584.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively mine and utilize information from large amounts of data in data management, especially in the big data era. The lack of high-quality data in enterprises leads to insufficient competitiveness, and existing text classification models cannot meet users' adaptive classification storage and fast retrieval needs.

Method used

Through multi-level classification management methods, combined with the data requirements of real scenarios, multi-level classification models and deep neural network models are adopted to retain key information, delete redundant information, and achieve end-to-end text classification management, supporting multi-level label classification and adaptive data management.

Benefits of technology

It improves the speed and accuracy of text classification, meets the classification management needs of different tasks, supports multi-level label classification management, improves data utilization efficiency and adaptability, and meets users' multi-dimensional retrieval needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687611A_ABST
    Figure CN120687611A_ABST
Patent Text Reader

Abstract

The data classification management method comprises the steps that a to-be-classified text is input into a first classification model, and the first classification model determines key information according to the to-be-classified text; obtaining semantic features according to the key information; the semantic features comprise word vectors of a plurality of key words and feature vectors of word frequency features; attribute information of the to-be-classified text is determined according to the semantic features, the attribute information comprises a first attribute, and the first attribute is a preset category. Therefore, key information highly related to semantics in the text is reserved, redundant invalid information is deleted, the semantic features of the original text can be reserved while the length of the text is shortened, and the classification speed of the long text is greatly increased on the premise that the classification precision is not reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a data classification management method, apparatus, and computing device. Background Art

[0002] On the one hand, government agencies, enterprises, hospitals, and other organizations generate massive amounts of data during their daily operations. However, this data is managed through basic categorization based on department or attribute, and the vast amount of information contained in the data is not properly mined and utilized. In the era of big data, fully mining and applying the information contained in this vast amount of data has always been a challenge in the field of natural language processing (NLP) and a research area that has been of great interest to the industry.

[0003] On the other hand, current large language models (LLMs) have demonstrated outstanding performance in natural language processing. However, training a good LLM requires massive amounts of high-quality data. Therefore, high-quality data is one of the keys to widening the gap between companies and their competitors.

[0004] Based on the above background, tasks such as data classification and data management have considerable research value. Summary of the Invention

[0005] The data classification management method, device and computing equipment provided in the embodiments of the present application, combined with the data and needs of real scenarios, propose multi-level classification of data, which can better help enterprises to perform multi-level management of file data, and at the same time provide stronger protection for data security and data compliance.

[0006] In a first aspect, an embodiment of the present application provides a data classification management method, the method comprising: inputting a text to be classified into a first classification model, the first classification model determining key information based on the text to be classified; obtaining semantic features based on the key information; the semantic features are feature vectors including word vectors of multiple key words and word frequency features; determining attribute information of the text to be classified based on the semantic features, the attribute information including a first attribute, and the first attribute being a preset category.

[0007] In this way, the embodiment of the present application retains key information in the text that is strongly related to semantics and deletes redundant invalid information. It can shorten the length of the text while retaining the semantic features of the original text, and greatly improve the classification speed of long texts without reducing the classification accuracy.

[0008] In one feasible implementation, the key information includes multiple key words, and the key information is determined based on the input text to be classified, including: segmenting the text to be classified to obtain multiple words; calculating the word frequency entropy of each word in the text to be classified, the word frequency entropy being the proportion of each word appearing in each preset category; determining one or more key words in the text to be classified based on the word frequency entropy; and determining the key information based on the one or more key words.

[0009] In this way, the embodiment of the present application can accurately screen and retain key information that is strongly related to the preset category through the word frequency entropy of the word, thereby shortening the length of the text. For example, when the word frequency entropy of the word is greater than or equal to the set parameter value, the word is determined to be a key word in the text to be classified. When the word frequency entropy of the word is less than the set parameter value, the word is determined to be a non-key word in the text to be classified. Key words are valid information in the text and participate in the classification, while non-key words are redundant information in the text that is weakly related or irrelevant to the preset category and can be deleted and does not need to participate in the classification.

[0010] In one feasible implementation, the key information includes multiple key words, and obtaining semantic features based on the key information includes: determining the word vector of each key word; determining the word frequency feature of each key word; merging the word vector and the word frequency feature to obtain the word feature of each key word; and merging the word features of each key word to obtain the semantic feature.

[0011] In this way, the embodiment of the present application obtains semantic features through key information. Since the semantic features are feature vectors composed of word vectors and word frequency features of each key word, although the length of the text input into the classification model is shortened, the semantic features of the original text are basically retained, which can greatly improve the classification speed of long texts without reducing the classification accuracy.

[0012] In one feasible implementation, determining the attribute information of the text to be classified based on the semantic feature includes: determining a first attribute of the text to be classified when the semantic feature conforms to a preset category.

[0013] In this way, the embodiment of the present application realizes end-to-end text classification management by outputting preset categories of text through the classification model, improving the text classification speed under massive data, thereby meeting the classification management requirements of basic tasks.

[0014] In one feasible embodiment, the attribute information also includes a second attribute, and the attribute information of the text to be classified is determined based on the semantic features, including: when the semantic feature vector does not conform to the preset category, the second attribute of the text to be classified is determined based on the distance between the word features of each keyword in the feature vector and the sample features of the user corpus; the second attribute is a newly defined category, and the sample features of the user corpus are obtained by extracting the features of the custom category.

[0015] In this way, for texts that do not appear in the preset categories, the embodiment of the present application calculates the distance between the word features of the key words and the sample features of the user corpus, matches the user's custom category requirements through feature aggregation, reasonably mines and infers the attribute information contained in the text, improves the adaptability and classification speed of text classification, and meets the classification management needs of different tasks.

[0016] In one feasible embodiment, the method further includes: when the confidence of the first attribute is lower than the threshold requirement, inputting the text to be classified into a second classification model; the classification accuracy of the second classification model is higher than that of the first classification model; the second classification model outputs the optimized attribute of the text to be classified, and the optimized attribute is the first attribute; the confidence of the optimized attribute meets the threshold requirement.

[0017] In this way, the embodiment of the present application also provides a second classification model with higher classification accuracy. If the attribute information output by the first classification model does not meet the confidence threshold, the second classification model with higher accuracy can be used to classify the text to be classified. The optimized attributes meet the confidence threshold requirements. The second classification model improves the classification accuracy and optimizes the classification results of the first classification model.

[0018] In one feasible implementation, the second classification model learns the semantic information of the text to be classified, and determines the optimized attributes of the text to be classified according to the semantic information of the text.

[0019] Therefore, the second classification model used in the embodiments of the present application can be a deep neural network model such as TextCNN, BERT, XLNet, or Fewshot algorithm. This type of deep neural network model uses a small number of labeled samples for training and can produce good classification results. The second classification model can fully learn the semantic information of the text, improving the classification accuracy of complex and long texts, but the speed is slow.

[0020] In one feasible embodiment, the attribute information also includes a third attribute, and the method also includes: extracting one or more hot words in the text to be classified; outputting the third attribute of the text to be classified based on the one or more hot words; the third attribute includes one or more levels of labels, and the one or more levels of labels are obtained by sorting one or more hot words according to the retrieval frequency.

[0021] Thus, the embodiment of the present application provides a function for automatically generating tags for the text to be classified, and realizes multi-level tag classification management of the text. Because the multi-level tags are obtained by sorting one or more hot words in the text based on the search frequency, this multi-level tag classification management can record and optimize the classification storage and retrieval management strategy.

[0022] In an achievable embodiment, the method further includes: receiving a search request from a user, determining a search term; searching attribute information of the text to be classified based on the search term, and outputting a search result.

[0023] In this way, the attribute information of the embodiment of the present application can help the retrieval system quickly retrieve the target file in the upper-level application. The attribute information performs classification management on the classified text in multiple dimensions. By matching the search terms and label structures, it can meet the user's data retrieval needs for files of different dimensions in different scenarios and has strong scalability.

[0024] In one feasible implementation, determining the tag structure of the text to be classified includes: saving the text to be classified; and generating the tag structure according to the first attribute, the second attribute, and / or the third attribute of the text to be classified.

[0025] In this way, the embodiment of the present application performs multi-level classification management on the classified text through the tag structure, which can simultaneously improve the applicability of the classification tags in upper-level applications and meet the user's data classification management needs for files with different hot spots in different scenarios, and has strong scalability.

[0026] In one feasible implementation, searching in the tag structure based on the search term includes: matching the tag structure according to the priority level of the search term.

[0027] In this way, since the priority of the search term is obtained based on the user's historical search records, the embodiment of the present application can optimize the search management strategy by matching the tag structure according to the priority of the search term, improve the search efficiency, and meet the needs of fast search under massive data.

[0028] In one feasible embodiment, the method further includes: obtaining labeled samples of corresponding categories according to the prompt template; the labeled samples include labeled simulated samples and labeled real samples; and saving the labeled simulated samples and / or labeled real samples as training set corpus.

[0029] In this way, the embodiment of the present application can adaptively generate sample data to meet the customer's data collection work with no or few samples, quickly provide massive, high-quality training set corpus, and save manpower.

[0030] In one feasible implementation, obtaining labeled samples of corresponding categories according to the prompt template includes: based on a generative language large model, taking the prompt template as input, and outputting labeled simulated samples.

[0031] In this way, when the user does not have real sample data, the embodiment of the present application generates labeled simulated data based on the generative language model, saving manpower for collecting data.

[0032] In one feasible implementation, obtaining labeled samples of corresponding categories according to the prompt template includes: based on a fine-tuned large language classification model, taking unlabeled real samples as input, and outputting labeled real samples.

[0033] In this way, the embodiment of the present application classifies unlabeled real corpus based on a fine-tuned large language classification model, without the need for manual data annotation, thus saving manpower for data annotation.

[0034] In the second aspect, an embodiment of the present application provides a data classification management device, which includes at least: a classification module, used to input the text to be classified into a first classification model, the first classification model determines key information based on the input text to be classified; obtains a semantic feature vector based on the key information; the semantic feature vector includes the word vector and word frequency feature of the key information; determines the first attribute of the text to be classified based on the semantic feature vector, and the first attribute is a preset category.

[0035] In one feasible embodiment, the first classification model includes a text screening module, which is used to segment the text to be classified to obtain multiple words; calculate the word frequency entropy of each word in the text to be classified; determine the key words in the text to be classified based on the word frequency entropy; the word frequency entropy is the proportion of each word appearing in each preset category; and determine key information based on multiple key words.

[0036] In one feasible embodiment, the first classification model includes a word vector extraction module, a word frequency extraction module and a feature fusion module; the word vector extraction module is used to extract the word vector of each keyword; the word frequency extraction module is used to extract the word frequency feature of each keyword; the feature fusion module is used to merge the word vector and word frequency feature to obtain the word feature of each keyword; the word feature of each keyword is merged to obtain the semantic feature of the text to be classified.

[0037] In one feasible implementation, when the semantic feature vector conforms to a preset category, the first classification model determines a first attribute of the text to be classified.

[0038] In one feasible embodiment, when the semantic feature vector does not conform to the preset category, the first classification model determines the second attribute of the text to be classified based on the distance between the word features of each keyword in the feature vector and the sample features of the user corpus; the sample features of the user corpus are obtained by extracting the features of the custom category.

[0039] In one feasible embodiment, the classification module also includes a second classification model; the classification accuracy of the second classification model is higher than that of the first classification model; when the confidence of the first attribute is lower than the threshold requirement, the second classification model outputs an optimized attribute based on the text to be classified, and the optimized attribute is a preset category; the confidence of the optimized attribute meets the threshold requirement.

[0040] In one feasible implementation, the second classification model learns the semantic information of the text to be classified, and determines the optimized attributes of the text to be classified according to the semantic information of the text.

[0041] In one feasible embodiment, the classification module also includes a label extraction model; the label extraction model is used to extract one or more hot words in the text to be classified; the third attribute of the text to be classified is output based on the one or more hot words; the third attribute includes one or more levels of labels, and the one or more levels of labels are obtained by sorting one or more hot words according to the retrieval frequency.

[0042] In one feasible embodiment, the device also includes a retrieval module; the retrieval module is used to receive a user's retrieval request and determine a search term; based on the search term, a search is performed in a tag structure of the text to be classified, and a retrieval result is output; the tag structure includes a first attribute, a second attribute and / or a third attribute of the text to be classified.

[0043] In one feasible implementation, the search module matches the tag structure according to the priority level of the search term.

[0044] In a feasible implementation, the retrieval module is further configured to save the text to be classified; and generate a tag structure according to the first attribute, the second attribute and / or the third attribute of the text to be classified.

[0045] In one feasible embodiment, the device also includes a data generation module; the data generation module is used to obtain labeled samples of corresponding categories according to the prompt template; the labeled samples include labeled simulated samples and labeled real samples; and the labeled simulated samples and / or labeled real samples are saved as training set corpus.

[0046] In one feasible embodiment, the data generation module includes a generative language model, which is used to take a prompt template as input and output a labeled simulated sample.

[0047] In one feasible implementation, the data generation module includes a fine-tuned large language classification model, where the fine-tuned large language classification model is configured to take unlabeled real samples as input and output labeled real samples.

[0048] In a third aspect, embodiments of the present application provide a computing device comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the programs stored in the memory are executed, the processor is configured to perform any of the methods described in the first aspect. The beneficial effects thereof can be found in the description of the first aspect and are not further elaborated here.

[0049] In a fourth aspect, embodiments of the present application provide a computing device cluster, comprising at least one computing device, each computing device including a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, thereby causing the computing device cluster to perform any of the methods described in the first aspect. The beneficial effects thereof can be found in the description of the first aspect and are not further elaborated here.

[0050] In a fifth aspect, an embodiment of the present application provides a program product that runs computer program instructions to perform the method provided in the first aspect. The beneficial effects thereof can be referred to the relevant description of the first aspect and will not be repeated here.

[0051] In a sixth aspect, embodiments of the present application provide a computer storage medium having instructions stored therein. When the instructions are executed on a computer, the computer executes the method provided in the first aspect. The beneficial effects thereof can be found in the description of the first aspect and are not further elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings described below are only the multiple embodiments disclosed in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0053] The following is a brief introduction to the drawings required for describing the embodiments or prior art.

[0054] Figure 1 This is a system architecture diagram of data classification management proposed in an embodiment of the present application;

[0055] Figure 2 Schematic diagram of a data classification management device provided in an embodiment of the present application;

[0056] Figure 3is a schematic diagram of a classification module in a data classification management device;

[0057] Figure 4 Schematic diagram of the semantic fusion classification model in the multi-level classification module;

[0058] Figure 5 A schematic diagram of a retrieval module in a data classification management device;

[0059] Figure 6 A schematic diagram of an analysis module in a data classification management device;

[0060] Figure 7 A schematic diagram of a data generation module in a data classification management device;

[0061] Figure 8 A flowchart of the data classification management method provided in an embodiment of the present application;

[0062] Figure 9 A flowchart of a data classification management method provided in Example 1 of the present application;

[0063] Figure 10 A flowchart of data generation in the data classification management method provided in Example 1 of the present application;

[0064] Figure 11a Schematic diagram of the settings interface in the client device provided in Example 2 of the present application;

[0065] Figure 11b Schematic diagram of the label generation interface provided in Example 2 of this application;

[0066] Figure 12 A flow chart of a data classification management method provided in Example 3 of the present application;

[0067] Figure 13 A schematic diagram of a computing device provided in an embodiment of the present application;

[0068] Figure 14 A schematic diagram of a computing device cluster is also provided for the embodiment of the present application;

[0069] Figure 15 A possible connection method. DETAILED DESCRIPTION

[0070] In the description of the embodiments of the present application, words such as "exemplary," "for example," or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary," "for example," or "for example" in the embodiments of the present application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary," "for example," or "for example" is intended to present the relevant concepts in a concrete manner.

[0071] In the description of the embodiments of this application, the term "and / or" is simply a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent the following three situations: A exists alone, B exists alone, and A and B exist at the same time. In addition, unless otherwise specified, the term "plurality" means two or more. For example, "multiple systems" refers to two or more systems, and "multiple terminals" refers to two or more terminals.

[0072] In the description of the embodiments of the present application, the terms "first\second\third, etc." or module A, module B, module C, etc. are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that the specific order or sequence can be interchanged where permitted so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0073] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly identifying the technical features being referred to. Thus, features specified as "first" or "second" may explicitly or implicitly include one or more of such features. The terms "include," "comprising," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.

[0074] In the description of the embodiments of the present application, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it can be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict.

[0075] In the description of the embodiments of the present application, the numbers representing the steps, such as S110, S120, etc., do not necessarily mean that the steps must be executed in this manner. If permitted, the order of the previous and next steps can be interchanged, or they can be executed simultaneously.

[0076] Related terms involved in the embodiments of this application:

[0077] Prompt is a piece of text or a set of vectors added to the input, allowing the model to perform masked language modeling (MLM) based on the input and the added prompt.

[0078] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0079] With the development of natural language processing technology, text classification, as one of the most basic research in the field of natural language processing, has a very important research value in recent years for tasks such as multi-level classification management of files.

[0080] Text classification methods include statistical and neural network-based approaches. Traditional statistical methods primarily use term frequency-inverse document frequency (TF-IDF) features to represent text and perform classification using methods such as Naive Bayes or Support Vector Machines. In recent years, with the development of deep learning, neural network-based methods have gradually become mainstream. Common neural network methods include convolutional neural networks (CNNs), recurrent neural networks (RNNs), and Transformer-based methods.

[0081] In addition to general text classification, some work specializes in text classification for specific scenarios. For example, sentiment recognition and patent classification methods based on patent document abstracts are examples. These methods focus on classifying text content in specific scenarios and have strong industry relevance, but they also essentially use mainstream statistical and neural network methods. Furthermore, some work utilizes generative neural networks to construct text data.

[0082] There are some problems with the text classification methods used in the above-mentioned text classification work. First, there is a lack of classification data corresponding to the text classification task. Even if researchers have collected data for the corresponding task, labeling the categories of this data requires a lot of manpower, which brings great difficulty to the implementation of actual document classification tasks. Existing technologies often work based on existing labeled samples and rarely involve strategies for constructing labeled samples. In some specific scenarios, users often have special classification needs. Existing classification models cannot meet the training requirements for specific scenarios, and new categories of data sets must be collected to retrain the classification model.

[0083] Secondly, in actual data storage and management, text classification requirements are not static and often vary depending on the specific task. In current applications, text classification categories are generally fixed, which cannot meet users' needs for adaptive classification storage. In some retrieval scenarios, full-text retrieval is too slow, and the number of category labels for stored files is too small, making it impossible to quickly retrieve the target file. Even if files are classified using multiple labels, the number of category labels is still limited, and when there are too many category labels, retrieval efficiency is still too low.

[0084] An embodiment of the present application provides a data classification management system. In addition to outputting the basic categories of samples in a semantic fusion classification model, it also combines user-defined semantic features, matches the user's new category requirements through feature aggregation, and rationally mines and applies the information contained in the data, thereby realizing end-to-end, scalable, and multi-level text label classification management. In the case of massive data, it improves the adaptability and classification speed of text classification, thereby meeting the classification management needs of different tasks.

[0085] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.

[0086] Figure 1 The following is a diagram showing a system architecture for data classification management proposed in an embodiment of the present application. Figure 1 As shown in the system diagram 100, the data generation device 16 is used to collect data and store it in the database 13 as training set samples. The training device 12 trains the classification module 111 in the data classification management device 11 based on the training set samples and the labeled sample data stored in the database 13. The classification module 111 inputs the text to be classified into a first classification model. The first classification model determines key information based on the input text to be classified; obtains a semantic feature vector based on the key information; the semantic feature vector includes the word vector and word frequency features of the key information; and determines the attribute information of the text to be classified based on the semantic feature vector. The attribute information includes preset categories, semantic features, etc.

[0087] The following will describe in more detail how the training device 12 trains the classification module 111 based on the labeled sample data stored in the database 13.

[0088] Because it is desired that the text category information output by the classification module 111 be as consistent as possible with the actual attributes, the attribute information output by the current classification module 111 can be compared with the confidence level of the attributes of the actual corpus, and the weight vectors of each layer of the network can be updated based on the difference between the two. Of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the network. For example, if the attribute confidence value output by the network is high, the weight vector is adjusted to make it output lower, and the adjustment is continued until the network can output the attributes of the actual corpus. The member classification module 111 obtained by the training device 22 can be applied to different systems or devices.

[0089] The retrieval module 112 is used to store the text to be categorized and generate a tag structure based on the first, second, and / or third attributes of the text to be categorized. It receives a user's search request, determines a search term, searches the tag structure of the text to be categorized based on the search term, and outputs the search results. The first attribute is a preset category, the second attribute is a newly defined category, and the third attribute is a multi-level tag containing multiple hot words.

[0090] The data classification management device 11 can call user-defined data, codes, etc. stored in the data storage system 15, and can also store data, text, text attribute information and / or tag structure, etc. in the data storage system 15.

[0091] exist Figure 1 In the example, both the retrieval module 112 and the analysis module 113 are configured with application layer interfaces for data exchange with the external client device 14. Users can input their required data through the application layer interface of the client device 14. This required data can be pre-set category data, text to be classified, or search terms. The analysis module 13 determines the user's classification requirements and historical search records. The system can optimize the classification storage and retrieval management strategy based on the user's requirements and historical search records. The retrieval module 112 returns the classification / retrieval results to the client device 14 and provides them to the user.

[0092] In the attached Figure 1 In the case shown in , the user can manually specify the custom category data input into the data classification management device 11, for example, by operating in the interface provided by the client device 14. In another case, the client device 14 can automatically input preset category data into the application layer interface and obtain the results. If the client device 14 needs to obtain member authorization to automatically input preset category data, the user can set the corresponding permissions in the client device 14. The user can view the results output by the data classification management device 11 on the client device 14, and the specific presentation form can be specific methods such as display, text, and search directory. The data classification management device 11 can also serve as a data collection terminal to store the collected data in the database 13.

[0093] It is worth noting that Figure 1 This is only a schematic diagram of a system architecture provided by the embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in the attached Figure 1 In the embodiment, the data storage system 15 is an external memory relative to the data classification management device 11. In other cases, the data storage system 15 can also be placed in the data classification management device 11. The data classification management device 11, as an execution device, can be any device, equipment, platform, or device cluster with computing and processing capabilities.

[0094] Similarly, the data generation device 16 is an external device relative to the data classification management apparatus 11 and can be any device, equipment, platform, or device cluster with computing and processing capabilities. In other cases, the data generation device 16 can also be placed in the data classification management apparatus 11 as a data generation module.

[0095] Figure 2 Schematic diagram of the data classification management device provided in the embodiment of this application. Figure 2 As shown, the data classification management device includes a classification module 111, a retrieval module 112, an analysis module 113 and a data generation module 114. Among them, the classification module 111 inputs the text to be classified into the first classification model, and the first classification model determines the key information based on the input text to be classified; obtains a semantic feature vector based on the key information; the semantic feature vector includes the word vector and word frequency feature of the key information; determines the attribute information of the text to be classified based on the semantic feature vector, and the attribute information includes a first attribute, which is a preset category. The retrieval module 112 receives the user's search request and determines the search term; searches the tag structure of the text to be classified based on the search term and outputs the search result; the tag structure includes the first attribute of the text to be classified. The analysis module 113 determines the sample category according to the user's classification needs and presets the prompt template; and counts the search term records of the user's search documents according to a fixed period to determine the search frequency of different search terms. The data generation module 114 obtains labeled samples of the corresponding category based on the generative large model according to the prompt template; the labeled samples include labeled simulated samples and labeled real samples, and the labeled simulated samples and / or labeled real samples are used as training set corpus; the training set corpus is used to train the classification module 111.

[0096] The data management device provided in the embodiment of the present application is further described below from the aspects of the classification module 111, the retrieval module 112, the analysis module 113 and the data generation module 114.

[0097] Figure 3 Figure 1 is a schematic diagram of a classification module in a data classification management device. Figure 3As shown, the classification module 111 includes a text screening module 21, a semantic fusion classification model 22 and a small sample learning model 23; the semantic fusion classification model 22 can be recorded as the first classification model, and the small sample learning model 23 can be recorded as the second classification model.

[0098] The text screening module 21 performs word segmentation on the text to be classified to obtain multiple words; calculates the word frequency entropy of each word in the text to be classified; determines the key words in the text to be classified based on the word frequency entropy; the word frequency entropy is the proportion of each word appearing in each preset category.

[0099] The semantic fusion classification model 22 determines key words based on the input text to be classified; obtains the word features of each key word, merges the word features of each key information to obtain semantic features; outputs the first attribute and / or second attribute of the text to be classified based on the semantic features; the first attribute is a preset category, and the second attribute is a text feature.

[0100] For example, Figure 4 This is a schematic diagram of the semantic fusion classification model in the multi-level classification module. Figure 4 As shown, the semantic fusion classification model 22 includes a word vector extraction module 221, a word frequency extraction module 222 and a feature fusion module 223.

[0101] Among them, the word vector extraction module 221 extracts the word vector of each keyword; the word frequency extraction module 222 extracts the word frequency feature of each keyword; the feature fusion module 223 merges the word vector and the word frequency feature to obtain the word feature of each keyword; the word feature of each keyword is spliced ​​to obtain the semantic feature of the text to be classified.

[0102] In some possible implementations, word vector extraction module 221 includes a FastText model, and word frequency extraction module 222 includes a Naive Bayes model. Shallow neural network models such as the FastText model and the Naive Bayes model are fast and work well for most simple samples. When the confidence level of the output results of such models is high, the text attribute information can be directly output based on the output results, and the text category can be determined based on the text attribute information.

[0103] The semantic fusion classification model 22 outputs attribute information of the text to be classified based on the semantic features.

[0104] In some feasible implementations, when the semantic feature conforms to a preset category, the semantic fusion classification model 22 outputs a first attribute of the text to be classified according to the semantic feature.

[0105] In some possible implementations, when the semantic features do not conform to a preset category, the semantic fusion classification model 22 outputs a feature vector including word features of each key word.

[0106] In some possible implementations, the text to be classified also includes text that does not appear in the preset categories. For text that does not appear in the preset categories, the classification module 111 calculates the distance between the word features of each key word in the feature vector output by the semantic fusion classification model 22 and the sample features of the user corpus; and determines the second attribute of the text to be classified based on the distance. The sample features of the user corpus are obtained by extracting the features of the custom category under the specified path, and the second attribute is the newly defined category.

[0107] When the confidence of the first attribute output by the semantic fusion classification model 22 is lower than the set probability threshold, the small sample learning model 23 can be used to output optimized attributes based on the text to be classified, and the optimized attributes are preset categories; the confidence of the optimized attributes meets the threshold requirements.

[0108] Confidence is also called reliability, or confidence level, or confidence coefficient. It uses a probability statement method to indicate the corresponding probability that the estimated value and the population parameter are within a certain allowable error range. This corresponding probability is called confidence.

[0109] In some feasible implementations, when the confidence of the first attribute or semantic feature is lower than the threshold requirement, the small sample learning model 23 learns the semantic information of the text to be classified based on the input text to be classified, and outputs the optimized attribute of the text to be classified based on the semantic information of the text; the confidence of the optimized attribute meets the threshold requirement.

[0110] Small-sample learning models23 include deep neural network models such as TextCNN, BERT, XLNet, or the Fewshot algorithm. These deep neural network models are trained using a small number of labeled samples and can achieve good classification results, fully learning the semantic information of the text. They offer high classification accuracy but are slow, making them suitable for classifying long texts with complex semantics.

[0111] For example, a Fewshot algorithm trained with a small amount of training data is fed into the text to be classified. The algorithm learns the semantic information of the text, classifies it based on the semantic information, and outputs the optimized first attribute of the text to be classified. The Fewshot algorithm can produce good classification results, and the confidence level of its output results can meet the set threshold requirements.

[0112] The classification module 111 also includes a label extraction model 24. The label extraction model 24 extracts one or more hot words from the text to be classified and outputs a third attribute of the text to be classified based on the one or more hot words. The third attribute includes one or more hierarchical labels, which are obtained by sorting the one or more hot words according to search frequency. Hot words are search words with a search frequency higher than a set threshold.

[0113] The classification module 111 can be embedded in any storage hardware platform and can also be applied to any data management application software.

[0114] The above is an introduction to the application of the classification module 111. Next, the retrieval module 112 will be introduced.

[0115] Figure 5 Figure 1 is a schematic diagram of a retrieval module in a data classification management device. Figure 5 As shown, the retrieval module 112 includes a management module 1121 and a query module 1122 .

[0116] The management module 1121 saves the tag structure of the text to be classified in a specified path. The tag structure includes multiple levels of tags such as the first attribute, the second attribute, the third attribute and / or the third attribute of the text to be classified.

[0117] The query module 1122 searches the multi-level tag structure under the specified path in the search task and outputs the search results.

[0118] Exemplarily, the query module 1122 searches the multi-level tag structure under the specified path according to the search term input by the user, and outputs the searched text.

[0119] In some implementations, the text to be classified is stored hierarchically in a corresponding text attribute list according to the tag structure.

[0120] In some implementations, the query module may determine priority levels based on the search frequencies of different search terms, and match multi-level tag structures according to the priority levels.

[0121] The above is an introduction to the application of the search module 112. In some specific scenarios, users often have special classification requirements and data management needs. Based on this, the embodiment of the present application provides an analysis module 113, which will be described below.

[0122] Figure 6 Figure 1 is a schematic diagram of the analysis module in the data classification management device. Figure 6As shown, the analysis module 113 includes a demand analysis module 1131 and a retrieval analysis module 1132. The demand analysis module 1131 determines the sample category according to the user's classification needs and sets a prompt template according to the sample category. The prompt template is used to guide the generative language model to generate labeled simulation samples.

[0123] In the data classification management system, the demand analysis module 1131 can obtain the sample category set by the user according to the classification needs through the user-oriented application layer interface, set a new prompt template according to the sample category or call the prompt template of the existing preset category, and define the corresponding save path of the generated data.

[0124] Sample categories include basic categories and custom categories. When the classification needs to be based on the basic categories of the text, a prompt template for the basic categories is set. When the classification needs to be based on the custom categories of the text, a prompt template for the custom categories is set.

[0125] Basic categories are used to meet the most basic data classification needs, including the classification of common data such as government and enterprise, hospitals, markets, human resources, sound fields, and processing. Prompt templates for basic categories include prompt templates for common data such as government and enterprise, hospitals, markets, human resources, sound fields, and processing. Basic categories can be pre-set and recorded as preset categories.

[0126] Custom categories are used to reclassify a specific scenario, including emotion recognition and classification based on document abstracts. Custom category prompt templates include emotion recognition prompt templates and document abstract-based prompt templates.

[0127] The search analysis module 1132 counts the search term records of users searching for documents at a fixed period, determines the search frequency of different search terms, and determines hot words based on the search frequency.

[0128] In the data classification management device provided in the present application, the analysis module 113 determines the user's classification needs and historical retrieval records, and the system can optimize the classification storage and retrieval management strategy based on the user's needs and historical retrieval records.

[0129] In some application scenarios, the existing classification model cannot meet the new classification requirements, and the classification model must be retrained by collecting new datasets. The data generation module 114 realizes the adaptive generation of classification data based on user needs and the large language model. The data generation module 114 is described below.

[0130] Figure 7 This is a schematic diagram of the data generation module in the data classification management device. Figure 7As shown, the data generation module 114 includes a generative language large model 141 and / or a large language classification model 142, wherein the generative language large model 141 generates labeled simulated samples according to the input prompt template, and uses the labeled simulated samples to train a fine-tuned large language classification model 142. The fine-tuned large language classification model 142 infers unlabeled real samples to obtain labeled real samples, and saves the labeled simulated samples and / or labeled real samples in a set save path / database 13 to obtain a training set corpus.

[0131] For example, generative language model 141 can be an open-source generative model such as GPT4 or CHATGLM. Generative language model 141 takes prompt templates as input and outputs labeled simulation samples. These labeled simulation samples are used to train and fine-tune large language classification model 142.

[0132] The large language classification model 142 may be a LAMA large model, which takes unlabeled real text as input and outputs labeled real samples.

[0133] When users put forward new text classification requirements, the data generation module 114 can adaptively meet the customer's data collection work with no or few samples, quickly provide massive, high-quality training set corpus to train the classification model, and save manpower.

[0134] The data generation module 114 is based on the technology of the large language model and combines the data of real scenarios with user needs to construct a training set corpus. The training set corpus includes labeled simulated samples and / or labeled real samples. The text screening module 21, semantic fusion classification model 22, small sample learning model 23 and label extraction model 24 trained using the labeled real samples in the training set corpus are more stable in actual tests.

[0135] In some application scenarios, the data generation module 114 can be set in other devices. The device is an external device relative to the data classification management device 11, and can be any device, equipment, platform, or device cluster with computing and processing capabilities, such as Figure 1 The data generating device 16 in.

[0136] The above is an introduction to the data generation module 114. The following describes the training process of each classification model in the classification module 111 based on the training set corpus.

[0137] The text screening module 21, semantic fusion classification model 22, and small sample learning model 23 in the classification module 111 are trained based on the training set corpus; the training goal is to ensure that the attributes and semantic features of the output sample are consistent with the target attributes and semantic features corresponding to the label of the sample.

[0138] In the training state, the text screening module 21 performs word segmentation on the labeled real samples in the training set corpus to obtain multiple words. The word segmentation can adopt any word segmentation method, such as Jieba word segmentation, etc. Based on the word segmentation results of the sample, the number of times n that each word i in the sample appears in category c is counted. c,i , the total number of words in each category of corpus N c , based on which we can calculate the proportion p of word i in each category c c,i , and then calculate the word frequency entropy E of word i according to the proportion of appearance in different categories c,i :

[0139]

[0140] Among them, category c is the label of the real sample, which is used to indicate the basic category of the text; based on the word frequency entropy E of each word c,i , words are divided into key words or non-key words, where key words participate in classification training and non-key words are discarded and do not participate in classification training.

[0141] The word vector extraction module 221 extracts the word vector of each keyword in the labeled real sample, the word frequency extraction module 222 extracts the word frequency feature of each keyword, and the feature fusion module 223 merges the word vector and the word frequency feature to obtain the word feature of each keyword in the labeled real sample, and merges the word features of each keyword to obtain the semantic feature of the labeled real sample. The basic category and semantic feature are output according to the semantic feature, and the trained semantic fusion classification model 22 is obtained when the text category and semantic feature output by the semantic fusion classification model 22 are consistent with the target attribute and semantic feature corresponding to the label of the sample.

[0142] The small sample learning model 23 is trained using a small number of real labeled samples to learn the semantic information of the labeled samples and output attribute information of the samples based on the semantic information. When the output attribute information matches the sample label, the trained small sample learning model 23 is obtained. The small sample learning model 23 can produce good classification results when trained using a small number of real labeled samples.

[0143] The training state also includes the training of the label extraction model 24. The label extraction model 24 extracts hot words from the labeled real samples, hierarchizes the hot words according to the search frequency, and obtains multi-level labels for the labeled real samples.

[0144] The tag extraction model 24 includes a hot word extraction model 231 and a tag stratification module 232 .

[0145] The training of the hot word extraction model 231 uses labeled samples, and the input is each keyword of the labeled sample. The training goal is to output multiple hot words in the sample. Hot words are keywords with a search frequency higher than a certain threshold.

[0146] The label stratification module 232 counts the search frequencies of different hot words, and stratifies the multiple hot words according to the search frequencies to obtain one or more hierarchical labels.

[0147] In some application scenarios, the training state of each classification model in the classification module 111 can be executed in the data classification management device 11 or in Figure 1 The training device 12 is shown as being executed, which is an external device relative to the data classification management device 11. The training device 12 can be any device, equipment, platform, or device cluster with computing and processing capabilities.

[0148] In the data classification management device 11 provided in the embodiment of the present application, the retrieval module 112 and the analysis module 113 are both provided with a user-oriented application layer interface, which can meet the end-to-end operation of "demand analysis-data collection and processing-data classification-application".

[0149] As an example of a software functional unit, the classification module 111 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the classification module 111 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.

[0150] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0151] As an example of a hardware functional unit, the classification module 111 may include at least one computing device, such as a server. Alternatively, the classification module 111 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0152] The multiple computing devices included in the classification module 111 can be distributed in the same region or in different regions. The multiple computing devices included in the classification module 111 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the classification module 111 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.

[0153] The above is an introduction to the functions and training scenarios of the data management device provided in the embodiment of the present application. Next, based on the above content, the data classification management method provided in the embodiment of the present application is introduced. Among them, the relevant introductions to the concepts, modules, units, etc. involved in the following content can be found above.

[0154] The data classification management method provided in the embodiment of the present application can be deployed in a server or a unified server cluster. The above-mentioned server generally refers to a general-purpose computer system installed with a mainstream operating system (Unix, Windows, etc.).

[0155] Figure 8 This is a flowchart of the data classification management method provided in the embodiment of the present application. It is understood that the method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. In addition, some or all steps of the data classification management method can refer to the aforementioned Figure 1-Figure 7 Related description in . Figure 8 As shown, the data classification management method includes the following steps:

[0156] S81: Input the text to be classified into the classification module 111.

[0157] Among them, the classification module 111 is a model combination trained based on the training set corpus, including a text screening module 21, a semantic fusion classification model 22 and a small sample learning model 23.

[0158] S82: Determine key information based on the text to be classified.

[0159] In some feasible implementations, S82 may determine the key words through the following steps S821 - S823 .

[0160] S821, performing word segmentation on the text to be classified by the text screening module 21 to obtain each word in the text.

[0161] S822, calculating the word frequency entropy of each word in the text to be classified; the word frequency entropy is the proportion of each word appearing in each preset category.

[0162] S823, determining key words in the text to be classified based on word frequency entropy, and determining key information based on the key words.

[0163] The words in the text to be classified include key words and non-key words. Key words are valid information in the text, and non-key words are redundant and invalid information in the text.

[0164] In some feasible implementations, a parameter value can be set. When the word frequency entropy of a word is greater than or equal to the set parameter value, the word is determined to be a key word in the text to be classified. When the word frequency entropy of the word is less than the set parameter value, the word is determined to be a non-key word in the text to be classified. The non-key word is deleted from the text to be tested, thereby determining the key information of the text to be classified.

[0165] Through step S823, only key words are retained for classification, and non-key words are discarded and do not participate in classification. In this way, key information in the text to be classified is screened out and redundant invalid information is deleted. The classification speed of long texts can be greatly improved without reducing the classification accuracy.

[0166] In some feasible implementations, the key information of the text to be classified also includes numbers, titles, punctuation marks, symbols, etc.

[0167] S83, obtaining semantic features based on the key information; the semantic features are feature vectors including word vectors of multiple key words and word frequency features.

[0168] The word features of each key word are obtained through the semantic fusion classification model 22, and the word features of each key word are merged to obtain the semantic features. The semantic fusion classification model 22 includes shallow neural network models such as the fasttext model and the naive Bayes model.

[0169] In some feasible implementations, semantic features may be obtained through the following steps S831 - S834 .

[0170] S831, extract the word vector of each keyword.

[0171] For example, the fasttext model can be used to extract the word vector of each keyword.

[0172] S832: Extract the frequency feature of each keyword.

[0173] For example, a naive Bayes model may be used to extract the word frequency feature of each keyword.

[0174] S833: Merge the word vector and the word frequency feature to obtain the word feature of each key word.

[0175] S834: Merge the word features of each keyword to obtain the semantic features of the text to be classified.

[0176] In some feasible implementations, the word features of each key word are input into a trained feature fusion module, and the semantic features of each text to be classified are output.

[0177] S84: Determine attribute information of the text to be classified based on the semantic features.

[0178] In some feasible implementations, the attribute information includes at least a preset category. When the semantic feature meets the feature of the preset category, the semantic fusion classification model 22 outputs the first attribute of the text to be classified according to the semantic feature, and the first attribute is the preset category.

[0179] When the semantic feature vector does not conform to the preset category, the second attribute of the text to be classified is determined based on the distance between the word feature of each keyword in the feature vector and the sample feature of the user corpus; the sample feature of the user corpus is obtained by extracting the features of the custom category, and the second attribute is the newly defined category.

[0180] In some feasible implementations, the data classification management method also includes extracting one or more hot words from the text to be classified, and outputting a third attribute of the text to be classified based on the one or more hot words; the third attribute includes one or more hierarchical labels, and the one or more hierarchical labels are obtained by sorting one or more hot words according to retrieval frequency.

[0181] In some feasible implementations, the data classification management method further includes saving a multi-level tag structure of the text to be classified in a specified path, where the multi-level tag structure includes a first attribute, a second attribute, and / or a third attribute of the text to be classified.

[0182] In some feasible implementations, the data classification management method further includes, in a retrieval task, performing a search based on a multi-level label structure under the specified path and outputting the search results; the multi-level label structure includes the first attribute, the second attribute and / or the third attribute of the text to be classified.

[0183] It is understandable that the size of the serial number of each step in the above-mentioned embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. In addition, in some possible implementations, the steps in the above-mentioned embodiments can be selectively executed according to actual conditions, and can be partially executed or fully executed, which is not limited here. In addition, all or part of any features of any of the above-mentioned embodiments can be freely and arbitrarily combined without contradiction; the combined technical solutions are also within the scope of this application.

[0184] The following further introduces the above data classification management method through some specific application scenarios.

[0185] Example 1

[0186] Embodiment 1 of the present application provides a data classification management method that uses a combination of multiple classification models to determine the attribute information of a text and collects training set corpus to train multiple classification models based on the user's new classification requirements. Embodiment 1 of the present application can quickly achieve multi-level classification of text while maintaining the same classification accuracy.

[0187] Figure 9 A data classification management method is provided in Example 1 of this application. It can be understood that the method can be executed by any device, equipment, platform, or equipment cluster with computing and processing capabilities. Figure 9 As shown, the data classification management method includes the following steps:

[0188] S901, in the semantic fusion classification model 22, the jieba segmentation is used to segment the text to be classified to obtain each word i.

[0189] S902: Count the number of times n each word i appears in category c. c,i , the total number of words in each category of corpus N c , we can calculate the proportion p of word i in each category based on this c,i , and then calculate the word frequency entropy E of word i according to the proportion of occurrence in different categories c,i , the calculation formula is as follows:

[0190]

[0191] S903: Determine key information in the text to be classified based on word frequency entropy.

[0192] In Example 1 of the present application, a parameter value can be set. When the word frequency entropy of a word is greater than or equal to the set parameter value, the word is determined to be a key word in the text to be classified. When the word frequency entropy of the word is less than the set parameter value, the word is determined to be a non-key word in the text to be classified, and the non-key word is deleted from the text to be tested.

[0193] S904: Use the fasttext model to extract the word vector of each keyword.

[0194] S905: Use the naive Bayes model to extract the word frequency feature of each keyword.

[0195] S906: Merge the word vector and the word frequency feature to obtain the word feature of each key word.

[0196] S907: Input the word features of each key word into the trained feature fusion module, and output the semantic features of each text to be classified.

[0197] S908 , when the semantic feature meets the first attribute, output the first attribute of the text to be classified according to the semantic feature, and execute S909 , where the first attribute is a preset category.

[0198] In the case where the semantic feature does not conform to the first attribute, a feature vector including word features of each key word is output, and S912 is executed.

[0199] S909: Determine the confidence level of the output result. If the confidence level is lower than the threshold, execute S910. If the confidence level is higher than or equal to the threshold, execute S911.

[0200] S910: Input the text to be classified into the small sample learning model 23. The small sample learning model 23 learns the semantic information of the text to be classified and outputs the optimized attributes of the text to be classified based on the semantic information of the text; the confidence of the optimized attributes meets the threshold requirement; and execute S911.

[0201] S911: directly determine the first attribute of the text according to the output result.

[0202] S912 , a distance is calculated between the word features of each key word in the feature vector output by the semantic fusion classification model 22 and the sample features of the user corpus; a second attribute of the text to be classified is determined based on the distance, where the second attribute is the newly defined category. The sample features of the user corpus are obtained by extracting the features of the custom category under the specified path.

[0203] In the method provided in Example 1 of this application, valid information in the text to be classified is screened based on word frequency entropy, and redundant invalid information is deleted. This can significantly improve the classification speed of long texts without reducing classification accuracy, thus achieving user-adaptive rapid text classification based on word frequency entropy. For user-defined text that does not appear in the preset categories, user-defined multi-level and scalable rapid text classification is supported.

[0204] In some specific scenarios, users often have special classification requirements. The above existing models cannot meet the new classification requirements, and the training set corpus must be re-collected to retrain the classification model. Therefore, the data classification management method provided in Example 1 of this application also provides the following data generation steps for the user's new classification requirements.

[0205] Figure 10 This is a flow chart of data generation provided in Example 1 of this application. Figure 10 As shown, when the user does not have real sample data, based on the large language model, a large amount of unlabeled text is used to generate labeled samples as the training set corpus of the model. When the user does not have real sample data, data generation includes steps S101-S105.

[0206] S101, a prompt template of a category needs to be set according to the user's classification. The prompt template is used to guide the generative language model 141 to generate labeled simulation samples.

[0207] In some feasible implementations, the sample category set by the user according to the classification needs can be obtained through the user-oriented application layer interface, a new prompt template can be set according to the sample category or a prompt template of an existing preset category can be called, and the corresponding save path of the generated sample data can be defined.

[0208] S102 , using the generative language model 141 to generate labeled simulation samples according to the prompt template of the category.

[0209] The generative language big model 141 includes any open source generative big model or language big model API. In some feasible implementations, a certain open source generative big model or language big model API can be selected according to the hardware / platform configuration. The language big model API includes GPT4 or CHATGLM.

[0210] The generative language model 141 takes the category prompt template as input and outputs labeled simulated samples.

[0211] For example, the hospital general data prompt template is input into GPT4 or CHATGLM to generate a simulated sample with the label "hospital"; the emotion recognition prompt template is input into GPT4 or CHATGLM to generate a simulated sample with the label "emotion recognition".

[0212] In the case where a large number of unlabeled real samples are provided by the user, the data generation method includes steps S103-S105.

[0213] S103: Use the labeled simulated samples generated in step S102 to train and fine-tune the large language classification model 142 to obtain a trained large language classification model 142. The large language classification model 142 is used to classify unlabeled real samples.

[0214] S104: A large number of unlabeled real samples are input into the large language classification model 142, and labeled real samples are output, where the labels indicate the classification of the real samples.

[0215] S105: Save the labeled simulated samples and / or labeled real samples in a set storage path / database to obtain a training set corpus.

[0216] Next, the text screening module 21, the semantic fusion classification model 22, the small sample learning model 23 and the label extraction model 24 can be trained using the training set corpus. The semantic fusion classification model 22, the small sample learning model 23 and the label extraction model 24 trained using the real samples with labels in the training set corpus are more stable in actual tests. The specific training process and steps in Example 1 of this application can be referred to above. Figure 7 The relevant description will not be repeated here.

[0217] Example 2

[0218] In the application scenario given in Example 2, the data generation process is Figure 1 The data generation device 16 is shown as executing the data classification management apparatus 11. This apparatus is external to the data classification management apparatus 11 in the embodiment. The data generation device 16 can be any device, equipment, platform, or device cluster with computing and processing capabilities. The data generation device 16 is provided with an interface with the client-facing device 14 for data exchange.

[0219] In text classification, the following strategies are usually adopted to collect samples according to different scenarios: (1) Use labeled samples when there are labeled samples; (2) Use manually labeled samples when there are no labeled samples but manual labeling is easier; (3) Use generative large models to generate labeled simulated samples when there are no labeled samples and manual labeling is time-consuming.

[0220] Example 2 of the present application proposes a data generation method. When a user does not have available training corpus, data generated by a generative language model is used to form training data for text classification. If the user has unlabeled corpus, the generated data can be used to train a fine-tuned large language classification model 142, such as a LAMA model. The fine-tuned large language classification model 142 can then be used to label the unlabeled corpus to obtain labeled real data.

[0221] like Figure 11a As shown in the diagram of the interface setting in the client device 14, the user can set the corresponding prompt template through the application layer interface of the "Data Generation Settings" interface on the client device 14, define the corresponding generated data storage path, call the generative language model 141, and directly generate the corresponding data. Figure 11a As shown, on the "Data Generation Settings" interface, set multiple operation parameters and operation prompts including prompt template, save path, generation tool, generation data list, etc.

[0222] For example, the operation prompt is "Please configure the prompt template." The user can click the operation tab to determine the sample category to be set, set a new prompt template based on the sample category or call an existing preset prompt template, and define the corresponding save path for the generated sample data.

[0223] When the user does not have available training corpus, the specific steps of the process of using the generative language model to generate data are as follows:

[0224] S1001, a prompt template of a category needs to be set according to the user's classification. The prompt template is used to guide the generative language model 141 to generate labeled simulation samples.

[0225] S1002, specify the storage path of the corresponding generated data.

[0226] S1003: Generate labeled simulation samples according to the prompt template of the category using the generative language large model 141, and save the labeled simulation samples in a specified path.

[0227] S1004: Generate a data list to display the labeled simulation samples containing the prompt word under the specified path.

[0228] The generative language big model 141 includes any open source generative big model or language big model API. In some feasible implementations, a certain open source generative big model or language big model API can be selected according to the hardware / platform configuration. The language big model API includes GPT4 or CHATGLM.

[0229] The generative language model 141 takes the category prompt template as input and outputs labeled simulated samples.

[0230] For example, the hospital general data prompt template is input into GPT4 or CHATGLM to generate a simulated sample with the label "hospital"; the emotion recognition prompt template is input into GPT4 or CHATGLM to generate a simulated sample with the label "emotion recognition".

[0231] Figure 11b This is a schematic diagram of the label generation interface provided in Example 2 of this application. Figure 11b As shown in Figure 2, a fine-tuned large language classification model is trained based on labeled simulated samples in the training set corpus.

[0232] In the case of a large number of unlabeled real samples provided by the user, the process of training and fine-tuning the large language classification model to label the data includes steps S1004-S1006.

[0233] S1004: Using the labeled simulated samples generated in step S1003 to train and fine-tune the large language classification model, a trained large language classification model is obtained. The large language classification model is used to classify unlabeled real samples.

[0234] S1005 , a large number of unlabeled real samples are input into the large language classification model 142 , and labeled real samples are output, where the labels indicate the classification of the real samples.

[0235] S1006: Save the labeled simulated samples and / or labeled real samples in the set storage path / database 13 to obtain the training set corpus.

[0236] The training set corpus is used to train the text screening module, semantic fusion classification model, small-sample learning model, and label extraction model. The semantic fusion classification model, small-sample learning model, and label extraction model trained using real labeled samples from the training set corpus perform more stably in actual tests.

[0237] The data generation method proposed in Example 2 of the present application is based on a generative large model to adaptively construct the data required for file classification according to user needs and actual conditions, quickly collect classification samples, and can meet the user's sample data collection work with no samples or few samples.

[0238] Example 3

[0239] like Figure 12As shown, Example 3 of the present application also provides a data classification management method that can optimize classification storage and retrieval management strategies based on user needs and historical retrieval records. It is understood that this method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. It includes the following steps:

[0240] S1201: Count the search term records of users searching for documents according to a fixed period.

[0241] S1202: Determine the search frequencies of different search terms based on the search term records, and determine hot words based on the search frequencies. Hot words are search terms with a search frequency higher than a set threshold.

[0242] S1203: sorting different hot words into different levels based on the search frequency ranking of the multiple hot words.

[0243] S1204 , determining a multi-level label of the text to be classified based on one or more hot words in the text to be classified, where the multi-level label is obtained by sorting the one or more hot words according to search frequency.

[0244] In some feasible implementations, the text to be classified may be input into the tag extraction model 24 , and the tag extraction model 24 extracts one or more hot words from the text to be classified.

[0245] S1205 , saving the preset categories, semantic features and / or multi-level tags of the text in the tag of the text in order of priority to obtain a tag structure; the tag structure includes the preset categories, newly defined categories and multi-level tags of the text.

[0246] Among them, the preset categories and newly defined categories of the text can be obtained based on the method of Example 1 of the present application, which will not be repeated here.

[0247] S1206: Save the label structure of the text to be classified in the specified path.

[0248] In some feasible implementations, the data classification management method provided in Example 3 of the present application further includes a retrieval strategy:

[0249] S1207, searching the tag structure under the specified path according to the search term input by the user in the search task, and outputting the search results; the tag structure includes the preset categories, newly defined categories and multi-level tags of the text to be classified.

[0250] In actual retrieval tasks, high-priority tags in the tag structure can be matched first according to the search terms entered by the user, and then low-priority tags in the tag structure can be matched. Based on this, results for the retrieval task can be output more quickly.

[0251] The embodiment of the present application provides a computing device 1000. Figure 13 As shown, computing device 1000 includes a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. Processor 1004, memory 1006, and communication interface 1008 communicate with each other via bus 1002. Computing device 1000 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1000.

[0252] The bus 1002 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 13 The bus 1004 may include a path for transmitting information between various components of the computing device 1000 (eg, the memory 1006, the processor 1004, and the communication interface 1008).

[0253] The processor 1004 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0254] The memory 1006 may include volatile memory, such as random access memory (RAM). The processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0255] The memory 1006 stores executable program codes, and the processor 1004 executes the executable program codes to respectively implement the aforementioned Figure 13The functions of the classification module 111, the retrieval module 112, the analysis module 113 and the data generation module 114 shown in the figure can realize all or part of the steps of the method in the above embodiment. That is, the memory 106 stores instructions for executing all or part of the steps in the method in the above embodiment.

[0256] Alternatively, the memory 1006 stores executable code, and the processor 1004 executes the executable code to respectively implement the functions of the aforementioned data classification management device 11, thereby implementing all or part of the steps in the above-mentioned embodiment method. In other words, the memory 1006 stores instructions for executing all or part of the steps in the above-mentioned embodiment method.

[0257] The communication interface 1003 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1000 and other devices or a communication network.

[0258] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0259] like Figure 14 As shown, the computing device cluster includes at least one computing device 1000. The memory 1006 in one or more computing devices 1000 in the computing device cluster may store the same instructions for executing all or part of the steps in the above embodiment method.

[0260] In some possible implementations, the memory 1006 of one or more computing devices 1000 in the computing device cluster may also store partial instructions for executing all or part of the steps in the above-described method. In other words, the combination of one or more computing devices 1000 can jointly execute instructions for executing all or part of the steps in the above-described method.

[0261] It should be noted that the memory 1006 in different computing devices 1000 in the computing device cluster can store different instructions, which are respectively used to execute part of the functions of the data classification management device 11. That is, the instructions stored in the memory 1006 in different computing devices 1000 can implement the aforementioned Figure 13 The functions of one or more modules among the classification module 111, the retrieval module 112, the analysis module 113 and the data generation module 114 shown in FIG.

[0262] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network, which may be a wide area network or a local area network. Figure 15 A possible connection method is shown. Figure 15 As shown, two computing devices 1000A and 1000B are connected via a network. Specifically, the connection to the network is achieved through a communication interface within each computing device. In this possible implementation, memory 1006 within computing device 1000A stores instructions for executing the functions of classification module 111 and retrieval module 112. Simultaneously, memory 1006 within computing device 1000B stores instructions for executing the functions of analysis module 113 and data generation module 114.

[0263] It should be understood that Figure 15 The functionality of the computing device 1000A shown in FIG. 1 may also be implemented by multiple computing devices 1000. Similarly, the functionality of the computing device 1000B may also be implemented by multiple computing devices 1000.

[0264] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.

[0265] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the method in the above embodiment.

[0266] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0267] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0268] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).

[0269] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.

Claims

1. A data classification management method, characterized in that: The method comprises: Inputting the text to be classified into a first classification model; The first classification model determines key information based on the text to be classified; obtains semantic features based on the key information; the semantic features are feature vectors including word vectors and word frequency features of multiple key words; and determines attribute information of the text to be classified based on the semantic features, the attribute information including at least a first attribute, which is a preset category.

2. The method according to claim 1, characterized in that The key information includes a plurality of key words, and determining the key information based on the input text to be classified includes: Segmenting the text to be classified to obtain a plurality of words; Calculating the word frequency entropy of each word in the text to be classified, where the word frequency entropy is the proportion of each word appearing in each of the preset categories; Determine one or more key words in the text to be classified based on the word frequency entropy; The key information is determined according to the one or more key words.

3. The method according to any one of claims 1 or 2, characterized in that The key information includes a plurality of key words, and obtaining a semantic feature vector according to the key information includes: Determine the word vector of each key word; Determining the word frequency feature of each key word; The word vector and the word frequency feature are combined to obtain the word feature of each of the key words; and the word features of each of the key words are combined to obtain the semantic feature.

4. The method according to any one of claims 1 to 3, characterized in that Determining attribute information of the text to be classified according to the semantic features includes: In a case where the semantic feature vector conforms to the preset category, a first attribute of the text to be classified is determined.

5. The method according to any one of claims 1 to 4, characterized in that The attribute information further includes a second attribute, and determining the attribute information of the text to be classified according to the semantic feature includes: When the semantic feature does not conform to the preset category, the second attribute of the text to be classified is determined based on the distance between the word feature of each keyword in the feature vector and the sample feature of the user corpus; the sample feature of the user corpus is obtained by extracting the features of the custom category, and the second attribute is a newly defined category.

6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: When the confidence level of the first attribute is lower than a threshold requirement, inputting the text to be classified into a second classification model; the classification accuracy of the second classification model is higher than that of the first classification model; The second classification model outputs an optimized attribute of the text to be classified, where the optimized attribute is a first attribute; and the confidence level of the optimized attribute meets the threshold requirement.

7. The method according to claim 6, characterized in that The second classification model learns the semantic information of the text to be classified, and determines the optimized attributes of the text to be classified according to the semantic information of the text.

8. The method according to any one of claims 1 to 7, characterized in that The attribute information further includes a third attribute, and the method further includes: Extracting one or more hot words from the text to be classified; A third attribute of the text to be classified is output according to the one or more hot words; the third attribute includes one or more levels of labels, and the one or more levels of labels are obtained by sorting the one or more hot words according to retrieval frequency.

9. The method according to any one of claims 1 to 8, characterized in that The method further comprises: Receive the user's search request and determine the search term; A search is performed in the attribute information of the text to be classified based on the search term, and a search result is output.

10. The method according to any one of claims 1 to 9, characterized in that: The attribute information includes a tag structure, and the method further includes: Save the text to be classified; A tag structure is generated according to the first attribute, the second attribute and / or the third attribute of the text to be classified.

11. The method according to claim 10, characterized in that The searching the attribute information of the to-be-classified text based on the search term includes: matching the tag structure according to the priority level of the search term.

12. The method according to any one of claims 1 to 11, characterized in that The method further comprises: Obtain labeled samples of the corresponding category according to the prompt template; the labeled samples include labeled simulated samples and labeled real samples; The labeled simulated samples and / or labeled real samples are saved as training set corpus.

13. The method according to claim 12, characterized in that The method of obtaining labeled samples of corresponding categories according to the prompt template includes: based on a generative language large model, taking the prompt template as input, and outputting labeled simulation samples.

14. The method according to claim 12, characterized in that The method of obtaining labeled samples of corresponding categories according to the prompt template includes: based on a fine-tuned large language classification model, taking unlabeled real samples as input, and outputting labeled real samples.

15. A data classification management device, characterized in that: The device at least comprises: A classification module is used to input the text to be classified into a first classification model, the first classification model determines key information based on the input text to be classified; obtains a semantic feature vector based on the key information; the semantic feature vector includes the word vector and word frequency features of the key information; and determines the attribute information of the text to be classified based on the semantic feature vector, the attribute information includes at least a first attribute, and the first attribute is a preset category.

16. The device according to claim 15, characterized in that The first classification model includes a text screening module, which is used to segment the text to be classified to obtain multiple words; Calculating the word frequency entropy of each word in the text to be classified; The word frequency entropy is the proportion of each word appearing in each of the preset categories; Determine one or more key words in the text to be classified based on the word frequency entropy; The key information is determined according to one or more key words.

17. The device according to any one of claims 15-16, characterized in that In a case where the semantic feature conforms to the preset category, the first classification model determines a first attribute of the text to be classified.

18. The device according to any one of claims 15 to 17, characterized in that: When the semantic feature does not conform to the preset category, the second attribute of the text to be classified is determined based on the distance between the word feature of each keyword and the sample feature of the user corpus; the second attribute is a newly defined category, and the sample feature of the user corpus is obtained by extracting the features of the custom category.

19. The device according to any one of claims 15 to 18, characterized in that The classification module further includes a second classification model; the classification accuracy of the second classification model is higher than that of the first classification model; When the confidence of the first attribute is lower than the threshold requirement, the second classification model outputs an optimized attribute based on the text to be classified, and the optimized attribute is the first attribute; the confidence of the optimized attribute meets the threshold requirement.

20. The device according to any one of claims 15 to 19, characterized in that The classification module also includes a label extraction model; The tag extraction model is used to extract one or more hot words in the text to be classified; output a third attribute of the text to be classified based on the one or more hot words; the third attribute includes one or more levels of labels, and the one or more levels of labels are obtained by sorting the one or more hot words according to the retrieval frequency.

21. The device according to any one of claims 15 to 20, characterized in that The device also includes a data generation module; The data generation module is used to obtain labeled samples of corresponding categories according to the prompt template; the labeled samples include labeled simulated samples and labeled real samples; The labeled simulated samples and / or labeled real samples are saved as training set corpus.

22. A computing device, characterized in that include: at least one memory for storing a program; At least one processor is used to execute the program stored in the memory, and when the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1 to 14.

23. A computing device cluster, characterized in that: comprising at least one computing device, each computing device including a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in a memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 14.

24. A computer-readable storage medium storing a computer program, wherein when the computer program is executed on a processor, the processor is caused to execute the method according to any one of claims 1 to 14.

25. A computer program product, characterized in that When the computer program product is run on a processor, the processor is caused to perform the method according to any one of claims 1 to 14.