Intelligent text classification system and method supporting playback and multi-model management
Through intelligent text classification systems and methods that support playback and multi-model management, the problems of insufficient adaptability of text classification models to new concepts and events in the existing technology, high technical thresholds and diversified business scenarios are solved, and the flexible adaptation and diversified application needs of text classification models are achieved.
Patent Information
- Application Number
- CN202411888531.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-06
AI Technical Summary
The existing text classification technology faces the problems of classification forgetting caused by data and user preference migration, high threshold and low technical background of users, as well as the lack of diverse business scenarios and unified tools.
It provides an intelligent text classification system and method that supports playback and multi-model management, including label processing module, corpus management module, algorithm management module and model management module. Through the collaborative work of these modules, a multi-classification, multi-label, and multi-level label system is built, text data is collected and preprocessed, classification algorithms are registered and managed, and the algorithm and corpus are associated, trained and deployed.
It improves the flexible adaptability of the text classification model to new concepts and events, lowers the technical threshold, and allows ordinary users to conveniently use advanced text classification technology, enhances the scalability and flexibility of the text classification model, and can effectively meet the application needs of different fields and multiple levels.
Smart Images

Figure CN119938920A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to an intelligent text classification system and method supporting replay and multi-model management. Background Art
[0002] Text classification is a core branch in the field of natural language processing (NLP). It aims to map unstructured or semi-structured text data to a predefined set of categories to achieve effective organization and retrieval of information. At present, text classification technology has been widely used in many fields such as information retrieval, sentiment analysis, topic modeling, spam filtering, news detection, social media monitoring, etc., showing its powerful ability and practicality in processing large-scale text data. According to the implementation approach, text classification includes classification based on rule models and classification based on training models. Rule-based text classification methods focus on using predefined rules and patterns to identify and classify texts. These rules are usually manually constructed based on the knowledge and experience of domain experts and can accurately match specific text features such as keywords, phrases or grammatical structures. The advantage of this method lies in its interpretability and controllability, and it is particularly effective for some text classification tasks with strong structure and regularity. However, the limitations of rule models are also obvious. They often have difficulty in dealing with the ambiguity and complexity of language, and the cost of formulating and maintaining rules is high; text classification based on training models includes text classification based on machine learning and text classification based on deep learning. Machine learning-based text classification methods can learn feature weights from annotated text data to predict the category to which new text belongs, including Naive Bayes, Support Vector Machine (SVM), Decision Tree, Random Forest, etc. The advantage of machine learning models is that they can handle large-scale data sets and are generally more generalizable than rule-based models. However, they may require a large amount of annotated data, and the quality of feature engineering directly affects the final classification performance. Deep learning-based text classification can automatically learn complex representations of text without explicit feature engineering, such as convolutional neural networks (CNN), recurrent neural networks (RNN), long short-term memory networks (LSTM), and deep learning models of Transformer architecture. Deep learning models are particularly good at capturing long-term dependencies and semantic structures of text, thus achieving breakthrough performance on many NLP tasks.
[0003] However, even with the promotion of deep learning, text classification technology still faces a series of challenges, which not only limit the further development of the technology, but also hinder its effective application in specific fields, especially in professional fields such as military and finance. The first is classification forgetfulness caused by data and user preference migration. Current text classification models are mostly trained based on fixed label systems and static data sets, which cannot meet the dynamic classification needs of actual business in the real world. Especially in the military field, new vocabulary for tactics, weapon systems or enemy situation analysis continues to emerge, requiring classification systems to have higher flexibility and adaptability. The second is the high threshold of deep learning models and the low technical background of users. The training and deployment of deep learning models usually require the intervention of professional algorithm engineers, and involve complex parameter adjustment and optimization processes, making it difficult for ordinary users to implement them by themselves, increasing the difficulty and cost of popularizing text classification. Finally, there are diverse business scenarios and a lack of unified tools. Different application scenarios require text classification systems to have multiple classification capabilities such as multi-label, multi-classification, and multi-level. However, existing systems are often designed for a specific scenario, lack versatility and configurability, and are difficult to meet diversified business needs. Summary of the invention
[0004] The present invention aims to solve at least one of the above-mentioned technical problems existing in the prior art.
[0005] To this end, a first aspect of the present invention provides an intelligent text classification system that supports replay and multi-model management.
[0006] A second aspect of the present invention provides an intelligent text classification method that supports replay and multi-model management.
[0007] The present invention provides an intelligent text classification system supporting replay and multi-model management, comprising:
[0008] The label processing module is used to build and edit labels and their hierarchical structures, and to form a text classification label system by building a multi-classification, multi-label, and multi-level label library;
[0009] The corpus management module is used to collect, organize and pre-process text data, and process the sample data associated with the classification label system to form a classification corpus through corpus replay, data pre-processing, sample equalization and data enhancement functions; wherein, by constructing a corpus, the text classification corpus data is pre-processed and managed to realize the data association between the label system and the algorithm model;
[0010] An algorithm management module is used to register and manage classification association algorithms within the system; wherein the algorithm management module realizes integrated management of classification algorithms within the system and third-party classification algorithms by registering multi-classification, multi-label, and multi-level classification model algorithms, and configuring algorithm parameters, algorithm types, and algorithm addresses to form a classification algorithm library;
[0011] The model management module is used to associate, train and deploy the classification algorithm and corpus to form a classification service; wherein the model management module associates, integrates and publishes the algorithm model, classification corpus and label system by configuring, verifying and deploying the text classification model to form a text classification prediction service.
[0012] The intelligent text classification system supporting replay and multi-model management according to the above technical solution of the present invention may also have the following additional technical features:
[0013] In the above technical solution, users customize the creation and editing of tags through the tag processing module, and the edited content includes at least the name, description and attributes of the tag; wherein, the tag processing module allows users to create parent tags and child tags to form a tag tree structure; the tag processing module supports tag nesting and batch operations.
[0014] In the above technical solution, the corpus management module uploads text classification training sample corpus according to the text classification label, performs denoising, encoding conversion, text segmentation preprocessing, sample equalization processing, data enhancement and feedback classification data replay to form classification training corpus.
[0015] In the above technical solution, the corpus replay includes introducing feedback data based on historical classified corpus data by building a data storage mechanism combined with a classified corpus screening strategy and a replay strategy, so that the classification model can be continuously updated and trained;
[0016] The replay strategy is formulated based on the time window and classification accuracy index.
[0017] In the above technical solution, the data preprocessing includes performing denoising, encoding detection and conversion, text segmentation and text full-width and half-width conversion operations on the classified corpus data to ensure the quality of the corpus;
[0018] Furthermore, the sample balancing includes balancing the number of samples by repeatedly sampling samples, synthesizing new samples or deleting samples to address the problem of label sample imbalance in the classification corpus;
[0019] Furthermore, the data enhancement includes increasing the training data by synonym replacement, context transformation or random insertion / deletion / exchange for problems with limited training data.
[0020] In the above technical solution, the classification algorithm library includes machine learning algorithms and deep learning algorithms stored in the system and algorithms provided by third-party services called through an API interface.
[0021] In the above technical solution, the model management module includes a model configuration unit;
[0022] The model configuration unit is used to create and configure basic information of the text classification model;
[0023] Wherein, the model configuration unit configures the model name, model association algorithm, model type, and model association corpus through a human-computer interaction interface to generate model configuration information;
[0024] The model types include training models, rule models and fusion models.
[0025] In the above technical solution, the model management module also includes a model verification unit;
[0026] The model verification unit is used to perform consistency verification on the model association algorithm and the classification corpus to generate a classification basic model;
[0027] The consistency check includes: consistency check of text classification corpus and associated algorithms in the model-associated corpus, consistency check of associated corpus labels and defined classification models, and consistency check of associated classification corpus and classification labels.
[0028] In the above technical solution, the model management module also includes a model deployment unit:
[0029] The model deployment unit performs model training, business configuration and model publishing on the classification basic model to generate classification services; the model deployment unit has training models and business models deployed therein;
[0030] Among them, the model deployment unit calls the model training algorithm to fine-tune and evaluate the training model, and adjusts the replay strategy based on the time window and classification accuracy index to achieve continuous training and updating of the replayed classification data; configures the business model through the regularized expression of text content features and business logic association, configures the model name, model weight, and model quantity of the fusion model, and publishes the model to form a text classification prediction service to support business system calls.
[0031] An intelligent text classification method supporting replay and multi-model management provided by the present invention is applied to the intelligent text classification system supporting replay and multi-model management described in any one of the above technical solutions, and the method comprises:
[0032] Build and edit labels and their hierarchical structures, and form a text classification label system by building a multi-classification, multi-label, and multi-level label library;
[0033] Collect, organize and preprocess text data, and process the sample data associated with the classification label system to form a classification corpus through corpus replay, data preprocessing, sample balancing and data enhancement functions; wherein, by constructing a corpus, the text classification corpus data is preprocessed and managed to achieve data association between the label system and the algorithm model;
[0034] Register classification association algorithms, register multi-classification, multi-label, and multi-level classification model algorithms, and configure algorithm parameters, algorithm types, and algorithm addresses to form a classification algorithm library, thereby achieving integrated management of the system's internal and third-party classification algorithms;
[0035] The classification algorithm and corpus are associated, trained and deployed to form a classification service. Specifically, the algorithm model, classification corpus and label system are associated, integrated and published by configuring, verifying and deploying the text classification model to form a text classification prediction service.
[0036] In summary, due to the adoption of the above technical features, the beneficial effects of the present invention are:
[0037] The present invention performs text classification configuration with scenario-based and label system-configured methods as the core. Application personnel can realize continuous training and release of classification models by replaying corpus data and label systems in classification application scenarios according to specific tasks of actual applications, thereby improving the device's flexible adaptability to emerging concepts, events, or label changes caused by business changes, especially meeting the needs for classification of rapidly changing professional terms in the military, finance and other fields, and greatly enhancing the scalability and flexibility of text classification models in different fields.
[0038] The present invention adopts a multi-model association method to classify labels for complex and diverse text data. On the one hand, it integrates the mainstream classification model of the industry to realize one-click training and release of training models for classification scenarios, which reduces the technical threshold; on the other hand, it fully considers the actual business and adds business experience in multiple dimensions including the title, text, keywords, etc. of the material data. Through integration with the training model, it absorbs the advantages of different models, fully realizes the complementarity and applicability of the text classification model, and enables ordinary users to easily use advanced text classification technology without paying too much attention to the technology itself.
[0039] The present invention improves the ability of text classification to cope with complex and changeable business application scenarios by providing classification applications of multi-classification, multi-level, multi-label and other classification scenarios and the integrated deployment of training models and rule engines, and can effectively meet the multi-level and diversified application requirements of different fields and different businesses, and enhance the versatility of the tool itself. At the same time, through the feedback enhancement device of the tool, a continuous iteration and upgrade mechanism for the classification model is established, so that rapid model iteration can be achieved without the user's feeling, and the model effect is improved, thereby improving the practical application value of text classification technology in engineering.
[0040] Additional aspects and advantages of the present invention will become apparent from the following description or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0042] Figure 1 It is a schematic diagram of the operation flow of an intelligent text classification system supporting replay and multi-model management according to an embodiment of the present invention. DETAILED DESCRIPTION
[0043] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0044] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.
[0045] Refer to the following Figure 1 To describe an intelligent text classification system and method supporting replay and multi-model management provided according to some embodiments of the present invention.
[0046] Some embodiments of the present application provide an intelligent text classification system that supports replay and multi-model management.
[0047] like Figure 1 As shown, the first embodiment of the present invention proposes an intelligent text classification system that supports replay and multi-model management, including: a label processing module, a corpus management module, an algorithm management module and a model management module.
[0048] The label processing module is used to construct and edit labels and their hierarchical structures, and to form a text classification label system by building a multi-category, multi-label, and multi-level label library.
[0049] Specifically, users can customize the creation and editing of tags through the tag processing module, and the edited content includes at least the name, description and attributes of the tag, so as to build a tag system that meets specific business needs. Among them, the tag processing module allows users to create parent tags and child tags to form a tag tree structure, which is convenient for managing and organizing tags; through the customized classification tag system, the text classification tags and the display of the classification tag system are dynamically updated according to actual business needs, forming a multi-classification, multi-tag, and multi-level tag library.
[0050] In some embodiments, the label processing module supports label nesting and batch operations to adapt to diverse business scenarios; and through the batch creation and editing functions of labels, the maintenance of large-scale label systems is simplified.
[0051] The corpus management module is used to collect, organize and pre-process text data, and process the sample data associated with the classification label system to form a classification corpus through corpus replay, data pre-processing, sample equalization and data enhancement functions; wherein, by constructing a corpus, the text classification corpus data is pre-processed and managed to achieve data association between the label system and the algorithm model.
[0052] Specifically, the corpus management module uploads text classification training sample corpus according to text classification labels, performs denoising, encoding conversion, text segmentation preprocessing, sample equalization processing, data enhancement, and classification data replay for feedback to form classification training corpus.
[0053] In some embodiments, the corpus replay includes introducing feedback data based on historical classified corpus data by constructing a data storage mechanism in combination with a classified corpus screening strategy and a replay strategy, so that the classification model can be continuously updated and trained; wherein the replay strategy is formulated based on a time window and a classification accuracy index, for example, replaying the corpus according to a set period or replaying according to the classification accuracy in a text classification service fed back by users.
[0054] The data preprocessing includes performing denoising, encoding detection and conversion, text segmentation, and text full-width and half-width conversion on the classified corpus data to ensure the quality of the corpus. The sample balancing includes balancing the number of samples by repeated sampling, synthesizing new samples, or deleting samples to address the problem of unbalanced label samples in the classified corpus; the data enhancement includes increasing the training data by synonym replacement, context transformation, or random insertion / deletion / exchange to address the problem of limited training data.
[0055] The algorithm management module is used to register and manage the classification association algorithms in the system; wherein, the algorithm management module realizes the integrated management of the classification algorithms in the system and those of third parties by registering multi-classification, multi-label, and multi-level classification model algorithms, and configuring algorithm parameters, algorithm types, and algorithm addresses to form a classification algorithm library. Specifically, the classification algorithm library includes machine learning algorithms and deep learning algorithms stored in the system, as well as algorithms provided by third-party services through API interfaces.
[0056] Through the above algorithm management module, not only the functions of the algorithm library are expanded, but also flexibility is provided, so that users can choose the optimal algorithm combination according to actual needs, thereby improving the accuracy and practicality of the classification model.
[0057] The model management module is used to associate, train and deploy the classification algorithm and corpus to form a classification service; wherein the model management module associates, integrates and publishes the algorithm model, classification corpus and label system by configuring, verifying and deploying the text classification model to form a text classification prediction service. The model management module supports the joint use of multiple models to improve the accuracy and flexibility of classification.
[0058] In some embodiments, the model management module includes a model configuration unit, a model verification unit, and a model deployment unit.
[0059] The model configuration unit is used to create and configure the basic information of the text classification model; wherein the model configuration unit configures the model name, model associated algorithm, model type, and model associated corpus through a human-computer interaction interface to generate model configuration information; specifically, the model types include training models, rule models, and fusion models. In some embodiments, the model configuration unit can also configure the model weight and model priority.
[0060] The model verification unit is used to perform consistency verification on the model association algorithm and classification corpus to generate a classification basic model; wherein the consistency verification includes: consistency verification of the text classification corpus and the association algorithm in the model association corpus, consistency verification of the associated corpus label and the defined classification model, and consistency verification of the associated classification corpus and the classification label.
[0061] The model deployment unit performs model training, business configuration and model publishing on the classification basic model to generate classification services; the model deployment unit deploys training models and business models; wherein the model deployment unit calls the model training algorithm to perform model fine-tuning and model evaluation on the training model, and adjusts the replay strategy based on the time window and classification accuracy index to achieve continuous training and updating of the replayed classification data; the business model is configured through the regularized expression of text content features and business logic association, and the model name, model weight, and model quantity of the fusion model are configured and the model is published to form a text classification prediction service to support business system calls.
[0062] Some other embodiments of the present invention provide an intelligent text classification method supporting replay and multi-model management, which is applied to the intelligent text classification system supporting replay and multi-model management described in any of the above embodiments, and the method includes steps S1-S5.
[0063] S1. Build and edit labels and their hierarchical structures, and form a text classification label system by building a multi-classification, multi-label, and multi-level label library.
[0064] S2. Collect, organize and preprocess text data, and process the sample data associated with the classification label system through corpus replay, data preprocessing, sample equalization and data enhancement functions to form a classification corpus; wherein, by constructing a corpus, the text classification corpus data is preprocessed and managed to achieve data association between the label system and the algorithm model.
[0065] S3. Register classification association algorithms. By registering multi-classification, multi-label, and multi-level classification model algorithms, and configuring algorithm parameters, algorithm types, and algorithm addresses to form a classification algorithm library, integrated management of system and third-party classification algorithms is achieved.
[0066] S4. Associate, train and deploy the classification algorithm and corpus to form a classification service; wherein, by configuring, verifying and deploying the text classification model, the algorithm model, classification corpus and label system are associated, integrated and published to form a text classification prediction service.
[0067] S5. According to user feedback from the text classification prediction service, the corpus replay strategy in step S2 is adjusted based on the time window and classification accuracy index to achieve continuous training and updating of the replayed classification data.
[0068] In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0069] Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. An intelligent text classification system supporting replay and multi-model management, characterized in that: include: The label processing module is used to build and edit labels and their hierarchical structures, and to form a text classification label system by building a multi-classification, multi-label, and multi-level label library; The corpus management module is used to collect, organize and pre-process text data, and process the sample data associated with the classification label system to form a classification corpus through corpus replay, data pre-processing, sample equalization and data enhancement functions; wherein, by constructing a corpus, the text classification corpus data is pre-processed and managed to realize the data association between the label system and the algorithm model; An algorithm management module is used to register and manage classification association algorithms within the system; wherein the algorithm management module realizes integrated management of classification algorithms within the system and third-party classification algorithms by registering multi-classification, multi-label, and multi-level classification model algorithms, and configuring algorithm parameters, algorithm types, and algorithm addresses to form a classification algorithm library; The model management module is used to associate, train and deploy the classification algorithm and corpus to form a classification service; wherein the model management module associates, integrates and publishes the algorithm model, classification corpus and label system by configuring, verifying and deploying the text classification model to form a text classification prediction service.
2. The intelligent text classification system supporting replay and multi-model management according to claim 1, characterized in that: The user customizes the creation and editing of tags through the tag processing module, and the edited content includes at least the name, description and attributes of the tag; wherein, the tag processing module allows the user to create parent tags and child tags to form a tag tree structure; the tag processing module supports tag nesting and batch operations.
3. The intelligent text classification system supporting replay and multi-model management according to claim 1, characterized in that: The corpus management module uploads text classification training sample corpus according to the text classification label, performs denoising, encoding conversion, text segmentation preprocessing, sample equalization processing, data enhancement and feedback classification data replay on the training corpus to form classification training corpus.
4. The intelligent text classification system supporting replay and multi-model management according to claim 3, characterized in that: The corpus replay includes introducing feedback data based on historical classified corpus data by building a data storage mechanism combined with a classified corpus screening strategy and a replay strategy, so that the classification model can be continuously updated and trained; The replay strategy is formulated based on the time window and classification accuracy index.
5. The intelligent text classification system supporting replay and multi-model management according to claim 3, characterized in that: The data preprocessing includes performing denoising, encoding detection and conversion, text segmentation, and text full-width and half-width conversion operations on the classified corpus data to ensure the quality of the corpus; Furthermore, the sample balancing includes balancing the number of samples by repeatedly sampling samples, synthesizing new samples or deleting samples to address the problem of label sample imbalance in the classification corpus; Furthermore, the data enhancement includes increasing the training data by synonym replacement, context transformation or random insertion / deletion / exchange for problems with limited training data.
6. The intelligent text classification system supporting replay and multi-model management according to claim 1, characterized in that: The classification algorithm library includes machine learning algorithms and deep learning algorithms stored in the system, as well as algorithms provided by third-party services called through an API interface.
7. The intelligent text classification system supporting replay and multi-model management according to claim 1, characterized in that: The model management module includes a model configuration unit; The model configuration unit is used to create and configure basic information of the text classification model; Wherein, the model configuration unit configures the model name, model association algorithm, model type, and model association corpus through a human-computer interaction interface to generate model configuration information; The model types include training models, rule models and fusion models.
8. The intelligent text classification system supporting replay and multi-model management according to claim 7, characterized in that: The model management module also includes a model verification unit; The model verification unit is used to perform consistency verification on the model association algorithm and the classification corpus to generate a classification basic model; The consistency check includes: consistency check of text classification corpus and associated algorithms in the model-associated corpus, consistency check of associated corpus labels and defined classification models, and consistency check of associated classification corpus and classification labels.
9. The intelligent text classification system supporting replay and multi-model management according to claim 8, characterized in that: The model management module also includes a model deployment unit: The model deployment unit performs model training, business configuration and model publishing on the classification basic model to generate classification services; the model deployment unit has training models and business models deployed therein; Among them, the model deployment unit calls the model training algorithm to fine-tune and evaluate the training model, and adjusts the replay strategy based on the time window and classification accuracy index to achieve continuous training and updating of the replayed classification data; configures the business model through the regularized expression of text content features and business logic association, configures the model name, model weight, and model quantity of the fusion model, and publishes the model to form a text classification prediction service to support business system calls.
10. An intelligent text classification method supporting replay and multi-model management, characterized in that: An intelligent text classification system supporting replay and multi-model management as applied to any one of claims 1 to 9, the method comprising: Build and edit labels and their hierarchical structures, and form a text classification label system by building a multi-classification, multi-label, and multi-level label library; Collect, organize and preprocess text data, and process the sample data associated with the classification label system to form a classification corpus through corpus replay, data preprocessing, sample balancing and data enhancement functions; wherein, by constructing a corpus, the text classification corpus data is preprocessed and managed to achieve data association between the label system and the algorithm model; Register classification association algorithms, register multi-classification, multi-label, and multi-level classification model algorithms, and configure algorithm parameters, algorithm types, and algorithm addresses to form a classification algorithm library, thereby achieving integrated management of the system's internal and third-party classification algorithms; The classification algorithm and corpus are associated, trained and deployed to form a classification service. Specifically, the algorithm model, classification corpus and label system are associated, integrated and published by configuring, verifying and deploying the text classification model to form a text classification prediction service.