Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

307 results about "Text categorization" patented technology

Text categorization (a.k.a. text classification) is the task of assigning predefined categories to free-text documents. It can provide conceptual views of document collections and has important applications in the real world.

Method and system for large language model (LLM)-selection for response generation to user queries

Disclosed herein, is a method and system for selecting a LLM for response generation to user queries. The method includes receiving a user query from a user device. The method includes determining, for the user query, a query type from a set of query types through a fine-tuned text classification model. The method includes retrieving a plurality of document embeddings based on the user query and the query type from a vector database through a semantic search technique. The method includes preparing a prompt using the user query and the plurality of document embeddings. The method includes inputting the prompt to an LLM selected from a set of LLMs based on the query type. The method includes generating, via the selected LLM, a response to the user query based on the prompt.
Owner:L&T TECH SERVICES LTD

Data-free knowledge amalgamation for text classification

PendingUS20260030511A1Biological modelsPseudo dataText categorization
A method, computer system, and a computer program product for data-free knowledge amalgamation are provided. Multiple pre-trained teacher machine learning models are obtained. Each is trained on a respective different set of training data. Pseudo-data samples that mimic original training data of the teacher models are generated. A block-wise amalgamation with a self-regulative strategy to integrate knowledge from the multiple teacher models is implemented by inputting the pseudo-data samples into the teacher models and into a student machine learning model. The implementing also includes aligning intermediate representations of the student model with a unified representation capturing relevant features from the teacher models.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

News industry classification method and device based on large language model active learning

The invention discloses a news industry classification method and device based on large language model active learning, relates to the technical field of text classification, and can remarkably reduce the manual annotation cost while ensuring the news text classification precision. According to the scheme, the method comprises the following steps: calling a large language model to classify each news text for multiple times for a plurality of news texts to obtain at least one corresponding tag; based on a majority voting mode, dividing the plurality of news texts and the corresponding labels into high-confidence samples and low-confidence samples; performing label labeling on the low-confidence sample to obtain a labeled sample; and taking the labeled sample and the high-confidence sample as a training set, carrying out iterative training on the classification model by using the training set until the model precision of the classification model meets a preset condition, obtaining a target classification model, and carrying out prediction classification on the news text by using the target classification model.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Cross-language text classification and processing method and system based on deep transfer learning

The invention provides a cross-language text classification and processing method and system based on deep transfer learning, and relates to the technical field of text processing, and the method comprises the steps: extracting feature representations of a source language text and a target language text at different linguistic levels through a multi-level semantic transfer network; and determining an optimal alignment path, performing nonlinear mapping alignment to obtain fusion features, propagating category semantics by using a semantic bridging function, iteratively updating pseudo-tag confidence distribution of the target language text, and completing classification in combination with a multi-task learning model. According to the method, the problem of text classification in a cross-language scene is effectively solved, and the accuracy and efficiency of low-resource language text processing are improved.
Owner:SHANGHAI XIRUAN TECH CO LTD

Multi-label text classification method based on semantic representation enhancement and dynamic weighted depolarization contrast learning and application thereof

ActiveCN121858739AEnhanced Semantic RepresentationImprove the problem of insufficient semantic representationBiological modelsSpecial data processing applicationsData setSemantic representation
The invention relates to the technical field of natural language processing text classification, in particular to a multi-label text classification method based on semantic representation enhancement and dynamic weighted depolarization contrast learning and application of the multi-label text classification method. Preprocessing data in the training data set to obtain input tensor representation; enhancing text semantic representation by fusing multi-layer hidden semantic representation of a depth model; secondly, an improved label graph convolutional network is constructed, regularization, layer normalization, residual connection and label perception attention pooling are introduced into a graph neural network, fine-grained representation of the relation between labels is achieved, and label-text interaction is enhanced; and finally, introducing depolarization weighted contrast loss to construct dynamic weighted depolarization contrast learning, endowing a negative sample with a higher weight, and reducing false negative sample interference, thereby overcoming label semantic overlapping, and aiming at solving the problem of how to enhance the characterization capability of the model in a multi-label semantic overlapping and label incomplete scene.
Owner:YUNNAN NORMAL UNIV

Sample generation method and device, text classification model training method and device and medium

The invention provides a sample generation method, a text classification model training method and device and a medium, and relates to the technical field of artificial intelligence, in particular to the technical field of text classification, natural language processing and deep learning. According to the implementation scheme, a target text unit is recognized from an original sample set used for training a text classification model; determining a target category of an enhanced sample to be generated for the target text unit in the plurality of categories; obtaining a first semantic scene rule corresponding to a target category of the target text unit; and based on the target text unit and the first semantic scene rule, utilizing the large model to generate a first enhanced sample for training a text classification model.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Text classification method and device, electronic equipment and storage medium

The invention provides a text classification method and device, electronic equipment and a storage medium, relates to the technical field of data processing, in particular to the field of natural language processing and the like, and can be used for searching text classification and other application scenes. According to the implementation scheme, an initial similar list is extracted from a semantic vector library according to a user query word; performing information enhancement on each available query word in the initial similar list by utilizing a knowledge graph to obtain an information enhanced similar list; performing dynamic weighting according to the category of each available query word in the information enhancement similar list to obtain the weight of each available query word; according to the weight, determining the weighted similarity between each available query word and the user query word, and taking the available query word with the highest weighted similarity value as a candidate query word; and when the weighted similarity value of the candidate query word is greater than or equal to a preset similarity threshold value, taking the tag corresponding to the candidate query word as the tag of the user query word.
Owner:BAIDU COM TIMES TECH (BEIJING) CO LTD

Policy text intelligent classification system and method based on enhanced multi-layer attention mechanism

The invention provides a policy text intelligent classification system and method based on an enhanced multilayer attention mechanism, and belongs to the technical field of natural language processing and artificial intelligence. The system comprises a data preprocessing module, a multi-model prediction module, a feature fusion preprocessing module, a multi-layer attention fusion module, an uncertainty perception optimization module, an intelligent training control module and a classification result output module. The data preprocessing module is responsible for policy text standardization processing and BERT feature extraction; the multi-layer attention fusion module realizes dynamic weighted fusion of a model prediction result and an original feature through a double-layer attention mechanism; and the uncertainty perception optimization module calculates prediction uncertainty based on a GHM Loss mechanism. According to the method, a technical support and practice framework is provided for effectively supporting policy document automatic management and intelligent decision making when an artificial intelligence NLP technology is applied to practice innovation of an intelligent government and tasks with huge policy text classification cardinal number, complex tasks and high reliability requirements are processed.
Owner:HARBIN INST OF TECH

Text classification model training method, text classification method and equipment

The invention relates to a training method of a text classification model and a text classification method and device.The training method of the text classification model comprises the steps that text samples carrying the same classification labels are determined, the text samples are clustered to obtain M clustering clusters, a target text sample is determined from the M clustering clusters, and the target text sample is selected from the M clustering clusters; the method comprises the steps of obtaining a target text sample, splicing the target text sample to obtain a combined text sample, and training a to-be-trained text classification model based on the text sample and the combined text sample to obtain a target text classification model. By means of the method, the classification accuracy of the text classification model can be improved, and then the accuracy of text classification is improved.
Owner:MASHANG CONSUMER FINANCE CO LTD

A text classification method based on a bidirectional long short-term memory model and a knowledge graph

The application provides a text classification method based on combination of bidirectional long short-term memory model and knowledge graph retrieval, which retrieves relevant prior supporting facts from the knowledge graph according to the task by using an attention mechanism, and incorporates the prior supporting facts into a feature space together with features learned from training data to classify the text. It firstly generates a word embedding model of a sentence by using a GloVe tool, and then respectively puts the word embedding model into a knowledge graph retrieval module and a bidirectional long short-term memory network BiLSTM, and then splices the output of the retrieval model and the output of the BiLSTM model to obtain a final classification. Compared with a traditional method, the accuracy is obviously improved by using the method of the knowledge graph. Finally, the model is evaluated on a 20Newsgroups text classification data set, and the experimental results prove the effectiveness of the model.
Owner:SHANDONG UNIV OF SCI & TECH

Text classification method and device, computer equipment and storage medium

The invention relates to a text classification method and device, computer equipment and a storage medium, and the method comprises the steps: inputting a sample text sequence into an encoder module of a to-be-trained text classification model, and obtaining a text feature matrix; processing the text feature matrix, a preset first fusion weight and a preset second fusion weight through a feature fusion module of the to-be-trained text classification model to obtain a first fusion feature and a second fusion feature; inputting the first fusion feature into a first classification head of a to-be-trained text classification model to obtain a text classification prediction result; inputting the second fusion feature into a second classification head of the to-be-trained text classification model to obtain a text sequence generation result; and according to the text classification prediction result and the text sequence generation result, training the to-be-trained text classification model to obtain a trained text classification model, thereby improving the text classification precision of the model.
Owner:JUHAOKAN TECH CO LTD +1

Search text classification method and device, computer device and storage medium

The application discloses a search text classification method and device, computer equipment and a storage medium, and belongs to the technical field of computers. The method comprises the following steps: acquiring a search text to be classified and a plurality of search results corresponding to the search text, wherein the plurality of search results comprise historical search results of executed interactive operations; determining reference category information of the search text based on preset categories to which the plurality of search results belong; acquiring fusion features of the search text based on the search text and the reference category information; and classifying the search text based on the fusion features to obtain predicted category information of the search text. The application considers that the possibility that the preset category to which the search text belongs is the same as the preset category to which the historical search results of the executed interactive operations belong is relatively large, so the preset category to which the historical search results of the executed interactive operations belong is additionally considered in addition to the search text itself, and the accuracy of classifying the search text is improved.
Owner:TENCENT TECH (BEIJING) CO LTD

Multi-level harmful text classification method

The invention discloses a multilevel harmful text classification method, and relates to the technical field of natural language processing and text classification, and the method comprises the steps: firstly collecting harmful and normal texts to construct a positive and negative sample library; establishing a harmful category hierarchical tree; blocking the seed data, generating harmful examples by using a cue word template, and screening the harmful examples; then, a model is built for comparative learning training; finally, dynamic threshold adjustment is performed, a threshold is determined by combining real-time data, and whether the sample belongs to a specific harmful category or not is judged accordingly. Through a dynamic harmful category hierarchical tree and hierarchical activeness matching algorithm, it is ensured that a classification system is comprehensive in timeliness and accurate in positioning; meanwhile, seed text partitioning, sample expansion and the like are utilized to generate high-quality training data and flexibly adjust a classification threshold, so that the accuracy, adaptability and practicability of harmful text detection are remarkably improved.
Owner:HANGZHOU ANQUAN DIGITAL INTELLIGENCE TECH CO LTD

Model training method, device and equipment and computer storage medium

The embodiment of the invention provides a model training method and device, equipment and a computer storage medium, and relates to the technical field of artificial intelligence. The method comprises the steps that when a multi-label text classification model is trained, the relevance between category labels marked by training data is analyzed, exclusiveness and relevance between category pairs are determined, the multi-label text classification model learns the relevance relation between the exclusiveness and the relevance between the category labels, and the classification efficiency of the multi-label text classification model is improved. And the text feature vector is fused with the text feature vector of the training data, so that the predicted multi-class result output by the multi-label text classification model is improved. Besides, since the prediction of the multi-class result learns exclusiveness between class labels, the multi-label text classification model can be prevented from outputting class pairs with exclusiveness at the same time, so that the accuracy of the model output result is further improved.
Owner:CHINA MOBILE M2M +1

Big language model fused reasoning enhanced text classification method and system

The invention relates to the technical field of text classification, and particularly discloses a reasoning enhanced text classification method and system fused with a large language model.The method comprises the steps that chained reasoning data containing the step-by-step reasoning process from an input text to a classification label is generated for a training sample of a text classification task through the large language model; based on the chain reasoning data, a lightweight fine tuning method based on LoRA is adopted to perform fine tuning on a target large language model, and a text classification model is obtained; and for a to-be-classified text, generating a plurality of reasoning paths by using the fine-tuned text classification model, obtaining a final classification label through a majority voting mechanism based on a classification result output by each reasoning path, and performing multi-path consistency majority voting through mechanisms such as high-quality reasoning chain generation, lightweight fine tuning based on the reasoning chain, multi-path consistency majority voting and the like. The multi-step reasoning ability and classification stability of the model are effectively improved, and the technical bottleneck of inconsistency of reasoning blind spots and results in the prior art is overcome.
Owner:JILIN UNIVERSITY

A label semantic enhanced weakly supervised text classification method and system

This invention discloses a weakly supervised text classification method and system with enhanced label semantics, belonging to the field of machine learning. Based on the BERT weakly supervised text classification framework, in the category vocabulary construction stage, weighted category representation based on Zipf's law is used to denoise category words, and irrelevant words in the category vocabulary are removed by utilizing the decreasing semantic similarity property. In the sample labeling stage, word category labeling is based on the MASK mechanism, and then the classification model is optimized based on a self-training module. Using category indicator words in the samples as bridges, a cross-level semantic association of "sample sentence - indicator word - category label" is established. This invention introduces more algorithms to reduce label noise in the vocabulary construction and sample weak label generation stages to achieve the effect of enhanced label semantics, significantly improving text classification performance in different language environments.
Owner:INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES

Student model training method and text classification system based on pre-trained language model

A method for training a student model based on a pre-trained language model (PLM) and a text classification system are disclosed. The method includes: constructing cue-based training samples; adjusting the pre-trained language model using the cue-based training samples to obtain a cue-adjusted teacher model; and training the student model using the processed training samples, wherein during training, the student model simultaneously learns the classification probability vectors output by the cue-adjusted teacher model and the original teacher model. This invention requires the student model to learn from two teacher models simultaneously, thereby alleviating the overfitting problem of the student model in small-sample scenarios by adding a distillation path that learns from the original PLM teacher model with unsupervised data. Furthermore, by transferring the intermediate layer representation of the PLM through knowledge probes and stabilizing the performance of knowledge distillation through comparative learning, the student model can learn higher-order dependencies from the intermediate layer representation of the teacher model, improving the accuracy and efficiency of knowledge distillation.
Owner:ALIBABA (CHINA) CO LTD

A text classification method, device, equipment and storage medium

The application discloses a text classification method and device, equipment and a storage medium. The method comprises the following steps: obtaining text data; determining the feature vectors of the document nodes, concept nodes and word nodes corresponding to the text data based on the text data; constructing a text heterogeneous graph based on the feature vectors of the document nodes, concept nodes and word nodes; determining the weights of the edges between the nodes in the text heterogeneous graph; obtaining the text feature vector corresponding to the text data based on the text heterogeneous graph; and classifying the text feature vector by using a classification function to determine the text category. In this way, the feature vectors of the concept nodes are obtained, the prior knowledge in the text is obtained, and the concept nodes are fused when the text heterogeneous graph is constructed, so that the feature sparsity problem caused by the lack of context in the short text can be relieved to some extent, the text feature vector extracted based on the text heterogeneous graph can more accurately represent the features of the text, and the accuracy of the text classification is improved.
Owner:CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1

Urban land utilization identification method based on BERT model text classification algorithm

The invention relates to an urban land utilization recognition method based on a BERT model text classification algorithm, and the method is characterized in that the method comprises the following steps: 1, determining an urban land utilization recognition type, determining an urban land utilization recognition region range, obtaining POI data in the region range, and carrying out the data screening; step 2, establishing an urban land utilization identification unit in an urban land utilization identification model based on a BERT model text classification algorithm, and associating the screened POI data with the urban land utilization identification unit; step 3, carrying out geographic text mining on the POI data; and step 4, performing urban land utilization identification of the urban land utilization identification model based on the BERT model text classification algorithm to obtain a high-precision urban land utilization identification result. According to the method, high-precision urban land utilization identification can be timely and accurately carried out by utilizing the available POI data.
Owner:TIANJIN UNIV

Text classification mapping method based on hierarchical semantic matching and fact data fusion

The invention provides a text classification mapping method based on hierarchical semantic matching and fact data fusion, and relates to the technical field of text classification. The problems that in an existing text classification mapping technology, cross-category logic errors exist, the matching accuracy rate and the coverage rate are difficult to balance, and noise-containing external fact data cannot be effectively purified and fused are solved. According to the method, text classification mapping is realized through a two-stage process: firstly, data is preprocessed and vectorized, then full-text and high-level semantic similarity is subjected to weighted fusion, and prediction mapping is generated by using double thresholds; and then purifying external fact data, finally fusing the two types of mappings and labeling sources, and outputting a high-quality mapping table. According to the method, the mapping accuracy and logic consistency are effectively improved, noise data are filtered to enhance the reliability of a result, the accuracy and the coverage rate are balanced through a double-threshold strategy, high automation of the process is realized to improve the efficiency, and the technical framework is high in universality and can adapt to multi-field text classification mapping requirements.
Owner:BEIJING HUIWEN JIECHUANG TECHNOLOGY CO LTD

Techniques for providing explanations about text classification

The chatbot system is configured to execute code to cause the chatbot system to determine a classification result for the utterance and one or more anchors, each anchor of the one or more anchors corresponding to one or more anchor words of the utterance. For each anchor of the one or more anchors, one or more synthetic utterances are generated, and one or more classification results for the one or more synthetic utterances are determined. A report including a representation of a particular anchor of the one or more anchors is generated by the chatbot system, the particular anchor corresponding to the highest confidence value among the one or more anchors. The one or more synthetic utterances may be used to generate a new training dataset for training a machine learning model. The training dataset may be refined according to a threshold confidence value to filter out datasets for training.
Owner:ORACLE INT CORP

A loan purpose text classification method and system based on a convolutional neural network

The application discloses a loan purpose text classification method based on a convolutional neural network, which comprises the following steps: S1, establishing a green loan identification text classification model based on a text convolutional neural network, wherein the green loan identification text classification model based on the text convolutional neural network is an end-to-end text classification model, which directly learns relevant features of a green loan from original text and identifies green loan purpose text based on the relevant features; and S2, identifying whether a loan purpose text is a green loan according to the green loan identification text classification model based on the text convolutional neural network. Corresponding systems, electronic devices and computer readable storage media are also disclosed. By constructing an end-to-end green loan identification text classification model based on TextCNN, the relevant features of a green loan are directly learned from original text, and efficient and accurate identification of green loan purpose text is realized.
Owner:BEIJING DADAO ZHIJIAN TECH CO LTD

An unsupervised sentence learning method based on redundancy information reduction

The present application relates to a kind of unsupervised sentence learning method based on redundancy information reduction, belong to natural language processing technical field.The present application focuses on the redundancy information identification and reduction problem in sentence embedding Sentence Embedding technology, aims at optimizing the quality of sentence semantic representation by the way of unsupervised learning.This method is applicable to text classification, information retrieval, semantic similarity calculation and other tasks, and provides an efficient solution for low-resource language and cross-language application scenarios.The technical scheme of the present application combines high-frequency vocabulary analysis, dynamic dimension screening and contrast learning regularization strategy, which can effectively alleviate the over-smoothing problem and redundant information representation problem existing in the sentence embedding process of pre-training model.By constructing, screening and separating redundant information from Token level, this method significantly improves the distinguishing ability of model in semantic feature capture, and provides a new research idea and technical means for sentence representation learning under unsupervised learning framework.
Owner:YUNNAN POWER GRID CO LTD +1

Few-sample hierarchical text classification method based on pre-training language model

The invention relates to the technical field of hierarchical text classification, and particularly provides a few-sample hierarchical text classification method based on a pre-training language model. The method comprises the following steps: generating a prompt of an adaptive task and a classification feature representation through a dynamic prompt generator; a semantic sharing vocabulary mapper is adopted, and semantic consistency is kept through coding father-son relations and cross-level label dependence; the dynamic weighting loss function is utilized, the loss contribution is adjusted according to sample difficulty, label hierarchy and category imbalance, the problem of hierarchical semantic confusion is solved, semantic sharing between labels is achieved, and dependence of a model on hierarchical labeling is reduced.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)

Knowledge-Enhanced Document-Label Attention Method for Multi-Label Text Classification

This invention provides a knowledge-enhanced document-tag attention method for multi-tag text classification. First, it innovatively mines and selects external knowledge from multi-tag documents to enrich document content, and jointly encodes and trains documents and knowledge to improve the interactivity of latent semantics between documents and knowledge. Simultaneously, it embeds the constructed tag sets to capture the contextual relationships between the tag sets corresponding to each document. Then, based on a document-knowledge-tag global attention mechanism, a weighted attention mechanism is applied to document-tag pairs and knowledge-tag pairs to fuse global information between documents, knowledge, and tags, assigning weights to obtain dependent and independent tag representations, thereby capturing the interaction features between documents, knowledge, and tag sets respectively. Finally, all tags for each document are predicted based on the global representations of documents, knowledge, and tags. This method solves the problems of insufficient document richness and tag dependency in multi-tag text classification.
Owner:FUDAN UNIVERSITY

Text classification method and device, electronic equipment and storage medium

The application discloses a text classification method and device, electronic equipment and a storage medium. The method comprises the following steps: obtaining a first text; preprocessing the first text to obtain a word vector included in the first text and at least one word vector corresponding to the word vector; performing weighted fusion processing on each word vector in the word vector and the at least one word vector based on a lattice-LSTM model to obtain a word lattice vector of each position corresponding to each word vector; performing semantic aggregation on the word lattice vector based on a capsule network model to obtain a first semantic vector of each position; determining a category label corresponding to the first semantic vector of each position, and taking the category label as the classification of the first text.
Owner:CHINA MOBILE COMM LTD RES INST +1

A text classification method, device, computer equipment and storage medium

Embodiments of the present application provide a text classification method and device, computer equipment and a storage medium, wherein the method comprises: obtaining a text to be classified; performing text analysis on a first character set included in the text to be classified to obtain a first vector corresponding to the text to be classified; performing text analysis on a second character set included in the text to be classified to obtain a second vector corresponding to the text to be classified; the lengths of the characters included in the first character set and the characters included in the second character set are different; performing analysis on the text to be classified according to a reference vector set to obtain an auxiliary vector corresponding to the text to be classified, the reference vector set being obtained according to the text to be classified and a plurality of reference texts associated with the text to be classified; and performing classification processing on the text to be classified based on the first vector, the second vector and the auxiliary vector to obtain a target category to which the text to be classified belongs, thereby improving the accuracy of text classification.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method and apparatus for training a generative model, generating training samples for a text classifier

ActiveCN116484968Bresolve confusionPattern recognitionSemantic vector
Embodiments of the present specification provide a method and apparatus for training a generation model and generating training samples for a text classifier. In the method for training the generation model, first processing and second processing are performed on a first text sample. The first processing includes determining a semantic vector of the first text sample by a first encoder. A first category of the first text sample is predicted based on the semantic vector by a text classifier, and a first prompt text corresponding to the first category is constructed. The second processing includes determining a first discrete vector corresponding to the first text sample in a target vector space by a second encoder. A reconstructed text of the first text sample is determined based on the first prompt text and the first discrete vector by a decoder. The generation model is trained based on a reconstruction loss determined based on the first text sample and the reconstructed text.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Text classification method and device, computer device and computer readable storage medium

The application discloses a text classification method and device, computer equipment and a computer readable storage medium, and belongs to the technical field of artificial intelligence. The application fully obtains the relationship information of entities and concepts in the target text by representing the association relationship between the entities and the concepts corresponding to the target text by applying a semantic graph, determines first classification information based on the semantic graph, directly determines second classification information based on the context information of the target text, and determines the category to which the target text belongs in combination with the first classification information and the second classification information, that is, in the text classification process, the information of the relationship between the entities in the target text and the context of the target text is comprehensively considered, the category to which the target text belongs is determined based on more comprehensive text information, and therefore the accuracy of the text classification result is effectively improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Text classification model training method, text classification method and device

The present disclosure provides a text classification model training method, a text classification method and a device, and relates to the technical field of artificial intelligence, in particular to the technical field of natural language processing, machine learning and the like. The implementation scheme is as follows: a sample set is obtained; the parameters of a first text classification model are adjusted at least once based on the sample set to obtain a trained second text classification model, each adjustment comprising: adjusting the parameters of the current first text classification model using a first subset of the current sample set to obtain an adjusted text classification model; determining a first evaluation value of each first output category of the adjusted text classification model using a second subset of the current sample set; in response to the first evaluation value of any first output category being less than a threshold value, deleting samples with a category label of the first output category from the current sample set; and in response to the first evaluation values of all first output categories being greater than or equal to the threshold value, determining the adjusted text classification model as the second text classification model.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD