Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

94 results about "Text categorization" patented technology

Text categorization (a.k.a. text classification) is the task of assigning predefined categories to free-text documents. It can provide conceptual views of document collections and has important applications in the real world.

Student model training method and text classification system based on pre-trained language model

A method for training a student model based on a pre-trained language model (PLM) and a text classification system are disclosed. The method includes: constructing cue-based training samples; adjusting the pre-trained language model using the cue-based training samples to obtain a cue-adjusted teacher model; and training the student model using the processed training samples, wherein during training, the student model simultaneously learns the classification probability vectors output by the cue-adjusted teacher model and the original teacher model. This invention requires the student model to learn from two teacher models simultaneously, thereby alleviating the overfitting problem of the student model in small-sample scenarios by adding a distillation path that learns from the original PLM teacher model with unsupervised data. Furthermore, by transferring the intermediate layer representation of the PLM through knowledge probes and stabilizing the performance of knowledge distillation through comparative learning, the student model can learn higher-order dependencies from the intermediate layer representation of the teacher model, improving the accuracy and efficiency of knowledge distillation.
Owner:ALIBABA (CHINA) CO LTD

Text classification method and device, electronic equipment and storage medium

The application discloses a text classification method and device, electronic equipment and a storage medium. The method comprises the following steps: obtaining a first text; preprocessing the first text to obtain a word vector included in the first text and at least one word vector corresponding to the word vector; performing weighted fusion processing on each word vector in the word vector and the at least one word vector based on a lattice-LSTM model to obtain a word lattice vector of each position corresponding to each word vector; performing semantic aggregation on the word lattice vector based on a capsule network model to obtain a first semantic vector of each position; determining a category label corresponding to the first semantic vector of each position, and taking the category label as the classification of the first text.
Owner:CHINA MOBILE COMM LTD RES INST +1

Method and apparatus for training a generative model, generating training samples for a text classifier

ActiveCN116484968Bresolve confusionPattern recognitionSemantic vector
Embodiments of the present specification provide a method and apparatus for training a generation model and generating training samples for a text classifier. In the method for training the generation model, first processing and second processing are performed on a first text sample. The first processing includes determining a semantic vector of the first text sample by a first encoder. A first category of the first text sample is predicted based on the semantic vector by a text classifier, and a first prompt text corresponding to the first category is constructed. The second processing includes determining a first discrete vector corresponding to the first text sample in a target vector space by a second encoder. A reconstructed text of the first text sample is determined based on the first prompt text and the first discrete vector by a decoder. The generation model is trained based on a reconstruction loss determined based on the first text sample and the reconstructed text.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Complaint text classification method and device, computer device and storage medium

The application relates to a customer complaint text classification method and device, computer equipment and a storage medium. The method comprises the following steps: preprocessing text data to be processed to obtain a plurality of word segmentation results; performing vectorization processing on the plurality of word segmentation results to obtain first vector values of the word segmentation results; inputting the plurality of word segmentation results into a pre-trained topic analysis model to obtain second vector values of each customer complaint topic to which the text data to be processed belongs and third vector values of keywords of the text data to be processed; performing splicing processing on the first vector values, the second vector values and the third vector values to obtain splicing features, and taking the splicing features as features of the text data to be processed; and inputting the features of the text data to be processed into a topic classification model and a perception classification model of a pre-trained classification model respectively to obtain a customer complaint topic category and a customer complaint perception category to which the text data to be processed belongs. The method can deeply analyze specific customer complaint contents under a field category.
Owner:SHANGHAI PUDONG DEVELOPMENT BANK

Text classification method and device, electronic equipment and computer readable storage medium

The application provides a text classification method and device, electronic equipment and a computer readable storage medium. The method comprises: clustering each text data in a text data set according to each category number in a plurality of pre-set category numbers, to obtain a category to which each text data belongs under different category numbers; wherein each text data corresponds to one belonging category under one category number; dividing the plurality of text data into a plurality of groups of text data based on the category to which each text data belongs under different category numbers, and the same group of text data belongs to the same category under each category number; and determining a category division result of each group of text data according to the category to which each group of text data belongs under different category numbers. According to the embodiments of the application, the accuracy of the text classification result can be improved.
Owner:MASHANG CONSUMER FINANCE CO LTD

Detection of Sensitive Information in a Text Document

An apparatus (300) for detecting sensitive information in a first text document representative of a first topic is provided. The apparatus (300) is configured to generate a first updated text document by tagging a segment of text in the first text document using a list of one or more types of sensitive information for a second topic: train a language model on text representative of the first topic and on a list of one or more types of sensitive information for a third topic, wherein the language model is a transformer-based machine learning model; and generate a second updated text document by classifying as sensitive a segment of text in the first updated text document using the trained language model representative of relationships between the tagged segment, one or more types of sensitive information for the third topic, and the text representative of the first topic.
Owner:TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)

An agricultural knowledge large model training data construction method, system and device

The application provides an agricultural knowledge large model training data construction method, system and device, the method comprises the following steps: S1, constructing a Mamba classifier based on an agricultural text data classification standard and performing text classification; S2, constructing a text structured rule generation model based on the text classification result and performing knowledge extraction; S3, constructing agricultural large model training data based on the knowledge extraction result; S4, evaluating the agricultural large model training data and adjusting the Mamba classifier and the text structured rule generation model based on the evaluation result. The application can accurately match the diversified training needs of the agricultural large model, ensure that the data is highly adapted to the training needs of the agricultural field through multi-dimensional data screening, and thus provide reliable data for large model multi-scene training tasks.
Owner:CHINA AGRI UNIV

Image classification method, image classification apparatus, electronic device, and storage medium

Embodiments of the present application provide an image classification method, an image classification device, an electronic device and a storage medium, and belong to the technical field of artificial intelligence. The method comprises: obtaining a target image to be processed; performing text recognition on the target image to obtain original text data; performing structured processing on the original text data according to a preset algorithm to obtain line text data; performing classification processing on the line text data through a preset text classification model to obtain first classification data; performing classification processing on the original text data through a preset regular matching mode to obtain second classification data; and obtaining target classification data according to the first classification data and the second classification data. The embodiments of the present application can improve the accuracy of image classification.
Owner:CHINA PING AN LIFE INSURANCE CO LTD

Machine learning-based text classification

A system and method include training a classification model to classify data based on first data associated with a first usage scenario, receiving second data associated with a second usage scenario inputting the second data to the classification model and receiving a likelihood of a first classification from the classification model, determining a similarity between the second data and a plurality of data associated with the second usage scenario, modifying the likelihood based on the determined similarity, determining a second classification of the second data based on the modified likelihood, and processing the second data according to the second classification of the second data.
Owner:SAP SE

A network information collection method and system based on real-time text analysis

The application discloses a network information collection method and system based on real-time text analysis, comprising the following steps: acquiring real-time text data, preprocessing the real-time text data, acquiring first data and second data according to the preprocessed real-time text data, calculating the similarity of the first data and the similarity of the second data, weighting the similarity of the first data and the similarity of the second data to obtain a classification target, constructing a text classification model according to the classification target, inputting the real-time text data into the text classification model to obtain classification data, and outputting the classification data as collected network information. The method can not only improve the accuracy of network information collection, but also has good interpretability and can be directly applied to a network information collection system.
Owner:CHINA NAT INST OF STANDARDIZATION

A text classification method and system based on multi-task learning

This invention belongs to the field of natural language processing and artificial intelligence technology, and discloses a text classification method and system based on multi-task learning. Existing text classification methods suffer from problems such as isolated task processing and inconsistent category order during training and inference in campus public opinion scenarios. This invention uses a shared encoder to extract semantic features of the text, which are then processed by a subsequent regularization module. At least two parallel task classification modules simultaneously complete topic classification and sentiment classification. This multi-task text classification model uses joint loss for optimized training and introduces a category alignment mechanism to ensure deployment stability. This invention significantly improves processing efficiency and reduces resource consumption by outputting multi-dimensional results in a single inference by a single model; it achieves joint performance optimization by utilizing task relevance; and it fundamentally avoids prediction errors through the category alignment mechanism, greatly enhancing the reliability and practicality of the system.
Owner:FUJIAN AGRI VOCATIONAL & TECH COLLEGE

Semantic-based text classification method and device, computer device and storage medium

ActiveCN117874234BEnhance feature expressionImprove classification efficiencyData setLinguistic model
This application belongs to the fields of artificial intelligence and finance, and relates to a semantic-based text classification method. The method includes inputting a training sample set and a knowledge graph into a knowledge-enhanced language model to obtain knowledge-enhanced text semantic feature vectors; inputting the text semantic feature vectors into a capsule network model to output classification prediction results; calculating the loss value between the predicted classification result and the classification label; adjusting the model parameters based on the loss value to output a model to be validated; validating the model to be validated using a test sample set to obtain a text semantic classification model; and inputting the text to be classified into the text semantic classification model for classification. This application also provides a semantic-based text classification device, computer equipment, and storage medium. Furthermore, this application relates to blockchain technology, allowing the classified text dataset to be stored in the blockchain. This application can effectively identify multi-labeled text, improving the efficiency and accuracy of text classification.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

A video dialogue style recognition method based on multi-modal emotion fusion

ActiveCN119068561BMedicineText categorization
A video dialogue style recognition method based on multi-modal emotion fusion is used to predict and identify the dialogue style of characters in a movie clip: different feature extraction models are used to extract visual, auditory and text features from the video, and a pre-trained multi-modal emotion model is used to extract visual emotion features, auditory emotion features and text emotion features; a multi-head attention mechanism is used to fuse visual features with visual emotion features, auditory features with auditory emotion features, and text features with text emotion features; the processed emotional visual features, emotional auditory features and emotional text features are input into the corresponding classification network to obtain visual classification results, auditory classification results and text classification results; finally, these results are fused to obtain the final dialogue style prediction result.
Owner:NANJING UNIV

A personalized learning path recommendation method based on reinforcement learning

ActiveCN116521997BSolve the problem of poor learning resultsPersonalized learningText categorization
The application provides a kind of personalized learning path recommendation method based on reinforcement learning, it is related to educational data mining technical field.The method first constructs learner simulator according to the learning record of scholar, the simulator can judge the learning level of learner;Then the knowledge relationship graph between the knowledge points contained in the exercise is automatically constructed by the concept map automatic construction model based on text classification and association rule mining;Based on the knowledge relationship graph and the cognitive diagnosis model, an exercise navigation module is designed to select potential candidate exercises;After the reinforcement learning agent selects the action in the action space, the state transition is determined in the state space, the model parameters are updated according to the loss function and optimization strategy, and the reinforcement learning model is optimized;Finally, the designed reinforcement learning model is used to recommend exercises for learners, and the reinforcement learning model parameters are updated according to the learning situation of learners.The method can recommend efficient and reasonable learning path for learners.
Owner:NORTHEASTERN UNIV CHINA

A text classification method based on document similarity

PendingCN122285904AFeature vectorDocument similarity
This invention discloses a text classification method based on document similarity, belonging to the field of document processing technology. The method iteratively optimizes the word segmentation mechanism to update the word set of document sentences and generates a first sentence vector. A first similarity vector is generated by combining sentence weights and the first sentence vector. Multiple topic words are extracted from the word set based on entity weights to generate topic feature vectors. Knowledge supplement vectors are extracted from a knowledge base based on the context vectors of each topic word and entity weights. A topic enhancement vector is generated by combining the topic feature vector and the knowledge supplement vector. A second similarity vector is generated based on the topic enhancement vector. The first and second similarity vectors are then optimized through difference comparison to obtain the document similarity vector, and the text category is output. This invention, through iterative word segmentation optimization, topic-guided knowledge enhancement, and a similarity comparison mechanism, can effectively improve the reliability and accuracy of text classification.
Owner:JIANGXI MECHANICAL & ELECTRICAL VOCATIONAL & TECH COLLEGE

Text classification method and device, electronic equipment, storage medium and program product

This application relates to the field of computer technology, providing a text classification method, apparatus, electronic device, storage medium, and program product. The method includes: inputting text to be classified into a text classification model to obtain an initial predicted probability distribution of text categories output by the text classification model; outputting an initial classification result based on the text category corresponding to the highest predicted probability in the initial predicted probability distribution; receiving feedback information based on the initial classification result; correcting the initial predicted probability distribution based on the feedback information to obtain a target predicted probability distribution; and outputting a target classification result based on the text category corresponding to the highest predicted probability in the target predicted probability distribution. This application incorporates user feedback information to correct the initial predicted probability distribution, which can increase the probability of potential correct text categories based on the initial predicted probability distribution, thereby improving the accuracy of the final classification result.
Owner:MIDEA GROUP CO LTD

Hierarchical text classification method based on label guided semantic interaction and multi-dimensional contrast learning

The application belongs to the field of text classification, and discloses a hierarchical text classification method based on label guided semantic interaction and multi-dimensional contrast learning, proposes a label guided semantic interaction module, adopts a multi-head attention mechanism to model dynamic semantic interaction between a text and a label, and then generates unique context-aware embedding representation for each label. At the sample level, a hierarchical hard negative sample construction method is proposed, the parent-child relationship and sibling relationship in the label hierarchy are used to construct hard negative samples, and then the representation quality of long-tail labels is improved. At the label level, a hierarchical distance-aware optimization method is proposed, the embedding distance between labels is adaptively adjusted based on the label hierarchy structure, and then the discrimination ability of label embedding is improved.
Owner:ZHEJIANG UNIV OF FINANCE & ECONOMICS +1

Text classification method and device, computer device and storage medium

The application relates to a text classification method and device, computer equipment, a storage medium and a computer program product. The method comprises the following steps: obtaining a text to be classified, obtaining a target topic feature vocabulary; calculating a text feature vector corresponding to the text to be classified; obtaining a topic feature vector corresponding to each candidate topic, the topic feature vector being calculated based on a feature word set corresponding to the candidate topic; calculating the similarity between the text feature vector to be classified and each topic feature vector; obtaining a target topic feature vector corresponding to the text to be classified based on the similarity; and obtaining a candidate topic corresponding to the target topic feature vector as a target topic corresponding to the text to be classified. The method can improve the accuracy of short text classification.
Owner:ZHAOLIAN CONSUMER FINANCE CO LTD +1

A text classification method and system based on guided map contrast learning

The application provides a text classification method and system based on guided graph contrastive learning. The text classification method comprises: performing upper and lower semantic modeling on an input original text, converting a discrete text sequence into a continuous vector representation; mining potential theme semantic structures in the text based on a neural network to obtain distribution characteristics of the document in a theme space; representing the text data in a graph structure form to depict multi-level semantic relationships and structural relationships among the document, words and themes; generating multiple enhanced text graph views with consistent semantics but different structures through guided disturbance on the graph structure; and training in a graph contrastive learning mode to maximize the consistency of the representation of the same document under different enhanced text graph views and to reduce the distance between the representations of different documents. The method significantly improves the generalization performance and robustness of the model in the text classification task and reduces the dependence on artificial feature design and large-scale labeled data.
Owner:INNER MONGOLIA UNIVERSITY

Text classification method and device, storage medium and electronic equipment

PendingCN122132563ASemantic analysisBiological modelsMulti-label classificationText categorization
Embodiments of the present specification disclose a text classification method and device, a storage medium and an electronic device. After a first label set corresponding to a text to be classified is recalled, a second label set of the text to be classified is obtained through a preset co-occurrence probability matrix. In a selection stage for a target label corresponding to the text to be classified, a multi-label classification problem is converted into a matching problem between a label and the text to be classified. The first label set recalled and the second label set obtained through the co-occurrence probability matrix are matched with the text to be classified for binary identification, so as to finally determine the target label corresponding to the text to be classified.
Owner:CHONGQING ANT CONSUMER FINANCE CO LTD

A power technology standard entity relationship extraction method based on multi-round automatic question answering

This invention proposes a method for extracting entity relations from power technology standards based on multi-turn automatic question answering, comprising: Step 1. Constructing a corpus of power technology standards; Step 2. Constructing information extraction element templates for different types of power technology standards corpus; Step 3. Constructing a question-and-answer corpus of power technology standards based on the information extraction element templates; Step 4. Constructing a text classification module for the power technology standards corpus described in Step 1; Step 5. Automatically constructing a multi-turn question-and-answer and question module by matching the text classification results with the information extraction element templates in Step 4. This module breaks down complex questions into simpler ones and provides step-by-step reasoning answers; Step 6. Constructing a machine reading comprehension module for the power technology standards question-and-answer corpus, which provides step-by-step reasoning answers to the questions automatically constructed in Step 5, thus completing the extraction method of this invention. This invention can effectively alleviate the phenomena of overlapping relations and cross-sentence dependencies in complex texts.
Owner:STATE GRID LIAONING ELECTRIC POWER CO LTD +2

Text classification method, apparatus, device, medium, and product

The present disclosure relates to the technical field of machine learning, and particularly relates to a text classification method, device, equipment, medium and product. The method comprises: obtaining a word vector of text data to be classified; extracting a first text feature of the text data to be classified based on the word vector; extracting a second text feature of the text data to be classified based on the word vector; extracting a third text feature of the text data to be classified based on the first text feature; obtaining a fusion feature by calculating a gating score of the first text feature, the second text feature and the third text feature; and obtaining a classification result of the text data to be classified based on the fusion feature. By extracting the first text feature, the second text feature and the third text feature, long-distance dependency features, key local features and context attention features of the text data to be classified can be obtained at the same time, which is conducive to improving the processing capability of the text classification model for long text and the understanding capability of the text classification model for complex semantic structures, and improving the classification accuracy.
Owner:CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1

Active few-shot learning method and system for chinese text classification

PendingCN122364450AText categorizationInformation quantity
This invention relates to the field of natural language processing technology, and particularly to an active few-shot learning method and system for Chinese text classification. The method comprises the following steps: S1: Constructing an unlabeled sample pool; S2: Employing an active few-shot sampling strategy based on "hierarchical clustering uncertainty-diversity" to iteratively select information-rich candidate samples from the unlabeled pool; S3: Manually labeling the candidate samples and adding them to the support set; S4: Fine-tuning the pre-trained language model using the updated support set; S5: Repeating steps S2 to S4 until a preset labeling budget is reached or the model performance meets the requirements. The active few-shot sampling strategy includes: hierarchically dividing the unlabeled pool based on pseudo-labels, and performing clustering uncertainty sampling incorporating diversity constraints within each category layer. This invention effectively solves the semantic redundancy and class imbalance problems in Chinese few-shot text classification, significantly improving model performance and accelerating convergence under low labeling budgets.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

Text classification model training method, text classification method, apparatus, device, storage medium and computer program product

The disclosure provides a text classification model training method, a text classification method, an apparatus, an electronic device, and a computer-readable storage medium, and relates to artificial intelligence technology. The text classification model training method includes: performing machine translation on a plurality of first text samples in a first language to obtain a plurality of second text samples in a second language different from the first language; training a first text classification model for the second language based on a plurality of third text samples in the second language and corresponding class labels; performing confidence-based filtering on the plurality of second text samples by the trained first text classification model; and training a second text classification model for the second language based on the filtered second text samples.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Oil and gas exploration business process generation method and device

This invention discloses a method and apparatus for generating oil and gas exploration business processes. The method includes: acquiring selection results of oil and gas exploration data and exploration business requirements; breaking down the exploration business requirements and identifying keywords; inputting the identified keywords sequentially into a text classification model and outputting classification results; searching a module library for modules whose keyword matching probability with the module name is greater than a preset probability based on the classification results; determining the input-output relationship of the modules; configuring parameters in the modules that are not pre-configured based on the selection results and exploration business requirements; and generating and displaying an oil and gas exploration business process based on the module's input-output relationship and the selected oil and gas exploration data, using the configured module as a node. This method can improve the generation efficiency of oil and gas exploration business processes and the accuracy of oil and gas exploration data processing, and can quickly respond to changes in exploration business requirements.
Owner:RICHFIT INFORMATION TECH +1