Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

41 results about "Categorical models" patented technology

Classification model training method, device and equipment based on large language model

The invention provides a classification model training method, device and equipment based on a large language model, and relates to the technical field of natural language processing and reinforcement learning. The method comprises the steps of performing classification prediction on a first training set through a base large language model to generate an initial prediction category, and performing fine adjustment on the base large language model according to the initial prediction category and a corresponding real category to obtain a fine-adjusted large language model; through the fine-tuned large language model, screening difficult sample texts which are wrongly classified from the first training set, and constructing a second training set based on the difficult sample texts; and based on the second training set, a reinforcement learning strategy is adopted to update strategy parameters in the fine-tuned large language model, and a target classification model is obtained.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Prompt engineering and in-context example selection for large language models

Various embodiments of the present disclosure provide prompt engineering and iterative, feedback-based generative techniques that improve traditional LLM technology, including extractive LLM techniques. The techniques may include generating, using a machine learning classification model, a resolution capability classification for an input data object that comprises an input question and an input document; generating using a large language model (LLM), an initial predictive output for the input data object based on an initial generative model prompt for the input data object; identifying a classification model output divergence based on a comparison between the resolution capability classification and the initial predictive output; and in response to the classification model output divergence: generating an augmented generative model prompt by modifying the initial generative model prompt with a representation of the resolution capability classification, and generating, using the LLM, an updated predictive output based on the augmented generative model prompt.
Owner:OPTUM INC

Detection and prevention of adversarial attacks against large language models

Systems, methods, and apparatuses are disclosed for detection and prevention of adversarial attacks against large language models. Techniques may include receiving an input associated with a target large language model, analyzing the input with a pre-trained classification algorithm to determine a first deconstruction process to be applied to the input, and modifying the input with a first deconstruction model using the determined first deconstruction process. Techniques may also include determining a score of a likelihood of the input being adversarial based on an output of the first deconstruction model and by applying a classification model and updating at least one of the first deconstruction model or the classification model based on the score.
Owner:CYBER ARK SOFTWARE LTD

Emotion classification method based on visual language model and conditional reasoning

The invention provides a sentiment classification method based on a visual language model and conditional reasoning, which comprises the following steps of: firstly, acquiring a text picture pair and labeling sentiment labels on the text picture pair to form a sentiment label set; then, a visual language model is used as a strategy model, general reasoning and conditional reasoning are carried out on the text picture pairs respectively, reasoning characterization, emotion prediction labels and conditional reasoning results are generated, and response samples are formed after the reasoning characterization, the emotion prediction labels and the conditional reasoning results are combined; and calculating a reward value and an advantage estimation value based on the response sample to optimize the strategy model, finally utilizing the optimized strategy model to carry out sentiment prediction on a new text picture pair, and outputting a final sentiment classification result. According to the method, the defect that a traditional multi-classification model is easily interfered by noise texts or complex visual contents is overcome, the classification precision is improved, a general reasoning process and a conditional reasoning process of a strategy model are recombined into a group of response samples, it is guaranteed that each group of response contains different classification labels, and the problem of advantage collapse is solved.
Owner:HANGZHOU DIANZI UNIV

Big language model illusion detection method and system

The invention discloses a big language model illusion detection method and system, and relates to the technical field of natural language processing and artificial intelligence. Comprising the steps that 1, a question and answer data set is selected, twenty internal state vectors generated when a large language model generates answers to an input question are extracted, and the internal state vectors are semantic coding vectors of a hidden layer in the model reasoning process; 2, the twenty internal state vectors are constructed into a sample matrix, and a covariance matrix of the sample matrix is calculated; 3, calculating a standardized determinant of the covariance matrix, wherein the standardized determinant is a twentieth-power root of a determinant value of the covariance matrix; 4, taking the obtained standardized determinant value as a feature F1, and taking the token number of the answer output by the large language model as a feature F2; 5, training a support vector machine dichotomy model by using the question and answer data set, wherein labels of training samples are illusion or non-illusion; and step 6, inputting F1 and F2 into the trained support vector machine dichotomy model, and outputting a hallucination or non-hallucination detection result.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Psychological state text classification method based on large model generative data enhancement

The invention discloses a psychological state text classification method based on large model generative data enhancement, and relates to the technical field of natural language processing. According to the method, patient psychological semantic clusters are automatically found by clustering a limited number of original samples, and then a natural language semantic template of each patient psychological semantic cluster is extracted; then, under the dual control of emotional polarity and psychological themes by utilizing a large language model, according to the natural language semantic templates, psychological state texts with consistent themes and diversified expressions are generated as enhanced samples; according to the method, high-quality sample generation is carried out by utilizing a natural language semantic template obtained based on clustering through a large language model so as to expand a model training sample, and the authenticity and diversity of the generated enhanced sample can be ensured through the large model generation type data enhancement method; therefore, the classification accuracy and generalization ability of the psychological state text classification model obtained through training are improved.
Owner:JIANGNAN UNIV

Explanatable text classification method based on large model concept generation

The invention provides an interpretable text classification method based on large model concept generation, and relates to the field of natural language processing. In the training stage, firstly, stable and interpretable concept features and corresponding dimensions are extracted and marked for each sample through a large model, so that a task-aware concept system is constructed; and then coding each labeled sample to carry out prediction modeling to obtain a text classification model with high interpretability. In the reasoning stage, firstly, experience samples are screened for new samples on the basis of a screening strategy combining dual semantic consistency and sample diversity enhancement; and then marking concepts of new samples and dimensions corresponding to the concepts through a large model by utilizing a task-aware concept system and an experience sample, and then feeding back a marking result to a text classification model to finish final prediction. According to the method, the processing efficiency and the prediction accuracy of the text classification task are improved, the transparency, the stability and the business controllability of classification model output are enhanced, and application scene requirements with relatively high interpretability requirements are met.
Owner:HEFEI UNIV OF TECH +1

Model drift detection techniques

Techniques for detecting machine learning model drift are described. Model drift can result in the model misclassifying inputs. A system for detecting drift in natural language processing (NLP) models involves determining high-dimensional embeddings of inputs and high-dimensional embeddings of training samples, reducing the high-dimensional embeddings to low-dimensional embeddings, and comparing the low-dimensional embeddings to determine whether the inputs are statistically different than the training samples. When the inputs are statistically different than the training samples, model drift is detected, and retraining of the model may be performed. The system can detect drift in other classification models as well and can process with respect to other types of inputs (e.g., audio, image, etc.).
Owner:AMAZON TECH INC

Detection and prevention of adversarial attacks against large language models

PendingUS20260189585A1AlgorithmCategorical models
Systems, methods, and apparatuses are disclosed for detection and prevention of adversarial attacks against large language models. Techniques may include receiving an input associated with a target large language model, analyzing the input with a pre-trained classification algorithm to determine a first deconstruction process to be applied to the input, and modifying the input with a first deconstruction model using the determined first deconstruction process. Techniques may also include determining a score of a likelihood of the input being adversarial based on an output of the first deconstruction model and by applying a classification model and updating at least one of the first deconstruction model or the classification model based on the score.
Owner:CYBER ARK SOFTWARE LTD

Machine learning-based text classification

A system and method include training a classification model to classify data based on first data associated with a first usage scenario, receiving second data associated with a second usage scenario inputting the second data to the classification model and receiving a likelihood of a first classification from the classification model, determining a similarity between the second data and a plurality of data associated with the second usage scenario, modifying the likelihood based on the determined similarity, determining a second classification of the second data based on the modified likelihood, and processing the second data according to the second classification of the second data.
Owner:SAP SE

A case retrieval method based on multiple models

The present application relates to the field of artificial intelligence, in particular to a case retrieval method based on multiple models. Various multi-source data are collected and integrated; the semantic dependency relationship between the text fragments of the contradiction is captured by using Bert, and the similarity of the case data is calculated by using the BM25 algorithm; a model is trained by Bert according to twelve types of classifiers, and at the same time, a model is retrained for the overall similarity model training. The similarity model constructed by combining the case classification model in the prior art and the present application completes the accurate case retrieval of diversified and long-length contradiction dispute data. The legal case data of the present application is more specific and perfect, contains sufficient legal knowledge, can cope with the change of legal rules, the unification of diversified legal data, accurate collection, efficient auxiliary analysis of legal cases and resolution work. The present application is more accurate in calculating the similarity, and the relevant category of legal cases is recommended accurately.
Owner:UNIV OF SCI & TECH OF CHINA

Hierarchical fine-grained human trajectory activity type inference method and related device

The invention belongs to the field of urban calculation and intelligent transportation, and discloses a hierarchical fine-grained human trajectory activity type inference method and related devices.The method comprises the steps that firstly, rigid activity types are recognized through predefined anchor point rules, and efficient labeling of regular segments is ensured; subsequently, the remaining staying segments mark transactional or leisure motivations through a motivation classification model to differentiate internal heterogeneity of non-rigid activities; and finally, performing structured inference on the marked fragment by a specific large language model of the motivator, and outputting a fine-grained non-rigid activity type. By the adoption of the method, the overall accuracy and robustness of fine-grained activity inference are remarkably improved, the performance is optimized especially on identification of non-rigid activities, adaptability to complex multi-constraint scenes is enhanced, and reliable fine-grained classification requirements are supported.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Utilizing a large language model (LLM) to automatically construct a machine learning (ML) classification model

A computerized method includes: obtaining a first dataset of pre-labeled textual items, wherein each pre-labeled textual item is associated with a pre-label; feeding each of the pre-labeled textual items into a Large Language Model (LLM), and prompting it to generate textual reasoning that supports the pre-label of each pre-labeled textual item; collating the generated textual reasonings, and generating therefrom a textual instruction prompt; obtaining a second dataset of not-yet- labeled textual items; feeding each of the not-yet-labeled textual items into the LLM, and commanding it to utilize the textual instruction prompt and to generate a textual label for each of the not-yet-labeled textual items; collecting those textual items, that were labeled by the LLM, into a third dataset of LLM-labeled textual items; automatically training a Machine Language (ML) classification model on that third dataset of LLM-labeled textual items; deploying that ML classification model in a platform for classification of textual items.
Owner:VARONIS SYSTEMS INC

Psychological assessment method and system based on dialogue content

The invention discloses a psychological assessment method and system based on dialogue content, and relates to the technical field of semantic processing, and the method comprises the steps: obtaining question and answer dialogue data of a user and a system in a multi-round interaction process; encoding each round of question sentences and answer sentences by using a semantic analysis model, and extracting semantic embedding vectors and topic feature vectors; calculating a semantic matching degree between adjacent question and answer pairs, and marking the semantic matching degree as a semantic deviation sample when the matching degree continuously decreases; constructing a topic embedding sequence of multiple rounds of dialogues and counting a fuzzy word proportion to form a topic and a statistical feature vector; inputting the semantic deviation samples, the themes and the statistical features into a machine learning classification model to obtain a psychological situation classifier; and outputting an evaluation result during dialogue operation. The problem that psychological reasons for topic avoidance of users in multiple rounds of dialogues are difficult to quantify in real time is solved.
Owner:ANLIZHI INTELLIGENT ROBOT TECH (BEIJING) CO LTD

Configuration and Training of Classification Models

Methods, systems, devices, and non-transitory computer readable media for training machine-learning models are provided. The disclosed technology can include receiving input samples associated with classification concepts. Based on inputting the input samples into a first plurality of machine-learned models, classification outputs comprising labels and confidence scores can be generated. The first plurality of machine-learned models can comprise one or more multimodal large language models (LLMs) and one or more domain-specific models. Annotated input samples comprising the input samples, the classification outputs, and identifiers that identify each of the first plurality of machine-learned models that generated each of the classification outputs can be generated. Furthermore, based on the annotated input samples, one or more second machine-learned models can be trained. The training can comprise modifying parameters of the one or more second machine-learned models based on the confidence scores.
Owner:GOOGLE LLC

A classification model construction method and system based on brain network normative modeling

This invention belongs to the field of brain network technology and relates to a method and system for constructing a classification model based on standardized brain network modeling. The construction method includes: acquiring electroencephalogram (EEG) signal data of a subject; determining the partial orientation coherence adjacency matrix corresponding to the EEG signal data; processing the partial orientation coherence adjacency matrix according to a first formula to determine the individualized causal directed brain network partial orientation coherence matrix corresponding to the EEG signal data to be tested, wherein the feature path length and network density in the individualized causal directed brain network partial orientation coherence matrix are in an optimal balance state; predicting the individualized directed brain network partial orientation coherence matrix to be tested using a standardized benchmark model, and determining the deviation between the predicted value of the standardized benchmark model and the true value of the individualized directed brain network partial orientation coherence matrix; the standardized benchmark model is a model established based on the brain network adjacency matrix of healthy subjects; and constructing a classification model using the deviation as a classification feature.
Owner:BRAIN-COMPUTER INTERACTION & HUMAN-COMPUTER INTEGRATION HAIHE LAB

Large language model training method and target classification model generation method and device

PendingCN121660025AFinanceBiological modelsLinguistic modelCategorical models
The invention discloses a large language model training method and device and a target classification model generating method and device. The method comprises the steps of determining a target training task and at least one historical training task to be subjected to memory training; obtaining a first synthetic instance set corresponding to the target training task; obtaining a second synthetic instance set corresponding to the historical training tasks in the at least one historical training task, and determining a target synthetic instance set according to the first synthetic instance set and the second synthetic instance set; and training a to-be-trained large language model according to the target synthesis instance set to obtain a target large language model. The technical problem that old knowledge of the model is maintained depending on real training data of a large language model in related technologies, and new and old knowledge of the large language model cannot be considered when the real training data is unavailable is solved.
Owner:ALIBABA CLOUD COMPUTING CO LTD

Sentiment classification and model training method and device, medium, product and equipment

The invention discloses an emotion classification and model training method and device, a medium, a product and equipment, and the method comprises the steps: independently converting obtained text information into first prompt information corresponding to a generation model and second prompt information corresponding to a classification model, the first prompt information is input into a first model containing a generative model to obtain a first feature, so that the first feature can contain rich semantic and context information learned by the generative model, and the second prompt information is input into a second model containing a classification model to obtain a second feature; according to the method, the first feature and the second feature are combined, so that the second feature can contain the preliminary classification information learned by the classification model, and the third feature with richer information content can be obtained through the first feature and the second feature, so that the richness of the feature information received by the prediction model is improved, and the accuracy of an emotion prediction result output by the prediction model is improved.
Owner:CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1

A mental state text classification method based on large model generative data enhancement

ActiveCN121935378BPsychological statusMental state
This application discloses a method for classifying psychological state texts based on large-scale generative data augmentation, relating to the field of natural language processing technology. This method automatically discovers patient psychological semantic clusters by clustering a limited number of original samples, then extracts the natural language semantic template for each patient's psychological semantic cluster. Subsequently, under the dual control of emotional polarity and psychological theme, a large-scale language model is used to generate psychological state texts with consistent themes and diverse expressions according to these natural language semantic templates as augmented samples. This method utilizes the large-scale language model to generate high-quality samples based on the natural language semantic templates obtained from clustering to expand the model's training samples. This large-scale generative data augmentation method can ensure the authenticity and diversity of the generated augmented samples, thereby improving the classification accuracy and generalization ability of the trained psychological state text classification model.
Owner:JIANGNAN UNIV

Supply chain risk quantification method and system based on large language model weak supervised learning

The invention relates to the technical field of resource management, and provides a supply chain risk quantification method based on large language model weak supervised learning. The method comprises the following steps: obtaining a multi-time-period operation disclosure text of a target enterprise and cleaning clauses to form a structured statement set; screening a potential risk statement subset based on the supply chain risk keyword library; calling a large language model to carry out weak supervision labeling to generate a risk category label; training a small risk sentiment classification model by using a label, reasoning a full amount of statements, and outputting a probability value of each risk category; aggregating the probability according to a time period to obtain a semantic distribution vector to represent a risk semantic state; and constructing a semantic residual image for the adjacent periodic vector difference values, and calculating a supply chain risk trend index according to the semantic residual image to realize risk evolution trend quantification.
Owner:GUANGXI UNIVERSITY OF FINANCE AND ECONOMICS

Verb metaphor identification method based on multi-verb extraction and thought chain prompts

A verb metaphor recognition method based on multi-verb extraction and thought chain prompting belongs to the field of natural language processing in deep learning and is used for metaphor recognition of Chinese sentences. The key points are: extracting verbs from texts in a data set; performing thought chain prompting on a large model to obtain a prompt result of whether the verb is a metaphor; splicing the sentence containing the verb in the text with the verb and encoding them to obtain a sentence feature vector; encoding the prompt result as a thought chain feature vector; splicing the sentence feature vector with the thought chain feature vector at a certain weight ratio to obtain a spliced ​​feature vector; inputting the spliced ​​feature vector into a classification model for metaphor judgment to obtain a judgment result of whether the sentence contains a metaphor. The present invention can significantly improve the accuracy and interpretability of metaphor recognition.
Owner:DALIAN UNIV OF TECH

User emotion classification method and device, server, medium and program product

The application relates to the technical field of artificial intelligence, in particular to a user emotion classification method and device, a server, a medium and a program product, the method comprising the following steps: obtaining a word vector, a sentence vector and a syntactic dependency relation graph of corpus of a user, and obtaining a heterogeneous graph between words and sentences in a training corpus sample; inputting the word vector, the sentence vector and the heterogeneous graph into a sentence emotion classification model to obtain a first hidden state vector; inputting the word vector and the syntactic dependency relation graph into a user emotion classification model to obtain a second hidden state vector; and determining an emotion classification result of the user according to the first hidden state vector and the second hidden state vector. The method can improve the prediction accuracy of the emotion category.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Systems and methods for time-series classification through residual learning

PendingUS20250371306A1Neural learning methodsTime series classificationCategorical models
Methods and systems for enhancing the performance of time series classification models through the introduction of a joint residual-classification framework. This framework aims to address class imbalance issues by effectively integrating residuals with classification model embeddings. In embodiments, categorical ground-truth data is converted into continuous data, and a time series forecasting model is trained to predict residuals that are subsequently integrated into the embeddings of a classifier model. This integration facilitates more accurate model predictions by incorporating additional context specific to the data's characteristics.
Owner:ROBERT BOSCH GMBH

Power system semantic framework analysis method and system based on pre-training language model

The invention provides an electric power system semantic framework analysis method and system based on a pre-trained language model, and the method comprises the steps: recognizing a sentence in an electric power system text through a pre-trained framework recognition model, and obtaining a target word in the sentence and a framework type activated by the target word; performing sequence labeling on the sentence by utilizing a pre-trained argument recognition model to obtain an argument range of the target word; utilizing a pre-trained argument role classification model to predict and classify arguments in the sentences based on the frame type and the argument range to obtain semantic role tags of the arguments; determining a power system event contained in the sentence based on a frame type and a semantic role tag corresponding to the argument; according to the method, the key steps of the semantic framework analysis task are subjected to task association, so that the whole semantic analysis process is more accurate and efficient, the finally obtained event analysis result of the electric power system is more accurate, and the decision of the electric power system and the accuracy of related instructions are ensured.
Owner:GLOBAL ENERGY INTERCONNECTION RES INST CO LTD +2

Explainable text classification method based on large model concept generation

The application provides an interpretable text classification method based on large model concept generation, and relates to the field of natural language processing. In the training stage, first, a stable and interpretable concept feature and its corresponding dimension are extracted and labeled for each sample by a large model to construct a task-aware concept system; then each labeled sample is encoded for prediction modeling to obtain a high-interpretable text classification model. In the reasoning stage, first, an experienced sample is selected for a new sample based on a screening strategy combining double semantic consistency and sample diversity enhancement; then the concept and its corresponding dimension of the new sample are labeled by a large model using the task-aware concept system and the experienced sample, and the labeling result is fed back to the text classification model to complete the final prediction. The application improves the processing efficiency and prediction accuracy of the text classification task, enhances the transparency, stability and business controllability of the classification model output, and meets the demand of application scenarios with high requirements for interpretability.
Owner:HEFEI UNIV OF TECH +1

Multi-agent oriented collaborative system ethical risk prediction method, system, device and medium

PendingCN122334973ARisk indicatorCategorical models
This application relates to a method, system, device, and medium for predicting ethical risks in multi-agent collaborative systems. The method includes: acquiring real-time interaction logs of multiple agents in the education field, obtaining a dynamic change sequence of interactions through time-series analysis; dividing time stages based on interaction frequencies exceeding a preset threshold, calculating a stage risk evolution index, and quantifying the mean value by combining it with core ethical risk indicators to generate an educational ethical risk early warning vector; based on this vector, training a Bayesian network using Bayesian estimation to obtain a joint probability distribution of risk evolution; generating risk evolution paths from the distribution using the Monte Carlo method, and identifying propagation path types using a machine learning classification model; inputting this type of time-series data into a pre-trained long short-term memory network, and weightedly fusing it with real-time interaction data to generate educational ethical risk prediction parameters. This method can improve the timeliness and accuracy of risk early warning, avoid related ethical risks, and contribute to the standardized and safe development of artificial intelligence education.
Owner:GUANGZHOU UNIVERSITY

Banking flow classification method and system based on large language model

The application discloses a bank flow classification method based on a large language model, comprising: preprocessing bank flow data to generate standardized samples; inputting the samples into a rule engine for matching, outputting a rule result if successful, and inputting a classification model for prediction if failed; if the model prediction confidence is lower than a threshold, calling a large language model for auxiliary classification; inputting subject information into the large language model to generate a candidate category array, and inputting transaction information and the candidate array into the large language model together, so that the large language model outputs a prediction result in the candidate array; obtaining existing results in the rules, the model and the large language model, calculating a comprehensive score by weighted voting, and selecting the highest score as the final classification result. The application improves the accuracy and adaptability of bank flow classification.
Owner:SHENZHEN ZHIYU INTELLIGENT BUSINESS CO LTD

Text classification-based bidding life cycle and announcement type identification method

The invention discloses a bidding life cycle and announcement type identification method adopting text classification, and the method comprises the steps: obtaining bidding announcement information from an open channel, analyzing an original bidding announcement, carrying out the data preprocessing, extracting key information according to preset rules and labels, balancing the label distribution of sample data, and carrying out the fine adjustment based on ERNIE. According to the method, life cycle and type identification of bidding announcements is performed by combining a rule method and deep learning, categories of multi-source public bidding announcement information are summarized in a unified manner, information integration is facilitated, classification and identification are performed based on the deep learning model, and the classification efficiency is improved. According to the method, complexity and maintenance difficulty of a rule-based method are avoided, only a small part of data is manually labeled for sample data of model training, labor cost is saved, and the accuracy and efficiency of identifying bidding life cycles and announcement types are improved.
Owner:ANHUI ZHIYIXIN INFORMATION TECH CO LTD