Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

18 results about "Predictive text" patented technology

Predictive text is an input technology used where one key or button represents many letters, such as on the numeric keypads of mobile phones and in accessibility technologies. Each key press results in a prediction rather than repeatedly sequencing through the same group of "letters" it represents, in the same, invariable order. Predictive text could allow for an entire word to be input by single keypress. Predictive text makes efficient use of fewer device keys to input writing into a text message, an e-mail, an address book, a calendar, and the like.

A text-guided face spoofing detection method and system

This application provides a text-guided method and system for detecting face forgery. The method includes obtaining a face image to be detected (whether it is genuine or fake), inputting the face image into a face detection model to obtain the face authenticity detection result. The face detection model includes: constructing a text prompt lexicon covering multiple granularities and generating multi-dimensional text prototypes; extracting visual features, optimizing the feature distribution of visual features to obtain global visual features; performing feature separation and enhancement on the global visual features; mapping the global visual features after feature separation and enhancement to predicted text features; applying similarity constraints to obtain cross-modal prototype matching results; applying discriminative constraints on different-dimensional text prototypes to obtain feature measurement learning results; and outputting the detection result based on the cross-modal prototype matching results and feature measurement learning results. This method solves the problem of poor generalization in face forgery detection and improves detection accuracy and generalization ability.
Owner:NANJING UNIV OF POSTS & TELECOMM

End-to-end speech recognition method and system based on keyword attention enhancement mechanism

The invention relates to the technical field of speech recognition, and provides an end-to-end speech recognition method and system based on a keyword attention enhancement mechanism, and the method comprises the steps: extracting entity keywords through a keyword searcher; the voice cache manager receives continuous air traffic control audio clips, converts the continuous air traffic control audio clips into an audio Mel spectrogram and then converts the audio Mel spectrogram into an audio embedded sequence; the keyword encoder unit maps the lexical element embedding representation sequence into a keyword embedding sequence; the audio transliteration decoder unit performs keyword attention enhancement calculation and converts splicing vectors of the audio embedding sequence, the transliteration start mark and the lexical element embedding representation sequence into a prediction text sequence; and inputting the new lexical element embedded representation sequence into a text translator through autoregression until the audio transwriting decoder unit outputs a transwriting end mark, and outputting the predicted text sequence as a speech recognition text. According to the invention, real-time identification of the streaming input voice is realized, and the method has the advantages of high identification precision, low response delay, flexible deployment and the like.
Owner:NANKAI UNIV +1

System and method for artificial intelligence based web form completion

PendingUS20260187356A1Data platformText entry
A computer-implemented method comprises providing previous form data to a training and validation module, filtering the previous data to create a subset of training data meeting a validation threshold, for each of a plurality of web-based input forms training at least one artificial intelligence model relating to at least one field in at least one web-based input form using the subset of training data, deploying at least one model endpoint corresponding to a model based on the trained model and a template of the form, transmitting a web-based input form web page to a client device, wherein the web page is displayed on a user interface, wherein the web page includes an AI-based form field prediction text input portion used for the AI-based form field prediction and an interactive representation of the web-based input form separate and distinct from the free text input portion; receiving, from the client device, a first user input at the AI-based form field prediction text input portion of the web-based input form; requesting a plurality of suggested answers to the web-based input form based on the user input; selecting a respective model endpoint corresponding to the web-based input form; generating at least one predicted user input using the model endpoint and the first user input; modifying the interactive representation of the web-based input form to include the at least one predicted user input; transmitting the modified interactive representation of the web-based input form including the at least one predicted user input to the client device; receiving, from the client device, a second user input at certain fields of the modified interactive representation of the web-based input form relating to a user modification of the suggested answers; and storing the inputs at the fields of the completed web-based input form to the data platform.
Owner:ENABLON SAS

Semantic effect evaluation method and related apparatus

ActiveCN114492461BData miningMachine learning
The application discloses a semantic effect evaluation method and related device, the semantic effect evaluation method comprises: obtaining a to-be-evaluated dialogue; wherein the to-be-evaluated dialogue comprises predicted text related to a user intention; inputting the to-be-evaluated dialogue into a multi-turn dialogue test set, verifying the predicted text of different nodes in the to-be-evaluated dialogue by using the multi-turn dialogue test set; in response to at least one error node with an identification error existing in the to-be-evaluated dialogue, reconstructing content after the error node in the to-be-evaluated dialogue based on the multi-turn dialogue test set to obtain a first dialogue; and evaluating the first dialogue based on all nodes in the first dialogue. In this way, the semantic effect can be verified in an offline manner, and when a node with a semantic identification error is encountered in the test process, the remaining part after the error node is reconstructed in real time, so that the next sentence of dialogue can continue to flow into the next node for evaluation, and finally the evaluation of the semantic effect of all nodes is completely realized.
Owner:IFLYTEK SOUTH CHINA ARTIFICIAL INTELLIGENCE RES INST GUANGZHOU CO LTD

End-to-end speech recognition method and system based on keyword enhanced attention mechanism

ActiveCN122090829BAir traffic controlSpeech sound
The application relates to the technical field of speech recognition, and provides an end-to-end speech recognition method and system based on a keyword-enhanced attention mechanism, which comprises the following steps: a keyword retriever extracts entity keywords; a speech buffer manager receives continuous air traffic control audio segments, converts the audio segments into audio mel spectrum graphs, and then converts the audio mel spectrum graphs into audio embedding sequences; a keyword encoder unit maps a word embedding representation sequence into a keyword embedding sequence; an audio transcription decoder unit performs keyword-enhanced attention calculation, converts an audio embedding sequence, a transcription start flag and a splicing vector of a word embedding representation sequence into a predicted text sequence; a new word embedding representation sequence is input into a text transcriptioner through self-recurrence until a transcription end flag is output by the audio transcription decoder unit; and the predicted text sequence is output as a speech recognition text. The application realizes real-time recognition of streaming input speech, and has the advantages of high recognition accuracy, low response delay and flexible deployment.
Owner:NANKAI UNIV +1

A single-stage gating multi-modal fusion method with two-stage cross-attention

ActiveCN120654812BAlgorithmComputer vision
The present application relates to the technical field of artificial intelligence, in particular to a two-stage cross-attention single-stage gating multimodal fusion method, comprising: inputting text through an embedding layer and a text encoder to obtain a text vector; inputting an image through an image encoder to obtain an image vector; inputting the text vector and the image vector into a modal feature fusion module; the modal feature fusion module adopts a two-stage cross-attention single-stage gating modal fusion mechanism to output a fusion vector; and inputting the fusion vector through a decoder to obtain a predicted text. The present application improves the interaction effect between modes, reduces the hallucinations generated by the model, effectively reduces the calculation parameters of the model, and has good generality and practicability.
Owner:JIANGNAN UNIV

Adjustment method and apparatus for text description of multimodal large model, and device

The present application relates to an adjustment method and apparatus for a text description of a multimodal large model, and a device. The method comprises: determining a first sample image and configuring a first text description thereof as a second text description; adding an image trigger in the first sample image to obtain second sample images; by means of each third sample image and each second sample image, adjusting parameters of the added image trigger and a context generator; and obtaining image feature vectors from the sample images by means of an image encoder, obtaining, by means of a text encoder, text feature vectors from prediction text obtained by means of the context generator and text descriptions corresponding to the sample images, performing feature alignment on the basis of the image feature vectors and the text feature vectors so as to obtain output text of the multimodal large model for the sample images, and, on the basis of the similarity between the image feature vectors and the text feature vectors, determining a loss function. The present application keeps parameters of the multimodal large model unchanged as much as possible while specifically adjusting the output text of the multimodal large model, thereby improving adjustment efficiency.
Owner:CHINA TELECOM NETWORK SECURITY TECH CO LTD

An imbalanced sentiment analysis method and system based on counterfactual theory

The application provides an imbalanced sentiment analysis method and system based on counterfactual theory, and the method comprises the following steps: obtaining an imbalanced sentiment analysis training data set, wherein the training data set comprises original text and sentiment polarity; preprocessing the training data set; constructing an imbalanced sentiment analysis model based on counterfactual theory, training the imbalanced sentiment analysis model based on counterfactual theory based on the preprocessed data, and obtaining a trained imbalanced sentiment analysis model based on counterfactual theory; inputting the preprocessed text to be predicted into the trained imbalanced sentiment analysis model based on counterfactual theory, and obtaining the sentiment tendency of the text to be predicted. The imbalanced sentiment analysis method based on counterfactual theory is adopted, counterfactual learning is used, the problem of imbalanced sentiment categories in sentiment analysis is solved through oversampling, and thus the problem of low accuracy of the existing imbalanced sentiment analysis method is solved.
Owner:CHONGQING UNIV

Scene text extraction method and system without fine-grained detection

ActiveCN115879462BFeature vectorData set
The application provides a scene text extraction method without fine-grained detection. First, the obtained text image is input into a pre-trained text block detector to enable the text block detector to detect and crop the text image to form a text block image. Then, a pre-trained text block recognizer is used to obtain a semantic feature vector and a position feature vector of the text block image based on a text block feature map. Feature fusion and splicing are performed based on the semantic feature vector and the position feature vector to obtain a predicted feature, and a predicted text corresponding to the predicted feature is obtained. This framework combining coarse-grained detection and multi-instance recognition can reduce the detection burden, and the rich context information can be used for recognition. The text block detector can be trained by a text block level data set generated by a heuristic text block generation method based on a real data set, and high-precision text extraction can be achieved without fine-grained detection.
Owner:COMMUNICATION UNIVERSITY OF CHINA +1

Training method of video understanding model, processing method and device of video data, equipment, storage medium and program product

The application provides a video understanding model training method, a video data processing method, an apparatus, a device, a storage medium and a program product. The method comprises: obtaining a first visual feature sample and a first audio feature sample of a video sample, and obtaining a text feature sample of a prompt word sample; generating a first thinking process sample based on the first visual feature sample, the first audio feature sample and the text feature sample through a first video understanding model; generating a first predicted text sample based on the first thinking process sample through the first video understanding model; dividing the first predicted text sample into a plurality of predicted text segment samples, determining a first reward value based on the plurality of predicted text segment samples and at least one real text segment, and updating parameters of the first video understanding model based on the first reward value. Through the application, the accuracy of the video understanding model in generating structured transcription content corresponding to a video can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Text labeling model acquisition and text labeling method, related apparatus

The application discloses a text labeling model acquisition method and a text labeling method and related devices, and relates to the technical field of data processing. A first classifier and a second classifier of a first labeling model are trained using a first sample set, and the second classifier of the first labeling model is trained using a second sample set, to obtain a second labeling model. The first classifier is used to predict the category of a text, and the second classifier is used to predict the keyword of the text, so that the model is trained from two angles to improve the accuracy. In addition, the noise amount of the first sample set is lower than that of the second sample set, and the noise amount of any text sample set is inversely related to the accuracy of the category label of the text, so that the adverse effect of noise on classification is reduced. The second labeling model is used to obtain the keyword of the text in a third sample set. The third sample set and the keyword are input into an LLM to obtain the category label of the text in the third sample set output by the LLM. The second labeling model is trained based on the third sample set and the category label of the text in the third sample set to obtain a text labeling model. The cost can be reduced without using the LLM.
Owner:HYTERA COMM CORP

Dialogue assistance method and device, electronic equipment and computer readable storage medium

PendingCN122332505ATimed textEngineering
This application discloses a dialogue assistance method, apparatus, electronic device, and computer-readable storage medium, relating to the field of real-time text communication technology. The method includes: receiving RTT text sent by an RTT peer; parsing the RTT text to predict the corresponding dialogue mood and scenario; generating and displaying multiple styles of response text based on the dialogue mood and scenario; and sending the selected response text to the RTT peer in response to selecting any response text. Thus, this solution can assist users in quickly and accurately responding to RTT messages with appropriate mood and expression.
Owner:TCL COMM TECH (CHENGDU) LTD

Information processing apparatus and control method

An information processing apparatus includes: an input unit capable of detecting handwritten input; a display unit capable of displaying a trajectory of the handwritten input; a text recognition unit which recognizes text based on the handwritten input detected by the input unit; a prediction processing unit which predicts text candidates following the text based on the text recognized by the text recognition unit; a handwriting synthesis unit which synthesizes the text candidates predicted by the prediction processing unit into handwritten characters by combining first generation processing to generate handwritten characters based on tokens, each of which is a set of one or more handwriting strokes, collected from the handwritten input by a user, and second generation processing to generate handwritten characters based on a conversion model prepared in advance for converting text characters to handwritten characters; and a display processing unit which displays, on the display unit, the handwritten characters.
Owner:LENOVO (SINGAPORE) PTE LTD

Conversion of free-text natural language questions to structured data queries for database searching using large language models

A question processing system and methods are provided that are configured to intelligently generate structured data queries from natural language questions using a large language model (LLM). The system includes a processor and a computer readable medium operably coupled thereto, the computer readable medium comprising a plurality of instructions stored in association therewith that are accessible to, and executable by, the processor, to perform operations which include receiving a natural language question for structured data including columns having discrete text values, generating, using an LLM, a first structured data query having a predicted text value, searching a data store for similar ones of the discrete text values, generating, using the LLM, a second structured data query that refines the first structured data query based on the similar discrete text values, and querying a structured database based on the first and / or second structured data query.
Owner:NICE LTD

Model training method, text generation method, and related device

PendingCN122347195Aimprove accuracyEffective learning of fitnessSample graphLinguistic model
The application relates to the computer technical field and discloses a model training method, a text generation method and related equipment, which comprises the following steps: respectively extracting features of sample images by a basic feature network and N different expert feature networks in a visual language model to obtain basic visual features and N expert visual features; predicting the adaptation degrees of the expert feature networks according to the basic visual features by a gate network; weighting and fusing the N expert visual features according to the corresponding adaptation degrees to obtain expert fusion features; performing residual knowledge transfer on the expert fusion features and performing residual fusion on the expert fusion features and the basic visual features to obtain target visual features; performing text decoding on the target visual features and text embedding representations of instruction prompt texts by a text decoding network to obtain predicted texts; and adjusting parameters of the visual language model according to the predicted texts and sample texts until a training end condition is reached. The scheme can improve the accuracy of visual and text perception.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Picture recognition model generation, picture recognition method, device, equipment and medium

The disclosure provides a picture recognition model generation, picture recognition method, device, equipment and medium, image processing technical field, especially in the field of artificial intelligence and computer vision. The specific implementation scheme is: obtaining a plurality of training data, the training data including a sample picture, a model instruction and a labeled text of the sample picture, the model instruction being used for obtaining a text of a specified attribute, and the labeled text being a text of the specified attribute associated with a visual feature and text information of the sample picture; inputting the sample picture and the model instruction into a first multi-modal large model to obtain a predicted text; determining a model loss value according to the labeled text and the predicted text; and using the model loss value to fine-tune the first multi-modal large model to obtain a picture recognition model. The picture recognition model is used to understand a target picture to obtain a picture recognition result. The scheme can be applied to diversified and complex application scenarios, and the accuracy of picture content picture recognition is improved.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Training text processing model

PendingUS20260195618A1Data miningText processing
In a method for training a text processing model, a sample text set is obtained. The sample text set includes a plurality of independent sample text and sample labels of the plurality of independent sample text. At least two independent sample text in the plurality of independent sample text are concatenated for each of a plurality of concatenated sample text. A plurality of concatenated sample features is obtained from the plurality of concatenated sample text through a text processing model. A feature corresponding to each independent sample text in the plurality of concatenated sample features is masked to obtain a plurality of predicted text. A loss function is determined based on differences between the plurality of predicted text and the sample labels. Parameter updating on the text processing model is performed based on the loss function.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD