Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

160 results about "Predictive text" patented technology

Predictive text is an input technology used where one key or button represents many letters, such as on the numeric keypads of mobile phones and in accessibility technologies. Each key press results in a prediction rather than repeatedly sequencing through the same group of "letters" it represents, in the same, invariable order. Predictive text could allow for an entire word to be input by single keypress. Predictive text makes efficient use of fewer device keys to input writing into a text message, an e-mail, an address book, a calendar, and the like.

Emotion prediction and disease derivation method and system based on multi-modal fusion

The invention discloses an emotion prediction and disease derivation method and system based on multi-modal fusion, and the system comprises a data collection and preprocessing module, an emotion fusion module, an abnormal condition detection and cloud uploading module, and a disease possibility derivation module. The data acquisition and preprocessing module comprises a video part, a text part and an audio part, and the video part comprises face emotion recognition and prediction and human motion recognition and prediction; the text part comprises text content emotion recognition and prediction; the audio part comprises voice-to-text and voice tone emotion recognition and prediction, the system comprehensively captures an emotion state by fusing multi-mode information such as video, text and voice, and the accuracy and prediction capability of emotion recognition are improved; and by predicting the future emotion trend, the abnormal condition is warned in advance, and the response timeliness is improved.
Owner:JIANGSU UNIV OF SCI & TECH IND TECH RES INST OF ZHANGJIAGANG

Building vision-language models using masked distillation from foundation models

The present disclosure relates to systems, non-transitory computer-readable media, and methods for training and implementing a vision-language model using masked distillation and contrastive image-text training. In particular, in one or more embodiments, the disclosed systems generate, utilizing a vision encoder, an image embedding from a masked digital image comprising a digital image with one or more masked patches. In some embodiments, the disclosed systems generate, utilizing a text encoder, a text embedding from a masked text phrase. In one or more embodiments, the disclosed systems generate, utilizing the vision-language model from the image embedding and the text embedding, a predicted text reconstruction of the text description and a predicted image reconstruction of the digital image. In some embodiments, the disclosed systems modify parameters of the vision-language model according to a masked distillation loss between the predicted text reconstruction and a text reconstruction generated by a pretrained large language model.
Owner:ADOBE INC

Voice cloning method and device based on emotion enhancement and related medium

The invention discloses a speech cloning method and device based on emotion enhancement, and a related medium. The method comprises the following steps: respectively obtaining a reference audio and a prediction text corresponding to the reference audio; respectively preprocessing the reference audio and the prediction text to generate standardized audio data for feature extraction; inputting the standardized audio data into an emotion enhancement module so as to generate acoustic features matched with a target voice style through multi-round feature fusion and an autoregression generation mechanism; and inputting the acoustic features into a speech synthesis module for decoding processing, and outputting predicted speech. According to the invention, by introducing the emotion enhancement module, accurate modeling of the emotion style of the speaker in the acoustic feature generation process is realized, so that the emotion simulation degree and the personality restoration capability of the synthesized speech are significantly improved.
Owner:AFIRSTSOFT CO LTD

Text summary generation model training method, text summary generation method and device

The present application relates to a text summarization model training method, a text summarization method, an apparatus, a computer device, a storage medium, and a computer program product. The method comprises: inputting a training text and a corresponding training image set into an initial text summarization model to obtain a predicted text summary, and generating a target loss based on the difference between the predicted text summary and the labeled text summary; inputting masked training data corresponding to the training text and first training data into the initial text summarization model to obtain masked predicted data, and generating a reconstruction loss based on the difference between the masked labeled data and the masked predicted data; adjusting model parameters of the initial text summarization model based on the target loss and the reconstruction loss until convergence conditions are met, thereby obtaining a target text summarization model; and using the target text summarization model to generate a text summary of the text. This method can improve the prediction accuracy of the model and the quality of the generated text summary.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1

Video content understanding method and device based on structured grammar information, electronic equipment and storage medium

The invention belongs to the technical field of computer application, and discloses a video content understanding method and device based on structured grammar information, electronic equipment and a storage medium, and the method comprises the steps: inputting a training sample into a target model for content understanding processing, obtaining a prediction text, and constructing a syntactic tree corresponding to the prediction text; calculating a syntactic tree editing distance between the syntactic tree and the syntactic tree of the reference text; calculating language structure loss by using the syntactic tree editing distance, and updating model parameters of the target model by using the language structure loss; under the condition that the model parameters of the target model are trained, obtaining a target video; and inputting the target video into the target model for processing to obtain a content text of the target video. In the application, the target model is trained based on the language structure loss calculated based on the syntax tree, and the target video can be understood, so that the content text with accurate grammar and reasonable and natural sentence structure is obtained.
Owner:CHINA UNIV OF PETROLEUM (BEIJING)

Contract risk auditing method and device

The invention relates to the field of artificial intelligence, and particularly provides a contract risk auditing method and device, and the method comprises the following steps: S1, splitting a contract content paragraph, and decomposing a contract text into a plurality of independent paragraph units with a logic relation; s2, converting the contract text into a structured data unit, and performing text analysis and risk factor extraction; s3, table textualization: converting table data in the contract into a plain text format which can be processed by a large model; s4, constructing a list table to systematically organize key information in the contract; s5, performing text semantic analysis to deeply understand the contract text; s6, element checking: carefully reviewing and verifying key elements in the contract text; and S7, risk information summarization: summarizing and analyzing potential risks in the contract text. Compared with the prior art, the method has the advantages that potential risks in the text can be recognized and predicted through learning of a large amount of text data, and therefore the efficiency and accuracy of contract risk auditing are remarkably improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Data generation method and device based on multi-modal large language model

The invention discloses a data generation method based on a multi-modal large language model, and the method comprises the steps: obtaining inquiry text information used for representing a business task in a business application scene, and image information associated with the business task, the inquiry text information comprises at least one of task definition information used for representing a service task needing to be completed and task logic information used for representing service logic, reasoning the obtained inquiry text information and the image information based on the trained multi-mode large language model, and obtaining the inquiry text information and the image information. Generating prediction text information used for representing response information of the inquiry text information based on the image information, the prediction text information comprises at least one of entity information used for representing a target image in the image information and image position information where an entity is located, attribute information used for representing additional information of the entity, and rule information used for representing a rule disassembled based on task logic. According to the invention, the generated data has pyramid dimension information.
Owner:HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Method, apparatus and electronic device for determining word representation vector

Embodiments of the present disclosure provide a method, an apparatus and an electronic device for determining a word representation vector, and a computer-readable storage medium, which belong to a field of processing natural languages. The method includes obtaining a set of glyph units of a text; obtaining a context vector of the text based on the set of glyph units; and predicting next word of the text based on the context vector. The method for determining the word representation vector of the present disclosure can effectively obtain a corresponding set of glyph units even for hieroglyphics in which hyperbolic characters are prone to appear or languages evolved from the hieroglyphics, thereby improving an accuracy of determining the word representation vector.
Owner:SAMSUNG ELECTRONICS CO LTD

Augmentative and alternative communication (AAC) solutions

Method, software, and apparatus for improved Augmentative and Alternative Communication (AAC) solutions. In one aspect, a user interface is provided with a set of suggestions comprising text, phrases, etc., and navigation buttons that enable users to select words and phrases to add to be written and / or spoken in a manner that reduces the number of user inputs. The suggestions are displayed in alphabetical order in rows with navigation buttons adjacent to the rows, with activation of a navigation button resulting in generation of updated suggestions having alphabetical ranges that are bounded by suggestions in associated rows. This approach may be combined with predictive text means to enable users to easily formulate text and / or speech content.
Owner:ANSELL PETER JOHN

Detection model training method and apparatus, and text bounding box detection method and apparatus

PCT designated stageWO2025162436A1Biological modelsText detectionAlgorithm
The present disclosure relates to a detection model training method and apparatus, and a text bounding box detection method and apparatus. The training method comprises: inputting into a detection model to be trained a sample image carrying a text label, wherein the text label comprises a reference text bounding box and a reference probability value of a preset deletion symbol being drawn on text within the reference text bounding box; performing detection on the sample image by means of the detection model, so as to determine a text detection result of the sample image, wherein the text detection result comprises a predicted text bounding box and a predicted probability value of the preset deletion symbol being drawn on text in the predicted text bounding box; and on the basis of a target loss function, converging the text detection result and the text label, so as to obtain a trained detection model, wherein the target loss function comprises a first loss function used for evaluating the accuracy of the predicted text bounding box, and when the overlap ratio between the predicted text bounding box and the reference text bounding box is in different overlap ratio intervals, the first loss function is configured with penalty terms of different weights.
Owner:SHENZHEN XINGTONG TECH CO LTD

Handwritten text recognition method and device, equipment and storage medium

The invention discloses a handwritten text recognition method and device, equipment and a storage medium, and belongs to the field of text recognition. The method comprises the following steps: acquiring a handwritten text image; extracting handwriting style characteristics according to the handwritten text image, wherein the handwriting style characteristics are used for reflecting a writing style corresponding to the handwritten text image; calibrating a feature extraction position of the handwritten text image according to the handwritten style feature, and extracting a text feature of the handwritten text image according to the calibrated feature extraction position, the text feature being used for reflecting a text in the handwritten text image; and predicting a prediction text corresponding to the handwritten text image according to the text features. According to the method, the accuracy of extracting the text features of the handwritten text image can be improved, so that the recognition performance of handwritten texts of different writing styles is improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Personalized aphasia communication assistant system

Methods and systems for improving communication involving a personal with aphasia are disclosed. The methods and systems include: obtaining an audio input indicative of speech of the person with aphasia; providing the audio input to a personalized aphasia translation assistant, wherein the personalized aphasia translation assistant was trained to recognize and translate speech of the person with aphasia using a general dataset of aphasia-speech; determining a plurality of words from the audio input using an aphasia-specific recognition model; inputting the plurality of words into an aphasia generative model, the aphasia generative model comprising a natural language processing machine learning model trained using a general dataset of aphasia sentences and a corresponding dataset of translated sentences; generating one or more formulated and contextual sentences using the personalized aphasia generative assistant; and outputting the one or more formulated and contextual sentences. A further method provides the training of a personalized aphasia communication assistant, in particular the adaptation of a pre-trained speech recognition model and of a pre-trained generative speech model based on predicted text and a user confirmation feedback of the accuracy of the predicted text.
Owner:UNIV OF SOUTH FLORIDA

System and method for automatically generating code for software code updates

A system includes a memory and processors operably coupled to the memory and configured to receive a request for a software code change to one or more of a first instance of a software application or a plurality of second instances of a software application. The first instance of the software application is represented as a first node of a directed acyclic graph (DAG) and each instance of the plurality of second instances of the software application is represented as one of a plurality of second nodes of the DAG. The processors are further configured to identify relationships between each node impacted by the requested software code change. The processors are further configured to execute a first machine-learning model and a second machine-learning model to generate a prediction of textual software code based on a knowledge graph. The textual software code is to be deployed for updating each node.
Owner:BANK OF AMERICA CORP

Aligned vision-language model for text-rich image understanding

The present disclosure relates to systems, non-transitory computer-readable media, and methods for generating and implementing a vision-language model that identifies and understands text-rich content depicted in digital images. For example, the disclosed systems determine, from among a plurality of digital images with at least a threshold probability of depicting text-rich content, a subset of digital images corresponding to a set of text-rich image classifications. In some embodiments, the disclosed systems generate a ground truth text phrase utilizing an optical character recognition model to process a digital image from the subset of digital images. In certain embodiments, the disclosed systems also generate a predicted text phrase utilizing a vision-language model and compare the ground truth text phrase with the predicted text phrase. In some embodiments, the disclosed systems modify parameters of the vision-language model based on comparing the ground truth text phrase and the predicted text phrase.
Owner:ADOBE INC

Generating Clinical Documentation Using Large Language Models and Artificial Intelligence

Systems and methods generate clinical documentation using large language models and artificial intelligence (AI). A template management module is provided to create customizable templates. A processing unit can receive input data from various sources and use AI to generate transcripts, summarize sessions, and produce clinical documentation such as clinical notes. The processing unit may also generate Current Procedural Terminology (CPT) and diagnosis codes, generate after-visit summaries, and generate referral letters. The AI may be trained on past clinical notes and can adapt to the clinician's style over time, with a feedback loop for continuous improvement. Additional features include cohort-based training, real-time language translation, predictive text, and analytics for documentation trends. The system supports customization of note length, style, and keywords, as well as integration with external medical databases and patient portals.
Owner:ORCHID EXCHANGE INC

Machine learning and multi-stage prompting techniques for generating target classification signatures

Various embodiments of the present disclosure provide machine learning architectures and data processing techniques for improving computer-based text comprehension. The techniques include generating, using a trained classifier model, target classification probabilities for labelled text-based objects from a testing portion of a labelled training dataset and identifying predictive text-based objects from the labelled text-based objects based on the target classification probabilities. The techniques include applying a staged prompting mechanism with a generative extraction model to identify a target set of explanatory text segments from the predictive text-based objects that may be clustered into semantic segment clusters. The techniques include generating explanatory summary segments respectively corresponding to the semantic segment clusters and generating a target classification signature based on a plurality of terms from the one or more explanatory summary segments.
Owner:OPTUM INC

Large model method and device for AVI detection defect data judgment and medium

The invention discloses a large model method and device for AVI detection defect data judgment and a medium. The large model method comprises the steps that a defect detection large model is constructed; the method comprises the following steps: acquiring an industrial product photo, generating an image text data pair after text labeling, inputting the image text data pair into a defect detection large model, generating intermediate feature representation by matching an image encoder and a text encoder, and generating a prediction mask by a decoder based on the intermediate feature representation; performing multi-modal alignment on the predicted text mask and the text feature information to obtain modal feature representation; inputting the modal feature representation and the predicted image mask into a cue word encoder to generate a Prompt vector code with the same dimension as the image feature information; and inputting the Prompt vector code into the LLM model, and outputting a semantic level judgment result after language questioning. The large model method has good mobility, expansibility, strong generalization ability and perception ability, gets rid of threshold dependence, and realizes accurate redetermination of AVI detection defect data.
Owner:JIANGSU PROVISION ELECTRONICS CO LTD

Speech translation method and device

The invention provides a speech translation method and device, and the method comprises the steps: translating source language speech data based on a speech translation model, and obtaining a target language text; a training target of the speech translation model comprises minimizing a difference between a first target language prediction text generated based on source language sample speech data and a translation label corresponding to the source language sample speech data, and minimizing a difference between speech features of the source language sample speech data and text features of a first source language sample text. And minimizing the difference between a second target language prediction text generated based on the pseudo source language speech features and a translation label corresponding to the second source language sample text. According to the method and the device, the problem of scarcity of annotated voice data is solved by efficiently utilizing relatively rich text data in a low-resource language scene, so that the performance of a voice translation model is improved.
Owner:IFLYTEK CO LTD

Text generation method and apparatus, electronic device, and storage medium

The present application is suitable for the technical field of natural language processing, and provides a text generation method and apparatus, an electronic device, and a storage medium. The method comprises: acquiring a predicted token and a first hidden representation corresponding to text to be predicted, wherein the predicted token includes a new token obtained by performing inference on the text to be predicted, and the first hidden representation is an embedding representation obtained by performing, by means of a transformer layer, feature extraction on the text to be predicted; using the predicted token and the first hidden representation as an input of a trained acceleration network to obtain a candidate token outputted by the acceleration network, wherein the acceleration network is used for inferring the new token on the basis of the first hidden representation and the predicted token to obtain the candidate token; and on the basis of the predicted token and the candidate token, determining first target text corresponding to the text to be predicted. The present application can improve the quality of generated text.
Owner:SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD

Method and related device for discovering generalized intent in text based on deviation self-correction calibration

The present invention discloses a method and system for discovering generalized intent in text based on self-correction calibration of deviations, comprising: obtaining a predicted text sample and inputting it into a preset generalized intent discovery model to obtain the original output of the model; using the Softmax function to classify the original output of the model, and taking the category to which the Softmax maximum value belongs as the sample prediction category of the text sample. Among them, a biased branch and a trainable branch are set inside the generalized intent discovery model. First, a biased branch with a bias is obtained through pre-training, and the model parameters are fixed. Then, the training text sample is input into the pre-trained biased branch and the trainable branch respectively, and two original outputs are output. The original output of the biased branch is used to adjust the original output of the trainable branch, and the deviation of the model for known categories is used to alleviate category deviation and category confusion, which has a role in alleviating category deviation and category confusion, and effectively improves the recognition accuracy of the model for new categories.
Owner:PEOPLE CN CO LTD +1

Intelligent conversion method and device based on video and text, electronic equipment and medium

ActiveCN115205758BGuarantee generation premiseGuaranteed conversion effectSemantic analysisCharacter and pattern recognitionFeature vectorEngineering
The application relates to the field of artificial intelligence, and discloses an intelligent conversion method based on videos and texts, which comprises the following steps: acquiring a training video and corresponding video text of the training video, and extracting training pictures in the training video; using an encoder in a pre-constructed text-video conversion model to perform feature vector coding, vector masking and vector splicing on the training pictures and the video text to obtain picture-text splicing vectors; using a semantic analysis network in the pre-constructed text-video conversion model to identify predicted pictures and predicted texts of the picture-text splicing vectors, and then decoding to obtain predicted videos and predicted video texts; calculating model loss of the pre-constructed text-video conversion model according to the predicted videos and the predicted video texts, and the training videos and the video texts, so as to generate a trained text-video conversion model, realize scene conversion on to-be-converted scene data, and obtain a scene conversion result. The application can improve the scene conversion efficiency between videos and texts.
Owner:CHINA MERCHANTS FINANCE HLDG CO LTD

A text-guided face spoofing detection method and system

This application provides a text-guided method and system for detecting face forgery. The method includes obtaining a face image to be detected (whether it is genuine or fake), inputting the face image into a face detection model to obtain the face authenticity detection result. The face detection model includes: constructing a text prompt lexicon covering multiple granularities and generating multi-dimensional text prototypes; extracting visual features, optimizing the feature distribution of visual features to obtain global visual features; performing feature separation and enhancement on the global visual features; mapping the global visual features after feature separation and enhancement to predicted text features; applying similarity constraints to obtain cross-modal prototype matching results; applying discriminative constraints on different-dimensional text prototypes to obtain feature measurement learning results; and outputting the detection result based on the cross-modal prototype matching results and feature measurement learning results. This method solves the problem of poor generalization in face forgery detection and improves detection accuracy and generalization ability.
Owner:NANJING UNIV OF POSTS & TELECOMM

Speech recognition method, training method, device, electronic device and storage medium

The present disclosure provides a speech recognition method, training method, device, electronic device, and storage medium. The method includes obtaining a target speech, determining a recognized text of the target speech based on a speech recognition model, and training the model in the following manner: using the model to predict the predicted text of the speech sample, and updating the model parameters of the model when the performance evaluation index of the model meets the iteration condition. The index parameters include the recognition quantization parameters and semantic difference parameters of the statistical objects contained in the predicted text. When the statistical object is correctly recognized, the recognition quantization parameters include the recognized correct quantization data. When the recognition is incorrect, the recognition quantization parameters include the edited quantization data of the editing operation, and the semantic difference parameters are used to correct the edited quantization data. In this case, the performance evaluation index can differentially measure the impact of the predicted text on the user's understanding, avoiding the recognition text with large semantic deviation from being mistakenly output as the correct text.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Method of proficient typing using a limited number of classes

Systems and methods for proficient typing using a limited number of classes are disclosed. In one embodiment, capturing keystrokes and returning predictive text on a display includes creating a dictionary of words and representative keypresses associated with each, capturing a series of keypresses of keys on a keyboard, where at least one key represents assignments of more than two characters, for each captured keypress, generating probabilities for the likelihood each of a plurality of words from the dictionary to be next by providing a language model with all previous text in the series up to but not including the last token, storing the generated probabilities with the word in the dictionary, applying a filter to the dictionary to obtain a subset of words, and displaying the subset of words on a display, obtaining input that selects a word from the subset, and indicating the selected word as committed.
Owner:RGT UNIV OF CALIFORNIA

Video time retrieval method based on bidirectional semantic enhancement

A video time retrieval method based on bidirectional semantic enhancement comprises the following steps: firstly, respectively extracting video features and querying text features by using a pre-trained video encoder and a pre-trained text encoder; secondly, respectively inputting the extracted features into a feature alignment module to obtain aligned video features and text features; then, designing a TGVM module, dynamically enhancing video features according to global features of a text mode, and reducing irrelevant information in the video features; thirdly, designing a Feedback Decoder module and feeding back comparison loss, shortening the distance between decoded text features and original video features, and enhancing text feature representation through the original video features; then, using a cross attention mechanism to obtain joint features after interaction of the video and the text; finally, a corresponding time slice in the video is queried through a decoder prediction text, and a model is trained under the supervision of time retrieval joint loss; according to the method, the multi-modal reasoning capability of the multi-modal interaction part is enhanced, so that video and text information can be better aligned.
Owner:XIDIAN UNIV

Machine customer service training system and method, voice reply method, and electronic device

The application provides a machine customer service training system and method, a voice reply method and an electronic device. The machine customer service training system comprises a machine customer service model, a user model, a reward parameter configuration component and a termination component. The user model is used to generate a plurality of first predicted texts according to a first text output by the machine customer service model and a historical communication text of the first text. The machine customer service model is used to randomly determine a target predicted text from the plurality of first predicted texts, and generate a second predicted text according to the target predicted text and the historical communication text. The reward parameter configuration component is used to configure a first positive reward parameter for the machine customer service model when the current conversation between the user model and the machine customer service model ends successfully. The termination component is used to terminate the training of the machine customer service model when the number of training times of the machine customer service model is greater than a threshold value, so as to obtain a trained machine customer service model. The application can train a high-quality machine customer service model.
Owner:ALIBABA DAMO (HANGZHOU) TECH CO LTD

Multi-hop reading understanding method and system based on paragraph pair loss and replication mechanism

PendingCN120670586ASemantic analysisBiological modelsAlgorithmCopying mechanism
The invention discloses a multi-hop reading understanding method based on a paragraph pair loss and replication mechanism, and the method comprises the steps: respectively splicing a preset question with a plurality of paragraphs, and obtaining a spliced paragraph; screening the plurality of spliced paragraphs to obtain high-score paragraphs; splicing the high-score paragraph with a preset question again to obtain a high-score spliced paragraph; generating word list distribution and replication distribution based on cross attention scores for the high-score spliced paragraphs through a preset decoder, and dynamically fusing the two types of distribution to generate a prediction distribution probability; and determining a maximum value in the prediction distribution probability, and determining a corresponding prediction token in a preset word list according to an index corresponding to the maximum value so as to obtain a prediction text.
Owner:SUN YAT SEN UNIV

Voice representation model training method and device, equipment, storage medium and product

The invention discloses a voice representation model training method, device and equipment, a storage medium and a product, and the method comprises the steps: when a voice representation model is trained, carrying out the coding of obtained voice features, carrying out the token boundary prediction through employing the voice features, generating a token boundary, generating prediction text data according to the token boundary, carrying out the discretization of the prediction text data, and carrying out the training of a voice representation model. And discretized speech representation is obtained. Due to the fact that the token-level discretized speech features can be obtained, cross-modal expression with higher coupling degree of the speech mode and the text mode and closer information expression can be obtained, and semantic joint expression with more information is achieved, the coupling degree with a multi-modal language model is further improved, and the precision of a large model is prevented from being affected.
Owner:CHINA MOBILE COMM LTD RES INST +1

A text input information processing method, apparatus and storage medium

This application discloses a text input information processing method, apparatus, and storage medium. The method involves acquiring user text input information, including complete words already entered by the user and / or characters representing incomplete words currently being entered by the user; generating a first semantic feature based on the text input information to characterize the semantics of the text input information; wherein the first semantic feature further characterizes the positional information of complete words and / or input characters in the text input information, the positional information including word-level positional information corresponding to words and / or character-level positional information corresponding to characters; and generating output information corresponding to the text input information based on the first semantic feature, wherein the output information includes predicted words corresponding to the text input information. This method improves the accuracy of prediction results corresponding to the text input information and enhances compatibility with different types of prediction tasks, thereby improving applicability.
Owner:BEIJING YUANSHI TECHNOLOGY CO LTD