Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

397 results about "First language" patented technology

A first language, native language or mother/father/parent tongue (also known as arterial language or L1), is a language that a person has been exposed to from birth or within the critical period. In some countries, the term native language or mother tongue refers to the language of one's ethnic group rather than one's first language.

Personally identifiable information scrubber with language models

Sanitizing data can be a cumbersome task, particularly when the volume of data is large, the content is sensitive, and / or the type of sanitation requires contextual determinations. Sanitizing large amounts of data is tedious and may often require highly trained personnel with clearances and / or other qualifications. In the systems and methods of the present disclosure, language models (LMs) are used to solve these and other technical issues with tools that may allow sanitizing data easily, with high versatility, context awareness, and / or low demand for computational resources. In particular, some of the disclosed systems and methods use a first language model and a second language model (being less resource-intensive than the first language model) to generate sanitized output data with improved efficiency and accuracy. This dual-model approach ensures that sensitive information is handled appropriately while optimizing computer resource usage.
Owner:OPENAI OPCO LLC

Generating speaker video and audio in multiple languages for videoconferencing

Systems and methods for generating speaker video and audio in multiple languages for videoconferencing are provided. For example, a computing device can access a speaker speech audio signal that includes a speaker speech in a first language, a video of the speaker and a translated speech audio signal of the speaker speech in a second language. The computing device generates, based on the translated speech audio signal, a converted translated speech audio signal that includes a speech in the second language having voice characteristics in the speaker speech. The computing device further generates a lip-synched speaker video based on the video of the speaker and the converted translated speech audio signal. Lip movements in the lip-synched speaker video correspond to the converted translated speech audio signal. The converted translated speech audio signal and the lip-synched speaker video are transmitted to a video conference provider configured to host the video conference.
Owner:ZOOM VIDEO COMM INC

Data extraction method and system for entity evaluation

The present invention provides a data extraction method for entity evaluation. The method is used for extracting target data on the basis of a predefined indicator, and comprises: an underlying data processing step: at least comprising using a first language model to identify identical events in underlying data; and an indicator data extraction step: at least comprising, on the basis of the predefined indicator, using the first language model to extract target data from among the underlying data, wherein a base large model of the first language model comprises a large language model based on a Transformer model architecture, and the first language model is obtained by training on the basis of the base large model. The technology of the present invention can significantly improve the accuracy of data extraction and reduce training costs.
Owner:YINGTOU INFORMATION & TECH SHANGHAI CO LTD

Multi-lingual text-to-speech controlling

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for performing text-to-speech modeling. In some implementations, a computer device receives input text in a first language to convert to a desired speech in a second language. The computer device receives one or more criteria for modifying the desired speech and converts the input text to a desired text in the second language. The computer device generates audio representations of the desired text in the second language and predicts, for each of the audio representations of the desired text and using the one or more received criteria, a pitch value and a duration value. The computer device generates the desired output speech using the predicted pitch value and the predicted duration value for each of the audio representations and provides the desired speech for output.
Owner:MURF INC

Systems and methods for multilingual data processing and arrangement on a multilingual user interface

A method for generating a first case dataset in a first language. The method includes receiving adverse event data. The method further includes determining case data including general case data and regional case data and providing the case data to a translator computing device to enable display on a user interface including multiple duolingual text fields with a first language text field including at least a portion of the text data in the first language and a second language text field adjacent the first language text field. The method further includes receiving the text data in the second language from a translator computing device. The text data in the second language is received via the second language text fields of the plurality of duolingual text fields. The method further includes generating and outputting the first case dataset including the text data in the first language.
Owner:VEEVA SYSTEMS INC

Systems and methods for improving results from language machine learning models

Provided herein are systems, methods, and computer-readable media for improving results from machine learning models. An example method may include obtaining a user input query; generating first logits from a first language model by applying the first language model to the user input query; generating second logits from a second language model by applying the second language model to the user input query; combining the first logits and the second logits; determining probabilities associated with tokens from the combined first logits and second logits; and generating an output token based on the determined one or more probabilities.
Owner:SHOPEE IP SINGAPORE PTE LTD

Filtering Content for Automated User Interactions Using Language Models

Methods, apparatus, and processor-readable storage media for filtering content for automated user interactions using language models are provided herein. An example method includes obtaining a plurality of portions of content based on a query corresponding to one or more topics related to an organization, where the plurality of portions of content is retrieved from at least one content source corresponding to the organization, and configuring a first language model instance to generate a score for each portion of content in the plurality of portions of content based on its relevancy to the query. The method includes filtering the plurality of portions of content based at least in part on one or more filtering criteria and the score generated for each portion, and generating, using a second language model instance, a response to the query, where the response is based on the portions of content resulting from the filtering.
Owner:PLANETART LLC

Unlearning data from language models

Devices and techniques are generally described for unlearning information from large language models (LLMs). In various examples, a first language model (LM) trained on a first training corpus D may be determined. First data F that is a subset of D may be determined. A first auxiliary LM may be trained using the first training corpus D and a second auxiliary LM may be trained using a second training corpus D / F, where the second training corpus D / F represents the first training corpus D without the first data F. A first text input may be determined. The first LM may be updated based at least in part on a first prediction difference between predictions the first LM and the second auxiliary LM for a first set of inputs and a second prediction difference between the predictions of the first LM and the first auxiliary LM for the first set of inputs.
Owner:AMAZON TECH INC

Translation model training method, text translation method and generative model training method

The invention relates to a training method of a translation model, a text translation method and a training method of a generative model. The method comprises the steps that a training text of a first language is input into an initial translation model, a model output text is obtained, and the model output text comprises a model translation text of a second language; obtaining a format reward corresponding to the model output text based on the format of the model output text, wherein the format reward is used for representing whether the model output text accords with a preset format; obtaining a measurement reward corresponding to the model translation text based on the training text and the model translation text, wherein the measurement reward is used for representing the translation quality of the model translation text; obtaining a target reward according to the format reward and the measurement reward; and training the initial translation model based on the target reward and a preset reinforcement learning algorithm to obtain a target translation model. According to the training method, a format and measurement mixed reward mode is adopted, a guidance signal with rich and effective information is provided for reinforcement learning, and the translation quality of the translation model can be improved.
Owner:SHUXING TECH (BEIJING) CO LTD

Systems and methods for using machine-learning to extract and process audio data

A method for extracting and processing audio data may include receiving one or more packets of multimedia content. The one or more packets of multimedia content may comprise audio data. The method may further include extracting the audio data from the one or more packets of multimedia content. The audio data may comprise verbal speech in a first language. The method may further include converting the audio data into first text data in the first language based on the verbal speech in the first language. The method may further include providing the first text data to a generative machine-learning model. The generative machine-learning model may have been trained to translate the first text data in the first language to a second language and generate second text data in the second language. The method may further include transmitting, to a user interface, the second text data in the second language.
Owner:STATS LLC

Language migration detection method and device, equipment and storage medium

The invention discloses a language migration detection method and device, equipment and a storage medium, and relates to the technical field of computers. The method comprises the steps of obtaining a first language source code and a second language migration code, wherein the second language migration code is obtained by migrating the first language source code to a second language environment; analyzing the first language source code to obtain a first abstract syntax tree, and analyzing the second language migration code to obtain a second abstract syntax tree; respectively encoding each first node of the first abstract syntax tree and each second node of the second abstract syntax tree through a pre-trained semantic encoder to obtain a first semantic vector of each first node and a second semantic vector of each second node; and inputting the first semantic vector and the second semantic vector into a pre-trained logic detection model, and determining whether the logic of the first language source code and the logic of the second language migration code are consistent or not through the logic detection model, thereby realizing comprehensive detection of the business logic of the source code and the migration code.
Owner:广州三七极创网络科技有限公司

Information processing method, model training method, device, equipment, medium and product

The invention provides an information processing method and device, a model training method and device, equipment, a medium and a product, and relates to the technical field of computers. The public opinion information processing method comprises the steps that on the basis of a first language model, a crawled original language text is coded to obtain an original language vector, on the basis of a second language model, a target language text is coded to obtain a target language vector, the target language text is generated by translating the original language text, and the first language model is a multi-language model; generating a sentence-level first fusion vector and a word-level second fusion vector based on the original language vector and the target language vector; obtaining corresponding theme features and emotion features based on the first fusion vector; performing entity extraction on the second fusion vector to obtain entity features; and generating a public opinion analysis result based on the subject features, the emotion features and the entity features. According to the technical scheme, the public opinion association capture capability is enhanced through multi-feature joint analysis, and the public opinion analysis efficiency and accuracy can be improved.
Owner:CHINA TELECOM GLOBAL LTD

LLM powered security product facade

An AI-based application uses AI models to expeditiously identify a product or service most relevant to an issue input to the application instead of a user navigating a platform or catalogue of products / services. The application prompts a first language model to convert the input description into a structured description of the issue. The application then, for each product / service category, prompts a second language model to identify a product / service most relevant to the structured description within each product / service category. The application then returns an identification of a product / service determined to be most relevant to the issue as represented by the structured description, confidence in the identification, and a reason the product / service was identified.
Owner:PALO ALTO NETWORKS INC

Speech translation method and device of end cloud translation system, equipment and medium

The invention relates to a speech translation method, device and equipment of an end cloud translation system and a medium, and relates to the technical field of intelligent speech translations, the end cloud translation system comprises an audio receiving end and a cloud, the method is applied to the audio receiving end, and the method comprises the following steps: obtaining a first language audio of the audio receiving end; sending the first language audio to a cloud; and receiving and playing a second language audio, wherein the second language audio is translated by the cloud according to the first language audio and then is sent out. The method can be widely applied to cross-border conference scenes, and the cross-language communication efficiency and experience are improved.
Owner:SHENZHEN XINZHILIAN SOFTWARE CO LTD

Efficient and effective system and method to build multi-lingual large language models

A method of enabling a language model trained in a first language to support a second language includes extending an existing vocabulary of the language model to include additional tokens for text in the second language. The method also includes initializing the additional tokens for text in the second language based on subtokens of tokens for text in the first language from the existing vocabulary. The method further includes training the language model using a mixed language dataset that includes a first language corpus and a second language corpus. In addition, the method includes performing instruction tuning using a dataset that includes (i) instruction and response pairs involving the first language and (ii) instruction and response pairs involving the second language.
Owner:SAMSUNG ELECTRONICS CO LTD

Language model-based question and answer method and device, equipment and storage medium

One or more embodiments of the invention provide a language model-based question and answer method, apparatus and device, and a storage medium. The method comprises the steps of obtaining a to-be-solved question; determining a task type corresponding to the question, and selecting a target first language model corresponding to the task type from at least one first language model; the question is input into the target first language model, a thinking chain reasoning step corresponding to the question is generated through the target first language model, the thinking chain reasoning step and the question are input into a second language model, and the second language model conducts thinking chain reasoning on the basis of the thinking chain reasoning step. Generating an answer corresponding to the question; wherein the number of the model parameters of the first language model is smaller than the number of the model parameters of the second language model. According to the method, the calculation cost and the response delay of the model can be effectively reduced while the generation precision of the model is ensured.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Methods and devices for performing iterative nash policy optimization on language models

The present disclosure describes various methods, systems, and storage medium for training a language model to obtain an iterative Nash policy optimized (INPO) language model. The method includes initializing a first language model by a reference language model; for N-th iteration with N starting from 1 to M, generating a plurality of responses using the N-th language model for each prompt in a plurality of prompts, and constructing a preference dataset using a preference oracle based on the plurality of responses for each prompt, wherein the preference dataset comprises a winning response and a losing response; training the N-th language model to obtain a (N+1)-th language model by minimizing a value of an INPO function for all preference dataset for the plurality of prompts, wherein the INPO function comprises an expectation term, a regularization term, and a modification term; and outputting the (M+1)-th language model.
Owner:TENCENT AMERICA LLC

AUTOMATIC TRANSCRIPT-ASSISTED SPEECH LANGUAGE TRANSLATION USING LANGUAGE MODELS

Devices, systems, and techniques are disclosed that implement the training and deployment of automatic transcription-based translation systems using language models. The techniques include: processing, using a first speech-to-text (S2T) model, an initial input that includes spoken language in a first language to generate a transcription of the spoken language; and processing, using a second S2T model, a second input to generate a translation of the spoken language into a second language. The second input includes at least a representation of the spoken language and the transcription of the spoken language.
Owner:NVIDIA CORP

Dialogue system

A dialogue system, comprising:an input, configured to receive input data from a user, wherein the input data comprises one or more of text data, speech data, image data and motion data; an output, configured to output data to the user; and one or more processors, configured to:obtain information identifying a skill and obtain information identifying a proficiency level of the user for the identified skill from stored proficiency level information; and execute at least one iteration of a coaching session, each iteration comprising performing one or more dialogue interactions, wherein each dialogue interaction comprises:receiving first input data from the user via the input;generating a first language model prompt and providing the first language model prompt to a language model, said first language model prompt comprising the first input data, the information identifying a skill, the information identifying a proficiency level of the user for the identified skill and a request to generate coaching information based on the first input data, the information identifying a skill and the information identifying a proficiency level; and generating first output data based on a first language model response to the first language model prompt and outputting, via the output, the first output data to the user;wherein the at least one iteration of the coaching session further comprises, after the one or more dialogue interactions:generating a second language model prompt and providing the second language model prompt to the language model, said second language model prompt comprising the information identifying a skill, the information identifying a proficiency level of the user for the identified skill, the first input data and the first output data, and a request to generate at least one proficiency update assessment based on the first input data, the first output data, the identified skill and the information identifying a proficiency level; generating second output data based on a second language model response to the second language model prompt and outputting, via the output, the second output data to the user; receiving second input data from the user via the input; determining a revised proficiency level of the user for the identified skill based on the second input data; andupdating the stored proficiency level information based on the revised proficiency level.
Owner:MARAHTA AMRICK LAL

Estimation method, recording medium, and estimation device

An estimation method includes: obtaining a first voice feature group of a plurality of persons who speak a first language; obtaining a second voice feature group of a plurality of persons who speak a second language; obtaining a voice feature of a subject; correcting the voice feature of the subject according to a relationship between the first voice feature group and the second voice feature group; estimating, from the voice feature of the subject that has been corrected, an oral function or a cognitive function of the subject by using an estimation process for an oral function or a cognitive function based on the second language; and outputting a result of estimation of the oral function or the cognitive function of the subject.
Owner:PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD

Dialogue system and a dialogue method

A computer-implemented method of controlling an output from a dialogue system, the method comprising:receiving, by way of an input, first input data relating to speech or text provided by a user;selecting a first state from a plurality of states of a deterministic model, at least some states of the plurality of states being associated with a corresponding portion of a language model prompt including an instruction to call a corresponding function;responsive to the first state being associated with a corresponding portion, generating a first language model prompt comprising at least part of the corresponding portion associated with the selected first state;providing the first language model prompt as input to a language model to generate a first language model output;determining whether to execute a function based on the first language model output;responsive to determining to execute a first function based on the first language model output, executing the determined first function to generate a first function output;selecting a second state from the plurality of states based on the first function output;responsive to the second state being associated with a corresponding portion of a language model prompt, generating a second language model prompt comprising at least part of the corresponding portion associated with the selected second state;providing the second language model prompt as input to the language model to generate a second language model output;determining whether to provide an output to the user based on the second language model output; andresponsive to determining to provide an output to the user based on the second language model output, outputting, by way of an output, speech or text to the user.
Owner:POLYAI LTD

Translation method, target information determining method, related apparatus, and storage medium

A translation method is provided, including: encoding to-be-processed text information to obtain a source vector representation sequence, the to-be-processed text information belonging to a first language; obtaining a source context vector corresponding to a first instance according to the source vector representation sequence, the source context vector indicating to-be-processed source content in the to-be-processed text information at the first instance; determining a translation vector according to the source vector representation sequence and the source context vector; and decoding the translation vector and the source context vector, to obtain target information of the first instance, the target information belonging to a second language.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Personally identifiable information scrubber with language models

Sanitizing data can be a cumbersome task, particularly when the volume of data is large, the content is sensitive, and / or the type of sanitation requires contextual determinations. Sanitizing large amounts of data is tedious and may often require highly trained personnel with clearances and / or other qualifications. In the systems and methods of the present disclosure, language models (LMs) are used to solve these and other technical issues with tools that may allow sanitizing data easily, with high versatility, context awareness, and / or low demand for computational resources. In particular, some of the disclosed systems and methods use a first language model and a second language model (being less resource-intensive than the first language model) to generate sanitized output data with improved efficiency and accuracy. This dual-model approach ensures that sensitive information is handled appropriately while optimizing computer resource usage.
Owner:OPENAI OPCO LLC

Audio translation with preserved speaker characteristics

A method includes receiving an audio stream from a first user associated with a first client device, wherein the audio stream is spoken in a first language by the first user. The method further includes retrieving translation data associated with a second user. The method further includes converting a first portion of the audio stream received from the first user into a plurality of phonemes of a second language, wherein the second language is defined by the language preference. The method further includes predicting a respective duration of each of the phonemes in the plurality of phonemes. The method further includes outputting, by a synthesizer, a first portion of output speech that includes the plurality of phonemes where each of the phonemes in the plurality of phonemes has the respective duration. The method further includes providing the first portion of output speech to the second user device.
Owner:ROBLOX CORP

Method of recognizing speech, device, and medium

A method of recognizing a speech, a device, and a medium. The method includes: processing, by using an acoustic model, speech data to be recognized and a first text segment obtained by recognition to obtain respective acoustic probabilities of a plurality of candidate text segments; processing the first text segment by using a first language sub-model to obtain respective initial language probabilities of the plurality of candidate text segments; processing the first text segment by using a constraint sub-model to obtain extendibility relationships of the plurality of candidate text segments with respect to the first text segment; adjusting the initial language probabilities of the candidate text segments according to the extendibility relationships to obtain respective first language probabilities of the plurality of candidate text segments; and determining a target text segment from the plurality of candidate text segments according to the first language probabilities and the acoustic probabilities.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Generating subset(s) of candidate languages for translation application(s)

Various implementations include initiating, at a client device, a translation application for translation of a dialog session between a first user speaking in a first language and a second user speaking in a second language. In many implementations, the first language, spoken by the first user, can be determined based on one or more features of the client device. Additional or alternative implementations include determining a subset of candidate second languages from a plurality of languages available to the translation application. In a variety of implementations, the system can render output based on the subset of candidate second languages, and can process received input from the second user indicative of one or more of the candidate second languages in the subset of candidate second languages.
Owner:GOOGLE LLC

Data processing method and apparatus

The present application provides a data processing method and apparatus. The method comprises: acquiring a natural language processing (NLP) training data set; inputting the NLP training data set into a first large language model to train the first large language model, the trained first large language model comprising a first language modeling parameter; acquiring the first language modeling parameter; obtaining first text data on the basis of the NLP training data set; and obtaining a first prompt word set on the basis of the first text data, the first language modeling parameter, and a second large language model, wherein the first prompt word set and the first text data are used for training a third large language model, and the first large language model and the second large language model are based on the same model architecture. The method provided by the present application can obtain the first prompt word set of which the data volume is much smaller than that of the NLP training data set. In this way, using the first prompt word set to train the third large language model can improve the speed and accuracy of large language model training.
Owner:HUAWEI TECH CO LTD

Multilingual support for natural language processing applications

A data processing system implements obtaining textual content in a first language from a first client device and segmenting the textual content into a plurality of first tokens. The system also implements translating the first tokens from the first language to a second language using a bilingual dictionary, extracting features information from the second tokens to create a features vector, providing the feature vector to a first natural language processing model trained to analyze textual input in the second language and to output contextual information indicating one or more topics or subject matter of the first textual content, and providing the contextual information to a first machine learning model configured to analyze the contextual information and to identify one or more content items predicted to be relevant to the contextual information. The system further implements providing the information identifying the one or more content items to the first client device.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Knowledge refinement for language model enhancement

Certain embodiments of the disclosure provide techniques for knowledge refinement for language model fine-tuning. A method generally includes obtaining a raw information item; partitioning the raw information item into a plurality of first contextual units, wherein each first contextual unit comprises a first portion of the raw information item; generating, via a first language model, first synthetic data based on: the plurality of first contextual units; the raw information item; and at least one of fine-grained synthesis, interleaved generation, or assembly augmentation; and fine-tuning a second language model based on the first synthetic data.
Owner:INTUIT INC

Automated generation of targeted feedback using speech characteristics extracted from audio samples to address speech defects

To provide a system and method for providing speech instructions based on classification of speech of language communication from a user.SOLUTION: A computing system can generate speech characteristics of a first verbal communication using a first audio sample from a user. The computing system can determine a first speech classification of the first verbal communication based on the speech characteristics from a plurality of speech classifications. The computing system can select one action from a plurality of actions that includes modifying one or more speech characteristics that define a user's speech based on the first speech classification. The computing system can provide instructions to present a message prompting the user to perform the speech defined by this one action selected from the plurality of actions. The efficacy of the medication that the user is taking to address the condition can be increased.SELECTED DRAWING: Figure 1
Owner:CLICK THERAPEUTICS INC