Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

328 results about "Spoken language" patented technology

A spoken language is a language produced by articulate sounds, as opposed to a written language. Many languages have no written form and so are only spoken. An oral language or vocal language is a language produced with the vocal tract, as opposed to a sign language, which is produced with the hands and face. The term "spoken language" is sometimes used to mean only vocal languages, especially by linguists, making all three terms synonyms by excluding sign languages. Others refer to sign language as "spoken", especially in contrast to written transcriptions of signs.

Techniques for determining conversational intent

The present disclosure relates to systems and methods for enhancing the interaction between users and automated agents, such as digital assistants, by employing Large Language Models (LLMs) to infer the intent of spoken language. The invention involves continuously monitoring ambient audio, converting speech to text, and utilizing LLMs to determine whether spoken language is intended for the automated agent. A structured prompt, including the converted text and specific instructions, is sent to the LLM, which is fine-tuned to process domain-specific prompts. The LLM provides a structured output in a standardized format, indicating the user's intent. The system may involve multiple prompts to perform separate tasks, such as identifying intent and generating additional context-specific data. This approach facilitates a more natural and intuitive user experience by eliminating the need for wake words and allowing seamless conversational interaction with virtual assistants across various platforms and devices.
Owner:SNAP INC

Spoken language pronunciation training correction system based on intelligent equipment

The invention belongs to the technical field of intelligent voice processing, and particularly relates to a spoken language pronunciation training and correcting system based on intelligent equipment, which acquires a user rhythm feature set including pitch change rate, accent intensity, pause duration and intonation contour by acquiring spoken language audio data sent by a user for a target text, and corrects the spoken language pronunciation training and correcting system. The method comprises the following steps: acquiring spoken language audio data, converting the spoken language audio data into a phoneme sequence aligned with target text time, generating a rhythm deviation degree report according to comparison evaluation of a user rhythm feature set and the phoneme sequence, an execution intonation mode, accent distribution, speech stream sound change and speech speed rhythm, and determining the rhythm deviation degree according to a rhythm problem type in the report in combination with user historical learning data. And forming and outputting a correction scheme including a text prompt, a targeted minimum contrast training unit and a listen-and-read simulation task, solving the problem of weak capability of correcting hyper-phoneme in a second language spoken language of an adult, and improving the authentic and fluency of the spoken language, thereby realizing efficient personalized learning.
Owner:HUNAN DIGITAL TECHNOLOGY CO LTD

Self-adaptive interactive language learning system capable of multimodal emotion calculation and matching method of self-adaptive interactive language learning system

The invention discloses a multi-modal emotion calculation enabling adaptive interactive language learning system and a matching method thereof, and belongs to the technical field of artificial intelligence, and the system comprises a multi-modal emotion real-time perception and quantification module MHAE-Net, a personalized adaptive dialogue and intervention strategy module ADIS-RL, and a virtual dialogue partner interactive interface module EC-NLG. The multi-modal emotion real-time perception and quantification module MHAE-Net is composed of a voice signal acquisition and processing unit, a voice emotion analysis unit, a text emotion analysis unit, a facial expression emotion analysis unit, a physiological signal emotion analysis unit and a multi-modal emotion fusion and decision-making unit. The interactive strategy is dynamically adjusted, the oral anxiety of the language is effectively relieved, the oral confidence is improved, the system adopts the multi-mode emotion calculation technology, the accuracy and robustness of emotion recognition are improved, the targeted interactive strategy is designed for different emotion states, and personalized emotion support and language practice are provided for learners.
Owner:SHENZHEN XIXING INTELLIGENT TECHNOLOGY CO LTD

Personalized Russian spoken language practice recommendation method and system based on artificial intelligence

The invention relates to the technical field of artificial intelligence education, in particular to a Russian spoken language practice personalized recommendation method and system based on artificial intelligence, and the method comprises the steps: 1, outputting a phoneme sequence with a timestamp through Russian automatic voice recognition; 2, collecting an exercise interruption position and repeated read-after behavior data; 3, generating a dynamic learner portrait; 4, mapping high-frequency errors in the learner portrait into abnormal path weights of map nodes; 5, a lattice tail error option and a non-matching body verb interference item are injected; 6, when the voice fluency attenuation of the learner exceeds a dynamic threshold value, the sentence complexity is reduced; and 7, calculating an error rate descent gradient based on the exercise completion data, and dynamically adjusting the abnormal path weight of the knowledge graph. Through audio stream analysis and syntax tree construction, the system can accurately identify errors of the learner in grammar, pronunciation and other aspects, and the learning efficiency is improved.
Owner:HARBIN UNIV

Spoken language understanding joint model method based on label attention and window mechanism

The invention discloses a spoken language understanding joint model method based on label attention and a window mechanism. The spoken language understanding joint model method is used for improving the effects of intention recognition and slot filling. The method comprises the following steps: firstly, performing semantic coding on a statement input by a user by adopting a self-attention mechanism and Bi-LSTM coding to generate basic semantic representation; then, dynamically adjusting attention distribution of each lexical element through a tag attention mechanism, extracting sentence-level intention and slot tag semantics, and constructing an overall semantic context, so as to form an intention tag attention module and a slot tag attention module; the intention module preliminarily predicts the intention of a statement by using a block-level sliding window, and the slot module preliminarily predicts slot position information by means of a slot classifier. And then, through a graph convolution layer, carrying out adaptive fusion on the preliminarily predicted intention and slot position information, and realizing information interaction between nodes to obtain an updated label embedding representation. And finally, decoding the embedded representation through a classifier module, and generating a final intention and slot position recognition result.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Artificial intelligence (AI)-driven mixed-initiative dialogue digital medical assistant

In accordance with at least one aspect of this disclosure, an artificial intelligence driven bi-directional medical assistant is provided. The assistant comprises an input module configured to recognize, in real time, spoken language and convert the spoken language to a computer readable form to generate a patient embedding. The computer readable form includes, in certain embodiments, a mathematical vector associated with the patient embedding, and the spoken language includes a conversation between a clinician and a patient during a patient visit.
Owner:QUANTUM AI LLC

Correction system, method and equipment for oral vocal training

The invention discloses a correction system, method and device for spoken language vocal training, and relates to the technical field of speech recognition, and the system comprises an environment data analysis module which is used for collecting environment noise parameters in real time, preliminarily judging whether a training environment meets the requirements or not according to the environment noise parameters, and optimizing the training environment; the multi-mode environment adaptability module is used for collecting optimized environment noise parameters and adjusting voice judgment weights and visual judgment weights according to the optimized environment noise parameters; and the spoken language accuracy judgment module is used for collecting pronunciation data and facial image data of a pronunciator, performing pronunciation accuracy evaluation in combination with a voice judgment weight and a visual judgment weight, and providing targeted correction suggestions according to an evaluation result, so that the problem of evaluation deviation caused by failure of a single mode in the prior art is solved, and the accuracy of the pronunciation is improved. The comprehensive evaluation of the pronunciation accuracy is realized, the overall environmental adaptability is improved, and the pronunciation evaluation accuracy is enhanced.
Owner:王海骄

AUTOMATIC TRANSCRIPT-ASSISTED SPEECH LANGUAGE TRANSLATION USING LANGUAGE MODELS

Devices, systems, and techniques are disclosed that implement the training and deployment of automatic transcription-based translation systems using language models. The techniques include: processing, using a first speech-to-text (S2T) model, an initial input that includes spoken language in a first language to generate a transcription of the spoken language; and processing, using a second S2T model, a second input to generate a translation of the spoken language into a second language. The second input includes at least a representation of the spoken language and the transcription of the spoken language.
Owner:NVIDIA CORP

Estimation method, recording medium, and estimation device

An estimation method includes: obtaining a first voice feature group of a plurality of persons who speak a first language; obtaining a second voice feature group of a plurality of persons who speak a second language; obtaining a voice feature of a subject; correcting the voice feature of the subject according to a relationship between the first voice feature group and the second voice feature group; estimating, from the voice feature of the subject that has been corrected, an oral function or a cognitive function of the subject by using an estimation process for an oral function or a cognitive function based on the second language; and outputting a result of estimation of the oral function or the cognitive function of the subject.
Owner:PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD

A multi-intent spoken language understanding method based on syntax analysis

The application discloses a multi-intent spoken language understanding method based on syntax analysis, which comprises the following steps: firstly, obtaining an intent feature matrix and a slot feature matrix according to a user input sentence, and constructing a multi-level intent feature from the intent feature matrix; secondly, obtaining initial intent labels and initial slot prediction labels by using an intent decoding module and a slot decoding module respectively on the intent feature matrix and the slot feature matrix; then, inputting the multi-level intent feature and the initial slot prediction labels into a slot-intent interaction module to obtain an enhanced intent feature matrix, and inputting the slot feature matrix, the multi-level intent feature and the initial intent labels into an intent-slot interaction module to obtain an enhanced slot feature matrix; finally, inputting the enhanced intent feature matrix and the slot feature matrix into the intent decoding module and the slot decoding module respectively to obtain intent labels and slot sequence labels of the spoken language understanding task. The application improves the accuracy of intent recognition and slot sequence labeling and the accuracy of multi-intent recognition.
Owner:HANGZHOU DIANZI UNIV

Spoken language evaluation method and device based on deep learning and medium

The invention discloses a spoken language evaluation method and device based on deep learning and a medium, and relates to the technical field of spoken language evaluation, and the method comprises the steps: collecting spoken language audio signals, carrying out the acoustic feature extraction of the spoken language audio signals through Mel-frequency cepstrum coefficient transformation, and generating an acoustic feature vector sequence; constructing a deep learning pronunciation diagnosis model, inputting the acoustic feature vector sequence into the deep learning pronunciation diagnosis model, calculating a multi-dimensional distance between each voice segment in the acoustic feature vector sequence and the phoneme prototype in a measurement space, and generating a pronunciation diagnosis result; converting the acoustic feature vector sequence into a text sequence through a speech recognition conversion method; and constructing a deep learning role analysis model, and inputting the text sequence into the deep learning role analysis model to generate a semantic role graph. According to the method, the acoustic deviation between the quantized speech segment and the standard phoneme in the measurement space is calculated through the phoneme prototype distance, and accurate space-time positioning and quantitative guidance of the pronunciation defect are realized.
Owner:CHANGCHUN VOCATIONAL INST OF TECH

Industrial automation design environment prompt engineering for generative AI

An integrated development environment (IDE) for designing, programming, and configuring aspects of an industrial automation system uses a generative artificial intelligence (AI) model and associated neural networks to generate portions of an industrial automation project in accordance with functional requirements provided to the industrial IDE system in intuitive formats, such as spoken or written plain language text. The system uses generative AI to translate plain language requests or functional specifications into industrial control code, human-machine interface (HMI) applications, device configuration settings, or other aspects of an industrial control project.
Owner:ROCKWELL AUTOMATION TECH INC

Text understanding model training method and device, equipment and storage medium

The invention discloses a text understanding model training method and device, equipment and a storage medium. Comprising the steps of obtaining a sample text and a sample label set; inputting the plurality of sample intention tags and the plurality of sample slot tags into a preset text understanding model for semantic relationship coding processing to obtain an intention tag relationship representation and a slot tag relationship representation; performing intention decoding processing on the sample text representation and the intention label relation representation based on a preset text understanding model to obtain a predicted intention result, and performing slot decoding processing on the sample text representation and the slot label relation representation to obtain a predicted slot result; determining target loss according to the prediction intention result, the target intention label, the prediction slot position result and the target slot position label; and performing iterative training on a preset text understanding model based on the target loss to obtain a trained text understanding model. The method can improve the prediction accuracy of intention prediction and slot prediction, and can be used for various scenes such as artificial intelligence, spoken language understanding, task-based dialogues and the like.
Owner:腾讯医疗健康(深圳)有限公司

Artificial intelligence and machine learning for transcription and translation for media editing

Dialog in a language unfamiliar to an editor poses obvious challenges during the media editing process. It is nearly impossible to edit media containing spoken dialog without a clear comprehension of the underlying language. The methods described here use a combination of artificial intelligence and machine learning models to generate a language proxy in which the dialog is translated into a language that is familiar to the media editor. The editor is then able to edit the media composition in their own language. To generate an edited media composition with spoken dialog in the original language, the edited language proxy is synchronized with and linked back to the original media. The methods combine automatic speech recognition, translation, speech to text, and voice cloning together with existing non-AI technologies such as captioning and media relinking.
Owner:AVID TECHNOLOGY INC

System and method for utterance language understanding

To provide a system and method for performing utterance language understanding with fewer computing resources.SOLUTION: A computer implemented method for carrying out utterance language understanding executes: receiving data representing audio including a voice; processing data using a model for determining a text corresponding to a content of the voice; receiving input for carrying out a language understanding task including a presentation based on a text of one or more semantic labels; processing input using at least a portion of the model to extract semantic information from the text corresponding to the content of the voice; and acquiring the semantic information extracted in connection with the language understanding task.SELECTED DRAWING: Figure 1
Owner:KK TOSHIBA

Sign language animation generation method and device based on semantic analysis, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a sign language animation generation method, device, equipment and medium based on semantic parse. The sign language animation generation method comprises the steps that voice input is received and recognized as text content, field semantic parse is conducted on the text content to generate a field semantic template, and the field semantic template is used for generating a sign language animation; and converting the domain semantic template into a sign language intermediate representation sequence, generating a three-dimensional sign language action sequence based on the sign language intermediate representation sequence, rendering the three-dimensional sign language action sequence into a virtual image sign language animation, and displaying the virtual image sign language animation. According to the invention, by fusing speech recognition, semantic analysis and three-dimensional action rendering, direct conversion from spoken language content to sign language animation is realized, and a complete visual expression link from speech to sign language is formed, so that a user can intuitively understand the speech content in a sign language form through a virtual image, and the user experience is improved. Therefore, the barrier-free performance of human-computer interaction and the accuracy of information transmission are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

System and method for speech-to-text conversion of passenger announcements on board of an aircraft

A system for speech-to-text conversion of passenger announcements on board of an aircraft comprises a cabin management control configured to provide a speech signal related to an announcement to passengers on board of the aircraft; a cabin application server comprising a processing unit configured to convert the speech signal into a text message containing text corresponding to spoken words of the speech signal, wherein the processing unit is configured to convert the speech signal locally without accessing online computing resources; and a network interface configured to provide the text message to at least one passenger device on board of the aircraft.
Owner:AIRBUS OPERATIONS GMBH

Techniques for determining conversational intent

The present disclosure relates to systems and methods for enhancing the interaction between users and automated agents, such as digital assistants, by employing Large Language Models (LLMs) to infer the intent of spoken language. The invention involves continuously monitoring ambient audio, converting speech to text, and utilizing LLMs to determine whether spoken language is intended for the automated agent. A structured prompt, including the converted text and specific instructions, is sent to the LLM, which is fine-tuned to process domain-specific prompts. The LLM provides a structured output in a. standardized format, indicating the user's intent. The system may involve multiple prompts to perform separate tasks, such as identifying intent and generating additional context-specific data. This approach facilitates a more natural and intuitive user experience by eliminating the need for wake words and allowing seamless conversational interaction with virtual assistants across various platforms and devices.
Owner:SNAP INC

Intelligent English teaching method and system and storage medium

The invention belongs to the technical field of English teaching methods, and particularly relates to an intelligent English teaching method and system and a storage medium, and the method comprises the steps: constructing a multi-dimensional linguistic feature analysis model, and extracting lexical features, syntactic relationship features and semantic deviation features in a text input by a student through a natural language processing technology; dynamic student portraits are established based on the cognitive psychology theory, cognitive level labels are updated according to real-time learning data, and the data comprise grammar error clustering distribution, spoken language fluency indexes and vocabulary association response time; and generating a personalized teaching path, matching teaching materials from the hierarchical resource library according to the cognitive level label, and dynamically adjusting the complexity and presentation form of a teaching strategy. Lexical, syntactic and semantic deviation features are extracted through a natural language processing technology, student error types can be accurately positioned, the problem that traditional error correction only stays on surface modification is avoided, and cognitive tags are updated based on real-time learning data.
Owner:HUBEI UNIV OF ARTS & SCI

System and method for bidirectional automatic sign language translation and production

An embodiment relates to a system and method for bidirectional automatic sign language translation and visualization. of the system includes at least two communication-capable devices for receiving and processing information from the system's input and / or output and showing the output of the system. At least two individuals are communicating with each other, both using different modes of communication such as sign language and spoken language. The individuals are able to utilize two separate computing devices, with the system installed or disposed, to translate the information the individuals are signing.
Owner:SAMSUNG ELECTRONICS CO LTD

Text-driven end-to-end AI virtual digital human generation method and system

The invention discloses an end-to-end AI virtual digital human text broadcast video generation method, which comprises the following steps: configuring the tone and speed of a speaker and the spoken language connection, laughter, pause and emotional degree of audio, converting text content into a corresponding audio file, and carrying out alignment operation on the audio file and the text content; extracting a human face frame to obtain potential features of the human face, and splicing all image frames containing the human face and the potential features of the human face in an inverted order to obtain image features; obtaining coding features of the audio, and obtaining audio features corresponding to each frame in combination with the frame rate of the audio; splicing the audio features and the image features of each frame, predicting and decoding potential features, and obtaining a face picture of each frame after prediction; and covering each frame of face picture after prediction back to the original frame, combining the audio and the video, adding the generated subtitles, and generating a final AI virtual digital human text broadcast video. According to the invention, a real-time high-quality digital human image can be obtained.
Owner:SUZHOU AEROSPACE INFORMATION RES INST +1

System and method for providing an interactive virtual classmate to promote student participation in a virtual reality classroom

A virtual learning system is introduced herein that provides a virtual reality classroom environment having a virtual agent. The virtual agent plays the role of an active student and promotes classroom participation through verbal and nonverbal interactions with the teacher and other students. The virtual agent, which is embodied as a 3D virtual avatar in virtual reality, interacts with an actual teacher and actual students with both spoken language and body gestures. The behaviors of the virtual agent avatar help to encourage other students to participate more actively in the virtual classroom environment.
Owner:PURDUE RES FOUND

Model training method, written language spoken language conversion method, device and product

The invention provides a model training method, a written language spoken language conversion method, equipment and a product, which are applied to the field of natural language processing. The model training method comprises the steps that supervised training is conducted on a large language model based on first training data and task cues indicating to convert written languages into spoken languages, a supervised training model is obtained, and the first training data comprises supervised data pairs formed by written language texts and spoken language texts; based on second training data and a supervised training model, multiple rounds of reinforcement learning training are carried out, a written language spoken language model is obtained, and the second training data comprises a non-labeled written language data set and preference data used for generating reward signals. Therefore, the spoken language model of the written language is obtained by training after reinforcement learning of the large language model, and the quality of the spoken text transferred by the spoken language model of the written language is improved.
Owner:IFLYTEK CO LTD

Spoken language understanding method based on multi-view expert fusion and interest word selection

This invention discloses a spoken language understanding method based on multi-view expert fusion and interest lexical selection, belonging to the field of natural language understanding and semantic parsing technology. The method includes: extracting hidden state sequences from the input utterance using a shared encoder; constructing a multi-view expert fusion module containing utterance view experts, block view experts, and lexical view experts to generate multi-granularity expert features; generating intent aggregation features and slot aggregation features through a task decoupling gating mechanism; performing interest lexical selection on the intent aggregation features, calculating lexical importance scores, generating weighted intent representations, and predicting multi-intent labels; and inputting the slot aggregation features into a decoder with diagonal mask constraints to generate a position-aligned slot label sequence. This invention enhances the model's dynamic focusing ability on key intent signals and can be applied to intelligent dialogue systems, virtual assistants, and vertical domain semantic parsing scenarios.
Owner:JIANGNAN UNIV

Methods and systems for support of multi-language user sessions and fulfillments

Described herein are methods, systems, and media for supporting multi-language user sessions and fulfillments comprising: maintaining a repository of fulfillment objects each comprising a language and a region; establishing a user session with a user; determining a user region for the user in association with establishing the user session; identifying one or more fulfillment objects in the repository available for the user region; processing the user session, the user session comprising one or more user requests; applying a language detection model to each request to determine a user request spoken language; applying an understanding module to each request to recommend one or more of the fulfillment objects matching the region for the user; and rendering a response to each request to the user, utilizing the one or more of the fulfillment objects matching the region for the user, in the user request spoken language.
Owner:AUTOMATION ANYWHERE INC

Spoken language assessment methods, apparatuses, related devices, and computer program products

The application discloses a spoken language evaluation method and device, related equipment and a computer program product. The method comprises the following steps: obtaining the answer data of a testee, wherein the answer data comprises a question, the answer audio of the testee and a reference answer; identifying the answer text corresponding to the answer audio; obtaining the reasoning score of the testee by combining the answer text and the answer data and through a configured reasoning scoring model; obtaining a configured calibration model, wherein the calibration model is obtained by pre-training based on the answer text of a calibration testee, the reasoning score of the calibration testee and an expert score; the calibration testee is part of the testees participating in the current oral test; and obtaining the final score of each testee by using the calibration model to score according to the answer text and the reasoning score of each testee. Compared with the prior art which determines the score by simply calculating the similarity between the answer text and the reference answer, the oral evaluation result obtained by the application is more accurate.
Owner:IFLYTEK CO LTD

A brain-computer interface system for recognizing the intention of Chinese oral language based on a sound-meaning integration double model

ActiveCN121560160BSpoken languageStereotaxis
The application provides a Chinese spoken language intention recognition brain-computer interface system based on a sound-meaning integration double model, belongs to the technical field of biomedical engineering, and relates to language brain-computer interface technology. Taking sound-meaning integration as the core, the stereotactic intracranial electroencephalogram (sEEG) technology is adopted to collect neural signals of the brain articulatory motor coding area and the semantic concept organization coding area. The system comprises a voice initiation decoder, a speech decoder, a semantic decoder and a Chinese word speech-semantic fusion synthesizer, the target decoder is constructed by extracting high gamma band features of key brain areas of the frontal lobe (left inferior frontal gyrus, premotor cortex, etc.), the temporal lobe (anterior temporal lobe, dorsolateral temporal lobe, etc.). At the same time, a visual and auditory induction training paradigm is matched, three tasks of listening to sound to group words, looking at words to group words and word association are set, and the subjects are supported to generate words independently. The system effectively solves the homonym and near homonym word ambiguity problem in Chinese spoken language recognition, and provides a precise interactive tool for ALS and other speech disorder patients.
Owner:BEIJING TIANTAN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV

Chunk-wise attention for longform ASR

A method includes receiving training data including a corpus of multilingual unspoken textual utterances, a corpus of multilingual un-transcribed non-synthetic speech utterances, and a corpus of multilingual transcribed non-synthetic speech utterances. For each un-transcribed non-synthetic speech utterance, the method includes generating a target quantized vector token and a target token index, generating contrastive context vectors from corresponding masked audio features, and deriving a contrastive loss term. The method also includes generating an alignment output, generating a first probability distribution over possible speech recognition hypotheses for the alignment output, and determining an alignment output loss term. The method also includes generating a second probability distribution over possible speech recognition hypotheses and determining a non-synthetic speech loss term. The method also includes pre-training an audio encoder based on the contrastive loss term, the alignment output loss term, and the non-synthetic speech loss term.
Owner:GOOGLE LLC