Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

17 results about "Speech comprehension" patented technology

Speech comprehension starts with the identification of the speech signal against an auditory background and its transformation to an abstract representation, also called decoding. Speech sounds are perceived as phonemes, which form the smallest unit of meaning.

Film and television play table book extraction method and device, storage medium and computer equipment

According to the movie and television play table book extraction method and device, the storage medium and the computer equipment provided by the invention, after an audio and video file of a movie and television play is split into a video file and an audio file, feature recognition is performed on the video file to obtain a subtitle text, speaker face information and a video understanding text; performing voice understanding on the audio file to obtain a voice transcription text and a voice understanding text; wherein the voice transcription text can be corrected into the standard transcription text with high accuracy through the subtitle text. Therefore, based on the face information of the speaker, the line segment of each speaker in the standard transcriptional text and the audio and video file is aligned, so that speaker information with accurate segmentation and semantic coherence can be obtained; and then, through combination with a character side-writing text generated by side-writing analysis on the speaker based on the video, the voice understanding text and the speaker information, table book information is constructed, and related feature description of the character can be covered on the basis of containing the line content, so that the content and depth of the table book are enriched.
Owner:GUANGZHOU QUWAN NETWORK TECH CO LTD +1

Intelligent voice interaction system and method based on streaming multi-mode fusion and equipment control protocol

PendingCN121260156ASpeech recognitionSpeech synthesisSpeech comprehensionEngineering
The embodiment of the invention discloses an intelligent voice interaction system and method based on streaming multi-mode fusion and an equipment control protocol, the system comprises a voice input processing module, a voice understanding and generating module and a voice synthesis module, the voice input processing module is used for converting an audio signal into a first token sequence, and the first token sequence is used for converting the audio signal into a second token sequence; the voice understanding and generating module is used for determining a response token sequence according to the first token sequence on the basis of a multi-modal Transform architecture so as to realize voice understanding and generation; and the voice synthesis module is used for synthesizing the response token sequence into an output audio so as to carry out at least one of the following adjustments on the converted audio of the response token sequence: emotion parameter adjustment, tone adjustment and rhythm adjustment. By adopting the embodiment of the invention, low-delay and high-naturalness intelligent voice interaction can be realized, multi-modal fusion and equipment control are supported, and the user experience is remarkably improved.
Owner:SHENZHEN HUANZHI TECHNOLOGY CO LTD

Multi-modal voice interaction large model training method and system based on voice acoustic feature regulation and control, terminal equipment and medium

ActiveCN120954388ASpeech recognitionSpeech synthesisSpeech comprehensionModal voice
The invention discloses a multi-modal voice interaction large model training method and system based on voice acoustic feature regulation and control, terminal equipment and a medium, and relates to the technical field of multi-modal voice interaction.The method comprises the steps that a text token of a text training sample is obtained, a corresponding voice token is constructed, and pre-training data used for converting the text token into the voice token is obtained; in combination with the multi-modal input sample and the pre-training data, constructing fine-tuning training data for speech understanding and dialogue generation; constructing and pre-training a basic model by using the pre-training data; a multi-modal voice interaction large model is constructed based on a pre-training basic model, and fine tuning data is used for training, so that voice acoustic features can be regulated and controlled based on multi-modal input, and voice is output. According to the method, through alignment and staged training of the text token and the voice token, fine regulation and control of voice acoustic characteristics are realized, long voice continuity and interaction naturalness are improved, and a model is efficiently endowed with voice interaction capability of controllable timbre and emotion.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Method for voice recognition, in particular in a motor vehicle

PCT designated stageWO2026008826A1Speech recognitionSpeech comprehensionEngineering
The invention relates to a method for assisting voice recognition, in particular in a motor vehicle, adapted to a predetermined objective of a user and using: - at least one voice sensor arranged in the vehicle, in particular a microphone, connected to a voice assistant at least carrying out natural voice comprehension processing, - a predetermined list of a number N of questions Qi=1...N previously asked and learned according to the predetermined objective and / or the driving context and in-vehicle context, each question corresponding to a state, the list being translated into a system of states, and - a transition matrix representing the possible transitions between each state of the system of states, the transitions having been previously learned.
Owner:AMPERE SAS

Robot voice interaction optimization method based on voice recognition

InactiveCN121641025ASpeech recognitionSpeech comprehensionEngineering
The invention discloses a robot voice interaction optimization method based on voice recognition, and the method comprises the following steps: collecting the voice of a user, associating the voice with an interaction stage, and obtaining the execution risk of a robot; based on robot self-sounding reference, extracting echo, continuity and time-frequency stability evidence, and generating voice availability description according to an interaction stage; performing recognition and semantic analysis on the voice, and outputting recognition uncertainty and semantic ambiguity in combination with a voice availability constraint candidate result; voice availability, recognition uncertainty, semantic ambiguity and execution risk are fused, interaction risk assessment is generated, and an interaction strategy is selected; and executing the interaction strategy and updating the rule and the strategy based on execution feedback to realize closed-loop optimization of voice interaction. Through the interaction reliability evidence compiling and risk fusion assessment method, voice understanding and risk collaborative decision execution are realized, and the method has the advantages of low false execution, high interaction controllability and high self-optimization capability.
Owner:YUNZHOU INNOVATION TECH (GUANGZHOU) CO LTD

over-ear hearing aids for age-related hearing loss

This invention relates to a hearing aid, specifically a hearing aid for elderly people to compensate for and improve hearing loss. The user can select an amplification gain value that matches their hearing level from multiple programs of divided amplification gain values ​​via a selection unit. The user listens by covering their ear with the main body of the hearing aid, thereby improving speech comprehension and sound quality.
Owner:株式会社MP公司

Speech recognition-based image generation system and method

PendingDE102025130325A1Semantic analysis2D-image generationSpeech comprehensionDisplay device
A system and a method for generating an image based on speech recognition are disclosed. The system for generating an image based on speech recognition includes: a speech recognition device configured to capture a user's speech information and convert the captured speech information into textual user request information; a speech understanding device electrically connected to the speech recognition device and configured to analyze the textual user request information and extract semantic information from the user; and a cloud image generation device connected to the speech understanding device and configured to generate image data based on the extracted semantic information from the user using a stability diffusion algorithm.and a display device that communicates with the cloud image generation device and is configured to receive an image generated by the cloud image generation device and to decode and display the received image.
Owner:HYUNDAI MOTOR CO LTD +1

Voice recognition technology, particularly in a motor vehicle

ActiveFR3164311A1Speech recognitionSpeech comprehensionEngineering
A speech recognition method, particularly in a motor vehicle. A speech recognition assistance method, particularly in a motor vehicle, adapted to a predetermined user objective and using: - at least one voice sensor located in said vehicle, in particular a microphone, connected to a voice assistant implementing at least natural speech understanding processing, - a predetermined list of a number N of questions Qi=1…N previously asked and learned according to the predetermined objective and / or the driving and vehicle cabin context, each question corresponding to a state, said list being translated into a system of states, and - a transition matrix representing the possible transitions between each state of said system of states, the transitions having been previously learned. Figure for the abstract: Fig. 5
Owner:AMPERE SAS

A method and device for extracting a script of a movie or TV series, a storage medium and a computer device

The film and television script extraction method and device, the storage medium and the computer device provided by the application, after the audio and video file of the film and television is split into a video file and an audio file, the video file is subjected to feature recognition to obtain a subtitle text, a speaker face information and a video understanding text; and the audio file is subjected to speech understanding to obtain a speech transcription text and a speech understanding text; wherein the speech transcription text can be corrected into a standard transcription text with higher accuracy through the subtitle text. Therefore, based on the speaker face information, the standard transcription text and the dialogue segment of each speaker in the audio and video file are aligned, the speaker information with accurate segmentation and coherent semantics can be obtained, then the script information is constructed by combining the character profile text generated by the profile analysis of the speaker based on the video, the speech understanding text and the speaker information, the script information can cover the related feature description of the character on the basis of the dialogue content, thereby enriching the content and depth of the script.
Owner:GUANGZHOU QUWAN NETWORK TECH CO LTD +1

An off-line interest point recommendation method based on multi-source sensors and speech understanding

This invention belongs to the field of navigation and location services technology, and discloses an offline point of interest (POI) recommendation method based on multi-source sensors and speech understanding. The method includes the following steps: collecting a user's voice request and performing offline speech recognition to obtain text request information; performing intent understanding and preference extraction based on the obtained text request information; acquiring environmental perception data related to the current location and historical movement trajectory based on multi-source sensors; calculating the similarity between the user's needs and local POI data using cosine similarity, and sorting and filtering to obtain a set of target POIs; finally, outputting the set of target POIs as the recommendation result. This invention can, in an environment without network connection, calculate and comprehensively score candidate POIs based on the user's subjective experience needs expressed through voice, combined with environmental data collected by multi-source sensors and an offline POI database, thereby outputting POI recommendation results that meet the user's personalized preferences.
Owner:CHENGDU SKYSCANNER MICROSATELLITE TECH CO LTD

Offline interest point recommendation method based on multi-source sensor and voice understanding

The invention belongs to the technical field of navigation and location services, and discloses an off-line interest point recommendation method based on a multi-source sensor and voice understanding, which comprises the following steps of: acquiring a voice request of a user and performing off-line voice recognition to obtain text request information; performing intention understanding and preference extraction on the acquired text request information; acquiring environment sensing data related to the current position and the historical movement track based on a multi-source sensor; and calculating the similarity between the user demand and the local POI data by adopting cosine similarity, sorting and screening to obtain a target interest point set, and finally outputting the target interest point set as a recommendation result. According to the method, experience index calculation and comprehensive scoring can be carried out on the candidate interest points in the environment without network connection according to subjective experience requirements expressed by the user through voice and in combination with the environment data collected by the multi-source sensor and the offline POI database, and therefore the interest point recommendation result meeting the personalized preference of the user is output.
Owner:CHENGDU SKYSCANNER MICROSATELLITE TECH CO LTD

Speech understanding method and system based on text alignment and electronic equipment

The embodiment of the invention provides a speech understanding method and system based on text alignment and electronic equipment. The method comprises the steps that training data are input into a voice understanding model, the training data comprise a training text, and the voice understanding model comprises a CTC posterior simulation module; in a CTC posteriori simulation module, the training text is converted into a pseudo CTC posteriori simulating real audio distribution characteristics; and a pseudo CTC posteriori is utilized to generate a pseudo posteriori supervision projection module, the pseudo posteriori supervision projection module is used for performing projection reasoning on the input voice to obtain a structured semantic alignment posteriori representation, and a large language model is utilized to determine a voice understanding result of the semantic alignment posteriori representation. According to the embodiment of the invention, clean text symbol labels are converted into noise multi-frame false posteriors, the false posteriors are close to the distribution characteristics of real voice, and the real voice is kept compressed efficiently, so that the over-fitting is relieved and the training / reasoning is accelerated. The multi-task expandability and maintainability are improved, and multi-task zero sample generalization is supported.
Owner:AISPEECH CO LTD

AI voice interactive ordering terminal system

PendingCN122658306ASpeech comprehensionVoice activity
The application discloses an AI voice interactive meal ordering terminal system and relates to the technical field of voice dialogue, and comprises the following steps: an audio interface module collects meal ordering voice to perform voice activity segmentation; a voice understanding module performs language detection and standardization processing; a dialogue reasoning module generates a natural language response and a structured action list; an operation verification module verifies the legality of the structured action according to a menu knowledge base, and sends the action failing in the verification to a rollback action recovery module for repair; an order state management module applies the structured action passing the verification to a shopping cart state machine, executes constraint conditions, calculates a price, and generates an order summary; and a voice output module generates voice feedback in a target language, and through the combination of the structured action list and the shopping cart state machine, the conversion from voice to a deterministic order is realized, the consistency of an order state in a multi-round dialogue is ensured, and the fault tolerance is improved through rollback recovery when the action verification fails.
Owner:GUANGZHOU CHUANGTU INNOVATION TECHNOLOGY CO LTD

Speech understanding device, speech understanding method, and program

PCT designated stageWO2025257891A1Speech analysisSpeech comprehensionLearning unit
This speech understanding device understands speech using an input speech, an input sentence, and a speech understanding model, and comprises: a correct answer data generation unit that generates correct answer data including a speech factor phrase, a delimiter, and a correct answer output sentence that represent speech information; and a training unit that trains the speech understanding model using training data including the input speech, the input sentence, and the correct answer data.
Owner:NT T INC

Effective speech detection method and apparatus

The application provides an effective speech detection method and device, the method comprises the following steps: extracting audio features of an audio signal to be detected based on a feature extraction model; determining effective speech signals in the audio signal based on a first effective speech recognition model and applying the audio features; the feature extraction model and the first effective speech recognition model constitute a first detection model, the first detection model is jointly trained with a speech understanding model in a training stage, the speech understanding model takes the audio features extracted by the feature extraction model as input and is used for predicting speech content, and a total loss value of the joint training includes an effective speech detection loss value of the first detection model and a speech understanding loss value of the speech understanding model. The first detection model can be trained with the aid of the speech understanding task, so that the first detection model can avoid missing the effective speech, that is, the ability of the first detection model to detect the effective speech is improved.
Owner:HEFEI IFLY DIGITAL TECH CO LTD

A multi-modal speech interaction large model training method and system based on speech acoustic feature regulation, a terminal device, and a medium

ActiveCN120954388BSpeech recognitionSpeech synthesisSpeech comprehensionModal voice
The application discloses a kind of based on multi-modal speech interaction big model training method, system, terminal equipment and medium of voice acoustic feature regulation, it is related to multi-modal speech interaction technical field, the method includes: obtaining the text token of text training sample and constructs corresponding voice token, obtains the pre-training data for converting text token into voice token;Combining multi-modal input sample and pre-training data, construct fine-tuning training data for speech understanding and dialogue generation;Using pre-training data constructs and pre-trains basic model;Based on pre-training basic model, build multi-modal speech interaction big model, with fine-tuning data training, so that it can be based on multi-modal input regulation voice acoustic feature and output voice.The application is trained by the alignment and stage of text token and voice token, realize the fine regulation of voice acoustic feature, improve long speech coherence and interactive naturalness, efficiently give model controllable timbre, the speech interaction ability of emotion.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

A method and system for generating a multi-dimensional speech training program

ActiveCN121687374BAchieve precise quantitative analysisImprove targetingSpeech trainingSpeech comprehension
The application discloses a kind of multi-dimension speech training plan generation method and system.The method first collects patient age, gender and speech sample, obtains seven-dimensional evaluation parameters including sound pressure, amplitude perturbation, maximum vocalization duration, fundamental frequency perturbation, vowel space area, tone impairment and speech comprehension score.Subsequently, according to the preset logic, these parameters are sequentially determined based on the parameters, and dynamically combine different training modules such as loudness, breath, pitch, vowel, glide, tone and consonant into personalized speech training plan.The application solves the problem of existing technology training scheme solidification through multi-dimensional evaluation and dynamic module matching, significantly improves the individualization degree and rehabilitation effect of speech training.
Owner:BEIJING REHABILITATION HOSPITAL CAPITAL MEDICAL UNIVERSITY(BEIJING WORKERS SANATORIUM)