Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

7 results about "Speech comprehension" patented technology

Speech comprehension starts with the identification of the speech signal against an auditory background and its transformation to an abstract representation, also called decoding. Speech sounds are perceived as phonemes, which form the smallest unit of meaning.

Robot voice interaction optimization method based on voice recognition

InactiveCN121641025ASpeech recognitionSpeech comprehensionEngineering
The invention discloses a robot voice interaction optimization method based on voice recognition, and the method comprises the following steps: collecting the voice of a user, associating the voice with an interaction stage, and obtaining the execution risk of a robot; based on robot self-sounding reference, extracting echo, continuity and time-frequency stability evidence, and generating voice availability description according to an interaction stage; performing recognition and semantic analysis on the voice, and outputting recognition uncertainty and semantic ambiguity in combination with a voice availability constraint candidate result; voice availability, recognition uncertainty, semantic ambiguity and execution risk are fused, interaction risk assessment is generated, and an interaction strategy is selected; and executing the interaction strategy and updating the rule and the strategy based on execution feedback to realize closed-loop optimization of voice interaction. Through the interaction reliability evidence compiling and risk fusion assessment method, voice understanding and risk collaborative decision execution are realized, and the method has the advantages of low false execution, high interaction controllability and high self-optimization capability.
Owner:YUNZHOU INNOVATION TECH (GUANGZHOU) CO LTD

over-ear hearing aids for age-related hearing loss

This invention relates to a hearing aid, specifically a hearing aid for elderly people to compensate for and improve hearing loss. The user can select an amplification gain value that matches their hearing level from multiple programs of divided amplification gain values ​​via a selection unit. The user listens by covering their ear with the main body of the hearing aid, thereby improving speech comprehension and sound quality.
Owner:株式会社MP公司

A method and device for extracting a script of a movie or TV series, a storage medium and a computer device

The film and television script extraction method and device, the storage medium and the computer device provided by the application, after the audio and video file of the film and television is split into a video file and an audio file, the video file is subjected to feature recognition to obtain a subtitle text, a speaker face information and a video understanding text; and the audio file is subjected to speech understanding to obtain a speech transcription text and a speech understanding text; wherein the speech transcription text can be corrected into a standard transcription text with higher accuracy through the subtitle text. Therefore, based on the speaker face information, the standard transcription text and the dialogue segment of each speaker in the audio and video file are aligned, the speaker information with accurate segmentation and coherent semantics can be obtained, then the script information is constructed by combining the character profile text generated by the profile analysis of the speaker based on the video, the speech understanding text and the speaker information, the script information can cover the related feature description of the character on the basis of the dialogue content, thereby enriching the content and depth of the script.
Owner:GUANGZHOU QUWAN NETWORK TECH CO LTD +1

An off-line interest point recommendation method based on multi-source sensors and speech understanding

This invention belongs to the field of navigation and location services technology, and discloses an offline point of interest (POI) recommendation method based on multi-source sensors and speech understanding. The method includes the following steps: collecting a user's voice request and performing offline speech recognition to obtain text request information; performing intent understanding and preference extraction based on the obtained text request information; acquiring environmental perception data related to the current location and historical movement trajectory based on multi-source sensors; calculating the similarity between the user's needs and local POI data using cosine similarity, and sorting and filtering to obtain a set of target POIs; finally, outputting the set of target POIs as the recommendation result. This invention can, in an environment without network connection, calculate and comprehensively score candidate POIs based on the user's subjective experience needs expressed through voice, combined with environmental data collected by multi-source sensors and an offline POI database, thereby outputting POI recommendation results that meet the user's personalized preferences.
Owner:CHENGDU SKYSCANNER MICROSATELLITE TECH CO LTD

Offline interest point recommendation method based on multi-source sensor and voice understanding

The invention belongs to the technical field of navigation and location services, and discloses an off-line interest point recommendation method based on a multi-source sensor and voice understanding, which comprises the following steps of: acquiring a voice request of a user and performing off-line voice recognition to obtain text request information; performing intention understanding and preference extraction on the acquired text request information; acquiring environment sensing data related to the current position and the historical movement track based on a multi-source sensor; and calculating the similarity between the user demand and the local POI data by adopting cosine similarity, sorting and screening to obtain a target interest point set, and finally outputting the target interest point set as a recommendation result. According to the method, experience index calculation and comprehensive scoring can be carried out on the candidate interest points in the environment without network connection according to subjective experience requirements expressed by the user through voice and in combination with the environment data collected by the multi-source sensor and the offline POI database, and therefore the interest point recommendation result meeting the personalized preference of the user is output.
Owner:CHENGDU SKYSCANNER MICROSATELLITE TECH CO LTD

Speech understanding method and system based on text alignment and electronic equipment

The embodiment of the invention provides a speech understanding method and system based on text alignment and electronic equipment. The method comprises the steps that training data are input into a voice understanding model, the training data comprise a training text, and the voice understanding model comprises a CTC posterior simulation module; in a CTC posteriori simulation module, the training text is converted into a pseudo CTC posteriori simulating real audio distribution characteristics; and a pseudo CTC posteriori is utilized to generate a pseudo posteriori supervision projection module, the pseudo posteriori supervision projection module is used for performing projection reasoning on the input voice to obtain a structured semantic alignment posteriori representation, and a large language model is utilized to determine a voice understanding result of the semantic alignment posteriori representation. According to the embodiment of the invention, clean text symbol labels are converted into noise multi-frame false posteriors, the false posteriors are close to the distribution characteristics of real voice, and the real voice is kept compressed efficiently, so that the over-fitting is relieved and the training / reasoning is accelerated. The multi-task expandability and maintainability are improved, and multi-task zero sample generalization is supported.
Owner:AISPEECH CO LTD

A method and system for generating a multi-dimensional speech training program

ActiveCN121687374BAchieve precise quantitative analysisImprove targetingSpeech trainingSpeech comprehension
The application discloses a kind of multi-dimension speech training plan generation method and system.The method first collects patient age, gender and speech sample, obtains seven-dimensional evaluation parameters including sound pressure, amplitude perturbation, maximum vocalization duration, fundamental frequency perturbation, vowel space area, tone impairment and speech comprehension score.Subsequently, according to the preset logic, these parameters are sequentially determined based on the parameters, and dynamically combine different training modules such as loudness, breath, pitch, vowel, glide, tone and consonant into personalized speech training plan.The application solves the problem of existing technology training scheme solidification through multi-dimensional evaluation and dynamic module matching, significantly improves the individualization degree and rehabilitation effect of speech training.
Owner:BEIJING REHABILITATION HOSPITAL CAPITAL MEDICAL UNIVERSITY(BEIJING WORKERS SANATORIUM)