Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

39 results about "Lip feature" patented technology

Voice interaction method and device based on lip language enhancement, equipment and storage medium

The invention discloses a voice interaction method and device based on lip language enhancement, equipment and a storage medium, and the method comprises the steps: extracting lip language features based on an image sequence of a lip region, and carrying out the feature extraction of a voice signal, and obtaining an audio feature; performing cross-modal fusion coding on the lip language features and the audio features to generate mixed features containing audio-visual information; inputting the mixed features into a large language model, understanding the intention of the interaction object and generating a corresponding semantic reply; and finally, synthesizing into voice and / or converting into characters. According to the invention, by introducing the lip features, additional visual clues are provided for speech recognition, and the robustness and accuracy of speech recognition can be significantly improved; effective fusion coding is carried out on the lip language features and the sound features, and semantic information splitting caused by simple and independent recognition is avoided; and the capability of the large model is fully utilized, so that more natural and more intelligent interaction experience is realized.
Owner:SHENZHEN WANRUI INTELLIGENT TECH CO LTD

Lip language recognition method based on multi-modal attention fusion and Transform model

The invention provides a lip language recognition method based on multi-modal attention fusion and a Transform model, and relates to the technical field of artificial intelligence and computer vision, and the method comprises the steps: carrying out the preprocessing of a collected video stream and an audio stream, obtaining a lip ROI region image sequence and audio features, carrying out the label calibration of the lip ROI region image sequence, and carrying out the recognition of a lip language. Generating a lip language text label with a timestamp, and taking the lip language text label with the timestamp as supervision information to train an end-to-end model based on Transform; based on the lip ROI region image sequence, utilizing a deep network architecture constructed by an adaptive residual attention module and a hierarchical feature extraction mechanism to perform multi-level feature extraction on the lip image to obtain lip features, and based on the lip features and the audio features, adopting an attention mechanism-based strategy to perform fusion to obtain multi-modal features; and based on the multi-modal features, using a trained Transform-based end-to-end model to perform lip language sequence identification so as to obtain a lip language identification result text.
Owner:INSPUR WORLDWIDE SERVICES LTD

3D digital human facial expression synthesis method and visualization system based on audio driving

The invention discloses a 3D digital human facial expression synthesis method based on audio driving and a visualization system. The method comprises the following steps: acquiring a data set; inputting the original audio in the data set into an audio encoder in the model, and extracting audio features; inputting the audio features into a KAN-based decoder in the model to generate 3D digital human facial expression actions; inputting 3D digital human facial expression actions into a graph encoder in the model, and extracting lip features; collaborative optimization and model training of multi-modal features are realized by constructing a joint loss function fusing audio features, 3D facial expression parameters and lip features; and inputting an original audio to be tested into the trained model, and generating corresponding 3D digital human facial expression actions through the audio encoder and the KAN decoder in sequence. According to the invention, the generated 3D digital human facial expression action is more vivid and closer to the real character expression.
Owner:SOUTH CHINA UNIV OF TECH

Low-resource language lip language recognition method and device based on pre-training fine tuning

The invention relates to the technical field of computer vision, in particular to a low-resource language lip language recognition method and device based on pre-training fine tuning. The method comprises the following steps: pre-training a model by using a large amount of English video data sets to ensure that the model obtains strong generalization ability and effective lip feature expression ability; and then, after the weight of the pre-training model is loaded, full-parameter fine tuning is performed on the model through a small amount of Tibetan lip language data sets, so that the challenge of scarcity of Tibetan video data is overcome. In the inference decoding stage, a Transform language model specially aiming at Tibetan text training is introduced, the problem of homonym confusion possibly occurring in the lip language recognition process is effectively reduced, and therefore the accuracy of sentence-level Tibetan lip language recognition is improved. The overall architecture is improved through the innovative structure and method, and effective pure vision lip language recognition of low-resource languages is successfully achieved.
Owner:MINZU UNIVERSITY OF CHINA

Speech recognition and speech synthesis optimization method and system based on large model

The invention provides a voice recognition and voice synthesis optimization method and system based on a large model, and the method comprises the steps: extracting lip motion features, lip feature timestamps and audio feature timestamps through obtaining a real-time voice input signal and an image frame sequence of a user face region, so as to generate an initial space-time offset sequence; generating a dynamic offset compensation parameter sequence based on the initial space-time offset sequence in combination with a reference alignment template in the voice and vision synchronization data set; and performing joint processing on the dynamic offset compensation parameter sequence, the lip motion characteristics and the real-time voice input signal by using a large model to reconstruct a target voice segment, generating a corrected phoneme sequence in combination with the lip motion characteristics, retrieving a mouth shape parameter group corresponding to the corrected phoneme sequence from a phoneme mouth shape mapping rule base, and performing mouth shape correction on the mouth shape parameter group. To generate a voice waveform in phase synchronization with the lip motion; according to the invention, the naturalness and immersion of man-machine interaction and the robustness in a voice missing or delay scene are improved.
Owner:LUSTER LIGHTWAVE CO LTD

Multi-modal speaker authentic identification model training method and electronic equipment

ActiveCN120148521ASpeech recognitionLip featureVideo sequence
The invention discloses a multi-mode speaker authentic identification model training method and electronic equipment, and relates to the field of speaker authentic identification, and the method comprises the steps: obtaining an audio and video sequence sample corresponding to a given speaker; performing lip feature coding on each audio frame in sequence in the audio and video sequence sample to determine a corresponding audio sample lip potential vector sequence; feature coding is carried out on the mouth area of the speaker in each sequential video frame in the audio and video sequence sample so as to determine a corresponding video sample lip potential vector sequence; calculating the audio and video sequence similarity corresponding to the audio sample lip potential vector sequence and the video sample lip potential vector sequence; and training the speaker authentic identification model according to the audio and video sequence similarity and the audio and video speaker alignment label. Therefore, through the similarity comparison that the speaking audio and the lip chart of the to-be-detected person are matched into the same dimension vector through the potential vector, the accuracy of an authentic identification result can be more effectively improved.
Owner:AISPEECH CO LTD

Voice operation method of device, apparatus, and electronic device

A voice operation method of a device, comprising: acquiring a video collected by a camera; acquiring voice information collected by a microphone; detecting a face image in the video; extracting a lip feature and a face feature of the face image; determining a time interval according to the lip feature; intercepting a corresponding audio segment in the voice information according to the time interval; acquiring voiceprint information according to the face feature; and performing voice recognition on the audio segment according to the voiceprint information to acquire voice information. The voice operation method of the device does not need to pre-determine a target user and pre-record voiceprint information of the target user, can autonomously extract voiceprint information of multiple users when multiple users use the device at the same time, and can separate voices one by one. Meanwhile, the voiceprint can be autonomously updated and registered. The voice operation method of the device can significantly improve the voice recognition effect of the device in a noisy or multi-user speaking scene.
Owner:HUAWEI TECH CO LTD

Sound and lip synchronization detection method and device, electronic equipment and storage medium

The invention provides a sound and lip synchronous detection method and device, electronic equipment and a storage medium, relates to the technical field of artificial intelligence, and is suitable for the financial field and the medical field. The method comprises the following steps: encoding a target voice to obtain an initial voice feature, and encoding a target object lip in a target face video to obtain an initial lip movement feature; intercepting audio-visual feature extraction sub-models from N feature converters cascaded in the audio-visual association model; performing voice feature extraction on the initial voice feature and the video zero matrix through an audio-visual feature extraction sub-model to obtain a target voice feature; performing lip feature extraction on the initial lip movement feature and the audio zero matrix through an audio-visual feature extraction sub-model to obtain a target lip movement feature; and carrying out voice and lip synchronization classification on the target voice feature and the target lip movement feature to obtain a voice and lip synchronization category. According to the method, the feature learning difficulty can be remarkably reduced, the audio-visual association modeling precision is improved, and thus the accuracy of audio-lip synchronous detection is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Video generation method, model training method, device and computer program product

The invention discloses a video generation method and device, a model training method and device and a computer program product, and the video generation method comprises the steps: obtaining a target audio and a reference picture used for generating a video, and the reference picture comprises a sound production object; determining a global visual feature of each to-be-generated video frame corresponding to the audio clip according to the clip features of one or more audio clips corresponding to the target audio and the reference image; according to the pronunciation feature of each audio frame of the target audio and the lip feature of the sounding object in the reference picture, determining the lip feature of the sounding object in a to-be-generated video frame corresponding to the audio frame; and generating each video frame according to the lip feature and the global visual feature corresponding to the to-be-generated video frame. Through the scheme provided by the invention, the expression of characters in the generated video is more vivid and natural, the lip action and the audio can be accurately synchronized, and the visual experience of a user is improved.
Owner:BEIJING AUTONAVI YUNMAP TECH CO LTD

Semantic recognition method and device fusing audio and video, equipment and medium

The invention provides an audio and video fused semantic recognition method and device, and the method comprises the steps: carrying out the feature extraction of audio information, and obtaining an audio feature; performing feature extraction on the video information to obtain lip features; aligning and splicing the audio features and the lip features to obtain joint features; performing signal-to-noise ratio estimation on the spectrum feature map to obtain a signal-to-noise ratio, and obtaining an audio weight and a video weight according to the signal-to-noise ratio and a preset weight distribution function; performing feature splitting on the joint feature based on the audio weight and the video weight to obtain an audio branch feature and a video branch feature, and performing feature fusion on the audio branch feature, the video branch feature and the joint feature to obtain a fusion feature; and performing semantic recognition on the fused features to obtain a recognition result. By means of the semantic recognition method and device fusing the audio and the video, the technical problem that the voice recognition effect is poor in a complex acoustic environment is solved.
Owner:HEFEI UNIV OF TECH

Voice translation screen control method and system for interaction

The invention relates to a voice translation screen control method and system for interaction, and relates to the field of intelligent interaction technology, and the method comprises the steps: obtaining a region detection image; performing feature recognition in the region detection image to determine user lip features, and determining a user lip position according to the user lip features; determining an effective pickup distance according to the lip position of the user and the position of each pickup device; determining a pickup sensitive range corresponding to the effective pickup distance according to the pickup matching relationship; randomly selecting a pickup sensitivity value from each pickup sensitivity range, defining the selected pickup sensitivity value as an effective sensitivity value, and controlling each pickup device to work according to the effective sensitivity value to obtain the external voice volume; and determining a user representative volume according to each external voice volume, determining a sensitivity adjustment coefficient according to the user representative volume and the effective recognition volume, and adjusting and updating each effective sensitivity value according to the sensitivity adjustment coefficient. The method and the device have the effect of facilitating subsequent analysis of the collected voice.
Owner:NINGBO LANKE INTELLIGENT ENG

Method for training multi-modal speaker verification model and electronic device

ActiveCN120148521BSpeech recognitionSpeaker verificationFeature coding
The application discloses a training method of a multi-modal speaker authentication model and an electronic device, relates to the field of speaker authentication, and comprises the following steps: obtaining an audio-video sequence sample corresponding to a given speaker; performing lip feature coding on each audio frame in the audio-video sequence sample in sequence to determine a corresponding audio sample lip latent vector sequence; performing feature coding on a speaker mouth region in each video frame in the audio-video sequence sample in sequence to determine a corresponding video sample lip latent vector sequence; calculating an audio-video sequence similarity corresponding to the audio sample lip latent vector sequence and the video sample lip latent vector sequence; and training a speaker authentication model according to the audio-video sequence similarity and an audio-video speaker alignment label. Thus, the similarity of the speaker audio and the lip shape graph of the person to be detected is compared by matching the speaker audio and the lip shape graph into the same dimension vector through the latent vector, so that the accuracy of the authentication result can be effectively improved.
Owner:AISPEECH CO LTD

Multimodal duration-controlled TTS (dc-TTS) system for automatic dubbing

PCT designated stageWO2025186682A8Natural language translationSpeech recognitionData setLip feature
A system and method for multimodal duration-controlled text to speech (TTS) is provided. The system receives a dataset including reference video clip, reference audio, and text. The system generates: text features based on application of a text encoder on the text, visual features based on application of a visual feature extractor on the reference video clip, a first speaker embedding associated with the human speaker based on the reference audio, and lip features associated with the human speaker based on the reference video clip. The system further generates audio tokens based on application of a neural language model on the text features, the visual features, the first speaker embedding, and the lip features. The system generates an audio based on application of a neural vocoder on the audio tokens and trains the neural language model based on the generated audio and a loss function.
Owner:SONY GROUP CORP

A voice translation screen control method and system for interaction

The application relates to a voice translation screen control method and system for interaction, and relates to the field of intelligent interaction technology, which comprises the following steps: acquiring a region detection image; performing feature recognition on the region detection image to determine user lip features, and determining user lip positions according to the user lip features; determining effective sound pickup distances according to the user lip positions and positions of sound pickup devices; determining sound pickup sensitive ranges corresponding to the effective sound pickup distances according to sound pickup matching relationships; randomly selecting a sound pickup sensitive value in each sound pickup sensitive range to define an effective sensitive value, and controlling the sound pickup devices to work at the effective sensitive value to acquire external voice volumes; determining a user representative volume according to the external voice volumes, determining a sensitive adjustment coefficient according to the user representative volume and an effective recognition volume, and adjusting and updating the effective sensitive values according to the sensitive adjustment coefficient. The application has the effect of facilitating subsequent analysis of collected voices.
Owner:NINGBO LANKE INTELLIGENT ENG

A voice-driven talking portrait generation method and an electronic device

The application discloses a speech-driven talking portrait generation method and electronic equipment, wherein the method comprises the following steps: acquiring audio features, spatial features and image features; inputting the audio features and the spatial features into an audio-spatial coding module, aligning the audio features and the spatial features by dynamically adjusting the spatial position of the audio features and the facial region, and obtaining spatially perceived lip features; inputting the audio features and the image features into an audio-guided dynamic fusion module, calculating audio-perceived channel attention and spatial attention by using a dynamic linear layer, and obtaining dynamic fusion features; acquiring reference key points, inputting the reference key points and the audio features into an expression optimization module, fusing the reference key points and the audio reference points, and obtaining key point features; and inputting the lip perception features, the dynamic fusion features and the key point features into a trained generation network to generate video frames. The method realizes a higher-quality, more natural and stable speech-driven portrait synthesis effect.
Owner:SOUTH CHINA UNIV OF TECH

Virtual anchor real-time driving system based on facial motion capture

ActiveCN121842342BResolve driver conflictsImprove viewing experienceTelevision system detailsColor television detailsPhoneme recognitionEngineering
The application discloses a virtual anchor real-time driving system based on facial motion capture, aims to solve the problems of insufficient precision, lack of naturalness and weak scene adaptation of virtual anchor facial motion in the prior art, and through multi-thread parallel collection of audio and video streams, optimizes data quality through adaptive preprocessing; adopts multi-dimensional facial feature collaborative extraction, double-branch phoneme recognition, face function area semantic segmentation and confidence dynamic weight adjustment technology, realizes accurate representation and collaborative fusion of expression and lip feature; finally, through the lip-expression collaborative driving mechanism, the real-time driving characteristics of the adaptive virtual image are output; thereby effectively improving the precision and naturalness of virtual anchor facial motion, enhancing the complex scene adaptability, giving consideration to real-time response and low-cost deployment, and being applicable to virtual live broadcast, online education and other scenes, and having good application value.
Owner:GUIZHOU NORMAL UNIVERSITY

Anti-spoofing authentication system

An anti-spoofing authentication system includes an image capture device that converts received light into an image signal representing a captured image; a microphone that converts sound wave into a voice signal; a face detection device that detects a human face on the captured image, the face detection device including a lip detection device that detects lip features associated with the detected human face; a voice activity detection (VAD) device that detects presence of human voice in the voice signal; and an anti-spoofing device that detects a spoofing attack according to the detected lip features and the detected human voice.
Owner:HIMAX TECH LTD

Gaussian sputtering conversation face generation method based on lip features and head posture guidance

The invention provides a Gaussian sputter conversation face generation method based on lip features and head posture guidance, which can effectively generate matched lip motion and stable posture action for different input audios under the condition of giving a speech video of a target object. The method comprises the following steps: firstly, constructing a three-dimensional pre-training data set through a face reconstruction technology, and learning a mapping relation between audio and lip movement by adopting a grid generation module; in the lip feature generation process, the modules are finely adjusted to generate lip features adaptive to the talking style of the target object; in the process of extracting the head posture parameters, optimizing the head posture parameters by adopting a feature point matching algorithm and a beam adjustment method, and smoothing the parameters by adopting a filter; and finally, introducing a dynamic three-dimensional Gaussian renderer, and synthesizing a vivid mouth shape synchronization video by taking the output of the two processes as a guidance condition. Compared with a previous method, the method has higher precision under the driving of intra-domain audio and cross-domain audio.
Owner:SICHUAN UNIV

Speaking face generation method and device, equipment and medium

The invention relates to an artificial intelligence technology, can be applied to business system platforms of medical health, financial science and technology and the like, and discloses a speaking face generation method, device, equipment and medium, and the method comprises the steps: obtaining a reference face image and an original video frame sequence, carrying out the coding, and generating a potential representation, performing lip mask processing on each frame of image in the original video frame sequence to obtain a local mouth image, and encoding the local mouth image to generate lip features; obtaining random sampling noise, and generating joint potential features based on the potential characterization, the lip features and the random sampling noise; on the basis of the reference face image features, combining the potential features and the input audio features, generating lip alignment features; and after the audio and lip alignment features are decoded, lip synchronization frames are generated and assembled, and a speaking face video is generated. According to the method, high-resolution visual quality can be kept, meanwhile, deep alignment of audio content, facial expressions and time sequence dynamic states is achieved, and the speaking face generation quality and generalization ability are improved.
Owner:平安科技(上海)有限公司

Face animation generation method and device and storage medium

The invention discloses a face animation generation method and device and a storage medium, and relates to the technical field of digital humans, and the face animation generation method comprises the steps: coding an input audio based on a whisper model and a wav2cev model, and obtaining an audio feature vector; public features, face features, head posture features and lip features are extracted based on the input face image, and the public features comprise face grid features, tooth features and eye features; the first combined feature and the second combined feature are fused to obtain a fused feature, the first combined feature comprises a common feature, a face feature and a head posture feature, and the second combined feature comprises a common feature and a lip feature; and taking the fusion feature and the audio feature vector as condition information, and injecting the condition information into a diffusion model through a cross attention mechanism, so that the diffusion model generates a face animation, and accurate synchronization of the audio information and the face action in the face animation is realized.
Owner:EMDOOR INFORMATION CO LTD

A voice interaction recognition enhancement method, device and storage medium

ActiveCN119207393BSpeech recognitionModal voiceLip feature
The application provides a speech interaction recognition enhancement method, which comprises the following steps: collecting a video of a speaker facing a camera for speech interaction, splitting the video into N segments of to-be-identified speech and N frames of to-be-identified images, and forming to-be-identified data; inputting the to-be-identified data into a preset speech interaction recognition enhancement model, wherein the model comprises a lip feature extraction network, a speech feature extraction network, a time feature extraction network and an activation network, the time feature extraction network is used for adding time sequence information to a lip feature matrix extracted by the lip feature extraction network and a speech feature matrix extracted by the speech feature extraction network, the activation network is used for simulating an activation-inhibition mechanism of visual information on an auditory nerve circuit in a biological brain, realizing the interaction of two modalities of vision and hearing, and obtaining a speech recognition result. The application realizes the accurate recognition of lip language and speech in a noisy environment by using the dual-modality speech recognition of vision and hearing, and the response capability and accuracy meet the requirements.
Owner:TSINGHUA UNIVERSITY

Video keyword retrieval method and device based on multiple encoders and electronic equipment

The invention provides a video keyword retrieval method and device based on multiple encoders and electronic equipment. The method comprises the following steps: acquiring a lip region image sequence of a to-be-detected video; the lip region image sequence is input into a visual feature extraction network, the visual feature extraction network comprises multiple visual encoders arranged in parallel, different types of visual features are extracted and fused, and multi-encoder visual features are obtained; inputting the to-be-detected keyword into a text encoder, and extracting to obtain a text feature; and performing alignment processing on the text feature and the multi-encoder visual feature to generate a joint feature, inputting the joint feature into a classifier, and outputting the probability of occurrence of the to-be-detected keyword in the video. A multi-encoder structure is introduced to carry out modeling on video lip features, cross-modal alignment is realized in combination with text features, and the problem that keyword detection cannot be carried out in a noisy environment or a silent scene in the prior art is effectively solved.
Owner:BEIJING YUANJIAN INFORMATION TECH CO LTD

3D digital human multi-mode real-time driving method and system based on low-computing-power equipment

The invention provides a 3D digital human multi-mode real-time driving method and system based on low-computing-power equipment, and relates to the technical field of digital humans, and the method comprises the steps: collecting original emotion data sets of a 3D digital human under various expressions through 3D scanning equipment; preprocessing the original emotion data set to obtain multi-modal emotion data, and performing feature extraction on the multi-modal emotion data to obtain voice features, lip features and micro-expression features; constructing a neural radiation field function based on the multi-modal emotion data and the 3D model space parameters, and performing cross-modal feature alignment on the neural radiation field function and the knowledge migration model to construct a lightweight recognition model; and transmitting real-time emotion data collected by the 3D scanning device to the lightweight recognition model, obtaining an expression driving signal of the current 3D digital person, and driving the 3D digital person to execute a corresponding response behavior according to the expression driving signal. According to the invention, the richness and naturalness of the emotional expression of the 3D digital human can be improved.
Owner:WUHAN DINGSEN ELECTRONIC TECH CO LTD

Method and device for detecting other people to answer on behalf, computer equipment and storage medium

The invention discloses a detection method and device for other people to answer on behalf, computer equipment and a storage medium, and relates to the technical field of face-to-face detection in the fields of finance, insurance, medical treatment, banking and the like, and the method comprises the steps: obtaining a to-be-detected video stream and a corresponding audio stream; extracting a face region image from each frame of video image of the video stream; face key point detection is carried out on the face region image, lip key points are positioned, a lip 2D coordinate sequence is obtained, and the lip 2D coordinate sequence is converted into a lip feature map; extracting a face feature map of the face region image, and generating a visual feature vector based on the lip feature map and the face feature map; extracting an audio feature vector of the audio stream, and performing multi-modal alignment on the visual feature vector and the audio feature vector to obtain a fusion feature vector; and detecting the fusion feature vector based on a pre-trained classifier to obtain a detection result of answering by others on behalf. According to the invention, the accuracy of sound and lip synchronous detection can be greatly improved, so that other people's pickup behaviors can be effectively identified.
Owner:PING AN TECH (SHENZHEN) CO LTD

Virtual anchor real-time driving system based on facial motion capture

ActiveCN121842342Aexact matchPrecise linkageTelevision system detailsColor television detailsPhoneme recognitionEngineering
The invention discloses a virtual anchor real-time driving system based on facial motion capture, and aims at solving the problems that in the prior art, virtual anchor facial motions are insufficient in accuracy, lack of natural sense and weak in scene adaptation. Audio and video streams are collected in parallel through multiple threads, and the data quality is optimized through self-adaptive preprocessing; the technologies of multi-dimensional facial feature collaborative extraction, double-branch phoneme recognition, face functional region semantic segmentation and confidence coefficient dynamic weight adjustment are adopted, and accurate representation and collaborative fusion of expression and mouth shape features are achieved; and finally, outputting real-time driving characteristics matched with the virtual image through a mouth shape-expression cooperative driving mechanism. Therefore, the accuracy and the natural sense of the face action of the virtual anchor are effectively improved, the adaptability of complex scenes is enhanced, real-time response and low-cost deployment are both considered, and the method is suitable for scenes such as virtual live broadcast and online education and has good application value.
Owner:GUIZHOU NORMAL UNIVERSITY

A large model-based speech recognition and speech synthesis optimization method and system

ActiveCN121034309BSpeech recognitionSpeech synthesisData setLip feature
The application provides a speech recognition and speech synthesis optimization method and system based on a large model, which extracts lip movement features, lip feature time stamps and audio feature time stamps by acquiring real-time speech input signals and image frame sequences of a user's face area to generate an initial space-time offset sequence; generates a dynamic offset compensation parameter sequence based on the initial space-time offset sequence and in combination with a reference alignment template in a speech visual synchronization data set; jointly processes the dynamic offset compensation parameter sequence, the lip movement features and the real-time speech input signals by using a large model to reconstruct a target speech segment, generates a corrected phoneme sequence in combination with the lip movement features, retrieves a mouth shape parameter group corresponding to the corrected phoneme sequence from a phoneme mouth shape mapping rule library to generate a speech waveform synchronized with the lip movement phase; and the application improves the naturalness, immersion and robustness in a speech missing or delayed scenario of human-computer interaction.
Owner:LUSTER LIGHTWAVE CO LTD

Lip speech recognition method and system based on subspace sparse attention mechanism and medium

The present application relates to a kind of lip reading method based on subspace sparse attention mechanism, system and medium, method includes: obtaining lip region image sequence, based on the lip region image sequence extraction obtains lip feature sequence;The lip feature sequence is input to the preset training complete phoneme sequence extraction model, obtains the pronunciation phoneme sequence corresponding to the lip feature sequence;The pronunciation phoneme sequence is input to the sentence inference model with the subspace sparse self-attention mechanism built in and obtains target sentence sequence.The present application enhances the context information by constructing a special attention mechanism, realizes the prediction long sentence sequence in a forward operation, so that the inference rate and accuracy are greatly improved.
Owner:WUHAN UNIV OF TECH CHONGQING RES INST

Lip-reading-based voice interaction methods, devices, equipment, and storage media

ActiveCN120600019BLip featureHuman–computer interaction
This invention discloses a lip-reading-enhanced voice interaction method, apparatus, device, and storage medium. The lip-reading-enhanced voice interaction method includes: extracting lip-reading features from image sequences of the lip region; extracting audio features from the speech signal; performing cross-modal fusion encoding of the lip-reading features and audio features to generate hybrid features containing audiovisual information; inputting the hybrid features into a large language model to understand the intent of the interactive object and generate corresponding semantic responses; and finally synthesizing the speech and / or converting it into text. This invention, by introducing lip features, provides additional visual cues for speech recognition, significantly improving the robustness and accuracy of speech recognition; effectively fusing and encoding lip-reading features and audio features avoids the semantic information fragmentation caused by simple independent recognition; and fully utilizes the capabilities of a large model to achieve a more natural and intelligent interactive experience.
Owner:SHENZHEN WANRUI INTELLIGENT TECH CO LTD

Multimodal duration-controlled TTS (DC-TTS) system for automatic dubbing

PCT designated stageWO2025186682A1Natural language translationSpeech recognitionData setLip feature
A system and method for multimodal duration-controlled text to speech (TTS) is provided. The system receives a dataset including reference video clip, reference audio, and text. The system generates: text features based on application of a text encoder on the text, visual features based on application of a visual feature extractor on the reference video clip, a first speaker embedding associated with the human speaker based on the reference audio, and lip features associated with the human speaker based on the reference video clip. The system further generates audio tokens based on application of a neural language model on the text features, the visual features, the first speaker embedding, and the lip features. The system generates an audio based on application of a neural vocoder on the audio tokens and trains the neural language model based on the generated audio and a loss function.
Owner:SONY GROUP CORP

Lip reading method and device, electronic equipment, storage medium and program product

PendingCN122337199ANerve networkNetwork output
This invention relates to the field of lip-reading technology, providing lip-reading recognition methods, devices, electronic devices, storage media, and program products. The lip-reading recognition method includes: acquiring audio data and facial data including lip images; obtaining a lip feature map based on the facial data and an audio feature sequence based on the audio data; inputting the lip feature map and the audio feature sequence into a multi-scale interactive modal fusion network to obtain multi-modal lip-reading features output by the multi-scale interactive modal fusion network; the multi-scale interactive modal fusion network integrates audio and visual modal information; inputting the multi-modal lip-reading features into a lip-reading recognition neural network to obtain the recognition result output by the lip-reading recognition neural network; the lip-reading recognition neural network performs result recognition on the input data through a self-attention mechanism. This invention designs a multi-scale interactive modal fusion network that integrates audio and visual features, effectively improving the accuracy and stability of lip-reading recognition.
Owner:CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1