Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

18 results about "Lip feature" patented technology

Voice operation method of device, apparatus, and electronic device

A voice operation method of a device, comprising: acquiring a video collected by a camera; acquiring voice information collected by a microphone; detecting a face image in the video; extracting a lip feature and a face feature of the face image; determining a time interval according to the lip feature; intercepting a corresponding audio segment in the voice information according to the time interval; acquiring voiceprint information according to the face feature; and performing voice recognition on the audio segment according to the voiceprint information to acquire voice information. The voice operation method of the device does not need to pre-determine a target user and pre-record voiceprint information of the target user, can autonomously extract voiceprint information of multiple users when multiple users use the device at the same time, and can separate voices one by one. Meanwhile, the voiceprint can be autonomously updated and registered. The voice operation method of the device can significantly improve the voice recognition effect of the device in a noisy or multi-user speaking scene.
Owner:HUAWEI TECH CO LTD

Sound and lip synchronization detection method and device, electronic equipment and storage medium

The invention provides a sound and lip synchronous detection method and device, electronic equipment and a storage medium, relates to the technical field of artificial intelligence, and is suitable for the financial field and the medical field. The method comprises the following steps: encoding a target voice to obtain an initial voice feature, and encoding a target object lip in a target face video to obtain an initial lip movement feature; intercepting audio-visual feature extraction sub-models from N feature converters cascaded in the audio-visual association model; performing voice feature extraction on the initial voice feature and the video zero matrix through an audio-visual feature extraction sub-model to obtain a target voice feature; performing lip feature extraction on the initial lip movement feature and the audio zero matrix through an audio-visual feature extraction sub-model to obtain a target lip movement feature; and carrying out voice and lip synchronization classification on the target voice feature and the target lip movement feature to obtain a voice and lip synchronization category. According to the method, the feature learning difficulty can be remarkably reduced, the audio-visual association modeling precision is improved, and thus the accuracy of audio-lip synchronous detection is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Semantic recognition method and device fusing audio and video, equipment and medium

The invention provides an audio and video fused semantic recognition method and device, and the method comprises the steps: carrying out the feature extraction of audio information, and obtaining an audio feature; performing feature extraction on the video information to obtain lip features; aligning and splicing the audio features and the lip features to obtain joint features; performing signal-to-noise ratio estimation on the spectrum feature map to obtain a signal-to-noise ratio, and obtaining an audio weight and a video weight according to the signal-to-noise ratio and a preset weight distribution function; performing feature splitting on the joint feature based on the audio weight and the video weight to obtain an audio branch feature and a video branch feature, and performing feature fusion on the audio branch feature, the video branch feature and the joint feature to obtain a fusion feature; and performing semantic recognition on the fused features to obtain a recognition result. By means of the semantic recognition method and device fusing the audio and the video, the technical problem that the voice recognition effect is poor in a complex acoustic environment is solved.
Owner:HEFEI UNIV OF TECH

A voice translation screen control method and system for interaction

The application relates to a voice translation screen control method and system for interaction, and relates to the field of intelligent interaction technology, which comprises the following steps: acquiring a region detection image; performing feature recognition on the region detection image to determine user lip features, and determining user lip positions according to the user lip features; determining effective sound pickup distances according to the user lip positions and positions of sound pickup devices; determining sound pickup sensitive ranges corresponding to the effective sound pickup distances according to sound pickup matching relationships; randomly selecting a sound pickup sensitive value in each sound pickup sensitive range to define an effective sensitive value, and controlling the sound pickup devices to work at the effective sensitive value to acquire external voice volumes; determining a user representative volume according to the external voice volumes, determining a sensitive adjustment coefficient according to the user representative volume and an effective recognition volume, and adjusting and updating the effective sensitive values according to the sensitive adjustment coefficient. The application has the effect of facilitating subsequent analysis of collected voices.
Owner:NINGBO LANKE INTELLIGENT ENG

Virtual anchor real-time driving system based on facial motion capture

ActiveCN121842342BResolve driver conflictsImprove viewing experienceTelevision system detailsColor television detailsPhoneme recognitionEngineering
The application discloses a virtual anchor real-time driving system based on facial motion capture, aims to solve the problems of insufficient precision, lack of naturalness and weak scene adaptation of virtual anchor facial motion in the prior art, and through multi-thread parallel collection of audio and video streams, optimizes data quality through adaptive preprocessing; adopts multi-dimensional facial feature collaborative extraction, double-branch phoneme recognition, face function area semantic segmentation and confidence dynamic weight adjustment technology, realizes accurate representation and collaborative fusion of expression and lip feature; finally, through the lip-expression collaborative driving mechanism, the real-time driving characteristics of the adaptive virtual image are output; thereby effectively improving the precision and naturalness of virtual anchor facial motion, enhancing the complex scene adaptability, giving consideration to real-time response and low-cost deployment, and being applicable to virtual live broadcast, online education and other scenes, and having good application value.
Owner:GUIZHOU NORMAL UNIVERSITY

Gaussian sputtering conversation face generation method based on lip features and head posture guidance

The invention provides a Gaussian sputter conversation face generation method based on lip features and head posture guidance, which can effectively generate matched lip motion and stable posture action for different input audios under the condition of giving a speech video of a target object. The method comprises the following steps: firstly, constructing a three-dimensional pre-training data set through a face reconstruction technology, and learning a mapping relation between audio and lip movement by adopting a grid generation module; in the lip feature generation process, the modules are finely adjusted to generate lip features adaptive to the talking style of the target object; in the process of extracting the head posture parameters, optimizing the head posture parameters by adopting a feature point matching algorithm and a beam adjustment method, and smoothing the parameters by adopting a filter; and finally, introducing a dynamic three-dimensional Gaussian renderer, and synthesizing a vivid mouth shape synchronization video by taking the output of the two processes as a guidance condition. Compared with a previous method, the method has higher precision under the driving of intra-domain audio and cross-domain audio.
Owner:SICHUAN UNIV

Video keyword retrieval method and device based on multiple encoders and electronic equipment

The invention provides a video keyword retrieval method and device based on multiple encoders and electronic equipment. The method comprises the following steps: acquiring a lip region image sequence of a to-be-detected video; the lip region image sequence is input into a visual feature extraction network, the visual feature extraction network comprises multiple visual encoders arranged in parallel, different types of visual features are extracted and fused, and multi-encoder visual features are obtained; inputting the to-be-detected keyword into a text encoder, and extracting to obtain a text feature; and performing alignment processing on the text feature and the multi-encoder visual feature to generate a joint feature, inputting the joint feature into a classifier, and outputting the probability of occurrence of the to-be-detected keyword in the video. A multi-encoder structure is introduced to carry out modeling on video lip features, cross-modal alignment is realized in combination with text features, and the problem that keyword detection cannot be carried out in a noisy environment or a silent scene in the prior art is effectively solved.
Owner:BEIJING YUANJIAN INFORMATION TECH CO LTD

Method and device for detecting other people to answer on behalf, computer equipment and storage medium

The invention discloses a detection method and device for other people to answer on behalf, computer equipment and a storage medium, and relates to the technical field of face-to-face detection in the fields of finance, insurance, medical treatment, banking and the like, and the method comprises the steps: obtaining a to-be-detected video stream and a corresponding audio stream; extracting a face region image from each frame of video image of the video stream; face key point detection is carried out on the face region image, lip key points are positioned, a lip 2D coordinate sequence is obtained, and the lip 2D coordinate sequence is converted into a lip feature map; extracting a face feature map of the face region image, and generating a visual feature vector based on the lip feature map and the face feature map; extracting an audio feature vector of the audio stream, and performing multi-modal alignment on the visual feature vector and the audio feature vector to obtain a fusion feature vector; and detecting the fusion feature vector based on a pre-trained classifier to obtain a detection result of answering by others on behalf. According to the invention, the accuracy of sound and lip synchronous detection can be greatly improved, so that other people's pickup behaviors can be effectively identified.
Owner:PING AN TECH (SHENZHEN) CO LTD

Virtual anchor real-time driving system based on facial motion capture

ActiveCN121842342Aexact matchPrecise linkageTelevision system detailsColor television detailsPhoneme recognitionEngineering
The invention discloses a virtual anchor real-time driving system based on facial motion capture, and aims at solving the problems that in the prior art, virtual anchor facial motions are insufficient in accuracy, lack of natural sense and weak in scene adaptation. Audio and video streams are collected in parallel through multiple threads, and the data quality is optimized through self-adaptive preprocessing; the technologies of multi-dimensional facial feature collaborative extraction, double-branch phoneme recognition, face functional region semantic segmentation and confidence coefficient dynamic weight adjustment are adopted, and accurate representation and collaborative fusion of expression and mouth shape features are achieved; and finally, outputting real-time driving characteristics matched with the virtual image through a mouth shape-expression cooperative driving mechanism. Therefore, the accuracy and the natural sense of the face action of the virtual anchor are effectively improved, the adaptability of complex scenes is enhanced, real-time response and low-cost deployment are both considered, and the method is suitable for scenes such as virtual live broadcast and online education and has good application value.
Owner:GUIZHOU NORMAL UNIVERSITY

A large model-based speech recognition and speech synthesis optimization method and system

ActiveCN121034309BSpeech recognitionSpeech synthesisData setLip feature
The application provides a speech recognition and speech synthesis optimization method and system based on a large model, which extracts lip movement features, lip feature time stamps and audio feature time stamps by acquiring real-time speech input signals and image frame sequences of a user's face area to generate an initial space-time offset sequence; generates a dynamic offset compensation parameter sequence based on the initial space-time offset sequence and in combination with a reference alignment template in a speech visual synchronization data set; jointly processes the dynamic offset compensation parameter sequence, the lip movement features and the real-time speech input signals by using a large model to reconstruct a target speech segment, generates a corrected phoneme sequence in combination with the lip movement features, retrieves a mouth shape parameter group corresponding to the corrected phoneme sequence from a phoneme mouth shape mapping rule library to generate a speech waveform synchronized with the lip movement phase; and the application improves the naturalness, immersion and robustness in a speech missing or delayed scenario of human-computer interaction.
Owner:LUSTER LIGHTWAVE CO LTD

Lip speech recognition method and system based on subspace sparse attention mechanism and medium

The present application relates to a kind of lip reading method based on subspace sparse attention mechanism, system and medium, method includes: obtaining lip region image sequence, based on the lip region image sequence extraction obtains lip feature sequence;The lip feature sequence is input to the preset training complete phoneme sequence extraction model, obtains the pronunciation phoneme sequence corresponding to the lip feature sequence;The pronunciation phoneme sequence is input to the sentence inference model with the subspace sparse self-attention mechanism built in and obtains target sentence sequence.The present application enhances the context information by constructing a special attention mechanism, realizes the prediction long sentence sequence in a forward operation, so that the inference rate and accuracy are greatly improved.
Owner:WUHAN UNIV OF TECH CHONGQING RES INST

Lip-reading-based voice interaction methods, devices, equipment, and storage media

ActiveCN120600019BLip featureHuman–computer interaction
This invention discloses a lip-reading-enhanced voice interaction method, apparatus, device, and storage medium. The lip-reading-enhanced voice interaction method includes: extracting lip-reading features from image sequences of the lip region; extracting audio features from the speech signal; performing cross-modal fusion encoding of the lip-reading features and audio features to generate hybrid features containing audiovisual information; inputting the hybrid features into a large language model to understand the intent of the interactive object and generate corresponding semantic responses; and finally synthesizing the speech and / or converting it into text. This invention, by introducing lip features, provides additional visual cues for speech recognition, significantly improving the robustness and accuracy of speech recognition; effectively fusing and encoding lip-reading features and audio features avoids the semantic information fragmentation caused by simple independent recognition; and fully utilizes the capabilities of a large model to achieve a more natural and intelligent interactive experience.
Owner:SHENZHEN WANRUI INTELLIGENT TECH CO LTD

Lip reading method and device, electronic equipment, storage medium and program product

PendingCN122337199ANerve networkNetwork output
This invention relates to the field of lip-reading technology, providing lip-reading recognition methods, devices, electronic devices, storage media, and program products. The lip-reading recognition method includes: acquiring audio data and facial data including lip images; obtaining a lip feature map based on the facial data and an audio feature sequence based on the audio data; inputting the lip feature map and the audio feature sequence into a multi-scale interactive modal fusion network to obtain multi-modal lip-reading features output by the multi-scale interactive modal fusion network; the multi-scale interactive modal fusion network integrates audio and visual modal information; inputting the multi-modal lip-reading features into a lip-reading recognition neural network to obtain the recognition result output by the lip-reading recognition neural network; the lip-reading recognition neural network performs result recognition on the input data through a self-attention mechanism. This invention designs a multi-scale interactive modal fusion network that integrates audio and visual features, effectively improving the accuracy and stability of lip-reading recognition.
Owner:CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1

Method for detecting pickup based on synthetic lips and related equipment

The embodiment of the invention belongs to the technical field of artificial intelligence, and relates to a pickup detection method and related equipment based on synthetic lips, and the method comprises the steps: carrying out the lip sequence extraction processing of original video data according to a CNN face detector, and obtaining a lip feature sequence; performing synthetic lip sequence generation processing on the original audio data according to a pre-trained lip generation model to obtain a synthetic lip feature sequence; calculating the difference between the lip feature sequence and the synthesized lip feature sequence in the low-dimensional embedding space to obtain a difference feature sequence; performing action segmentation processing on the difference feature sequence according to a multi-scale time convolution network to obtain an action probability distribution sequence with the same length as the input time step; and performing pickup recognition processing on the action probability distribution sequence according to the classifier to obtain a pickup detection result. The method can be used for carrying out related video processing in a financial science and technology business system, and can effectively detect pickup behaviors in scenes such as online examinations, remote interviews, video conferences and the like.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent voice recognition and interaction system and method fusing AI visual information

The invention discloses an intelligent voice recognition and interaction system and method fusing AI visual information, and relates to the technical field of voice recognition. The method comprises the following steps: synchronously acquiring audio and RGB-D video streams, and extracting acoustic features, visual lip language features and three-dimensional space interaction features in parallel by using a double-flow convolutional neural network; the method comprises the following steps: constructing a cross-modal dynamic gating fusion network, combining a real-time signal-to-noise ratio and visual confidence, dynamically distributing a sound visual weight, and realizing high-robustness recognition based on lip language in a high-noise environment; aiming at anaphora ambiguity in a natural language, realizing intention analysis of what you see is what you control by utilizing projection and collision detection of sight lines and gesture rays in a three-dimensional semantic map; according to the invention, the problems of low recognition rate, unknown reference and tedious wake-up in a complex sound field environment are effectively solved.
Owner:FUJIAN CLOUD INTELLIGENT TECH CO LTD

A low-resource language lip reading method and device based on pre-training fine-tuning

The present application relates to the technical field of computer vision, and particularly relates to a low-resource language lip reading method and device based on pre-training fine-tuning. The method comprises: pre-training a model by using a large English video dataset to ensure that the model has strong generalization ability and effective lip feature expression ability; then, after loading the pre-trained model weight, fine-tuning the model by using a small amount of Tibetan lip reading dataset to overcome the challenge of Tibetan video data scarcity. In the decoding stage, a Transformer language model specially trained for Tibetan text is introduced, which effectively reduces the homonym confusion problem that may occur in the lip reading process, thereby improving the accuracy of sentence-level Tibetan lip reading. The overall architecture is improved by the above innovative structure and method, and effective pure visual lip reading of low-resource languages is successfully realized.
Owner:MINZU UNIVERSITY OF CHINA

A Deep Learning-Based Lip Correction System

This invention discloses a deep learning-based lip correction system, comprising an image acquisition module, an image preprocessing module, a lip key region segmentation model design module, a lip feature detection model design module, and a lip correction module. This invention belongs to the field of image processing, specifically referring to a deep learning-based lip correction system. This scheme introduces a dynamic grayscale weighted mask to significantly amplify the gradient in low grayscale regions; it suppresses weights in high grayscale regions through dynamic weights, adds an edge-aware term to the loss function to enhance sensitivity to lip line edges, and improves the accuracy of lip key region segmentation; thereby improving the reliability of subsequent lip correction; it combines rotation region representation with endpoint fitting to accurately characterize the key lip region; and it introduces endpoint offset penalty loss and angle fitting loss to precisely penalize corner point errors, eliminate fitting jumps, improve the positioning accuracy of key lip control points, and provide precise anchor points; thus improving the final lip correction effect.
Owner:JINYU ZHIDA TECHNOLOGY (TIANJIN) CO LTD

Lip sound synchronous detection and calibration method and device, computer equipment and storage medium

The invention belongs to the interdisciplinary field of computer vision, audio signal processing and artificial intelligence, and relates to a lip sound synchronous detection and calibration method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining to-be-detected audio and video data, and carrying out the separation and preprocessing of the audio and video data; performing lip region detection and extraction on the video frame sequence to obtain a lip image sequence; performing multi-view lip shape reconstruction and view normalization on the lip image sequence to generate a front lip map sequence; constructing an audio and video spatio-temporal feature tensor, and generating an audio tensor and a video tensor; performing bimodal heterogeneous three-dimensional convolutional coupling network feature mapping, and calculating a measurement distance between the audio feature vector and the video feature vector; and judging whether the audio and video data has the voice lip movement consistency or not according to the relationship between the measurement distance and a preset judgment threshold. The problems of lip feature deformation and difficult extraction caused by non-front view angles such as side face and head lowering in a digital human teaching video are systematically solved.
Owner:GUANGDONG OPEN UNIV (GUANGDONG POLYTECHNIC VOCATIONAL COLLEGE)