Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

32 results about "Lip feature" patented technology

Voice interaction method and device based on lip language enhancement, equipment and storage medium

The invention discloses a voice interaction method and device based on lip language enhancement, equipment and a storage medium, and the method comprises the steps: extracting lip language features based on an image sequence of a lip region, and carrying out the feature extraction of a voice signal, and obtaining an audio feature; performing cross-modal fusion coding on the lip language features and the audio features to generate mixed features containing audio-visual information; inputting the mixed features into a large language model, understanding the intention of the interaction object and generating a corresponding semantic reply; and finally, synthesizing into voice and / or converting into characters. According to the invention, by introducing the lip features, additional visual clues are provided for speech recognition, and the robustness and accuracy of speech recognition can be significantly improved; effective fusion coding is carried out on the lip language features and the sound features, and semantic information splitting caused by simple and independent recognition is avoided; and the capability of the large model is fully utilized, so that more natural and more intelligent interaction experience is realized.
Owner:SHENZHEN WANRUI INTELLIGENT TECH CO LTD

Low-resource language lip language recognition method and device based on pre-training fine tuning

The invention relates to the technical field of computer vision, in particular to a low-resource language lip language recognition method and device based on pre-training fine tuning. The method comprises the following steps: pre-training a model by using a large amount of English video data sets to ensure that the model obtains strong generalization ability and effective lip feature expression ability; and then, after the weight of the pre-training model is loaded, full-parameter fine tuning is performed on the model through a small amount of Tibetan lip language data sets, so that the challenge of scarcity of Tibetan video data is overcome. In the inference decoding stage, a Transform language model specially aiming at Tibetan text training is introduced, the problem of homonym confusion possibly occurring in the lip language recognition process is effectively reduced, and therefore the accuracy of sentence-level Tibetan lip language recognition is improved. The overall architecture is improved through the innovative structure and method, and effective pure vision lip language recognition of low-resource languages is successfully achieved.
Owner:MINZU UNIVERSITY OF CHINA

Speech recognition and speech synthesis optimization method and system based on large model

The invention provides a voice recognition and voice synthesis optimization method and system based on a large model, and the method comprises the steps: extracting lip motion features, lip feature timestamps and audio feature timestamps through obtaining a real-time voice input signal and an image frame sequence of a user face region, so as to generate an initial space-time offset sequence; generating a dynamic offset compensation parameter sequence based on the initial space-time offset sequence in combination with a reference alignment template in the voice and vision synchronization data set; and performing joint processing on the dynamic offset compensation parameter sequence, the lip motion characteristics and the real-time voice input signal by using a large model to reconstruct a target voice segment, generating a corrected phoneme sequence in combination with the lip motion characteristics, retrieving a mouth shape parameter group corresponding to the corrected phoneme sequence from a phoneme mouth shape mapping rule base, and performing mouth shape correction on the mouth shape parameter group. To generate a voice waveform in phase synchronization with the lip motion; according to the invention, the naturalness and immersion of man-machine interaction and the robustness in a voice missing or delay scene are improved.
Owner:LUSTER LIGHTWAVE CO LTD

Voice operation method of device, apparatus, and electronic device

A voice operation method of a device, comprising: acquiring a video collected by a camera; acquiring voice information collected by a microphone; detecting a face image in the video; extracting a lip feature and a face feature of the face image; determining a time interval according to the lip feature; intercepting a corresponding audio segment in the voice information according to the time interval; acquiring voiceprint information according to the face feature; and performing voice recognition on the audio segment according to the voiceprint information to acquire voice information. The voice operation method of the device does not need to pre-determine a target user and pre-record voiceprint information of the target user, can autonomously extract voiceprint information of multiple users when multiple users use the device at the same time, and can separate voices one by one. Meanwhile, the voiceprint can be autonomously updated and registered. The voice operation method of the device can significantly improve the voice recognition effect of the device in a noisy or multi-user speaking scene.
Owner:HUAWEI TECH CO LTD

Sound and lip synchronization detection method and device, electronic equipment and storage medium

The invention provides a sound and lip synchronous detection method and device, electronic equipment and a storage medium, relates to the technical field of artificial intelligence, and is suitable for the financial field and the medical field. The method comprises the following steps: encoding a target voice to obtain an initial voice feature, and encoding a target object lip in a target face video to obtain an initial lip movement feature; intercepting audio-visual feature extraction sub-models from N feature converters cascaded in the audio-visual association model; performing voice feature extraction on the initial voice feature and the video zero matrix through an audio-visual feature extraction sub-model to obtain a target voice feature; performing lip feature extraction on the initial lip movement feature and the audio zero matrix through an audio-visual feature extraction sub-model to obtain a target lip movement feature; and carrying out voice and lip synchronization classification on the target voice feature and the target lip movement feature to obtain a voice and lip synchronization category. According to the method, the feature learning difficulty can be remarkably reduced, the audio-visual association modeling precision is improved, and thus the accuracy of audio-lip synchronous detection is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Video generation method, model training method, device and computer program product

The invention discloses a video generation method and device, a model training method and device and a computer program product, and the video generation method comprises the steps: obtaining a target audio and a reference picture used for generating a video, and the reference picture comprises a sound production object; determining a global visual feature of each to-be-generated video frame corresponding to the audio clip according to the clip features of one or more audio clips corresponding to the target audio and the reference image; according to the pronunciation feature of each audio frame of the target audio and the lip feature of the sounding object in the reference picture, determining the lip feature of the sounding object in a to-be-generated video frame corresponding to the audio frame; and generating each video frame according to the lip feature and the global visual feature corresponding to the to-be-generated video frame. Through the scheme provided by the invention, the expression of characters in the generated video is more vivid and natural, the lip action and the audio can be accurately synchronized, and the visual experience of a user is improved.
Owner:BEIJING AUTONAVI YUNMAP TECH CO LTD

Semantic recognition method and device fusing audio and video, equipment and medium

The invention provides an audio and video fused semantic recognition method and device, and the method comprises the steps: carrying out the feature extraction of audio information, and obtaining an audio feature; performing feature extraction on the video information to obtain lip features; aligning and splicing the audio features and the lip features to obtain joint features; performing signal-to-noise ratio estimation on the spectrum feature map to obtain a signal-to-noise ratio, and obtaining an audio weight and a video weight according to the signal-to-noise ratio and a preset weight distribution function; performing feature splitting on the joint feature based on the audio weight and the video weight to obtain an audio branch feature and a video branch feature, and performing feature fusion on the audio branch feature, the video branch feature and the joint feature to obtain a fusion feature; and performing semantic recognition on the fused features to obtain a recognition result. By means of the semantic recognition method and device fusing the audio and the video, the technical problem that the voice recognition effect is poor in a complex acoustic environment is solved.
Owner:HEFEI UNIV OF TECH

Method for training multi-modal speaker verification model and electronic device

ActiveCN120148521BSpeech recognitionSpeaker verificationFeature coding
The application discloses a training method of a multi-modal speaker authentication model and an electronic device, relates to the field of speaker authentication, and comprises the following steps: obtaining an audio-video sequence sample corresponding to a given speaker; performing lip feature coding on each audio frame in the audio-video sequence sample in sequence to determine a corresponding audio sample lip latent vector sequence; performing feature coding on a speaker mouth region in each video frame in the audio-video sequence sample in sequence to determine a corresponding video sample lip latent vector sequence; calculating an audio-video sequence similarity corresponding to the audio sample lip latent vector sequence and the video sample lip latent vector sequence; and training a speaker authentication model according to the audio-video sequence similarity and an audio-video speaker alignment label. Thus, the similarity of the speaker audio and the lip shape graph of the person to be detected is compared by matching the speaker audio and the lip shape graph into the same dimension vector through the latent vector, so that the accuracy of the authentication result can be effectively improved.
Owner:AISPEECH CO LTD

Multimodal duration-controlled TTS (dc-TTS) system for automatic dubbing

PCT designated stageWO2025186682A8Natural language translationSpeech recognitionData setLip feature
A system and method for multimodal duration-controlled text to speech (TTS) is provided. The system receives a dataset including reference video clip, reference audio, and text. The system generates: text features based on application of a text encoder on the text, visual features based on application of a visual feature extractor on the reference video clip, a first speaker embedding associated with the human speaker based on the reference audio, and lip features associated with the human speaker based on the reference video clip. The system further generates audio tokens based on application of a neural language model on the text features, the visual features, the first speaker embedding, and the lip features. The system generates an audio based on application of a neural vocoder on the audio tokens and trains the neural language model based on the generated audio and a loss function.
Owner:SONY GROUP CORP

A voice translation screen control method and system for interaction

The application relates to a voice translation screen control method and system for interaction, and relates to the field of intelligent interaction technology, which comprises the following steps: acquiring a region detection image; performing feature recognition on the region detection image to determine user lip features, and determining user lip positions according to the user lip features; determining effective sound pickup distances according to the user lip positions and positions of sound pickup devices; determining sound pickup sensitive ranges corresponding to the effective sound pickup distances according to sound pickup matching relationships; randomly selecting a sound pickup sensitive value in each sound pickup sensitive range to define an effective sensitive value, and controlling the sound pickup devices to work at the effective sensitive value to acquire external voice volumes; determining a user representative volume according to the external voice volumes, determining a sensitive adjustment coefficient according to the user representative volume and an effective recognition volume, and adjusting and updating the effective sensitive values according to the sensitive adjustment coefficient. The application has the effect of facilitating subsequent analysis of collected voices.
Owner:NINGBO LANKE INTELLIGENT ENG

A voice-driven talking portrait generation method and an electronic device

The application discloses a speech-driven talking portrait generation method and electronic equipment, wherein the method comprises the following steps: acquiring audio features, spatial features and image features; inputting the audio features and the spatial features into an audio-spatial coding module, aligning the audio features and the spatial features by dynamically adjusting the spatial position of the audio features and the facial region, and obtaining spatially perceived lip features; inputting the audio features and the image features into an audio-guided dynamic fusion module, calculating audio-perceived channel attention and spatial attention by using a dynamic linear layer, and obtaining dynamic fusion features; acquiring reference key points, inputting the reference key points and the audio features into an expression optimization module, fusing the reference key points and the audio reference points, and obtaining key point features; and inputting the lip perception features, the dynamic fusion features and the key point features into a trained generation network to generate video frames. The method realizes a higher-quality, more natural and stable speech-driven portrait synthesis effect.
Owner:SOUTH CHINA UNIV OF TECH

Virtual anchor real-time driving system based on facial motion capture

ActiveCN121842342BResolve driver conflictsImprove viewing experienceTelevision system detailsColor television detailsPhoneme recognitionEngineering
The application discloses a virtual anchor real-time driving system based on facial motion capture, aims to solve the problems of insufficient precision, lack of naturalness and weak scene adaptation of virtual anchor facial motion in the prior art, and through multi-thread parallel collection of audio and video streams, optimizes data quality through adaptive preprocessing; adopts multi-dimensional facial feature collaborative extraction, double-branch phoneme recognition, face function area semantic segmentation and confidence dynamic weight adjustment technology, realizes accurate representation and collaborative fusion of expression and lip feature; finally, through the lip-expression collaborative driving mechanism, the real-time driving characteristics of the adaptive virtual image are output; thereby effectively improving the precision and naturalness of virtual anchor facial motion, enhancing the complex scene adaptability, giving consideration to real-time response and low-cost deployment, and being applicable to virtual live broadcast, online education and other scenes, and having good application value.
Owner:GUIZHOU NORMAL UNIVERSITY

Anti-spoofing authentication system

An anti-spoofing authentication system includes an image capture device that converts received light into an image signal representing a captured image; a microphone that converts sound wave into a voice signal; a face detection device that detects a human face on the captured image, the face detection device including a lip detection device that detects lip features associated with the detected human face; a voice activity detection (VAD) device that detects presence of human voice in the voice signal; and an anti-spoofing device that detects a spoofing attack according to the detected lip features and the detected human voice.
Owner:HIMAX TECH LTD

Gaussian sputtering conversation face generation method based on lip features and head posture guidance

The invention provides a Gaussian sputter conversation face generation method based on lip features and head posture guidance, which can effectively generate matched lip motion and stable posture action for different input audios under the condition of giving a speech video of a target object. The method comprises the following steps: firstly, constructing a three-dimensional pre-training data set through a face reconstruction technology, and learning a mapping relation between audio and lip movement by adopting a grid generation module; in the lip feature generation process, the modules are finely adjusted to generate lip features adaptive to the talking style of the target object; in the process of extracting the head posture parameters, optimizing the head posture parameters by adopting a feature point matching algorithm and a beam adjustment method, and smoothing the parameters by adopting a filter; and finally, introducing a dynamic three-dimensional Gaussian renderer, and synthesizing a vivid mouth shape synchronization video by taking the output of the two processes as a guidance condition. Compared with a previous method, the method has higher precision under the driving of intra-domain audio and cross-domain audio.
Owner:SICHUAN UNIV

Speaking face generation method and device, equipment and medium

The invention relates to an artificial intelligence technology, can be applied to business system platforms of medical health, financial science and technology and the like, and discloses a speaking face generation method, device, equipment and medium, and the method comprises the steps: obtaining a reference face image and an original video frame sequence, carrying out the coding, and generating a potential representation, performing lip mask processing on each frame of image in the original video frame sequence to obtain a local mouth image, and encoding the local mouth image to generate lip features; obtaining random sampling noise, and generating joint potential features based on the potential characterization, the lip features and the random sampling noise; on the basis of the reference face image features, combining the potential features and the input audio features, generating lip alignment features; and after the audio and lip alignment features are decoded, lip synchronization frames are generated and assembled, and a speaking face video is generated. According to the method, high-resolution visual quality can be kept, meanwhile, deep alignment of audio content, facial expressions and time sequence dynamic states is achieved, and the speaking face generation quality and generalization ability are improved.
Owner:平安科技(上海)有限公司

A voice interaction recognition enhancement method, device and storage medium

ActiveCN119207393BSpeech recognitionModal voiceLip feature
The application provides a speech interaction recognition enhancement method, which comprises the following steps: collecting a video of a speaker facing a camera for speech interaction, splitting the video into N segments of to-be-identified speech and N frames of to-be-identified images, and forming to-be-identified data; inputting the to-be-identified data into a preset speech interaction recognition enhancement model, wherein the model comprises a lip feature extraction network, a speech feature extraction network, a time feature extraction network and an activation network, the time feature extraction network is used for adding time sequence information to a lip feature matrix extracted by the lip feature extraction network and a speech feature matrix extracted by the speech feature extraction network, the activation network is used for simulating an activation-inhibition mechanism of visual information on an auditory nerve circuit in a biological brain, realizing the interaction of two modalities of vision and hearing, and obtaining a speech recognition result. The application realizes the accurate recognition of lip language and speech in a noisy environment by using the dual-modality speech recognition of vision and hearing, and the response capability and accuracy meet the requirements.
Owner:TSINGHUA UNIVERSITY

Video keyword retrieval method and device based on multiple encoders and electronic equipment

The invention provides a video keyword retrieval method and device based on multiple encoders and electronic equipment. The method comprises the following steps: acquiring a lip region image sequence of a to-be-detected video; the lip region image sequence is input into a visual feature extraction network, the visual feature extraction network comprises multiple visual encoders arranged in parallel, different types of visual features are extracted and fused, and multi-encoder visual features are obtained; inputting the to-be-detected keyword into a text encoder, and extracting to obtain a text feature; and performing alignment processing on the text feature and the multi-encoder visual feature to generate a joint feature, inputting the joint feature into a classifier, and outputting the probability of occurrence of the to-be-detected keyword in the video. A multi-encoder structure is introduced to carry out modeling on video lip features, cross-modal alignment is realized in combination with text features, and the problem that keyword detection cannot be carried out in a noisy environment or a silent scene in the prior art is effectively solved.
Owner:BEIJING YUANJIAN INFORMATION TECH CO LTD

Method and device for detecting other people to answer on behalf, computer equipment and storage medium

The invention discloses a detection method and device for other people to answer on behalf, computer equipment and a storage medium, and relates to the technical field of face-to-face detection in the fields of finance, insurance, medical treatment, banking and the like, and the method comprises the steps: obtaining a to-be-detected video stream and a corresponding audio stream; extracting a face region image from each frame of video image of the video stream; face key point detection is carried out on the face region image, lip key points are positioned, a lip 2D coordinate sequence is obtained, and the lip 2D coordinate sequence is converted into a lip feature map; extracting a face feature map of the face region image, and generating a visual feature vector based on the lip feature map and the face feature map; extracting an audio feature vector of the audio stream, and performing multi-modal alignment on the visual feature vector and the audio feature vector to obtain a fusion feature vector; and detecting the fusion feature vector based on a pre-trained classifier to obtain a detection result of answering by others on behalf. According to the invention, the accuracy of sound and lip synchronous detection can be greatly improved, so that other people's pickup behaviors can be effectively identified.
Owner:PING AN TECH (SHENZHEN) CO LTD

Virtual anchor real-time driving system based on facial motion capture

ActiveCN121842342Aexact matchPrecise linkageTelevision system detailsColor television detailsPhoneme recognitionEngineering
The invention discloses a virtual anchor real-time driving system based on facial motion capture, and aims at solving the problems that in the prior art, virtual anchor facial motions are insufficient in accuracy, lack of natural sense and weak in scene adaptation. Audio and video streams are collected in parallel through multiple threads, and the data quality is optimized through self-adaptive preprocessing; the technologies of multi-dimensional facial feature collaborative extraction, double-branch phoneme recognition, face functional region semantic segmentation and confidence coefficient dynamic weight adjustment are adopted, and accurate representation and collaborative fusion of expression and mouth shape features are achieved; and finally, outputting real-time driving characteristics matched with the virtual image through a mouth shape-expression cooperative driving mechanism. Therefore, the accuracy and the natural sense of the face action of the virtual anchor are effectively improved, the adaptability of complex scenes is enhanced, real-time response and low-cost deployment are both considered, and the method is suitable for scenes such as virtual live broadcast and online education and has good application value.
Owner:GUIZHOU NORMAL UNIVERSITY

A large model-based speech recognition and speech synthesis optimization method and system

ActiveCN121034309BSpeech recognitionSpeech synthesisData setLip feature
The application provides a speech recognition and speech synthesis optimization method and system based on a large model, which extracts lip movement features, lip feature time stamps and audio feature time stamps by acquiring real-time speech input signals and image frame sequences of a user's face area to generate an initial space-time offset sequence; generates a dynamic offset compensation parameter sequence based on the initial space-time offset sequence and in combination with a reference alignment template in a speech visual synchronization data set; jointly processes the dynamic offset compensation parameter sequence, the lip movement features and the real-time speech input signals by using a large model to reconstruct a target speech segment, generates a corrected phoneme sequence in combination with the lip movement features, retrieves a mouth shape parameter group corresponding to the corrected phoneme sequence from a phoneme mouth shape mapping rule library to generate a speech waveform synchronized with the lip movement phase; and the application improves the naturalness, immersion and robustness in a speech missing or delayed scenario of human-computer interaction.
Owner:LUSTER LIGHTWAVE CO LTD

Lip speech recognition method and system based on subspace sparse attention mechanism and medium

The present application relates to a kind of lip reading method based on subspace sparse attention mechanism, system and medium, method includes: obtaining lip region image sequence, based on the lip region image sequence extraction obtains lip feature sequence;The lip feature sequence is input to the preset training complete phoneme sequence extraction model, obtains the pronunciation phoneme sequence corresponding to the lip feature sequence;The pronunciation phoneme sequence is input to the sentence inference model with the subspace sparse self-attention mechanism built in and obtains target sentence sequence.The present application enhances the context information by constructing a special attention mechanism, realizes the prediction long sentence sequence in a forward operation, so that the inference rate and accuracy are greatly improved.
Owner:WUHAN UNIV OF TECH CHONGQING RES INST

Lip-reading-based voice interaction methods, devices, equipment, and storage media

ActiveCN120600019BLip featureHuman–computer interaction
This invention discloses a lip-reading-enhanced voice interaction method, apparatus, device, and storage medium. The lip-reading-enhanced voice interaction method includes: extracting lip-reading features from image sequences of the lip region; extracting audio features from the speech signal; performing cross-modal fusion encoding of the lip-reading features and audio features to generate hybrid features containing audiovisual information; inputting the hybrid features into a large language model to understand the intent of the interactive object and generate corresponding semantic responses; and finally synthesizing the speech and / or converting it into text. This invention, by introducing lip features, provides additional visual cues for speech recognition, significantly improving the robustness and accuracy of speech recognition; effectively fusing and encoding lip-reading features and audio features avoids the semantic information fragmentation caused by simple independent recognition; and fully utilizes the capabilities of a large model to achieve a more natural and intelligent interactive experience.
Owner:SHENZHEN WANRUI INTELLIGENT TECH CO LTD

Multimodal duration-controlled TTS (DC-TTS) system for automatic dubbing

PCT designated stageWO2025186682A1Natural language translationSpeech recognitionData setLip feature
A system and method for multimodal duration-controlled text to speech (TTS) is provided. The system receives a dataset including reference video clip, reference audio, and text. The system generates: text features based on application of a text encoder on the text, visual features based on application of a visual feature extractor on the reference video clip, a first speaker embedding associated with the human speaker based on the reference audio, and lip features associated with the human speaker based on the reference video clip. The system further generates audio tokens based on application of a neural language model on the text features, the visual features, the first speaker embedding, and the lip features. The system generates an audio based on application of a neural vocoder on the audio tokens and trains the neural language model based on the generated audio and a loss function.
Owner:SONY GROUP CORP

Lip reading method and device, electronic equipment, storage medium and program product

PendingCN122337199ANerve networkNetwork output
This invention relates to the field of lip-reading technology, providing lip-reading recognition methods, devices, electronic devices, storage media, and program products. The lip-reading recognition method includes: acquiring audio data and facial data including lip images; obtaining a lip feature map based on the facial data and an audio feature sequence based on the audio data; inputting the lip feature map and the audio feature sequence into a multi-scale interactive modal fusion network to obtain multi-modal lip-reading features output by the multi-scale interactive modal fusion network; the multi-scale interactive modal fusion network integrates audio and visual modal information; inputting the multi-modal lip-reading features into a lip-reading recognition neural network to obtain the recognition result output by the lip-reading recognition neural network; the lip-reading recognition neural network performs result recognition on the input data through a self-attention mechanism. This invention designs a multi-scale interactive modal fusion network that integrates audio and visual features, effectively improving the accuracy and stability of lip-reading recognition.
Owner:CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1

Method for detecting pickup based on synthetic lips and related equipment

The embodiment of the invention belongs to the technical field of artificial intelligence, and relates to a pickup detection method and related equipment based on synthetic lips, and the method comprises the steps: carrying out the lip sequence extraction processing of original video data according to a CNN face detector, and obtaining a lip feature sequence; performing synthetic lip sequence generation processing on the original audio data according to a pre-trained lip generation model to obtain a synthetic lip feature sequence; calculating the difference between the lip feature sequence and the synthesized lip feature sequence in the low-dimensional embedding space to obtain a difference feature sequence; performing action segmentation processing on the difference feature sequence according to a multi-scale time convolution network to obtain an action probability distribution sequence with the same length as the input time step; and performing pickup recognition processing on the action probability distribution sequence according to the classifier to obtain a pickup detection result. The method can be used for carrying out related video processing in a financial science and technology business system, and can effectively detect pickup behaviors in scenes such as online examinations, remote interviews, video conferences and the like.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent voice recognition and interaction system and method fusing AI visual information

The invention discloses an intelligent voice recognition and interaction system and method fusing AI visual information, and relates to the technical field of voice recognition. The method comprises the following steps: synchronously acquiring audio and RGB-D video streams, and extracting acoustic features, visual lip language features and three-dimensional space interaction features in parallel by using a double-flow convolutional neural network; the method comprises the following steps: constructing a cross-modal dynamic gating fusion network, combining a real-time signal-to-noise ratio and visual confidence, dynamically distributing a sound visual weight, and realizing high-robustness recognition based on lip language in a high-noise environment; aiming at anaphora ambiguity in a natural language, realizing intention analysis of what you see is what you control by utilizing projection and collision detection of sight lines and gesture rays in a three-dimensional semantic map; according to the invention, the problems of low recognition rate, unknown reference and tedious wake-up in a complex sound field environment are effectively solved.
Owner:FUJIAN CLOUD INTELLIGENT TECH CO LTD

A low-resource language lip reading method and device based on pre-training fine-tuning

The present application relates to the technical field of computer vision, and particularly relates to a low-resource language lip reading method and device based on pre-training fine-tuning. The method comprises: pre-training a model by using a large English video dataset to ensure that the model has strong generalization ability and effective lip feature expression ability; then, after loading the pre-trained model weight, fine-tuning the model by using a small amount of Tibetan lip reading dataset to overcome the challenge of Tibetan video data scarcity. In the decoding stage, a Transformer language model specially trained for Tibetan text is introduced, which effectively reduces the homonym confusion problem that may occur in the lip reading process, thereby improving the accuracy of sentence-level Tibetan lip reading. The overall architecture is improved by the above innovative structure and method, and effective pure visual lip reading of low-resource languages is successfully realized.
Owner:MINZU UNIVERSITY OF CHINA

A Deep Learning-Based Lip Correction System

This invention discloses a deep learning-based lip correction system, comprising an image acquisition module, an image preprocessing module, a lip key region segmentation model design module, a lip feature detection model design module, and a lip correction module. This invention belongs to the field of image processing, specifically referring to a deep learning-based lip correction system. This scheme introduces a dynamic grayscale weighted mask to significantly amplify the gradient in low grayscale regions; it suppresses weights in high grayscale regions through dynamic weights, adds an edge-aware term to the loss function to enhance sensitivity to lip line edges, and improves the accuracy of lip key region segmentation; thereby improving the reliability of subsequent lip correction; it combines rotation region representation with endpoint fitting to accurately characterize the key lip region; and it introduces endpoint offset penalty loss and angle fitting loss to precisely penalize corner point errors, eliminate fitting jumps, improve the positioning accuracy of key lip control points, and provide precise anchor points; thus improving the final lip correction effect.
Owner:JINYU ZHIDA TECHNOLOGY (TIANJIN) CO LTD

Lip movement detection method and device, electronic equipment and readable storage medium

The invention discloses a lip movement detection method and device, electronic equipment and a readable storage medium, and belongs to the field of image processing. The method comprises the steps that an image frame of a target face of a detection target is acquired, face feature points of the target face in the image frame and space coordinates of the face feature points are determined, and the face feature points at least comprise eye feature points and lip feature points; calculating the face reference distance of the target face according to the space coordinates of the eye feature points, and calculating the lip distance of the target face according to the space coordinates of the lip feature points; and performing dynamic proportion calculation based on the face reference distance and the lip distance, and determining the lip state of the target face in the image frame based on a dynamic proportion calculation result. According to the invention, the cost of lip movement detection is reduced.
Owner:GRG INTELLIGENT TECH SOLUTION CO LTD

Photovoltaic power station monitoring method and system based on voice assistance

The invention relates to the technical field of photovoltaic power station monitoring, and particularly discloses a photovoltaic power station monitoring method and system based on voice assistance, and the method comprises the steps: respectively training a lip information diagnosis model and a voice information recognition model; obtaining current voice information of an operator on duty of the photovoltaic power station and a current lip image corresponding to the current voice information, and diagnosing current lip features in the current lip image by using the trained lip information diagnosis model, so as to obtain standby initial information of each Chinese character in the current voice information; recognizing each sensitive word in the current voice information by using the trained voice information recognition model, and matching the initial consonant of each sensitive word with the standby initial consonant information set; when the matching is successful, matching all sensitive words in the current voice information with a task library; and executing the task with the highest matching rate with all the sensitive words in the current voice information in the task library. According to the invention, the monitoring efficiency, accuracy and user experience of the photovoltaic power station are obviously improved.
Owner:NANTONG ALPHA ESS CO LTD