Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

160 results about "Facial motion" patented technology

Emotion detection system based on facial recognition

The invention discloses an emotion detection system based on facial recognition, and relates to the technical field of computer vision and emotion calculation. A video stream time sequence analysis module is used for extracting a facial micro-expression image sequence of continuous frames, a time sequence feature vector containing a micro-expression intensity gradient, an illumination robustness coefficient and a facial action unit cooperation feature is generated, and a multi-mode dynamic sensing module is combined to carry out real-time analysis on an emotion classification probability, voice emotion parameters and physiological signals. And the fusion decision module performs dynamic weighted fusion on the multi-modal data based on the scene adaptive weight, and finally generates a comprehensive emotion score. Through multi-modal time sequence modeling and a dynamic weight optimization mechanism, the accuracy and environmental adaptability of emotion recognition are remarkably improved, and real-time perception and accurate decision making of customer emotion are realized in a target scene.
Owner:NORTHEAST FORESTRY UNIV

Audio and video depth forgery detection method based on quality perception and multi-scale alignment

The invention discloses an audio and video depth forgery detection method based on quality perception and multi-scale alignment, and the method comprises the following steps: coding a synchronous audio and video sequence, and obtaining a frame-level visual feature, a facial action unit and a phoneme-level voice representation; a visual quality evaluation module is introduced to generate a spatial reliability mask, and quality weighting is carried out on the visual features; designing a global-local multi-scale cross-modal alignment mechanism, performing bidirectional cross-attention modeling on voice and face dynamic synchronization globally, and performing physiological coupling alignment on phonemes and face action units locally; and an uncertainty perception reasoning and calibration scheme is provided, adaptive temperature scaling is carried out according to quality and consistency, and uncertainty calibration is carried out by self-supervision loss. According to the method, the problems of insufficient robustness and excessive self-confidence misjudgment of an existing method in a low-quality video and high-synchronization counterfeit scene are solved, and the cross-dataset generalization capability and the actual deployment reliability are remarkably improved.
Owner:NANJING UNIV OF SCI & TECH

Method for driving emotion interaction of intelligent device based on multi-modal understanding

The invention relates to the technical field of data processing, in particular to a method for driving emotion interaction of an intelligent device based on multi-modal understanding, and aims to eliminate illumination and noise interference and output a standardized face video stream, an effective voice segment and a touch thermodynamic diagram through an environment adaptive acquisition module. The feature extraction module extracts facial action optical flow features, voice Mel-frequency cepstral coefficient vectors and tactile pressure gradient parameters. The cross-modal correlation model adopts a tensor decomposition algorithm to calculate a space-time correlation matrix of visual and voice features, and the tactile feature weight is dynamically adjusted in combination with environmental parameters. According to the response strategy, an intervention scheme is retrieved based on a graph database, emotion confirmation statements, guide statements and behavior suggestions are fused to generate multi-mode response, and PID adjustment of the temperature control device and tactile pulse output of the vibration device are synchronously driven. And the feedback evaluation module verifies the emotion recognition consistency through a Pearson's correlation coefficient, triggers conflict sample separation storage and model increment training, and realizes closed-loop optimization.
Owner:BEIJING HAOXINQING MOBILE MEDICAL TECH CO LTD

Man-machine physical twin system based on multi-modal video analysis and adaptive mapping and control method

The invention belongs to the technical field of man-machine interaction and robot control, and particularly relates to a man-machine physical twin system based on multi-modal video analysis and adaptive mapping and a control method, and the system comprises a multi-modal collection unit which is used for collecting original video streams of human body actions, expressions and voices, and compensating and repairing dynamic shielding; the semantic analysis unit comprises an action analysis module used for extracting a human skeleton key point set object interaction track from the original video stream, and an expression / voice analysis module used for outputting a facial action unit coefficient and a voice text with an emotion label; the self-adaptive mapping unit is used for establishing human body-robot joint kinematics mapping, distributing control bandwidth in real time according to task types and degrading non-key modes based on network states to guarantee action continuity, and non-contact action capture is achieved through a pure vision scheme by eliminating dependence on a wearable sensor; and the problems of failure and the like of a traditional visual scheme under shielding and illumination variation are solved.
Owner:AIMI (BEIJING) ROBOT CO LTD

Facial action unit recognition and model training method and device, and electronic equipment

The invention provides a facial action unit recognition and model training method and device and electronic equipment. The method comprises the steps of obtaining a training sample set and inputting the training sample set into an initial facial action unit recognition model; wherein the training sample set comprises multiple frames and face action unit (AU) real labels corresponding to each frame; performing multi-level feature extraction on each frame by the facial action unit recognition model, generating a visual token according to the extracted features, performing semantic reasoning based on the visual token and a prompt text of a facial action unit recognition task, and generating an AU prediction probability corresponding to each frame; determining a loss value of the loss function according to the AU real label and the AU prediction probability, and reversely adjusting model parameters of the facial action unit recognition model based on the loss value; when a preset training condition is met, training is stopped, and a trained facial action unit recognition model is obtained. According to the invention, semantic reasoning can be carried out by using visual features to complete AU identification, and the accuracy of a facial action unit detection result can be improved.
Owner:BEIZHI TECHNOLOGY (ANJI) CO LTD

Emotion recognition method and device for road and bridge engineering scenes

The present invention provides an emotion recognition method and device for road and bridge engineering scenarios, relating to the field of emotion recognition technology. The present invention pre-processes the acquired user's facial video and physiological information, then performs feature extraction to obtain the user's facial motion units and physiological features; then selects and fuses the optimal feature subset of the user's facial motion units and physiological features; and finally, based on the multi-channel feature inverse reasoning of a dynamic Bayesian network and an interpretable emotion recognition model, obtains the mapping relationship between the fused optimal feature subset and the emotion component, thereby obtaining the user's emotional state recognition result. Compared with the existing technology, the emotion recognition system of the present invention is more reliable and robust, and can achieve multi-level, all-round, and stable emotion recognition of users in the complex environment of road and bridge engineering.
Owner:HEFEI UNIV OF TECH

Human body state detection method, system and device based on facial action and medium

The invention relates to a human body state detection method, system and device based on facial actions and a medium, and the method comprises the steps: obtaining an image sequence of a human face, carrying out the human face detection and mark point positioning of the image sequence, and obtaining the coordinates of a face mark point; performing micro-expression action extraction on the image sequence based on the facial mark point coordinates to obtain facial micro-expression action features; performing time sequence analysis on the facial action features to obtain time sequence facial action features; performing multi-scale illumination adjustment on the time sequence facial action features to obtain illumination facial action features; performing multi-modal analysis on the illumination facial action features through a pre-trained long-short term memory network to obtain comprehensive human body state features; and performing real-time state analysis based on the comprehensive human body state characteristics to obtain a human body state detection result. According to the method, different types of facial action features can be comprehensively processed, and the reliability of a detection result is enhanced.
Owner:SHENZHEN ELM TECH CO LTD

Motion capture combined movie and television animation character expression generation method and system

The invention relates to the technical field of movie and television animation production, in particular to a movie and television animation character expression generation method and system combined with motion capture, and the method comprises the steps: firstly collecting a real-time motion data set of a target actor through motion capture equipment, including face and body motion data subsets; the facial action data is composed of facial key point displacement data with multiple continuous timestamps, then calling a preset expression mapping model to decouple the facial action data to obtain a basic and personalized expression feature set, then adjusting the basic expression feature set according to preset emotion parameters of an animation role, generating an emotion adaptive expression feature set, and generating a personalized expression feature set according to the emotion adaptive expression feature set; the method comprises the steps of obtaining a personalized expression feature set, fusing with the personalized expression feature set to obtain a target expression feature set, finally driving a three-dimensional face model to generate an expression animation by using the target expression feature set, and binding body action data to a three-dimensional body model to generate a complete character animation, thereby improving the authenticity and expressive force of animation character expressions.
Owner:CHENGDU LIFANG FANTASY TECHNOLOGY CO LTD

TTS voice and 3D mouth shape synchronous generation method, apparatus and device, and medium

The invention relates to the technical field of finance, medical health and artificial intelligence, and provides a TTS voice and 3D mouth shape synchronous generation method, device, equipment and medium, a TTS voice and 3D mouth shape synchronous generation model is obtained by using a small sample adaptive training engine and based on transfer learning framework training, and the generalization ability and training efficiency of the model are improved; multi-modal joint coding and a dynamic weight distribution module are used for joint coding, so that the precision of a generated result is improved; a 3D mouth shape dynamic sequence is obtained based on a diffusion-autoregression hybrid generation module, the detail generation capability of a diffusion model and the time sequence coherence of an autoregression model are fused, and the generation precision and real-time performance are improved; and the real-time synchronization module based on cross-modal attention guidance obtains the synchronously adjusted 3D mouth shape parameters and the facial action sequence matched with the emotion and generates a target video, so that multi-modal dynamic fusion is realized, the lip movement-voice synchronization precision is improved, and the authenticity of emotion expression is enhanced.
Owner:PING AN TECH (SHENZHEN) CO LTD

Vehicle-mounted expression data construction method, system and device based on GAN and vehicle

The invention provides a GAN-based vehicle-mounted expression data construction method, system and device and a vehicle, and relates to the technical field of intelligent driving. A key point detection technology is utilized to analyze and extract a facial action unit and expression intensity and duration thereof to generate a structured action description. And then, calling a generative adversarial network model to generate a synthetic video corresponding to the target expression type, and carrying out dynamic cutting and frame extraction according to a preset interval on the videos to obtain a multi-frame synthetic image. And finally, performing dual-model filtering on the synthesized image to ensure data quality, and executing privacy anonymization processing, thereby generating a high-quality vehicle-mounted expression data set. According to the method, the data acquisition cost is reduced, the diversity and accuracy of a data set are improved, and the performance and reliability of an expression recognition algorithm in an actual vehicle-mounted environment are effectively enhanced.
Owner:CHINA FAW CO LTD

Face forgery identification method, face forgery identification model, equipment and medium

The invention relates to a face forgery identification method, a face forgery identification device, equipment and a medium. The method comprises the following steps: determining a first frame image from a to-be-identified video; wherein the first frame image is a frame image of which the face action amplitude change exceeds a threshold value; acquiring first optical flow motion information of the first frame image and second optical flow motion information of the second frame image; wherein the sampling time of the second frame image is before the sampling time of the first frame image; performing gradient calculation on the first frame image to obtain first gradient information corresponding to the first frame image; and based on the first optical flow motion information, the second optical flow motion information and the first gradient information, obtaining a face forgery identification result of the to-be-identified video.
Owner:ACADEMY OF BROADCASTING SCI STATE ADMINISTATION OF PRESS PUBLICATION RADIO FILM & TELEVISION

Facial emotion analysis method and system based on multi-modal alignment training

The invention relates to the technical field of face recognition, and discloses a multi-modal alignment training-based face sentiment analysis method and system, and the method comprises the steps: obtaining multi-source data, carrying out the preprocessing of the multi-source data, obtaining a training set, constructing a basic model for the face sentiment analysis, selecting a sample with an inference text to generate a small number of high-quality sentiment analysis samples, and carrying out the recognition of the high-quality sentiment analysis samples; supervising and finely adjusting the model; selecting a sample with a facial action unit label and an emotion label, and performing reinforcement learning training on the model in combination with the predicted accuracy of the facial action unit label, the emotion label and the reasoning text; and training the basic model by using the training set, expanding the original training set by using the output of the trained basic model to obtain a new training set, continuously training the model, stopping training until a preset condition is met to obtain a final basic model, and performing facial sentiment analysis by using the final basic model. According to the method, illusion can be controlled, the accuracy of a facial emotion analysis result is improved, and the data set construction cost is reduced.
Owner:SUZHOU UNIV

Multi-modal face restoration and expression recognition system and method based on semantic guidance of facial action unit

ActiveCN121998875AOvercome the problem of physiological distortion in repair resultsGet rid of dependenceImage enhancementSemantic analysisVisual technologyLinguistic model
The invention belongs to the technical field of image restoration and computer vision, and particularly relates to a multi-mode face restoration and expression recognition system and method based on semantic guidance of a facial action unit. According to the system, multi-scale features are extracted through visual coding, and an AU activation probability is detected by using a graph neural network; the semantic conversion module converts the numerical probability into an interpretable biomechanical structured text; the multi-modal reasoning module fuses vision and text information, introduces the common sense reasoning ability of a multi-modal large language model, and improves the student network performance through knowledge distillation; and finally, the conditional generation module realizes image restoration by taking the semantic features as guidance. The facial action unit is used as a biomechanical medium, the multi-modal reasoning ability is converted into restoration constraint, the defects that in the prior art, restoration of physiology is distorted, recognition depends on image quality, and two tasks are isolated are overcome, collaborative enhancement of face restoration and expression recognition is achieved, and it is ensured that the restoration result is clear in vision and conforms to the physiological law of facial muscles.
Owner:YANGTZE RIVER DELTA RES INST OF NPU TAICANG +1

Digital human video rendering method, device and medium

The invention discloses a digital human video rendering method and device and a medium, and relates to the crossing field of computer graphics and generative adversarial networks, and the method comprises the steps: constructing a generative adversarial model based on a generative adversarial network of a single fusion architecture; performing multi-modal preprocessing on the original audio and video data based on the digital human reference image; extracting dual-granularity speech features through a speech feature extraction module of a generative adversarial model, and fusing the dual-granularity speech features; determining a local deformation field of the digital human reference image in a UV parameterized space based on a reference key point corresponding to the digital human reference image and the fused speech features; sampling the identity texture of the digital human reference image to generate a target digital human face image; and verifying the digital human face image based on multiple scales through a discriminator of a generative adversarial model. Through dual-granularity speech feature fusion and generator multi-resolution injection, deep alignment of speech semantics and facial actions is realized.
Owner:山东浪潮智慧建筑科技有限公司

Multi-modal behavior data processing system

The invention discloses a multi-modal behavior data processing system, which comprises a data acquisition module, a feature extraction layer, a dynamic attention weight layer, a multi-modal fusion layer and a downstream task decision-making layer, the data acquisition module guides human-computer interaction through an international neurological and mental interview tool matched with a DSM-5 standard and acquires audio and video stream data; the feature extraction layer extracts a video feature vector (including facial action unit activation intensity and the like), an audio feature vector (including Mel frequency cepstrum coefficient and the like) and a text feature vector (generated by a deep language model after automatic speech recognition transcription) in parallel; the dynamic attention weight layer is combined with data quality, symptomatic priori knowledge and cross-modal correlation to generate a dynamic fusion weight; weighting, splicing and dimensionality reduction are carried out on the multi-modal fusion layer to obtain a fusion feature vector; and the downstream task decision-making layer completes evaluation and generates a multi-modal behavioral index evaluation report. The system is deployed in a non-intrusive manner, the risk is controllable, and the evaluation robustness and accuracy can be improved.
Owner:NEW MAYO HEALTH MANAGEMENT RESEARCH INSTITUTE (CHONGQING) CO LTD +1

Training instances of machine learning model for facial expression prediction and generating new avatars used in training

For each avatar, testing images are rendered for different facial expressions that each have ground truth facial action units. An instance of a machine learning model is applied to the testing images to generate predicted facial action units for each testing image. A predictive performance of the instance is calculated for each avatar based on the predicted and ground truth facial action units for the testing images of the avatar. A first set of features common to the avatars for which the predictive performance was better than a first threshold, and a second set of features common to the avatars for which the predictive performance was worse than a second threshold, are identified. The features present only in the second set are identified, as difference features. New avatars having the difference features are generated.WO
Owner:PURDUE RES FOUND +1

Setting a region of interest of a head-mounted camera based on facial movements

Utilization of windowing to set a region of interest (ROI) of a camera used for tracking facial expressions. In one embodiment, a system includes an inward-facing head-mounted camera that captures images of a region on a user's head utilizing a sensor that supports changing of its ROI. The system also includes a computer that detects, in a first subset of the images, a first sub-region in which changes due to a first facial movement reach a first threshold and reads from the camera a first ROI that covers at least a portion of the first sub-region. The computer detects, in a second subset of the images, a second sub-region in which changes due to a second facial movement reach a second threshold, and then reads from the camera a second ROI that covers at least a portion of the second sub-region, with the first and second ROIs being different.
Owner:FACENSE LTD

Bullet screen generation method and device, electronic equipment and storage medium

The invention relates to a bullet screen generation method and device, electronic equipment and a storage medium, and the method comprises the steps: collecting the current face data of a watching user in real time through camera equipment in a process of playing a target media content, analyzing a face action unit combination from the face data, and generating a bullet screen. And determining current emotional state data of the watching user according to the facial action unit combination, and generating and displaying bullet screen content in real time according to the emotional state data. The bullet screen content highly matched with the current emotional state of the user can be generated in real time according to the objective face data of the user, the physiological signals are dynamically converted into the bullet screen content, the bullet screen real-time performance and the emotional expression rate are improved, and the real emotional state of the user is transmitted.
Owner:SHANGHAI ZHONG YUAN NETWORK CO LTD

Cranial nerve motor function evaluation system and method based on terminal with camera

The invention relates to a cranial nerve motor function evaluation system and method based on a terminal with a camera. The system comprises a camera used for collecting a high-definition video sequence of facial actions; the model is used for identifying positioning points of the face according to the high-definition video sequence acquired by the camera; the analysis and evaluation module is used for quantitatively evaluating the cranial nerve motor function state according to the change of the positioning points of the face acquired by the model; the analysis and evaluation module performs analysis and evaluation according to the change quantitative indexes of the positioning points of the face; the camera, the model and the analysis and evaluation module are all deployed at the terminal; and the model and the analysis and evaluation module are deployed in an operating system of the terminal. The evaluation system and method are deployed on a common terminal, and are used for collecting a video sequence of a user completing a standard evaluation action, automatically identifying the movement of a positioning point of a face, and outputting a quantitative score according to a medical standard so as to evaluate human face muscle movement and a cranial nerve function state.
Owner:BEIJING NEURORIENT TECHNOLOGY CO LTD +1

Driver fatigue behavior detection method and system based on video pose invariance

The application discloses a driver fatigue behavior detection method and system based on video posture invariance, and relates to the technical field of computer vision. The application proposes a key frame selection model based on facial geometric information and a head and face action information fusion space-time network. First, the driver video captured by the vehicle-mounted camera is subjected to sequential processing, and image preprocessing is performed. Then, the key frame selection model based on facial geometric information is constructed based on the geometric features of the facial key points and a two-stage decision mechanism, and the key frames in the video sequence are extracted. Finally, the facial action modalities under any posture are extracted based on the facial forward processing, and the head posture attributes obtained based on the head posture estimation are combined to construct the head and face action information fusion space-time network, which is used for detecting the yawning, speaking, normal and other driver states. The application fully considers the head posture attributes, has high posture robustness, and can effectively distinguish the yawning and other fatigue behaviors from other driver states.
Owner:SHANDONG UNIV

Method and system for matching facial expressions and actions of digital cat puppet image through human voice

The invention discloses a method for matching facial expressions and actions of digital cat puppet images through human voice, which comprises the following steps of: S1, acquiring voice data of a user, and preprocessing the voice data; s2, converting the preprocessed voice data into text data, performing sentiment analysis, and extracting sentiment information; s3, constructing an action library, pre-defining various facial actions of the digital cat puppet, and distributing an emotion label for each action; according to a sentiment analysis result, selecting proper facial expressions and actions for mapping; and S4, generating a corresponding facial action according to a mapping result, and rendering the generated facial action to a model of the digital cat puppet in real time. According to the method, dynamic interaction between a virtual character and user voice input is realized through voice emotion driven digital cat puppet face action matching and real-time rendering, based on voice preprocessing and emotion analysis technologies, emotion information in voice can be accurately extracted, and the naturalness and coherence of digital cat puppet face actions are ensured.
Owner:CHINA UNICOM WO MUSIC & CULTURE CO LTD +1

A method, apparatus, and electronic device for detecting attacks targeting identity authentication.

This application provides a method, apparatus, and electronic device for detecting attacks on identity authentication, relating to the field of identity authentication technology. The method for detecting attacks on identity authentication includes: performing specified segmentation processing on target audio and target video to obtain multiple data segments, wherein the multiple data segments include various audio segments and various video segments, and the target audio and target video are content recorded during user authentication based on reading specified verification text; extracting speech features from each audio segment to obtain feature vectors for each audio segment, and extracting facial motion features from each video segment to obtain feature vectors for each video segment; determining the deviation result corresponding to each target data segment of a specified media type based on the obtained feature vectors; and determining the detection result based on the deviation result corresponding to each target data segment. Therefore, this solution can improve the accuracy of attack detection.
Owner:HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Voice-driven lip shape generation method, device and apparatus

The application discloses a speech-driven lip shape generation method, device and equipment, and relates to the technical field of artificial intelligence. The method comprises the following steps: extracting facial motion parameters and face identification features based on target image frames of a video sequence; encoding a driving audio sequence to obtain a time sequence audio feature sequence; performing prediction based on the facial motion parameters and the time sequence audio feature sequence to obtain target implicit key point expression coefficients; and generating target lip shape video frames based on the target implicit key point expression coefficients and the time sequence audio feature sequence. The application further generates target lip shape video frames by calculating target implicit key point expression coefficients, simplifies the overall process of lip shape generation, and improves the efficiency of speech-driven lip shape generation.
Owner:CHINA MERCHANTS BANK

ADHD multi-feature extraction and fusion classification method and system based on original video

The application discloses an ADHD multi-feature extraction and fusion classification method based on original videos, and comprises the following steps: S1, a network camera is used to collect video records of a subject watching videos, the videos are preprocessed, and preprocessed image frames are obtained; S2, the preprocessed image frames are analyzed to obtain behavior modes including facial actions, eye movements and head movements; S3, feature components of the behavior modes are extracted and fused; and S4, a deep learning network is constructed to classify the fused features. The application further discloses an ADHD multi-feature extraction and fusion classification system based on original videos. The application classifies ADHD patients based on video sequences, avoids invasive influence, reduces cost and is easy to popularize. Through multi-modal feature fusion, the limitation of single modal data is reduced, better accuracy and effectiveness are achieved, and in addition, the application can also be used for classifying autism cases.
Owner:ANHUI MEDICAL UNIV

Fatigue driving detection method and device, computer equipment and storage medium

The invention relates to the technical field of face recognition, and discloses a fatigue driving detection method and device, computer equipment and a storage medium, and the method comprises the steps: collecting video stream data in a driving process, and employing a preset target detection model to position an initial face region based on the video stream data; fitting face feature points based on the initial face region, constructing a face feature triangle based on the face feature points, constructing a face feature vector based on the face feature triangle, and constructing a driver state analysis data set according to a time sequence; based on the state analysis data sets and a preset sliding window, obtaining facial feature vectors in all the state analysis data sets, and projecting all the facial feature vectors to a preset face projection reference plane to obtain a facial motion feature point set; and calculating a facial motion information entropy based on the facial motion feature point set to detect the driving state of the driver. The problems that the driving state of a driver cannot be detected in real time in a complex scene and the detection accuracy is low are solved.
Owner:GUANGDONG COMM POLYTECHNIC

A bionic robot facial movement debugging method and related equipment

The present application provides a method and related equipment for debugging facial movements of a bionic robot, and relates to the field of intelligent robot technology. The present application can automatically adjust the motor movement amplitude of a bionic robot to be debugged by using a manually fine-tuned reference bionic robot as a reference object, ensuring that the reference bionic robot and the adjusted bionic robot to be debugged can directly display the same or similar facial expression effects when performing the same facial movements, thereby effectively reducing the manual adjustment workload and the deviation in the robot's facial movement performance during the robot mass production process, and improving the robot's adjustment efficiency and the consistency of the robot's facial movement performance, thereby meeting the timeliness requirements and facial movement performance consistency requirements during robot mass production.
Owner:UBTECH ROBOTICS CORP LTD

Pressure detection model training method and device, electronic equipment, and storage medium

The application relates to the fields of artificial intelligence and medical health, and discloses a training method and device of a stress detection model, an electronic device and a storage medium, which comprises the following steps: constructing a stress detection model, wherein the stress detection model comprises a feature extraction layer, a feature fusion layer and a result prediction layer which are connected in sequence; the feature extraction layer comprises an action unit feature recognition network, an emotion feature recognition network, a performance feature recognition network and a latent feature recognition network; constructing a training set, each data sample in the training set is a facial image of an identified object, and the data sample contains annotations of facial action unit features, facial emotion features, facial performance features and facial image coding features; inputting the training set into the stress detection model for training to obtain the stress detection model. The application greatly reduces the stress detection cost, improves the detection result efficiency, improves the operability, facilitates user use, and has high practicability.
Owner:PING AN TECH (SHENZHEN) CO LTD

A data processing method and system for multimodal facial motion point data and vocal cord motion data

The present invention discloses a data processing method and system for multimodal facial motion point data and vocal cord motion data. The method includes providing text, collecting continuous facial images or videos and laryngeal vibration data of a normal person speaking, preprocessing the data, extracting temporal and spatial features, establishing a facial and neck motion model for Chinese pronunciation, and having a deaf-mute person imitate the pronunciation according to the model and obtain feedback. The system includes a depth camera, a laryngeal vibration sensor, and a microphone. By comprehensively utilizing multimodal data, it provides instant feedback to the deaf-mute, lowers the learning threshold, and improves communication efficiency. It is applicable to deaf-mute groups worldwide. This invention promotes speech vocalization training and has broad application prospects and social significance.
Owner:ZHEJIANG UNIV

Peripheral facial paralysis rehabilitation device

The invention provides a peripheral facial paralysis rehabilitation device. The peripheral facial paralysis rehabilitation device comprises a host and at least one stimulation patch electrically connected with the host, the stimulation patch is prefabricated into a special-shaped geometric structure matched with the anatomical direction of facial muscles, and at least integrates a myoelectricity acquisition part, an electrical stimulation part and a heating part; the host is configured to collect a surface electromyographic signal sequence of a user under a specified facial action through the electromyographic collection part of the worn stimulation patch, and the stimulation patch worn by the user is in one-to-one correspondence with target facial muscles involved in the specified facial action; determining at least one muscle function state index according to the surface electromyogram signal sequence; determining a working mode according to the muscle function state index; and controlling the electrical stimulation part and / or the heating part of the stimulation patch worn by the user to work according to the working parameters corresponding to the determined working mode.
Owner:SHUGUANG HOSPITAL AFFILIATED WITH SHANGHAI UNIV OF T C M

Face driving model training method and device, equipment, medium and product

The invention discloses a face-driven model training method and device, equipment, a medium and a product. The method comprises the following steps: extracting a target face motion coefficient from three-dimensional model training data, wherein the face motion coefficient is used for representing face motion information corresponding to the three-dimensional model training data; performing motion synthesis on audio training data corresponding to the three-dimensional model training data through a first generator of a cyclic generative adversarial network to obtain a generated face motion coefficient, the first generator at least comprising a U-Net network; the generated face motion coefficient and the target face motion coefficient are discriminated through a first discriminator of a cyclic generative adversarial network, a first discrimination result is obtained, and the first discriminator at least comprises a pyramid network; and iteratively updating model parameters of the initial face driving model based on the first judgment result to obtain a target face driving model. According to the scheme provided by the invention, the training efficiency of the face driving model is improved while the training precision of the face driving model is improved.
Owner:CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1