Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

843 results about "Facial affect" patented technology

Multi-sensory autonomous multimodal emotion-synchronized environmental control architecture and regulation system (amesecar)

An autonomous environmental regulation and behavioral monitoring system is disclosed, configured to adapt temperature, lighting, and acoustic conditions based on real-time emotional and physiological data. The system includes a dual-redundant central processor, hierarchical communication networks, multi-angle visual acquisition units, infrared thermometers, and modular environmental subsystems. It detects posture, gestures, facial expressions, and thermal signals to classify user states and apply individualized airflow, light, and sound modulation without relying on external internet connectivity. The system also monitors connected appliances using voltage-based pressure analysis to forecast device degradation. With integrated gesture recognition, privacy-preserving data handling, and predictive adaptation, the invention enables multi-user personalization, long-term learning, and uninterrupted operation within residential, administrative, or healthcare infrastructures.
Owner:SEYEDKHAMOUSHI FAEZEHALSADAT +1

Psychological accompanying method based on multi-modal emotion recognition

The invention discloses a psychological accompanying method based on multi-modal emotion recognition, and belongs to the technical field of psychological health services, and the method comprises the steps: S1, synchronously collecting physiological signals, voice features, facial expressions and interactive behavior data of a user through a multi-modal sensor; s2, performing fusion analysis on the multi-modal data by using a deep learning model, and identifying a current emotional state and an emotional intensity level of the user; and S3, dynamically generating an adaptive psychological accompanying intervention scheme based on a preset emotion-intervention strategy mapping rule in combination with historical emotion data and personalized preferences of the user. According to the method, the physiological signals, the voice features, the facial expressions and the interactive behavior data are synchronously collected through the multi-modal sensor, the deep learning model is used for fusion analysis, and compared with single-modal recognition, the emotion state and the intensity level of the user can be judged more comprehensively and accurately, the emotion misjudgment risk is reduced, and a reliable basis is provided for subsequent intervention.
Owner:刘梓宸

Employment and entrepreneurship support system based on artificial intelligence

The invention discloses an employment and entrepreneurship support system based on artificial intelligence, relates to the technical field of occupational planning, and aims to solve the problem that a traditional support mode is insufficient in individuation and accuracy. The system comprises a multi-modal data acquisition module used for acquiring facial expressions, whole body dynamics and voice data of a user in at least 30 minutes of video dialogue in real time; a dialogue text is processed through a large language model based on a Transform architecture, continuous time sequence modeling is carried out on multi-modal dynamic behaviors by applying a liquid time constant network, and deep fusion reasoning is carried out in combination with multiple groups of multi-head Transform attention mechanisms. The core analyzes the real thought, psychological state, behavior pattern and core ability of the user through consistency verification, generates a structured dynamic user insight abstract, and customizes personalized vocational development or entrepreneurship planning according to the structured dynamic user insight abstract. According to the method, the potential of the user can be deeply informed, high-precision personalized planning is provided, the decision-making quality and success rate are improved, and the method has dynamic adaptation and continuous learning capabilities.
Owner:青岛市军队离休退休干部活动中心

System for real-time analysis of emotional feedback during motivational presentations

A system for real-time analysis of emotional feedback during motivational speeches, consisting of: a series of multimodal sensors, including at least one visual sensor configured to capture facial expressions of spectators, at least one directional microphone configured to capture the audio responses of the audience, and optionally one or more physiological sensors configured to capture biometric signals from spectators; an edge-based processing unit that is communicatively coupled to the arrangement of multimodal sensors, wherein the edge-based processing unit comprises the following: (a) a feature extraction module configured to extract visual features from captured facial images, acoustic features from voice responses, and physiological features from biometric signals; (b) an emotion inference machine configured to process the features using a deep learning-based emotion recognition model comprising a convolutional neural network (CNN) for classifying facial expressions, a recurrent neural network (RNN) for classifying voice emotions, and a multimodal late fusion layer configured to compute a composite emotion state vector representing the aggregated emotions of the audience; (c) a timestamp and speech alignment module configured to correlate the calculated composite emotion state vector with segmented portions of a live motivational speech based on real-time speech-to-text transcription and semantic analysis; and (d) a session-based storage unit configured to log time-indexed emotional state vectors and corresponding speech segments for post-event analysis; A speaker feedback interface comprising a portable display or a podium-mounted visualization panel, wherein the interface is configured to display visual indicators of emotional feedback in real time, the indicators being derived from the emotional state vector and including at least emotional trend graphs, threshold alerts, or engagement indices.
Owner:1XL LLC FZ +2

Three-dimensional digital human generation method and system capable of voice interaction

The invention belongs to the technical field of three-dimensional reconstruction, and discloses a three-dimensional digital human generation method and system capable of voice interaction. According to the invention, brand new speaking audios in different languages are automatically generated according to different languages of the input target text and the sampled human voice audios; the sequential stability and detail reduction capability of three-dimensional human motion are guaranteed by using multi-model joint estimation and a sequential loss function, and facial expression details and hand postures in the image can be accurately estimated. After the high-precision three-dimensional human body model is obtained through estimation, human body action and expression generation is carried out based on voice driving, accurate synchronization of actions and expressions generated through voice is achieved, and facial expression movement and body posture movement, namely a whole-body three-dimensional human body model, conforming to brand-new speaking audio are accurately generated; and finally, rendering the whole-body three-dimensional human body model into a real digital human capable of voice interaction by using a three-dimensional neural rendering model. According to the invention, the realization of single person picture input, high-precision three-dimensional digital person generation and voice interaction is facilitated.
Owner:NANJING UNIV OF SCI & TECH

Conversation marketing strategy optimization method and system based on artificial intelligence

The invention relates to the technical field of intelligent dialogues, and discloses a dialogue marketing strategy optimization method based on artificial intelligence, and the method comprises the following steps: collecting the voice, text, facial expression and physiological signals of a user in real time through a multi-modal perception assembly of a terminal device, constructing a dynamic emotion map, and extracting a multi-modal feature vector; a cross-modal information processing module is utilized to align the multi-modal data through a comparative learning algorithm, and causal relationship description of user behaviors and strategies and strategy risk scores are generated; cooperatively training a global strategy model through differential privacy and homomorphic encryption technologies; combining the dynamic emotion map and a causal model to generate an emotion adaptive dialogue script, and optimizing strategy selection through a reinforcement learning algorithm; and dynamically updating the emotion map, the risk score and the global model through a closed-loop feedback mechanism to form a real-time optimized strategy generation system. According to the invention, the practicability of artificial intelligence to real-time services can be improved.
Owner:SHENZHEN SKYCRANE TECH CO LTD

System for delivering personalized motivational content using biometric signals

A system for the real-time delivery of personalized motivational content based on biometric information; the system includes: a biometric acquisition module configured to capture a variety of physiological signals from a user, wherein the physiological signals include at least heart rate variability, electrodermal activity, facial expressions and electroencephalographic (EEG) signals; a preprocessing module that is operationally coupled with the biometric acquisition module, wherein the preprocessing module is configured to remove noise, normalize and extract signal features from the physiological signals in real time; a multimodal biometric fusion engine configured to temporally align and synchronize the extracted features across signal modalities using dynamic time distortion and confidence-weighted interpolation; a motivational state inference model with a hybrid neural architecture comprising a Convolutional Neural Network (CNN) for spatial pattern recognition and a Recurrent Neural Network (RNN) for temporal sequence modeling, wherein the inference model is configured to output a motivational input score and an affective state classification; an engine for recommending motivational content, configured to select and prioritize content from a content repository based on motivational uptake score, user profile metadata, contextual signals including time of day and geolocation, and historical content effectiveness profiles; and a content delivery subsystem comprising one or more output modalities selected from an acoustic actuator, a visual display, a haptic actuator or an environmental controller, wherein the content delivery subsystem is capable of presenting the selected motivational content in a modality that is dynamically adapted to the user's current psychophysiological state.
Owner:1XL LLC FZ +3

Qt interactive game role action response control method and system fused with AI emotion recognition

The invention relates to the technical field of game development and artificial intelligence, and discloses a Qt interactive game role action response control method fused with AI emotion recognition, which comprises the following steps: S1, acquiring multi-modal data such as facial expression, voice intonation and limb action of a player by using sensors such as a camera and a microphone on game equipment; s2, the collected multi-modal data are transmitted to an AI emotion recognition module, a deep learning algorithm is adopted to analyze and process the data, and the current emotion state of the player is recognized; and S3, obtaining current game scene information and other operation instructions input by the player through the Qt framework. According to the Qt interactive game role action response control method and system fused with AI emotion recognition, through the AI emotion recognition technology, a game role can sense the emotion state of a player and make the emotion interaction action matched with the game role, the emotion resonance between the game role and the player is enhanced, and the interaction experience and immersion of a game are greatly improved.
Owner:张敏飞

Automated nonverbal analysis system

Examples relate to computer-implemented methods for analyzing communication in digital evaluation. A computing device accesses multimodal data comprising video and audio information of human subjects and configures a computational model using this data to identify patterns in communication that correlate with assessment metrics. The configuring implements processing techniques that preserve relationships between features across different modalities. When a video recording of a candidate is received, the computing device processes the video using the configured computational model to extract communication features. These features may include facial expressions, gestures, eye movements, posture, vocal tone, and speech patterns. The device generates an evaluation of the candidate based on the extracted communication features and outputs a representation of the evaluation.
Owner:LIGHT STEVEN PATRICK

Sign language recognition method and device based on multi-modal deep learning

The invention discloses a sign language recognition method and device based on multi-modal deep learning, and the method comprises the steps: multi-modal data input: capturing hand motions, gesture tracks and facial expressions at the same time through a camera, a motion capture sensor and other devices, and forming multi-modal data input; sign language action recognition: precise recognition of sign language actions is realized through a combined model of a deep convolutional neural network and a long-short-term memory network; facial expression and gesture track combined recognition: realizing understanding and translation of complex sign language sentences by combining the captured facial expressions and gesture tracks; and context natural language processing: generating a target statement in combination with context semantic understanding, and outputting a translation result. Through hand motion capture, facial expression analysis, gesture trajectory tracking and context natural language processing, complex sign language motions can be recognized more accurately and translated into characters or voices in real time, and low delay and high accuracy are achieved.
Owner:MIANYANG CITY UNIV

Server, display device and digital human processing method

The embodiment of the invention provides a server, display equipment and a digital human processing method. The method comprises the following steps: receiving voice data input by a user and sent by the display equipment; broadcast voice is determined based on the voice data; extracting voice features of the broadcast voice; determining mouth shape parameters based on the voice features; determining emotion parameters and acquiring user image data; generating digital human image data based on the user image data, the emotion parameters and the mouth shape parameters; and sending the broadcast voice and the digital human image data to the display device, so that the display device plays the broadcast voice and displays a digital human image based on the digital human image data. According to the embodiment of the invention, the expression parameters and the mouth shape parameters are determined according to the voice data input by the user, the expression parameters and the mouth shape parameters are combined to generate the digital human image with better facial expression expression, and emotion customization and control are realized.
Owner:HISENSE VISUAL TECH CO LTD

Self-adaptive interactive language learning system capable of multimodal emotion calculation and matching method of self-adaptive interactive language learning system

The invention discloses a multi-modal emotion calculation enabling adaptive interactive language learning system and a matching method thereof, and belongs to the technical field of artificial intelligence, and the system comprises a multi-modal emotion real-time perception and quantification module MHAE-Net, a personalized adaptive dialogue and intervention strategy module ADIS-RL, and a virtual dialogue partner interactive interface module EC-NLG. The multi-modal emotion real-time perception and quantification module MHAE-Net is composed of a voice signal acquisition and processing unit, a voice emotion analysis unit, a text emotion analysis unit, a facial expression emotion analysis unit, a physiological signal emotion analysis unit and a multi-modal emotion fusion and decision-making unit. The interactive strategy is dynamically adjusted, the oral anxiety of the language is effectively relieved, the oral confidence is improved, the system adopts the multi-mode emotion calculation technology, the accuracy and robustness of emotion recognition are improved, the targeted interactive strategy is designed for different emotion states, and personalized emotion support and language practice are provided for learners.
Owner:SHENZHEN XIXING INTELLIGENT TECHNOLOGY CO LTD

Government affair service digital human intelligent interaction method and system based on deep learning

The invention discloses a government affair service digital human intelligent interaction method and system based on deep learning, and the method comprises the steps: firstly obtaining government affair service specification policy data and interaction behavior data, including voice, text, facial expression, limb movement and the like, of a user; performing spatio-temporal feature extraction and fusion on the original interaction data to generate a user behavior semantic vector, performing feature extraction on the policy data to generate a standard vector, analyzing the user behavior semantic vector based on a government affair policy cognition map to obtain a government affair handling intention vector, and optimizing an intention generation service execution vector in combination with the standard vector; and multi-modal behavior decoding and cross-modal alignment loss function optimization are carried out to obtain an interaction behavior sequence conforming to the specification, and finally, the digital human is controlled to carry out interaction feedback with the user according to the sequence. According to the invention, pertinence and normalization of interactive feedback of the government affair service digital human and the integrating degree with the actual demand of the user can be improved.
Owner:JIANGSU LIANBANG INFORMATION TECH CO LTD

Intelligent accounting teaching method combined with behavior recognition

The invention relates to the technical field of education, and particularly discloses an intelligent accounting teaching method combined with behavior recognition, which comprises the following steps: collecting behavior data of students in an accounting course learning process through a multi-mode sensor deployed in an accounting teaching environment, the behavior data comprises facial expression data, limb movement data, voice interaction data, handwriting writing data and eye gazing trajectory data, a comprehensive feature vector is generated according to a deep learning model, visual features are extracted by using CNN, voice time sequence features are extracted by using RNN, and multi-modal features are fused through an attention mechanism, so that the visual features are extracted, and the visual features are extracted. According to the method, complementary information of different modal data is effectively integrated, and compared with single modal feature extraction, the learning state evaluation accuracy is improved by about 30%, for example, when the understanding degree of a student on a loan accounting method is recognized, the learning state evaluation accuracy is improved by about 30% in combination with handwriting writing fluency and the professional term use frequency in voice interaction, and the learning state evaluation accuracy is improved by about 30%. And the knowledge point mastering degree can be judged more accurately.
Owner:GUANGZHOU HUAXIA VOCATIONAL COLLEGE

Teaching interaction-oriented digital human three-dimensional reconstruction system

The invention provides a teaching interaction-oriented digital human three-dimensional reconstruction system, and relates to the technical field of human-computer interaction, and the system comprises a data collection module which is used for collecting a voice signal, gesture action data, eye tracking data and facial expression data of a user in real time through a multi-modal sensor array, and obtaining a teaching type identifier, an interaction stage state parameter and a participant role identifier of the current teaching scene through a preset teaching scene database. According to the invention, accurate adaptation of teaching scenes and intelligent generation and dynamic optimization of interaction strategies are realized, the accuracy and suitability of teaching interaction are improved, and personalized and high-quality teaching interaction experience is brought to users.
Owner:XIAMEN YIXUE SOFTWARE CO LTD

Customizable system for managing personalized communications using ai-generated video

Systems and methods for generating customized media content are provided. Data regarding engagement with customized media content may be tracked and used to train a neural network to generate a content generation module that optimize for increased engagement by adjusting parameters (weights). A new content generation module may be generated by the trained neural network based on a selected set of content attributes and generative artificial intelligence (AI) protocols. New customized media content may thereafter be generated by using the new video content generation module to incorporate multi-modal fusion of facial expression data into video content for the new customized media content based on the selected set of content attributes, use voice matching algorithms to generate an audio track for the video content, synchronize the audio track to the video content, and integrate one or more of the selected set of content attributes into the new customized media content.
Owner:HOOT HEALTH INC

Campus psychological assessment multi-modal emotion recognition and privacy protection method and system

The invention discloses a campus psychological assessment multi-mode emotion recognition and privacy protection method and system, and the method comprises the steps: collecting physiological signal data, voice signal data and facial expression data in a non-contact manner, carrying out the preprocessing of the collected data, extracting the feature vector of the signal data, and carrying out the recognition of the feature vector of the signal data; the method comprises the following steps of: inputting a modal-invariant basic model to carry out multi-modal fusion, then inputting the modal-invariant basic model into a lightweight multi-modal emotion recognition model to carry out emotion recognition, generating a visual report of a user emotion state and emotion intensity according to an emotion recognition result, calculating a DASS-21 index according to the visual report, generating a standard evaluation scale, and providing emotion guidance. According to the method, the physiological signals, the voice signals and the facial expression data are fused, the emotion features are captured from multiple angles, the emotion recognition accuracy is improved, the psychological state can be more comprehensively described through the multi-modal fusion mode, and the root of the psychological problem can be more accurately positioned in an auxiliary mode.
Owner:ANHUI NORMAL UNIV

Voice-driven facial expression control method and system for humanoid robot

The invention discloses a humanoid robot voice-driven facial expression control method and system. The extracted voice audio features are subjected to face corresponding key point prediction through an audio emotion key point prediction model, and the audio emotion key point prediction model is a regression model based on a long short-term memory network and is used for learning a nonlinear mapping relation between an audio feature sequence and face key point coordinates; outputting a predicted key point position difference value or a relative distance; the predicted face key points are input into a steering engine angle mapping model, and the steering engine angle mapping model maps the geometrical relationship of the face key points into angle instruction parameters for controlling a robot face micro steering engine based on a pre-trained nonlinear regression model; and according to the predicted angle instruction parameters, driving a robot face mechanism to make human expression simulating motion synchronous with the voice content. The problem that multi-channel interaction of a traditional robot is not coordinated is solved, and the naturalness and emotional expressive force of man-machine interaction are remarkably improved.
Owner:WUHAN UNIV

Multi-modal information fusion emotion detection method based on visible light and voiceprint

The invention relates to the technical field of artificial intelligence, and particularly provides a multi-modal information fusion emotion detection method based on visible light and voiceprint. The method comprises the following steps: respectively acquiring video images and environment sounds of old people in a monitoring environment through a visible light image module and an audio voiceprint acquisition module; a facial expression detection module is used for recognizing a human face in the video image, and a visual anomaly signal is output; recognizing an audio voiceprint signal in the environmental sound according to a voiceprint detection module, and outputting an audio abnormal signal; according to the method, the detection accuracy is effectively improved, the missing report is reduced, the method is suitable for home and old-age care institution environments, and the method is of great significance to guarantee the safety of old people and alleviate serious consequences caused by falling down.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)

Emotion rehabilitation intervention method based on face recognition

The invention discloses an emotion rehabilitation intervention method based on face recognition. The method comprises the steps that S1, a user face image sequence is collected through a front camera; s2, performing image preprocessing to generate a standardized image tensor set; s3, facial dynamic features are extracted through an improved BiFormer network, emotional states are recognized, and a multi-dimensional emotional state sequence and a facial expression coding sequence are generated; s4, jointly analyzing the emotion change trend and the expression behavior mode of the user through a multi-mode RETNet network, and outputting a psychological state grade and a future emotion fluctuation score sequence of the user; s5, matching an emotion rehabilitation intervention strategy through a preset rule engine; s6, executing the emotion rehabilitation intervention strategy, collecting user data in real time, and generating a user response scoring sequence; and S7, performing incremental updating on the multi-modal RETNet network and a preset rule engine based on the user response score sequence. According to the invention, the accuracy of emotional state recognition and the intelligent level of intervention strategy pushing are improved.
Owner:HEBEI SANYI INFORMATION TECHNOLOGY CO LTD

Driver fatigue detection system and method applied to intelligent driving

The invention relates to the technical field of intelligent driving, and discloses a driver fatigue detection system and method applied to intelligent driving, and the driver fatigue detection system applied to intelligent driving comprises an identity library building unit, a map route unit, a face evaluation unit, a driving evaluation unit and an intelligent driving management unit. The identity library establishing unit is used for establishing a driver identity library and collecting driver image data of a driving position and driving image data of a vehicle driving direction in real time; the map route unit is used for accessing a satellite map through the vehicle machine module; the fatigue state of the driver can be timely and accurately evaluated by monitoring the facial expression and the driving behavior of the driver in real time; when the fatigue degree is detected to be high, the system can automatically start a corresponding intelligent driving working state, for example, in a first-level intelligent driving state, a vehicle is guided to a safe position to stop, traffic accidents caused by fatigue of a driver are avoided, and driving safety is greatly guaranteed.
Owner:SCI & TECH CO LTD HEFEI INTELLIGENT VEHICLE TECH CO LTD

Multi-modal interaction method, system and device based on intelligent cabin and vehicle

The invention provides a multi-mode interaction method, system and device based on an intelligent cabin and a vehicle, and relates to the technical field of intelligent cabins, and the method comprises the steps: displaying a main interface comprising a multi-mode interaction instruction setting control, a user authority management control and a real-time feedback display area; and flexible configuration and real-time feedback of the interaction mode of the intelligent cabin are realized. According to the method, the safety and personalized setting of operation authorities of different users are ensured by utilizing cross validation binding of the biological characteristic data and the authority levels. Meanwhile, through fusion processing of multi-mode input signals such as voice, gestures, eye movement and facial expressions, the weight of each signal is dynamically adjusted according to a signal priority rule, an interaction instruction conforming to the intention of the user is generated, and the naturalness and efficiency of interaction are improved. Finally, the execution mechanism is controlled to complete the corresponding operation by performing matching verification on the interaction intention instruction and the user permission level, and the safety and accuracy of the operation are further ensured.
Owner:CHINA FAW CO LTD

Facial expression-based emotion real-time identification and long-term monitoring method

The invention relates to the technical field of computer vision and emotion calculation, in particular to an emotion real-time recognition and long-term monitoring method based on facial expressions. According to the method, an emotion recognition result is obtained by recognizing a high-definition facial image, and the emotion recognition result, environment information and physiological state data are fused to obtain time-space aligned multi-modal data; performing emotional causal analysis based on the multi-modal data, and judging emotional causes by combining a rule engine and a machine learning model: outputting a real-time emotional state recognition result and a periodic emotional report according to the emotional causes, and performing differentiated feedback according to the emotional causes. According to the method, through multi-source data fusion and a causal inference mechanism, the accuracy and interpretability of emotion recognition are effectively improved, and the technical span from passive recognition to personalized active intervention is realized.
Owner:HUAZHONG UNIV OF SCI & TECH

Counterfeit face video detection method and device based on multi-modal behavior consistency, electronic equipment, storage medium and program product

The invention provides a forged face video detection method and device based on multi-modal behavior consistency, electronic equipment, a storage medium and a program product. The method comprises the following steps: extracting voice features, facial expression features and head action features from a video signal to be detected; recognizing voice emotion, facial emotion and semantic emotion; based on the VAD value sequences of the various emotions, emotion consistency features and emotion synchronism features among the various emotions are calculated, and emotion semantic consistency features between the semantic content and the facial emotion and between the semantic content and the voice emotion are calculated; constructing a cross-modal time dependence graph to obtain interaction features; processing the voice features, the facial expression features and the head action features by using a hierarchical attention network to obtain time sequence features; a multi-dimensional fusion feature vector is formed; and processing the fusion feature vector by using a preset binary classifier to obtain a classification result indicating whether the to-be-detected video signal is a fake face video or not.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Language barrier execution type intervention effect evaluation method based on deep learning

The invention relates to the technical field of language barrier evaluation, and discloses a deep learning-based language barrier executive intervention effect evaluation method. The method comprises the following steps: acquiring real-time voice data and facial expression data of a language barrier patient in intervention training through a multi-modal data acquisition device to form an original behavior feature set; performing acoustic feature hierarchical analysis on the voice data by adopting a time sequence feature extraction network to generate a voice time sequence feature vector; performing micro-expression dynamic capture on the expression data through a three-dimensional convolutional neural network to generate an expression state feature vector; inputting the two types of vectors into a multi-modal feature fusion layer to carry out cross-modal correlation analysis, and generating a comprehensive behavior evaluation matrix; on the basis of the matrix, an intervention effect analysis model driven by an attention mechanism is adopted, a behavior improvement degree index of the current intervention stage is calculated, accurate evaluation of the intervention effect is achieved, and support is provided for dynamic adjustment of language barrier rehabilitation intervention.
Owner:SHANDONG VOCATIONAL COLLEGE OF SPECIAL EDUCATION

Facial expression implementation control method for man-machine interaction robot

The invention relates to the field of human-computer interaction, and discloses a human-computer interaction robot facial expression implementation control method, which comprises the following steps: dynamically capturing the face, sound and posture of a user in a human-computer interaction scene, and generating an initial emotion perception sequence; performing hierarchical cross mapping on the initial emotion perception sequence, and performing combinatorial analysis on facial expressions, voices and intonations and body movement features to form an emotion matching vector; based on the emotion matching vector, utilizing a priority regulation model to predict the emotion trend of the user at the next moment, and identifying an expression enhancement point and an emotion attenuation area; mapping the generated facial micro-expression adjusting instruction to a robot facial driving unit, and performing real-time correction and conflict resolution on an expression action sequence; and reversely fusing a user fixation point, facial muscle micro-motion and voice emotion which are acquired in real time into an emotion matching vector, and dynamically updating a perception weight and expression regulation and control parameters. The method has the advantage of improving the emotion matching degree of the robot expressions.
Owner:BEIJING HAIBAICHUAN TECH CO LTD

Portrait image reconstruction method, portrait video generation method and electronic equipment

The invention discloses a portrait image reconstruction method, a portrait video generation method and electronic equipment, and relates to the technical field of virtual digital faces, the method comprises the following steps: obtaining a portrait source image, a portrait target image and at least one portrait reference image, each portrait reference image having corresponding speaker reference information; extracting a first target portrait motion feature corresponding to the portrait target image, and performing deformation alignment on the portrait source image and each portrait reference image according to the first target portrait motion feature to obtain a corresponding source image deformation texture feature and each reference image deformation texture feature; and reconstructing a portrait target image according to the source image deformation texture features and the reference image deformation texture features. Therefore, the reference information of different visual angles, head postures and facial expressions is integrated, the alignment texture coverage of the source image is optimized, and the naturalness and vividness of virtual character generation are improved.
Owner:AISPEECH CO LTD

Virtual image generation method based on machine learning

The invention relates to the technical field of virtual image generation, and discloses a virtual image generation method based on machine learning, and the method comprises the steps: obtaining a reference image and text description information, and synthesizing the reference image and the text description information into a data set, thereby obtaining a feature set; according to the feature set, analyzing dominant hue distribution features and judging a style tendency, calculating texture complexity according to a judgment result, and analyzing a cartoonalization degree to obtain a cartoonalization degree quantitative index; and when the quantitative index is greater than a threshold value, setting a character proportion exaggeration coefficient as a high-magnification amplification parameter, and if the quantitative index is less than the threshold value, setting the character proportion exaggeration coefficient as a standard proportion parameter, performing real-time adjustment on the facial expression within a preset dynamic expression amplitude range according to the facial control point coordinates, and generating a control instruction. According to the expression change control instruction, synchronously calculating to obtain a coordination parameter; and performing dynamic rendering adjustment on the clothes according to the coordination parameters to obtain virtual image generation data. The method can solve the problem that the generation result deviates from the target style.
Owner:HANGZHOU ZDJOYS TECH CO LTD

Live broadcast method and system based on intelligent digital human model, and medium

The invention relates to the technical field of live broadcast display, in particular to a live broadcast method and system based on an intelligent digital human model, and a medium. The method comprises the following steps: acquiring user live broadcast environment data; facial expression data, voice data and limb movement data are collected through a multi-modal sensor, and feature fusion processing is carried out to obtain user dynamic portrait data; generating a digital human model in real time based on the dynamic portrait data of the user, and constructing an intelligent driving engine comprising an emotion response module and a semantic understanding module for the digital human model; generating a personalized digital human image and a digital human behavior decision tree based on an intelligent driving engine; analyzing the live broadcast interaction data stream in real time; through multi-modal perception, an intelligent driving engine and a dynamic interaction strategy, a personalized, high-interactivity and self-adaptive flow pushing intelligent digital human live broadcast system is realized, and the technical bottlenecks of existing virtual live broadcast in the aspects of interactivity, personalization, scene fusion and the like are broken through.
Owner:HUNAN TESCO E-COMMERCE CO LTD

Wearable facial movement tracking devices

This technology provides systems and methods for tracking facial movements and reconstructing facial expressions by learning skin deformation patterns and facial features. Frontal view images of a user making a variety of facial expressions are acquired to create a data training set for use in a machine-learning process. Head-mounted or neck-mounted wearable devices are equipped with one or more camera(s) or acoustic device(s) in communication with a data processing system. The cameras capture images of contours of the users face from either the cheekbone or the chin profile of the user. The acoustic devices transmit and receive signals to calculate a representation of the skin deformation. A data processing system uses the images, the profile of the contours, or skin deformation to track facial movement or to reconstruct facial expressions of the user based on the data training set.
Owner:CORNELL UNIVERSITY