Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

635 results about "Facial affect" patented technology

Multi-sensory autonomous multimodal emotion-synchronized environmental control architecture and regulation system (amesecar)

An autonomous environmental regulation and behavioral monitoring system is disclosed, configured to adapt temperature, lighting, and acoustic conditions based on real-time emotional and physiological data. The system includes a dual-redundant central processor, hierarchical communication networks, multi-angle visual acquisition units, infrared thermometers, and modular environmental subsystems. It detects posture, gestures, facial expressions, and thermal signals to classify user states and apply individualized airflow, light, and sound modulation without relying on external internet connectivity. The system also monitors connected appliances using voltage-based pressure analysis to forecast device degradation. With integrated gesture recognition, privacy-preserving data handling, and predictive adaptation, the invention enables multi-user personalization, long-term learning, and uninterrupted operation within residential, administrative, or healthcare infrastructures.
Owner:SEYEDKHAMOUSHI FAEZEHALSADAT +1

System for real-time analysis of emotional feedback during motivational presentations

A system for real-time analysis of emotional feedback during motivational speeches, consisting of: a series of multimodal sensors, including at least one visual sensor configured to capture facial expressions of spectators, at least one directional microphone configured to capture the audio responses of the audience, and optionally one or more physiological sensors configured to capture biometric signals from spectators; an edge-based processing unit that is communicatively coupled to the arrangement of multimodal sensors, wherein the edge-based processing unit comprises the following: (a) a feature extraction module configured to extract visual features from captured facial images, acoustic features from voice responses, and physiological features from biometric signals; (b) an emotion inference machine configured to process the features using a deep learning-based emotion recognition model comprising a convolutional neural network (CNN) for classifying facial expressions, a recurrent neural network (RNN) for classifying voice emotions, and a multimodal late fusion layer configured to compute a composite emotion state vector representing the aggregated emotions of the audience; (c) a timestamp and speech alignment module configured to correlate the calculated composite emotion state vector with segmented portions of a live motivational speech based on real-time speech-to-text transcription and semantic analysis; and (d) a session-based storage unit configured to log time-indexed emotional state vectors and corresponding speech segments for post-event analysis; A speaker feedback interface comprising a portable display or a podium-mounted visualization panel, wherein the interface is configured to display visual indicators of emotional feedback in real time, the indicators being derived from the emotional state vector and including at least emotional trend graphs, threshold alerts, or engagement indices.
Owner:1XL LLC FZ +2

Three-dimensional digital human generation method and system capable of voice interaction

The invention belongs to the technical field of three-dimensional reconstruction, and discloses a three-dimensional digital human generation method and system capable of voice interaction. According to the invention, brand new speaking audios in different languages are automatically generated according to different languages of the input target text and the sampled human voice audios; the sequential stability and detail reduction capability of three-dimensional human motion are guaranteed by using multi-model joint estimation and a sequential loss function, and facial expression details and hand postures in the image can be accurately estimated. After the high-precision three-dimensional human body model is obtained through estimation, human body action and expression generation is carried out based on voice driving, accurate synchronization of actions and expressions generated through voice is achieved, and facial expression movement and body posture movement, namely a whole-body three-dimensional human body model, conforming to brand-new speaking audio are accurately generated; and finally, rendering the whole-body three-dimensional human body model into a real digital human capable of voice interaction by using a three-dimensional neural rendering model. According to the invention, the realization of single person picture input, high-precision three-dimensional digital person generation and voice interaction is facilitated.
Owner:NANJING UNIV OF SCI & TECH

System for delivering personalized motivational content using biometric signals

A system for the real-time delivery of personalized motivational content based on biometric information; the system includes: a biometric acquisition module configured to capture a variety of physiological signals from a user, wherein the physiological signals include at least heart rate variability, electrodermal activity, facial expressions and electroencephalographic (EEG) signals; a preprocessing module that is operationally coupled with the biometric acquisition module, wherein the preprocessing module is configured to remove noise, normalize and extract signal features from the physiological signals in real time; a multimodal biometric fusion engine configured to temporally align and synchronize the extracted features across signal modalities using dynamic time distortion and confidence-weighted interpolation; a motivational state inference model with a hybrid neural architecture comprising a Convolutional Neural Network (CNN) for spatial pattern recognition and a Recurrent Neural Network (RNN) for temporal sequence modeling, wherein the inference model is configured to output a motivational input score and an affective state classification; an engine for recommending motivational content, configured to select and prioritize content from a content repository based on motivational uptake score, user profile metadata, contextual signals including time of day and geolocation, and historical content effectiveness profiles; and a content delivery subsystem comprising one or more output modalities selected from an acoustic actuator, a visual display, a haptic actuator or an environmental controller, wherein the content delivery subsystem is capable of presenting the selected motivational content in a modality that is dynamically adapted to the user's current psychophysiological state.
Owner:1XL LLC FZ +3

Server, display device and digital human processing method

The embodiment of the invention provides a server, display equipment and a digital human processing method. The method comprises the following steps: receiving voice data input by a user and sent by the display equipment; broadcast voice is determined based on the voice data; extracting voice features of the broadcast voice; determining mouth shape parameters based on the voice features; determining emotion parameters and acquiring user image data; generating digital human image data based on the user image data, the emotion parameters and the mouth shape parameters; and sending the broadcast voice and the digital human image data to the display device, so that the display device plays the broadcast voice and displays a digital human image based on the digital human image data. According to the embodiment of the invention, the expression parameters and the mouth shape parameters are determined according to the voice data input by the user, the expression parameters and the mouth shape parameters are combined to generate the digital human image with better facial expression expression, and emotion customization and control are realized.
Owner:HISENSE VISUAL TECH CO LTD

Campus psychological assessment multi-modal emotion recognition and privacy protection method and system

The invention discloses a campus psychological assessment multi-mode emotion recognition and privacy protection method and system, and the method comprises the steps: collecting physiological signal data, voice signal data and facial expression data in a non-contact manner, carrying out the preprocessing of the collected data, extracting the feature vector of the signal data, and carrying out the recognition of the feature vector of the signal data; the method comprises the following steps of: inputting a modal-invariant basic model to carry out multi-modal fusion, then inputting the modal-invariant basic model into a lightweight multi-modal emotion recognition model to carry out emotion recognition, generating a visual report of a user emotion state and emotion intensity according to an emotion recognition result, calculating a DASS-21 index according to the visual report, generating a standard evaluation scale, and providing emotion guidance. According to the method, the physiological signals, the voice signals and the facial expression data are fused, the emotion features are captured from multiple angles, the emotion recognition accuracy is improved, the psychological state can be more comprehensively described through the multi-modal fusion mode, and the root of the psychological problem can be more accurately positioned in an auxiliary mode.
Owner:ANHUI NORMAL UNIV

Voice-driven facial expression control method and system for humanoid robot

The invention discloses a humanoid robot voice-driven facial expression control method and system. The extracted voice audio features are subjected to face corresponding key point prediction through an audio emotion key point prediction model, and the audio emotion key point prediction model is a regression model based on a long short-term memory network and is used for learning a nonlinear mapping relation between an audio feature sequence and face key point coordinates; outputting a predicted key point position difference value or a relative distance; the predicted face key points are input into a steering engine angle mapping model, and the steering engine angle mapping model maps the geometrical relationship of the face key points into angle instruction parameters for controlling a robot face micro steering engine based on a pre-trained nonlinear regression model; and according to the predicted angle instruction parameters, driving a robot face mechanism to make human expression simulating motion synchronous with the voice content. The problem that multi-channel interaction of a traditional robot is not coordinated is solved, and the naturalness and emotional expressive force of man-machine interaction are remarkably improved.
Owner:WUHAN UNIV

Multi-modal information fusion emotion detection method based on visible light and voiceprint

The invention relates to the technical field of artificial intelligence, and particularly provides a multi-modal information fusion emotion detection method based on visible light and voiceprint. The method comprises the following steps: respectively acquiring video images and environment sounds of old people in a monitoring environment through a visible light image module and an audio voiceprint acquisition module; a facial expression detection module is used for recognizing a human face in the video image, and a visual anomaly signal is output; recognizing an audio voiceprint signal in the environmental sound according to a voiceprint detection module, and outputting an audio abnormal signal; according to the method, the detection accuracy is effectively improved, the missing report is reduced, the method is suitable for home and old-age care institution environments, and the method is of great significance to guarantee the safety of old people and alleviate serious consequences caused by falling down.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)

Emotion rehabilitation intervention method based on face recognition

The invention discloses an emotion rehabilitation intervention method based on face recognition. The method comprises the steps that S1, a user face image sequence is collected through a front camera; s2, performing image preprocessing to generate a standardized image tensor set; s3, facial dynamic features are extracted through an improved BiFormer network, emotional states are recognized, and a multi-dimensional emotional state sequence and a facial expression coding sequence are generated; s4, jointly analyzing the emotion change trend and the expression behavior mode of the user through a multi-mode RETNet network, and outputting a psychological state grade and a future emotion fluctuation score sequence of the user; s5, matching an emotion rehabilitation intervention strategy through a preset rule engine; s6, executing the emotion rehabilitation intervention strategy, collecting user data in real time, and generating a user response scoring sequence; and S7, performing incremental updating on the multi-modal RETNet network and a preset rule engine based on the user response score sequence. According to the invention, the accuracy of emotional state recognition and the intelligent level of intervention strategy pushing are improved.
Owner:HEBEI SANYI INFORMATION TECHNOLOGY CO LTD

Facial expression-based emotion real-time identification and long-term monitoring method

The invention relates to the technical field of computer vision and emotion calculation, in particular to an emotion real-time recognition and long-term monitoring method based on facial expressions. According to the method, an emotion recognition result is obtained by recognizing a high-definition facial image, and the emotion recognition result, environment information and physiological state data are fused to obtain time-space aligned multi-modal data; performing emotional causal analysis based on the multi-modal data, and judging emotional causes by combining a rule engine and a machine learning model: outputting a real-time emotional state recognition result and a periodic emotional report according to the emotional causes, and performing differentiated feedback according to the emotional causes. According to the method, through multi-source data fusion and a causal inference mechanism, the accuracy and interpretability of emotion recognition are effectively improved, and the technical span from passive recognition to personalized active intervention is realized.
Owner:HUAZHONG UNIV OF SCI & TECH

Language barrier execution type intervention effect evaluation method based on deep learning

The invention relates to the technical field of language barrier evaluation, and discloses a deep learning-based language barrier executive intervention effect evaluation method. The method comprises the following steps: acquiring real-time voice data and facial expression data of a language barrier patient in intervention training through a multi-modal data acquisition device to form an original behavior feature set; performing acoustic feature hierarchical analysis on the voice data by adopting a time sequence feature extraction network to generate a voice time sequence feature vector; performing micro-expression dynamic capture on the expression data through a three-dimensional convolutional neural network to generate an expression state feature vector; inputting the two types of vectors into a multi-modal feature fusion layer to carry out cross-modal correlation analysis, and generating a comprehensive behavior evaluation matrix; on the basis of the matrix, an intervention effect analysis model driven by an attention mechanism is adopted, a behavior improvement degree index of the current intervention stage is calculated, accurate evaluation of the intervention effect is achieved, and support is provided for dynamic adjustment of language barrier rehabilitation intervention.
Owner:SHANDONG VOCATIONAL COLLEGE OF SPECIAL EDUCATION

Facial expression implementation control method for man-machine interaction robot

The invention relates to the field of human-computer interaction, and discloses a human-computer interaction robot facial expression implementation control method, which comprises the following steps: dynamically capturing the face, sound and posture of a user in a human-computer interaction scene, and generating an initial emotion perception sequence; performing hierarchical cross mapping on the initial emotion perception sequence, and performing combinatorial analysis on facial expressions, voices and intonations and body movement features to form an emotion matching vector; based on the emotion matching vector, utilizing a priority regulation model to predict the emotion trend of the user at the next moment, and identifying an expression enhancement point and an emotion attenuation area; mapping the generated facial micro-expression adjusting instruction to a robot facial driving unit, and performing real-time correction and conflict resolution on an expression action sequence; and reversely fusing a user fixation point, facial muscle micro-motion and voice emotion which are acquired in real time into an emotion matching vector, and dynamically updating a perception weight and expression regulation and control parameters. The method has the advantage of improving the emotion matching degree of the robot expressions.
Owner:BEIJING HAIBAICHUAN TECH CO LTD

Multi-modal emotion recognition method and system based on knowledge distillation

The invention provides a multi-modal emotion recognition method and system based on knowledge distillation. The method comprises a pre-training stage, a knowledge distillation stage, a fine adjustment stage and a prediction stage. The pre-training stage comprises the following steps: independently training a preset neural network by using electroencephalogram, electrocardiogram and facial expression data to obtain three independent teacher models; the knowledge distillation stage comprises the step of migrating the representation learned by the teacher model to the student model through a knowledge distillation technology; the fine tuning stage comprises the step of carrying out joint optimization on the student model by utilizing a small amount of labeled multi-modal data; the prediction stage comprises the steps of dynamically fusing the multi-modal features by adopting an attention mechanism, and inputting the fused features into a classifier to obtain an emotion recognition result. The method is oriented to three modes of electroencephalogram, electrocardio and facial expression, and the characterization ability of the model to complex emotion and state change is improved; and a knowledge distillation mechanism is introduced, so that the complexity of the model is reduced, and the deployment feasibility and the operation efficiency of the model in practical application are improved.
Owner:ZHENGZHOU UNIV

Student psychological risk perception method based on multiple modes

The invention discloses a student psychological risk perception method based on multiple modes, and relates to the technical field of emotion calculation and intelligent education. The method comprises the following steps: firstly, extracting a facial expression feature vector and a voice intonation feature vector respectively by using a convolutional neural network and Fourier transform through a collected video stream and an audio stream; then adaptive denoising processing is carried out on environmental interference, timestamp alignment and dynamic time warping are carried out on the denoised multi-modal data, time sequence synchronization is ensured, and corrected multi-modal sequence data are formed; then, dynamic emotion track features are extracted from the sequence data, a preliminary emotion state label is generated by comparing the dynamic emotion track features with a baseline threshold value, and the threshold value is adaptively updated in combination with historical data so as to improve the judgment accuracy; and finally, aggregating the emotional state labels of a plurality of students to generate a visual group emotional thermodynamic diagram so as to realize macroscopic perception of group psychological risks. The accuracy, robustness and visualization degree of student psychological state analysis are effectively improved, and an efficient technical means is provided for campus psychological early warning.
Owner:景安大数据科技有限公司

Infant state identification method and system based on Chinese medicine five-tone monitoring analysis

The invention discloses an infant state recognition method and system based on traditional Chinese medicine five-tone monitoring analysis, and the method comprises the steps: synchronously collecting the crying sound, physiological signals and behavior videos of an infant, extracting the Mel-frequency cepstral coefficient of the audio, the heart rate variability, galvanic skin response and respiratory rate of the physiological signals, and the facial expression and limb movement features, and carrying out the recognition of the state of the infant through the Mel-frequency cepstral coefficient of the audio, the heart rate variability, galvanic skin response and respiratory rate. Performing structured integration by using a multi-modal feature fusion model; in combination with the five-tone theory of traditional Chinese medicine, a corresponding relation between audio features and five-organ states is established, and the robustness of five-tone and five-organ mapping is improved through fuzzy reasoning and a Bayesian mechanism; the system dynamically adjusts the weight coefficient of each mode, adapts to the individual and emotion historical trend, achieves more accurate emotion recognition and physiological evaluation, achieves cross-mode and multi-layer information fusion, improves the accuracy and interpretation ability of infant emotion and five-internal-organ state recognition, and provides a scientific basis for clinical evaluation and health management.
Owner:DONGGUAN BINHAI BAY CENT HOSPITAL

Interaction method and device, electronic equipment and readable storage medium

The invention discloses an interaction method and device, electronic equipment and a readable storage medium, and belongs to the technical field of electronic equipment, and the method comprises the steps: obtaining a video stream of a user; determining a behavior state of the user according to the video stream; wherein the behavior state comprises at least one of the following items: an emotional state, an attention state, a facial expression, a body posture and an action; and outputting first interaction content according to the behavior state of the user.
Owner:VIVO MOBILE COMM CO LTD

Full-process closed-loop nursing assistant robot control system based on multi-modal interaction

PendingCN120985685AProgramme-controlled manipulatorNursing aidClosed loop
The invention discloses a full-process closed-loop nursing assistant robot control system based on multi-modal interaction. The system comprises a multi-modal sensing module, a facial expression analysis module, an action evaluation module, a closed-loop control module, an execution driving module and a data interaction module. The multi-modal sensing module integrates multi-source information to realize accurate feature fusion; the facial expression analysis module generates expression feature vectors; the action evaluation module calculates action normative parameters; the closed-loop control module generates a dynamic task queue and distributes instructions; the execution driving module responds to the instruction to complete operation; and the data interaction module realizes data transmission and storage. The system overcomes the defects that in the prior art, multi-modal information fusion is insufficient, and control closed-loop real-time performance and adaptability are poor, nursing accuracy and efficiency are improved, accurate and personalized nursing requirements are met, manual nursing defects are made up, and operation normalization is unified.
Owner:SHENZHEN NANSHAN DISTRICT PEOPLES HOSPITAL

Multi-modal emotion recognition method based on electroencephalogram signal and facial expression fusion

The invention belongs to the field of biomedical engineering, particularly relates to a multi-modal emotion recognition method based on electroencephalogram signal and facial expression fusion, and aims to improve the accuracy of emotion recognition. The method comprises the steps that electroencephalogram signals and a face video of a subject are collected, frequency band extraction, time window division and standardization processing and key frame extraction and face embedding coding are conducted respectively, time synchronization is achieved through data alignment, electroencephalogram time sequence fragments are input into a preset coding network to extract electroencephalogram time sequence features, and the electroencephalogram time sequence features are obtained. Inputting the facial spatial feature matrix into a preset convolution and time sequence fusion network to extract facial spatial and temporal features, then realizing dynamic interaction between modals through a cross attention mechanism to obtain preliminary fusion features, obtaining modal confidence degrees corresponding to the modals through a preset regression network, weighting the electroencephalogram feature matrix and the facial expression feature matrix, and obtaining an electroencephalogram feature matrix and a facial expression feature matrix; and outputting a final fusion feature matrix. And inputting the final fusion feature matrix into a preset classifier for emotion category prediction, and outputting an emotion recognition result.
Owner:TAIYUAN UNIVERSITY OF TECHNOLOGY

System and method for generating a real-time, interactive companion on a user device

The present invention relates to a system and method for generating a real-time, interactive companion on a user device. The system comprises one or more processors and a memory. The memory stores executable instructions that, when executed by the one or more processors, cause the system to receive at least one user input comprising at least one of text, voice, touch, or gesture data. The processor extracts semantic and emotional context from the user input using a natural-language processing module configured to generate embeddings representing at least one of linguistic content, conversational intent, or affective cues and processes the embeddings, via a motion mapping and emotional mapping module to generate animation parameters defining at least one of facial expressions, gestures, and full-body motion. Based on the animation parameter, a rendering engine renders an emotionally guided digital companion in temporal synchronization with the received user input. The motion synthesis and rendering inferences are performed on the user device without transmitting raw sensor inputs or generated animation parameters off-device. This on-device execution ensures low-latency, enhanced privacy, and reduced reliance on cloud infrastructure.
Owner:ANIMATION INC

Children pain intelligent evaluation system based on AI facial expression analysis

The invention relates to the field of medical auxiliary diagnosis, in particular to an intelligent child pain assessment system based on AI facial expression analysis, a facial detection module obtains a child facial image, and performs three-stage cascade detection by using a multi-task cascade convolutional neural network model to obtain facial key point pixel coordinate information; the key point positioning module carries out noise filtering processing on the coordinate information to obtain more accurate face key point coordinate information, the information is transmitted to the key point thermodynamic diagram representation module, and the key point thermodynamic diagram representation module generates a face key point thermodynamic diagram representing the degree of pains of children in combination with the face key point coordinate information and a thermodynamic diagram mechanism. The child pain intelligent evaluation module evaluates the degree of pain of the child based on the thermodynamic diagrams and obtains an evaluation result, and finally, the pain evaluation result visualization module visualizes the evaluation result, and the accuracy of pain characterization is improved through high-precision face detection and key point thermodynamic diagram technology, so that the accuracy of pain characterization is improved. And an effective method is provided for evaluating the pains of children.
Owner:AFFILIATED CHILDRENS HOSPITAL OF CAPITAL INST OF PEDIATRICS

Digital human expression generation method based on multi-modal feature fusion and emotion enhancement

The invention discloses a digital human expression generation method based on multi-modal feature fusion and emotion enhancement, and belongs to the technical field of generative artificial intelligence. Comprising the steps of extracting voice semantic features and voice emotion features based on driving voice, extracting text emotion features based on an emotion text, and extracting visual latent variables and image semantic features based on a reference image; fusing the voice emotion features and the text emotion features through a perception resampling mechanism to generate fused emotion control features; fusing emotion control features, voice semantic features, visual latent variables and image semantic features as conditions, inputting the conditions into a DiT video generation model, generating a denoised video potential vector, mapping the denoised video potential vector into a facial expression potential vector by a Transform adapter, and restoring the facial expression potential vector into a facial expression parameter sequence by a FaceVese decoder to drive a digital human. According to the invention, end-to-end generation from voice and text to high-fidelity and emotion-controllable facial expression parameters can be realized.
Owner:ZHEJIANG UNIV

Multi-modal fine-grained emotion recognition method oriented to human-computer interaction and based on large model

According to the man-machine interaction-oriented multi-modal fine-grained emotion recognition method based on the large model provided by the invention, cross-modal alignment from coarse granularity to fine granularity is realized through an attention pairing interaction module (APIM) on the basis of an aspect-driven vision-text alignment and fusion network (AVTAF); emotion-related visual features (such as facial expressions and gestures) in a robot scene can be accurately captured, and environmental noise is inhibited; meanwhile, the RD-GAT is enhanced, and the reasoning ability of a large model on multi-modal emotion semantics is improved by integrating external emotion knowledge (such as SenticNet). The technology provides a new normal form for intelligent upgrading of robot emotion interaction and multi-modal understanding of a large model, and is expected to promote breakthrough application in the fields of family service robots, medical accompanying assistants, multi-modal content generation and the like.
Owner:BEIJING INST OF TECH

Emotion analysis method based on multi-modal large model

The invention discloses an emotion analysis method based on a multi-modal large model. The method comprises the following steps: S1, extracting multi-modal emotion features; respectively designing special emotional feature extractors for three modes of facial expression, voice and text; s2, carrying out cross-modal emotion alignment and fusion; mapping the emotion features of different modes to a unified emotion semantic space; s3, an emotion inconsistency detection mechanism; the method is specially used for detecting the emotion inconsistency phenomenon between different modes. S4, fine-grained sentiment classification is carried out; the sentiment classifier comprises three sub-tasks of basic sentiment classification, complex sentiment recognition and sentiment intensity regression; s5, a model fine tuning training strategy; and optimizing the performance of the model by adopting a multi-stage fine-tuning training strategy. According to the method, deep fusion and accurate analysis of facial expressions, voice acoustic features and text semantic information are realized, so that a complex emotional state and an emotional inconsistency phenomenon are effectively recognized.
Owner:JIANGSU HOPERUN SOFTWARE CO LTD

Cognitive load assessment method based on fusion of behavior characteristics and heart rate variability in classroom video

The invention discloses a cognitive load assessment method based on fusion of behavior characteristics and heart rate variability in a classroom video. According to the method, face and behavior videos of students in a real classroom are collected, behavior characteristics such as eye movement tracks, sitting postures and facial expressions of the students are extracted through a computer vision algorithm, heart rate variability indexes are predicted in combination with a remote photoplethysmography technology, and a time sequence characteristic sequence is formed. And then, a time sequence deep learning model is adopted to carry out joint modeling on the multi-modal time sequence characteristics, and the cognitive load scale level is taken as a supervision signal to construct a classification model to realize cognitive load level prediction. The method has the advantages of non-contact, automation, high adaptability and the like, and can be applied to personalized teaching monitoring and intelligent teaching feedback.
Owner:SHAANXI NORMAL UNIV

Optimization algorithm model and method based on elder care emotion accompanying

The invention discloses an adjusting and optimizing algorithm model based on elder care emotion accompanying. The model comprises a multi-mode emotion perception module for receiving voice, images and physiological signals and extracting features such as timbre, facial expression and pulse to generate emotion feature vectors; the emotion recognition module is used for predicting the real-time emotion state of the old people through fusion of CNN and LSTM and multi-modal features; the strategy adjusting and optimizing module adopts a Bayesian optimization or genetic algorithm to automatically adjust and optimize the voice intonation, the interaction rhythm and the dialogue strategy of the accompanying system; the feedback learning module iteratively updates the model based on the old people feedback data, and optimizes the interaction effect; and the personalized accompanying generation module generates accompanying strategies such as voice consolation and music recommendation according to the tuning result. The multi-mode emotion perception module is composed of a voice monitoring sub-module, an expression monitoring sub-module and a physiological signal monitoring sub-module. The strategy tuning module comprises a self-adaptive parameter search sub-module and a multi-target optimization sub-module; and the feedback learning module comprises a user feedback collection and reinforcement learning sub-module, so that the interaction experience and the emotion adaptability of the accompanying system are improved.
Owner:闫中举

Student state visual analysis system for optimizing virtual human teaching

The invention discloses a student state visual analysis system for optimizing virtual human teaching. The student state visual analysis system comprises a front-end sensing layer, a multi-modal feature extraction module, a time sequence analysis layer and a teaching decision visualization layer, the front-end sensing layer collects facial expressions and eye movement data of students in real time, and locates facial areas of the students in real time by adopting a parameter-optimized SSD target detection algorithm; the multi-modal feature extraction module extracts seven types of expressions and four types of eye movement features based on a RepVGG model through a multi-modal fusion method, and constructs a 11-dimensional time series data stream; the time sequence analysis layer dynamically models a cognitive state based on an LSTM network model, and outputs the cognitive state through a continuous behavior rule; the teaching decision visualization layer generates an interactive instrument board, realizes accurate mapping of a cognitive state and a course time axis through a thermodynamic diagram and a trend curve, and automatically marks a high-doubt section. The method breaks through the limitation of a single mode, improves the feedback efficiency, and is suitable for a virtual human asynchronous teaching scene.
Owner:HUAIYIN INSTITUTE OF TECHNOLOGY

Public devices connected to augmented reality glasses having gaze and secondary input sensors

An augmented reality headset, with an eye tracking device, and a secondary user input device, such as, a brain wave sensor detecting the user's brain waves, a microphone detecting user sounds, or finger recognition sensor, or facial recognition sensor detecting facial gestures. When entering an elevator, the user views elevator input icons, displayed in mid-air. The user gazes at, an elevator floor icon, a cursor follows the gaze to the icon. The user thinks click, or says enter, or finger gestures, or facial gestures, which activates the icon. The activated floor icon, directs the elevator, to move to the floor. The input icons can be associated with activating, an internet web page, or a public device, wirelessly connected to the glasses, like, a multi-user door opener. The touch-free operation, enables the user to avoid bacteria, that may be on physical touch input buttons, of the public door opener.
Owner:CLEMENTS SIGMUND LINDSAY

Multi-modal digital art content generation and personalized recommendation system and method

According to the multi-modal digital art content generation and personalized recommendation system and method, a sequential dynamic knowledge graph is constructed, a pattern, color and narrative structure triple is extracted from an art database, a natural language processing technology is adopted to mine implicit relations, and the artistic content is recommended to the personalized recommendation system and method. Multi-modal feature fusion is realized by utilizing visual and text feature extraction and a gating attention mechanism, a style template is automatically loaded according to equipment performance, rendering parameters are dynamically adjusted, and a recommendation strategy is optimized through facial expression analysis, gaze point distribution calculation and emotional state inference. Rendering detail levels are adjusted based on data acquired by a touch screen sensor and a motion sensor, design contents conforming to culture specifications are generated by combining a conditional GAN correction network, real-time updating of culture symbols and related rules is supported, and the system has strong intelligent processing capability and is suitable for popularization and application. The efficient generation and personalized recommendation requirements of the digital art content in multiple scenes can be met, and the user experience and the interaction effect are remarkably improved.
Owner:南昌理工学院

Automatic rigging with 2d supervised learning

PendingUS20260080601A1AnimationMedicineAnimation
According to one aspect of the present disclosure, a method of training a deformation prediction model is provided. In some implementations, a method includes obtaining a neutral expression three-dimensional (3D) mesh and a set of facial action coding system (FACS) weights, wherein the set of FACS weights represent a target facial pose or a target facial expression. The method further includes obtaining a predicted 3D mesh from the deformation prediction model, wherein the predicted mesh is arranged to at least partially mimic the target facial pose or target facial expression, rendering a two-dimensional (2D) image from the predicted mesh, and adjusting the deformation prediction model based on one or more 2D loss functions, the one or more 2D loss functions being based on comparison of the 2D image with a groundtruth 2D image obtained from a pre-trained 2D animation model.
Owner:ROBLOX CORP

Verification method and device based on facial expression recognition and storage medium

The invention discloses a verification method and device based on facial expression recognition and a storage medium. Relates to the field of artificial intelligence, and the method comprises the steps: obtaining a face image of a target customer in a target transaction process of the target customer under the condition of obtaining the authorization of the target customer; extracting facial features from the facial image, inputting the facial features into the expression recognition model, and outputting an emotion category corresponding to the facial expression on the facial image; calculating a transaction risk value corresponding to the facial expression based on the emotion category corresponding to the facial expression; calculating a transaction risk value corresponding to the target transaction based on the transaction information of the target transaction; and determining a target verification mode according to the transaction risk value corresponding to the facial expression and the transaction risk value corresponding to the target transaction, and executing the target verification mode on the target customer. Through application of the method and the device, the problem of relatively low verification accuracy caused by relatively single verification mode for the client identity in the transaction process in related technologies is solved.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA