Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2728results about "Acquiring/recognising facial features" patented technology

Driving state monitoring and feedback method and system based on multi-modal human factors intelligent data analysis, and edge computing terminal device

A driving state monitoring and feedback method and system based on multi-modal human factors intelligent data analysis, and an edge computing terminal device. The method comprises: receiving multi-modal human factors data of a tested driver that is collected in real time (S110); pre-processing the multi-modal human factors data, wherein the pre-processing comprises denoising processing and data normalization processing (S120); sending the pre-processed multi-modal human factors data into a pre-trained first state recognition model, so as to obtain a real-time recognized driver state, wherein driver states include a normal state and abnormal states, and the types of the abnormal states include a plurality of states such as a fatigue state, a distracted state and an angry state (S130); and when it is recognized that the driver state is an abnormal state, generating, for different categories of abnormal states, driving state feedback instructions to a driving intervention system, such that the driving intervention system performs state feedback adjustments on the driver on the basis of the received driving state feedback instructions (S140). The method and system can recognize different driving states of a driver in real time and then perform processing on the basis of different driving states, thereby avoiding the occurrence of traffic accidents.
Owner:KINGFAR INTERNATIONAL INC

Digital human interaction control method and device fusing emotional semantics and logical reasoning and storage medium

The invention provides a digital human interaction control method and device fusing emotion semantics and logical reasoning and a storage medium. The method comprises the following steps: analyzing multi-modal input data of a user, constructing emotion-semantics joint representation, and generating a logic decision path; and through a cognitive fusion module, emotion-semantic representation and a logic decision path are fused, and an interaction response adapting to emotion and logic consistency is generated. The system optimizes an emotion semantic model and a logical reasoning rule on line according to user feedback and interaction history, and real-time interaction of emotion dynamic and logical rules is achieved. The system can dynamically adjust the logic decision path based on the multi-mode emotional state of the user, and improves the naturalness and situation adaptability of interaction. A dynamic time warping algorithm and a factorization machine are introduced to process a multi-modal feature fusion problem, and the accuracy and robustness of emotional state recognition are improved. The online optimization mechanism enables the model and the rule to be evolved continuously, and reasoning errors are corrected automatically through user feedback, so that error circulation is avoided.
Owner:HANGZHOU DIGITAL SPACE TECHNOLOGY CO LTD

Robot behavior mode dynamic adjustment method based on multi-mode perception

The invention relates to the technical field of man-machine interaction, in particular to a robot behavior mode dynamic adjustment method based on multi-mode perception, which comprises the following steps: S1, collecting a visual image, a voice signal and an environment parameter of a target scene in real time; s2, extracting each modal feature; s3, dynamically allocating weights, and generating a fusion feature vector; s4, identifying a current scene type and a user attribute; s5, matching a corresponding interaction mode from a preset strategy library based on the scene type and the user attribute identified in the step S4; and S6, executing the matched interaction mode in the S5, and updating the weight distribution rule according to the feedback data. According to the method, vision, voice and environment characteristics are fused through a multi-mode perception technology, the behavior mode of the robot is dynamically adjusted based on a self-adaptive weighted fusion algorithm, and the interaction strategy is optimized in combination with user feedback, so that the interaction accuracy and the intelligent level of the robot under different scenes and user attributes are improved.
Owner:SHENZHEN WEILIAN ELEPHANT TECH CO LTD

Using continuous gestures for selectively processing facial movements

Systems, methods, and computer program products are disclosed for selectively employing a processing mode based on a continuous gesture. Selectively employing a processing mode may include detecting an existence of a mode-selection gesture by an individual. Upon detecting the existence of the mode-selection gesture, a first mode for processing facial micromovements of the individual may be continuously implemented while continuously detecting the mode-selection gesture. Following the continuous detection of the mode-selection gesture, a cessation of the mode-selection gesture may be detected. In response, implementation of the first mode may cease, and implementation of a second mode different from the first mode may be initiated for processing facial micromovements of the individual.
Owner:APPLE INC

Method for real-time generation of empathy expression of virtual human based on multimodal emotion recognition and artificial intelligence system using the method

Provided are a conversational artificial intelligence (AI) system and method based on real-time multimodal emotion recognition. The system includes a model server configured to provide a machine learning-based conversational model, a terminal configured to perform a conversation with the machine learning-based conversational model through the model server, display a virtual human responding to a user during a conversation with the user, and capture a facial image of the user during the conversation, and a multimodal empathetic conversation-generation system configured to access the model server and receive a response to a question of the user from the terminal, and assess an emotion of the user from the facial image of the user and control, based on the assessed emotion, an expression of the virtual human displayed on the terminal.
Owner:SANGMYUNG UNIV IND ACAD COOP FOUND

Multi-sensory autonomous multimodal emotion-synchronized environmental control architecture and regulation system (amesecar)

An autonomous environmental regulation and behavioral monitoring system is disclosed, configured to adapt temperature, lighting, and acoustic conditions based on real-time emotional and physiological data. The system includes a dual-redundant central processor, hierarchical communication networks, multi-angle visual acquisition units, infrared thermometers, and modular environmental subsystems. It detects posture, gestures, facial expressions, and thermal signals to classify user states and apply individualized airflow, light, and sound modulation without relying on external internet connectivity. The system also monitors connected appliances using voltage-based pressure analysis to forecast device degradation. With integrated gesture recognition, privacy-preserving data handling, and predictive adaptation, the invention enables multi-user personalization, long-term learning, and uninterrupted operation within residential, administrative, or healthcare infrastructures.
Owner:SEYEDKHAMOUSHI FAEZEHALSADAT +1

AI-powered personalized advertising system

An AI-driven system for personalized advertising in real time, where: ◯ an analytics unit to monitor and analyze user behavior in real time across multiple digital platforms such as websites, mobile applications, social media, and smart devices; the unit collects data on user engagement, browsing patterns, time spent, and content preferences to enable targeted, personalized advertising; ◯ a prediction module that predicts preferences and interests of a user, operatively connected to the user interaction data collection unit, wherein the prediction module uses machine learning models such as deep learning, recurrent neural networks (RNNs) and transformer-based architectures to predict interests of users based on historical interactions and inferred preferences; ◯ an emotion and sentiment analysis unit that assesses the user's mood in real time through computer vision, natural language processing (NLP) and voice analysis, whereby the analysis of facial expressions, voice pitch and linguistic mood is used to determine emotional states and receptivity to advertising content; ◯ an embodiment of a context awareness component in operational communication with the emotion analysis and mood unit, in some cases further augmented by various environmental and situational data such as the device type and its physical location, date and time, and the content processed in the device, processing methods, etc., in establishing adaptability and automatic ad placement to be as relevant as possible to the user and their status as prescribed; ◯ Use reinforcement learning algorithms and generative AI models to drive advertising with personalization engines. Creative elements, messaging, and presentations are dynamically adjusted in real time based on user responses to ensure advertising is personalized and always optimized for best performance; o a privacy-focused AI system with federated learning, differential privacy methods, and on-device AI processing to reduce targeted advertising while complying with international data protection laws such as the General Data Protection Regulation (GDPR) and California consumer privacy laws; o an operationally adapted ad delivery mechanism to engage with real-time bidding (RTB) networks, programmatic advertising exchanges and demand-side platforms (DSPs) and place advertisements through digital advertising networks, which guarantees the delivery of tailored advertising to the most relevant audience in real time; and o a contextual feedback loop in which the machine learning models used in the preference and interest prediction module and in advertising personalization The engine is continuously updated to reflect the latest user engagement data, improving personalization over time and optimizing advertising performance.
Owner:AL-ABABNEH HASSAN ALI +3

Multi-modal interaction method and system of digital human intelligent agent

The invention relates to the field of multi-modal interaction analysis, in particular to a multi-modal interaction method and system of a digital human agent. The method comprises the following steps: acquiring a real-time face image and a voice signal input stream of an interactive user based on an intelligent agent; performing real-time micro-expression recognition and deep emotion analysis based on the real-time facial image to obtain real-time emotion features of the user; performing time sequence evolution analysis on the real-time emotion characteristics of the user, performing holographic user emotion deep mining, and constructing a user emotion holographic characteristic spectrum; carrying out adaptive acoustic gain processing on the voice signal input stream, and carrying out voice-emotion association analysis based on the user emotion holographic characteristic spectrum to generate a voice-emotion linkage mapping spectrum; and carrying out eyeball fixation point migration tracking based on the user emotion holographic feature map and the real-time face image, and generating a user interaction depth intention signal. Through the real-time deep semantic understanding and emotion perception ability, the intelligent agent interaction intelligence and response accuracy are improved.
Owner:GUANGDONG HUITONG INFORMATION TECH CO LTD

Real-time video stream behavior identification and early warning system

The invention relates to the technical field of video behavior recognition, and discloses a behavior recognition and early warning system for a real-time video stream. The system comprises a spatio-temporal feature modeling module, a behavior fragment extraction module, an anomaly propagation modeling module, a risk area positioning module and an early warning strategy generation module. The spatial-temporal feature modeling module builds a dynamic model based on historical data, captures a skeleton key point three-dimensional coordinate sequence, a motion optical flow vector field and a micro-expression intensity spectrum, and outputs a theoretical behavior mode vector; the behavior fragment extraction module generates a multi-modal difference feature tensor through cross-modal difference analysis; the exception propagation modeling module generates an exception propagation path risk probability distribution cloud picture in combination with spatial constraint and trajectory information; the risk area positioning module identifies a high-risk area and marks a boundary; and the early warning strategy generation module dynamically configures monitoring parameters, starts high-frame-rate micro-expression capture for a high-risk area, and performs a track disturbance test on an adjacent area.
Owner:GAOZI TECHNOLOGY (SHENZHEN) CO LTD

Self-learning multi-modal emotion recognition method based on multi-scale cavity attention

The invention provides a self-learning multi-modal emotion recognition method based on multi-scale cavity attention, and solves the problem of low recognition precision caused by different importance of basic action units of a face and different distances between key action units, and the problem of different modality confidence during decision-level fusion. The method comprises the following steps: preprocessing a facial expression image, inputting the facial expression image into a multi-scale cavity attention convolution module, extracting features through a parallel three-branch convolution structure, splicing the features, calibrating through an attention mechanism to obtain an enhanced feature map, and sending the enhanced feature map to a full connection layer to recognize emotion; an original electroencephalogram signal is input into a time-frequency-space three-dimensional feature extraction network, the signal is decomposed, differential entropy features are calculated and processed by a global attention module comprising a frequency spectrum attention module, a space attention module and a time attention module, time-frequency-space multi-dimensional feature representation is output, and emotions are recognized by a full connection layer; and finally, inputting the emotion recognition result of the facial expression and the electroencephalogram signal into a self-learning weight module, and obtaining a final emotion recognition result through dynamic weighted fusion.
Owner:DALIAN UNIV

Multi-mode-based AI digital human intelligent interaction method, system and equipment

The invention relates to the technical field of computer vision and human-computer interaction, and discloses an AI digital human intelligent interaction method, system and equipment based on multiple modalities, and the method comprises the steps: pre-awakening a digital human when a human face is detected, and further thoroughly awakening the digital human based on recognized preset voice information or preset gesture information; voice and video information of a user in the interaction process is obtained, a keyword extraction result, a gesture recognition result and an emotional state tag are generated, a pre-constructed knowledge base is utilized to retrieve related information, a big language generation model module is combined to generate an answer text, and the answer text is input into a preset voice synthesis model to generate emotional voice output. And based on the current emotional state label of the user, driving the digital human animation to be output in an emotional manner. According to the method and the system, the digital human for understanding the emotion of the user, generating personalized answers, providing voices with rich emotions and displaying natural expressions and actions can be created, better interaction with the user can be realized, and more humanized and effective services can be provided.
Owner:BEI JING WAN JIE SHU JU KE JI YOU XIAN ZE REN GONG SI WU HAN FEN GONG SI +1

Virtual human design and application platform and method based on artificial intelligence, equipment and medium

The invention provides a virtual human design and application platform, method and device based on artificial intelligence, and relates to the technical field of virtual digital humans. The method comprises the steps of performing local anonymization on multi-modal input data on user equipment, encoding generated anonymized multi-modal features to obtain a multi-modal feature vector, and inputting the multi-modal feature vector into an emotion calculation model to obtain a user emotion intensity quantized value; inputting the multi-modal feature vector into a context sensing model, and generating a user intention vector after context correction in combination with a knowledge graph; generating an updated personality parameter matrix according to the user emotion intensity quantized value and the user intention vector; and outputting voice waveform data, facial muscle motion parameters and skeleton joint coordinate data based on the personality parameter matrix, and driving the virtual digital human three-dimensional model to perform real-time rendering. According to the scheme, the naturalness, emotional resonance and long-term user retention rate of virtual digital human interaction can be improved, and user privacy data security is protected.
Owner:郑雯月

Vehicle-mounted emotion interaction method and device based on multi-dimensional recognition

The embodiment of the invention provides a vehicle-mounted emotion interaction method and device based on multi-dimensional recognition, and the method and device achieve the precise judgment of the emotion of a driver through innovatively constructing an emotion fusion recognition model and integrating the facial expression, voice emotion, driving behavior and physiological state features. And designing a scene-based self-adaptive interaction strategy, and establishing an interaction triggering threshold value for intelligent matching in combination with external environment data and a danger level. An interaction effect evaluation mechanism is introduced, an interaction strategy model is continuously optimized through an online learning module, and dynamic adjustment of personalized interaction content is achieved. According to the method, the defects of the traditional technology in the aspects of emotion recognition, interaction strategies, effect evaluation and the like are effectively overcome, and the intelligent level and the user experience of vehicle-mounted emotion interaction are remarkably improved.
Owner:SHENZHEN ZHI HUI LIN NETWORK TECH CO LTD

Personalized digital human generation method based on single video

The invention discloses a personalized digital human generation method based on a single video, and relates to the field of virtual digital human modeling and driving, and the method aims at the video and voice data of a target person, through introducing a multi-modal alignment constrained voice driving synchronization mechanism and combining semantic understanding and an emotion label expression generation model, a personalized digital human model is generated. High synchronization, nature and vividness of digital human facial expressions and voice contents are realized. According to the method, a self-supervised style consistency constraint is added in model training to ensure that the style of a generated character image is stable, and a 3D semantic mask is added in an image fusion stage to improve the fusion precision and realistic effect of a synthetic facial expression and a reference face. According to the method, the digital human can be rapidly cloned and driven to synthesize the expression through the voice only through a single video sample, the generation process is efficient, and the obtained digital human video has excellent sense of reality and interactivity.
Owner:LIANGSHENG DIGITAL CREATIVE DESIGN (HANGZHOU) CO LTD

Face recognition method and system for dynamic environment

The invention relates to the technical field of face recognition, in particular to a face recognition method and system for a dynamic environment, and the method comprises the steps: collecting a face video stream through a multispectral imaging device, and carrying out the preprocessing of dynamic noise reduction, distortion correction and the like; constructing a multi-scale space-time fusion feature extraction network to extract dynamic space-time features and fuse cross-modal features; establishing an environment disturbance simulation generation model to generate a virtual sample, and performing domain adaptive alignment; designing an online incremental feature updating mechanism to optimize parameters of the feature encoder; deploying a heterogeneous graph neural network to carry out multi-modal decision fusion; and a hierarchical verification architecture is adopted to complete identity recognition. The system comprises a data acquisition and preprocessing module, a multi-scale space-time fusion feature extraction module and the like. According to the method, the problem of face recognition in a dynamic environment is effectively solved, the recognition accuracy, robustness, real-time performance and reliability can be remarkably improved in the scenes of complex illumination, posture expression change, shielding, background noise and the like, and the method has a wide application prospect.
Owner:GUANGZHOU CHENGTA INFORMATION TECH CO LTD

Intelligent driving system and method and electronic equipment

The invention provides an intelligent driving system and method and electronic equipment, relates to the technical field of intelligent driving, and realizes personalized improvement of driver ability and safety guarantee of a driving environment through three cooperative mechanisms including a teaching strategy module, a safety perception module and a behavior management module. Analyzing the multi-dimensional student data through a teaching strategy module by using a multi-modal large language model, and generating a personalized teaching strategy adaptive to the ability of the student; a driving environment is monitored through a safety sensing module, when a specific complex scene is detected, a multi-source heterogeneous sensor data fusion strategy is adjusted according to a scene type and a preset intervention degree, and then a safety coping strategy matched with environment interference is generated; the driving behavior data are recognized through the behavior management module, and multi-mode interaction feedback is generated in combination with the physiological data and the state data of the driver and used for guiding and correcting the driving behavior.
Owner:YIXIAN INTELLIGENCE

Aircraft cabin environment personalized adjustment method based on sentiment analysis

The invention discloses an aircraft cabin environment personalized adjustment method based on sentiment analysis, and the method comprises the steps: collecting the facial expression, voice waveform and environment parameters of a passenger in real time through a cabin multi-source sensor, and generating a standardized physiological signal matrix and anonymized voice features through noise reduction and feature extraction; inputting the physiological signal matrix and the voice features into a pre-trained deep learning model, outputting an emotion index and a classification label, and updating model parameters through a federal learning framework; dynamically generating temperature, humidity and oxygen concentration adjusting instructions and dynamic weights by adopting a fuzzy reasoning system in combination with the emotion indexes and passenger preset preferences; environment adjustment is executed through a closed-loop control system, and environment parameter errors are fed back; synchronously updating a fuzzy inference system rule base and deep learning model parameters by utilizing reinforcement learning in combination with environmental parameter errors and emotion index changes, and completing optimization of a closed loop; according to the invention, real-time dynamic regulation and control and continuous adaptation optimization of the personalized cabin environment can be realized.
Owner:WENZHOU DOVER AVIATION IND GROUP CO LTD

Live broadcast system based on AI interaction

The invention discloses a live broadcast system based on AI interaction, and the system specifically comprises a data collection module which is used for obtaining the real-time bullet screen data and interactive behavior data of a user in a live broadcast process; the emotion judgment module is used for determining an emotion tendency state judgment result of the user in the live broadcast process through the real-time bullet screen data and the interactive behavior data; the parameter adjusting module is used for dynamically adjusting an emotion expression parameter of the AI anchor according to a live broadcast content theme and commodity recommendation key information on the basis of the emotion tendency state judgment result; and the style migration module is used for inputting the emotion expression parameters into a style migration model, generating dynamic expression image data and mapping the dynamic expression image data into an image model of an AI anchor. According to the method and the device, dynamic matching between AI anchor image generation and user emotion requirements is realized, the image generation capability of the AI anchor is improved, and the emotion interaction effect between the AI anchor and the user is remarkably enhanced.
Owner:东莞市三奕电子科技股份有限公司

Employment and entrepreneurship support system based on artificial intelligence

The invention discloses an employment and entrepreneurship support system based on artificial intelligence, relates to the technical field of occupational planning, and aims to solve the problem that a traditional support mode is insufficient in individuation and accuracy. The system comprises a multi-modal data acquisition module used for acquiring facial expressions, whole body dynamics and voice data of a user in at least 30 minutes of video dialogue in real time; a dialogue text is processed through a large language model based on a Transform architecture, continuous time sequence modeling is carried out on multi-modal dynamic behaviors by applying a liquid time constant network, and deep fusion reasoning is carried out in combination with multiple groups of multi-head Transform attention mechanisms. The core analyzes the real thought, psychological state, behavior pattern and core ability of the user through consistency verification, generates a structured dynamic user insight abstract, and customizes personalized vocational development or entrepreneurship planning according to the structured dynamic user insight abstract. According to the method, the potential of the user can be deeply informed, high-precision personalized planning is provided, the decision-making quality and success rate are improved, and the method has dynamic adaptation and continuous learning capabilities.
Owner:青岛市军队离休退休干部活动中心

System and method for AI-powered narrative analysis of video content

A system, a method and a processor are for AI-powered generation and delivery of video clips. The processor is configured to: load a first video file of a first video content item, the first video file comprising video frames associated with timestamps; load a first subtitle file of the first video content item, the first subtitle file comprising subtitle text associated with the timestamps; execute a natural language processing (NLP) model with the subtitle text as input, the NLP model including language pre-processing steps for classifying words, names or phrases in the subtitle text and associating initial classifiers with the subtitle text, the NLP model including one or more of a recurrent neural network (RNN), a Bidirectional Encoder Representations from Transformers (BERT) model, or a generative pre-trained transformer (GPT) model for a dialogue analysis comprising processing sequences of dialogue in the subtitle text in view of the initial classifiers to associate one or more portions of the dialogue with one or more first classifiers of first narrative elements; execute an image recognition model with at least some of the video frames as input, the image recognition model including a convolutional neural network (CNN) for an object detection analysis and a facial recognition analysis comprising processing video sequences to associate one or more of the video frames with one or more second classifiers of second narrative elements; generate a narrative map of the first video content item by temporally aligning the first narrative elements with the second narrative elements based on the timestamps associated with the video frames and the first subtitle file; and generate a video clip including at least one segment of the first video content item, the at least one segment including selected video frames associated with at least one of the first or second narrative elements identified from the narrative map and selected for inclusion in the video clip.
Owner:PARAMOUNT GLOBAL INC

Expression recognition method and system based on multi-scale features and spatial attention

The present invention relates to the technical field of expression recognition, and in particular, to an expression recognition method and system based on multi-scale features and spatial attention. The method includes: performing feature extraction on acquired facial image data by using an HNFER neural network model to obtain an original input feature map; performing pooling and concatenation on extracted features based on a CoordAtt attention mechanism to obtain a feature map; performing deep convolution processing on the feature map to obtain an attention map, and then performing element-by-element multiplication to obtain a final feature map; and performing feature transformation and normalization on the final feature map to obtain an expression category probability and output the expression category probability. In the present invention, by integrating scale perception and spatial attention technologies, the model can recognize and classify different emotional states more accurately and maintain high performance even under complex environmental conditions.
Owner:YANTAI UNIV

Visual sensor of a tactical gear to use facial recognition technology to identify a person of interest and to cause a responsive device on the tactical gear to notify a wearer

Disclosed are an apparatus, system, and method of a visual sensor of a tactical gear to use facial recognition technology to identify a person of interest and to cause a responsive device on the tactical gear to notify a wearer. In one embodiment, a personal protective equipment includes a responsive device integrated in a tactical gear and a visual sensor. The visual sensor of the tactical gear identifies a target person using an identity artificial intelligence model. The responsive device notifies a wearer of the tactical gear when the visual sensor identifies the target person. Further, the responsive device notifies an additional person when the visual sensor identifies a threat to a protectee of the wearer and the additional person. The responsive device may vibrate when the visual sensor of the tactical gear detects an ambient threat to the protectee of the wearer.
Owner:GOVERNMENTGPT INC

Craniofacial dynamic reconstruction method and system based on multi-modal data fusion

The invention relates to the technical field of medical image processing, and discloses a craniofacial dynamic reconstruction method and system based on multi-modal data fusion, and the method comprises the steps: arranging a multi-modal data collection device in a target craniofacial region, and obtaining a static CT image, a static MRI image, a dynamic expression video sequence and a surface electromyogram signal; preprocessing the static CT image, the static MRI image, the dynamic expression video sequence and the surface electromyogram signal; inputting the preprocessed data into a multi-scale finite element model, and simulating a coupling relationship between muscle contraction force and skin deformation by adopting a biomechanical driving strategy to generate a dynamic craniofacial model; and fusing the geometric error and the motion consistency score of the real data by adopting a linear regression method, and outputting a comprehensive reconstruction quality index. According to the method, the problems of low craniofacial dynamic modeling accuracy and poor robustness in the prior art can be solved.
Owner:青峰宇

AI multi-mode fusion interaction method, device, system and equipment

The invention discloses an AI multi-mode fusion interaction method, device, system and equipment. An edge device obtains multi-mode information of a user in a current scene mode; preprocessing the multi-modal information, and outputting result data conforming to the current scene mode; when the edge device opens the uploading authority and the AI function module cannot meet the multi-modal information processing requirement, the result data is sent to the cloud service device, so that the cloud service device carries out processing according to the priority processing strategy of the multi-modal information to generate multi-modal fusion data, and the multi-modal fusion data is returned to the edge device for storage. And the edge device determines whether to publish the multi-modal fusion data according to the service requirement and synchronizes the multi-modal fusion data to the mobile terminal device. According to the method and the device, the corresponding scene mode can be flexibly switched according to different scene requirements, and each AI function module is integrated on the edge equipment, so that the corresponding AI function module is triggered to process the multi-modal information and send the multi-modal information to the cloud service equipment for fusion processing, and efficient processing and fusion of the multi-modal information are realized.
Owner:XIAMEN RGBLINK SCI & TECH CO LTD

Multi-modal emotion recognition method and system based on cross-modal alignment and matching enhancement

The invention discloses an emotion recognition method and system based on cross-modal alignment and matching enhancement. According to the method, firstly, feature extraction is carried out on text, audio and video modalities in a data set, and then a text and audio cross-modal emotion alignment module and a text and video cross-modal emotion alignment module are constructed respectively, so that cross-modal semantic alignment is realized. Constructing an emotion label matching module based on an alignment result, generating modal pairs with similar emotions but different labels by using a difficult negative sample mining strategy, and paying attention to cross-modal emotion consistency through a dichotomy task guide model; performing modal feature fusion on the three modals through a six-layer attention crossing mechanism, finally splicing feature vectors, inputting the spliced feature vectors into a long-sequence context fusion modeling module for deep modal fusion, and capturing cross-modal interaction information; and the fused features are sent to an emotion classification module, and a final emotion category recognition result is output.
Owner:NANJING UNIV OF POSTS & TELECOMM

AI digital human interaction system

The invention discloses an AI (artificial intelligence) digital human interaction system, which comprises a dialogue management module for receiving emotional state information transmitted by an emotional recognition module and other information input by a user feedback processing module; tracking a dialogue state, selecting a proper response mode according to a dialogue strategy, and generating a corresponding system behavior; the system behavior instruction is transmitted to an action and expression generation module; the user feedback processing module is used for collecting and analyzing feedback data of the user; and the feedback module analyzes and processes the feedback data, extracts valuable information, and transmits an analysis result to the dialogue management module and the reinforcement learning module for guiding the improvement and optimization of the system. Comprehensive information acquisition is realized through the multi-modal information acquisition module, continuity and individuation can be enhanced through the dialogue management module, an interaction strategy is optimized through the reinforcement learning module, and continuous improvement and real-time feedback are ensured through the user feedback processing module; and the interaction effect of the whole system is more natural and vivid.
Owner:SHANGHAI YEKE INTELLIGENT TECHNOLOGY CO LTD

Old people emotion recognition method and device based on multi-modal perception

The embodiment of the invention provides an elderly emotion recognition method and device based on multi-modal perception, and the method and device achieve the optimization and enhancement of the signal quality through innovatively constructing a multi-modal data preprocessing mechanism and integrating the facial expression, voice and posture features. And designing a personalized feature mapping model based on historical emotion expression data, and establishing an adaptive feature fusion strategy for intelligent matching in combination with a cross-modal attention network. A hierarchical time sequence classification mechanism is introduced, dynamic modeling of the emotional development trend is realized through a long-short term memory network, and accurate prediction of the emotional state is supported. According to the method, the defects of the traditional technology in the aspects of multi-modal processing, personalized modeling, time sequence analysis and the like are effectively overcome, and the accuracy and reliability of sentiment recognition of the old people are remarkably improved.
Owner:SHENZHEN ZHI HUI LIN NETWORK TECH CO LTD

Optimization of overall editing vector to achieve target expression photo editing effect

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for automatically generating datasets for a particular target expression photo effect. In one aspect, a system comprises receiving a plurality of image pairs, each comprising an original face image and an expressive face image representative of a target expression photo editing effect, generating an initial overall editing vector, wherein generating the initial overall editing vector comprises processing each image pair using a style space encoder model to generate an embedding of the original face image and an embedding of the expressive face image in an embedding space, optimizing the initial overall editing vector in accordance with one or more optimization criteria to generate an optimized overall editing vector, and applying the optimized overall editing vector to an input face image to generate a target expression face image that has the target expression photo editing effect.
Owner:GOOGLE LLC

Detecting and utilizing facial micro-motion

Systems, methods, and non-transitory computer-readable media containing instructions are disclosed for detecting and utilizing facial skin micro-motion. In some non-limiting embodiments, detection of facial skin micro-motion occurs using a speech detection system that may include a wearable housing, a light source (coherent light source or incoherent light source), a light detector, and at least one processor. The one or more processors may be configured to analyze light reflections received from the facial region to determine facial skin micromotion, and extract meanings from the determined facial skin micromotion. Examples of meanings that may be extracted from the determined facial skin micromotion may include words spoken by the individual (silent or voiced), an identity of the individual, an emotional state of the individual, a heart rate of the individual, a respiratory rate of the individual, or any other biometric, emotion or speech related indicator.
Owner:APPLE INC

Personalized image generation using combined image features

Examples described herein relate to personalized image generation using combined image features. A plurality of input images is provided by a user of an interaction application. Each of the plurality of input images depicts at least part of a subject. Each input image is encoded to obtain an identity representation. The identity representations obtained from the plurality of input images are combined to obtain a combined identity representation associated with the subject. A personalized output image is generated via a generative machine learning model. The generative machine learning model processes the combined identity representation and at least one additional image generation control to generate the personalized output image. At a user device, the personalized output image is presented in a user interface of the interaction application.
Owner:SNAP INC