Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1999 results about "Emotionality" patented technology

Emotionality is the observable behavioral and physiological component of emotion. It is a measure of a person's emotional reactivity to a stimulus. Most of these responses can be observed by other people, while some emotional responses can only be observed by the person experiencing them. Observable responses to emotion (i.e., smiling) do not have a single meaning. A smile can be used to express happiness or anxiety, a frown can communicate sadness or anger, and so on. Emotionality is often used by psychology researchers to operationalize emotion in research studies.

Gait emotion recognition method, system, storage medium, and computer equipment based on spatiotemporal graph convolution.

This invention relates to a gait emotion recognition method, system, storage medium, and computer device based on spatiotemporal graph convolution. The method includes the following steps: S1, data augmentation by reversing the temporal direction of gait; S2, obtaining deep emotion features and prior emotion features respectively through a spatiotemporal graph convolutional network and prior feature statistical methods; S3, performing nonlinear mapping on the prior emotion features using a feature mapping layer; S4, inputting the fused features of the deep emotion features and prior emotion features into an emotion classifier to obtain the emotion category. The feature mapping layer of this invention achieves more effective feature fusion by performing nonlinear mapping on prior features; it also introduces causal temporal convolution to replace general temporal convolution, effectively extracting fine-grained temporal features by enhancing temporal correlation and cross-period feature fusion. Furthermore, a walking direction recognition auxiliary task is designed to accelerate the training and convergence speed of the model, enhancing the ability to extract temporal-dependent features and the performance of emotion recognition.
Owner:SOUTH CHINA UNIV OF TECH

Intelligent psychological intervention system based on multi-modal fusion

The invention discloses an intelligent psychological intervention system based on multi-modal fusion, which is characterized in that a three-dimensional evaluation system is constructed by integrating speech sentiment analysis, keyboard dynamics monitoring and physiological signal acquisition, and time sequence alignment and feature weighted fusion of multi-source data are realized by adopting a cross-modal Transform model. The core of the system comprises an adaptive intervention engine which defines a multi-dimensional state space based on a hierarchical reinforcement learning architecture, optimizes an intervention strategy through a PPO algorithm, and realizes dynamic emotion interaction in AR and VR scenes in combination with a digital twin training module; according to the clinical decision support system, physiological behavior characteristics and psychological assessment trends are integrated by using a multi-time scale risk prediction model, and a personalized early warning threshold system is constructed, so that the psychological state recognition accuracy is improved, the intervention intensity self-adaptive adjustment response time is shortened, and the high-risk signal early warning timeliness reaches the minute level; and the problems of evaluation hysteresis and strategy stiffness of traditional psychological intervention are obviously improved.
Owner:JIANGSU ZHUODUN INFORMATION TECH CO LTD

Digital human interaction control method and device fusing emotional semantics and logical reasoning and storage medium

The invention provides a digital human interaction control method and device fusing emotion semantics and logical reasoning and a storage medium. The method comprises the following steps: analyzing multi-modal input data of a user, constructing emotion-semantics joint representation, and generating a logic decision path; and through a cognitive fusion module, emotion-semantic representation and a logic decision path are fused, and an interaction response adapting to emotion and logic consistency is generated. The system optimizes an emotion semantic model and a logical reasoning rule on line according to user feedback and interaction history, and real-time interaction of emotion dynamic and logical rules is achieved. The system can dynamically adjust the logic decision path based on the multi-mode emotional state of the user, and improves the naturalness and situation adaptability of interaction. A dynamic time warping algorithm and a factorization machine are introduced to process a multi-modal feature fusion problem, and the accuracy and robustness of emotional state recognition are improved. The online optimization mechanism enables the model and the rule to be evolved continuously, and reasoning errors are corrected automatically through user feedback, so that error circulation is avoided.
Owner:HANGZHOU DIGITAL SPACE TECHNOLOGY CO LTD

Dialogue interaction system based on multi-modal emotion perception and knowledge graph dynamic enhancement

The invention belongs to the field of artificial intelligence, and provides a dialogue interaction system based on multi-mode emotion perception and knowledge graph dynamic enhancement. A user edge terminal obtains multi-modal data, a lightweight Transform fusion network is adopted, the multi-modal data is converted into a fusion feature vector through cross-modal attention fusion and dynamic weight adjustment, and the fusion feature vector and historical conversations in a preset round are compressed in real time; the cloud service platform inputs the compressed data into DKGE, mining and fusing feature vectors and entity knowledge and emotional relations implied in historical dialogues in real time in the dialogue interaction process to update a dynamic knowledge graph, performing knowledge enhancement processing based on the dynamic knowledge graph, and constructing an initial reply prototype of the current round of interaction of the user; inputting the fusion feature vector and the updated dynamic knowledge graph into a dialogue strategy model, and determining a response strategy and a knowledge calling direction of the current round of dialogue; and the edge terminal generates real-time interaction reply information according to the initial reply prototype, the response strategy and the knowledge calling direction.
Owner:LONGMA ZHIXIN (ZHUHAI HENGQIN) TECH CO LTD

Large-model complaint intention recognition method based on sentiment analysis

The invention discloses a large-model complaint intention recognition method based on sentiment analysis, and the method comprises the steps: obtaining call voice data of a customer, and converting the call voice data into text data; preprocessing the text data, and removing noise and marking components; performing emotion feature analysis on the text data, extracting an emotion feature value, and generating an emotion vector; the emotion vectors and the text data are input into a large model together, and the large model combines the emotion feature values and context semantics to generate intention feature vectors; constructing an emotion-intention state vector, inputting the emotion-intention state vector into an asynchronous dominant actor commentator algorithm model, and generating a corresponding complaint intention probability value; judging a complaint intention probability, and generating risk early warning; and preferentially distributing high-risk customers and responding to customer demands. Through combination of voice data preprocessing, text semantic feature extraction, emotion intensity quantitative analysis and a multi-modal fusion algorithm and efficient risk assessment based on an asynchronous dominant actor reviewer model, the early warning and response capabilities of customer complaint risks are significantly improved.
Owner:HENAN ZHONGYUAN CONSUMER FINANCE CO LTD

Robot anthropomorphic interaction method based on multi-modal emotion recognition and customized portrait generation

The invention discloses a robot anthropomorphic interaction method based on multi-modal emotion recognition and customized portrait generation. The method comprises the following steps: S1, dynamically fusing multi-modal emotions; the method comprises the following steps: S1, synchronously acquiring voice, visual and text signals through a multi-source heterogeneous sensor, capturing a user voice stream by a high-fidelity microphone array, and extracting acoustic characteristics such as intonation and speed, S2, performing cross-modal reasoning; s3, synchronously generating contents; step S4: style migration; step S5, anthropomorphic voice and expression generation; according to the method, man-machine interaction emotion is analyzed and generated by utilizing a large language model and multi-modal information fusion, the singleness of interaction emotion and the deficiency of emotional sharing ability are avoided, a strong emotion interaction characteristic is achieved, the image of the robot is obtained through a generative technology and can be migrated to any image, the limitation that a specific image is independently made is broken through, and the interaction effect of the robot is improved. The advantage that one robot can be suitable for different scenes is achieved.
Owner:JIANGSU YUNMU ZHIZAO TECH CO LTD

Scene-based emotional interactive accompanying doll system and method

The invention discloses a scene-based emotional interactive accompanying doll system and method, and relates to the technical field of artificial intelligence, and the system comprises a data collection module, a feature analysis module, a portrait construction module, an interactive decision module, an execution control module and a database module. The data acquisition module is used for acquiring interaction data and scene data; the feature analysis module generates an emotion feature vector and a scene feature vector by using a multi-modal feature network, and determines a user emotion state and a scene type; the portrait construction module is used for constructing a user portrait; the interaction decision module comprises an emotion evolution unit, a scene interaction unit and a physiological regulation unit and is used for generating an execution instruction sequence; the execution control module is used for controlling the doll to complete emotion interaction behaviors. Through the scene perception and sentiment analysis technology, accurate recognition and personalized interaction response of the doll to the user sentiment state are achieved, and intelligent sentiment accompanying service is provided.
Owner:BEIJING LEKAIWENYU TECHNOLOGY CO LTD

Method for real-time generation of empathy expression of virtual human based on multimodal emotion recognition and artificial intelligence system using the method

Provided are a conversational artificial intelligence (AI) system and method based on real-time multimodal emotion recognition. The system includes a model server configured to provide a machine learning-based conversational model, a terminal configured to perform a conversation with the machine learning-based conversational model through the model server, display a virtual human responding to a user during a conversation with the user, and capture a facial image of the user during the conversation, and a multimodal empathetic conversation-generation system configured to access the model server and receive a response to a question of the user from the terminal, and assess an emotion of the user from the facial image of the user and control, based on the assessed emotion, an expression of the virtual human displayed on the terminal.
Owner:SANGMYUNG UNIV IND ACAD COOP FOUND

AI multi-mode emotion interaction memory terminal

The invention relates to the technical field of AI interaction, and discloses an AI multi-modal emotion interaction memory terminal, which realizes microsecond-level synchronization of voice, facial expression and text data through a multi-thread acquisition engine, dynamically allocates each modal weight by adopting a multi-head cross attention mechanism, and adaptively adjusts modal importance based on a conversation context hidden state; when the cross-modal confidence difference exceeds a threshold value, a gating LSTM conflict resolution module is activated, and the multi-source data collaboration problem is solved; the emotional memory modeling constructs an emotional state transition topology based on a graph convolutional network, protects user privacy in combination with a differential privacy mechanism, and realizes associated event storage of millisecond backtracking of short-term memory and long-term memory. The technology integrates multi-modal dynamic perception, privacy security calculation and adaptive learning ability, significantly improves the real-time performance and personification degree of emotion interaction, and can be applied to the fields of intelligent customer service, emotion accompanying, health monitoring and the like.
Owner:SHENZHEN XINZHI FUTURE TECHNOLOGY CO LTD

User emotion recognition and psychological intervention system and method based on large language model

Aiming at the problems of insufficient language understanding depth, weak personalized dialogue generation ability, lack of continuous learning and long-term user state modeling and the like in the current emotion recognition and psychological intervention technology, the invention provides a user emotion recognition and psychological intervention method combined with a large language model (LLM). According to the method, the potential emotional state is identified by analyzing free text information input by a user by utilizing the powerful capabilities of a large language model in the aspects of natural language understanding, emotional modeling and text generation; constructing a multi-round dialogue context, and reasoning a psychological change trend of the user; in combination with a psychological knowledge base, personalized and mild psychological intervention dialogue content with a dredging effect is generated. The system supports recognition and classification of various emotional states such as depression, anxiety and alonity, is suitable for various interaction scenes (such as APPs, webpages and social robots), and can greatly improve the precision of emotion recognition and the timeliness and effectiveness of psychological intervention. The emotion recognition and psychological intervention method based on the large language model provides solid technical support for constructing an intelligent, continuous and personalized psychological health management system, and has wide application prospects and profound social significance.
Owner:CHANGCHUN UNIV OF TECH

Multi-modal interaction method and system of digital human intelligent agent

The invention relates to the field of multi-modal interaction analysis, in particular to a multi-modal interaction method and system of a digital human agent. The method comprises the following steps: acquiring a real-time face image and a voice signal input stream of an interactive user based on an intelligent agent; performing real-time micro-expression recognition and deep emotion analysis based on the real-time facial image to obtain real-time emotion features of the user; performing time sequence evolution analysis on the real-time emotion characteristics of the user, performing holographic user emotion deep mining, and constructing a user emotion holographic characteristic spectrum; carrying out adaptive acoustic gain processing on the voice signal input stream, and carrying out voice-emotion association analysis based on the user emotion holographic characteristic spectrum to generate a voice-emotion linkage mapping spectrum; and carrying out eyeball fixation point migration tracking based on the user emotion holographic feature map and the real-time face image, and generating a user interaction depth intention signal. Through the real-time deep semantic understanding and emotion perception ability, the intelligent agent interaction intelligence and response accuracy are improved.
Owner:GUANGDONG HUITONG INFORMATION TECH CO LTD

Multi-mode-based AI digital human intelligent interaction method, system and equipment

The invention relates to the technical field of computer vision and human-computer interaction, and discloses an AI digital human intelligent interaction method, system and equipment based on multiple modalities, and the method comprises the steps: pre-awakening a digital human when a human face is detected, and further thoroughly awakening the digital human based on recognized preset voice information or preset gesture information; voice and video information of a user in the interaction process is obtained, a keyword extraction result, a gesture recognition result and an emotional state tag are generated, a pre-constructed knowledge base is utilized to retrieve related information, a big language generation model module is combined to generate an answer text, and the answer text is input into a preset voice synthesis model to generate emotional voice output. And based on the current emotional state label of the user, driving the digital human animation to be output in an emotional manner. According to the method and the system, the digital human for understanding the emotion of the user, generating personalized answers, providing voices with rich emotions and displaying natural expressions and actions can be created, better interaction with the user can be realized, and more humanized and effective services can be provided.
Owner:BEI JING WAN JIE SHU JU KE JI YOU XIAN ZE REN GONG SI WU HAN FEN GONG SI +1

Personalized English education system and method based on multi-modal sentiment analysis

The invention provides a personalized English education system and method based on multi-modal sentiment analysis. The system comprises a multi-modal interaction module for receiving and processing multi-modal data of students, an emotion recognition module for performing emotion analysis on the input multi-modal data, and a personalized module for evaluating real-time learning states of the students based on historical learning data of the students, and the core brain module is used for dynamically adjusting interactive feedback according to the emotional state and the real-time learning state. The emotion recognition module comprises emotion information fusion, the emotion information fusion adopts a weighting strategy, and the final output emotion state is adjusted through emotion consistency constraint and a conflict correction mechanism. According to the invention, through an emotion consistency loss function, a conflict correction mechanism and a knowledge graph-based super-outline control mechanism, the emotion and cognitive states of the students are accurately identified.
Owner:XIAMEN UNIV

Virtual human design and application platform and method based on artificial intelligence, equipment and medium

The invention provides a virtual human design and application platform, method and device based on artificial intelligence, and relates to the technical field of virtual digital humans. The method comprises the steps of performing local anonymization on multi-modal input data on user equipment, encoding generated anonymized multi-modal features to obtain a multi-modal feature vector, and inputting the multi-modal feature vector into an emotion calculation model to obtain a user emotion intensity quantized value; inputting the multi-modal feature vector into a context sensing model, and generating a user intention vector after context correction in combination with a knowledge graph; generating an updated personality parameter matrix according to the user emotion intensity quantized value and the user intention vector; and outputting voice waveform data, facial muscle motion parameters and skeleton joint coordinate data based on the personality parameter matrix, and driving the virtual digital human three-dimensional model to perform real-time rendering. According to the scheme, the naturalness, emotional resonance and long-term user retention rate of virtual digital human interaction can be improved, and user privacy data security is protected.
Owner:郑雯月

Dynamic self-adaptive multi-modal sentiment analysis fusion method and system

The invention provides a dynamic self-adaptive multi-modal sentiment analysis fusion method and system, and relates to the technical field of multi-modal sentiment analysis. The method comprises the following steps: synchronously acquiring voice, text, facial expression and limb movement data of a target user to form a multi-modal data set; the method comprises the following steps: firstly, extracting emotional characteristics of each mode, and constructing a cross-mode correlation model to capture a collaborative and complementary relationship among different modes; and calculating a real-time confidence score and a complementarity index of each modal based on the weight matrix of the cross-modal correlation model. Then, according to the scores and the indexes, a weighted average or maximum entropy algorithm is dynamically selected to fuse multi-modal emotion features, and a comprehensive emotion feature vector is generated; and finally, inputting the vector into a pre-training deep learning model, and outputting an emotional state category of the user. According to the method, the user emotion can be accurately and comprehensively captured, efficient emotion recognition and classification are realized, and the robustness and flexibility of an emotion analysis system in a complex scene are improved.
Owner:HUNAN OPEN UNIV (HUNAN PROVINCIAL CADRE EDUCATION & TRAINING ONLINE COLLEGE)

Vehicle-mounted emotion interaction method and device based on multi-dimensional recognition

The embodiment of the invention provides a vehicle-mounted emotion interaction method and device based on multi-dimensional recognition, and the method and device achieve the precise judgment of the emotion of a driver through innovatively constructing an emotion fusion recognition model and integrating the facial expression, voice emotion, driving behavior and physiological state features. And designing a scene-based self-adaptive interaction strategy, and establishing an interaction triggering threshold value for intelligent matching in combination with external environment data and a danger level. An interaction effect evaluation mechanism is introduced, an interaction strategy model is continuously optimized through an online learning module, and dynamic adjustment of personalized interaction content is achieved. According to the method, the defects of the traditional technology in the aspects of emotion recognition, interaction strategies, effect evaluation and the like are effectively overcome, and the intelligent level and the user experience of vehicle-mounted emotion interaction are remarkably improved.
Owner:SHENZHEN ZHI HUI LIN NETWORK TECH CO LTD

Text user sentiment analysis method and system based on multiple modes and AI

The invention provides a text user sentiment analysis method and system based on multiple modes and AI, and relates to the technical field of text analysis, and the method comprises the steps: obtaining and preprocessing text, image and audio information of a user, and extracting semantic, visual and acoustic feature vectors; multi-modal features are fused through a cross-modal attention mechanism; utilizing a graph neural network to construct a user emotion social graph to calculate emotion propagation intensity; and finally, obtaining an analysis result containing emotion category and intensity through an emotion classifier. The emotion state of the user can be comprehensively captured, the emotion analysis accuracy is improved, and complex emotion expression is effectively recognized.
Owner:ZHEJIANG SHUXIN NETWORK CO LTD

Voice emotion recognition method and device based on context information, equipment and medium

PendingCN120636474ASpeech recognitionSingle sentenceSpeech sound
The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a context information-based voice emotion recognition method, device, equipment and medium, which comprises the following steps: receiving an original voice stream and generating an independent voice segment, recognizing a text and determining a speaker role type, and extracting an acoustic feature index; and generating a preliminary emotion label, generating context information in combination with the historical dialogue text, and inputting the context information, the preliminary emotion label, the speaker role type and the acoustic feature index into a multi-modal fusion module to generate an emotion judgment result. According to the method, multi-modal fusion is realized on the basis of context information by combining voice, text and role information, so that the emotion change of each role can be accurately recognized and understood in a complex dialogue scene, the problems of large single sentence emotion judgment error and neglect of the context information in a traditional method are avoided, and the accuracy and stability of emotion recognition are effectively improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Psychotherapy-healing-oriented large model dialogue agent

The invention relates to the field of healing dialogues, and particularly discloses a psychotherapy-healing-oriented large-model dialogue agent which realizes deep analysis of core emotions, cognitive states and potential demands of a user by integrating real-time input, historical dialogue abstracts and emotional memory and adopting a structured coding and feature fusion technology. Based on the emotion signal and the semantic analysis result, matching a predefined emotional strategy label, and ensuring that a response strategy accords with psychological intervention logic. The psychological theory is converted into an executable dialogue framework through key element recognition and instruction embedding, and the integrating degree and safety are improved in cooperation with multi-candidate generation and an emotional quality evaluation mechanism based on semantic embedding. According to the scheme, by continuously optimizing emotion representation and strategy matching, the defects of a traditional model in the aspects of emotion tracking, professional fusion and safety control are effectively overcome, and the depth and reliability of psychological support are remarkably improved.
Owner:WEST LAKE XINCHEN (HANGZHOU) TECH CO LTD

Emotion detection system based on facial recognition

The invention discloses an emotion detection system based on facial recognition, and relates to the technical field of computer vision and emotion calculation. A video stream time sequence analysis module is used for extracting a facial micro-expression image sequence of continuous frames, a time sequence feature vector containing a micro-expression intensity gradient, an illumination robustness coefficient and a facial action unit cooperation feature is generated, and a multi-mode dynamic sensing module is combined to carry out real-time analysis on an emotion classification probability, voice emotion parameters and physiological signals. And the fusion decision module performs dynamic weighted fusion on the multi-modal data based on the scene adaptive weight, and finally generates a comprehensive emotion score. Through multi-modal time sequence modeling and a dynamic weight optimization mechanism, the accuracy and environmental adaptability of emotion recognition are remarkably improved, and real-time perception and accurate decision making of customer emotion are realized in a target scene.
Owner:NORTHEAST FORESTRY UNIV

Methods and systems for speech emotion retrieval via natural language prompts

Methods and systems for generating training data for training a contrastive language-audio machine-learning model. A plurality of audio segments are retrieved from a speech emotion recognition (SER) database along with metadata associated with the audio segments. The metadata of each audio segment includes an emotion class. Words or terms associated with emotions are retrieved from a lexicon. A large language model (LLM) is executed on (i) the classes of emotion associated with the audio segments and (ii) the words or terms from the lexicon. This generates a plurality of text captions associated with emotion, which are stored in a caption pool. For each audio segment retrieved from the SER database, that audio segment is paired with one or more of the text captions from the caption pool that were generated based on the emotion class associated with that audio segment. This yields audio-text pairs for training a contrastive learning model.
Owner:ROBERT BOSCH GMBH

Digital human generation method based on multi-modal large model

The invention provides a digital human generation method based on a multi-modal large model. The method comprises the following steps: constructing a digital human basic model; generating a structured training set; generating a question and answer model supporting multi-channel interaction; semantic answers of the user questions are output, text emotional tendencies of the semantic answers are extracted, and emotional intensity parameters are output; generating facial muscle movement track data, and performing real-time rendering on the digital human basic model according to the facial muscle movement track data to output a digital human three-dimensional image with emotion expression. According to the embodiment of the invention, cross-modal alignment is carried out on text, image and audio data, and a multi-modal large model containing visual, voice and knowledge models is optimized by using a joint training method, so that more natural and smoother multi-channel interaction experience is realized; in addition, by introducing an emotion recognition model and a face interaction model, the emotion tendency contained in the semantic answer can be captured and reflected more accurately, so that a digital human three-dimensional image with real emotion expression is output.
Owner:CHINA NAT BUILDING MATERIALS TECH CO LTD +2

System and Method for Real-Time Identity-Free Personalization Using Fluid Emotional Trait Vectors, Modular Engine Mesh Architecture, Context-Aware Engagement Logic, and Adaptive Goal Mutation

A system and method for real-time, identity-free personalization using deformable emotional trait vectors to dynamically adapt digital and voice-based experiences. Each user session is modeled as a behavioral object known as a Vectra, composed of fluidic traits—such as mass, viscosity, temperature, volatility, and texture—that evolve continuously in response to live behavioral, contextual, environmental, and voice-derived signals. These Vectras traverse a dynamically warped emotional space, the Vectraverse, influenced by ambient conditions including time of day, noise level, inventory urgency, and engagement rhythm. Gravitational pull toward predefined emotional goal attractors modulates system behavior, while a goal mutation engine reclassifies session intent when confidence decays or friction spikes. Outputs include tone modulation, content pacing, offer framing, and gamified reward logic—all executed without storing identity, login credentials, or historical profiles. The system supports modular deployment across voice, screen, signage, mobile, and in-room environments, and integrates with large language models, AI agents, and third-party personalization stacks via privacy-safe APIs and federated learning. Designed for zero-ID personalization, the platform enables emotionally intelligent, context-aware engagement across any surface or session.
Owner:GINSBERG JUSTIN

Digital human interaction system and method based on multi-modal emotion recognition

ActiveCN121116129ASemantic analysisSpeech analysisInteractive modelingData stream
The embodiment of the invention provides a digital human interaction system and method based on multi-modal emotion recognition, and belongs to the technical field of digital human interaction. The system comprises a multi-modal sensing module used for collecting multi-modal data and preprocessing the multi-modal data to generate a standardized data stream; the cross-modal fusion and emotion recognition module is used for carrying out interactive modeling on the multi-modal features and outputting a current emotion label and emotion intensity; the reaction planning module is used for generating a composite reaction strategy; and the digital human rendering module is used for mapping the composite reaction strategy into control signals corresponding to the voice, the facial expression and the action respectively, and driving a digital human to execute corresponding voice output, facial expression change and limb action through the control signals so as to realize interaction. According to the method, multi-modal data are deeply fused through the cross-modal graph neural network and comparative learning, the weight is dynamically adjusted in combination with the modal confidence, and the emotion recognition accuracy and robustness are improved.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

Hierarchical emotional speech generation method and device, equipment and medium

The invention relates to the technical field of speech synthesis, can be applied to business scenes of medical health, financial science and technology and the like, and discloses a hierarchical emotional speech generation method, which comprises the following steps: acquiring an input text and an emotional speech sample, extracting a text embedding feature from the input text, extracting a Mel spectrum feature from the emotional speech sample, and obtaining a text embedding feature; the method comprises the following steps: extracting phoneme-level, word-level and statement-level sentiment distribution characteristics through a hierarchical sentiment distribution extraction module, carrying out time dimension alignment, generating a multi-level sentiment guidance matrix, and inputting the multi-level sentiment guidance matrix and text embedding characteristics into a sentiment synthesis module to generate a target Mel spectrum; and finally, converting the target Mel spectrum into target voice through a vocoder. The emotion control is expanded from the statement level to the fine-grained level of phonemes, words and the like, and the multi-level emotion guidance matrix is combined for generation, so that the generated speech is finer and more abundant in emotion expression and conforms to the context, and the naturalness of speech generation and the accuracy of emotion transmission are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Emotive text-to-speech with auto detection of emotions

A method of providing emotive text-to-speech includes obtaining input text characterizing a natural language response generated by an assistant LLM to a query input by a user during a conversation between the user and the assistant LLM, and processing, using the assistant LLM, the input text conditioned on an emotion detection task prompt to predict, as output from the assistant LLM, an emotional state of the natural language response. The method also includes determining, based on the emotional state of the natural language response predicted as output from the assistant LLM, an emotional embedding for the input text and instructing a TTS model to process the input text and the emotional embedding to generate a synthesized speech representation of the natural language response conveying the emotional state of the natural language response as specified by the emotional embedding.
Owner:GOOGLE LLC

Multi-modal emotion fusion analysis method and system

The invention discloses a multi-modal emotion fusion analysis method and system, and the method comprises the steps: carrying out the feature extraction of multi-modal emotion data modal by modal through a feature extraction module, and generating a text original feature, an audio original feature and a visual original feature; performing cross-modal alignment interactive fusion on the original text features, the original audio features and the original visual features based on a unified semantic alignment module, and constructing collaborative fusion features; performing mode and channel double-layer dynamic fusion optimization by adopting a dynamic fusion regulation and control module according to the text original feature, the audio original feature, the visual original feature and the collaborative fusion feature, and determining a unified fusion feature; performing hierarchical residual semantic gating enhancement based on the unified fusion features according to a high-order semantic abstraction module to generate semantic enhancement features; and inputting the semantic enhancement features into an emotion prediction module, and outputting an emotion analysis result. Based on the above scheme, a more stable, accurate and reliable emotion recognition result can be provided.
Owner:GUANGDONG UNIV OF TECH

Social robot identification method and system based on deep learning

The invention discloses a social robot identification method and system based on deep learning, and relates to the technical field of robots, and the method comprises the steps: carrying out the smooth processing of a tweet through a large language model, and generating a tweet fused with expression semantics in combination with a natural language processing model; capturing an emotion expression difference between a robot account and a real user; generating a comprehensive feature vector of global context sensing; outputting fusion features; the output fusion features are mapped to a high-dimensional space through a linear layer, and the detection probability of a robot account is output through an activation function; and based on adaptive moment estimation, gradient back propagation is carried out by using a cross entropy loss function, and parameters of the graph convolutional network model are optimized to obtain a final classification result. According to the method, the problems of sparse text expression and emotion information loss are effectively relieved, and the separability of the social robot and the real user in emotion behavior modes is enhanced.
Owner:曾卡芊

Multi-modal emotion recognition method based on heart and brain coupling and graph neural network

The invention relates to a multi-modal emotion recognition method based on heart and brain coupling and a graph neural network, and belongs to the field of artificial intelligence. Comprising the steps of data preprocessing, graph representation construction, multi-view graph convolutional network construction, fusion graph network construction and cross-domain joint optimization and sentiment classification. The method has the advantages that an adaptive adjacency matrix optimization strategy based on a triple constraint mechanism is proposed to solve the modal alignment and deviation problems represented by a multi-modal diagram in a data-driven branch, redundant noise is eliminated by adopting global regularization constraint, and unique feature representation in a modal is enhanced through modal specificity; a deep association rule is mined in combination with a cross-modal interaction module, the modeling ability of a heart and brain emotional state is improved, a multi-view image convolutional network is further designed, global features and local features are extracted, features of a cognitive heuristic branch and a data driven branch are combined by adopting an attention mechanism-based image fusion network, a domain confrontation strategy is introduced, and a cognitive network is constructed. And the generalization of the method is enhanced.
Owner:JILIN UNIVERSITY

Financial risk assessment system based on emotion perception and dynamic knowledge graph

The invention relates to the technical field of financial risk assessment, and discloses a financial risk assessment system based on emotion perception and a dynamic knowledge graph, and the technical scheme is characterized in that the system comprises a multi-source data collection module which is in butt joint with a plurality of data sources and collects structured data and unstructured data; carrying out standardization processing on the data to obtain standard data; the emotion perception engine is used for quantifying customer emotion features; the dynamic knowledge graph construction module is used for attenuating the weight of historical data according to the time dimension and updating the knowledge graph in real time when an entity extraction model is used for extracting an entity and a corresponding relation and incremental updating is carried out on the knowledge graph; the multi-modal risk assessment model is used for fusing structured data, emotion features and knowledge graph embedding and outputting a risk score; and the dynamic feedback and optimization module iteratively updates the risk assessment logic according to the risk score and the market feedback, and performs incremental training on the multi-modal risk assessment model through misjudgment data.
Owner:JIANGSU SUNING BANK CO LTD