Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

649 results about "Emotion recognition" patented technology

Emotion recognition is the process of identifying human emotion, most typically from facial expressions as well as from verbal expressions. This is both something that humans do automatically but computational methodologies have also been developed.

Digital human interaction system and method based on multi-modal emotion recognition

ActiveCN121116129ASemantic analysisSpeech analysisInteractive modelingData stream
The embodiment of the invention provides a digital human interaction system and method based on multi-modal emotion recognition, and belongs to the technical field of digital human interaction. The system comprises a multi-modal sensing module used for collecting multi-modal data and preprocessing the multi-modal data to generate a standardized data stream; the cross-modal fusion and emotion recognition module is used for carrying out interactive modeling on the multi-modal features and outputting a current emotion label and emotion intensity; the reaction planning module is used for generating a composite reaction strategy; and the digital human rendering module is used for mapping the composite reaction strategy into control signals corresponding to the voice, the facial expression and the action respectively, and driving a digital human to execute corresponding voice output, facial expression change and limb action through the control signals so as to realize interaction. According to the method, multi-modal data are deeply fused through the cross-modal graph neural network and comparative learning, the weight is dynamically adjusted in combination with the modal confidence, and the emotion recognition accuracy and robustness are improved.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

Multi-modal emotion recognition method and system based on cross-modal alignment and matching enhancement

The invention discloses an emotion recognition method and system based on cross-modal alignment and matching enhancement. According to the method, firstly, feature extraction is carried out on text, audio and video modalities in a data set, and then a text and audio cross-modal emotion alignment module and a text and video cross-modal emotion alignment module are constructed respectively, so that cross-modal semantic alignment is realized. Constructing an emotion label matching module based on an alignment result, generating modal pairs with similar emotions but different labels by using a difficult negative sample mining strategy, and paying attention to cross-modal emotion consistency through a dichotomy task guide model; performing modal feature fusion on the three modals through a six-layer attention crossing mechanism, finally splicing feature vectors, inputting the spliced feature vectors into a long-sequence context fusion modeling module for deep modal fusion, and capturing cross-modal interaction information; and the fused features are sent to an emotion classification module, and a final emotion category recognition result is output.
Owner:NANJING UNIV OF POSTS & TELECOMM

Multi-modal dialogue emotion recognition method and system based on cross-modal fusion and comparative learning

The invention relates to the technical field of natural language processing, in particular to a multi-modal dialogue emotion recognition method and system based on cross-modal fusion and comparative learning. According to the method, learnable residual scaling and pre-normalization are introduced through a cross-modal encoder, deep interaction of texts, voices and visual modals is stabilized, and gradient explosion is inhibited; a dialogue graph fusing semantic similarity and time proximity is constructed online through a semantic-time sequence graph enhancement module, and a graph attention network is used for explicitly modeling long-distance round dependence and cross-speaker emotion transmission; through an adaptive comparison and alignment module, a dynamic scheduling comparison loss and index moving average updated mode-emotion prototype library is adopted to realize data distribution adaptive cross-mode alignment; through cooperative work of the modules, the problems that in the prior art, cross-modal fusion is unstable, long-distance and cross-speaker dependence modeling is insufficient, and cross-dataset alignment capacity is weak are solved.
Owner:CHONGQING TELECOMM PLAN & DESIGN INST

Method of emotion recognition in cross-subject EEG signals

PendingUS20250384293A1Psychotechnic devicesSensorsMedicineAutologistic regression
A method of emotion recognition in cross-subject EEG signals, belonging to technical field of deep learning, includes the following steps: S1, constructing the extracted DE features into positive and negative samples by using a positive and negative sample generator; S2, sending the DE features of an anchor and the positive and negative samples into the encoder for coding, mapping the DE features to a latent space, performing regression prediction on the encoded anchor samples in the latent space by using an autoregressive model, training the encoder by using a probability supervision contrastive loss function; and S3, connecting the trained encoder to the classifier for fine tuning, and training the classifier through the cross entropy loss function; in this process, the encoder does not perform gradient propagation to complete cross-subject emotion recognition.
Owner:DALIAN UNIV

Multi-scene self-adaptive man-machine interaction system and method based on emotion recognition

The invention discloses a multi-scene self-adaptive man-machine interaction system and method based on emotion recognition, relates to the technical field of man-machine interaction, and solves the technical problems of realizing fusion perception of multi-modal emotion features and improving the accuracy of emotion judgment in complex scenes. According to the method, facial, voice and text emotion features are extracted by adopting a multi-modal fusion technology, the limitation of single-modal recognition is solved, an emotion-scene association rule base and a user portrait are constructed, real-time scene classification is combined, accurate mapping of emotions, scenes and demands is realized, one-step interaction of strategies is avoided, and the user experience is improved. Language interaction adaptation is designed from the form, content and style three-dimensional degree, it is ensured that languages are natural and fit scenes, functional response adaptation improves efficiency through priority ranking and execution mode optimization, environment linkage adaptation is combined with user emotion dynamic adjustment directions, collaborative linkage of languages, functions and environments is achieved, and strategy splitting is avoided.
Owner:NANJING LAOJIAJIA INTELLIGENT TECH CO LTD

Double-branch electroencephalogram emotion recognition method and system based on brain region topology and space-time

The invention belongs to the field of artificial intelligence and electroencephalogram emotion recognition, and provides a double-branch electroencephalogram emotion recognition method and system based on brain region topology and time-space, and the method comprises the steps: preprocessing a to-be-recognized electroencephalogram signal to obtain a plurality of electroencephalogram fragments, and extracting a difference entropy sequence of each electroencephalogram fragment and a Spearman correlation coefficient matrix between channels; based on the Spearman correlation coefficient matrix, utilizing a bridging dynamic graph attention network module to extract topological features of a brain region; processing the differential entropy sequence by using a multi-scale space-time mixed attention module to obtain multi-scale space-time features; carrying out residual mutual cross attention fusion on the topological features of the brain region and the multi-scale spatial-temporal features to obtain fusion features; and performing classification based on the fusion features, and determining an emotion recognition result corresponding to the electroencephalogram signal. According to the method, the accuracy and robustness of emotion recognition are improved by utilizing the spatial topology characteristics and the multi-topology time dynamic characteristics of the electroencephalogram signals, and the defects of modeling spatial dependence and time dynamic are overcome.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Intelligent accompanying robot system

The invention discloses an intelligent accompanying robot system, and the system comprises the following modules: a multi-mode interaction module which is composed of a voice recognition unit, a voice synthesis unit, an emotion recognition unit, and a visual recognition unit, and is used for collecting the voice, facial expression, motion, and environment information of a user; the localized AI decision module is used for deploying a lightweight DeepSeekR1 large language model based on an ESP32-S3 edge computing chip, performing real-time processing on the multi-modal data and generating a social guidance strategy and an emotion intervention instruction; the data management module comprises a user behavior database and a privacy protection unit and is used for desensitizing the data and realizing sensitive data isolation through local storage; according to the method, the single-person intervention cost is reduced, and the time consumption of manual scene simulation is reduced.
Owner:NANTONG UNIV

Intelligent customer service interaction content recommendation method and system based on artificial intelligence

The invention provides an intelligent customer service interaction content recommendation method and system based on artificial intelligence, and the method comprises the steps: obtaining power grid equipment state data and user behavior data, and carrying out the weighted fusion through an attention mechanism, and outputting a fusion feature vector; predicting a prediction vector set of a future time period based on the fused feature vector; screening a demand event set higher than a threshold in the prediction vector set, and generating a recommended content set; performing emotion recognition according to a current input text of the user, and calculating a recommended content score and a pushing priority score in combination with the recommended content set; a cross-modal consistency regular term is introduced to correct the push strategy vector, and an adjusted recommendation push score and a corresponding recommendation content set are obtained; and converting the adjusted recommendation push score and the corresponding recommendation content set into a scheduling task, and performing linkage execution with a scheduling system. According to the invention, efficient, active and intelligent upgrading of the intelligent customer service system in a power grid scene is realized.
Owner:HAINAN POWER GRID CO LTD

Multi-modal dialogue emotion recognition method and system

The invention relates to a multi-modal dialogue emotion recognition method and system, and belongs to the technical field of natural language processing, and the method comprises the steps: constructing a multi-modal dialogue emotion recognition network; the multi-modal dialogue emotion recognition network comprises a multi-modal feature extraction network, a semantic-guided global-local interactive fusion network, a three-modal comparison learning module and a modal dynamic balance optimizer; wherein the semantic-guided global-local interaction fusion network comprises a global semantic center construction module and a semantic-guided local interaction module; training a multi-modal dialogue emotion recognition network; and inputting the to-be-detected dialogue data into the trained multi-modal dialogue emotion recognition network for emotion recognition to obtain an emotion recognition result. According to the method, blind inter-modal interaction is converted into ordered information transmission based on global semantic understanding, the problem of limited interaction quality caused by lack of semantic guidance in a traditional method is solved, and the depth and accuracy of cross-modal understanding are improved.
Owner:HUBEI UNIV OF TECH

Multi-modal emotion recognition fusion method based on multi-head attention mechanism

The invention provides a multi-modal emotion recognition fusion method based on a multi-head attention mechanism, and the method comprises the steps: carrying out the feature extraction and fusion of various data, capturing the internal relation of a modal through the multi-head attention mechanism, achieving the information complementation between modals through cross-modal interaction, and dynamically adjusting the weight according to the quality of the modals. And a self-built database containing a large amount of Chinese data is constructed, and a culture adaptation optimization strategy is combined, so that the generalization performance and culture adaptability of the model in Chinese user groups are improved. The system is deployed on a cloud server, optimized hardware and software configuration is adopted, efficient model reasoning and multi-user concurrent processing are achieved, and the large-scale real-time application requirement is met. The emotion of the user can be monitored in real time and early warning can be provided in scenes such as psychological counseling and group emotion monitoring, professionals are assisted in better understanding the emotion state of the user, and the service effect is improved.
Owner:SHENZHEN SERUN HEALTH TECHNOLOGY CO LTD

Anti-fact multi-mode dialogue emotion causal reasoning method based on double-branch hypergraph

The invention discloses an anti-fact multi-mode dialogue emotion causal reasoning method based on a double-branch hypergraph. The method comprises the following steps: respectively extracting sentence level feature vectors of three modes of text, voice and vision from input multi-mode dialogue data; carrying out modeling on a high-order relationship in the modals and between the modals by utilizing a hypergraph structure, and constructing a dialogue hypergraph containing multi-modal nodes and emotion nodes; introducing a hypergraph attention network on the hypergraph, learning contribution weight of each modal node to a target emotion node, and selecting a candidate reason node set; the candidate reason nodes are intervened, an anti-fact branch is constructed, a fact situation and final node feature representation under the anti-fact situation are calculated, and a causal effect vector is obtained; and designing a joint optimization objective function, and carrying out joint training on emotion recognition loss and causal consistency loss to realize synchronous prediction of emotion categories and emotion reasons. According to the method, a high-order semantic relationship can be effectively modeled in a multi-modal dialogue scene, and a key reason for emotion formation is reasoned.
Owner:JIANGSU UNIV

Multi-modal sentiment analysis method and device based on feature decoupling and guide correction

The invention discloses a multi-modal sentiment analysis method and device based on feature decoupling and guide correction, and relates to the technical field of sentiment analysis, and the method comprises the steps: building an initial multi-modal sentiment analysis network, the initial multi-modal sentiment analysis network comprises a multi-modal feature extraction mapping block, a cross-modal interaction enhancement coding block, a feature decoupling block, a guide type gating semantic correction block and a low-rank fusion sentiment analysis block, and the total loss value of a training sentiment analysis result for training the multi-modal sentiment data is determined; and performing optimization iteration on the initial multi-mode sentiment analysis network according to the total loss value until a target multi-mode sentiment analysis network is determined, and outputting a target sentiment analysis result of the to-be-analyzed multi-mode sentiment data by adopting the target multi-mode sentiment analysis network. Based on the above scheme, the multi-modal fusion emotion recognition precision is improved.
Owner:GUANGDONG UNIV OF TECH

Emotion recognition method based on attention mechanism and capsule network

The invention discloses an emotion recognition method based on an attention mechanism and a capsule network. The emotion recognition method specifically comprises the following steps: inputting multi-channel original EEG signals; performing data truncation, data filling and data standardization on the original EEG signal, unifying the format, and generating a preprocessed EEG signal; extracting space-time joint features from the preprocessed EEG signals by using a three-dimensional convolutional neural network; the construction of the graph structure comprises modeling a complex relationship between channels; introducing a graph attention mechanism to strengthen information expression of important channels and relationships; and high-dimensional vector modeling and final classification of the features are carried out through a capsule network, and capture of deeper space structure information is completed.
Owner:XIAMEN UNIV

Prompt enhancement-based multi-modal emotion recognition method and system, medium and equipment

The invention discloses a multi-mode emotion recognition method and system based on prompt enhancement, a medium and equipment, and belongs to the technical field of emotion analysis, and the multi-mode emotion recognition method based on prompt enhancement comprises the steps: obtaining input data of text, visual and audio modes; the modal-based prompt encoder performs prompt template-based feature enhancement on each modal input data to generate prompt enhanced feature representation; establishing an inter-modal information exchange channel by adopting a cross-modal adaptive alignment mechanism based on prompt enhanced feature representation, and generating alignment dominant features; based on the aligned dominant features, dynamically fusing the high-quality features by adopting a multi-level quality evaluation mechanism to generate enhanced dominant features; and performing multi-modal fusion on the enhanced dominant feature and the other two modal features, and outputting an emotion prediction result. According to the method, different data missing modes are effectively dealt with, the limitation of a rigid interaction mode of a traditional method is solved, the quality and reliability differences of different modes are considered, and robust multi-mode sentiment analysis is achieved.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +2

Cross-subject electroencephalogram emotion recognition method and system based on space-time adaptive graph coding learning

The invention belongs to the field of deep learning, and provides a cross-subject brain electrical emotion recognition method and system for space-time adaptive graph coding learning, and the method comprises the steps: extracting features from different frequency bands through a sliding window technology, and constructing a feature matrix covering channels and time dimensions; fusing space-time hybrid embedding, time embedding and space embedding, and converting the feature matrix into a high-dimensional semantic vector; calculating channel characteristic difference and time trend change, and generating a dynamic topology matrix adaptive to cross-tested individual difference; combining the dynamic topological matrix to calculate spatial cross-channel and time stride length attention in parallel, and fusing and strengthening key spatial-temporal features; and the decoder maps the coding features by means of a domain adversarial decoding module, reduces the distribution difference between a source domain and a target domain through cross-domain adversarial training, and outputs an emotion classification result. Through the method, the cross-subject generalization ability and the recognition accuracy of electroencephalogram emotion recognition are improved.
Owner:ANHUI AGRICULTURAL UNIVERSITY

Multi-modal emotion calculation method and system

PendingCN121479676ASemantic analysisBiological modelsPersonalizationInteraction field
The invention discloses a multi-modal emotion calculation method and system, and mainly relates to the technical field of artificial intelligence and human-computer interaction. Comprising the following steps: synchronously collecting voice, visual and tactile multi-modal data of a user, and carrying out feature extraction on each modal data; performing cross-modal fusion on the extracted multi-modal features, including time sequence alignment and semantic association modeling, and generating a global emotion feature vector; performing dynamic emotion reasoning based on the global emotion feature vector, and outputting an emotion category and emotion intensity; generating a tactile feedback signal according to the emotion category and the emotion intensity, and driving an actuator to output; and updating the emotion memory graph based on user interaction data to complete personalized model self-adaption. The method has the beneficial effects that high-precision and low-delay emotion recognition and natural tactile feedback in a cross-culture scene are realized, and the core bottleneck of dynamic drift adaptation failure and emotion-behavior feedback decoupling in the prior art is solved.
Owner:ZHONGKE XINHE (BEIJING) TECHNOLOGY CO LTD

Multi-modal emotion recognition method based on TCN-GCN dynamic topology learning and two-stage fusion

The invention relates to a multi-modal emotion recognition method based on TCN-GCN dynamic topology learning and two-stage fusion, and belongs to the field of artificial intelligence. And acquiring a multi-modal emotion recognition data set, constructing an emotion recognition model based on TCN-GCN and two-stage fusion, inputting the training set into the model for iterative training, performing emotion recognition on the test set by using the trained model, and outputting an emotion recognition result. The method has the advantages that electroencephalogram space-time cooperation information can be completely reserved, the electroencephalogram channel mutual relation during emotion activation can be dynamically adapted, redundant information among multiple modes can be effectively filtered, complementarity of multi-mode features can be more comprehensively captured, and therefore the classification accuracy of emotion recognition is improved.
Owner:JILIN UNIVERSITY

Education resource recommendation method and system based on artificial intelligence

The invention discloses an educational resource recommendation method and system based on artificial intelligence. The method comprises the following steps: acquiring audio data, interaction data and task data acquired by a user terminal; performing feature extraction on the audio data, the interaction data and the task data to obtain an emotion feature vector, a learning rhythm vector and a content feature vector; inputting the emotion feature vector and the learning rhythm vector into a pre-constructed emotion recognition model to obtain a psychological state vector; the cognitive load is calculated based on the learning rhythm vector and the content feature vector, and then the cognitive load is corrected through the psychological state vector; matching a state interval of the corrected cognitive load according to a preset threshold interval; and according to the state interval, adjusting a difficulty coefficient of the recommended course, rearranging a course content sequence and an auxiliary learning prompt, generating structured data, and outputting the structured data as an intelligent auxiliary learning recommendation result. According to the invention, online learning interactivity and teaching quality in rural and remote areas are effectively improved.
Owner:NANJING NORMAL UNIVERSITY

Robust multi-mode emotion understanding method for intelligent customer service digital human

The invention relates to an intelligent customer service digital human-oriented robust multi-mode emotion understanding method, and belongs to the field of artificial intelligence and human-computer interaction. The method is implemented by an intelligent customer service system, and comprises the following steps: S1, acquiring user data in real time; s2, extracting modal features by using a deep learning network; s3, for the deficiency of modal features, adopting a noise condition fractional network based on a diffusion model for recovery; s4, using the reinforcement learning network to optimize the fused weight corresponding to the modal features; s5, combining the weight of the fusion, and achieving the fusion of the modal features through an attention mechanism; and S6, taking the fused modal features as input, establishing an emotion recognition model by using a deep learning network, and completing an emotion recognition task. The method can effectively solve the problem of data missing in real environment interaction of the intelligent customer service digital person, has high emotion recognition accuracy, provides stable and high-quality service for the user, and has good robustness.
Owner:CHONGQING INST OF GREEN & INTELLIGENT TECH CHINESE ACAD OF SCI

Emotion recognition method and system based on retrieval enhancement cross-modal conditional diffusion model

The invention discloses an emotion recognition method and system based on a retrieval enhancement cross-modal conditional diffusion model. The method comprises the following steps: receiving text, voice and image multi-modal data, detecting modal integrity and extracting existing modal features; if a missing mode exists, similar samples are retrieved in an external sample library based on existing mode features, and retrieval enhancement representation is generated through weighted fusion; a cross-modal condition diffusion model is constructed, and missing modal features are recovered by combining forward diffusion fusion information and a back diffusion and iteration unified cross-modal attention mechanism; and splicing a complete modal feature, and outputting an emotion category and intensity through Transform fusion network and joint loss function optimization. According to the method, the limitation that an existing generation model only depends on local information is broken through, the combination advantage of external retrieval and diffusion generation is fully utilized, the stable performance can be kept in multiple missing modes, and the wide application prospect is achieved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Emotion recognition method and system based on behavior and physiological modality of orthogonal fusion

The invention discloses a behavior and physiological modal emotion recognition method and system based on orthogonal fusion, and the method comprises the steps: obtaining multi-modal data comprising a video frame sequence, an audio signal and an electroencephalogram signal, and carrying out the preprocessing of the multi-modal data; constructing a preliminary emotion recognition network; training the constructed preliminary emotion recognition network by using the preprocessed multi-modal data to obtain a trained emotion recognition model; and obtaining a to-be-recognized video frame sequence, an audio signal and an electroencephalogram signal, and inputting the to-be-recognized video frame sequence, the audio signal and the electroencephalogram signal into the emotion recognition model to obtain a corresponding emotion recognition result. According to the method, a modularized emotion recognition model is constructed, all functional modules cooperate with one another, redundant information is reduced, complementary information of multiple modes is fused, emotion feature characterization accuracy is improved, and therefore the emotion recognition effect is improved.
Owner:NANJING MEDICAL UNIV

Man-machine interaction voice perception method and system based on gradient intelligent dispatch subnet pool

The invention relates to the technical field of voice emotion recognition, in particular to a man-machine interaction voice sensing method and system based on a gradient intelligent calling subnet pool. The method comprises the steps of obtaining an emotion data set; constructing a man-machine interaction voice perception model based on a gradient intelligent dispatching sub-network pool; the system comprises an acoustic clue sensing purification module, a layered acoustic essential coding module, a gradient harmony subnet pool module, a task specific feature extraction module, a focus and confidence joint calibration module, a self-adaptive optimization strategy module and a real-time reasoning and decision fusion module. Carrying out emotion decision making by utilizing the constructed human-computer interaction voice perception model; and outputting a decision result. According to the invention, through the acoustic clue sensing purification module and the layered acoustic essential coding, the problems of emotional information distortion and identity feature confusion caused by real environmental noise are fundamentally solved.
Owner:YANTAI UNIV

Emotion recognition method and device based on multi-modal consensus and diversity decoupling

The invention relates to an emotion recognition method and device based on multi-modal consensus and diversity decoupling. The method comprises the following steps: firstly, collecting multi-modal input data including language, vision and audio signals and carrying out corresponding preprocessing; then, constructing a multi-modal consensus and diversity decoupling emotion recognition model which comprises a multi-modal decoupling coding module, a prototype-Gram unification module, a feature enhancement module, a diversity classification module and an emotion prediction head; then, inputting the preprocessed multi-modal input data into the multi-modal consensus and diversity decoupling emotion recognition model, and performing model training optimization based on a total loss function formed by emotion prediction task loss, decoupling loss, unified target loss and diversity loss; and finally, inputting the multi-modal data to be recognized into the trained multi-modal consensus and diversity decoupling emotion recognition model, and outputting an emotion recognition result. And the accuracy, robustness and interpretability of the multi-modal emotion recognition system are improved.
Owner:SICHUAN UNIV

Multi-modal emotion recognition method based on Mama state space model and cross-modal self-distillation

The invention belongs to the technical field of artificial intelligence and multi-modal emotion calculation, and discloses a multi-modal emotion recognition method based on a Mama state space model and cross-modal self-distillation. Through the organic combination of the efficient sequence modeling capability of the Mamba state space model and the knowledge sharing mechanism of cross-modal self-distillation, the advantages of the state space model in the aspects of time sequence modeling and calculation efficiency are fully played, and meanwhile, the limitation of a single model architecture is made up through a cross-modal attention mechanism; the technical bottlenecks of an existing multi-modal emotion recognition method in the aspects of long sequence processing efficiency, cross-modal information fusion and knowledge transfer sufficiency are effectively solved, and an efficient and reliable technical solution is provided for further development and practical application of the multi-modal emotion recognition technology.
Owner:NORTHEASTERN UNIV CHINA

Electroencephalogram emotion recognition method and system based on multi-task self-supervision and dynamic graph fusion network

The invention belongs to the technical field of artificial intelligence and physiological signal processing, and discloses an electroencephalogram emotion recognition method and system based on a multi-task self-supervision and dynamic graph fusion network, and the method comprises the steps: obtaining and preprocessing electroencephalogram and other physiological signals, and extracting multi-band energy features; constructing a dynamic graph structure, taking electrodes as nodes, taking frequency band energy as characteristics, and fusing spatial distance and functional connectivity to generate a dynamic adjacency matrix; designing multi-task self-supervised pre-training, including spatial jigsaw, frequency jigsaw and cross-modal contrast learning tasks, to learn general characterization; a dynamic graph fusion network is adopted to carry out end-to-end training, and a shared feature extraction module of the dynamic graph fusion network realizes adaptive fusion of multi-modal features by utilizing Chebyshev graph convolution and embedding a cross-modal attention mechanism; the classification module sets an independent classification head for each task, and optimization is carried out through a joint loss function. According to the method, the accuracy and generalization ability of electroencephalogram emotion recognition are remarkably improved.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Voice emotion recognition method based on multiple scales and multiple features

The invention discloses a voice emotion recognition method based on multiple scales and multiple features, and belongs to the technical field of artificial intelligence. The method comprises the following steps: firstly, preprocessing an audio signal and extracting a spectrogram and a Mel-frequency cepstral coefficient; then, a residual network, a bidirectional long-short-term memory network and a HuBERT pre-training model are respectively utilized to extract spectrogram high-order spatial features, time sequence context features and voice semantic embedding features; secondly, inputting the first two features into a multi-dimensional multi-scale feature extraction module to extract richer time-frequency features, performing deep fusion by using a multi-layer cross attention mechanism, and performing weighted fusion with speech semantic embedded features; and finally, all the advanced features are spliced, and a final emotion category is recognized through a full-connection classifier. According to the invention, through combination of multi-scale feature extraction and an advanced fusion mechanism, the problem of insufficient complex emotion modeling ability in the prior art is effectively overcome, and the accuracy and robustness of voice emotion recognition are significantly improved.
Owner:NANJING INST OF TECH

Self-supervised emotion recognition method based on heart-brain joint codebook and related equipment

The embodiment of the invention provides a self-supervised emotion recognition method based on a heart and brain combined codebook and related equipment, and belongs to the technical field of physiological signal processing and artificial intelligence. The method comprises the following steps: respectively defining heart beats of electrocardiosignals and electroencephalogram signal segments with equal lengths as words, and constructing sentences; a shared heart and brain joint codebook is created and trained, and electrocardio and electroencephalogram words are mapped to a unified discrete semantic space through vector quantization so as to learn cross-subject general characterization; then, discretizing a signal sentence by using the codebook, and combining space and position embedding and inputting a Transform encoder to carry out mask pre-training so as to learn context semantics of the signal; and finally, finely tuning the pre-training model for an emotion recognition task. According to the method, deep semantic fusion of heart and brain signals is realized through signal structuring and codebook sharing, dependence on labeled data is effectively overcome, and emotion recognition accuracy and cross-subject generalization ability are remarkably improved.
Owner:SOUTH CHINA UNIV OF TECH

System for latency-aware orchestration and performance optimization in artificial intelligence telephone communication

A system for latency-aware orchestration and performance optimization in AI-driven telephone communication, consisting of: a speech capture unit configured to capture an analog audio signal from a telephone interface and convert the analog audio signal into a digital audio signal stream; a feature extraction unit that is operationally coupled with the speech acquisition unit and is configured to generate a feature representation of the digital audio signal stream through spectral decomposition, noise reduction, and temporal segmentation; an AI inference processor communicatively connected to the feature extraction unit, configured to run one or more AI models for automatic speech recognition, natural language understanding, and emotion recognition on the feature representation to generate intermediate results for inference; a latency orchestration controller coupled to the AI ​​inference processor, wherein the latency orchestration controller is configured to monitor latency across multiple processing stages, predict cumulative delay propagation using a hybrid latency estimation model, and orchestrate the execution scheduling of the AI ​​inference processor based on the predicted latency deviation; a performance optimization unit coupled with the latency orchestration controller and configured to dynamically adjust computational accuracy, inference batch size, and feature processing resolution based on latency thresholds and quality constraints set by the latency orchestration controller; and a transmission synchronization array configured to time-align the processed output generated by the AI ​​inference processor and transmit it to a remote communication node, with the transmission synchronization array maintaining deterministic time coordination between successive packets and the orchestrated inference results.
Owner:CHEEKURI KARTHIK CHAKRAVARTHY DULUTH

Brain heuristic multi-expert multi-modal emotion recognition method and system, equipment and medium

The invention discloses a brain heuristic multi-expert multi-mode emotion recognition method and system, equipment and a medium, and belongs to the technical field of artificial intelligence and biomedical signal processing. The method comprises the following steps: by simulating a brain function partitioning mechanism, dividing an electroencephalogram signal into a plurality of brain regions according to neuroanatomy prior, and designing a special expert network for each region; a global-local double-current encoder is adopted to cooperatively extract spatial-temporal characteristics of each brain region signal, and meanwhile, a multi-scale large-kernel convolution module is utilized to extract peripheral physiological signal characteristics; and finally, dynamically fusing multi-expert features through an adaptive routing network to realize sentiment classification. Expert load balancing and bifurcation regularization joint loss are introduced into the model in training, and effective cooperation and feature diversity of experts are ensured. According to the method, excellent recognition precision is obtained in practice, it is verified through interpretability analysis that the decision-making process conforms to neuroscience cognition, and a high-precision and high-reliability solution is provided for application of brain-computer interfaces, mental health monitoring and the like.
Owner:CHONGQING INST OF GREEN & INTELLIGENT TECH CHINESE ACAD OF SCI

Multi-modal emotion recognition method

The invention belongs to the technical field of multi-modal information processing, and particularly relates to a multi-modal emotion recognition method, which can give consideration to modal generality and modal difference at the same time, realizes adaptive fusion between modals, and is suitable for emotion understanding and analysis of video, voice and text multi-source data. Comprising the following steps: 1, preprocessing a CMU-MOSI data set and a CMU-MOSEI data set to generate training data; 2, processing data by adopting a pre-training encoder to obtain a high-quality initial feature sequence; 3, inputting each modal initial feature into a sharing and private coding unit through a feature decoupling module, and respectively obtaining a cross-modal public emotion sharing feature and a modal specific emotion private feature; according to the multi-modal emotion recognition method provided by the invention, the emotion recognition capability is remarkably improved, and effective data support can be provided for scenes such as psychological analysis, medical assistance and human-computer interaction.
Owner:CHANGCHUN UNIV OF SCI & TECH