Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

77 results about "Expressive communication ability" patented technology

Feature attention and bilinear gating fused speech emotion recognition method and device

The invention discloses a feature attention and bilinear gating fused speech emotion recognition method and device, and the method comprises the following steps: 1, collecting an audio file, obtaining corresponding label information, generating audio waveform and time frequency representation data through preprocessing, and constructing an audio waveform mask and a time frequency mask to mark an effective information region; 2, constructing a dual-path feature extraction module which comprises a time-frequency feature extraction module and a pre-training acoustic feature coding module; wherein the time-frequency feature extraction module models emotion correlation through local convolution and a multi-dimensional attention mechanism, and performs global time sequence modeling based on a bidirectional gated loop network; the pre-training acoustic feature coding module extracts high-level speech representation with high expression ability for emotion distinguishing by using a pre-training model; and step 3, constructing a feature fusion module and an emotion classification module, and combining with a dual-path feature extraction module to form a speech emotion recognition model.
Owner:SICHUAN UNIV

Teaching speech emotion recognition method based on dynamic time sequence modeling and multi-scale fusion

The invention discloses a teaching speech emotion recognition method based on dynamic time sequence modeling and multi-scale fusion. The method comprises the steps of speech data set preprocessing, speech feature extraction, speech emotion classification network construction, speech emotion classification network training, inputting a test set into the trained speech emotion classification network, and outputting the probability of each type of emotion. According to the method, the direction of the information flow is dynamically adjusted through the adaptive time displacement module, the features of different time scales are extracted by using the multi-scale convolution branch, the modeling capability of the model for a complex time sequence structure and variation data is improved, the feature expression is enhanced, the time sequence features are extracted by using the Wav2Vec2.0 pre-training model, and the time sequence features are extracted by using the Wav2Vec2.0 pre-training model. A speech emotion classification network comprising an AdaShiftFormer learning module and a multi-scale time sequence fusion module is constructed, training is carried out in combination with classification loss and comparison loss, and classification performance is optimized. The method is superior to the prior art in emotion classification accuracy and feature expression ability, and can assist in teacher speech behavior analysis and classroom interaction optimization in an intelligent education scene.
Owner:SHAANXI NORMAL UNIV

Voice synthesis method and system based on VITS improvement

The invention provides a voice synthesis method and system based on VITS improvement, and the method comprises the steps: optimizing a text encoder of a VITS model, introducing a large language model, and enabling the emotion, intention and speaking style of an input text to be captured when the text is encoded; a random disturbance item is introduced when the Q value is dynamically planned and solved, the alignment flexibility in the initial training stage is improved, meanwhile, monotonicity constraint is strictly kept, and it is avoided that a suboptimal solution is obtained through convergence too early; a ConvNeXt module is used as a basic backbone network of a decoder, and ISTFT is utilized to efficiently reconstruct a time domain signal, so that waveform up-sampling is realized, redundant calculation of traditional transpose convolution is avoided, and reasoning is accelerated. According to the method, the reasoning speed, the emotion expression ability and the style control flexibility of speech synthesis can be effectively improved, a new solution is provided for cross-language diversified speech synthesis, and a reference is provided for the more efficient and more intelligent development of the speech synthesis technology.
Owner:豫章师范学院

Sound authenticity identification method and system based on multi-modal feature deep interactive fusion

The invention discloses a sound authenticity identification method based on multi-modal feature deep interactive fusion. According to the method, a double-flow architecture is adopted, and a pre-trained BEATs model and a CNN14 network are respectively utilized to extract Transform sequence features and convolution time-frequency embedding features of an audio; an adaptive MobileFormer fusion device is innovatively proposed, an original MobileFormer structure used in the image field is transformed into a bidirectional cross-modal interaction module suitable for one-dimensional time sequence audio features, dynamic complementary modeling of local details and global semantic features is achieved through a cross attention mechanism, and dynamic nonlinearity is introduced to activate and enhance the expression ability; and the fused enhanced features are input into a hierarchical graph attention network ASSIST, and high-order semantic modeling and authenticity classification are completed by combining spectrogram and time sequence double-flow reasoning. Experimental results show that the performance of the method on an ASVspoof2021LA data set is superior to that of an existing baseline model. The method can effectively detect AI generation voice, replay attack and other forged voice, and is suitable for a voice authentication system.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

AI text recognition method and device based on ensemble learning and advanced semantic statistical feature analysis

The invention provides an AI text recognition method and device based on ensemble learning and advanced semantic statistical feature analysis, and the method comprises the steps: 1, respectively sending a to-be-recognized text into a Bert detector and a high-order natural language statistical feature detector for recognition, the high-order natural language statistical feature detector comprises a word logarithm probability detector, a word ranking logarithm detector, an Entropy detector and a confusion degree detector; and 2, performing election on detection results output by the Bert detector and the high-order natural language statistical feature detector by using an election module to obtain an AI text recognition result. According to the method, an integrated learning strategy is adopted, and a pre-training language model subjected to fine tuning is combined with high-order natural language statistical characteristics, so that when the model detects a large language model to generate a text, the strong expression ability of the pre-training language model can be fully utilized, and a deep rule of the text can be captured through the high-order statistical characteristics; and the detection accuracy is improved.
Owner:ZHENGZHOU XINDA ADVANCED TECH RES INST

Multi-mode knowledge representation learning method fusing multi-attention mechanism and semantic enhancement

The invention relates to a multi-modal knowledge representation learning method fusing a multi-attention mechanism and semantic enhancement, and aims to solve the problems that the multi-modal knowledge representation method is insufficient in multi-modal feature fusion and difficult to effectively model complex semantic relationships such as symmetric and anti-symmetric. The method comprises the following steps: firstly, extracting entity image and text features from a multi-modal knowledge graph by using a CLIP model, and extracting entity audio features by using a VGGish model; then image features are enhanced through spatial attention, and image-text feature fusion is realized through cross attention; designing a cross-modal fusion module to dynamically fuse the image-text features and the audio features to obtain multi-modal fusion features; and finally, extracting structured features based on a ComplEx model, integrating the structured features with multi-modal fusion features through an adaptive dual-channel scoring function, and optimizing a training process by adopting a contrast learning loss function. The multi-modal features can be fully fused, the semantic expression ability is enhanced, and the performance of downstream tasks such as intelligent question and answer is improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Incomplete multi-mode dialogue emotion recognition method and system based on speaker and time sequence information joint graph network

The invention discloses an incomplete multi-mode dialogue emotion recognition method and system based on a speaker and time sequence information joint graph network, and the method comprises the steps: obtaining the deep features of a text mode, a voice mode and a visual mode in a dialogue through a feature extraction module, and guaranteeing the high expression capability of the features through a pre-training model; the random mode missing simulation module effectively simulates the data incomplete condition in a real scene, and the robustness of the model is improved; a bi-directional long-short term memory network (Bi-LSTM) is combined with a time sequence diagram network (TGNN) to capture context and time dynamic characteristics of a dialogue, and meanwhile, an interaction relationship between speakers is modeled through a speaker influence matrix, so that joint modeling of a time sequence and speaker information is realized; the deep features are further extracted through the graph convolutional network, and the emotion discrimination of the features is enhanced; finally, the modal reconstruction and emotion classification module significantly improves the accuracy and robustness of incomplete multi-modal dialogue emotion recognition through reconstruction of missing modals and multi-class emotion prediction, and is suitable for man-machine interaction and emotion analysis application in a complex real scene.
Owner:SOUTHEAST UNIV

Multi-modal sentiment classification method and system based on corpus enhancement and cross-modal generation

The invention relates to the technical field of artificial intelligence and emotion recognition, and discloses a multi-modal emotion classification method and system based on corpus enhancement and cross-modal generation, and the method comprises the following steps: 1, carrying out the preprocessing and feature extraction of multi-modal data; wherein the multi-modal data comprises a video modal, a text modal and an audio modal; 2, projecting each modal feature to a low-dimensional shared space; wherein video, text and audio modal features are respectively mapped to a unified low-dimensional space through a full connection layer so as to eliminate dimension differences among modals, and an ReLU activation function is adopted to enhance expression ability; and step 3, performing feature enhancement based on a text library on each input modal feature to output the enhanced modal features. The scheme supports single-mode, dual-mode and three-mode input, adapts to diversity of data missing in practical application, and is high in model recognition accuracy.
Owner:ANHUI UNIV OF SCI & TECH

Method and device for evaluating speech expression ability of neurological and mental diseases

ActiveCN121483566AHealth-index calculationSemantic analysisMicrolinguisticsNeuropsychiatric disease
The invention discloses a method and device for evaluating speech expression ability of neurological and mental diseases. According to the method, speech is induced through a visual stimulation material, speech data are collected, and after speech recognition, semantic calibration and text error correction, evaluation indexes are calculated from multiple dimensions of fluency, microcosmic linguistics and macroscopic linguistics in combination with a natural language processing technology, a large language model and a vector embedding technology, so that the evaluation accuracy is improved. And the result credibility is guaranteed through a robustness verification mechanism of iteration reflection of a double-big language model. According to the method, the problems of large subjective deviation, low efficiency and shallow analysis dimension of traditional evaluation are solved, and objective, accurate, efficient and comprehensive evaluation is realized.
Owner:THE SECOND XIANGYA HOSPITAL OF CENT SOUTH UNIV

Multi-language voiceprint recognition method based on pre-training voice model

The invention discloses a multi-language voiceprint recognition method based on a pre-training voice model. The pre-training voice model WavLM and a traditional voiceprint recognition model ECAPA-TDNN are fused. According to the method, a multi-layer perceptron (MLP) module is introduced for further refining and converting the features extracted by the WavLM, so that the features are more suitable for the input requirement of an ECAPA-TDNN model, and the abstraction and expression ability of the model to the features is enhanced. In the aspect of multilingual voiceprint recognition, the model is finely adjusted by using a small amount of voice data sets, and the method comprises the following basic steps of: firstly, freezing parameters of a pre-trained voice model WavLM, so that the pre-trained voice model WavLM keeps learned knowledge; and then, parameters of the MLP module and the ECAPA-TDNN model are continuously adjusted in training, so that the multilingual voiceprint recognition capability is learned. In the application, a to-be-recognized voice passes through the fusion model to obtain a feature vector, and after the vector is subjected to judgment and decision making, a voiceprint recognition result is obtained.
Owner:BEIJING UNIV OF TECH

Method for realizing personalized singing synthesis model training through multi-mode voice driving

PendingCN120375806ASpeech synthesisPersonalizationModal voice
The invention relates to the technical field of voice signal processing, and discloses a method for realizing personalized singing synthesis model training through multi-modal voice driving, and the method comprises the following steps: obtaining multi-modal input data, including text, reference audio, speaker and emotional features; processing the data to obtain coding features; the inter-modal redundant information is evaluated and suppressed through redundant perception coding; compressing coding features by using an information bottleneck model, and retaining effective information; fusing the compressed features to generate personalized singing features; the input decoding module generates Mel spectrum features; and converting the Mel spectrum into an audio waveform through a vocoder, and outputting the audio waveform. And converting the Mel spectrum into an audio waveform through a vocoder, and outputting the audio waveform. According to the method, multi-modal data can be effectively fused, the individuation and emotion expression ability of the singing sound is improved, redundant information interference is reduced, the quality and efficiency of singing sound synthesis are improved, and finally, high-fidelity individualized singing sound is output.
Owner:SHENZHEN ZHONGLU CULTURE COMMUNICATION CO LTD

Classification method and system for data enhancement and hybrid expert mechanism feature selection

The invention discloses a classification method and system for data enhancement and hybrid expert mechanism feature selection. The system comprises an automatic speech recognition (ASR) module, a text-to-speech synthesis (TTS) module, a multi-modal feature extraction module, a hybrid expert mechanism (MoE) module, a common attention mechanism module, a feature fusion module and a classification module. According to the invention, a voice data enhancement module based on a voice-to-text (TTS) technology is utilized to improve data diversity and model generalization ability; by means of multi-level acoustic and text feature extraction, language changes are represented more comprehensively; a hybrid expert mechanism (MoE) is utilized to realize dynamic selection of multi-modal features, and the feature utilization efficiency is improved; according to the method, the fusion mode between different modal features is optimized by using a co-attention mechanism, the interaction expression ability between the features is enhanced, the recognition precision and the robustness of the system in a multi-modal environment are remarkably improved, and the defects are overcome.
Owner:SHANGHAI JIAOTONG UNIV

Estrus-sharing dialogue generation method and system fusing psychological information

The invention discloses an emotional dialogue generation method and system fused with psychological information, and the emotional expression capability of generating response can be improved through the obtained global emotional latent variable fused with a speaker and a listener. By determining a multi-hop knowledge reasoning path set, obtaining context semantic representation of each knowledge reasoning path, and fusing the context semantic representation with dialogue history global context representation to obtain knowledge-enhanced context representation, the knowledge reasoning ability and semantic relevance in a dialogue context are improved; by constructing a knowledge representation matrix, taking a global emotion latent variable as a bias item, and dynamically adjusting the knowledge representation matrix and a value vector generated by knowledge-enhanced context representation through semantic features, multi-scale attention fusion representation is obtained, and the context adaptive capacity and emotion representation naturalness of generated response are effectively improved. According to the method, the psychological state of the user can be identified more accurately, and the dialogue response with humanization, estrus-sharing expression and interaction adaptability is generated.
Owner:ZHEJIANG NORMAL UNIV

Multi-modal dialogue emotion recognition method based on uncertainty adaptive weighting

The invention belongs to the technical field of multi-modal emotion recognition, and particularly relates to a multi-modal dialogue emotion recognition method based on uncertainty adaptive weighting. The method comprises the following steps: processing original modal information of each modal to obtain an alignment feature vector of each modal; calculating a cross-modal fusion feature of each cross-modal path based on the alignment feature vector of each modal; obtaining an uncertainty score of each cross-modal path according to the cross-modal fusion features of each path; obtaining the optimized weight of each path based on the initial weight and the uncertainty score of each path; and obtaining a predicted emotion category based on the optimized weights of all the paths and the cross-modal fusion features. According to the method, the expression ability of the non-text mode in the key emotion scene is remarkably enhanced while the semantic advantages of the text are kept.
Owner:ZHEJIANG UNIV OF TECH

Method, device and equipment for training emotional speech synthesis model and storage medium

This invention relates to the field of artificial intelligence technology and discloses a training method, apparatus, device, and storage medium for an emotional speech synthesis model, which can be applied to intelligent voice dialogue scenarios in finance, insurance, and medical fields. This invention selects at least one layer of a pre-trained speech synthesis model as the target layer, loads a VB-LoRA fine-tuning module onto the target layer, and then fine-tunes it using emotional speech data. This enables the model to achieve emotional speech synthesis and output. No emotional information is added during the training phase; emotional information is only added during fine-tuning. This allows for fine-tuning by adding emotional information of different emotion categories, giving the model the ability to express different emotion categories, thus enhancing the model's scalability and flexibility. Furthermore, during fine-tuning, only the parameters of the target layer with the VB-LoRA fine-tuning module are adjusted, eliminating the need for full parameter fine-tuning of the entire model, reducing the workload of model fine-tuning and lowering computational costs.
Owner:PING AN TECH (SHENZHEN) CO LTD

Natural emotional singing method and emotional singing audio processing method

PendingCN122347963AAcoustic areaMusic class
The application discloses a natural emotional singing method and an emotional singing audio processing method, and belongs to the singing information field, and aims to solve the problems of fragmented training elements and single feedback system of a traditional vocal music teaching method. The technical scheme comprises six training modules combined by action guidance, emotion awakening and sound area training, and a training system for improving singing emotional expression capability based on a deep learning model for real-time extraction of user singing audio emotional indexes and generation of dynamic visual feedback.
Owner:BEIJING FABIAN EDUCATION TECH CO LTD

Voice keyword recognition method based on CNN-LSTM and online knowledge distillation

The invention discloses a voice keyword recognition method based on CNN-LSTM and online knowledge distillation, and relates to the technical field of voice processing. According to the method, an improved contraction residual attention module is introduced into a CNN-LSTM-based neural network model and is used for discovering and inhibiting redundant information and noise in features, and the expression ability of the features is enhanced. According to the method, a training method based on online knowledge distillation is introduced, the current training iteration is supervised by using information in the previous two training iterations, and an extra regularization effect is provided for the training process by using the difference in prediction probabilities output by a neural network model. The overfitting risk of the neural network model is reduced, and the generalization ability of the neural network model is enhanced.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Voice generation method and device based on preference alignment, equipment and medium

The invention relates to the technical field of speech synthesis, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a speech generation method and device based on preference alignment, equipment and a medium, and the method comprises the steps: obtaining a pre-training speech generation model and a preference training sample pair, constructing a preference alignment model and a non-preference alignment model, and taking the pre-training model as a reference; determining a loss function based on the preference samples and the non-preference samples, and updating model parameters to obtain a trained model; and receiving a target condition item and generating an agent prompt in a reasoning stage, fusing speed prediction results of the two types of models, and generating a target voice based on a fusion result in a stream matching process. According to the method, human preference signals are fused through a preference alignment mechanism, so that the naturalness and semantic consistency are considered in the speech generation process, and the speech quality and the personalized expression ability are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Open vocabulary point supervision time sequence action positioning method based on progressive optimization network

The invention discloses an open vocabulary point supervision time sequence action positioning method based on a progressive optimization network, and belongs to the field of video understanding. Firstly, visual features are preliminarily screened through semantic guidance early enhancement, so that background noise interference is inhibited, and the expression ability of category-related features is improved. And then, semantic association of modeling actions in different time periods is further established by utilizing context-semantic post-enhancement, the category identification degree of visual expression is enhanced, and the model is promoted to mine more unconstrained new category suggestions. And finally, the action boundary is optimized in combination with pseudo-label constraint, fine modeling of the boundary position is realized, and the generalization ability of the model is enhanced while the detection precision is improved. According to the method, a progressive optimization modeling strategy is combined, visual and semantic features are fused to construct a unified progressive optimization network, and the method aims at improving the action positioning performance while reducing mark dependence. The method has a wide application prospect in tasks such as intelligent monitoring, abnormal behavior detection and video question and answer.
Owner:BEIJING UNIV OF TECH

Sign language video generation method based on improved Transformer model

The present invention provides a method and device for generating sign language videos based on an improved Transformer model. The method of the present invention first extracts the skeletal posture sequence in the sign language video and removes redundant information to reduce the amount of calculation. In addition, considering the importance of spatiotemporal information to the accuracy of generating sign language videos, a semantically rich embedding module is designed to encode position and speed information into the same high-dimensional space as the input of the model, thereby improving the coordination of joint movements and improving the accuracy of feature representation. Finally, an encoder-decoder model with a pyramid structure is constructed. The encoder accepts a spoken sentence as input and encodes the information in the sequence into an intermediate representation. The decoder then decodes the intermediate representation into a target sign language posture sequence in a semi-autoregressive manner. The present invention can effectively improve the utilization rate of semantic information and the overall expression ability of movements, thereby significantly improving the accuracy and speed of sign language video generation.
Owner:HEBEI UNIVERSITY

Augmented reality content recommendation and expressive presentation method based on user behavior perception

ActiveCN121685902BImprove interactive experienceImprove the effect of information transmissionInformation transmissionThree-dimensional space
The application provides a kind of based on user behavior perception's augmented reality content recommendation and expressive presentation method, including constructing the state representation of multiple virtual roles in augmented reality scene;Joint objective function for evaluating virtual role layout rationality is constructed;According to the attribute information of virtual role, the adaptive weight is generated by the weight prediction model constructed in advance, and the adaptive weight is used to adjust the cost item related to the importance of role in the space cost function;Based on the joint objective function, the position and orientation of all virtual roles are iteratively optimized to obtain an optimized layout;Expressive presentation is carried out in the augmented reality scene The optimized layout is realized in three-dimensional space Automatic optimization of role position and orientation, avoid visual conflict, improve the naturalness, coordination and semantic expression ability of layout, thereby significantly enhance the interactive experience and information transmission effect of AR system.
Owner:BEIJING TECH & BUSINESS UNIV

Empathy-Based Wearable Systems for Perceptual and Communicative Assistance

PendingUS20260179644A1Speech recognitionHeadphonesEmpathy
The invention provides a unified assistive technology system, the ADA Empathy Wearable Bundle, enhancing perception, communication, and self-expression for individuals with visual, auditory, or speech impairments.It integrates Empathy Glasses, Empathy Headphones, Word Articulation Guidance (WAGS), Word Articulation Response Monitoring (W.A.R.M.), and Conversational MIDI (C-MIDI), using AI-driven narrative interpretation, augmented reality, and blockchain-based security.The system delivers vivid narrative audio for blind users, multimodal AR overlays for hearing-impaired users, and real-time visual and auditory articulation guidance for speech-impaired or non-verbal users. A tone-based trust protocol ensures operational integrity, while blockchain storage preserves privacy and consent.Grounded in ethical AI, empathy, and human dignity, the system enables users to perceive, understand, and engage with the world fully, offering a holistic, compassionate, and transformative solution in assistive technology.
Owner:BOWES DEAN

Toy (phonograph)

ActiveCN309423469SPhonographTesting Methods
1. Name of the product of this design: Toy (phonograph). 2. Purpose of this product: Through repeated practice with the repeat function of the phonograph, children's understanding of different sentence patterns can be strengthened and deepened, and their ability to express themselves can be cultivated. 3. The key point of the design of this product lies in its shape. 4. The picture or photo that best illustrates the design points: Stereoscopic drawing 1.
Owner:BEIJING XINTANG SICHUANG EDUCATIONAL TECH CO LTD

A sign language sentiment recognition and teaching feedback method based on machine learning

PendingCN122454639AComputer aided instructionComputer-aided
The application relates to the technical field of artificial intelligence, affective computing and computer-assisted teaching, in particular to a sign language emotion recognition and teaching feedback method based on machine learning, which collects learner sign language videos and extracts face, hand and posture holographic key points; action semantic features and emotion expression features are obtained in parallel by using a double-branch time sequence coding network; the sign language content and the emotion state containing continuous values of valence-arousal-dominance and discrete categories are obtained by a semantic recognition subnetwork and an emotion recognition subnetwork respectively; the recognition result is compared with standard semantics and emotion labels, and emotion intensity, naturalness, semantic matching degree scores and comprehensive quality scores are calculated; emotion correction instructions, reinforcement learning training sequence planning and key point trajectory visualization feedback are generated based on the scores; and teaching strategies are iteratively improved through closed-loop optimization and emotion resonance models. The application realizes synchronous recognition and quantitative evaluation of sign language semantics and emotions, and provides an adaptive feedback means for sign language emotion expression ability training.
Owner:杨莉红

Intelligent interactive writing system and method based on emotion adjustment

The invention discloses an intelligent interactive writing system and method based on emotion regulation. The intelligent interactive writing system based on emotion adjustment comprises a user registration and initialization module, a front end, a multi-mode large language model, a surface emotion recognition module, a core emotion information unlocking module, a staged emotion exploration module and a wound scene reproduction and guide emotion adjustment module. An important others interactive dialogue module; and a data storage and retrieval module. According to the method, various emotion adjusting technologies are integrated, various psychological emotion adjusting technologies such as writing, dialogue emotion adjusting and role playing are combined, an emotion vocabulary library is provided, the emotion expression ability of the user is enriched, the user is helped to deal with past psychological trauma through trauma scene writing, and unmet emotion requirements are met through important others' dialogues. According to the invention, a three-layer architecture design is adopted to ensure the stability and expansibility of the system; the responsibilities of the client and the server are clearly separated, and the data transmission efficiency is optimized.
Owner:上海丽瑗健康管理咨询中心

Coaching method, electronic equipment and computer readable storage medium

The invention discloses a tutoring method, electronic equipment and a computer readable storage medium, and the method comprises the steps: obtaining an emotional state of a user under the condition that first dialogue content of the user is received, and the first dialogue content comprises a tutoring demand of the user; based on the first dialogue content and the emotional state, a dialogue content generation model is used to generate second dialogue content responding to the first dialogue content, and the second dialogue content comprises copywriting content and dialogue mood; and outputting the second dialogue content so as to tutorize the user. By means of the scheme, the emotion expression ability in AI dialogue tutoring can be enhanced, the mechanical feeling of AI dialogue tutoring is effectively improved, and the naturalness and reality sense of the interaction process are improved.
Owner:WANGYIYOUDAO INFORMATION TECH BEIJING CO LTD

Key information generation method and system, electronic equipment, medium and product

The invention relates to a key information generation method and system, electronic equipment, a medium and a product. The key information generation method comprises the steps of obtaining a to-be-processed voice signal; preprocessing the voice signal to obtain a tensor feature corresponding to the voice signal; performing fusion processing on the tensor features in the channel dimension and the time dimension to obtain first fusion features corresponding to the voice signals; based on the first fusion feature, determining a semantic feature corresponding to the voice signal; using an external attention module of the trained information generation model to weight a common part of the semantic feature corresponding to the voice signal to obtain a common feature corresponding to the voice signal; a self-attention module of the information generation model is utilized to weight the personality part of the semantic feature corresponding to the voice signal, and a personality feature corresponding to the voice signal is obtained; and generating key information corresponding to the voice signal based on the common characteristics and the personality characteristics. The semantic expression capability of the features can be improved, so that the accuracy of generating the key information is improved.
Owner:CHONGQING CHANGAN AUTOMOBILE CO LTD

An emotion dialogue generation method based on an improved generative adversarial network

The emotion dialogue generation method based on improved generative adversarial network belongs to the dialogue system field under natural language processing, realizes the interaction between memory matrices through a self-attention mechanism, enhances the long-distance transmission energy between information, thereby improves the expression ability and feature extraction ability of the model, and solves the problem that the existing model has weak expression ability and generates short sentences. Meanwhile, a multi-class discriminator is set, which respectively discriminates the true and false of the text and the emotion category, calculates the category relative loss and the emotion information loss two parts to feed back and update the generator, so as to improve the consistency of the emotion information, and make the emotion expression of the reply generation sentence more obvious and clear. Finally, the iterative evolution algorithm is used in the generator update, the temperature parameter and the quality parameter are controlled respectively, and the child generator in the optimal direction of the evolution temperature is selected to complete the generation task. The dialogue generation method of the present application realizes the emotion embedding which considers the quality and diversity of the reply.
Owner:BEIJING UNIV OF TECH

Personalized synthesis and recognition enhancement method for dysarthria speech

ActiveCN120412540ASpeech synthesisDysarthric speechSpeech synthesis
The invention discloses a personalized synthesis and recognition enhancement method for dysarthria speech, a dysarthria speech synthesis model comprises a long-range dependence feature coding module, an unsteady feature coding module and a decoding module, the input of the dysarthria speech synthesis model comprises samples, and the output comprises synthesis of dysarthria speech. The sample is a dysarthria text sequence; the input of the long-range dependency feature coding module comprises a sample, and the output is an alignment vector z; the input of the unsteady state feature coding module comprises an alignment vector z, and the output of the unsteady state feature coding module is final embedding representation # imgabs0. The input of the unsteady state feature coding module is final embedding representation # imgabs1. The output of the unsteady state feature coding module is synthetic dysarthria voice. According to the dysarthria speech synthesis model, the ability of extracting personalized characteristics of the dysarthria speech, the speech synthesis performance and the refined expression ability of the characteristics of the dysarthria speech are improved.
Owner:TIANJIN UNIV

Emotion recognition reasoning modeling method and device, storage medium and program product

The invention provides an emotion recognition reasoning modeling method and device, a storage medium and a program product, and relates to the technical field of brain-like calculation and multi-modal intelligent perception crossing. The emotion recognition reasoning modeling method comprises the following steps: constructing a modal perception encoder, and encoding original multi-modal input including texts, voices and images into a pulse time sequence in a unified format; constructing a modal fusion module, in a pulse self-attention mechanism embedded with SNN logic, using LIF integral to obtain Spike features, and retaining event-driven features and a long-distance dependency relationship between modeling modals to obtain fusion features; and constructing an emotion predictor, performing statistical processing on the fusion features, realizing category mapping, and then obtaining emotion state output. According to the method, a structured multi-modal fusion mechanism is realized, interaction between different modals can be completed in a pulse stage, shallow fusion which is only spliced in a feature layer is avoided, and the collaborative expression ability between the modals is improved.
Owner:HUA DATA TECH (SHANGHAI) CO LTD