Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

33 results about "Speech Acoustics" patented technology

The acoustic aspects of speech in terms of frequency, intensity, and time.

Method, device, storage medium and electronic device for generating virtual voice

The present invention discloses a method, device, storage medium, and electronic device for generating virtual speech. The method comprises: obtaining multiple different speech text samples and speech attribute information, wherein each speech text sample in multiple different language speech text samples corresponds to a language and an object; inputting each speech text sample into a multi-stream encoder to obtain text features corresponding to each speech text sample; and training a preset speech acoustic model based on generative adversarial network modeling using the text features and speech features to obtain a target acoustic model for generating virtual speech. The present invention can support cross-language data training and the generation of cross-language speakers. The multi-stream encoder can better capture text features in different languages, improve the flexibility and reliability of virtual preset generation, and thus solve the technical problem of low flexibility and reliability in generating virtual speech in the prior art.
Owner:BEIJING UNISOUND INFORMATION TECH CO LTD

Traditional folk song singing voice acoustics and physiology multi-modal analysis method and system

The invention discloses a traditional folk song singing voice acoustics and physiology multi-modal analysis method and system, and the method comprises the steps: carrying out the synchronous calibration of a voice recording device, a voice monitoring device and a respiration sensor through a unified clock source, obtaining a multi-modal signal data flow, and obtaining a time-aligned original data set; according to the original data set subjected to time alignment, Fourier transform is adopted to process the voice signal part, acoustic feature vectors are determined, and a physiological signal sequence is extracted from the original data set subjected to time alignment; if the acoustic feature vectors are matched with the timestamps of the physiological signal sequences, Pearson coefficients between the acoustic feature vectors and the timestamps are calculated through correlation analysis, and quantitative correlation strength is judged; for the part of which the quantitative correlation strength is higher than a fourth preset threshold value, fitting a mapping relationship between the acoustic characteristics and the physiological mechanism by adopting a linear regression model to obtain a parameterized singing technique model; and according to the parameterized singing technique model, analyzing the relationship between traditional singing voice acoustics and physiological multiple modes.
Owner:NORTHWEST UNIVERSITY FOR NATIONALITIES

Computer-aided senile language erosion assessment method and assessment system

The invention discloses a computer-aided old-age language erosion assessment method and assessment system, and relates to the technical field of old-age health assessment, and the method comprises the following steps: S1, collecting the voice data, text input data and interactive behavior data of an old-age user; s2, preprocessing the data, and extracting voice acoustic features, language structure features and cognitive behavior features; s3, inputting the extracted features into a pre-trained language erosion evaluation model, and outputting a language ability score and an erosion type classification result; and S4, generating a visual evaluation report, wherein the visual evaluation report comprises the language ability degradation degree, key obstacle points and intervention suggestions. According to the method, the score and classification result is automatically output through the multi-modal data acquisition and pre-trained deep learning model, so that the evaluation time is greatly shortened, the subjective deviation is eliminated, the result objectivity is ensured, the problems of low manual evaluation efficiency and high subjectivity are solved, and the effect of automatic evaluation is realized.
Owner:BEIJING FOREIGN STUDIES UNIVERSITY

Multi-modal voice interaction large model training method and system based on voice acoustic feature regulation and control, terminal equipment and medium

ActiveCN120954388ASpeech recognitionSpeech synthesisSpeech comprehensionModal voice
The invention discloses a multi-modal voice interaction large model training method and system based on voice acoustic feature regulation and control, terminal equipment and a medium, and relates to the technical field of multi-modal voice interaction.The method comprises the steps that a text token of a text training sample is obtained, a corresponding voice token is constructed, and pre-training data used for converting the text token into the voice token is obtained; in combination with the multi-modal input sample and the pre-training data, constructing fine-tuning training data for speech understanding and dialogue generation; constructing and pre-training a basic model by using the pre-training data; a multi-modal voice interaction large model is constructed based on a pre-training basic model, and fine tuning data is used for training, so that voice acoustic features can be regulated and controlled based on multi-modal input, and voice is output. According to the method, through alignment and staged training of the text token and the voice token, fine regulation and control of voice acoustic characteristics are realized, long voice continuity and interaction naturalness are improved, and a model is efficiently endowed with voice interaction capability of controllable timbre and emotion.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

WIFI-based indoor voice positioning method and device

The application discloses a WIFI-based indoor voice positioning method and device. The method utilizes existing terminal equipment and is based on voice acoustic principles. The corresponding terminal equipment can be controlled to emit and receive voice information in a WIFI communication mode, so that the position of the terminal equipment is determined, and the positioning of a sound source is completed. The indoor positioning operation is simplified, the cost of system deployment is reduced, the accuracy of indoor positioning results is increased, and the convenience of indoor positioning in application is improved.
Owner:FOSHAN VIOMI ELECTRICAL TECH

Adversarial network optimization method and system for short-utterance speaker verification

ActiveCN114530156BSpeech analysisPattern recognitionSpeaker verification
The embodiment of the specification provides a generative adversarial network optimization method and system for short speech speaker verification, wherein the method comprises the following steps: acquiring a plurality of pairs of long and short speech acoustic feature samples; inputting the short speech acoustic feature sample into a generator for splicing to obtain a generated pseudo long speech acoustic feature sample; inputting the pseudo long speech acoustic feature sample and the acquired long speech acoustic feature sample into a speaker verification model respectively, outputting a pseudo identity feature sample and a true identity feature sample through the speaker verification model; inputting the true identity feature sample and the pseudo identity feature sample into a discriminator and a classifier, calculating the loss of the discriminator and the classifier through a loss function, and updating the parameters of the discriminator, the classifier and the generator through back propagation optimization. The problem that the discrimination effect of a speaker verification system becomes poor as the speech duration becomes shorter is solved.
Owner:STATE GRID CORPORATION OF CHINA +1

Intelligent interactive door control system

The invention discloses an intelligent interactive door control system, and relates to the technical field of door control systems. An image acquisition module is used for acquiring image information in front of a door in real time; the space processing module is used for processing and analyzing the collected images and human body movement tracks, automatically identifying the identity of the visitor and judging behaviors; the voice interaction module is used for generating a voice conversation with a visitor; the control module is used for calling the voice interaction module to communicate with the visitor according to the recognition result of the space processing module; according to the invention, active and intelligent noise suppression is realized, and the interaction reliability in a severe acoustic environment is ensured. The core of the system is to break the barriers of vision and hearing, and the system can understand the meaning of speech by performing feature-level fusion on the visual information and the voice acoustic features, so as to make distinct and highly personalized responses. By bypassing strong dependence on a final text recognition result, fusion is creatively carried out on a feature level.
Owner:BEIJING QIREN TECHNOLOGY CO LTD

Anomaly processing method and device based on multi-dimensional classifier chain, equipment and medium

This invention relates to the field of intelligent decision-making technology and can be applied to business scenarios such as fintech and healthcare. It discloses an anomaly handling method, apparatus, device, and medium based on a multi-dimensional classifier chain, comprising: acquiring a voice data stream and generating a text data sequence with role identifiers, while simultaneously extracting voice acoustic features; performing semantic parsing on the text data sequence and determining dialogue structure features; fusing three types of features to form a unified input feature vector; performing chain-like prediction based on label dependencies using a multi-dimensional classifier chain model and outputting multi-dimensional classification prediction results; determining the anomaly level based on the prediction results and matching the target strategy to execute the corresponding action. This invention improves the accuracy of multi-dimensional prediction by fusing multi-source features and utilizing a classifier chain model to model label dependencies, thereby more reliably identifying anomaly levels and triggering matching strategies, achieving more accurate customer service analysis and recovery processing.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Hearing-impaired voice conversion and generation method based on multi-language self-adaption

PendingCN121483225ABiological modelsSpeech recognitionIntelligibility (communication)Speech Acoustics
The invention relates to a hearing-impaired voice conversion and generation method based on multi-language self-adaption, and belongs to the field of voice processing and artificial intelligence. According to the invention, hearing-impaired voices in multiple languages are converted into standard voices, so that normal communication with hearing-impaired people is facilitated. According to the method, acoustic features of hearing-impaired speech are extracted and aligned with target language phonemes, the acoustic features are input into a speech adaptation model, standard speech features are generated, multi-language training is carried out based on the standard speech features and target language corpus, and the training result is used for generating target language natural speech. According to the invention, the intelligibility of hearing-impaired voice can be improved, so that the voice is closer to natural pronunciation. According to the invention, rapid adaptation of multiple languages can be realized, and the cost of cross-language training is reduced. The system can be applied to education, medical rehabilitation, cross-border communication and other scenes, and provides effective voice communication assistance for hearing-impaired people.
Owner:BEIJING INST OF COMP TECH & APPL

Smart home user behavior prediction method based on semantic analysis

The invention discloses a smart home user behavior prediction method based on semantic analysis, and relates to the technical field of natural language processing. The method comprises the following specific steps: data acquisition: respectively acquiring user interaction data, equipment operation historical data and user interaction feedback data through a voice acquisition device, a user input interface, a smart home gateway and a feedback interface, and uniformly storing the data after formatting processing; the speech acoustic features and the text features are extracted on the emotional semantic recognition level, the pre-training model is used for outputting the user emotional state, the defect that in the prior art, emotional perception is lacked is overcome, in the semantic ambiguity resolution aspect, a learning mechanism with user interaction feedback data as an optimization signal is adopted, and the user experience is improved. When the same number of times of correction of the same fuzzy instruction by the user reaches a preset threshold value, the semantic mapping rule is automatically generated and written into the rule base, and the semantic analysis priority can be adjusted according to the family structure change.
Owner:ZHEJIANG UNIV OF FINANCE & ECONOMICS

Voice conversion method, device, computer equipment, storage medium and program product

This application relates to a speech conversion method, apparatus, computer device, storage medium, and program product. The method includes: obtaining the speaker's body language and target speech information, where the target speech information represents speech information emitted through a specific vocalization state; determining the speaker's emotional state at the time the target speech information was emitted based on the body language; and performing speech conversion processing on the target speech information using an emotional speech acoustic model corresponding to the emotional state to obtain emotional speech information corresponding to the target speech information, where the emotional speech information represents speech conveying the emotional state. This method enables the recipient to correctly understand the meaning of the speaker's whispered expression.
Owner:YOUME TECH (SHENZHEN) CO LTD

Model training method, speech recognition method, device, medium, and program product

ActiveCN121565156BSpeech recognitionSpeech AcousticsAutomatic speech
This invention discloses a model training method, a speech recognition method, a device, a medium, and a program product. The method includes: decoupling sample speech information by an acoustic feature encoder to determine sample speech content features and sample speech acoustic style features; performing acoustic feature discrimination on the sample speech content features using an acoustic feature discriminator to determine adversarial loss; fusing the sample speech content features and sample speech acoustic style features using a feature fusion module to determine fused sample speech features; extracting text information from the fused sample speech features using an automatic speech recognition decoder, and determining the automatic speech recognition loss based on the extracted text and speech recognition training samples; and substituting the adversarial loss and the automatic speech recognition loss into a pre-constructed total loss function to complete the training of the speech recognition model. This enables the speech recognition model to recognize speech data acquired with different accents and in different environments.
Owner:SUZHOU KEDA TECH +1

A speech speaker conversion point detection method and a model training method thereof

PendingCN122369471APattern recognitionNoise
This invention discloses a method for detecting speaker transition points and its model training method. The detection method includes: extracting acoustic features of the speech, inputting them into a pre-trained multi-task coupled learning model to obtain a probability sequence of transition points, and determining the location of the transition points through post-processing. The innovation of this model lies in: employing a shared speaker embedding extraction network and introducing a frame-by-frame feature gating unit in its speaker transition point regression branch to suppress noise interference using speech activity probability; simultaneously, optimizing the model training using a regression loss based on bell-shaped distribution soft labels to achieve refined boundary localization. The corresponding training method clarifies the model construction, soft label generation, and multi-task joint training process. This invention solves the problems of low accuracy in speaker transition point detection, sensitivity to silence and noise, and inaccurate boundary localization in existing technologies, effectively improving the accuracy and robustness of detection.
Owner:JIANGSU UNIV OF SCI & TECH

A system and method for predicting the risk of alzheimer's disease

PendingCN122436204ABlood biomarkersData acquisition
The application discloses an Alzheimer's disease risk prediction system and method, and belongs to the field of biomedical detection and artificial intelligence technology, and the system comprises: a multi-modal data acquisition module for acquiring biological samples, images, voices and demographic information; a biomarker detection module for detecting the expression levels of p-tau217 and GFAP; an image analysis module for extracting fat metabolism parameters in PET / CT images; a voice feature extraction module for extracting acoustic features; and a risk prediction module with an integrated prediction model combining multiple machine learning models, which takes the above multi-modal data as input variables, comprehensively analyzes and outputs risk probabilities. The application combines blood biomarkers, image metabolism parameters, voice acoustic features and demographic information, and through multi-dimensional data complementation and correction, the accuracy, convenience and reliability of early risk prediction of Alzheimer's disease are significantly improved.
Owner:GUANGZHOU NANFANG COLLEGE

Homophone error correction method and device and storage medium

The invention provides a homophone error correction method and device and a storage medium, and belongs to the technical field of character error correction, and the method comprises the steps: importing to-be-corrected voice data, original voice data and real text data; performing voice recognition on the to-be-corrected voice data and the original voice data to obtain to-be-corrected text features and original text features; constructing a training model, and performing model analysis on the training model according to the original text features and the real text data to obtain a homophone error correction model; and performing error correction analysis on the to-be-corrected text features through the homophone error correction model to obtain a homophone error correction result. The method improves the context association capability, does not need to depend on shallow text matching, improves the adaptability of a low-resource scene and a dynamic scene, fully considers the association of voice acoustic features, improves the error-tolerant rate of a rare homophone combination, and also can remarkably improve the homophone error correction accuracy.
Owner:SI-TECH INFORMATION TECH CO LTD +1

Virtual companion interaction method and system based on multi-modal interaction memory graph

PendingCN122633028AEngineeringSpeech Acoustics
The application discloses a virtual companion interaction method and system based on a multi-modal interaction memory graph, and relates to the technical field of artificial intelligence and human-computer interaction. The method extracts text semantics, speech acoustics and environmental visual features of user interaction through a multi-modal fusion neural network, constructs a multi-modal interaction memory graph with a sentiment weight and a decay factor, supports user configuration and correction of memory parameters, calculates an intimacy score based on six-dimensional data and triggers a virtual companion persona phase transition, intelligently judges whether to initiate active interaction based on an unfinished high-weight memory, real-time context and configuration and correction rules, and synchronously adjusts multi-modal performance parameters according to the intimacy. The application can realize gradual emotional relationship development, and significantly improves the naturalness, immersion and emotional connection depth of virtual companion interaction.
Owner:HANGZHOU ARK OF HOPE NETWORK TECHNOLOGY CO LTD

Streaming end-to-end speech recognition method, apparatus and electronic device

A method, an apparatus, and an electronic device for streaming end-to-end speech recognition are described. The method includes: extracting and encoding speech acoustic features of a received voice stream in units of frames; performing block processing, and predicting a number of activation points included in a same block that need to be encoded and outputted; determining position(s) of activation point(s) that need(s) to be decoded and outputted according to a prediction result, to a decoder to perform decoding at the position(s) of the activation point(s) and output a recognition result. Through the embodiments of the present disclosure, the robustness of a streaming end-to-end speech recognition system to noise can be improved, thereby improving the performance and the accuracy of the system.
Owner:ALIBABA GROUP HOLDING LTD

Mental electrophysiology assessment method and system based on video follow-up visit system

The invention relates to the technical field of mental health assessment, in particular to a mental electrophysiology assessment method based on a video follow-up visit system, which comprises the following steps of: acquiring doctor-patient double-channel audio and video data through the video follow-up visit system; preprocessing the collected data and extracting multi-modal features including facial behavior features, voice acoustic features and motion dynamics features; and constructing a spatial model of the mental pathological state of the patient based on the extracted multi-modal features, wherein the spatial model describes a dynamic evolution process of a multi-dimensional state variable through a stochastic differential equation. According to the method, a multi-modal feature extraction system of doctor-patient dual-channel audio and video data is constructed, and then a psychiatric pathology state space model based on a stochastic differential equation is established, so that the problems that subjective scale and static observation are mostly adopted in a traditional mental assessment method, and due to lack of objective quantitative indexes and a dynamic tracking mechanism, the accuracy of mental assessment is poor are solved. Therefore, the evaluation result is greatly influenced by subjective experience of doctors, and the dynamic change of symptoms cannot be captured.
Owner:河南医药大学第二附属医院(河南省精神病医院)

Text-to-voice method and device, computer equipment and storage medium

The invention relates to the field of artificial intelligence, is applied to financial and medical scenes, and discloses a text-to-voice method and device, computer equipment and a storage medium, and the method comprises the steps: receiving a target text, a voice prompt and an emotion label; performing semantic coding on the target text to extract text semantic features; performing acoustic coding on the voice prompt to extract voice acoustic features; based on the emotion label and the text semantic feature, generating a dynamic emotion feature through a time sequence emotion model; performing cross-modal fusion on the voice acoustic features and the dynamic emotion features to obtain cross-modal alignment features, and inputting the cross-modal alignment features and the text semantic features into a pre-training language model for feature fusion and voice token prediction to obtain an initial voice token sequence; and performing semantic alignment and spectrum conversion on the initial voice token sequence to generate a target voice waveform. According to the invention, high-quality voice with natural emotional circulation, controllable timbre and high conformity with professional contexts can be synthesized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent multi-mode voice conference interaction method and system

The invention relates to the field of voice interaction, in particular to an intelligent multi-mode voice conference interaction method and system. The method comprises the following steps: collecting historical voice signal data, carrying out feature extraction and analysis on original voice signal data, generating a voice acoustic vector, analyzing the voice acoustic vector, and constructing a voice baseline vector; collecting conference voice data, analyzing the conference voice data, and constructing a conference semantic state diagram; acquiring real-time voice signal data, and analyzing the real-time voice signal data based on the voice baseline vector to obtain a voice deviation parameter; and on the basis of the voice deviation parameter and the voice baseline vector, adjusting voice recognition adaptation, generating an adjustment strategy, stopping recording the conference semantic state diagram before conference interruption when the conference interruption is detected, and performing reconnection processing on the conference semantic state diagram after conference reconnection. According to the invention, the voice recognition suitability and the semantic understanding continuity can be improved.
Owner:JIAN XIANGE ACOUSTIC ELECTRONIC CO LTD

Music generation method, electronic device, storage medium and computer program product

The embodiment of the invention discloses a music generation method, electronic equipment, a storage medium and a computer program product. The method comprises the following steps: performing feature extraction on voice information of a target object to obtain voice semantic features and voice acoustic features of the target object; generating a human voice audio based on the voice semantic features and the voice acoustic features; based on the voice audio, generating a composing and music-making audio; and generating target music based on the voice audio and the arrangement music audio.
Owner:MIGU CO LTD +1

Information processing device, information processing method, and program

There is provided an information processing device, an information processing method, and a program that can provide a user experience with a further improved sense of reality. When a second avatar associated with a second user is present in a scene or a plurality of areas associated with a virtual space in which a first avatar associated with a first user is present, a voice acquisition unit acquires a voice of the second user, an acoustic environment determination processing unit performs acoustic environment determination processing of determining an acoustic environment of the scene or the areas in which the first avatar is present based on a collider associated with the scene or the areas, and an acoustic characteristics application unit applies acoustic characteristics matching a processing result of the acoustic environment determination processing to the voice of the second user. The present technology can be applied to, for example, a system that provides a metaverse virtual space.
Owner:SONY GROUP CORP

Speech synthesis method and electronic equipment

PendingCN121938342ASpeech synthesisSynthesis methodsSpeech Acoustics
The invention provides a speech synthesis method and electronic equipment, and the method comprises the steps: obtaining a first feature inputted by a speech synthesis model, the first feature being obtained based on multi-modal data processing; obtaining a second feature generated by the speech synthesis model based on the first feature reasoning, wherein the second feature is a speech semantic feature associated with the speech acoustic attribute; based on the first feature and the second feature, performing iterative prediction through a reasoning prediction model to obtain at least one candidate acoustic unit of respective corresponding output positions of continuous target iteration prediction; based on the plurality of candidate acoustic units predicted by continuous target iterations, screening a plurality of target acoustic units with continuously adjacent output positions through a speech synthesis model to synthesize target speech; wherein the model parameters of the reasoning prediction model are smaller than the model parameters of the speech synthesis model.
Owner:LENOVO (BEIJING) LTD

An operator-oriented multi-modal AI large model voice call abnormal risk real-time quality inspection method

PendingCN122340215ACommunications securityAbnormal voice
This invention belongs to the field of communication security and artificial intelligence quality inspection technology, and relates to a real-time quality inspection method and system for abnormal voice calls using a multimodal AI large-scale model for telecom operators. The method receives call information, industry information, and compliance script information; aligns and verifies detailed call records with audio recordings; and performs speech recognition and feature extraction. It combines text semantic features, speech acoustic features, call behavior features, voiceprint biometric features, and deviation features of the actual call content from the reported information to construct multimodal evidence units and evidence sequences. It uses an AI large-scale model to identify candidate risk segments, and uses recording integrity identifiers and reported deviation features as pre-gating to determine the risk level of candidate risk segments and perform review and triage, outputting a real-time quality inspection evidence package. This invention can improve the evidence completeness, risk identification accuracy, and traceability of handling in the quality inspection of abnormal voice calls for telecom operators.
Owner:HANGZHOU AITA TECH CO LTD

Speech synthesis method and device for long text data, equipment and medium

PendingCN121811847ASpeech synthesisSynthesis methodsSpeech Acoustics
The invention relates to the technical field of speech semantics, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a speech synthesis method, device and equipment for long text data and a medium, and the method comprises the steps: obtaining the long text data and corresponding historical speech data, and extracting global text semantic features and speech acoustic features; re-characterizing the global text semantic feature and the voice acoustic feature to obtain a text potential representation and a voice potential representation; performing cross-modal fusion on the text potential representation and the voice potential representation to obtain fusion features; retrieving a target context feature from the historical context information, and fusing the target context feature with the long-term memory feature to obtain a target long-term memory feature; performing attention interaction processing on the target long-term memory feature and the local text semantic feature to obtain a context enhancement feature; and generating a target synthetic speech according to the context enhancement feature. According to the invention, the accuracy of long text data speech synthesis can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Text-guided speech synthesis methods, devices, computer equipment, and storage media

This application belongs to the field of artificial intelligence technology and relates to a text-guided speech synthesis method. The method includes: annotating a speech dataset with style tags and injecting scene noise to obtain a reference speech set; inputting the reference speech set and the text dataset into an acoustic model; encoding the style tags using a style encoder to obtain style encoding features; encoding the reference speech using a reference encoder to obtain reference speech encoding features; encoding the text using a text encoder to obtain text encoding features; inputting all encoded features into an acoustic structure to obtain speech acoustic features; and inputting the speech acoustic features into a vocoder to synthesize a waveform to obtain predicted synthesized speech for training, thereby obtaining a speech synthesis model. This application also provides a text-guided speech synthesis device, computer equipment, and storage medium. Furthermore, this application relates to blockchain technology, allowing the text to be converted to be stored in the blockchain. This application improves the efficiency and quality of speech synthesis.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech reconstruction method and device based on intracranial neural electrical signals and electronic equipment

PendingCN122392482ASpeech reconstructionEngineering
The present application relates to a kind of speech reconstruction methods, device and electronic equipment based on intracranial nerve electric signal, the method comprises: the neural feature time sequence stream corresponding to the intracranial nerve electric signal continuously collected by the language-related brain area of subject is acquired;Neural feature time sequence stream is input to a flow causal decoding model, to generate corresponding speech acoustic feature time sequence stream in real time;Wherein, flow causal decoding model is the causal deep neural network model trained based on sample data, and sample data includes sample neural signal feature sequence and its corresponding sample speech acoustic feature sequence;Speech acoustic feature time sequence stream is integrated into speech waveform stream and output.The present application realizes real-time speech synthesis of millisecond level delay by flow causal decoding model and dynamic block strategy, solves the high delay problem of existing scheme, simultaneously avoids auditory feedback interference by causal mask mechanism, can adapt to the application demand in silent scene.
Owner:AFFILIATED HUSN HOSPITAL OF FUDAN UNIV

Emotional dialogue speech synthesis method and system based on heterogeneous subgraph comparative learning

The invention belongs to the technical field of voice data processing, and particularly discloses an emotional dialogue voice synthesis method and system based on heterogeneous subgraph comparative learning. The method comprises the following steps: obtaining a dialogue voice data set and carrying out cleaning preprocessing, carrying out characterization processing on dialogue voice data, constructing a heterogeneous dialogue graph, carrying out context emotion modeling based on heterogeneous subgraph comparison learning, and carrying out voice synthesis based on emotion perception. According to the method, through heterogeneous subgraph comparative learning, context emotion clues can be focused, and the problem that emotion and context are disjointed in a traditional method is solved; continuous emotion intensity and dimension changes are supported through spherical coordinate coding, and the limitation of discrete emotion labels is broken through; the heterogeneous graph can dynamically integrate multiple rounds of dialogue information and can adapt to emotional evolution of a long dialogue scene; the emotion parameters and the voice acoustic features are directly mapped, so that mismatching of the emotion and the voice features is effectively avoided.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

An office worker emotion recognition method and system based on visible light and voice signals

PendingCN122451571AFacial movementVoice source
The application discloses an office staff emotion recognition method and system based on visible light and voice signals, relates to the technical field of computer vision and voice signal processing, and the method is characterized in that: face images and voice signals are synchronously collected, first, it is judged whether a staff is in a silent state or a speaking state, if the staff is in the silent state, emotion is recognized by analyzing facial motion unit features of the whole face, if the staff is in the speaking state, the voice source is confirmed by lip reading verification, and then emotion is recognized by fusing facial motion features based on only eyebrow and eye regions and voice acoustic features. The application effectively solves the interference problem of mouth movement on expression recognition when speaking, and improves the accuracy and robustness of emotion state monitoring.
Owner:HUNAN INST OF TECH