Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

107 results about "Conversational speech" patented technology

Conversational speech is also known as Informal speech. Speakers regularly apply informal speech with friends and relatives, in daily conversations and in personal letters. Informal speech can cover informal text messages and different types of written communication.

Multi-speaker dialogue voice analysis method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology and medical health, and discloses a multi-speaker dialogue voice analysis method, device and equipment and a medium, and the method comprises the steps: obtaining a to-be-analyzed multi-speaker dialogue voice, determining a naturalness score based on an acoustic feature, and obtaining a multi-speaker dialogue voice analysis result; determining a semantic consistency score based on voice embedding and semantic embedding corresponding to a preset text, determining a speaker consistency score based on embedding of a plurality of speakers of the same speaker, determining an interaction rationality score based on voice alternate overlapping duration, determining a diversity score based on a variance of voice features, and fusing the scores, the comprehensive mass fraction is obtained. According to the invention, through quantitative evaluation of five dimensions of naturalness, semantic consistency, speaker consistency, interaction rationality and diversity, a comprehensive quality scoring system is established, so that the evaluation result simultaneously reflects voice fluency, content matching degree, identity stability, interaction rhythm rationality and feature richness.
Owner:PING AN TECH (SHENZHEN) CO LTD

Automated patient referral management system and method

Disclosed is a system and method for standalone automated patient referral management. The system comprises an intake module that accepts and processes referrals from diverse referral communication channels as entry points for patient referrals. A processing module undertakes the analysis of patient information to pinpoint a healthcare provider based on a healthcare need, a preference, and a schedule availability. A scheduling module confirms the appointment, and engages in dynamic three-way communication between the patient, the referring entity, and the healthcare provider. A feedback module dispatches feedback to the referring entity, offering a report on an outcome of the appointment. A natural language processing (NLP) module extracts information from conversational speech. The system utilizes algorithms to match patients with the healthcare providers based on criteria including specialty, availability, and patient preferences. A feedback module communicates with the originating EHR systems and provides updates to maintain continuity of care.
Owner:INNOVACCER INC

Double-earphone sound effect adjusting method and system, storage medium and program product

The invention discloses a double-earphone sound effect adjusting method and system, a storage medium and a program product, and relates to the field of loudspeakers, and the method comprises the steps: collecting an environment audio signal through a binaural microphone, and executing spectrum analysis to obtain a frequency distribution characteristic of the environment audio signal; based on a preset human voice frequency range, extracting an audio signal conforming to the human voice feature from the frequency distribution feature as a target audio; calculating the time difference and the intensity difference of the target audio between the left ear and the right ear to obtain the human sound source direction and the relative distance in the environment; acquiring head attitude sensor data, and calculating a matching included angle between the head orientation of the user and the human sound source direction based on the head attitude sensor data; when the matching included angle is smaller than a preset threshold value, increasing the audio signal gain of the head facing the direction; and when the conversation voice of the user is detected, the audio signal gain is kept unchanged. By implementing the application, the interaction efficiency when the double earphones are used can be improved.
Owner:BESING TECH SHENZHEN CO LTD

Emotion analysis method and system

The invention relates to the technical field of sentiment analysis, in particular to a sentiment analysis method and system. The sentiment analysis method comprises the steps of collecting a dialogue voice signal, extracting an acoustic feature and a semantic feature from the dialogue voice signal, performing double-flow extraction of the acoustic feature and the semantic feature on the collected dialogue voice signal, performing normalization on the acoustic feature and the semantic feature respectively, and performing sentiment analysis on the acquired dialogue voice signal. The normalized acoustic features and semantic features are spliced to form a joint feature vector, the joint feature vector is mapped into a high-dimensional intermediate emotion vector through linear transformation, controllable Laplace noise is added to the intermediate emotion vector, and the intermediate emotion vector is mapped into a high-dimensional intermediate emotion vector; and outputting an encrypted intermediate emotion vector, inputting the encrypted intermediate emotion vector into a pre-trained emotion analysis neural network, outputting a nine-dimensional emotion vector spectrum, and generating visual feedback on a local device by using the nine-dimensional emotion vector spectrum. According to the invention, emotion understanding and communication optimization in a dialogue scene are greatly improved.
Owner:CHONGQING MINGYUEHU INTELLIGENT TECH DEV CO LTD

Dialogue processing method, device, terminal and storage medium based on emotion recognition

The embodiment of the present application discloses a conversation processing method, device, terminal and storage medium based on emotion recognition. In the process of communication between an agent and a customer based on a business scenario, this solution collects the customer's first voice information and the agent's second voice information, identifies the customer's emotion based on the first voice information, and if the customer's emotion is negative, generates a conversation voice based on the first voice information and the second voice information according to the time point when the customer's negative emotion appears, then performs voice analysis on the conversation voice to obtain a voice analysis result, and performs text analysis on the conversation voice to obtain a text analysis result, and generates customer soothing strategy information based on the voice analysis result and the text analysis result, so that the agent can soothe the customer's emotion according to the customer soothing strategy information, thereby improving the customer's business experience.
Owner:PING AN BANK CO LTD

Personality prediction method and device based on voice conversation, electronic equipment and medium

The invention provides a personality prediction method and device based on voice conversation, electronic equipment and a medium. The method comprises the following steps: acquiring conversation voice information between a to-be-predicted user and an intelligent agent; extracting a target dialogue feature and a dialogue log of the to-be-predicted user from the dialogue voice information; according to the dialogue log and the target dialogue feature, updating the dialogue log and dialogue summary information of the to-be-predicted user; based on the fused target dialogue features, the updated dialogue log and the dialogue summary information, obtaining at least one prediction label of the to-be-predicted user; and inputting the user standardized text description information obtained based on the at least one prediction label fusion into a pre-trained large language model, so that the large language model outputs a personality prediction result of the user to be predicted. Therefore, the accuracy of personality prediction of the to-be-predicted user can be improved.
Owner:SHENZHEN RES INST OF BIG DATA +1

Robot adaptive interaction method and system

The invention relates to the technical field of intelligent robots, in particular to a robot self-adaptive interaction method and system. The robot adaptive interaction system comprises a voice data acquisition module, a dialogue voice data construction module and a dialogue text generation model updating module. In the interaction process of the robot and the person, the voice data of the user is obtained, the text data converted from the voice data is processed through the dialogue text generation model, the dialogue text data used for constructing the dialogue voice data is generated, the robot replies based on the dialogue voice data, interaction of the robot is achieved, and the interaction efficiency of the robot is improved. In the dialogue text generation model, semantic feature vectors are extracted based on the speaking style of the user, and the semantic feature vectors are further updated through multiple rounds of dialogues, so that dialogue text data generated through the text data and the standard semantic feature vectors can conform to the speaking habits of the user; and the experience of the user during interaction is improved.
Owner:WUHAN DONGHU UNIV

Business demand analysis method and device and computer readable storage medium

The invention provides a service demand analysis method and device and a computer readable storage medium, and the method comprises the steps: obtaining all dialogue voice data discussed by a service demand, and converting all dialogue voice data into corresponding dialogue text documents; generating a preliminary demand document based on all dialogue text documents; performing information extraction on the preliminary demand document by using a big and small language model according to a set demand template to obtain specific information, and filling the specific information to a position corresponding to the set demand template to obtain a demand document; and performing function point analysis on the demand document by using the big and small language models to generate a function demand specification. The problem that in the prior art, work such as collection and summarization, key point extraction, induction analysis and content checking of massive multi-source heterogeneous original documents consumes a large amount of time cost is solved.
Owner:中国邮政储蓄银行股份有限公司

Immersive psychotherapy VR system based on dialogue real-time generation

The invention discloses an immersive psychotherapy VR system based on dialogue real-time generation, and the system comprises a voice recognition module which is used for collecting dialogue voice in a consultation process, and transcribing the dialogue voice into a text in real time; the natural language processing module is used for performing semantic understanding on the transcribed dialogue text and extracting scene elements; the scene generation module is used for dynamically generating a corresponding three-dimensional virtual scene according to the scene elements and adjusting scene details in real time along with updating of the dialogue content; the graphic rendering module is used for rendering the three-dimensional virtual scene in real time and then outputting the three-dimensional virtual scene to the VR equipment; and the user interaction module is used for realizing interaction operation among the patient, the therapist and the system. According to the invention, a modular system architecture is adopted, technologies such as speech recognition, large language model understanding, three-dimensional scene generation and graphic rendering are organically combined, and real-time conversion from dialogue content to a virtual scene is realized.
Owner:GUANGZHOU COLLEGE OF COMMERCE

Selective disablement of noise cancelation for conversations

A noise canceling disablement system is provided to enables a user to hear select conversational speech directed at the user while wearing a noise cancelling hearable device. The disablement system automatically at least partial disables the noise canceling feature of the hearable device in response to recognizing conversational speech of a speaking person within a detected conversation zone of the user. In some cases, triggering of the noise canceling disablement further requires the speaking person to be identified by the system as significant person of the user.
Owner:SONY GROUP CORP

Dialogue voice generation method and system based on emotion perception adapter and large model reasoning

The invention discloses a dialogue voice generation method and system based on an emotion perception adapter and large model reasoning. The dialogue voice generation method adopted by the invention comprises the following steps: extracting voice emotion characteristics in original dialogue voice data by using a voice encoder and a time and hierarchical attention network; aligning the speech emotion features with text features of the large language model through an emotion perception coding module based on a query converter network, and generating emotion embedding compatible with the large language model; generating text embedding from statements of the input dialogue text by using a word segmentation device of a large language model; performing emotion embedding and text embedding by adopting an emotion adapter and a text adapter based on a partial low-rank adaptive network, and reasoning a text reply and a reply emotion state of the dialogue; and in combination with the text reply and the reply emotional state, generating a target voice conforming to an emotional context by using a voice generation model. According to the invention, the difference between the emotion and the text is effectively reduced, and the target voice with emotion consistency is generated.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO MARKETING SERVICE CENT

Medical record generation method and related device, electronic equipment and storage medium

The invention discloses a medical record generation method, a related device, electronic equipment and a storage medium, and the method comprises the steps: carrying out the voice recognition based on the doctor-patient conversation voice in an inquiry process, and obtaining a doctor-patient conversation text; performing information extraction based on the doctor-patient dialogue text to obtain a time information element, a medical event corresponding to a time interval and a spatial information element; constructing a symptom evolution time sequence chain based on the time information element and the medical event corresponding to the time interval, and constructing a symptom spatial map based on the spatial information element; performing feature extraction based on the symptom evolution time sequence chain to obtain a first feature, performing feature extraction based on the symptom space map to obtain a second feature, and performing feature extraction based on the doctor-patient dialogue text to obtain a third feature; and generating an electronic medical record text based on a fusion feature of the first feature, the second feature and the third feature. According to the scheme, dynamic information of disease evolution can be reflected in medical record generation.
Owner:THE FIRST AFFILIATED HOSPITAL OF ANHUI MEDICAL UNIV +1

Voice quality inspection method and device, electronic equipment and computer readable storage medium

The voice quality inspection method and device, the electronic equipment and the computer readable storage medium provided in the application are related to the artificial intelligence technical field and are suitable for the financial technology field. The method comprises the following steps: obtaining target conversation voice; performing text conversion on the target conversation voice to obtain target conversation text; performing semantic division on the target conversation text to obtain target semantic blocks and semantic block categories; performing information searching on a preset knowledge graph according to the target semantic blocks to obtain entity triple text; performing sentence reading on a preset semantic sentence database according to the semantic block categories to obtain standard semantic sentences; performing text fusion according to the semantic block categories, the standard semantic sentences, the target semantic blocks and the entity triple text to obtain target quality inspection prompt text; and performing quality inspection on the target quality inspection prompt text to obtain a voice quality inspection result. The application embodiment can improve the voice quality inspection accuracy.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent follow-up visit voice robot multi-round dialogue state tracking system

The invention discloses a multi-round dialogue state tracking system for an intelligent follow-up voice robot. The multi-round dialogue state tracking system comprises a hierarchical semantic state decoder module, a context disambiguation state optimization model module, an intention evolution state tracking algorithm module, a follow-up dialogue streaming analysis platform, a multi-round dialogue state storage module and a robot interaction control module. The modules work cooperatively, dialogue voice signals are analyzed through the decoder to output semantic features, the semantic features are processed through the disambiguation model, an intention evolution trajectory is constructed through a tracking algorithm, dialogue state features are extracted through the streaming analysis platform, the storage module stores data in a classified mode and supports the calling and control module to generate interaction signals. The system can deeply integrate semantic processing and historical data, meanwhile, efficient cooperation of state tracking and data management is achieved through bidirectional interaction of modules, tracking deviation and lag are eliminated, follow-up conversation interaction coherence and information accuracy are guaranteed, the intelligent requirement of medical follow-up is met, and the follow-up working efficiency is improved.
Owner:ATRI (SHANDONG) DIGITAL TECHNOLOGY CO LTD

Selective disablement of noise cancelation for conversations

A noise canceling disablement system is provided to enable a user to hear select conversational speech directed at the user while wearing a noise cancelling hearable device. The disablement system automatically at least partially disables the noise canceling feature of the hearable device in response to recognizing conversational speech of a speaking person within a detected conversation zone of the user. In some cases, triggering of the noise canceling disablement further requires the speaking person to be identified by the system as a significant person of the user.
Owner:SONY GROUP CORP

Dialogue voice generation method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, the scheme can be applied to the fields of finance and medical treatment, and the invention provides a dialogue voice generation method, device and equipment and a medium, and the method comprises the steps: converting an input text abstract into a dialogue type text structure with a multi-role interaction feature through a large language model; distributing a unique label feature for each proxy role in the dialogue text structure; automatically matching acoustic characteristic parameters conforming to the agent roles from a preset voice library according to the label characteristics; and converting the dialogue text structure into dialogue voice through a voice synthesis model according to the acoustic characteristic parameters of each agent role, and outputting the dialogue voice. According to the embodiment of the invention, the input text abstract can be converted into the dialogue type text structure with the multi-role interaction characteristic, the requirements of audiences for deep discussion and professional insight are met, and the dialogue type text structure can be converted into dialogue voice with content depth and expressive force according to the acoustic characteristic parameter of each agent role.
Owner:PING AN TECH (SHENZHEN) CO LTD

A speech prosody recognition method, system, device and storage medium

This invention discloses a speech prosody recognition method, system, device, and storage medium. First, the voice from a customer service call is captured to obtain a dialogue speech signal. Then, the dialogue speech signal is preprocessed to obtain a preprocessed dialogue speech file. Next, the preprocessed dialogue speech file is vectorized to obtain a corresponding feature matrix. The feature matrix is ​​input into a trained prosody model to obtain the model calculation result. Then, based on Mandarin and dialect templates, corresponding template thresholds are obtained. Using the model calculation result and template thresholds, the prosody recognition result of the feature matrix is ​​obtained. Finally, based on the prosody recognition result, the feature matrix is ​​processed by text mapping to obtain the dialogue word order text. This invention effectively improves the recognition accuracy and efficiency for speech containing dialects.
Owner:BEIJING GARUI INTELLIGENT TECH GRP CO LTD

A method, system and apparatus for dubbing a novel content

PendingCN122658285AEmotional expressivityConversational speech
The application is suitable for the technical field of language synthesis, and provides a novel content dubbing method, which comprises the following steps: obtaining a novel text to be dubbed, preprocessing the novel text, identifying role dialogue sentences, environment description sentences and mixed sentences; extracting dialogue content and performing language emotion analysis to obtain language emotion labels; assigning corresponding general timbre embedding vectors to different roles, tracking state changes of the roles in plot development, and generating timbre gradient control parameters; based on the general timbre embedding vectors, the timbre gradient control parameters and the language emotion labels, performing voice synthesis on each dialogue content to generate dialogue voice waveforms; extracting environment semantic labels and intensity parameters from the environment description sentences to generate environment sound effect signals; mixing the dialogue voice waveforms and the environment sound effect signals to generate a dubbing audio file, so that the emotional expression of the text context and the smooth evolution of the same role in different age stages can be taken into account.
Owner:ANHUI SANQI JIYU NETWORK TECH CO LTD

system

We provide the system. [Solution] A motion detection means for detecting the user's actions, The motion detection means detects the user's actions and initiates a conversation, and the voice acquisition means acquires voice information. A data management system that digitizes acquired voice information and compares it with accumulated past conversation data, A response generation means that generates an appropriate response based on past conversation data, A response providing means that provides the generated response to the user, An abnormality notification means that notifies of an abnormality if no operation is detected for a predetermined period of time, A system that includes this.
Owner:SOFTBANK GROUP CORP

Children physiological education doll, method and program product

The invention belongs to the technical field of child dolls, and particularly discloses a child physiological education doll, a method and a program product, and the method comprises the steps: collecting the conversation voice of a child through the doll, preprocessing the conversation voice when the conversation voice triggers a conversation activation state, and uploading the preprocessed conversation voice to a cloud platform, the pre-trained large model is called through the cloud platform to carry out voice analysis and intelligent dialogue answering, corresponding physiological knowledge answering information is obtained and played to the child, and meanwhile, when the child touches the sensing device of the corresponding part of the doll, touch feedback voice information is played according to the corresponding part, so that early education of the physiological knowledge of the child is achieved. The child physiological knowledge education doll fills the blank of child physiological knowledge education dolls, dialogue education and touch feedback education of child physiological knowledge can be intelligently and visually achieved in the form of the dolls, and the child can effectively receive related physiological knowledge conveniently.
Owner:XINYI ANIMATION (HUIZHOU) CO LTD

Systems and methods for predicting mental health conditions based on passive processing of conversational speech and language

PendingUS20260100259A1Semantic analysisMedical automated diagnosisMedicineMore language
Described herein are systems and methods for identifying the severity of a mental health condition or symptoms of same by listening to a human-to-human conversation by receiving conversation data, processing the conversation data to generate a language model output and / or an acoustic model output using one or more language models and / or acoustic models. Further described herein are systems and methods for automatically tracking and providing analytics on self-report questionnaires administered during the conversation.
Owner:ELLIPSIS HEALTH INC

Audio processing method and system, training method, computing device and storage medium

The invention relates to an audio processing method, which comprises the following steps of: receiving first and second mixed audios which are respectively acquired by a first microphone and a second microphone of wearable equipment for conversation voice, the dialogue voice comprises a first voice signal sent by a first speaker wearing the wearable device, a second voice signal sent by a second speaker having a dialogue with the first speaker, and a third voice signal of text-to-voice played by the electronic device; inputting the reference audio, the first mixed audio and the second mixed audio of the electronic equipment into a trained separation model to obtain a first mask and a second mask of the first speaker and the second speaker; selecting one of the first mixed audio and the second mixed audio as a mixed audio to be processed, and obtaining a first separated voice signal and a second separated voice signal from the mixed audio to be processed by using the first mask and the second mask; translating the first and second separated voice signals to obtain first and second translated texts, and playing the first translated text on the electronic equipment; the second translated text is sent to the wearable device for display thereof.
Owner:湖北星纪魅族集团有限公司

Speech analysis method and device, computer equipment and readable storage medium

The invention relates to the technical field of voice processing, and provides a voice analysis method and device, computer equipment and a readable storage medium, and the method comprises the steps: obtaining a to-be-analyzed conversation voice in a preset business scene, and generating preprocessed voice data; converting the preprocessed voice data into text information, and extracting acoustic features and text features of the voice data; acquiring basic language understanding and pattern recognition capability according to the acoustic features and the text features, and optimizing the basic language understanding and pattern recognition capability according to a small amount of labeled sample data corresponding to the business scene; and dialogue feature representation is generated according to the optimized acoustic features and text features, the feature representation is analyzed, an intelligent decision corresponding to the business scene is generated, and analysis of dialogue voice is completed. Through deep coupling of acoustics and text features, shallow application of traditional speech recognition is broken through, and a general technical path of small data driven accurate decision is provided for the fields of finance, medical health, old-age care and the like.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Method, system, electronic device and storage medium for interactive speech recognition

This invention provides a method, system, electronic device, and storage medium for interactive speech recognition. The method includes: recognizing the dialogue speech input by a user in the current round; inputting text hypotheses into an interactive automated simulation framework; performing semantic-based intent allocation between the text hypotheses and the recognition results of the previous round; if the text hypotheses are determined to be correction instructions for the recognition results of the previous round, using the correction instructions to infer and correct the previous round's recognition results to obtain a corrected result for the previous round; if the text hypotheses are determined to be new dialogue content, using a user simulator to simulate the user's correction behavior to generate correction prompts, and the interactive speech recognizer using the correction prompts to infer and correct the text hypotheses of the current round to obtain a corrected recognition result for the current round. In modeling the speech recognition task, this invention utilizes a large text model and an interactive automated simulation framework to achieve real-time error correction based on user feedback.
Owner:SHANGHAI JIAOTONG UNIV

Service quality monitoring method and device

The invention provides a service quality monitoring method and device. The method comprises the following steps: acquiring user voice and seat voice of a current round; performing emotion recognition on the user voice and the seat voice of the current round based on an emotion recognition model to obtain a user emotion result and a seat emotion result; and performing service quality evaluation based on the user emotion result and the seat emotion result of the current round, and determining a service quality monitoring result. According to the method provided by the invention, emotion recognition is performed on the user voice and the seat voice of the current round through the emotion recognition model to obtain the user emotion result and the seat emotion result, and then service quality evaluation is performed based on the user emotion result and the seat emotion result of the current round and the historical emotion result of the dialogue voice of the historical round. And the service quality monitoring result of the current round of dialogue voice is determined, so that timely, efficient and accurate service quality monitoring is realized, an enterprise can find a service quality problem in time, and the service quality of an agent is improved.
Owner:YUANBAO TECH (BEIJING) TECH CO LTD

A high-dynamic-environment intercom voice enhancement method and system based on intelligent noise reduction

ActiveCN122050413BStationary noiseNoise
The application provides a high-dynamic-environment intercom voice enhancement method and system based on intelligent noise reduction. In response to a release event of a push-to-talk button, a two-state noise dictionary is constructed based on impulsive noise components and non-stationary noise components in background noise in a current high-dynamic environment. In response to a press event of the push-to-talk button, when it is detected that there is a transient region matching the impulsive noise components in the noisy intercom audio signal, online updating of the two-state noise dictionary is triggered to obtain an online updating dictionary. The noisy intercom audio signal is reconstructed based on the online updating dictionary to obtain an initial enhanced voice signal. Envelope reconstruction is performed on the voice segment with abnormal zero-crossing rate in the initial enhanced voice signal to generate a final enhanced voice signal. The technical scheme provided by the application can enhance conversation voice in a high-dynamic environment with non-stationary noise and impulsive noise.
Owner:SHENZHEN AIQISHI INTELLIGENT TECHNOLOGY CO LTD

Voice editing processing method and device

The embodiment of the invention provides a voice editing processing method and device.The voice editing processing method comprises the steps that in the voice editing processing process, based on a voice text of dialogue voice input by a user in an interaction component of an application program, interaction action data of the user for the voice text are firstly obtained, and the interaction action data are sent to the interaction component; and detecting whether a voice editing intention exists or not according to the interaction action data and the interaction environment data, further determining an intention type of the voice editing intention when detecting that the voice editing intention exists, and displaying a voice editing identifier corresponding to the intention type on the interaction component. And acquiring a voice editing fragment collected after the voice editing identifier is triggered, and performing editing adaptation processing on the voice text based on the voice editing fragment, thereby realizing editing processing on the voice text by perceiving the voice editing intention.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Speech data construction method and device, electronic equipment and storage medium

The present disclosure relates to a voice data construction method and device, electronic equipment and storage medium. The voice data construction method comprises: obtaining a first foreign language medical inquiry dialogue text data set, a second foreign language medical inquiry dialogue text data set, a first Chinese medical inquiry dialogue text data set and a target voice data set, the target voice data set being a collection of Chinese voice data sets in a non-medical inquiry scene; translating each first foreign language medical inquiry dialogue text data and each second foreign language medical inquiry dialogue text data to obtain a plurality of second Chinese medical inquiry dialogue text data; and performing voice synthesis on the first Chinese medical inquiry dialogue text data and the second Chinese medical inquiry dialogue text data based on each target voice data to obtain a plurality of medical inquiry dialogue voice data. Thus, high-quality medical inquiry voice data can be obtained while reducing labor costs and improving efficiency.
Owner:BEIJING ECON MEDICAL TECHNOLOGY CO LTD

Depression detection method and system based on emotional expression

The application relates to the technical field of information technology, and discloses a depression detection method and system based on emotional expression. Specifically disclosed are the following: obtaining voice data and text data from an original doctor-patient conversation voice segment, extracting voice features and text features by using a pre-training model, and automatically labeling voice emotion and text emotion based on the same. Then, voice feature embedding and text feature embedding are extracted, multi-head attention mechanism is used to extract emotional representations of the voice and the text, cross-attention mechanism is used to analyze the degree of mismatch between the voice and the text emotional representations, and emotional inconsistency feature embedding is obtained. The voice feature embedding, the text feature embedding and the emotional inconsistency feature embedding are fused, and depression categories are predicted according to the fused features. The automatic emotional labeling method improves the labeling efficiency; the phenomenon of inconsistent emotional expression is used to assist depression diagnosis, and the accuracy of depression detection is improved.
Owner:SHENZHEN INST OF ADVANCED TECH

Voice data construction method and device, electronic equipment and storage medium

The invention relates to a voice data construction method and device, electronic equipment and a storage medium. The voice data construction method comprises the steps that a first foreign language medical inquiry dialogue text data set, a second foreign language medical inquiry dialogue text data set, a first Chinese medical inquiry dialogue text data set and a target voice data set are acquired, and the target voice data set is an acquired Chinese voice data set in a non-medical inquiry scene; translating each first foreign language medical interrogation dialogue text data and each second foreign language medical interrogation dialogue text data to obtain a plurality of second Chinese medical interrogation dialogue text data; and performing voice synthesis on the first Chinese medical interrogation dialogue text data and the second Chinese medical interrogation dialogue text data based on each piece of target voice data to obtain multiple pieces of medical interrogation dialogue voice data, thereby obtaining high-quality medical interrogation voice data while reducing the labor cost and improving the obtaining efficiency.
Owner:BEIJING ECON MEDICAL TECHNOLOGY CO LTD