Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

77 results about "Conversational speech" patented technology

Conversational speech is also known as Informal speech. Speakers regularly apply informal speech with friends and relatives, in daily conversations and in personal letters. Informal speech can cover informal text messages and different types of written communication.

Multi-speaker dialogue voice analysis method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology and medical health, and discloses a multi-speaker dialogue voice analysis method, device and equipment and a medium, and the method comprises the steps: obtaining a to-be-analyzed multi-speaker dialogue voice, determining a naturalness score based on an acoustic feature, and obtaining a multi-speaker dialogue voice analysis result; determining a semantic consistency score based on voice embedding and semantic embedding corresponding to a preset text, determining a speaker consistency score based on embedding of a plurality of speakers of the same speaker, determining an interaction rationality score based on voice alternate overlapping duration, determining a diversity score based on a variance of voice features, and fusing the scores, the comprehensive mass fraction is obtained. According to the invention, through quantitative evaluation of five dimensions of naturalness, semantic consistency, speaker consistency, interaction rationality and diversity, a comprehensive quality scoring system is established, so that the evaluation result simultaneously reflects voice fluency, content matching degree, identity stability, interaction rhythm rationality and feature richness.
Owner:PING AN TECH (SHENZHEN) CO LTD

Automated patient referral management system and method

Disclosed is a system and method for standalone automated patient referral management. The system comprises an intake module that accepts and processes referrals from diverse referral communication channels as entry points for patient referrals. A processing module undertakes the analysis of patient information to pinpoint a healthcare provider based on a healthcare need, a preference, and a schedule availability. A scheduling module confirms the appointment, and engages in dynamic three-way communication between the patient, the referring entity, and the healthcare provider. A feedback module dispatches feedback to the referring entity, offering a report on an outcome of the appointment. A natural language processing (NLP) module extracts information from conversational speech. The system utilizes algorithms to match patients with the healthcare providers based on criteria including specialty, availability, and patient preferences. A feedback module communicates with the originating EHR systems and provides updates to maintain continuity of care.
Owner:INNOVACCER INC

Personality prediction method and device based on voice conversation, electronic equipment and medium

The invention provides a personality prediction method and device based on voice conversation, electronic equipment and a medium. The method comprises the following steps: acquiring conversation voice information between a to-be-predicted user and an intelligent agent; extracting a target dialogue feature and a dialogue log of the to-be-predicted user from the dialogue voice information; according to the dialogue log and the target dialogue feature, updating the dialogue log and dialogue summary information of the to-be-predicted user; based on the fused target dialogue features, the updated dialogue log and the dialogue summary information, obtaining at least one prediction label of the to-be-predicted user; and inputting the user standardized text description information obtained based on the at least one prediction label fusion into a pre-trained large language model, so that the large language model outputs a personality prediction result of the user to be predicted. Therefore, the accuracy of personality prediction of the to-be-predicted user can be improved.
Owner:SHENZHEN RES INST OF BIG DATA +1

Immersive psychotherapy VR system based on dialogue real-time generation

The invention discloses an immersive psychotherapy VR system based on dialogue real-time generation, and the system comprises a voice recognition module which is used for collecting dialogue voice in a consultation process, and transcribing the dialogue voice into a text in real time; the natural language processing module is used for performing semantic understanding on the transcribed dialogue text and extracting scene elements; the scene generation module is used for dynamically generating a corresponding three-dimensional virtual scene according to the scene elements and adjusting scene details in real time along with updating of the dialogue content; the graphic rendering module is used for rendering the three-dimensional virtual scene in real time and then outputting the three-dimensional virtual scene to the VR equipment; and the user interaction module is used for realizing interaction operation among the patient, the therapist and the system. According to the invention, a modular system architecture is adopted, technologies such as speech recognition, large language model understanding, three-dimensional scene generation and graphic rendering are organically combined, and real-time conversion from dialogue content to a virtual scene is realized.
Owner:GUANGZHOU COLLEGE OF COMMERCE

Medical record generation method and related device, electronic equipment and storage medium

The invention discloses a medical record generation method, a related device, electronic equipment and a storage medium, and the method comprises the steps: carrying out the voice recognition based on the doctor-patient conversation voice in an inquiry process, and obtaining a doctor-patient conversation text; performing information extraction based on the doctor-patient dialogue text to obtain a time information element, a medical event corresponding to a time interval and a spatial information element; constructing a symptom evolution time sequence chain based on the time information element and the medical event corresponding to the time interval, and constructing a symptom spatial map based on the spatial information element; performing feature extraction based on the symptom evolution time sequence chain to obtain a first feature, performing feature extraction based on the symptom space map to obtain a second feature, and performing feature extraction based on the doctor-patient dialogue text to obtain a third feature; and generating an electronic medical record text based on a fusion feature of the first feature, the second feature and the third feature. According to the scheme, dynamic information of disease evolution can be reflected in medical record generation.
Owner:THE FIRST AFFILIATED HOSPITAL OF ANHUI MEDICAL UNIV +1

Voice quality inspection method and device, electronic equipment and computer readable storage medium

The voice quality inspection method and device, the electronic equipment and the computer readable storage medium provided in the application are related to the artificial intelligence technical field and are suitable for the financial technology field. The method comprises the following steps: obtaining target conversation voice; performing text conversion on the target conversation voice to obtain target conversation text; performing semantic division on the target conversation text to obtain target semantic blocks and semantic block categories; performing information searching on a preset knowledge graph according to the target semantic blocks to obtain entity triple text; performing sentence reading on a preset semantic sentence database according to the semantic block categories to obtain standard semantic sentences; performing text fusion according to the semantic block categories, the standard semantic sentences, the target semantic blocks and the entity triple text to obtain target quality inspection prompt text; and performing quality inspection on the target quality inspection prompt text to obtain a voice quality inspection result. The application embodiment can improve the voice quality inspection accuracy.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent follow-up visit voice robot multi-round dialogue state tracking system

The invention discloses a multi-round dialogue state tracking system for an intelligent follow-up voice robot. The multi-round dialogue state tracking system comprises a hierarchical semantic state decoder module, a context disambiguation state optimization model module, an intention evolution state tracking algorithm module, a follow-up dialogue streaming analysis platform, a multi-round dialogue state storage module and a robot interaction control module. The modules work cooperatively, dialogue voice signals are analyzed through the decoder to output semantic features, the semantic features are processed through the disambiguation model, an intention evolution trajectory is constructed through a tracking algorithm, dialogue state features are extracted through the streaming analysis platform, the storage module stores data in a classified mode and supports the calling and control module to generate interaction signals. The system can deeply integrate semantic processing and historical data, meanwhile, efficient cooperation of state tracking and data management is achieved through bidirectional interaction of modules, tracking deviation and lag are eliminated, follow-up conversation interaction coherence and information accuracy are guaranteed, the intelligent requirement of medical follow-up is met, and the follow-up working efficiency is improved.
Owner:ATRI (SHANDONG) DIGITAL TECHNOLOGY CO LTD

A speech prosody recognition method, system, device and storage medium

This invention discloses a speech prosody recognition method, system, device, and storage medium. First, the voice from a customer service call is captured to obtain a dialogue speech signal. Then, the dialogue speech signal is preprocessed to obtain a preprocessed dialogue speech file. Next, the preprocessed dialogue speech file is vectorized to obtain a corresponding feature matrix. The feature matrix is ​​input into a trained prosody model to obtain the model calculation result. Then, based on Mandarin and dialect templates, corresponding template thresholds are obtained. Using the model calculation result and template thresholds, the prosody recognition result of the feature matrix is ​​obtained. Finally, based on the prosody recognition result, the feature matrix is ​​processed by text mapping to obtain the dialogue word order text. This invention effectively improves the recognition accuracy and efficiency for speech containing dialects.
Owner:BEIJING GARUI INTELLIGENT TECH GRP CO LTD

A method, system and apparatus for dubbing a novel content

PendingCN122658285AEmotional expressivityConversational speech
The application is suitable for the technical field of language synthesis, and provides a novel content dubbing method, which comprises the following steps: obtaining a novel text to be dubbed, preprocessing the novel text, identifying role dialogue sentences, environment description sentences and mixed sentences; extracting dialogue content and performing language emotion analysis to obtain language emotion labels; assigning corresponding general timbre embedding vectors to different roles, tracking state changes of the roles in plot development, and generating timbre gradient control parameters; based on the general timbre embedding vectors, the timbre gradient control parameters and the language emotion labels, performing voice synthesis on each dialogue content to generate dialogue voice waveforms; extracting environment semantic labels and intensity parameters from the environment description sentences to generate environment sound effect signals; mixing the dialogue voice waveforms and the environment sound effect signals to generate a dubbing audio file, so that the emotional expression of the text context and the smooth evolution of the same role in different age stages can be taken into account.
Owner:ANHUI SANQI JIYU NETWORK TECH CO LTD

system

We provide the system. [Solution] A motion detection means for detecting the user's actions, The motion detection means detects the user's actions and initiates a conversation, and the voice acquisition means acquires voice information. A data management system that digitizes acquired voice information and compares it with accumulated past conversation data, A response generation means that generates an appropriate response based on past conversation data, A response providing means that provides the generated response to the user, An abnormality notification means that notifies of an abnormality if no operation is detected for a predetermined period of time, A system that includes this.
Owner:SOFTBANK GROUP CORP

Children physiological education doll, method and program product

The invention belongs to the technical field of child dolls, and particularly discloses a child physiological education doll, a method and a program product, and the method comprises the steps: collecting the conversation voice of a child through the doll, preprocessing the conversation voice when the conversation voice triggers a conversation activation state, and uploading the preprocessed conversation voice to a cloud platform, the pre-trained large model is called through the cloud platform to carry out voice analysis and intelligent dialogue answering, corresponding physiological knowledge answering information is obtained and played to the child, and meanwhile, when the child touches the sensing device of the corresponding part of the doll, touch feedback voice information is played according to the corresponding part, so that early education of the physiological knowledge of the child is achieved. The child physiological knowledge education doll fills the blank of child physiological knowledge education dolls, dialogue education and touch feedback education of child physiological knowledge can be intelligently and visually achieved in the form of the dolls, and the child can effectively receive related physiological knowledge conveniently.
Owner:XINYI ANIMATION (HUIZHOU) CO LTD

Systems and methods for predicting mental health conditions based on passive processing of conversational speech and language

PendingUS20260100259A1Semantic analysisMedical automated diagnosisMedicineMore language
Described herein are systems and methods for identifying the severity of a mental health condition or symptoms of same by listening to a human-to-human conversation by receiving conversation data, processing the conversation data to generate a language model output and / or an acoustic model output using one or more language models and / or acoustic models. Further described herein are systems and methods for automatically tracking and providing analytics on self-report questionnaires administered during the conversation.
Owner:ELLIPSIS HEALTH INC

Speech analysis method and device, computer equipment and readable storage medium

The invention relates to the technical field of voice processing, and provides a voice analysis method and device, computer equipment and a readable storage medium, and the method comprises the steps: obtaining a to-be-analyzed conversation voice in a preset business scene, and generating preprocessed voice data; converting the preprocessed voice data into text information, and extracting acoustic features and text features of the voice data; acquiring basic language understanding and pattern recognition capability according to the acoustic features and the text features, and optimizing the basic language understanding and pattern recognition capability according to a small amount of labeled sample data corresponding to the business scene; and dialogue feature representation is generated according to the optimized acoustic features and text features, the feature representation is analyzed, an intelligent decision corresponding to the business scene is generated, and analysis of dialogue voice is completed. Through deep coupling of acoustics and text features, shallow application of traditional speech recognition is broken through, and a general technical path of small data driven accurate decision is provided for the fields of finance, medical health, old-age care and the like.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Method, system, electronic device and storage medium for interactive speech recognition

This invention provides a method, system, electronic device, and storage medium for interactive speech recognition. The method includes: recognizing the dialogue speech input by a user in the current round; inputting text hypotheses into an interactive automated simulation framework; performing semantic-based intent allocation between the text hypotheses and the recognition results of the previous round; if the text hypotheses are determined to be correction instructions for the recognition results of the previous round, using the correction instructions to infer and correct the previous round's recognition results to obtain a corrected result for the previous round; if the text hypotheses are determined to be new dialogue content, using a user simulator to simulate the user's correction behavior to generate correction prompts, and the interactive speech recognizer using the correction prompts to infer and correct the text hypotheses of the current round to obtain a corrected recognition result for the current round. In modeling the speech recognition task, this invention utilizes a large text model and an interactive automated simulation framework to achieve real-time error correction based on user feedback.
Owner:SHANGHAI JIAOTONG UNIV

A high-dynamic-environment intercom voice enhancement method and system based on intelligent noise reduction

ActiveCN122050413BStationary noiseNoise
The application provides a high-dynamic-environment intercom voice enhancement method and system based on intelligent noise reduction. In response to a release event of a push-to-talk button, a two-state noise dictionary is constructed based on impulsive noise components and non-stationary noise components in background noise in a current high-dynamic environment. In response to a press event of the push-to-talk button, when it is detected that there is a transient region matching the impulsive noise components in the noisy intercom audio signal, online updating of the two-state noise dictionary is triggered to obtain an online updating dictionary. The noisy intercom audio signal is reconstructed based on the online updating dictionary to obtain an initial enhanced voice signal. Envelope reconstruction is performed on the voice segment with abnormal zero-crossing rate in the initial enhanced voice signal to generate a final enhanced voice signal. The technical scheme provided by the application can enhance conversation voice in a high-dynamic environment with non-stationary noise and impulsive noise.
Owner:SHENZHEN AIQISHI INTELLIGENT TECHNOLOGY CO LTD

Voice editing processing method and device

The embodiment of the invention provides a voice editing processing method and device.The voice editing processing method comprises the steps that in the voice editing processing process, based on a voice text of dialogue voice input by a user in an interaction component of an application program, interaction action data of the user for the voice text are firstly obtained, and the interaction action data are sent to the interaction component; and detecting whether a voice editing intention exists or not according to the interaction action data and the interaction environment data, further determining an intention type of the voice editing intention when detecting that the voice editing intention exists, and displaying a voice editing identifier corresponding to the intention type on the interaction component. And acquiring a voice editing fragment collected after the voice editing identifier is triggered, and performing editing adaptation processing on the voice text based on the voice editing fragment, thereby realizing editing processing on the voice text by perceiving the voice editing intention.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Speech data construction method and device, electronic equipment and storage medium

The present disclosure relates to a voice data construction method and device, electronic equipment and storage medium. The voice data construction method comprises: obtaining a first foreign language medical inquiry dialogue text data set, a second foreign language medical inquiry dialogue text data set, a first Chinese medical inquiry dialogue text data set and a target voice data set, the target voice data set being a collection of Chinese voice data sets in a non-medical inquiry scene; translating each first foreign language medical inquiry dialogue text data and each second foreign language medical inquiry dialogue text data to obtain a plurality of second Chinese medical inquiry dialogue text data; and performing voice synthesis on the first Chinese medical inquiry dialogue text data and the second Chinese medical inquiry dialogue text data based on each target voice data to obtain a plurality of medical inquiry dialogue voice data. Thus, high-quality medical inquiry voice data can be obtained while reducing labor costs and improving efficiency.
Owner:BEIJING ECON MEDICAL TECHNOLOGY CO LTD

Depression detection method and system based on emotional expression

The application relates to the technical field of information technology, and discloses a depression detection method and system based on emotional expression. Specifically disclosed are the following: obtaining voice data and text data from an original doctor-patient conversation voice segment, extracting voice features and text features by using a pre-training model, and automatically labeling voice emotion and text emotion based on the same. Then, voice feature embedding and text feature embedding are extracted, multi-head attention mechanism is used to extract emotional representations of the voice and the text, cross-attention mechanism is used to analyze the degree of mismatch between the voice and the text emotional representations, and emotional inconsistency feature embedding is obtained. The voice feature embedding, the text feature embedding and the emotional inconsistency feature embedding are fused, and depression categories are predicted according to the fused features. The automatic emotional labeling method improves the labeling efficiency; the phenomenon of inconsistent emotional expression is used to assist depression diagnosis, and the accuracy of depression detection is improved.
Owner:SHENZHEN INST OF ADVANCED TECH

Voice data construction method and device, electronic equipment and storage medium

The invention relates to a voice data construction method and device, electronic equipment and a storage medium. The voice data construction method comprises the steps that a first foreign language medical inquiry dialogue text data set, a second foreign language medical inquiry dialogue text data set, a first Chinese medical inquiry dialogue text data set and a target voice data set are acquired, and the target voice data set is an acquired Chinese voice data set in a non-medical inquiry scene; translating each first foreign language medical interrogation dialogue text data and each second foreign language medical interrogation dialogue text data to obtain a plurality of second Chinese medical interrogation dialogue text data; and performing voice synthesis on the first Chinese medical interrogation dialogue text data and the second Chinese medical interrogation dialogue text data based on each piece of target voice data to obtain multiple pieces of medical interrogation dialogue voice data, thereby obtaining high-quality medical interrogation voice data while reducing the labor cost and improving the obtaining efficiency.
Owner:BEIJING ECON MEDICAL TECHNOLOGY CO LTD

Two-way interaction method, system and glasses for deaf-mute

The invention belongs to the technical field of interaction, and discloses a deaf-mute two-way interaction method, which comprises the steps of collecting voice of a dialogue person, converting the voice into first character information, visually displaying the first character information to a deaf-mute user, or converting the first character information into sign language actions or sign language schematic diagrams, displaying the sign language actions or sign language schematic diagrams to the deaf-mute user, capturing lip shape images of the dialogue person, and displaying the lip shape images to the deaf-mute user. Third character information is generated through matching based on the lip language database, the third character information is matched with the first character information to determine that the first character information is from the target dialogue person, a hand action image of the deaf-mute user is captured, the hand action image is converted into second character information, and the second character information is sent to the deaf-mute user. According to the method, the voice of the dialogue person is collected, the voice is converted into characters or sign language actions, the characters or sign language actions are visually displayed to the deaf-mute user, and the sign language actions are converted into the characters based on the sign language database by collecting the sign language actions of the deaf-mute user and are played in the form of voice.
Owner:INST OF MECHANICS CHINESE ACAD OF SCI

Demand information identification method and device based on voice information and large model

The embodiment of the invention discloses a demand information identification method and device based on voice information and a large model. A specific embodiment of the method comprises the steps that dialogue voice text information for a consultation scene is acquired from a recording storage database, and the dialogue voice text information is text information obtained after dialogue voice between different users is converted into texts; inputting the dialogue voice text information and a preset prompt information set into a prompt word embedding model to obtain an embedded structured text; generating a to-be-followed task information set according to the embedded structured text; and performing task demand identification on the follow-up task scoring report to generate follow-up task demand information, and controlling execution equipment to perform task execution on the follow-up task demand information. According to the embodiment, the completeness of key information capture is improved, and the response period of the demand information of the user is shortened.
Owner:BEIJING PENGLONGXING AUTOMOBILE TRADING CO LTD

Virtual avatar evolution system based on ai learning and observation and method thereof

A virtual avatar evolution system based on AI learning and observation and a method thereof are disclosed. In the system, when a virtual avatar triggers an conversation event, speech of the virtual avatar is continuously recorded to generate a complete conversational speech, and the complete conversational speech is converted into a complete conversation message via a speech-to-text technology, then the complete conversation message is used as new training data for re-training a pre-trained language model; the complete conversation message is input into the personality analysis model to obtain real-person personality data, and the virtual personality data of the virtual avatar is dynamically adjusted based on the real-person personality data, thereby achieving the technical effect of enhancing the interaction compatibility of the virtual avatar.
Owner:SQ TECH (SHANGHAI) CORP +1

Dialogue processing method, dialogue system, electronic device, and computer storage medium

The application provides a dialogue processing method, a dialogue system, an electronic device and a computer storage medium, comprising converting user input speech into corresponding dialogue states; determining corresponding initial dialogue skills, processing the initial dialogue skills, and determining target dialogue skills; selecting initial speech actions based on the target dialogue skills; processing the initial speech actions to determine target speech actions; and converting the target speech actions into corresponding dialogue speech and playing it. The scheme proposes a method of using hierarchical reinforcement learning to construct a two-layer dialogue strategy of top-level strategy and bottom-level strategy, selects initial dialogue skills through the top-level strategy, and then processes to obtain target dialogue skills; selects initial speech actions based on the target dialogue skills through the bottom-level strategy, and then processes to obtain target speech actions. Through the above method, corresponding dialogue skills can be called according to different dialogue scenes, and corresponding topic guidance can be performed, so as to achieve the purpose of smooth dialogue.
Owner:JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD

Speech recommendation method for emotion analysis based on voiceprint recognition and related equipment

The invention belongs to the technical field of artificial intelligence, and relates to a voiceprint recognition-based emotion analysis verbal skill recommendation method and related equipment, and the method comprises the steps: obtaining the real-time conversation voice of a target user side; extracting voice voiceprint features and dialogue semantic features; performing feature serialization processing; the voiceprint features and the semantic features are fused to obtain emotion semantic joint coding representation; according to an emotion quantitative index formula, identifying a user emotion level; and screening a verbal skill recommendation text corresponding to the latest user voice in combination with the user emotion level. On the basis of dialogue semantic features, voiceprint features during voice pronunciation are fully considered, and emotion analysis is performed in combination with voiceprint feature recognition, so that appropriate verbal skill texts are screened for dialogue according to real-time emotion changes of a user. When the method is applied to financial collection or medical condition notification scenes, financial or medical related employees can be more flexibly and scientifically assisted in business verbal skill selection, and emotional feelings of users can be better cared.
Owner:PING AN TECH (SHENZHEN) CO LTD

Man-machine conversation processing method and device, equipment and medium

The invention relates to the technical field of artificial intelligence and natural language processing, can be applied to the field of intelligent medical treatment and financial science and technology, and discloses a man-machine conversation processing method, device, equipment and medium. Obtaining a dialogue state map associated with the dialogue verbal skill stage and the dialogue content thereof; obtaining dialogue voice information of a current round, and performing semantic recognition according to the dialogue voice information of the current round to obtain a user intention and a verbal skill stage label corresponding to the dialogue of the current round; generating initial response content according to the user intention, and detecting and judging whether the initial response content is consistent with the response trend of the dialogue content in the corresponding verbal skill stage of the dialogue state graph or not by utilizing the dialogue state graph according to the verbal skill stage label corresponding to the current round of dialogue; if not, the response content is regenerated according to the user intention and the context dialogue voice information and the response trend.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent document question and answer and voice generation system and method

The invention relates to the technical field of artificial intelligence, in particular to an intelligent document question and answer and voice generation system and method.The method comprises the steps that corresponding text content is obtained according to a specific document identifier, cue words are generated in combination with the text content and user questions, a locally-deployed large language model is called to obtain answers, and the answers are sent to a server; the answer is returned to the user; receiving a voice generation instruction sent by a user for the question and answer history of the specific document identifier; obtaining a question and answer historical record according to the specific document identifier, and formatting the question and answer historical record into text data with a dialogue person distinguishing identifier; performing context analysis on the text data to generate corresponding speech synthesis control parameters; calling a locally deployed speech synthesis engine, and generating multi-role dialogue speech files with different timbres and expressions based on the text data and the control parameters; according to the invention, collaborative output and interaction of text and voice dual modes are realized, and information acquisition requirements of a user in different scenes are met.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Real-time full-duplex conversational speech-oriented streaming conversational state prediction method and system

Embodiments of the present application provide a kind of real-time full duplex voice dialogue-oriented streaming dialogue state prediction method and system.The method comprises: the processed speech mark sequence is input to the dialogue state prediction model of multimodal full duplex, carries out streaming dialogue state prediction, adopts staggered prediction mechanism, the current speech feature, text recognition result and dialogue state mark are modeled on time dimension, for realizing that text recognition result explicitly participates in the judgment process of dialogue state prediction, obtains the full duplex dialogue state mark of streaming continuous prediction;The interactive state of user is judged in real time based on the full duplex dialogue state mark of continuous prediction, to determine the control behavior of downstream dialogue.The embodiments of the present application reduce the overall computing overhead and real-time inference pressure of system, more accurately identify complex interactive behaviors such as user sentence completion, pause and interruption, etc.Improve the overall response performance and user experience of full duplex voice interaction system.
Owner:SHANGHAI JIAOTONG UNIV

Speech transcription model construction method and device, electronic equipment and storage medium

The invention relates to a voice transcription model construction method and device, electronic equipment and a storage medium. The voice transcription model construction method comprises the following steps: acquiring a medical inquiry dialogue text data set and a target voice data set; performing role segmentation processing on each medical interrogation dialogue text data to obtain doctor interrogation dialogue text data and patient interrogation dialogue text data; performing voice synthesis on the doctor interrogation dialogue text data and the patient interrogation dialogue text data based on each target voice data to obtain a plurality of doctor interrogation dialogue voice data and a plurality of patient interrogation dialogue voice data; a training data set is constructed based on the interrogation dialogue voice data of multiple doctors and the interrogation dialogue voice data of multiple patients, a to-be-trained voice transcription model is trained, and a trained voice transcription model is obtained, so that a large number of training data sets can be constructed for model training without annotation, and the training efficiency is improved. And the construction efficiency and accuracy of the model are improved.
Owner:BEIJING ECON MEDICAL TECHNOLOGY CO LTD

Medical record support apparatus, system, method, and program

To improve efficiency in creating a medical record on the basis of a medical interview.SOLUTION: A medical record support apparatus includes: first acquisition means for acquiring text information of an utterance content converted by speech recognition in which speakers are distinguished for conversational speech data between a medical worker and a patient during a medical interview; generation means for generating an instruction sentence for a language model to be output in a medical record format from the text information; and second acquisition means for acquiring medical record information of the patient by inputting the instruction sentence to the language model.SELECTED DRAWING: Figure 1
Owner:NATIONAL HEALTH CRISIS MANAGEMENT RESEARCH INSTITUTE