Speech data construction method and device, electronic equipment and storage medium
By translating and synthesizing foreign language medical consultation dialogue text data into Chinese, and utilizing Chinese voice data from non-medical consultation scenarios, the high equipment and labor costs of existing technologies are solved, enabling efficient acquisition of high-quality medical consultation voice data.
Patent Information
- Application Number
- CN202511019756.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2045-07-23
AI Technical Summary
Existing technologies require the installation of high-sensitivity recording equipment in each consultation room when constructing a speech transcription model for medical consultations. This results in high equipment costs and a large consumption of human resources. At the same time, manual review and annotation of consultation speech data are required, leading to low acquisition efficiency and high costs, and is also subject to restrictions on hospital privacy protection.
By acquiring foreign language and Chinese medical consultation dialogue text datasets, translating and synthesizing them into Chinese medical consultation dialogue speech data, and utilizing Chinese speech data from non-medical consultation scenarios for speech synthesis, the need to install recording equipment in each consultation room is avoided, reducing equipment costs, and data acquisition efficiency is improved through automated processing.
It enables the acquisition of high-quality medical consultation voice data at low cost and high efficiency, reduces equipment and labor costs, overcomes hospital privacy protection restrictions, and improves data acquisition efficiency.
Smart Images

Figure CN121148361B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of medical consultation technology, and in particular to a method, apparatus, electronic device and storage medium for constructing voice data. Background Technology
[0002] Currently, in the process of informatization and intelligentization of modern medicine, intelligent medical consultation is one of the core links for accurate diagnosis and treatment of patients. During the consultation process between doctors and patients, it is necessary to transcribe the doctor's and patient's voice dialogue into text, and then generate electronic medical record information based on the transcribed text, thereby providing a basic artificial intelligence-assisted tool for accurate diagnosis and treatment in modern medicine.
[0003] Current methods for speech transcription in the field of medical consultations primarily rely on building machine learning models for speech transcription. However, building these models requires a large amount of high-quality medical consultation speech data for training. Current methods for collecting this data involve using high-sensitivity and noise-resistant recording equipment in hospital clinics to overcome the noisy environment. This requires installing recording equipment in multiple clinics, leading to high equipment and manpower costs, as well as restrictions related to hospital privacy, making it difficult to collect large volumes of consultation speech data. Furthermore, the collected data requires manual review and processing to ensure accuracy and quality, and manual annotation is needed during the subsequent generation of training data, resulting in low training data acquisition efficiency and high labor costs.
[0004] Therefore, how to obtain high-quality medical consultation voice data at low cost and high efficiency is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] To address the aforementioned technical problems, this disclosure provides a method, apparatus, electronic device, and storage medium for constructing voice data.
[0006] A first aspect of this disclosure provides a method for constructing voice data, including:
[0007] Acquire a medical consultation dialogue text dataset in a first foreign language, a medical consultation dialogue text dataset in a second foreign language, a medical consultation dialogue text dataset in Chinese, and a target speech dataset. The target speech dataset is a collection of Chinese speech data collected in non-medical consultation scenarios.
[0008] Each medical consultation dialogue text in the first foreign language and each medical consultation dialogue text in the second foreign language were translated to obtain multiple second Chinese medical consultation dialogue texts.
[0009] Based on the target speech data in the target speech dataset, speech synthesis is performed on each first Chinese medical consultation dialogue text data and each second Chinese medical consultation dialogue text data to obtain multiple medical consultation dialogue speech data.
[0010] A second aspect of this disclosure provides a voice data construction apparatus, comprising:
[0011] The data acquisition unit is used to acquire a medical consultation dialogue text dataset in a first foreign language, a medical consultation dialogue text dataset in a second foreign language, a medical consultation dialogue text dataset in Chinese, and a target speech dataset. The target speech dataset is a collection of Chinese speech data collected in non-medical consultation scenarios.
[0012] The text translation unit is used to translate each medical consultation dialogue text data in the first foreign language and each medical consultation dialogue text data in the second foreign language, respectively, to obtain multiple second Chinese medical consultation dialogue text data.
[0013] The speech data construction unit is used to synthesize speech for each first Chinese medical consultation dialogue text data and each second Chinese medical consultation dialogue text data based on each target speech data in the target speech dataset, so as to obtain multiple medical consultation dialogue speech data.
[0014] A third aspect of this disclosure provides an electronic device, including:
[0015] processor;
[0016] Memory, used to store executable instructions;
[0017] The processor is used to read executable instructions from memory and execute the executable instructions to implement the voice data construction method provided in the first aspect above.
[0018] A fourth aspect of this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the voice data construction method provided in the first aspect.
[0019] A fifth aspect of this disclosure provides a computer program product comprising a computer program or instructions that, when executed by a processor, implement the voice data construction method of the first aspect described above.
[0020] The technical solution provided in this disclosure has the following advantages compared with the prior art:
[0021] The speech data construction method, apparatus, electronic device, and storage medium provided in this disclosure can acquire a first foreign language medical consultation dialogue text dataset, a second foreign language medical consultation dialogue text dataset, a first Chinese medical consultation dialogue text dataset, and a target speech dataset when constructing speech data. The target speech dataset is a collection of Chinese speech data collected in non-medical consultation scenarios. Then, each first foreign language medical consultation dialogue text dataset and each second foreign language medical consultation dialogue text dataset are translated to obtain multiple second Chinese medical consultation dialogue text datasets. Finally, based on each target speech dataset in the target speech dataset, speech synthesis is performed on each first Chinese medical consultation dialogue text dataset and each second Chinese medical consultation dialogue text dataset to obtain multiple medical consultation dialogue speech datasets. Therefore, it is possible to acquire multilingual medical consultation dialogue text data, and translate medical consultation dialogue text data in languages other than Chinese into Chinese format, thereby obtaining a large amount of Chinese medical consultation dialogue text data. Then, by using Chinese speech data from non-medical consultation scenarios, speech synthesis is performed on the obtained Chinese medical consultation dialogue text data, resulting in multiple medical consultation dialogue speech data. In the process of acquiring high-quality, large-scale medical consultation dialogue speech data, there is no need to install recording equipment in each consultation room, nor is it subject to hospital privacy protection restrictions, thus reducing equipment costs. At the same time, no manual review and processing is required, reducing labor costs and improving the efficiency of speech data acquisition. Attached Figure Description
[0022] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0023] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart of a voice data construction method provided in an embodiment of this disclosure;
[0025] Figure 2 This is a schematic diagram illustrating the processing of medical consultation dialogue text data according to an embodiment of the present disclosure;
[0026] Figure 3This is a schematic diagram illustrating a foreign language medical consultation text translation processing procedure provided in an embodiment of this disclosure;
[0027] Figure 4 This is a flowchart of a speech synthesis method provided in an embodiment of this disclosure;
[0028] Figure 5 This is a schematic diagram illustrating role segmentation processing of medical consultation dialogue text data according to an embodiment of this disclosure;
[0029] Figure 6 This is a schematic diagram of the structure of a speech synthesis model provided in an embodiment of this disclosure;
[0030] Figure 7 This is a schematic diagram of speech synthesis and splicing provided in an embodiment of this disclosure;
[0031] Figure 8 This is a schematic diagram of post-processing of medical consultation voice data provided in an embodiment of this disclosure;
[0032] Figure 9 This is a flowchart of another voice data construction method provided in this disclosure embodiment;
[0033] Figure 10 This is a schematic diagram of the structure of a voice data construction device provided in an embodiment of this disclosure;
[0034] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0035] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0036] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0037] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0038] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0039] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0040] In the process of informatization and intelligentization in modern medicine, intelligent medical consultation is one of the core links for accurate diagnosis and treatment of patients. During the consultation between doctors and patients, it is necessary to transcribe the doctor's and patient's voice dialogue into text, and then generate electronic medical record information based on the transcribed text, thereby providing a basic artificial intelligence-assisted tool for accurate diagnosis and treatment in modern medicine.
[0041] In the existing field of medical consultation, speech transcription mainly relies on building machine learning models for speech transcription. However, building these models requires a large amount of high-quality medical consultation speech data for training. Current methods for collecting this data involve using high-sensitivity and noise-resistant recording equipment in hospital clinics to overcome the noisy environment. This requires installing recording equipment in different clinics, leading to high equipment and manpower costs, as well as restrictions related to hospital privacy, making it difficult to collect large amounts of consultation speech data. Furthermore, the collected data requires manual review and processing to ensure accuracy and quality, and manual annotation is needed for subsequent training data generation, resulting in low training data acquisition efficiency and high labor costs. To address these issues, this disclosure provides a speech data construction method, which is described below with specific embodiments.
[0042] Figure 1This is a flowchart of a voice data construction method provided in an embodiment of the present disclosure. The method can be executed by a voice data construction device, which can be implemented in software and / or hardware. The voice data construction device can be configured in an electronic device, such as a server or terminal, wherein the terminal specifically includes a mobile phone, computer or tablet computer, etc.
[0043] like Figure 1 As shown, the voice data construction method provided in this embodiment includes the following steps.
[0044] S110. Obtain the first foreign language medical consultation dialogue text dataset, the second foreign language medical consultation dialogue text dataset, the first Chinese medical consultation dialogue text dataset, and the target speech dataset. The target speech dataset is a collection of Chinese speech data collected in non-medical consultation scenarios.
[0045] In this embodiment of the disclosure, the medical consultation dialogue text dataset can be understood as a collection of text data from conversations between doctors and patients in a medical consultation scenario, containing multiple medical consultation dialogue text datasets. This collection may include dialogue text data from different departments such as internal medicine, surgery, pediatrics, obstetrics and gynecology, and oncology. The data is complete and diverse.
[0046] A first-language medical consultation dialogue text dataset can be a medical consultation dialogue text dataset in a language other than a second foreign language. For example, it could be a French medical consultation dialogue text dataset, a German medical consultation dialogue text dataset, etc.
[0047] The second foreign language medical consultation dialogue text dataset can be an English medical consultation dialogue text dataset.
[0048] A target speech dataset can be understood as a collection of speech data collected from common everyday scenarios, containing multiple target speech datasets. This collection can include speech data sets with different sound frequencies, loudnesses, speech rates, and timbres to ensure the diversity of the speech data. For example, in terms of sound frequency, speech from different frequencies such as tenor, baritone, bass, soprano, mezzo-soprano, and contralto can be collected to ensure that the speech data reflects the broad spectrum of the human voice; in terms of speech rate, speech from different speeds such as fast, medium, and slow can be collected to reflect the natural variations in speech communication in various everyday scenarios; in terms of sound loudness, speech from different loudnesses such as high, medium, and low can be collected to expand the coverage of speech data across different loudness levels; and in terms of timbre, speech from different timbres such as deep, clear, nasal, vibrato, and hoarse voices can be collected to enrich the timbre diversity of the dataset.
[0049] In some embodiments of this disclosure, the electronic device can respond to a voice data construction instruction by parsing the instruction, and when the instruction contains data identification information, obtaining the data identification information corresponding to the instruction. Based on the data identification information, it can retrieve a first foreign language medical consultation dialogue text dataset, a second foreign language medical consultation dialogue text dataset, a first Chinese medical consultation dialogue text dataset, and a target voice dataset from a preset database. The voice data construction instruction can be understood as an instruction used to construct medical consultation voice data.
[0050] In other embodiments of this disclosure, the electronic device may respond to a voice data construction instruction, parse the voice data construction instruction, and when the voice data construction instruction does not contain data identification information, acquire data from a preset open source website to acquire a first foreign language medical consultation dialogue text dataset, a second foreign language medical consultation dialogue text dataset, a first Chinese medical consultation dialogue text dataset, and a target voice dataset. The preset open source website may be an open source website without copyright restrictions.
[0051] All operations involving the acquisition of information or data in this disclosure were conducted in accordance with the relevant data protection laws and hospital ethics policies of the country where the information or data was located, and with the authorization of the relevant information or data owner.
[0052] S120. Translate each medical consultation dialogue text in the first foreign language and each medical consultation dialogue text in the second foreign language respectively to obtain multiple second Chinese medical consultation dialogue texts.
[0053] Specifically, after acquiring the first foreign language medical consultation dialogue text dataset and the second foreign language medical consultation dialogue text dataset, the electronic device can translate each first foreign language medical consultation dialogue text dataset and each second foreign language medical consultation dialogue text dataset into Chinese, thereby obtaining the second Chinese medical consultation dialogue text data corresponding to each first foreign language medical consultation dialogue text dataset and each second foreign language medical consultation dialogue text dataset.
[0054] S130. Based on each target speech data in the target speech dataset, speech synthesis is performed on each first Chinese medical consultation dialogue text data and each second Chinese medical consultation dialogue text data to obtain multiple medical consultation dialogue speech data.
[0055] In this embodiment of the disclosure, speech synthesis can be understood as using target speech data as speech seed features to perform speech conversion processing on various Chinese medical consultation dialogue text data, converting the medical consultation dialogue text data into medical consultation dialogue speech data.
[0056] Specifically, after obtaining multiple first Chinese medical consultation dialogue text data and multiple second Chinese medical consultation dialogue text data, the electronic device can input each first and second Chinese medical consultation dialogue text data into a trained speech synthesis model for speech synthesis, thereby obtaining multiple corresponding medical consultation dialogue speech data. The trained speech synthesis model can be any machine learning model used for speech synthesis processing, and no restrictions are imposed here.
[0057] In this embodiment of the disclosure, the speech synthesis model can select any segment of target speech as the clone source of medical consultation dialogue speech, and then perform speech synthesis processing to obtain multiple medical consultation dialogue speech data.
[0058] Therefore, by using speech synthesis, a large amount of medical consultation dialogue voice data can be obtained with a certain amount of Chinese medical consultation dialogue text data, so as to meet the complexity and diversity of doctor and patient voice dialogue in actual medical consultation scenarios.
[0059] In this embodiment of the disclosure, when constructing voice data, a first foreign language medical consultation dialogue text dataset, a second foreign language medical consultation dialogue text dataset, a first Chinese medical consultation dialogue text dataset, and a target voice dataset can be obtained. The target voice dataset is a collection of Chinese voice data collected in non-medical consultation scenarios. Then, each first foreign language medical consultation dialogue text dataset and each second foreign language medical consultation dialogue text dataset are translated to obtain multiple second Chinese medical consultation dialogue text datasets. Finally, based on each target voice dataset in the target voice dataset, speech synthesis is performed on each first Chinese medical consultation dialogue text dataset and each second Chinese medical consultation dialogue text dataset to obtain multiple medical consultation dialogue voice datasets. Therefore, it is possible to acquire multilingual medical consultation dialogue text data, and translate medical consultation dialogue text data in languages other than Chinese into Chinese format, thereby obtaining a large amount of Chinese medical consultation dialogue text data. Then, by using Chinese speech data from non-medical consultation scenarios, speech synthesis is performed on the obtained Chinese medical consultation dialogue text data, resulting in multiple medical consultation dialogue speech data. In the process of acquiring high-quality, large-scale medical consultation dialogue speech data, there is no need to install recording equipment in each consultation room, nor is it subject to hospital privacy protection restrictions, thus reducing equipment costs. At the same time, no manual review and processing is required, reducing labor costs and improving the efficiency of speech data acquisition.
[0060] Based on the above embodiments of this disclosure, obtaining a first foreign language medical consultation dialogue text dataset, a second foreign language medical consultation dialogue text dataset, a first Chinese medical consultation dialogue text dataset, and a target speech dataset may specifically include: obtaining the original first foreign language medical consultation dialogue text dataset, the original second foreign language medical consultation dialogue text dataset, the original first Chinese medical consultation dialogue text dataset, and the original speech dataset; performing a first processing on each original first foreign language medical consultation dialogue text dataset, the original second foreign language medical consultation dialogue text dataset, and the original Chinese medical consultation dialogue text dataset, and performing a second processing on each original speech dataset to obtain the first foreign language medical consultation dialogue text dataset, the second foreign language medical consultation dialogue text dataset, the first Chinese medical consultation dialogue text dataset, and the target speech dataset.
[0061] In this embodiment of the disclosure, the first process may include data cleaning, target word removal, and structural transformation; the second process includes data noise reduction and invalid speech data trimming.
[0062] In this embodiment of the disclosure, the target word removal process may include removing sensitive information such as hospital information, patient information, and medical treatment information, as well as medical condition description information outside of the doctor-patient consultation dialogue, in order to protect patient privacy and ensure the effectiveness of the dialogue.
[0063] Structure transformation processing can be understood as converting text data into structured text data of alternating multi-turn dialogues between doctors and patients. The structured characteristics of multi-turn dialogues in text data can ensure that the alternating multi-turn dialogue pattern between doctors and patients during consultations can be realistically simulated when synthesizing speech data, thereby improving the realism and usability of speech datasets and providing data support for training speech transcription models that closely resembles real hospital consultation scenarios.
[0064] Data denoising can be understood as removing noise from the original speech data to reduce noise interference, ensure the clarity of the speech signal, and avoid excessive noise interfering with the human voice signal, which would make it difficult for the speech transcription model to learn and capture sound features, and ultimately the trained model would be unable to accurately transcribe speech.
[0065] Invalid speech data trimming can be understood as trimming invalid speech data, such as silent parts or non-human voice data, from the original speech data.
[0066] It should be noted that the execution order of the first processing of each original first foreign language medical consultation dialogue text data, the original second foreign language medical consultation dialogue text data, and the original Chinese medical consultation dialogue text data and the second processing of each original speech data in the original speech dataset is not restricted; they can be executed simultaneously or in any order.
[0067] Figure 2 This is a schematic diagram illustrating the processing of medical consultation dialogue text data according to an embodiment of this disclosure, such as... Figure 2 As shown, the original Chinese medical consultation dialogue text data is used as an example for explanation. The left side is the original medical consultation dialogue text data. The original medical consultation dialogue text data has been processed by removing information such as address, consultation time, patient, and disease description. At the same time, the format has been converted to obtain the processed medical consultation dialogue text data on the right.
[0068] In this embodiment, the original medical consultation dialogue text data corresponding to different languages can be processed by data cleaning, target word removal, and structural transformation. At the same time, the original speech data can be processed by data noise reduction and invalid speech data trimming. This improves the accuracy and usability of the processed medical consultation dialogue text data and target speech data corresponding to different languages, and also makes them conform to the format and structure of medical consultation dialogue, providing support for subsequent speech data construction.
[0069] In this embodiment of the disclosure, each first foreign language medical consultation dialogue text data and each second foreign language medical consultation dialogue text data are translated to obtain multiple second Chinese medical consultation dialogue text data. Specifically, this may include: for each first foreign language medical consultation dialogue text data, performing language conversion processing on the first foreign language medical consultation dialogue text data, converting the first foreign language corresponding to the first foreign language medical consultation dialogue text data into the second foreign language, to obtain the target medical consultation dialogue text data corresponding to the first foreign language medical consultation dialogue text data; translating the target medical consultation dialogue text data and the second foreign language corresponding to each second foreign language medical consultation dialogue text data into Chinese, to obtain multiple second Chinese medical consultation dialogue text data.
[0070] In some embodiments of this disclosure, language conversion processing is performed on the medical consultation dialogue text data in a first foreign language to convert the first foreign language corresponding to the medical consultation dialogue text data in a first foreign language into a second foreign language, thereby obtaining the target medical consultation dialogue text data corresponding to the medical consultation dialogue text data in the first foreign language. Specifically, this may include: inputting the medical consultation dialogue text data in the first foreign language and preset prompt words into a trained text translation model, and having the trained text translation model perform language conversion processing on the medical consultation dialogue text data in the first foreign language based on the preset prompt words to obtain the target medical consultation dialogue text data.
[0071] In this embodiment of the disclosure, the preset prompt word can be a word used to prompt for the source language and the language to be converted. For example, when converting French to English, the preset prompt word can be "Please translate French to English" or "translate French to English".
[0072] The trained text translation model can be based on an encoder-decoder network structure, and can be used to convert text data from a first foreign language to a second foreign language.
[0073] In this embodiment of the disclosure, obtaining the trained text translation model may specifically include: obtaining a training dataset, wherein each training data in the training dataset includes a first foreign language text and a second foreign language text corresponding to the first foreign language text; for each training data in the training dataset, inputting the training data and preset prompt words into the text translation model to be trained, whereby the encoder in the text translation model to be trained encodes the first foreign language text, and inputs the resulting foreign language feature vector sequence and preset prompt words into the decoder of the text translation model to be trained, whereby the decoder translates the foreign language feature vector sequence based on the preset prompt words and outputs predicted text; calculating the loss based on the predicted text, the second foreign language text, and a preset loss function to obtain a loss value; and adjusting the parameters of the text translation model to be trained based on the loss value until convergence, thereby obtaining the trained text translation model.
[0074] In this embodiment, the first and second foreign language texts included in the training data can be books, journal articles, etc., from all medical fields, and are not limited to medical consultation dialogue text data. Therefore, it can cover the main linguistic and semantic phenomena and knowledge structures in the medical field, and can be better applied to the training of the text translation model to be trained.
[0075] In this embodiment of the disclosure, the preset loss function can be the multi-class cross-entropy loss function, and the specific formula is as follows:
[0076]
[0077] Where Loss is the loss value; N is the number of samples in the training dataset; y i The predicted value for the i-th sample is the predicted text. Let be the true label of the i-th sample, i.e., the second foreign language text.
[0078] Specifically, before performing language conversion processing on the medical consultation dialogue text data in a first foreign language, converting the first foreign language corresponding to the first foreign language medical consultation dialogue text data into a second foreign language, and obtaining the target medical consultation dialogue text data corresponding to the first foreign language medical consultation dialogue text data, the electronic device can acquire a training dataset for training the text translation model to be trained. The first foreign language text in the training data is input into the encoder of the text translation model to be trained. The encoder encodes the first foreign language text, including positional encoding, to obtain a foreign language feature vector sequence. The text feature vector sequence is then input into the decoder through a cross-attention mechanism, and a preset prompt word is also input into the decoder. The decoder translates the foreign language feature vector sequence based on the preset prompt word and outputs the predicted text. Then, loss is calculated based on the predicted text, the second foreign language text in the training data, and a preset loss function to obtain the loss value. The parameters in the text translation model to be trained are adjusted based on the loss value until convergence, that is, the loss value reaches the preset loss threshold, resulting in the trained text translation model.
[0079] In other embodiments of this disclosure, language conversion processing is performed on the medical consultation dialogue text data in the first foreign language, converting the first foreign language corresponding to the medical consultation dialogue text data in the first foreign language into a second foreign language to obtain the target medical consultation dialogue text data corresponding to the medical consultation dialogue text data in the first foreign language. Specifically, this may include: performing entity recognition on the medical consultation dialogue text data in the first foreign language to obtain medical entities; querying a multilingual version of a medical domain knowledge graph based on the identified medical entities to recall knowledge graph data related to the medical consultation dialogue text data in the first foreign language; converting the format of the knowledge graph data to the same format as the medical consultation dialogue text data in the first foreign language; concatenating the converted knowledge graph data with the medical consultation dialogue text data in the first foreign language to obtain concatenated text data; inputting the concatenated text data into a target encoder for encoding processing to obtain a target feature vector corresponding to the concatenated text data; and inputting the target feature vector into a target decoder, whereby the target decoder obtains the target medical consultation dialogue text data corresponding to the medical consultation dialogue text data in the first foreign language based on the target feature vector and the context information of the second foreign language.
[0080] Among them, medical entities can include objects corresponding to language text, such as whether it is a doctor or a patient, keywords corresponding to the consultation, such as keywords related to the illness, such as cough, fever, etc.
[0081] The medical knowledge graph includes the relationships between various entities in different clinics. For example, querying the entity "fever" will provide detailed information about fever.
[0082] Knowledge graph data can be attribute information related to medical entities.
[0083] The contextual information of a second foreign language can include information such as the grammar and vocabulary rules of the second foreign language.
[0084] Furthermore, after obtaining the trained text translation model, the electronic device can evaluate the performance of the trained text translation model based on bilingual substitution scoring and multi-word recall rate.
[0085] The Bilingual Evaluation Understudy (BLEU) score is used to calculate the similarity between two sentences, measuring the similarity between the two texts. A higher BLEU score indicates a higher similarity between the two sentences, resulting in better performance of the trained text translation model.
[0086] Recall-Oriented Understudy for Gisting Evaluation (ROUGE-N) is used to calculate the similarity between multiple characters, and it measures the similarity between two texts based on the similarity between these characters. The higher the ROUGE-N value, the higher the similarity between the two sentences, and thus the better the performance of the trained text translation model.
[0087] Specifically, the electronic device can calculate the bilingual substitution score and multi-word recall rate between the text translation results output by the trained text translation model and the real labels. Based on the weights corresponding to the bilingual substitution score and multi-word recall rate, a weighted sum is performed to obtain the final score. If the final score is greater than a preset score threshold, it is determined that the performance of the trained text translation model meets the requirements. Otherwise, it is determined that the performance of the trained text translation model does not meet the requirements, and training continues until the performance of the trained text translation model meets the requirements.
[0088] In this embodiment of the disclosure, by evaluating the performance of the trained text translation model, the accuracy and reliability of text translation during the text translation process (i.e., converting the first foreign language into the second foreign language) are ensured.
[0089] In this embodiment, the target medical consultation dialogue text data and the corresponding second foreign language medical consultation dialogue text data for each second foreign language medical consultation dialogue text data are translated into Chinese to obtain multiple second Chinese medical consultation dialogue text data. The specific implementation method is similar, and a target text translation model for translating second foreign languages into Chinese can be used for processing. The specific acquisition method and performance evaluation of the target text translation model are similar to the specific implementation method and performance evaluation method for acquiring the trained text translation model in the above embodiments of this disclosure, and will not be repeated here.
[0090] Figure 3 This is a schematic diagram of a foreign language medical consultation text translation processing procedure provided in an embodiment of this disclosure, such as... Figure 3 As shown, the explanation uses French as the first foreign language and English as the second foreign language. First, the French medical consultation dialogue text data is translated into English, and then the obtained English medical consultation dialogue text data is translated into Chinese medical consultation dialogue text data, thereby improving the accuracy of the obtained Chinese medical consultation dialogue text data.
[0091] In this embodiment of the disclosure, when acquiring Chinese medical consultation dialogue text data, the scale and development history of medical fields in various countries are taken into account. Considering the structure, semantics and syntax rules between different languages, the first foreign language corresponding to the first foreign language medical consultation dialogue text data is first translated into a second foreign language such as English. Then, the target medical consultation dialogue text data in the second foreign language and the second foreign language medical consultation dialogue text data are translated into Chinese. This avoids the problem of low translation accuracy caused by directly translating the first foreign language into Chinese, thereby improving the accuracy of the obtained Chinese medical consultation dialogue text data.
[0092] Figure 4 This is a flowchart of a speech synthesis method provided in an embodiment of this disclosure, such as... Figure 4 As shown, speech synthesis is performed on each first Chinese medical consultation dialogue text data and each second Chinese medical consultation dialogue text data based on the target speech data in the target speech dataset to obtain multiple medical consultation dialogue speech data. Specifically, it may include the following steps:
[0093] S410. Perform deduplication on each first Chinese medical consultation dialogue text data and each second Chinese medical consultation dialogue text data to obtain multiple target Chinese medical consultation dialogue text data after deduplication.
[0094] In this embodiment, deduplication can be understood as removing duplicates from the first and second Chinese medical consultation dialogue text data that have the same text or a similarity higher than a preset similarity threshold. This avoids the waste of resources caused by repeatedly processing Chinese medical consultation dialogue text data that have the same text or a similarity higher than the preset similarity threshold.
[0095] S420. For each target Chinese medical consultation dialogue text data, perform role segmentation processing on the target Chinese medical consultation dialogue text data to obtain doctor consultation dialogue text data and patient consultation dialogue text data.
[0096] In this embodiment of the disclosure, role segmentation processing can be understood as segmenting the target Chinese medical consultation dialogue text data according to different roles, and dividing the dialogue between doctors and patients.
[0097] Figure 5 This is a schematic diagram illustrating role segmentation processing of medical consultation dialogue text data according to an embodiment of this disclosure, such as... Figure 5 As shown, the left side represents the Chinese medical consultation dialogue text data (i.e., the target Chinese medical consultation dialogue text data in this embodiment). After role segmentation processing, the medical consultation dialogue text data corresponding to the doctor and the medical consultation dialogue text data corresponding to the patient are obtained, i.e., the segmented medical consultation dialogue text data on the right. This allows for the separation of medical consultation dialogue text data for different roles, facilitating subsequent speech synthesis.
[0098] S430. The doctor's consultation dialogue text data and the patient's consultation dialogue text data are respectively combined with each target speech data in the target speech dataset to obtain multiple doctor's consultation dialogue speech data and multiple patient's consultation dialogue speech data corresponding to the doctor's consultation dialogue text data.
[0099] In this embodiment of the disclosure, the doctor's consultation dialogue text data is combined with various target speech data in the target speech dataset to perform speech synthesis, resulting in multiple doctor's consultation dialogue speech data corresponding to the doctor's consultation dialogue text data. Specifically, this may include: for each target speech data, inputting the doctor's consultation dialogue text data into a trained speech synthesis model, whereby the text encoder preprocessing network in the encoding module of the trained speech synthesis model performs a first processing on the doctor's consultation dialogue text data, and inputting the resulting text feature space vector into the shared encoder in the encoding module, whereby the shared encoder performs feature space transformation processing on the text feature space vector to obtain a first feature space vector, and inputting the first feature space vector into the trained speech synthesis model. The target speech data is input into the shared decoder in the decoding module of the trained speech synthesis model. The speech decoder preprocessing module performs a second processing on the target speech data and inputs the resulting speech feature space vector into the shared decoder in the decoding module. The shared decoder performs feature space transformation on the speech feature space vector to obtain a second feature space vector. The first and second feature space vectors are then processed a third time to obtain the target feature space vector. The target feature space vector is then input into the speech decoder postprocessing module for speech synthesis processing to obtain multiple doctor consultation dialogue speech data.
[0100] Figure 6 This is a schematic diagram of the structure of a speech synthesis model provided in an embodiment of this disclosure, such as... Figure 6 As shown, the speech synthesis model includes an encoding module and a decoding module. The encoding module includes a first speech input module, a first text input module, a speech encoder preprocessing network, a text encoder preprocessing network, and a shared encoder; the decoding module includes a second speech input module, a second text input module, a speech decoder preprocessing network, a text decoder preprocessing network, a shared decoder, a speech decoder postprocessing network, a text decoder postprocessing network, a speech output module, and a text output module.
[0101] The first voice input module and the second voice input module are used to input voice data, such as the target voice data in the embodiments of this disclosure.
[0102] The first text input module and the second text input module are used to input text data, such as doctor consultation dialogue text data and patient consultation dialogue text data in the embodiments of this disclosure.
[0103] The speech encoder preprocessing network and the speech decoder preprocessing network are used to preprocess the input speech data and extract features, respectively, converting the speech data into a speech feature space vector.
[0104] The text encoder preprocessing network and the text decoder preprocessing network are used to preprocess the input text data and extract features, respectively, converting the text data into a text feature space vector.
[0105] A shared encoder and a shared decoder are used to convert text feature space vectors and / or speech feature space vectors into a unified speech-text multimodal feature space vector.
[0106] The speech decoder post-processing network and the text decoder post-processing network are used to perform speech decoding or text decoding on the speech-text multimodal feature space vectors output by the shared decoder, respectively, and convert them into speech prediction data or text prediction data.
[0107] The speech output module is used to output the speech prediction data output by the speech decoder post-processing network.
[0108] The text output module is used to output the text prediction data output by the text decoder post-processing network.
[0109] In this embodiment of the disclosure, the first processing and the second processing are to preprocess and extract features from the input data (including voice data and / or text data).
[0110] The third processing can be understood as the process of converting feature space vectors into speech data, i.e., decoding processing.
[0111] It should be noted that the specific implementation method for synthesizing the patient consultation dialogue text data in each target Chinese medical consultation dialogue text data with the target speech data is similar to the specific implementation method for synthesizing the doctor consultation dialogue text data with the target speech data in the embodiments of this disclosure, and will not be repeated here.
[0112] In this embodiment of the disclosure, the method for obtaining the trained speech synthesis model, i.e., the specific method for training the speech synthesis model, is as follows: A speech-text multimodal dataset is obtained, wherein the speech-text multimodal dataset includes text-speech data pairs, speech-speech data pairs, text-text data pairs, and speech-text data pairs; target processing operations are performed on the speech-text multimodal dataset to obtain a processed speech-text multimodal dataset, wherein the target processing operations include word segmentation and text encoding processing of the text data in the dataset using a first algorithm, and feature extraction of the speech data in the dataset using a second algorithm; based on the processed speech-text multimodal dataset, self-supervised learning is performed on the encoding and decoding modules in the speech synthesis model to be trained to obtain a first speech synthesis model, wherein the first speech synthesis model can... The first speech synthesis model maps speech and text information to a shared multimodal feature space. Based on downstream tasks such as speech synthesis and speech transcription, supervised training is performed on the speech synthesis model to be trained. Specifically, the computational relationship between the first speech-text multimodal feature space obtained by the encoding module processing the first input data and the second speech-text multimodal feature space obtained by the decoding module processing the second input data is determined based on the cross-attention mechanism and contrastive learning. The first speech synthesis model is iteratively optimized by using a weighted average loss function of multiple downstream tasks until the model converges, resulting in the second speech synthesis model. After obtaining the second speech synthesis model, it is fine-tuned based on a speech dataset in the medical field to obtain the trained speech synthesis model.
[0113] S440. Based on multiple doctor consultation dialogue voice data and multiple patient consultation dialogue voice data, voice data is constructed to obtain multiple medical consultation dialogue voice data.
[0114] In this embodiment of the disclosure, multiple medical consultation dialogue voice data are constructed based on multiple doctor consultation dialogue voice data and multiple patient consultation dialogue voice data. Specifically, this may include: for each doctor consultation dialogue voice data, splicing the doctor consultation dialogue voice data and the corresponding patient consultation dialogue voice data to obtain spliced medical consultation dialogue voice data; adding a preset length of silence interval segment between each doctor consultation dialogue voice data and each patient consultation dialogue voice data in the spliced medical consultation dialogue voice data, and adding a start timestamp and an end timestamp to each doctor consultation dialogue voice data and each patient consultation dialogue voice data to obtain multiple medical consultation dialogue voice data.
[0115] In this embodiment of the disclosure, the preset length corresponding to the silence interval segment can be adaptively set according to the user's needs, and is not limited here.
[0116] Specifically, after acquiring the spliced medical consultation dialogue voice data, the electronic device adds a preset length of silent interval segment between each doctor's consultation dialogue voice and each patient's consultation dialogue voice to simulate the natural pauses in real medical consultation dialogue scenarios, thereby obtaining a coherent and complete medical consultation dialogue voice data.
[0117] Figure 7 This is a schematic diagram of a speech synthesis and splicing method provided in an embodiment of this disclosure, such as... Figure 7 As shown, the doctor's consultation dialogue voice data and the corresponding patient consultation dialogue voice data are concatenated. A silence interval segment is added after each doctor's consultation dialogue voice data, and similarly, a silence interval segment is added after each patient consultation dialogue voice data.
[0118] Furthermore, after adding silence interval segments, start and end timestamps were added to each doctor's consultation voice data and each patient's consultation voice data. For example... Figure 8 As shown, a segment of audio data, such as a medical consultation dialogue, can be pre-loaded. Using a pre-defined audio processing tool like Praat or other open-source tools, the audio data is tagged at the sentence level, recording the start and end timestamps of each sentence segment. This data is then mapped to corresponding text data, and stored. Figure 8 The storage method shown maps audio segments to corresponding text sentences one-to-one and adds start and end timestamps, which facilitates the construction of training datasets for speech transcription models.
[0119] In this embodiment of the disclosure, a large amount of medical consultation dialogue voice data can be obtained by synthesizing a large amount of Chinese medical consultation dialogue text data with Chinese voice data in non-medical consultation scenarios. This satisfies the complexity and diversity of voice dialogue between doctors and patients in actual medical consultation scenarios, and makes the constructed medical consultation dialogue voice data have significant data scale and data quality.
[0120] Furthermore, after acquiring multiple medical consultation dialogue voice data, the electronic device can construct a target training dataset based on the multiple medical consultation dialogue voice data, train the speech transcription model to be trained based on the target training dataset, and obtain the trained speech transcription model; and evaluate the performance of the trained speech transcription model.
[0121] Specifically, the performance evaluation of the trained speech transcription model can be based on the character error rate between the target predicted text data output by the trained speech transcription model and the real text data. If the character error rate is less than the preset error rate threshold, the performance evaluation result is determined to be passed; otherwise, the performance evaluation result is determined to be failed, and the model training operation continues until the performance evaluation result of the trained speech transcription model is passed.
[0122] Therefore, it is possible to train a speech transcription model using multiple medical consultation dialogue voice data, thereby improving the accuracy of the obtained speech transcription model.
[0123] Figure 9 This is a flowchart of another voice data construction method provided in this disclosure embodiment, such as... Figure 9 As shown, the voice data construction method may include the following steps:
[0124] S910. Obtain the medical consultation dialogue text dataset in the first foreign language, the medical consultation dialogue text dataset in the second foreign language, the medical consultation dialogue text dataset in the first Chinese language, and the target speech dataset. The target speech dataset is a collection of Chinese speech data collected in non-medical consultation scenarios.
[0125] S920. Input the medical consultation dialogue text data in the first foreign language and the preset prompt words into the trained text translation model. The trained text translation model performs language conversion processing on the medical consultation dialogue text data in the first foreign language based on the preset prompt words, converting the first foreign language corresponding to the medical consultation dialogue text data in the first foreign language into the second foreign language, and obtaining the target medical consultation dialogue text data.
[0126] S930. Translate the target medical consultation dialogue text data and the second foreign language corresponding to each second foreign language medical consultation dialogue text data into Chinese to obtain multiple second Chinese medical consultation dialogue text data.
[0127] S940. Perform role segmentation processing on each first Chinese medical consultation dialogue text data and each second Chinese medical consultation dialogue text data to obtain doctor consultation dialogue text data and patient consultation dialogue text data.
[0128] S950. The doctor's consultation dialogue text data and the patient's consultation dialogue text data are respectively combined with each target speech data in the target speech dataset to obtain multiple doctor's consultation dialogue speech data and multiple patient's consultation dialogue speech data corresponding to the doctor's consultation dialogue text data.
[0129] S960. For each doctor's consultation dialogue voice data, the doctor's consultation dialogue voice data and the corresponding patient's consultation dialogue voice data are spliced together to obtain the spliced medical consultation dialogue voice data.
[0130] S970. In the spliced medical consultation dialogue voice data, add a preset length of silence interval segment between each doctor consultation dialogue voice data and each patient consultation dialogue voice data, and add a start timestamp and an end timestamp to each doctor consultation dialogue voice data and each patient consultation dialogue voice data to obtain multiple medical consultation dialogue voice data.
[0131] It should be noted that the specific implementation methods of steps S910 to S970 are similar to those of the relevant steps in the above embodiments of this disclosure, and will not be repeated here.
[0132] In this embodiment, a large amount of Chinese medical consultation dialogue text data can be obtained by translating medical consultation dialogue text data in languages other than Chinese into Chinese. Then, based on Chinese speech data from non-medical consultation scenarios, speech synthesis is performed on the Chinese medical consultation dialogue text data to obtain a large amount of high-quality medical consultation dialogue speech data. This improves the richness of the obtained medical consultation dialogue speech data. Furthermore, it eliminates the need to install recording equipment in each consultation room and is not subject to hospital privacy protection restrictions, thus reducing equipment costs. Simultaneously, it eliminates the need for manual review and processing, reducing labor costs and improving the efficiency of speech data acquisition. At the same time, by adding a preset length of silence interval between each consultation dialogue speech data and adding start and end timestamps to each doctor's and patient's consultation dialogue speech data, the consultation dialogue speech data and consultation dialogue text data are aligned, providing support for subsequently obtaining a training dataset for speech transcription models.
[0133] Figure 10 This is a schematic diagram of the structure of a voice data construction device provided in an embodiment of this disclosure.
[0134] In this embodiment, the voice data construction device can be located within an electronic device and is understood as a functional module within the aforementioned electronic device. Specifically, the electronic device can be a server or a terminal, wherein the terminal specifically includes mobile phones, computers, or tablet computers, etc., without limitation.
[0135] like Figure 10 As shown, the voice data construction device 1000 may include a data acquisition unit 1010, a text translation unit 1020, and a voice data construction unit 1030.
[0136] The data acquisition unit 1010 can be used to acquire a first foreign language medical consultation dialogue text dataset, a second foreign language medical consultation dialogue text dataset, a first Chinese medical consultation dialogue text dataset, and a target speech dataset. The target speech dataset is a collection of Chinese speech data collected in non-medical consultation scenarios.
[0137] The text translation unit 1020 can be used to translate each first foreign language medical consultation dialogue text data and each second foreign language medical consultation dialogue text data respectively, to obtain multiple second Chinese medical consultation dialogue text data.
[0138] The speech data construction unit 1030 can be used to synthesize speech for each first Chinese medical consultation dialogue text data and each second Chinese medical consultation dialogue text data based on each target speech data in the target speech dataset, thereby obtaining multiple medical consultation dialogue speech data.
[0139] In this embodiment of the disclosure, when constructing voice data, a first foreign language medical consultation dialogue text dataset, a second foreign language medical consultation dialogue text dataset, a first Chinese medical consultation dialogue text dataset, and a target voice dataset can be obtained. The target voice dataset is a collection of Chinese voice data collected in non-medical consultation scenarios. Then, each first foreign language medical consultation dialogue text dataset and each second foreign language medical consultation dialogue text dataset are translated to obtain multiple second Chinese medical consultation dialogue text datasets. Finally, based on each target voice dataset in the target voice dataset, speech synthesis is performed on each first Chinese medical consultation dialogue text dataset and each second Chinese medical consultation dialogue text dataset to obtain multiple medical consultation dialogue voice datasets. Therefore, it is possible to acquire multilingual medical consultation dialogue text data, and translate medical consultation dialogue text data in languages other than Chinese into Chinese format, thereby obtaining a large amount of Chinese medical consultation dialogue text data. Then, by using Chinese speech data from non-medical consultation scenarios, speech synthesis is performed on the obtained Chinese medical consultation dialogue text data, resulting in multiple medical consultation dialogue speech data. In the process of acquiring high-quality, large-scale medical consultation dialogue speech data, there is no need to install recording equipment in each consultation room, nor is it subject to hospital privacy protection restrictions, thus reducing equipment costs. At the same time, no manual review and processing is required, reducing labor costs and improving the efficiency of speech data acquisition.
[0140] In some embodiments of this disclosure, the text translation unit 1020 may be specifically used to perform language conversion processing on each first foreign language medical consultation dialogue text data, converting the first foreign language corresponding to the first foreign language medical consultation dialogue text data into a second foreign language, thereby obtaining the target medical consultation dialogue text data corresponding to the first foreign language medical consultation dialogue text data.
[0141] The target medical consultation dialogue text data and the corresponding second foreign language medical consultation dialogue text data for each second foreign language medical consultation dialogue text data are translated into Chinese to obtain multiple second Chinese medical consultation dialogue text data.
[0142] In some embodiments of this disclosure, the voice data construction apparatus 1000 may further include a model acquisition unit.
[0143] The model acquisition unit can be used to acquire a training dataset before performing language conversion processing on the medical consultation dialogue text data in the first foreign language, converting the first foreign language corresponding to the medical consultation dialogue text data in the first foreign language into the second foreign language, and obtaining the target medical consultation dialogue text data corresponding to the medical consultation dialogue text data in the first foreign language. Each training data in the training dataset includes the first foreign language text and the second foreign language text corresponding to the first foreign language text.
[0144] For each training data in the training dataset, the training data and preset prompt words are input into the text translation model to be trained. The encoder in the text translation model to be trained encodes the first foreign language text, and the resulting foreign language feature vector sequence and preset prompt words are input into the decoder of the text translation model to be trained. The decoder translates the text based on the preset prompt words and outputs the predicted text.
[0145] Loss is calculated based on the predicted text, the second foreign language text, and a preset loss function to obtain the loss value;
[0146] The parameters of the text translation model to be trained are adjusted based on the loss value until convergence, resulting in a trained text translation model. The trained text translation model is used to convert text data from a first foreign language to a second foreign language.
[0147] In some embodiments of this disclosure, the text translation unit 1020 may also be specifically used to input the first foreign language medical consultation dialogue text data and preset prompt words into the trained text translation model, and the trained text translation model performs language conversion processing on the first foreign language medical consultation dialogue text data based on the preset prompt words to obtain the target medical consultation dialogue text data.
[0148] In some embodiments of this disclosure, the voice data construction unit 1030 may be specifically used to perform deduplication processing on each first Chinese medical consultation dialogue text data and each second Chinese medical consultation dialogue text data to obtain multiple target Chinese medical consultation dialogue text data after deduplication.
[0149] For each target Chinese medical consultation dialogue text data, role segmentation is performed on the target Chinese medical consultation dialogue text data to obtain doctor consultation dialogue text data and patient consultation dialogue text data.
[0150] The doctor's consultation dialogue text data and the patient's consultation dialogue text data are respectively combined with each target speech data in the target speech dataset to obtain multiple doctor's consultation dialogue voice data corresponding to the doctor's consultation dialogue text data and multiple patient's consultation dialogue voice data corresponding to the patient's consultation dialogue text data.
[0151] Voice data was constructed based on multiple doctor consultation dialogue voice data and multiple patient consultation dialogue voice data to obtain multiple medical consultation dialogue voice data.
[0152] In some embodiments of this disclosure, the speech data construction unit 1030 may also be specifically used to input the doctor's consultation dialogue text data into the trained speech synthesis model for each target speech data. The text encoder preprocessing network in the encoding module of the trained speech synthesis model performs a first processing on the doctor's consultation dialogue text data, and inputs the text feature space vector obtained after the first processing into the shared encoder in the encoding module. The shared encoder performs feature space transformation processing on the text feature space vector to obtain a first feature space vector, and inputs the first feature space vector into the shared decoder in the decoding module of the trained speech synthesis model.
[0153] The target speech data is input into the speech decoder preprocessing module of the trained speech synthesis model's decoding module. The speech decoder preprocessing module performs a second processing on the target speech data, and the resulting speech feature space vector is input into the shared decoder in the decoding module. The shared decoder performs feature space transformation on the speech feature space vector to obtain a second feature space vector. The first and second feature space vectors are then processed a third time to obtain the target feature space vector. The target feature space vector is then input into the speech decoder postprocessing module, which performs speech synthesis processing to obtain multiple doctor consultation dialogue speech data.
[0154] In some embodiments of this disclosure, the voice data construction unit 1030 may also be specifically used to splice the doctor's consultation dialogue voice data and the patient's consultation dialogue voice data corresponding to the doctor's consultation dialogue voice data for each doctor's consultation dialogue voice data, so as to obtain spliced medical consultation dialogue voice data.
[0155] In the spliced medical consultation dialogue voice data, a preset length of silent interval segment is added between each doctor consultation dialogue voice data and each patient consultation dialogue voice data. A start timestamp and an end timestamp are added to each doctor consultation dialogue voice data and each patient consultation dialogue voice data to obtain multiple medical consultation dialogue voice data.
[0156] It should be noted that, Figure 10 The voice data construction device 1000 shown can execute the various steps in the above method embodiments and realize the various processes and effects in the above method embodiments, which will not be elaborated here.
[0157] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.
[0158] In this embodiment of the disclosure, Figure 11 The electronic devices shown can be servers or terminals, and terminals specifically include mobile phones, computers, or tablets, etc., without limitation.
[0159] like Figure 11 As shown, the electronic device may include a processor 1110 and a memory 1120 storing computer program instructions.
[0160] Specifically, the processor 1110 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this disclosure.
[0161] Memory 1120 may include a mass storage device for information or instructions. For example, and not limitingly, memory 1120 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 1120 may include removable or non-removable (or fixed) media. Where appropriate, memory 1120 may be internal or external to the integrated gateway device. In a particular embodiment, memory 1120 is a non-volatile solid-state memory. In a particular embodiment, memory 1120 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0162] The processor 1110 reads and executes computer program instructions stored in the memory 1120 to perform the steps of the voice data construction method provided in this embodiment of the disclosure.
[0163] In one example, the electronic device may also include a transceiver 1130 and a bus 1140. Wherein, as... Figure 11 As shown, the processor 1110, memory 1120 and transceiver 1130 are connected via bus 1140 and communicate with each other.
[0164] Bus 1140 may include hardware, software, or both. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 1140 may include one or more buses.
[0165] This disclosure also provides a computer-readable storage medium that can store a computer program, which, when executed by a processor, enables the processor to implement the voice data construction method provided in this disclosure.
[0166] The aforementioned storage medium may, for example, include a memory 1120 containing computer program instructions, which can be executed by a processor 1110 of an electronic device to complete the voice data construction method provided in this embodiment. Optionally, the storage medium may be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), compact disc-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device.
[0167] This disclosure also provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, they implement the voice data construction method provided in this disclosure and can achieve the various processes and effects in the above embodiments of this disclosure, which will not be elaborated here.
[0168] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A voice data construction method characterized by comprising: The method comprises the following steps: obtain a first foreign language medical consultation dialogue text data set, a second foreign language medical consultation dialogue text data set, a first Chinese medical consultation dialogue text data set, and a target voice data set, wherein the target voice data set is a set of collected Chinese voice data in a non-medical consultation scenario; the first foreign language medical consultation dialogue text data set is a medical consultation dialogue text data set in a first foreign language, and the second foreign language medical consultation dialogue text data set is a medical consultation dialogue text data set in a second foreign language; the first foreign language is a foreign language other than English, and the second foreign language is English; translate each first foreign language medical consultation dialogue text data and each second foreign language medical consultation dialogue text data to obtain a plurality of second Chinese medical consultation dialogue text data; perform voice synthesis on each first Chinese medical consultation dialogue text data and each second Chinese medical consultation dialogue text data based on each target voice data in the target voice data set to obtain a plurality of medical consultation dialogue voice data.
2. The method of claim 1, wherein, The method further comprises the following steps before the step of translating each first foreign language medical consultation dialogue text data and each second foreign language medical consultation dialogue text data to obtain a plurality of second Chinese medical consultation dialogue text data: obtain a training data set, wherein each training data in the training data set comprises a first foreign language text and a second foreign language text corresponding to the first foreign language text; for each training data in the training data set, input the training data and a preset prompt word into a text translation model to be trained, encode the first foreign language text by an encoder in the text translation model to be trained, input the obtained foreign language feature vector sequence and the preset prompt word into a decoder in the text translation model to be trained, and translate the foreign language feature vector sequence based on the preset prompt word by the decoder to output a predicted text; 3. The method of claim 2, wherein, perform loss calculation based on the predicted text, the second foreign language text, and a preset loss function to obtain a loss value; adjust parameters of the text translation model to be trained based on the loss value until convergence is achieved to obtain a trained text translation model, wherein the trained text translation model is used to convert text data from a first foreign language to a second foreign language. 4. The method of claim 3, wherein, The language conversion processing is performed on the first foreign language medical inquiry dialogue text data based on the preset prompt word by using the trained text translation model, and the target medical inquiry dialogue text data is obtained. The language conversion processing is performed on the first foreign language medical inquiry dialogue text data based on the preset prompt word by using the trained text translation model, and the target medical inquiry dialogue text data is obtained.
5. The method of claim 1, wherein, The speech synthesis is performed on each first Chinese medical inquiry dialogue text data and each second Chinese medical inquiry dialogue text data based on each target speech data in the target speech data set, and a plurality of medical inquiry dialogue speech data is obtained. The speech synthesis is performed on each first Chinese medical inquiry dialogue text data and each second Chinese medical inquiry dialogue text data based on each target speech data in the target speech data set, and a plurality of medical inquiry dialogue speech data is obtained. The speech synthesis is performed on each first Chinese medical inquiry dialogue text data and each second Chinese medical inquiry dialogue text data based on each target speech data in the target speech data set, and a plurality of medical inquiry dialogue speech data is obtained. The speech synthesis is performed on each first Chinese medical inquiry dialogue text data and each second Chinese medical inquiry dialogue text data based on each target speech data in the target speech data set, and a plurality of medical inquiry dialogue speech data is obtained. The speech synthesis is performed on each first Chinese medical inquiry dialogue text data and each second Chinese medical inquiry dialogue text data based on each target speech data in the target speech data set, and a plurality of medical inquiry dialogue speech data is obtained.
6. The method of claim 5, wherein, The speech synthesis is performed on each first Chinese medical inquiry dialogue text data and each second Chinese medical inquiry dialogue text data based on each target speech data in the target speech data set, and a plurality of medical inquiry dialogue speech data is obtained. The speech synthesis is performed on each first Chinese medical inquiry dialogue text data and each second Chinese medical inquiry dialogue text data based on each target speech data in the target speech data set, and a plurality of medical inquiry dialogue speech data is obtained. The speech synthesis is performed on each first Chinese medical inquiry dialogue text data and each second Chinese medical inquiry dialogue text data based on each target speech data in the target speech data set, and a plurality of medical inquiry dialogue speech data is obtained. The speech synthesis is performed on each first Chinese medical inquiry dialogue text data and each second Chinese medical inquiry dialogue text data based on each target speech data in the target speech data set, and a plurality of medical inquiry dialogue speech data is obtained. The target voice data is input into a speech decoder preprocessing module in a decoding module of the trained speech synthesis model, the target voice data is secondly processed by the speech decoder preprocessing module, a speech feature space vector obtained after the second processing is input into a shared decoder in the decoding module, the speech feature space vector is processed by the shared decoder to obtain a second feature space vector, the first feature space vector and the second feature space vector are thirdly processed to obtain a target feature space vector, and the target feature space vector is input into a speech decoder post-processing module to obtain the plurality of doctor consultation dialogue voice data through speech synthesis processing of the speech decoder post-processing module.
7. The method of claim 5, wherein, The speech data construction based on the plurality of doctor consultation dialogue voice data and the plurality of patient consultation dialogue voice data obtains the plurality of medical consultation dialogue voice data, including: For each doctor consultation dialogue voice data, the doctor consultation dialogue voice data and the patient consultation dialogue voice data corresponding to the doctor consultation dialogue voice data are spliced to obtain spliced medical consultation dialogue voice data; In the spliced medical consultation dialogue voice data, a preset length of silence interval segment is added between each doctor consultation dialogue voice data and each patient consultation dialogue voice data, and a start timestamp and an end timestamp are added to each doctor consultation dialogue voice data and each patient consultation dialogue voice data to obtain the plurality of medical consultation dialogue voice data.
8. A voice data constructing apparatus characterized by comprising: Including: A data acquisition unit is configured to acquire a first foreign language medical consultation dialogue text data set, a second foreign language medical consultation dialogue text data set, a first Chinese medical consultation dialogue text data set, and a target voice data set, wherein the target voice data set is a set of collected Chinese voice data in a non-medical consultation scenario; the first foreign language medical consultation dialogue text data set is a medical consultation dialogue text data set in a first foreign language, the second foreign language medical consultation dialogue text data set is a medical consultation dialogue text data set in a second foreign language, the first foreign language is a foreign language other than English, and the second foreign language is English; A text translation unit is configured to translate each first foreign language medical consultation dialogue text data and each second foreign language medical consultation dialogue text data to obtain a plurality of second Chinese medical consultation dialogue text data; A speech data construction unit is configured to perform speech synthesis on each first Chinese medical consultation dialogue text data and each second Chinese medical consultation dialogue text data based on each target voice data in the target voice data set to obtain a plurality of medical consultation dialogue voice data.
9. An electronic device, comprising: Including: A processor; A memory configured to store executable instructions; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the speech data construction method of any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by the processor, the processor implements the voice data construction method in any one of claims 1-7.
Citation Information
Patent Citations
Medical inquiry data processing method and device
CN113555133A
Speech translation model modeling method and device based on speech synthesis data
CN115828943A