Medical science popularization live broadcast dialogue transcription method and device based on large language model
By using a large language model to perform structured and semantic optimization of medical science popularization live dialogues, the problem of insufficient accuracy in transcribed text was solved, achieving efficient and low-cost transcribed text generation and improving the quality and dissemination efficiency of medical science popularization content.
Patent Information
- Application Number
- CN202511532924.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2026-01-23
AI Technical Summary
In medical science popularization live streaming scenarios, existing technologies have problems such as high word error rate, confusion of medical concepts, unclear speaker roles and poor semantic coherence in the transcribed text generated by automatic speech recognition systems. As a result, the transcription results require a lot of manpower to calibrate and are very costly, making it difficult to meet the requirements of high accuracy and low cost.
The system employs a large language model to perform structuring and semantic optimization on the raw transcribed text generated by the automatic speech recognition system. It uses preset prompt templates to enhance punctuation, separate speakers, and correct errors. It also uses a medical natural language API to verify medical terminology and outputs high-quality transcribed text.
It significantly improves the accuracy and usability of transcribed texts, reduces the cost of manual calibration, enhances the quality and dissemination efficiency of medical science popularization content, and meets the needs of low cost and ease of use.
Smart Images

Figure CN121388136A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of medical information technology and natural language processing technology, specifically to a method and device for transcribing live medical science popularization dialogues based on a large language model. Background Technology
[0002] In the field of medical science popularization live streaming, accurately and efficiently converting the dialogue between experts and the audience into written text is a crucial step in knowledge preservation and dissemination. Currently, this task mainly relies on Automatic Speech Recognition (ASR) systems. However, the diverse accents, complex medical terminology, and impromptu nature of medical dialogues lead to a series of problems in the raw transcribed text generated by ASR systems, including high word error rates, confusion of medical concepts, unclear speaker roles, and poor semantic coherence. These inherent defects in the transcribed text seriously hinder the quality and efficiency of medical science popularization content and pose potential risks to its subsequent direct application. Therefore, how to effectively improve the accuracy and usability of ASR system output in medical science popularization scenarios has become an urgent technical challenge to be solved in this field.
[0003] To address these challenges, existing technical solutions primarily follow two paths. First, traditional ASR systems (such as Baidu Speech Recognition and iFlytek Open Platform) are directly employed. However, these systems commonly suffer from errors in recognizing drug names, dosages, and anatomical terms when processing specialized medical content, and lack effective speaker separation capabilities. This results in transcription results requiring significant manpower for calibration and modification, leading to high costs and low efficiency. Second, some research attempts to modify Large Language Models (LLMs) themselves, such as by equipping them with audio encoders or performing model transfer, aiming to endow LLMs with direct speech processing capabilities. However, these solutions typically require complex structural modifications to the model and rely on large-scale specialized corpora for training. Their high implementation costs and complexity make them difficult to widely apply in broad medical science popularization platforms.
[0004] In summary, existing technologies cannot simultaneously meet the stringent requirements of high accuracy, strong semantic coherence, and precise speaker role differentiation in medical science popularization live streaming scenarios, while also ensuring low cost and ease of use. There is an urgent need in this field for an innovative method that can overcome the aforementioned shortcomings to address the long-standing pain point of insufficient transcription accuracy in medical science popularization live streaming dialogues. This has become a pressing technical problem to be solved in this field. Summary of the Invention
[0005] The purpose of this application is to provide a method and device for transcribing medical science popularization live dialogues based on a large language model, so as to solve the problem of insufficient accuracy in transcribing medical science popularization live dialogues in the prior art.
[0006] To achieve the above objectives, the technical solution adopted in this application is as follows: According to one aspect of the embodiments of this application, a method for transcribing medical science popularization live dialogue based on a large language model is provided, comprising: receiving raw transcribed text about a medical dialogue generated by an automatic speech recognition system; inputting the raw transcribed text into a large language model, prompting the large language model to perform structuring and semantic optimization processing on the raw transcribed text through a preset prompt template; and outputting the target transcribed text obtained after processing by the large language model.
[0007] Based on the aforementioned technical means, the system first receives the original transcribed text of medical dialogue generated by an automatic speech recognition system, providing basic data for subsequent processing. Then, it inputs the text into a large language model, using preset prompt templates to allow the model to perform structuring and semantic optimization of the text. This effectively streamlines the text logic, standardizes expression, and corrects semantic issues. Finally, the target transcribed text is output, resulting in a higher-quality, clearer, and more semantically accurate medical science popularization live dialogue transcription. This greatly improves the usability and professionalism of the transcription results, facilitating the organization, dissemination, and subsequent research of medical knowledge.
[0008] Furthermore, the original transcribed text is input into the large language model, and the model performs structuring and semantic optimization on the original transcribed text through a preset prompt template. This includes: sequentially performing punctuation enhancement, speaker separation, and error correction on the original transcribed text; where punctuation enhancement is used to complete or correct punctuation marks in the transcribed text; speaker separation is used to identify and mark the speakers in the transcribed text; and error correction is used to correct textual errors in the transcribed text.
[0009] Based on the above technical means, punctuation enhancement, speaker separation and error correction operations are carried out sequentially on the original transcribed text. This can comprehensively optimize the text, complete or correct punctuation to make the text expression more standardized, identify and mark speakers to make the dialogue roles clear, correct text errors to improve the accuracy of the text, and finally obtain a higher quality and more reasonable transcribed text.
[0010] Furthermore, punctuation enhancement is achieved through the following methods: inputting the original transcribed text into a large language model; prompting the large language model to complete or correct punctuation marks in the text using a preset punctuation enhancement prompt template; and obtaining the first transcribed text with punctuation enhancement.
[0011] Based on the above technical means, the original transcribed text is input into the large language model, and punctuation marks are completed or corrected by using preset punctuation enhancement prompt templates. This can effectively solve the problem of missing or incorrect punctuation marks in the original transcribed text, enhance the readability and standardization of the text, and make the transcribed content more in line with normal language expression habits.
[0012] Furthermore, speaker separation is achieved through the following methods: inputting the first transcribed text with enhanced punctuation into a large language model; prompting the large language model to identify the statements of different speakers through a pre-set role recognition prompt template containing inference examples; classifying the statements into different speaker roles based on context, tone, emotion, or wording features; assigning role labels to the identified speaker statements and marking them in the first transcribed text to obtain the second transcribed text.
[0013] Based on the aforementioned technical means, the first transcribed text with enhanced punctuation is input into a large language model. Using a role recognition prompt template containing reasoning examples, different speakers are distinguished based on multiple features and assigned role labels. This clearly presents the speech of different roles in the dialogue, facilitates the understanding of the dialogue structure and logic, and enhances the information value and practicality of the transcribed text.
[0014] Furthermore, error correction is achieved through the following methods: inputting the second transcribed text after speaker separation into the large language model; prompting the large language model to identify and correct errors in medical terminology, semantic inconsistencies, or transcriptional omissions through a preset error correction prompt template; verifying medical terminology using a medical natural language API; and outputting the corrected target transcribed text.
[0015] Based on the above technical means, the second transcribed text after speaker separation is input into a large language model. By using a preset error correction prompt template to identify and correct errors in medical terminology, semantic inconsistencies, or transcription omissions, and by using a medical natural language API to verify medical terminology, the accuracy and professionalism of medical information in the transcribed text can be ensured, and misunderstandings or adverse effects caused by erroneous information can be avoided.
[0016] Furthermore, the original transcribed text is input into the large language model, and a preset prompt template prompts the large language model to perform structuring and semantic optimization processing on the original transcribed text. This also includes: in a single inference step, speaker separation and error correction processing are performed on the original transcribed text. Specifically, this involves: inputting the original transcribed text into the large language model; using a preset zero-shot prompt template, prompting the large language model to simultaneously perform speaker identification, error correction, and medical terminology consistency checks on the original transcribed text; and outputting the target transcribed text obtained after processing by the large language model.
[0017] Based on the above technical means, in a single inference step, the original transcribed text is input into a large language model, and speaker recognition, error correction, and medical terminology consistency checks are performed simultaneously through a preset zero-sample prompt template. This can efficiently integrate multiple processing tasks, reduce processing steps and time, and quickly output structured and semantically optimized target transcribed text, thereby improving overall processing efficiency.
[0018] Furthermore, before outputting the target transcribed text obtained after processing by the large language model, the method also includes: parsing the target transcribed text using regular expressions to extract the enhanced target transcribed text.
[0019] Based on the above technical means, before outputting the target transcribed text processed by the large language model, regular expressions are used to parse the target transcribed text to extract the enhanced content. This allows for further precise processing and optimization of the transcribed text, extracting key and effective information, making the final transcribed text more refined and accurate, and meeting the needs of different application scenarios.
[0020] According to another aspect of the embodiments of this application, a medical science popularization live dialogue transcription device based on a large language model is also provided, comprising: a speech recognition module for receiving raw transcribed text about medical dialogue generated by an automatic speech recognition system; a text processing module for inputting the raw transcribed text into a large language model and prompting the large language model to perform structuring and semantic optimization processing on the raw transcribed text through a preset prompt template; and a text output module for outputting the target transcribed text obtained after processing by the large language model.
[0021] According to another aspect of the embodiments of this application, an electronic device is also provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; wherein the memory is used to store computer programs; and the processor is used to execute the steps of the medical science popularization live dialogue transcription method based on a large language model in any of the above embodiments by running the computer programs stored in the memory.
[0022] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein the storage medium stores a computer program, wherein the computer program is configured to execute the steps of the medical science popularization live dialogue transcription method based on a large language model in any of the above embodiments when running.
[0023] The beneficial effects of this application are: This application employs a large language model to intelligently post-process the raw transcribed text generated by an automatic speech recognition system, significantly improving the final accuracy and usability of medical science popularization live dialogue transcription. This method effectively overcomes the inherent deficiencies of traditional ASR systems in medical terminology recognition and speaker identification, greatly reducing word error rates and medical concept error rates, thereby minimizing the risk of medical knowledge dissemination due to transcription errors and the high cost of manual calibration. Simultaneously, this solution eliminates the need for complex structural modifications or additional data training to the large language model, fully utilizing its inherent text understanding and generation capabilities. Through flexible prompting engineering techniques, it enhances professional performance, ensuring efficient and accurate processing while also offering significant advantages such as low cost, ease of implementation, and strong adaptability, greatly promoting the efficient production and dissemination of high-quality medical science popularization content. Attached Figure Description
[0024] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0025] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a schematic diagram of the hardware environment for an optional medical science popularization live dialogue transcription method based on a large language model, provided in an embodiment of this application. Figure 2 This is a flowchart illustrating an optional medical science popularization live dialogue transcription method based on a large language model, provided in an embodiment of this application. Figure 3 This is a structural block diagram of an optional medical science popularization live dialogue transcription device based on a large language model, provided in an embodiment of this application. Figure 4 This is a structural block diagram of an optional electronic device provided in an embodiment of this application. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] According to one aspect of the embodiments of this application, a method for transcribing live medical science popularization dialogues based on a large language model is provided. Optionally, in this embodiment, the above-mentioned method for transcribing live medical science popularization dialogues based on a large language model can be applied to a hardware environment consisting of a terminal and a server. The server is connected to the terminal via a network and can be used to provide services to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services to the server.
[0030] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal may not be limited to PC, mobile phone, tablet computer, etc.
[0031] The medical science popularization live dialogue transcription method based on a large language model, as described in this application, can be executed by a server, a terminal, or both. Specifically, the execution of the medical science popularization live dialogue transcription method based on a large language model, as described in this application, can also be performed by a client installed on the terminal.
[0032] Taking the medical science popularization live dialogue transcription method based on a large language model, which is jointly executed by a terminal and a server in this embodiment, as an example, please refer to [link to relevant documentation]. Figure 1 , Figure 1 This is a schematic diagram of the hardware environment for an optional medical science popularization live dialogue transcription method based on a large language model, as provided in an embodiment of this application. Figure 1As shown, the hardware environment of this medical science popularization live dialogue transcription method based on a large language model includes: a terminal 102 and a server 104 connected to the terminal 102 via a network. The server 104 is used to deploy a large language model, which executes the medical science popularization live dialogue transcription method based on a large language model according to this embodiment, performing structuring and semantic optimization processing on the transcribed text. The terminal 102 is used to deploy an automatic speech recognition system, generate the original transcribed text about the medical dialogue, and display the transcription processing results. These transcription processing results can be obtained by structuring and semantic optimization processing using the large language model deployed on the server 104.
[0033] The medical science popularization live broadcast dialogue transcription method based on a large language model in this embodiment can be applied to scenarios such as medical knowledge dissemination, online medical consultation record organization, medical education material compilation, and medical industry conference content archiving. For example, in health-related online live broadcasts, doctors and viewers engage in interactive Q&A sessions on the prevention and treatment of common diseases. This method can quickly convert the audio dialogue into accurate and standardized text, facilitating subsequent organization into popular science articles. On online medical consultation platforms, the audio of patient-doctor conversations transcribed using this method can form clear consultation records for doctors to review and for patients to retain. In medical education, the audio content of expert lectures can be transcribed into text and used as teaching materials. When the medical industry holds academic conferences, the discussion audio during the conference can be transcribed and archived using this method for easy reference and research later. This embodiment uses the transcription of an online live broadcast popular science dialogue on cardiovascular disease prevention as an example to illustrate the above-mentioned medical science popularization live broadcast dialogue transcription method based on a large language model.
[0034] In the field of medical science popularization live streaming, accurately and efficiently converting the dialogue between experts and the audience into written text is a crucial step in knowledge preservation and dissemination. Currently, this task mainly relies on Automatic Speech Recognition (ASR) systems. However, the diverse accents, complex medical terminology, and impromptu nature of medical dialogues lead to a series of problems in the raw transcribed text generated by ASR systems, including high word error rates, confusion of medical concepts, unclear speaker roles, and poor semantic coherence. These inherent defects in the transcribed text seriously hinder the quality and efficiency of medical science popularization content and pose potential risks to its subsequent direct application. Therefore, how to effectively improve the accuracy and usability of ASR system output in medical science popularization scenarios has become an urgent technical challenge to be solved in this field.
[0035] To address these challenges, existing technical solutions primarily follow two paths. First, traditional ASR systems (such as Baidu Speech Recognition and iFlytek Open Platform) are directly employed. However, these systems commonly suffer from errors in recognizing drug names, dosages, and anatomical terms when processing specialized medical content, and lack effective speaker separation capabilities. This results in transcription results requiring significant manpower for calibration and modification, leading to high costs and low efficiency. Second, some research attempts to modify Large Language Models (LLMs) themselves, such as by equipping them with audio encoders or performing model transfer, aiming to endow LLMs with direct speech processing capabilities. However, these solutions typically require complex structural modifications to the model and rely on large-scale specialized corpora for training. Their high implementation costs and complexity make them difficult to widely apply in broad medical science popularization platforms.
[0036] In summary, existing technologies cannot simultaneously meet the stringent requirements of high accuracy, strong semantic coherence, and precise speaker role differentiation in medical science popularization live streaming scenarios, while also ensuring low cost and ease of use. There is an urgent need in this field for an innovative method that can overcome the aforementioned shortcomings to address the long-standing pain point of insufficient transcription accuracy in medical science popularization live streaming dialogues. This has become a pressing technical problem to be solved in this field.
[0037] To address the aforementioned issues, this embodiment provides a medical science popularization live dialogue transcription method based on a large language model, running on the aforementioned server. Please refer to [link to relevant documentation]. Figure 2 , Figure 2 This is a flowchart illustrating an optional medical science popularization live dialogue transcription method based on a large language model, as provided in an embodiment of this application. Figure 2 As shown in the figure, the medical science popularization live dialogue transcription method based on a large language model in this application specifically includes the following steps: Step S201: Receive the original transcribed text of the medical conversation generated by the automatic speech recognition system; Step S202: Input the original transcribed text into the large language model, and prompt the large language model to perform structuring and semantic optimization processing on the original transcribed text through the preset prompt template; Step S203: Output the target transcribed text obtained after processing by the large language model.
[0038] Through steps S201 to S203, the system first receives the original transcribed text of the medical dialogue generated by the automatic speech recognition system, providing basic data for subsequent processing. Then, it inputs the text into a large language model, which uses preset prompt templates to perform structuring and semantic optimization on the text. This effectively organizes the text logic, standardizes the expression, and corrects semantic problems. Finally, the target transcribed text is output, resulting in a higher-quality, clearer, and more semantically accurate medical science popularization live dialogue transcription. This greatly improves the usability and professionalism of the transcription results, facilitating the organization, dissemination, and subsequent research of medical knowledge.
[0039] The following is combined with Figure 2 The medical science popularization live dialogue transcription method based on a large language model in the embodiments of this application will be explained.
[0040] In the technical solution of step S201, the original transcribed text of the medical dialogue generated by the automatic speech recognition system is received.
[0041] Automatic Speech Recognition (ASR) is a technology that converts the lexical content of human speech into computer-readable text input. In medical science popularization live streaming scenarios, this system recognizes and converts medical dialogues during live streams, either in real time or offline, generating raw transcribed text. However, due to the inherent limitations of speech recognition technology, such as background noise interference, unclear pronunciation, and difficulties in recognizing technical terms, the generated raw transcribed text may contain problems such as missing punctuation, speaker confusion, text errors, and semantic inaccuracies.
[0042] As an optional implementation, the original transcribed text can serve as an initial data source for subsequent processing, providing a foundation for further improvements in text quality and usability. For example, in subsequent processing based on a large language model, targeted optimizations can be performed to address any existing problems.
[0043] In the technical solution of step S202, the original transcribed text is input into the large language model, and the large language model is prompted to perform structuring and semantic optimization processing on the original transcribed text through a preset prompt template.
[0044] In this embodiment, the Large Language Model (LLM) is a massive neural network model based on deep learning algorithms. Through pre-training on massive amounts of text data, it learns rich linguistic knowledge, grammatical rules, and semantic information, possessing powerful language understanding and generation capabilities. The preset prompt templates are text instructions designed to guide the LLM to process the input text according to specific tasks and requirements. In this way, the capabilities of the LLM can be fully utilized to deeply optimize the original transcribed text.
[0045] Structured and semantic optimization processing refers to performing at least one of the following processing on the original transcribed text, including but not limited to punctuation correction, speaker role marking, and medical terminology error correction, in order to improve the readability, accuracy, and professionalism of the text.
[0046] As an optional implementation, diverse prompt templates can be designed according to different live streaming scenarios and needs to achieve more accurate text processing results. For example, for medical and disease science popularization dialogues, targeted prompt templates can be designed to enable the large language model to better understand the dialogue content and optimize it.
[0047] As an efficient and refined processing strategy, the Chain-of-Thought (CoT) approach simulates the logical flow of human problem-solving, breaking down complex text optimization tasks into multiple ordered and interconnected sub-tasks. In this embodiment, the Chain-of-Thought approach means that by constructing a logically rigorous and clearly defined processing chain, the large language model can gradually delve into the text according to a predetermined thought process, achieving comprehensive optimization from the surface to the depths and from the local to the global.
[0048] Specifically, when processing the original transcribed text, the thought chain approach first guides the large language model to perform punctuation enhancement, that is, to complete or correct missing or incorrect punctuation marks in the transcribed text, laying the foundation for subsequent processing; then, in the speaker separation stage, the model identifies and marks different speakers in the transcribed text based on context and language features to ensure the clarity of the dialogue structure; finally, in the error correction stage, the model precisely corrects textual errors in the transcribed text, including but not limited to medical terminology, grammatical errors, and semantic inconsistencies.
[0049] This chain-of-thought approach not only improves the accuracy and efficiency of text processing, but also enhances the logic and readability of the processing results, making it particularly suitable for medical science popularization live dialogue transcription scenarios with high accuracy requirements.
[0050] As an optional embodiment, the original transcribed text is input into a large language model, and the large language model is prompted by a preset prompt template to perform structuring and semantic optimization processing on the original transcribed text, including: sequentially performing punctuation enhancement, speaker separation, and error correction on the original transcribed text; wherein, punctuation enhancement is used to complete or correct punctuation marks in the transcribed text; speaker separation is used to identify and mark the speaker in the transcribed text; and error correction is used to correct textual errors in the transcribed text.
[0051] Punctuation marks play a crucial role in text, defining sentence structure, expressing tone, and indicating semantic pauses. Correct punctuation usage enhances clarity and accuracy. Speaker identification identifies and labels speakers in transcribed text. In medical science popularization live streams, which typically involve different roles such as doctors, patients, and hosts, accurate speaker identification helps in better understanding the logic and content of the dialogue. Error correction corrects textual errors in transcribed text, including spelling mistakes, grammatical errors, and textual deviations caused by speech recognition errors.
[0052] By using the above-mentioned technical means, punctuation enhancement, speaker separation, and error correction operations are performed on the original transcribed text in sequence. This can comprehensively optimize the text, complete or correct punctuation to make the text expression more standardized, identify and mark speakers to make the dialogue roles clear, and correct text errors to improve the accuracy of the text, ultimately resulting in a higher quality and more reasonable transcribed text.
[0053] As an optional embodiment, punctuation enhancement is achieved by: inputting the original transcribed text into a large language model; prompting the large language model to complete or correct punctuation marks in the text using a preset punctuation enhancement prompt template; and obtaining the first transcribed text with punctuation enhancement.
[0054] Pre-defined punctuation enhancement prompts can include examples and guidance on punctuation usage rules to help large language models better understand task requirements. For example, the prompts can explain the types of punctuation marks to use in different contexts and how to adjust punctuation based on the semantics and tone of the sentence.
[0055] In this embodiment, the content of the punctuation enhancement prompt template can be as shown in the following example: You are a helpful speech-to-text assistant. Your task is to correct punctuation in medical conversation transcripts, ensuring accurate reflection of natural pauses and speaker transitions. Here's a step-by-step guide to completing this task: 1. Contextual Interpretation: Analyze the natural pauses, transitions, and speaker shifts in each sentence.
[0056] 2. Sentence splitting: When the speaker changes, the continuous sentences are split into individual sentences.
[0057] 3. Question identification: Mark questions appropriately with question marks, and pay attention to tone and structure.
[0058] 4. Correction of continuous sentences: Break down continuous sentences into independent clauses with correct punctuation.
[0059] 5. Speaker transition: Use appropriate punctuation to separate clauses spoken by different speakers.
[0060] By using the above-mentioned technical means, the original transcribed text is input into the large language model, and punctuation marks are completed or corrected with the help of preset punctuation enhancement prompt templates. This can effectively solve the problem of missing or incorrect punctuation marks in the original transcribed text, enhance the readability and standardization of the text, and make the transcribed content more in line with normal language expression habits.
[0061] As an optional embodiment, speaker separation is achieved by: inputting the first transcribed text with enhanced punctuation into a large language model; prompting the large language model to identify the statements of different speakers through a preset role recognition prompt template containing inference examples; distinguishing different speaker roles based on context, tone, emotion, or wording features; assigning role labels to the identified speaker statements and marking them in the first transcribed text to obtain the second transcribed text.
[0062] Role recognition prompt templates containing reasoning examples provide known speaker roles and corresponding sentence examples, helping large language models learn how to identify roles based on text features. These templates include detailed reasoning guidelines such as context reading, sentence segmentation, reasoning processes, look-around strategies, labeling and justification, consistency attribution, and fine-grained attribution. Contextual information helps the model understand the overall logic and flow of the dialogue, thus more accurately judging speaker transitions; intonation features infer the speaker's emotions and attitudes by analyzing the use of interjections, exclamation marks, etc.; affective features involve analyzing the emotional tendencies expressed in the text; and phrasing features include specific vocabulary and expressions that different speakers might use.
[0063] In this embodiment, the content of the role recognition prompt template can be as shown in the following example: You are a helpful speech-to-text transcription assistant. Your current task is to tag a dialogue without speaker identifiers. You will use your deep understanding of medical terminology, dialogue structure, and context to accurately tag the text. Here's a step-by-step guide to completing this task: 1. Contextual reading: Read each sentence carefully, absorbing its content, tone, emotion, and vocabulary.
[0064] 2. Sentence Segmentation: When the speaker changes, actively break the sentence down into independent statements. Look for clues such as pauses, changes in speaking direction, thought conclusions, questions, and answers.
[0065] 3. Reasoning: Consider whether the language is a professional answer (hinting at a medical professional) or a medically related question (hinting at a host).
[0066] 4. Look-around strategy: Analyze the five sentences before and after the current sentence to understand the dialogue flow. You can then answer the question.
[0067] 5. Label and explain the reason: Label each sentence with "doctor" or "host" and provide a brief reason based on your analysis.
[0068] Make sure each reason applies to only one person.
[0069] 6. Attribution Consistency: Maintain a holistic approach throughout the recording process, giving equal attention to each sentence and handling each sentence meticulously.
[0070] 7. Extremely detailed attribution: Break the dialogue down into its smallest parts (questions, answers, utterances) for clarity and comprehension. Each clause should be accurately attributed to either the doctor or the host, and the speakers' identities should not overlap.
[0071] By using the above-mentioned technical means, the first transcribed text with enhanced punctuation is input into a large language model. Using a role recognition prompt template containing reasoning examples, different speakers are distinguished based on multiple features and role labels are assigned. This can clearly present the speech of different roles in the dialogue, making it easier to understand the dialogue structure and logic, and improving the information value and practicality of the transcribed text.
[0072] As an optional embodiment, error correction is achieved by: inputting the second transcribed text after speaker separation into a large language model; prompting the large language model to identify and correct errors in medical terminology, semantic inconsistencies, or transcriptional omissions using a preset error correction prompt template; validating the medical terminology using a medical natural language API; and outputting the corrected target transcribed text.
[0073] The preset error correction prompt templates can specify the types of errors to focus on and the direction of correction. For example, prompting the model to pay attention to the accurate use of medical professional terms such as drug names, dosages, and anatomical names, and to check the semantic coherence between sentences.
[0074] In the preset error correction prompt template, we can explicitly specify the types of errors that the model needs to pay special attention to and the corresponding correction directions. For example, the prompt model should ensure the accurate use of medical terminology and carefully check the semantic coherence between sentences to avoid meaning distortion caused by speech recognition errors. Below are some specific examples to guide the model to make more precise corrections: Correction tips for medical terminology recognition errors: "When processing the following transcribed text, please pay special attention to the accuracy of medical terminology. For example, if the original audio states 'the patient has a mild cough,' and it is incorrectly recognized as 'the patient has mild cough medicine,' please correct it to the correct expression." Correction tips for drug name confusion: "When encountering transcriptions related to drug names, be sure to carefully verify them. For example, if the original audio is 'The patient is using statins,' but it is incorrectly identified as 'The patient is using statin drugs,' please correct it to the accurate drug name based on the audio content." Correction tips for recognition errors caused by dialects or accents: "Considering that the speaker may have a local dialect or accent, please pay special attention to handling recognition errors caused by this. For example, if the original audio (with a southern accent) says 'This patient needs to have an ultrasound examination,' but it is incorrectly recognized as 'This patient needs to have an ultrasound peeling treatment,' please make the correct correction based on the actual content of the audio." Correction tips for abbreviations or omissions: "When processing transcribed text, it is also necessary to pay attention to whether there is excessive abbreviation or omission of information. For example, if the original audio is 'Liver function needs to be checked', but it is incorrectly identified as 'Stem cell function needs to be checked', please ensure that the missing or incorrectly abbreviated parts are supplemented and corrected to a complete and accurate expression." The correction tips for addressing terminological confusion in complex medical contexts are as follows: "When processing transcriptions involving complex medical contexts, be especially careful to avoid confusion between terms. For example, if the original audio is 'The patient has renal insufficiency, contrast agents are not recommended,' and it is incorrectly identified as 'The patient has renal insufficiency, sedatives are not recommended,' please correct it to the correct medical advice based on your medical knowledge and the audio content." The Medical Natural Language API (Application Programming Interface) is an interface specifically designed for processing natural language data in the medical field. It can provide functions such as accurate interpretation of medical terms, synonym replacement, and term classification. By verifying medical terms, it ensures that the corrected terms comply with medical standards.
[0075] In this embodiment, the content of the error correction prompt template can be as shown in the following example: You are a helpful speech-to-text transcription assistant. Your current task is to identify and correct errors in medical terminology, semantic inconsistencies, or transcription omissions in the input dialogue. Here's a step-by-step guide to accomplishing this task: 1. Contextual Interpretation: Analyze each sentence to find potential transcription errors, and consider medical terminology and context.
[0076] 2. Medical terminology verification: Pay special attention to medical terms, drug names, and procedures that may be misunderstood.
[0077] 3. Accent considerations: Different accents and potential misunderstandings should be taken into account during transcription.
[0078] 4. Homophone Analysis: Identify and correct potentially confusing words with similar pronunciations.
[0079] 5. Contextual coherence: Verify whether the corrected terminology is consistent with the medical background of the conversation.
[0080] By using the above-mentioned technical means, the second transcribed text after speaker separation is input into a large language model. By identifying and correcting errors in medical terminology, semantic inconsistencies, or transcription omissions through a preset error correction prompt template, and verifying medical terminology using a medical natural language API, the accuracy and professionalism of medical information in the transcribed text can be ensured, and misunderstandings or adverse effects caused by erroneous information can be avoided.
[0081] It should be noted that in order to reduce response variability, make the model output closer to the greedy decoding strategy, and ensure the accuracy and consistency of the correction, it is recommended to set the temperature parameter of LLM between 0.1 and 0.2.
[0082] In addition to the above-mentioned mind chain method for processing transcribed text, this embodiment also provides another method for processing transcribed text: the zero-shot method.
[0083] The core of the zero-shot method is that it does not rely on complex modifications to large language models or pre-training with large amounts of domain-specific data. Instead, it guides large language models to complete specific text processing tasks directly through well-designed prompt templates without any direct training examples.
[0084] Specifically, in this embodiment, the zero-shot method guides the large language model to complete both speaker identification and error correction tasks in one go through a comprehensive prompt template. This prompt template details the task requirements, contextual considerations, and expected output format, enabling the large language model to directly optimize the original transcribed text based on its language knowledge and reasoning abilities acquired during pre-training.
[0085] As an optional embodiment, the original transcribed text is input into a large language model, and a preset prompt template prompts the large language model to perform structuring and semantic optimization processing on the original transcribed text. The method also includes: in a single inference step, speaker separation and error correction processing are performed on the original transcribed text. Specific steps include: inputting the original transcribed text into the large language model; prompting the large language model to simultaneously perform speaker identification, error correction, and medical terminology consistency checks on the original transcribed text using a preset zero-shot prompt template; and outputting the target transcribed text obtained after processing by the large language model.
[0086] In this embodiment, the zero-shot prompt template is a special type of prompting that does not rely on a large amount of labeled data for training. Instead, it guides a large language model to directly perform a specific task through concise and clear instructions. In this case, the prompt template needs to clearly explain the multiple tasks to be completed simultaneously and their corresponding requirements.
[0087] In this embodiment, the content of the zero-sample prompt template can be as shown in the following example: You are a helpful speech-to-text transcription assistant. Your task is to review and correct transcription errors, focusing on accuracy and context. Consider the diverse accents of different speakers. Identify the role of each speaker based on intonation, emotion, and phrasing, whether it's a presenter or a doctor. Tag them accordingly in the transcription. Ensure the enhanced text reflects the original spoken content without adding new material. Your goal is to create accurate and context-consistent transcriptions, which will improve semantic clarity and reduce word error rates.
[0088] By employing the aforementioned technical means, the original transcribed text is input into a large language model in a single inference step. Speaker recognition, error correction, and medical terminology consistency checks are performed simultaneously using a preset zero-sample prompt template. This approach efficiently integrates multiple processing tasks, reduces processing steps and time, and quickly outputs structured and semantically optimized target transcribed text, thereby improving overall processing efficiency.
[0089] As an optional embodiment, before outputting the target transcribed text obtained after processing by the large language model, the method further includes: parsing the target transcribed text using regular expressions to extract the enhanced target transcribed text.
[0090] Regular expressions are powerful tools for matching, finding, and replacing specific patterns in text. They can accurately identify specific content within text by defining a set of rules. In the transcription and text processing of medical science popularization live broadcast dialogues, corresponding regular expressions can be designed to extract specific medical knowledge points, important suggestions, and other key information as needed. Specific regular expressions will not be explained here.
[0091] By using the above-mentioned technical means, before outputting the target transcribed text processed by the large language model, regular expressions are used to parse the target transcribed text to extract the enhanced content. This allows for further precise processing and optimization of the transcribed text, extracting key and effective information, making the final transcribed text more refined and accurate, and meeting the needs of different application scenarios.
[0092] It should be noted that in medical science popularization live-stream dialogue transcription, the zero-shot method can quickly adapt to different dialogue topics and styles, achieving instant text optimization without the need for time-consuming model fine-tuning or data collection. However, it is worth noting that the performance of the zero-shot method may be slightly inferior to that of a fully trained dedicated model, especially when dealing with extremely complex or specialized medical terminology. Therefore, in practical applications, the chain-of-thought method or the zero-shot method can be flexibly selected, or a combination of the advantages of both, depending on specific needs and resources, to achieve the best text processing results.
[0093] In the technical solution of step S203, the target transcribed text obtained after processing by the large language model is output.
[0094] In this embodiment, the target transcribed text output by LLM should be in plain text format, containing the corrected transcribed text, and each speaker's statement should be preceded by a clear role tag. This target transcribed text, after the preceding structuring and semantic optimization processing, possesses high quality and usability. It can accurately and clearly present the content of medical science popularization live-stream dialogues, providing a reliable textual foundation for the dissemination, research, and application of medical knowledge.
[0095] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0096] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM (Read-Only Memory) / RAM (Random Access Memory), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0097] According to another aspect of the embodiments of this application, an apparatus is also provided for implementing the above-described method for transcribing medical science popularization live dialogues based on a large language model. Please refer to... Figure 3 , Figure 3 This is a structural block diagram of an optional medical science popularization live dialogue transcription device based on a large language model, as provided in an embodiment of this application. Figure 3 As shown, the medical science popularization live dialogue transcription device 300 based on a large language model may include: The speech recognition module 301 is used to receive the raw transcribed text of a medical conversation generated by an automatic speech recognition system; The text processing module 302 is used to input the original transcribed text into the large language model and prompt the large language model to perform structuring and semantic optimization processing on the original transcribed text through a preset prompt template; The text output module 303 is used to output the target transcribed text obtained after processing by the large language model.
[0098] It should be noted that the speech recognition module 301 in this embodiment can be used to perform the above step S201, the text processing module 302 in this embodiment can be used to perform the above step S202, and the text output module 303 in this embodiment can be used to perform the above step S203.
[0099] Regarding the medical science popularization live dialogue transcription device based on a large language model in this embodiment, the specific manner in which its speech recognition module 301, text processing module 302, and text output module 303 execute the above-mentioned medical science popularization live dialogue transcription method based on a large language model has been described in detail in the embodiments related to this method, and will not be elaborated here.
[0100] It is understood that the technical solution provided in this embodiment, in the medical science popularization live dialogue transcription device based on a large language model, first receives the original transcribed text of the medical dialogue generated by the automatic speech recognition system, providing basic data for subsequent processing; then it is input into the large language model, and with the help of preset prompt templates, the model performs structuring and semantic optimization processing on the text, which can effectively sort out the text logic, standardize the expression and correct semantic problems; finally, the target transcribed text is output, which can obtain a medical science popularization live dialogue transcription content with higher quality, clearer organization and more accurate semantics, greatly improving the usability and professionalism of the transcription results, and facilitating the organization, dissemination and subsequent research of medical knowledge.
[0101] In addition to the modules described above, the apparatus in this embodiment may also include modules that execute any method as described in any of the embodiments of the medical science popularization live dialogue transcription method based on a large language model.
[0102] It should be noted that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules, as part of the device, can operate in ways such as... Figure 1 The method shown can be implemented in either software or hardware within a hardware environment, where the hardware environment includes a network environment.
[0103] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described medical science popularization live dialogue transcription method based on a large language model is also provided. The electronic device may be a server, a terminal, or a combination thereof.
[0104] According to another embodiment of this application, an electronic device is also provided; please refer to [link to relevant documentation]. Figure 4 , Figure 4 This is a structural block diagram of an optional electronic device provided in an embodiment of this application, such as... Figure 4 As shown, the electronic device may include: a processor 1501, a communication interface 1502, a memory 1503, and a communication bus 1504, wherein the processor 1501, the communication interface 1502, and the memory 1503 communicate with each other through the communication bus 1504.
[0105] Memory 1503 is used to store computer programs; When processor 1501 executes the program stored in memory 1503, it performs the following steps: Step S201: Receive the original transcribed text of the medical conversation generated by the automatic speech recognition system; Step S202: Input the original transcribed text into the large language model, and prompt the large language model to perform structuring and semantic optimization processing on the original transcribed text through the preset prompt template; Step S203: Output the target transcribed text obtained after processing by the large language model.
[0106] Understandably, the technical solution provided in this embodiment involves the electronic device's processor first receiving the original transcribed text of the medical dialogue generated by the automatic speech recognition system, providing basic data for subsequent processing. Then, it is input into a large language model, where a preset prompt template is used to allow the model to perform structuring and semantic optimization of the text. This effectively streamlines the text's logic, standardizes expression, and corrects semantic issues. Finally, the target transcribed text is output, resulting in a higher-quality, more organized, and semantically accurate medical science popularization live-stream dialogue transcription. This significantly improves the usability and professionalism of the transcription results, facilitating the organization, dissemination, and subsequent research of medical knowledge.
[0107] Optionally, in this embodiment, the communication bus can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used to represent it in the figure, but this does not mean that there is only one bus or one type of bus. The communication interface is used for communication between the aforementioned electronic device and other devices.
[0108] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0109] The processor mentioned above can be a general-purpose processor, including but not limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; it can also be DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0110] This application also provides a computer-readable storage medium, which includes a stored program, wherein the program executes the method steps of the above method embodiments when it runs.
[0111] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, ROMs, RAMs, portable hard drives, magnetic disks, or optical disks.
[0112] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0113] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods of the various embodiments of this application.
[0114] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0115] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or the indirect coupling or communication connection of units or modules may be electrical or other forms.
[0116] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the solution provided in this embodiment, depending on actual needs.
[0117] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0118] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for transcribing medical science popularization live dialogues based on a large language model, characterized in that, include: Receive the raw transcribed text of a medical conversation generated by an automatic speech recognition system; The original transcribed text is input into a large language model, and the model is prompted by a preset prompt template to perform structuring and semantic optimization processing on the original transcribed text. Output the target transcribed text obtained after processing by the large language model.
2. The medical science popularization live dialogue transcription method based on a large language model according to claim 1, characterized in that, The step of inputting the original transcribed text into a large language model and prompting the large language model to perform structuring and semantic optimization processing on the original transcribed text using a preset prompt template includes: The original transcribed text was sequentially subjected to punctuation enhancement, speaker separation, and error correction; The punctuation enhancement is used to complete or correct punctuation marks in the transcribed text; the speaker separation is used to identify and mark the speaker in the transcribed text; and the error correction is used to correct textual errors in the transcribed text.
3. The medical science popularization live dialogue transcription method based on a large language model according to claim 2, characterized in that, The punctuation enhancement is achieved through the following methods: The original transcribed text is input into a large language model; The pre-set punctuation enhancement prompt template prompts the large language model to complete or correct punctuation marks in the text; The first transcribed text with punctuation enhancement is obtained.
4. The medical science popularization live dialogue transcription method based on a large language model according to claim 3, characterized in that, The speaker separation is achieved through the following methods: The first transcribed text, enhanced with punctuation, is input into the large language model; The large language model identifies the statements of different speakers by using a pre-set role recognition prompt template that includes reasoning examples; Based on context, tone, emotion, or wording, different speaker roles are distinguished. The identified speaker statements are assigned role labels and marked in the first transcribed text to obtain the second transcribed text.
5. The medical science popularization live dialogue transcription method based on a large language model according to claim 4, characterized in that, The error correction is achieved through the following methods: The second transcribed text, after speaker separation, is input into the large language model; The system uses preset error correction prompts to help the large language model identify and correct errors in medical terminology, semantic inconsistencies, or transcriptional omissions. Use a medical natural language API to validate medical terms; Output the corrected target transcribed text.
6. The medical science popularization live dialogue transcription method based on a large language model according to claim 1, characterized in that, The step of inputting the original transcribed text into a large language model and prompting the large language model to perform structuring and semantic optimization processing on the original transcribed text using a preset prompt template further includes: in a single inference step, performing speaker separation and error correction processing on the original transcribed text, specifically including the following steps: The original transcribed text is input into a large language model; Using a preset zero-sample prompt template, the large language model simultaneously performs speaker recognition, error correction, and medical terminology consistency checks on the original transcribed text. Output the target transcribed text obtained after processing by the large language model.
7. The medical science popularization live dialogue transcription method based on a large language model according to claim 6, characterized in that, Before outputting the target transcribed text obtained after processing by the large language model, the method further includes: The target transcribed text is parsed using regular expressions to extract the enhanced target transcribed text.
8. A medical science popularization live dialogue transcription device based on a large language model, characterized in that, include: A speech recognition module is used to receive the raw transcribed text of a medical conversation generated by an automatic speech recognition system; The text processing module is used to input the original transcribed text into the large language model and prompt the large language model to perform structuring and semantic optimization processing on the original transcribed text through a preset prompt template; The text output module is used to output the target transcribed text obtained after processing by the large language model.
9. An electronic device comprising a processor, a communication interface, a memory, and a communication bus, wherein, The processor, the communication interface, and the memory communicate with each other via the communication bus, characterized in that... The memory is used to store computer programs; The processor is configured to execute the medical science popularization live dialogue transcription method based on a large language model as described in any one of claims 1 to 7 by running the computer program stored in the memory.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the medical science popularization live dialogue transcription method based on a large language model as described in any one of claims 1 to 7 when it is run.