Report generation method, electronic equipment, device, storage medium and program product
By obtaining the consultation text content of doctors and patients in the report generation method and using problem correction rules to generate consultation reports, the problems of low efficiency and poor accuracy when doctors manually record the recovery of the disease are solved, and efficient and accurate generation of consultation reports is achieved.
Patent Information
- Application Number
- CN202510266244.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, doctors need to manually record the patient's recovery during follow-up, resulting in long recording time, low efficiency and easy recording errors, affecting the accuracy of the report.
Provide a report generation method, by obtaining the consultation text content between the doctor and the patient, correcting the initial question through the question correction rules based on the predetermined consultation questions, determining the reply content of the consultation questions, and generating a consultation report.
It improves the efficiency and accuracy of getting consultation reports, reduces the time for doctors to record manually, and ensures that even if the call time is limited, it can communicate fully and avoid missing important issues.
Smart Images

Figure CN120199402A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and in particular, to a report generation method, an electronic device, a device, a storage medium, and a program product. Background Art
[0002] In order to improve the service level for patients, medical institutions can track and observe the recovery status of patients after medical treatment through a follow-up mechanism, so as to understand the patient's condition, medication situation, recovery status, etc. in real time. And it is possible to count the recovery status of a large number of patients, conduct disease analysis to accumulate treatment experience, and patients can consult doctors about their conditions and health problems in a timely manner. For example, doctors and patients can establish a network call. Based on the patient's condition and historical treatment situation, the doctor determines the health problems that need to be asked of the patient to inquire about the patient's condition recovery. Thus, the patient can answer the condition recovery of himself or herself according to the questions asked by the doctor, and the doctor can record the patient's condition recovery and fill out a consultation report.
[0003] However, there are various problems with this follow-up method. For example, since doctors need to manually record the patient's condition recovery, the recording process takes a lot of time and there may be recording errors. Therefore, when current doctors conduct follow-up on patients to understand the patient's condition recovery, the efficiency of obtaining the consultation report is low and the accuracy is poor. Summary of the Invention
[0004] On the one hand, a report generation method, an electronic device, a device, a storage medium, and a program product are provided to improve the efficiency and accuracy of obtaining a consultation report when a doctor conducts follow-up on a patient to understand the patient's condition recovery.
[0005] The report generation method includes: obtaining the consultation text content between a doctor and a patient; determining the reply content of the consultation questions from the consultation text content based on pre-determined consultation questions; the consultation questions are obtained by correcting the initial questions based on a question correction rule, and the initial questions are questions related to the abnormal physical condition of the patient; generating a consultation report based on the consultation questions and the reply content of the consultation questions.
[0006] In view of this, the embodiments of the present application provide a report generation method, which can determine the reply content of the inquiry questions from the inquiry text content between doctors and patients based on the obtained inquiry text content and in combination with the pre-determined inquiry questions. Thus, an inquiry report is generated based on the inquiry questions and the reply content of the inquiry questions. Importantly, the pre-determined inquiry questions are obtained by modifying the initial questions based on the question modification rules, and the initial questions are questions related to the abnormal physical conditions of the patients. Therefore, when a doctor conducts a follow-up visit after diagnosing a patient's condition, initial questions can be set in advance for the patient's abnormal physical conditions, and the initial questions can be modified to obtain inquiry questions targeted at the patient's condition. Thus, during the follow-up visit, based on the inquiry text content corresponding to the follow-up conversation between the doctor and the patient, the reply content of the patient to the pre-determined inquiry questions can be determined, thereby generating an inquiry report for the doctor. In this way, there is no need for the doctor to manually record the patient's condition recovery, and there is no need to spend a lot of recording time. Thus, even if the call duration between the doctor and the patient is limited, sufficient communication can be carried out between the doctor and the patient, and important questions related to the patient's condition can be avoided from being missed. Based on this method, when a doctor conducts a follow-up visit to a patient to understand the patient's condition recovery, the efficiency and accuracy of obtaining the inquiry report can be improved.
[0007] In some embodiments, the question modification rules include at least one of the following: expanding the proper nouns in the initial question; annotating the core words in the initial question; expanding the initial question and determining the triggering conditions for the extended extended questions.
[0008] Based on the above technical solution, when the present application modifies the initial questions related to the abnormal physical conditions of the patients, it can expand the proper nouns in the initial questions, annotate the core words in the initial questions, expand the initial questions, and determine the triggering conditions for the extended extended questions. In this way, targeted inquiry questions related to the patient's condition can be obtained based on the initial questions related to the patient's abnormal physical conditions. Thus, the inquiry questions corresponding to the patient can be accurately determined to accurately understand the patient's condition recovery when the doctor conducts a follow-up visit to the patient.
[0009] In some embodiments, the report generation method further includes: obtaining the initial questions determined for the abnormal physical conditions of the patient; determining at least one proper noun from the initial questions and determining the extended words of at least one proper noun, where the extended words are any one of the following: synonyms of the proper noun, alternative names of the proper noun, abbreviations of the proper noun, foreign names of the proper noun; and determining the inquiry questions based on the extended words and the initial questions.
[0010] Based on the above technical solution, after obtaining the initial problem determined for the abnormal physical condition of the patient, the present application can determine at least one proper noun from the initial problem and determine the extended words of at least one proper noun. Then, based on the extended words and the initial problem, the interrogation questions are determined. In this way, in the case where a proper noun corresponds to multiple names, the extended names with the same meaning can be determined based on one name. Then, based on these extended names, the corresponding interrogation questions are determined, so that when the doctor asks the patient about the symptoms corresponding to a certain proper noun, no matter what the form of the question is, the corresponding reply content can be determined from the patient's answer, so as to accurately determine the reply content of the interrogation question and accurately generate an interrogation report.
[0011] In some embodiments, the report generation method further includes: determining the extended question of the initial question and the trigger condition of the extended question, the correlation degree between the extended question and the initial question is greater than the preset correlation degree, and the trigger condition is that the reply content of the initial question matches the preset content; based on the extended question and the initial question, determining the interrogation questions.
[0012] Based on the above technical solution, the present application can further determine the extended question of the initial question and determine the trigger condition of the extended question, so as to determine the interrogation questions based on the extended question and the initial question. This is because for some questions, the reply content of the patient may have multiple results, and different results will lead the doctor to ask different questions. Therefore, it is necessary to determine the corresponding extended questions for the initial question under different reply contents. Thus, the interrogation questions asked by the doctor to the patient can be further expanded to improve the doctor's understanding of the patient's condition and enhance the accuracy of generating the interrogation report.
[0013] In some embodiments, the report generation method further includes: determining the abnormal questions from the interrogation questions, the abnormal questions include at least one of the following: the questions for which the reply content cannot be determined from the interrogation text content, the questions for which the reply content determined from the interrogation text content is abnormal, the questions not mentioned in the interrogation text content, and the abnormal reply content includes at least one of the following: the reply content does not match the question, and the parameter value in the reply content exceeds the parameter range matching the question; marking and displaying the abnormal questions in the interrogation report.
[0014] Based on the above technical solution, when the present application determines the response content of the interrogation questions from the interrogation text content, if there are some questions for which the response content cannot be determined from the interrogation text content, or the response content of some questions determined from the interrogation text content does not match the questions or the parameter values in the response content exceed the parameter range, or there are questions not asked by the doctor, these abnormal questions can be marked and displayed in the generated interrogation report. Thereby prompting the doctor that these abnormal questions need to reconfirm the response content with the patient or need to further inquire about these questions. In this way, the accuracy of the generated interrogation report can be improved.
[0015] In some embodiments, based on the interrogation questions and the response content of the interrogation questions, an interrogation report is generated, including: determining the output format of the response content, where the output format includes at least one of the following: text format, parameter unit, parameter order; generating an interrogation report through the output format based on the interrogation questions and the response content of the interrogation questions.
[0016] Based on the above technical solution, the present application can further determine the output format of the response content so that when generating an interrogation report, the interrogation report can be generated through the set text format, parameter unit, and parameter order. In this way, by setting the output format of the response content, the interrogation report can be generated through a unified output format, thereby improving the readability of the interrogation report.
[0017] In some embodiments, the report generation method further includes: identifying the interrogation conversation between the doctor and the patient to obtain the initial text content; preprocessing the initial text content to obtain the interrogation text content, and the preprocessing includes at least one of the following: adjusting the correspondence between each text content and the doctor / patient, text correction processing, and dialogue segmentation processing, and the text correction processing includes at least one of the following: text deletion, text error correction, text replacement, and text completion.
[0018] Based on the above technical solution, the present application can generate the initial text content corresponding to the interrogation conversation by identifying the interrogation conversation between the doctor and the patient. And further preprocess the initial text content to adjust the correspondence between each text content and the doctor / patient, perform text correction processing on the initial text content, and perform dialogue segmentation processing on the initial text content, so as to obtain the required interrogation text content. In this way, by preprocessing the initial text content corresponding to the interrogation conversation, the incorrect text in the text content can be adjusted to obtain the corresponding interrogation text content. Thereby, when determining the response content of the interrogation questions from the interrogation text content, the accuracy of the determined response content can be improved.
[0019] In some embodiments, the preprocessing includes adjusting the correspondence between each piece of text content and the doctor / patient. The initial text content includes multiple texts. Preprocessing the initial text content to obtain the consultation text content includes: determining the voice characteristics of the doctor and the voice characteristics of the patient. The voice characteristics include at least one of the following: tone, speech rate, intonation, timbre, and speech clarity. Based on the voice characteristics of the doctor and the voice characteristics of the patient, determining the speaker corresponding to each text in the multiple texts to obtain the consultation text content, and the speaker is the doctor or the patient.
[0020] Based on the above technical solution, the present application can determine the speaker corresponding to each text in the multiple texts included in the initial text content corresponding to the consultation dialogue according to the voice characteristics of the doctor and the voice characteristics of the patient, so as to accurately mark the speaker of each text as the doctor or the patient in the obtained consultation text content. Thus, it is possible to avoid confusing the content spoken by the doctor and the patient and obtaining incorrect consultation text content, so that when determining the reply content of the consultation question from the consultation text content subsequently, the reply content can be accurately determined.
[0021] In some embodiments, the preprocessing includes dialogue segmentation processing. The initial text content includes multiple texts. Preprocessing the initial text content to obtain the consultation text content includes: determining the coding information of each text in the multiple texts and the semantics of each text. Based on the coding information of each text and the semantics of each text, dividing the multiple texts into at least two text segments to obtain the consultation text content.
[0022] Based on the above technical solution, the present application can divide the multiple texts into at least two text segments to obtain the consultation text content by determining the coding information of each text in the multiple texts included in the initial text content and the semantics of each text, and based on the coding information of each text and the semantics of each text. In this way, when the initial text content includes conversations on multiple topics, the two parts of the conversation with relatively weak association in the initial text content can be subjected to dialogue segmentation processing to separately obtain the consultation text content corresponding to each topic. Thus, when determining the reply content of the consultation question from the consultation text content subsequently, the mutual interference of the conversation content between topics can be avoided, so as to accurately determine the reply content of the consultation question.
[0023] In view of this, embodiments of the present application provide an electronic device, which is used to execute a report generation method. The electronic device includes: a processor and a communication interface, where the processor and the communication interface are coupled; the communication interface is configured to: obtain the consultation text content between a doctor and a patient; the processor is configured to: determine the reply content of the consultation questions from the consultation text content based on the pre-determined consultation questions; the consultation questions are obtained by modifying the initial questions based on a question modification rule, and the initial questions are questions related to the abnormal physical condition of the patient; the processor is further configured to: generate a consultation report based on the consultation questions and the reply content of the consultation questions.
[0024] In some embodiments, the question modification rule includes at least one of the following: expanding proper nouns in the initial question; annotating core words in the initial question; expanding the initial question and determining the triggering conditions for the extended questions obtained by the expansion.
[0025] In some embodiments, the communication interface is further configured to: obtain the initial questions determined for the abnormal physical condition of the patient; the processor is further configured to: determine at least one proper noun from the initial questions and determine the extended words of the at least one proper noun, where the extended words are any one of the following: synonyms of the proper noun, alternative names of the proper noun, abbreviations of the proper noun, foreign names of the proper noun; and determine the consultation questions based on the extended words and the initial questions.
[0026] In some embodiments, the processor is further configured to: determine the extended questions of the initial questions and the triggering conditions for the extended questions, where the degree of association between the extended questions and the initial questions is greater than a preset degree of association, and the triggering condition is that the reply content of the initial question matches the preset content; and determine the consultation questions based on the extended questions and the initial questions.
[0027] In some embodiments, the processor is further configured to: determine abnormal questions from the consultation questions, where the abnormal questions include at least one of the following: questions for which no reply content is determined from the consultation text content, questions for which the determined reply content is abnormal, questions not mentioned in the consultation text content, and abnormal reply content includes at least one of the following: the reply content does not match the question, and the parameter value in the reply content exceeds the parameter range matching the question; and mark and display the abnormal questions in the consultation report.
[0028] In some embodiments, the processor is specifically configured to: determine the output format of the reply content, where the output format includes at least one of the following: text format, parameter unit, parameter order; and generate a consultation report through the output format based on the consultation questions and the reply content of the consultation questions.
[0029] In some embodiments, the processor is further configured to: identify the consultation dialogue between the doctor and the patient to obtain the initial text content; preprocess the initial text content to obtain the consultation text content, and the preprocessing includes at least one of the following: adjusting the correspondence between each text content and the doctor / patient, text correction processing, and dialogue segmentation processing. The text correction processing includes at least one of the following: text deletion, text error correction, text replacement, and text completion.
[0030] In some embodiments, the preprocessing includes adjusting the correspondence between each text content and the doctor / patient. The initial text content includes multiple texts. The processor is specifically configured to: determine the voice features of the doctor and the voice features of the patient, and the voice features include at least one of the following: tone, speech rate, intonation, timbre, and speech clarity; based on the voice features of the doctor and the voice features of the patient, determine the speaker corresponding to each text in the multiple texts to obtain the consultation text content, and the speaker is the doctor or the patient.
[0031] In some embodiments, the preprocessing includes dialogue segmentation processing. The initial text content includes multiple texts. The processor is specifically configured to: determine the encoding information of each text in the multiple texts and the semantics of each text; based on the encoding information of each text and the semantics of each text, divide the multiple texts into at least two text segments to obtain the consultation text content.
[0032] In another aspect, a report generation device is provided, including a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to run a computer program or instruction to implement the report generation method according to any one of the above embodiments.
[0033] In another aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer program instructions. When the computer program instructions run on the computer, the computer is caused to execute the report generation method according to any one of the above embodiments.
[0034] In another aspect, a computer program product is provided. The computer program product includes computer program instructions. When the computer program instructions are executed on the computer, the computer program instructions cause the computer to execute the report generation method according to any one of the above embodiments.
[0035] In another aspect, a computer program is provided. When the computer program is executed on the computer, the computer program causes the computer to execute the report generation method according to any one of the above embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To more clearly illustrate the technical solutions in the present disclosure, the following will briefly introduce the drawings required for some embodiments of the present disclosure. Obviously, the drawings in the following description are only the drawings of some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings. In addition, the drawings in the following description can be regarded as schematic diagrams and do not limit the actual dimensions of the products involved in the embodiments of the present disclosure, the actual processes of the methods, the actual timings of the signals, etc.
[0037] Figure 1 Structural diagram of an electronic device according to some embodiments;
[0038] Figure 2 Flowchart of a report generation method according to some embodiments;
[0039] Figure 3 Structural schematic diagram of a problem processing module according to some embodiments;
[0040] Figure 4 Schematic diagram of prompt word display according to some embodiments;
[0041] Figure 5 Schematic diagram of a consultation report according to some embodiments;
[0042] Figure 6 Structural schematic diagram of a dialogue processing module according to some embodiments;
[0043] Figure 7 Structural schematic diagram of a text chunking model according to some embodiments;
[0044] Figure 8 Schematic diagram of text preprocessing according to some embodiments;
[0045] Figure 9 Structural schematic diagram of a large language model according to some embodiments;
[0046] Figure 10 Schematic diagram of an aggregation result according to some embodiments;
[0047] Figure 11 Structural schematic diagram of another large language model according to some embodiments;
[0048] Figure 12 Structural schematic diagram of a comprehensive model according to some embodiments;
[0049] Figure 13 Structural schematic diagram of a report generation device according to some embodiments;
[0050] Figure 14Schematic structural diagram of another report generation device according to some embodiments. Detailed implementation manners
[0051] The technical solutions in some embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are only a part rather than all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided in the present disclosure shall fall within the protection scope of the present disclosure.
[0052] Unless the context requires otherwise, throughout the specification and claims, the term "comprise" and its other forms, such as the third-person singular form "comprises" and the present participle form "comprising", are interpreted as open and inclusive, that is, "including, but not limited to". In the description of the specification, the terms "one embodiment", "some embodiments", "exemplary embodiments", "example", "specific example" or "some examples", etc. are intended to indicate that the specific features, structures, materials or characteristics related to the embodiment or example are included in at least one embodiment or example of the present disclosure. The schematic representations of the above terms are not necessarily referring to the same embodiment or example. In addition, the specific features, structures, materials or characteristics may be included in any one or more embodiments or examples in any appropriate manner.
[0053] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present disclosure, unless otherwise stated, the meaning of "a plurality" is two or more.
[0054] "At least one of A, B, and C" has the same meaning as "at least one of A, B, or C", and both include the following combinations of A, B, and C: only A, only B, only C, the combination of A and B, the combination of A and C, the combination of B and C, and the combination of A, B, and C.
[0055] "A and / or B" includes the following three combinations: only A, only B, and the combination of A and B.
[0056] As used herein, depending on context, the term "if" is optionally interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting". Similarly, depending on context, the phrases "if it is determined that..." or "if [stated condition or event] is detected" are optionally interpreted to mean "when it is determined that..." or "in response to determining..." or "when [stated condition or event] is detected" or "in response to detecting [stated condition or event]".
[0057] The use of "is applicable to" or "is configured to" in this document means open and inclusive language, which does not exclude devices that are applicable to or configured to perform additional tasks or steps.
[0058] Additionally, the use of "based on" means open and inclusive, because a process, step, calculation, or other action "based on" one or more conditions or values can in practice be based on additional conditions or values beyond those stated.
[0059] In recent years, in order to improve the service level for patients, medical institutions can, through a follow-up mechanism, track and observe the recovery status of patients after they seek medical treatment, so as to understand the patient's condition, medication situation, recovery status, etc. in real time. This follow-up mechanism facilitates doctors to track and observe the patient's condition, obtain first-hand information for statistical analysis and experience accumulation. And it can also statistically analyze the recovery status of a large number of patients, analyze the condition to accumulate treatment experience, and patients can timely consult doctors about their condition and health problems. For example, doctors and patients can establish a network call. Based on the patient's condition and historical treatment situation, the doctor determines the health problems that need to be asked of the patient to inquire about the patient's condition recovery. Thus, the patient can answer the corresponding condition recovery according to the questions asked by the doctor, and the doctor can record the patient's condition recovery and fill out a medical history report.
[0060] However, when the network call is interrupted or the patient's description of their condition recovery is not clear and accurate enough, the doctor cannot accurately record the patient's condition recovery; or, limited by this way of network call, the call duration between the doctor and the patient is limited, the doctor will miss some questions that need to be asked, and the patient cannot feedback more health conditions to the doctor. Therefore, when current doctors conduct follow-up on patients to understand the patient's condition recovery, the efficiency of obtaining a medical history report based on the conversation content between the doctor and the patient is low and the accuracy is poor.
[0061] With the emergence of large language models, the inference and summarization capabilities of large language models can be utilized to combine the conversations between doctors and patients with follow-up questions and input them into the large language model for unified processing. However, when the follow-up conversation duration is long or there are many follow-up questions, this solution cannot guarantee the accuracy of the output results of the large language model and the output format, nor can it provide real-time feedback on the problems existing during the communication between doctors and patients. Currently, there are existing solutions that have been tried in the fields of dialogue and medical follow-up, but none of them can be applied to the scenario of real-time medical follow-up.
[0062] In view of this, the embodiments of the present application provide a report generation method. Based on the obtained medical consultation text content between doctors and patients and in combination with pre-determined medical consultation questions, the reply content of the medical consultation questions can be determined from the medical consultation text content. Then, based on the medical consultation questions and the reply content of the medical consultation questions, a medical consultation report is generated. Importantly, the pre-determined medical consultation questions are obtained by correcting the initial questions based on a question correction rule, and the initial questions are questions related to the abnormal physical conditions of the patients. Therefore, when a doctor conducts a post-diagnosis follow-up on a patient's condition, initial questions can be set in advance for the patient's abnormal physical conditions, and the initial questions can be corrected to obtain medical consultation questions that are targeted at the patient's condition. Thus, during the post-diagnosis follow-up, based on the medical consultation text content corresponding to the follow-up conversation between the doctor and the patient, the reply content of the patient to the pre-determined medical consultation questions can be determined, and thus a medical consultation report for the doctor can be generated. In this way, there is no need for the doctor to manually record the patient's condition recovery, and there is no need to spend a large amount of recording time. Therefore, even if the call duration between the doctor and the patient is limited, sufficient communication can be carried out between the doctor and the patient, and important questions related to the patient's condition can be avoided from being omitted. Based on this method, when a doctor conducts a post-diagnosis follow-up on a patient to understand the patient's condition recovery, the efficiency and accuracy of obtaining the medical consultation report can be improved.
[0063] The following will describe in detail the implementation manners of the embodiments of the present application in conjunction with the accompanying drawings of the specification.
[0064] As Figure 1 shown, Figure 1 This is a structural diagram of an electronic device 100 provided by an embodiment of the present application. The electronic device 100 can be a device with communication capabilities and data processing capabilities, such as a computer. The electronic device 100 can include a communication interface 101 and a processor 102, and the processor 102 is coupled to the communication interface 101.
[0065] Exemplarily, the computer in the embodiments of the present application may be a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, as well as a personal digital assistant (PDA), an augmented reality (AR) / virtual reality (VR) device, etc. The embodiments of the present application do not impose special restrictions on the specific form of the device.
[0066] The execution subject of the report generation method provided by the present application may be the central processing unit (CPU) of the computer, or the processing module for implementing report generation in the computer, or the application system for implementing report generation in the computer.
[0067] The computer in the embodiments of the present application may be a server, a desktop computer, a notebook computer, a mobile phone, a cloud device or a virtual machine, etc. It can provide a way to access service logic for use by the programs (systems) of client applications. The computer can provide a simple and manageable access mechanism to system resources for application programs. It also provides services such as the implementation of the hypertext transfer protocol (HTTP) and database connection management.
[0068] In the embodiments of the present application, the electronic device 100 may obtain the consultation text content between the doctor and the patient through the communication interface 101.
[0069] It should be noted that in the embodiments of the present application, the consultation text content between the doctor and the patient may be obtained by recognizing the consultation conversation between the doctor and the patient; the consultation text content may also be obtained from the text chat content between the doctor and the patient.
[0070] Specifically, the consultation conversation between the doctor and the patient may be recorded by a recording device, and then the voice may be converted into the corresponding text content through speech recognition technology to obtain the consultation text content between the doctor and the patient.
[0071] The electronic device 100 can be connected to other devices through the communication interface 101. Exemplary other devices can be computer devices, recording devices, digital tablets, or mice, etc. For example, the computer device sends the consultation dialogue monitored by the recording device to the electronic device 100 through the communication interface 101, so that the electronic device 100 can obtain the consultation text content between the doctor and the patient through the communication interface 101. Another example is that the user can input the consultation text content between the doctor and the patient through an input device such as a digital tablet or a mouse, and transmit the consultation text content between the doctor and the patient to the electronic device 100 through the communication interface 101.
[0072] The processor 102 in the electronic device 100 can be used to determine the reply content of the consultation question from the consultation text content based on a pre-determined consultation question; the consultation question is obtained by correcting the initial question based on the question correction rule, and the initial question is a question related to the abnormal physical condition of the patient; the processor is also configured to: generate a consultation report based on the consultation question and the reply content of the consultation question.
[0073] In the embodiments of the present application, the question correction rule includes at least one of the following: expanding the proper nouns in the initial question; annotating the core words in the initial question; expanding the initial question and determining the trigger condition of the extended question obtained by the expansion.
[0074] The communication interface 101 in the electronic device 100 can also obtain the initial question determined for the abnormal physical condition of the patient.
[0075] The processor 102 in the electronic device 100 can also determine at least one proper noun from the initial question and determine the expansion words of the at least one proper noun. The expansion words are any one of the following: synonyms of the proper noun, alternative names of the proper noun, abbreviations of the proper noun, foreign names of the proper noun; based on the expansion words and the initial question, determine the consultation question.
[0076] The processor 102 in the electronic device 100 can also determine the extended question of the initial question and the trigger condition of the extended question. The correlation degree between the extended question and the initial question is greater than the preset correlation degree, and the trigger condition is that the reply content of the initial question matches the preset content; based on the extended question and the initial question, determine the consultation question.
[0077] The processor 102 in the electronic device 100 can also determine abnormal questions from the consultation questions. The abnormal questions include at least one of the following: questions for which no answer content is determined from the consultation text content, questions with abnormal answer content determined from the consultation text content, and questions not mentioned in the consultation text content. Abnormal answer content includes at least one of the following: the answer content does not match the question, and the parameter value in the answer content exceeds the parameter range matching the question; mark and display the abnormal questions in the consultation report.
[0078] The processor 102 in the electronic device 100 can specifically determine the output format of the answer content. The output format includes at least one of the following: text format, parameter unit, and parameter order; generate a consultation report based on the consultation questions and the answer content of the consultation questions through the output format.
[0079] The processor 102 in the electronic device 100 can also identify the consultation dialogue between the doctor and the patient to obtain the initial text content; preprocess the initial text content to obtain the consultation text content. The preprocessing includes at least one of the following: adjusting the correspondence between each text content and the doctor / patient, text correction processing, and dialogue segmentation processing. The text correction processing includes at least one of the following: text deletion, text error correction, text replacement, and text completion.
[0080] When the preprocessing includes adjusting the correspondence between each text content and the doctor / patient, and the initial text content includes multiple texts, the processor 102 in the electronic device 100 can specifically determine the voice characteristics of the doctor and the voice characteristics of the patient. The voice characteristics include at least one of the following: tone, speech rate, intonation, timbre, and speech clarity; based on the voice characteristics of the doctor and the voice characteristics of the patient, determine the speaking object corresponding to each text in the multiple texts to obtain the consultation text content, and the speaking object is the doctor or the patient.
[0081] When the preprocessing includes dialogue segmentation processing and the initial text content includes multiple texts, the processor 102 in the electronic device 100 can specifically determine the coding information of each text in the multiple texts and the semantics of each text; based on the coding information of each text and the semantics of each text, divide the multiple texts into at least two text segments to obtain the consultation text content.
[0082] In the embodiments of the present application, the processor 102 may be a chip. The chip may include five major categories: logic chips, storage chips, sensor chips, power chips, and communication chips. Among them, the processor category mainly includes chips that undertake specific computing and control tasks in the system, such as microcontroller units (MCUs), also known as single-chip microcomputers or single-chip microcontrollers, central processing units (CPUs), graphics processing units (GPUs), network processors (NPUs), etc. The storage category mainly includes chips that undertake data storage in the system, as well as some storage controller chips, such as dynamic random access memories (DRAMs), static random access memories (SRAMs), Flash, etc. The sensing category mainly includes chips that undertake information collection, presentation, and interaction in the system, such as input / output devices and some signal processing chips. The communication category (wired and wireless) mainly includes chips that undertake communication functions in the system. For example, some Ethernet chips, switching chips, wide-area and local-area network, point-to-point and ad-hoc network chips, as well as devices such as filters, amplifiers, and power supplies for assisting communication can all belong to this category. What the public usually knows, such as Wi-Fi, Bluetooth, 5G basebands, global positioning systems (GPS), narrow-band Internet of Things (NB-IoT), network cards, switches, etc., can all be classified into this category.
[0083] The methods in the following embodiments can all be implemented in the electronic device 100 with the above hardware structure. In the following embodiments, the above electronic device 100 is taken as an example to illustrate the methods of the embodiments of the present application.
[0084] The following describes in detail the report generation method provided by the embodiments of the present application with reference to the accompanying drawings.
[0085] The report generation method of the embodiments of the present application can be applied to the scenario where a doctor conducts a post-diagnosis follow-up visit on a patient. As Figure 2 shown, the report generation method may include S201 - S203. The following provides a detailed description of S201 - S203.
[0086] S201. The electronic device obtains the text content of the consultation between the doctor and the patient.
[0087] In the embodiments of the present application, when a doctor conducts a post - diagnosis follow - up on a patient, the doctor can communicate with the patient over a long distance through a network call (voice call), so that the doctor can ask relevant questions based on the patient's condition and historical treatment situation to understand the patient's condition recovery.
[0088] It should be noted that the post - diagnosis follow - up is a process in which the doctor asks about the patient's condition again after a period of recovery after targeted treatment for the patient's condition. For example, after a patient is hospitalized for their illness in the hospital and is discharged home for recuperation when the condition is stable. The doctor can make a phone call to the patient after a certain period of time (such as one week, half a month, one month, etc.) after the patient is discharged to understand the patient's condition recovery.
[0089] Optionally, the text content of the medical consultation between the doctor and the patient can be obtained by processing the recording of the network call between the doctor and the patient, or can also be obtained by processing the recording with a recording device when the doctor and the patient are having a face - to - face conversation, or can also be obtained by processing the communication records (such as text messages or voice messages, etc.) of the communication between the doctor and the patient through social software.
[0090] S202. The electronic device determines the reply content of the interrogation question from the text content of the interrogation based on the pre - determined interrogation question.
[0091] Among them, the interrogation question is obtained by correcting the initial question based on the question correction rule, and the initial question is a question related to the patient's abnormal physical condition.
[0092] Optionally, the doctor can preset the initial questions related to the patient's condition for different patients, and then the electronic device corrects the initial questions based on the question correction rule to obtain the interrogation questions.
[0093] In some embodiments, the question correction rule includes at least one of the following:
[0094] 1 - 1. Expand the proper nouns in the initial question.
[0095] Among them, the proper nouns can refer to disease names, treatment methods, drug names, etc.
[0096] In some embodiments, obtain the initial questions determined for the patient's abnormal physical condition; determine at least one proper noun from the initial questions, and determine the expansion words of at least one proper noun, where the expansion words are any one of the following: synonyms of the proper noun, alternative names of the proper noun, abbreviations of the proper noun, foreign names of the proper noun; based on the expansion words and the initial questions, determine the interrogation questions.
[0097] It should be noted that for the interrogation questions asked by doctors to patients, the same concept can be expressed in multiple ways. Therefore, the expressions and focuses of each interrogation question are different. If the initial question is directly recognized by the model without modification to determine the response content, the expected result is often not obtained. For example, the model cannot understand the key points of the interrogation question, the medical terms are not clear, the format of the patient's answer is not standardized, etc. Therefore, it is necessary to correct the initial question through the set question correction rules to obtain the interrogation question.
[0098] It can be understood that a proper noun corresponds to multiple expressions (such as Chinese name, English name, abbreviation, common name, etc.). Therefore, it is necessary to expand multiple expressions that represent the same feature (concept). For example, "Transient Ischemic Attack" can also be expressed by the English name "Transient Ischemic Attack", or by the English abbreviation "TIA", or by the common name "mini-stroke".
[0099] 1-2. Mark the core words in the initial question.
[0100] Among them, the core words can be feature words closely related to the patient's condition to clearly point out the key points of concern in the question. For example, the core words can be blood sugar, blood pressure, recurrence, date, etc. For example: The following content in the question needs to be focused on: "Pay attention to the patient's blood sugar in the recent 7 days", "If the patient has relapsed multiple times, pay attention to the date of the latest relapse", etc.
[0101] 1-3. Expand the initial question and determine the trigger conditions for the extended extended questions.
[0102] Among them, the trigger conditions are used to determine whether there are extended questions for the initial question.
[0103] In some embodiments, determine the extended questions of the initial question and the trigger conditions for the extended questions. The correlation degree between the extended questions and the initial question is greater than the preset correlation degree. The trigger condition is that the response content of the initial question matches the preset content; based on the extended questions and the initial question, determine the interrogation questions.
[0104] Optionally, the initial question can be expanded, and the triggering conditions for the extended follow-up questions can be determined. It can be understood that some questions have corresponding sub-questions. When the response content of the question is different, different sub-questions need to be further asked. Therefore, based on the response content of the patient to the previous question (i.e., the triggering condition), the corresponding sub-question (i.e., the follow-up question) needs to be determined. For example, if the previous question is "Have you had a recurrence after discharge?", and the patient's response is "Yes, there has been a recurrence", then this response content corresponds to a new sub-question "When did the recurrence occur?". If the patient's response is "No, there has been no recurrence", then this response content has no corresponding new sub-question.
[0105] In some embodiments, the electronic device can correct the initial question through the question processing module. As Figure 3 shown, the question processing module includes: a proper noun expansion unit, a key concern index analysis unit, a question expansion unit, and an output format setting unit. By inputting the initial question into the proper noun expansion unit, the key concern index analysis unit, the question expansion unit, and the output format setting unit in sequence for processing, the consultation questions can be obtained.
[0106] It should be noted that the proper noun expansion unit is used to expand the proper nouns in the initial question; the key concern index analysis unit is used to mark the core words in the initial question; the question expansion unit is used to expand the initial question and determine the triggering conditions for the extended follow-up questions; the output format setting unit is used to determine the output format of the response content (specifically, reference can be made to the following embodiments).
[0107] Exemplarily, as Figure 4 shown, based on the pre-determined consultation questions, during the communication between the doctor and the patient, the electronic device can display prompt words to prompt the doctor with the consultation questions, follow-up questions (triggering conditions), proper nouns and their expansions, core words, etc. that need to be asked.
[0108] S203. The electronic device generates a consultation report based on the consultation questions and the response content of the consultation questions.
[0109] Optionally, as Figure 5 shown, in the consultation report, the consultation questions that the doctor has asked and the consultation questions that have not been asked can be displayed, and some precautions can also be displayed. For example: certain physical characteristics of the patient have not been asked (such as whether the patient can take care of himself / herself, how the patient recovered after discharge, blood pressure, blood sugar, etc.), the health parameters of the patient are abnormal (such as blood sugar of 20.0 mmol / L, far exceeding the normal range), and the patient's questions have not been answered (such as the patient reported dizziness, but the doctor did not ask further).
[0110] Optionally, the generation format of the consultation report and the display format of relevant parameters (reply content) appearing in the consultation report can be set in advance.
[0111] In some embodiments, the above S203 specifically includes: determining the output format of the reply content, where the output format includes at least one of the following: text format, parameter unit, parameter order; generating a consultation report through the output format based on the consultation question and the reply content of the consultation question.
[0112] Exemplarily, the format and specification of the model output consultation report can be set in advance through the output format setting unit in the question processing module. For example: setting the blood pressure value to be displayed in the format of "low pressure / high pressure", the unit of weight is kg, and the unit of height is cm, etc.
[0113] In the embodiments of the present application, by setting the output format of the reply content, the consultation report can be displayed in a unified standard display format, so that when the doctor views the consultation report, it is convenient for the doctor to understand the content displayed in the consultation report and avoid the situation where the doctor misinterprets the consultation report.
[0114] In some embodiments, determine the abnormal questions from the consultation questions, and mark and display the abnormal questions in the consultation report. For example, any of the following methods can be used for marking and display: bold text, highlighted text, colored text, etc.
[0115] Among them, the abnormal questions include at least one of the following: questions for which no reply content is determined from the consultation text content, questions with abnormal reply content determined from the consultation text content, questions not mentioned in the consultation text content, and abnormal reply content includes at least one of the following: the reply content does not match the question, and the parameter value in the reply content exceeds the parameter range matching the question.
[0116] Optionally, when generating a consultation report based on the consultation question and the reply content of the consultation question, the electronic device can further analyze the consultation question and the reply content of the consultation question to determine whether corresponding reply content is determined from the consultation text content for all consultation questions, or whether all consultation questions are mentioned in the consultation text content. If certain consultation questions do not have corresponding reply content determined from the consultation text content or are not mentioned, these consultation questions are determined as abnormal questions.
[0117] And, after determining the reply content corresponding to the consultation question from the consultation text content, further analyze whether the reply content matches the consultation question, or whether the parameter value in the reply content exceeds the parameter range matching the consultation question. If the reply content does not match the consultation question, or the parameter value in the reply content exceeds the parameter range matching the consultation question, the consultation question is determined as an abnormal question.
[0118] Exemplarily, assume the inquiry question is: "What medicine is taken for blood sugar control?", and the patient's reply is: "Took medicine A (for anti - inflammation)"; Obviously, medicine A is not a medicine for controlling blood sugar, so it is determined that the reply content does not match the inquiry question. Another example, the inquiry question is: "What are the blood pressure values?", and the patient's reply is: "The value of the high blood pressure is fifty, and the value of the low blood pressure is one hundred"; And the value of the high blood pressure being fifty is significantly beyond the normal parameter range of the high blood pressure (for example, 90 mmHg - 140 mmHg), so it is determined that the reply content does not match the inquiry question.
[0119] In this way, the doctor can clarify the inquiry questions for the follow - up visit, highlighting the key concerns and what answers are expected, and expand the details of the inquiry questions and concerns as much as possible to help improve the accuracy rate of the large - model during intelligent analysis. The doctor also needs to closely pay attention to the prompt content in the interface displayed on the electronic device, such as not repeating the questions that have already been asked, controlling the rhythm of the inquiry, and generally following the track of the set questions for the follow - up visit. For the real - time reminder content, the doctor can repeat the inquiry or further expand the inquiry. After the call ends, click the report generation button to generate an inquiry report.
[0120] The embodiment of the present application provides a report generation method. It can determine the reply content of the inquiry question from the inquiry text content between the doctor and the patient based on the obtained inquiry text content and in combination with the pre - determined inquiry questions. Then, based on the inquiry questions and the reply content of the inquiry questions, an inquiry report is generated. Importantly, the pre - determined inquiry questions are obtained by correcting the initial questions based on the question correction rules, and the initial questions are questions related to the abnormal physical condition of the patient. Therefore, when the doctor conducts a post - diagnosis follow - up on the patient's condition, the initial questions can be set in advance for the patient's abnormal physical condition, and the initial questions are corrected to obtain inquiry questions targeted at the patient's condition. Thus, during the post - diagnosis follow - up, based on the inquiry text content corresponding to the follow - up conversation between the doctor and the patient, the reply content of the patient to the pre - determined inquiry questions can be determined, and then the doctor's inquiry report can be generated. In this way, there is no need for the doctor to manually record the patient's condition recovery, and there is no need to spend a large amount of recording time. Therefore, even if the call duration between the doctor and the patient is limited, sufficient communication can be carried out between the doctor and the patient, avoiding missing important questions related to the patient's condition. Based on this method, when the doctor conducts a post - diagnosis follow - up on the patient to understand the patient's condition recovery, the efficiency and accuracy of obtaining the inquiry report can be improved.
[0121] In some embodiments, taking the post - consultation follow - up by means of an online call between a doctor and a patient as an example, when an online call is made between a doctor and a patient, the consultation dialogue between the doctor and the patient can be recognized to obtain the initial text content; then, the initial text content is pre - processed to obtain the consultation text content, and the pre - processing includes at least one of the following: adjusting the correspondence between each text content and the doctor / patient, text correction processing, and dialogue segmentation processing. The text correction processing includes at least one of the following: text deletion, text error correction, text replacement, and text completion.
[0122] Optionally, when an online call is made between a doctor and a patient, the Automatic Speech Recognition (ASR) technology can be used to convert the acquired call voice into text content (i.e., the initial text content) in real time.
[0123] It should be noted that since the call content between a doctor and a patient often includes some filler words, dialects, and interrupted speaking situations during a voice call, the directly obtained initial text content is often very rough, including a large number of low - quality sentences or words. Moreover, as the dialogue between the doctor and the patient deepens, the obtained initial text content becomes longer and longer. If the initial text content is not segmented and all is input into the model, it is likely to result in a lower recognition accuracy.
[0124] Therefore, it is necessary to modify and process the quality of the obtained initial text content. Additionally, the initial text content can be intelligently segmented so that after being input into the model, the model can process each text content separately and recognize the consultation questions and reply contents included in that segment, thereby improving the accuracy of the recognized content.
[0125] Optionally, in the initial text content obtained by performing ASR recognition on the voice call between a doctor and a patient, there may be a situation where the roles of the doctor and the patient are reversed, resulting in an identification error. That is, the content spoken by the doctor is identified as the content spoken by the patient, or the content spoken by the patient is identified as the content spoken by the doctor. Therefore, the user information recognized in the early stage of the dialogue can be used to adjust the initial text content based on the features recognized from the previous dialogue (such as tone, speech rate, clarity of pronunciation, etc.), and reasonably distribute the text content to each speaking user (doctor or patient).
[0126] Optionally, in the initial text content obtained by performing ASR recognition on the voice call between a doctor and a patient, there may also be a large number of filler words, such as "um... that... well". Such meaningless filler words also need to be processed by text deletion, otherwise it will affect the effect of subsequent text processing.
[0127] Optionally, in the initial text content obtained by performing ASR recognition on the voice call between the doctor and the patient through the model, there may also be error texts (such as "lower medicine") due to too fast speech rate or errors occurring during the ASR recognition conversion process. Therefore, it is necessary to perform text error correction on the text to correct "lower medicine" to "antihypertensive medicine".
[0128] Optionally, in the initial text content obtained by performing ASR recognition on the voice call between the doctor and the patient, there may also be colloquial expressions of the condition or drug names due to the habits or dialects of the speaking users. Therefore, in order to improve the recognition accuracy of the model, it is necessary to perform text replacement or text completion processing on the text. For example, replace "Drug A" with "a name".
[0129] Optionally, as the call between the doctor and the patient progresses, the entire call content will include multiple topic contents, and some topic contents unrelated to the condition are not necessary to be input into the model for analysis and processing. Therefore, the obtained initial text content can be segmented, and the dialogue content unrelated to the condition can be removed, so that the model focuses on the dialogue content related to the condition.
[0130] Exemplarily, in the content of the call between the doctor and the patient, the topic contents included in the dialogue can be: condition topic (communicating the condition recovery situation, current physical condition), medication topic (patient's medication situation, dosage, medication duration, etc.), diet topic (patient's diet, food intake, etc.), and other topics unrelated to the condition (such as family situation, income, etc.).
[0131] In some embodiments, the electronic device can preprocess the dialogue between the doctor and the patient through the dialogue processing module. As Figure 6 shown, the dialogue processing module includes: an ASR recognition unit, a speaker logical consistency correction unit, a text correction unit, a dialogue segmentation unit, and an interrogation text generation unit. By collecting the dialogue between the doctor and the patient and inputting it into each unit in the dialogue processing module for sequential processing, the interrogation text content can be obtained. The text correction unit has functions such as text deletion, text error correction, text completion, and text replacement.
[0132] It should be noted that the ASR recognition unit is used to recognize the dialogue between the doctor and the patient to generate the initial text content; the speaker logical consistency correction unit is used to adjust the corresponding relationship between each text content and the doctor / patient; the text correction unit is used to perform text deletion, text error correction, text replacement, and text completion functions; the dialogue segmentation unit is used to segment the text content; and the interrogation text generation unit is used to generate the interrogation text content.
[0133] In some embodiments, the preprocessing includes adjusting the correspondence between each piece of text content and the doctor / patient. The initial text content includes multiple texts. The preprocessing of the initial text content to obtain the consultation text content includes: determining the voice characteristics of the doctor and the patient. The voice characteristics include at least one of the following: tone, speech rate, intonation, timbre, and speech clarity. Based on the voice characteristics of the doctor and the patient, determining the speaker corresponding to each text in the multiple texts to obtain the consultation text content, where the speaker is the doctor or the patient.
[0134] Exemplarily, at the beginning of a voice call between a doctor and a patient, the patient and the doctor usually greet each other (for example, the patient may say "Hello, Doctor Wang", and the doctor may say "Hello, Xiaohong. How is your recovery?" etc.). Based on this content and identifying the characteristics such as the tone, speech rate, intonation, timbre, and speech clarity of the voice corresponding to this content, the identity information of each speaker (doctor or patient) can be determined.
[0135] In some embodiments, the preprocessing includes dialogue segmentation processing. The initial text content includes multiple texts. The preprocessing of the initial text content to obtain the consultation text content includes: determining the coding information of each text in the multiple texts and the semantics of each text. Based on the coding information of each text and the semantics of each text, dividing the multiple texts into at least two text segments to obtain the consultation text content.
[0136] Optionally, the dialogue segmentation processing can be performed through a text chunking model. Specifically, an Encoder model based on Transformer can be used to perform text segmentation semantically to ensure accurate segmentation of the text after a dialogue topic discussion is completed.
[0137] Exemplarily, such as Figure 7As shown, it is a schematic structural diagram of a text chunking model, which includes: a multi-head self-attention mechanism layer, a layer normalization layer, and a feed-forward neural network layer. The text chunking model is based on the structure and input format of the Transformer-based Encoder model. By combining the text of the same speaker using cls text1 and sep text2, and then performing token embedding processing. Further, by combining the correct segment embedding and position embedding, the token embedding is fed into the encoder structure of the Transformer. Finally, the output [CLS] token is obtained, and through a fully connected (FC) layer and a Softmax function, the identification of whether to segment is obtained.
[0138] Optionally, when training the model, NPS Loss is used for fitting optimization.
[0139] It should be noted that segment embedding plays an important role in the BERT model, mainly used to distinguish different sentences or paragraphs in the input text. When the BERT model processes an input containing multiple sentences, it uses segment embedding to distinguish these different sentences, so as to better understand the overall structure and semantics of the text. Also, position embedding plays an important role in natural language processing (NLP), especially in Transformer-based models. It is used to encode the position information for each position in the input sequence, helping the model understand the order relationship of the elements in the sequence, thereby improving the model's processing ability for sequence data.
[0140] Exemplarily, such as Figure 8 As shown, by preprocessing the initial text content (i.e., the original conversation) generated from the conversation between the doctor and the patient, the correspondence between the text content and the doctor / patient can be adjusted (for example, "about a hundred" is what the patient said, but is recorded as what the doctor said in the initial text content), text correction processing (for example, deleting filler words such as "um... that..."), and dialogue segmentation processing. Thus, the preprocessed consultation text content is obtained.
[0141] In some embodiments, the electronic device can determine the reply content of the consultation question from the consultation text content through the model analysis module based on the pre-determined consultation question. Exemplarily, such as Figure 9As shown, by inputting the consultation text content and consultation questions into a model (such as a large language model (LLM)), the consultation text content is combined with different consultation questions (prompts) and sent to the model for reasoning in a parallel processing manner to determine the answer content of each consultation question from the consultation text content.
[0142] It should be noted that parallel reasoning can greatly shorten the reasoning time while achieving the same reasoning effect. Specifically, you can use reasoning acceleration frameworks such as vLLM and TensorRT-LLM. vLLM (Virtual LargeLanguage Model) is an open source code library designed to help large language models (LLM) perform large-scale calculations more efficiently. It is mainly used to define, optimize and execute reasoning of large language models in production environments. TensorRT-LLM is built based on the TensorRT deep learning compilation framework, drawing on the efficient Kernels implementation in FastTransformer, and using NCCL to complete communication between devices. The framework aims to provide high-performance reasoning solutions and support multi-card or multi-machine reasoning, supporting the processing of large-scale models through two parallel mechanisms: Tensor Parallelism and Pipeline Parallelism.
[0143] Optionally, the output results after model inference (i.e., the answers to the determined medical questions) need to be post-processed and verified. For example, the output unit of weight is kg, and the cost is expressed in Arabic numerals. In addition, if the output result contains the prompt message "not mentioned", it means that the medical question has not been mentioned by the doctor; if there is a clear answer in the output result, the result is recorded and the report is generated; if the answer in the output result is unclear or exceeds the normal logical range, the doctor is reminded in real time and can ask the question again or confirm the follow-up question.
[0144] In some embodiments, the above models need to be trained. Currently, there are many open-source large language models, and their capabilities in question-answering, reasoning, mathematics, etc. are also remarkable. However, the ability to follow instructions for specific occasions is often insufficient. Based on this, it is necessary to fine-tune the instructions to make the large models better adapt to these tasks. Different "instruction" instructions will be assigned to different tasks when performing different tasks.
[0145] Exemplarily, the data organization form can be set to the json format. And through "instruction", it is indicated that the task of this model is: the information extraction task in medical follow-up conversations. Thus, based on the content of the medical interview text, when the input interview question is "Doctor: When was the last hospitalization? What was the reason for hospitalization?", the corresponding reply content (output) can be analyzed and determined from the medical interview text content as "Patient: It was in October 2018, and the reason for hospitalization was a cerebral apoplexy."
[0146] Optionally, fine-tuning the large model instructions often requires mixing general domain and specific domain data to avoid catastrophic forgetting of the original knowledge of the large model. For the construction of specific domain instruction data, a problem processing module (problem prompt engineering) can be used to determine the interview questions.
[0147] Specifically, for the construction of general domain instruction data, look for open-source instruction data sets and perform necessary screening and filtering. Then, feature selection can be performed on the instruction data set to be processed, and the TF-IDF method or the Sentence-Embedding method can be used. And, clustering is performed using the Kmeans or DBSCAN aggregation method. The Kmeans and DBSCAN aggregation methods each have their own advantages. The former can clearly specify the number of clusters after aggregation, and the latter can cluster texts that are semantically similar together without being restricted by the number of clusters.
[0148] Exemplarily, the aggregation method of the k-means clustering algorithm (Kmeans) or the density-based clustering algorithm (Density-Based Spatial Clustering of Applications with Noise, DBSCAN), the aggregation results on the same data set are as Figure 10 shown. For each cluster after clustering, n instructions can be selected (n is determined according to the required number of general instructions).
[0149] Furthermore, a model can be constructed, such as Figure 11As shown in the figure, a model is constructed based on the decoder part of the Transformer. The main components in each Transformer block are: Attention Layer (attention layer), RMSNorm (Root Mean Square Normalization), which is a normalization technique for deep learning models, and Multilayer Perceptron (MLP). Among them, the Attention layer adds Rotary Position Encoding (RoPE) to the QK tensors. This position encoding formally depends on absolute position encoding and can become relative position encoding during calculation, which is very helpful for learning the context semantics of long texts.
[0150] It should be noted that in the Feed Forward, two linear transformations need to be learned, and the ReLU (Rectified Linear Unit) activation function is used between these two transformations: FFN(x, W1, W2, b1, b2) = max(0, xW1 + b1)W2 + b2. The calculation method of the Gate unit in the model is: According to the above formula, the activation function SwiGLU can be obtained as: where Swish β (x) = xσ(βx), β is a specified constant, W1 and W2 represent matrices, b1 and b2 represent bias terms, and x represents the input.
[0151] In summary, as Figure 12 shown in the figure, by combining the question processing module, the dialogue processing module, and the large prediction model, on the one hand, during the communication between doctors and patients, the dialogue can be real-time voice-to-text converted, and the colloquial and elided content in the text can be corrected into text suitable for the understanding of the large language model. On the other hand, relying on the logical reasoning ability of the large language model and combining the question processing module and the dialogue processing module, it is possible to find answers to pre-set questions, form reports for certain answers, and give real-time reminders for uncertain or unmentioned questions.
[0152] Based on the above embodiments of the present application, the present application can solve problems such as low efficiency of doctor-patient communication, rigid Q&A forms, and inability to provide real-time feedback during the current medical follow-up process. By carrying out reasonable follow-up question prompt engineering (question processing), speech-to-text conversion technology, dialogue text chunking technology, and large language model technology, and integrating the large language model into the medical follow-up process, the pain points in the above traditional follow-up process can be well solved.
[0153] It should be pointed out that the embodiments of the present application can learn from or refer to each other. For example, for the same or similar steps, method embodiments, system embodiments, and device embodiments can all refer to each other without limitation.
[0154] In the embodiments of the present application, the report generation device may be divided into functional modules or functional units according to the above method examples. For example, each functional module or functional unit may be corresponding to each function, or two or more functions may be integrated into one processing module. The above integrated module may be implemented in the form of hardware, or may be implemented in the form of a software functional module or functional unit. Among them, the division of modules or units in the embodiments of the present application is illustrative, and is only a logical function division. In actual implementation, there may be other division methods.
[0155] As Figure 13 shown, it is a schematic structural diagram of a report generation device provided by an embodiment of the present application. The device is applied to an electronic device, and the electronic device includes: a communication interface and a processor, and the processor is coupled to the communication interface; the device includes: a communication unit 1301 and a processing unit 1302.
[0156] The communication unit 1301 is configured to obtain the consultation text content between the doctor and the patient.
[0157] The processing unit 1302 is configured to determine the reply content of the consultation question from the consultation text content based on the pre-determined consultation question; the consultation question is obtained by correcting the initial question based on the question correction rule, and the initial question is a question related to the abnormal physical condition of the patient; based on the consultation question and the reply content of the consultation question, a consultation report is generated.
[0158] In a possible implementation manner, the question correction rule includes at least one of the following: expanding the proper nouns in the initial question; marking the core words in the initial question; expanding the initial question and determining the trigger condition of the extended question obtained by the expansion.
[0159] In a possible implementation manner, the communication unit 1301 is further configured to obtain the initial question determined for the abnormal physical condition of the patient.
[0160] The processing unit 1302 is further configured to determine at least one proper noun from the initial question, and determine the expansion words of the at least one proper noun. The expansion words are any one of the following: synonyms of the proper noun, alternative names of the proper noun, abbreviations of the proper noun, foreign names of the proper noun; based on the expansion words and the initial question, the consultation question is determined.
[0161] In a possible implementation manner, the processing unit 1302 is further configured to determine the extended question of the initial question and the trigger condition of the extended question. The correlation degree between the extended question and the initial question is greater than the preset correlation degree, and the trigger condition is that the reply content of the initial question matches the preset content; based on the extended question and the initial question, the consultation question is determined.
[0162] In a possible implementation, the processing unit 1302 is further configured to determine abnormal questions from the consultation questions, where the abnormal questions include at least one of the following: questions for which no reply content is determined from the consultation text content, questions for which the determined reply content is abnormal from the consultation text content, questions not mentioned in the consultation text content, and abnormal reply content includes at least one of the following: the reply content does not match the question, and the parameter value in the reply content exceeds the parameter range matching the question; mark and display the abnormal questions in the consultation report.
[0163] The processing unit 1302 is specifically configured to determine the output format of the reply content, where the output format includes at least one of the following: text format, parameter unit, parameter order; generate a consultation report based on the consultation questions and the reply content of the consultation questions through the output format.
[0164] In a possible implementation, the processing unit 1302 is further configured to identify the consultation conversation between the doctor and the patient to obtain the initial text content; preprocess the initial text content to obtain the consultation text content, where the preprocessing includes at least one of the following: adjusting the correspondence between each text content and the doctor / patient, text correction processing, and dialogue segmentation processing, and the text correction processing includes at least one of the following: text deletion, text error correction, text replacement, and text completion.
[0165] In a possible implementation, the preprocessing includes adjusting the correspondence between each text content and the doctor / patient, and the initial text content includes multiple texts; the processing unit 1302 is specifically configured to determine the voice characteristics of the doctor and the voice characteristics of the patient, where the voice characteristics include at least one of the following: tone, speech rate, intonation, timbre, and speech clarity; determine the speaking object corresponding to each text in the multiple texts based on the voice characteristics of the doctor and the voice characteristics of the patient to obtain the consultation text content, and the speaking object is the doctor or the patient.
[0166] In a possible implementation, the preprocessing includes dialogue segmentation processing, and the initial text content includes multiple texts; the processing unit 1302 is further configured to determine the encoding information of each text in the multiple texts and the semantics of each text; divide the multiple texts into at least two text segments based on the encoding information of each text and the semantics of each text to obtain the consultation text content.
[0167] When implemented by hardware, the communication unit 1301 in the embodiments of the present application may be integrated on a communication interface, and the processing unit 1302 may be integrated on a processor. The specific implementation is as Figure 14 shown.
[0168] Figure 14Shows another possible structural schematic diagram of the report generation device involved in the above embodiments. The report generation device includes: a processor 1402 and a communication interface 1403. The processor 1402 is used to control and manage the operations of the device. For example, it executes the steps performed by the above-mentioned processing unit 1302, and / or is used to execute other processes of the technologies described herein. The communication interface 1403 is used to support the communication of the device with other network entities. For example, it executes the steps performed by the above-mentioned communication unit 1301. The device may further include a memory 1401 and a bus 1404. The memory 1401 is used to store the program code and data of the device.
[0169] Among them, the memory 1401 may be the memory in the device, etc. The memory may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk or solid-state drive; the memory may further include a combination of the above types of memory.
[0170] The above-mentioned processor 1402 may be to implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present application. The processor may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0171] The bus 1404 may be an extended industry standard architecture (EISA) bus, etc. The bus 1404 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 14 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0172] Figure 14 The device in the figure may also be a chip. The chip includes one or more than two (including two) processors 1402 and a communication interface 1403.
[0173] Optionally, the chip further includes a memory 1405. The memory 1405 may include read-only memory and random access memory, and provides operation instructions and data to the processor 1402. A part of the memory 1405 may further include non-volatile random access memory (NVRAM).
[0174] In some embodiments, the memory 1405 stores elements, execution modules, or data structures, or subsets thereof, or extended sets thereof.
[0175] In the embodiments of the present application, by invoking the operation instructions stored in the memory 1405 (which may be stored in the operating system), corresponding operations are executed.
[0176] Some embodiments of the present disclosure provide a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium) storing computer program instructions, which, when running on a computer (e.g., a receiving node), cause the computer to execute the synchronization method in any one of the above embodiments.
[0177] Exemplarily, the above computer-readable storage medium may include, but is not limited to: magnetic storage devices (e.g., hard disks, floppy disks, or magnetic tapes, etc.), optical disks (e.g., compact disks (CDs), digital versatile disks (DVDs), etc.), smart cards, and flash memory devices (e.g., erasable programmable read-only memories (EPROMs), cards, sticks, or key drives, etc.). The various computer-readable storage media described in the present disclosure may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.
[0178] Some embodiments of the present disclosure also provide a computer program product. For example, the computer program product is stored on a non-transitory computer-readable storage medium. The computer program product includes computer program instructions, which, when executed on a computer (e.g., a receiving node), cause the computer to execute the synchronization method in the above embodiments.
[0179] Some embodiments of the present disclosure also provide a computer program. When the computer program is executed on a computer (e.g., a receiving node), the computer program causes the computer to execute the synchronization method in the above embodiments.
[0180] The beneficial effects of the above computer-readable storage medium, computer program product, and computer program are the same as those of the synchronization method in some of the above embodiments, and will not be elaborated here.
[0181] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0182] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0183] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0184] The above is only the specific implementation manner of this disclosure, but the protection scope of this disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by this disclosure, thinking of changes or substitutions, should be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be subject to the protection scope of the claims.
Claims
1. A report generation method, characterized in that: The method comprises: Get the text content of the consultation between the doctor and the patient; Based on a predetermined medical question, the answer content of the medical question is determined from the medical question text content; the medical question is obtained by correcting the initial question based on the question correction rule, and the initial question is a question related to the abnormal physical condition of the patient; A medical consultation report is generated based on the medical consultation questions and the answers to the medical consultation questions.
2. The report generation method according to claim 1, characterized in that: The problem correction rule includes at least one of the following: Expanding the proper nouns in the initial question; Marking the core words in the initial question; The initial question is expanded, and a trigger condition of the expanded extended question is determined.
3. The report generation method according to claim 2, characterized in that: The report generation method further comprises: obtaining the initial problem determined for the abnormal physical condition of the patient; Determine at least one proper noun from the initial question, and determine an extended word of the at least one proper noun, wherein the extended word is any one of the following: a synonym of the proper noun, an alias of the proper noun, an abbreviation of the proper noun, or a foreign name of the proper noun; The diagnostic question is determined based on the expanded word and the initial question.
4. The report generation method according to claim 2, characterized in that: The report generation method further comprises: Determine the extended question of the initial question and the triggering condition of the extended question, the correlation between the extended question and the initial question is greater than a preset correlation, and the triggering condition is that the answer content of the initial question matches the preset content; The diagnostic question is determined based on the extended question and the initial question.
5. The report generation method according to claim 1, characterized in that: The report generation method further comprises: Determine abnormal questions from the medical questions, wherein the abnormal questions include at least one of the following: questions whose answers are not determined from the medical question text content, questions whose answers are abnormal as determined from the medical question text content, and questions not mentioned in the medical question text content, wherein the abnormal answers include at least one of the following: the answers do not match the questions, and the parameter values in the answers exceed the parameter range matching the questions; The abnormal problem is marked and displayed in the medical consultation report.
6. The report generation method according to any one of claims 1 to 5, characterized in that: The generating of a medical consultation report based on the medical consultation questions and the answers to the medical consultation questions includes: Determine an output format of the reply content, the output format including at least one of the following: text format, parameter unit, and parameter order; The medical consultation report is generated in the output format based on the medical consultation questions and the answers to the medical consultation questions.
7. The report generation method according to any one of claims 1 to 5, characterized in that: The report generation method further comprises: Identify the medical consultation dialogue between the doctor and the patient to obtain initial text content; The initial text content is preprocessed to obtain the consultation text content, wherein the preprocessing includes at least one of the following: adjusting the correspondence between each text content and the doctor / patient, text correction processing, and dialogue segmentation processing; the text correction processing includes at least one of the following: text elimination, text error correction, text replacement, and text completion.
8. The report generation method according to claim 7, characterized in that: The preprocessing includes adjusting the correspondence between each text content and the doctor / patient, and the initial text content includes multiple texts; The preprocessing of the initial text content to obtain the medical consultation text content includes: Determine the voice characteristics of the doctor and the patient, wherein the voice characteristics include at least one of the following: tone, speech speed, intonation, timbre, and speech clarity; Based on the voice features of the doctor and the voice features of the patient, the speaking object corresponding to each of the multiple texts is determined to obtain the consultation text content, and the speaking object is the doctor or the patient.
9. The report generation method according to claim 7, characterized in that: The pre-processing includes the dialogue segmentation processing, and the initial text content includes a plurality of texts; The preprocessing of the initial text content to obtain the medical consultation text content includes: Determining encoding information of each of the plurality of texts and semantics of each of the texts; Based on the encoding information of each text and the semantics of each text, the multiple texts are divided into at least two text segments to obtain the medical consultation text content.
10. An electronic device, characterized in that: The electronic device comprises a processor and a communication interface, wherein the processor is coupled to the communication interface; The communication interface is configured to: obtain the text content of the consultation between the doctor and the patient; The processor is configured to: determine the answer content of the medical question from the medical question text content based on the predetermined medical question; the medical question is obtained by correcting the initial question based on the question correction rule, and the initial question is a question related to the abnormal physical condition of the patient; The processor is further configured to generate a medical consultation report based on the medical consultation questions and the answers to the medical consultation questions.
11. A report generating device, characterized in that: include: A processor and a communication interface; the communication interface is coupled to the processor, and the processor is used to run a computer program or instruction to implement the report generation method as described in any one of claims 1-9.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions. When a computer executes the instructions, the computer executes the report generating method according to any one of claims 1 to 9.
13. A computer program product, characterized in that The computer program product comprises instructions, and when the instructions are executed on a computer, the computer performs the report generating method according to any one of claims 1 to 9.