Voice inquiry processing method and device
By separating and recognizing the voice streams of users and medical personnel during consultations, and combining this with medical insurance interface queries, the convenience of medical insurance data and cost settlement in online medical services has been solved, thus improving the user's medical experience.
Patent Information
- Application Number
- CN202511696309.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-03-03
AI Technical Summary
In online healthcare services, healthcare providers face competitive pressures, and users find it difficult to easily access medical insurance data and medical expense settlement information.
By acquiring the voice streams of consultations between users and medical personnel, performing voice separation and recognition, and combining user information and medical texts, the system calls the medical insurance interface to query medical insurance data and calculate costs, thus constructing structured consultation information.
It enables real-time querying of medical insurance data and calculation of medical expenses in online medical services, improving users' medical experience and the convenience of service provision.
Smart Images

Figure CN121601208A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of data processing technology, and in particular to a voice consultation processing method and device. Background Technology
[0002] With the continuous development and promotion of internet technology, artificial intelligence, and mobile terminals, online services based on the internet have emerged, such as online medical services. With the diversification of online medical services, more and more users are using smart devices to meet various service needs, providing users with convenient means of medical services. However, with the continuous increase in the number of medical service providers, these providers are also facing certain competition and pressure. Summary of the Invention
[0003] This specification provides one or more embodiments of a voice-based medical consultation processing method, comprising: acquiring a voice stream of a user and a medical professional collected by a service program in consultation mode. The consultation mode is activated after a consultation detection based on the user's gesture image sequence passes. The voice stream is subjected to speech separation and speech recognition to obtain user text and the medical professional's diagnosis text. A medical insurance interface is invoked to perform medical insurance data query and medical expense calculation based on the user's registration information in the service program and medical keywords in the diagnosis text to obtain medical expenses. Consultation information is constructed based on the user text, the diagnosis text, and the medical expenses, and the constructed structured consultation information is returned to the service program.
[0004] This specification provides one or more embodiments of another voice-based medical consultation processing method, including: performing a consultation detection based on a user's gesture image sequence, and activating the consultation mode of a service program after the consultation detection is successful; invoking a voice acquisition component to acquire the consultation voice stream between the user and a medical professional while the service program is in the consultation mode, and sending the consultation voice stream to the server; rendering and displaying structured consultation information returned by the server based on user text, the medical professional's diagnosis text, and medical expenses. The medical expenses are obtained by calling a medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text.
[0005] This specification provides one or more embodiments of a voice consultation processing device, comprising: a voice acquisition module configured to acquire a voice stream of a user and a medical professional collected by a service program in consultation mode. The consultation mode is activated after a consultation detection based on the user's gesture image sequence passes. A voice recognition module configured to perform voice separation and voice recognition on the consultation voice stream to obtain user text and the medical professional's treatment text. An interface calling module configured to call a medical insurance interface to perform medical insurance data query and medical expense calculation based on the user's registration information in the service program and medical keywords in the treatment text to obtain medical expenses. An information construction module configured to construct consultation information based on the user text, the treatment text, and the medical expenses, and return the constructed structured consultation information to the service program.
[0006] This specification provides one or more embodiments of another voice-based medical consultation processing device, comprising: a consultation detection module configured to perform consultation detection based on a user's gesture image sequence, and to initiate the consultation mode of a service program after the consultation detection is successful; a voice transmission module configured to call a voice acquisition component to acquire the consultation voice stream between the user and a medical professional in the consultation mode of the service program, and to send the consultation voice stream to a server; and a rendering module configured to render and display structured consultation information returned by the server based on user text, the medical professional's diagnosis text, and medical expenses. The medical expenses are obtained by calling a medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text.
[0007] This specification provides one or more embodiments of a voice consultation processing device, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to: acquire a user-medical personnel consultation voice stream collected by a service program in consultation mode. The consultation mode is activated after a consultation detection based on the user's gesture image sequence passes. The processor performs speech separation and speech recognition on the consultation voice stream to obtain user text and the medical personnel's diagnosis text. It calls a medical insurance interface to perform medical insurance data queries and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text. Based on the user text, the diagnosis text, and the medical expenses, it constructs consultation information and returns the constructed structured consultation information to the service program.
[0008] This specification provides one or more embodiments of another voice-based medical consultation processing device, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to: perform a medical consultation detection based on a user's gesture image sequence, and, upon successful detection, initiate a medical consultation mode of a service program; invoke a voice acquisition component to acquire the voice stream of the user and medical personnel during the medical consultation mode of the service program, and send the voice stream to a server; render and display structured medical consultation information returned by the server, based on user text, medical personnel's diagnosis text, and medical expenses. The medical expenses are obtained by invoking a medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text.
[0009] This specification provides one or more embodiments of a computer-readable storage medium for storing computer-executable instructions, which, when executed, perform the following steps: acquiring a user-medical personnel consultation voice stream collected by a service program in consultation mode. The consultation mode is activated after a consultation detection based on the user's gesture image sequence passes. Performing speech separation and speech recognition on the consultation voice stream to obtain user text and the medical personnel's diagnosis text. Calling a medical insurance interface to perform medical insurance data query and medical expense calculation based on the user's registration information in the service program and medical keywords in the diagnosis text to obtain medical expenses. Constructing consultation information based on the user text, the diagnosis text, and the medical expenses, and returning the constructed structured consultation information to the service program.
[0010] This specification provides one or more embodiments of another computer-readable storage medium for storing computer-executable instructions that, when executed, perform the following steps: performing a consultation detection based on a user's gesture image sequence; and, upon successful consultation detection, initiating the consultation mode of a service program. The system then invokes a voice acquisition component to acquire the consultation voice stream between the user and medical personnel while the service program is in consultation mode, and sends the consultation voice stream to the server. Finally, the system renders and displays structured consultation information returned by the server, based on user text, the medical personnel's diagnosis text, and medical expenses. The medical expenses are obtained by invoking a medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Figure 1 A schematic diagram illustrating the implementation environment of a voice consultation processing method provided in one or more embodiments of this specification; Figure 2 A flowchart of a voice consultation processing method provided in one or more embodiments of this specification; Figure 3 A schematic diagram of a consultation page provided for one or more embodiments of this specification; Figure 4 A schematic diagram of a consultation recording page provided for one or more embodiments of this specification; Figure 5 A schematic diagram illustrating a diagnostic annotation reminder on a medical consultation recording page, provided for one or more embodiments of this specification; Figure 6 A schematic diagram of a report generation reminder page provided for one or more embodiments of this specification; Figure 7 A schematic diagram of a consultation report page provided for one or more embodiments of this specification; Figure 8 This diagram illustrates a consultation exit reminder on a consultation recording page, provided for one or more embodiments of this specification. Figure 9 A flowchart illustrating a voice consultation processing method for offline consultation scenarios, provided in one or more embodiments of this specification. Figure 10 A flowchart of another voice consultation processing method provided in one or more embodiments of this specification; Figure 11 A schematic diagram of an embodiment of a voice consultation processing device provided in one or more embodiments of this specification; Figure 12 A schematic diagram of another embodiment of a voice consultation processing device provided in one or more embodiments of this specification; Figure 13 A schematic diagram of the structure of a voice consultation processing device provided in one or more embodiments of this specification; Figure 14 This is a schematic diagram of another voice consultation processing device provided in one or more embodiments of this specification. Detailed Implementation
[0012] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.
[0013] The voice consultation processing method provided in one or more embodiments of this specification is applicable to the voice consultation implementation environment. (Refer to...) Figure 1 The implementation environment includes at least: The service program includes a server-side component 101; additionally, the implementation environment may also include a service program 102. The server 101 of the service program is used to separate the voice streams of the user and medical personnel and perform voice recognition to obtain the user text and the medical personnel's diagnosis text. It calls the medical insurance interface to query medical insurance data and calculate medical expenses based on user information and diagnosis text to obtain medical expenses. It combines user text, diagnosis text and medical expenses to construct structured consultation information and return it to the service program. The server 101 can run on a server, which can be one or more servers, a server cluster composed of several servers, or a cloud server of a cloud computing platform. The service program 102 is used to send the voice stream of the user and medical personnel collected in the consultation mode to the server and receive the structured text information returned by the server. The service program 102 can run on the user terminal, which can be a mobile phone, personal computer, tablet computer, e-book reader, wearable device, device for information interaction based on AR (Augmented Reality) / VR (Virtual Reality), and laptop computer, etc. In addition, the implementation environment may also include a medical insurance platform 103, which is used to query medical insurance data and calculate medical expenses to obtain medical expenses, and to send the medical expenses back to the server through the medical insurance interface. In this implementation environment, server 101 obtains the consultation voice stream of the user and medical personnel collected by service program 102 in consultation mode, performs voice separation and voice recognition on the consultation voice stream to obtain user text and medical personnel diagnosis text, calls the medical insurance interface to query medical insurance data and calculate medical expenses based on user information and diagnosis text to obtain medical expenses, constructs structured consultation information based on user text, diagnosis text and medical expenses and returns it to service program 102. In this way, by analyzing the consultation voice stream of the user and medical personnel, diagnosis text is obtained, and then the diagnosis text is used to realize the real-time calculation of medical expenses and the real-time generation of consultation information.
[0014] One or more embodiments of a voice-based diagnostic processing method provided in this specification are as follows: Reference Figure 2 The voice consultation processing method provided in this embodiment can be applied to the server side of the service program, specifically including steps S202 to S208.
[0015] Step S202: Obtain the voice stream of the user and medical personnel during the consultation, collected when the service program is in consultation mode.
[0016] The service program described in this embodiment can be an application, subprogram (app), or web application that provides medical services; that is, the service program can be a medical service program, a medical application, a medical subprogram, or a medical web application. The user can be any user, such as a patient, a patient's family member, the patient's previous attending physician, and / or the patient's first patient. The medical personnel refer to personnel engaged in medical, preventive, and health care professional and technical work in medical institutions, such as staff, doctors, nurses, pharmacists, and / or laboratory personnel. The medical institution refers to an institution that provides medical services, such as a hospital, clinic, pharmacy, or school medical room.
[0017] The consultation voice stream refers to the voice communication between medical personnel and users during the consultation process. When the service program is in consultation mode, it means that the service program has started recording the consultation voice stream between the user and the medical personnel. The consultation voice stream can be collected using a preset sampling rate, such as 16kHz. Specifically, if the service program has not performed a consultation test or the consultation test has failed, it can be in general medical mode. When the service program is in general medical mode, users can access health records and have conversations with medical intelligent agents within the service program.
[0018] In specific implementation, the service program acquires the voice stream of the user and medical personnel during the consultation mode. Specifically, the service program can perform consultation detection based on the user's gesture image sequence. After the consultation detection is passed, the consultation mode is activated, and the voice acquisition component is called to collect the voice stream of the user and medical personnel during the consultation mode and send it to the server. Correspondingly, the server can acquire the voice stream of the user and medical personnel during the consultation mode acquired by the service program. Optionally, the consultation mode is activated after the consultation detection based on the user's gesture image sequence is passed.
[0019] In practice, users can submit access commands while inside a medical institution. Based on these commands, they can be redirected to the consultation page of the service program. Optionally, the service program can also redirect users to the consultation page while they are inside a medical institution. After redirection, the service program calls an image acquisition component to capture a sequence of the user's gesture images. The consultation scenario can be online or offline. The consultation page can be any page within the service program, such as the homepage. Figure 3 The consultation page shown includes a consultation reminder: "Wave your hand up to start recording, and wave down to automatically generate a consultation report." At this time, the image acquisition component can start capturing the user's gesture image sequence. Medical institutions refer to institutions that provide medical services, such as hospitals, clinics, pharmacies, or school clinics. The image acquisition component can be an image sensor, such as a camera, specifically a front-facing camera and / or a rear-facing camera. A gesture image sequence refers to a sequence of one or more gesture images.
[0020] After redirecting to the consultation page, the aforementioned service program calls the image acquisition component to collect the user's gesture image sequence. After passing the consultation detection based on the gesture image sequence, the service program can initiate consultation mode. In consultation mode, the service program can call the voice acquisition component to collect the consultation voice stream between the user and medical personnel and send the consultation voice stream to the server; for example... Figure 4 The shown consultation recording page displays the recording duration as "00:05" and the text message "Consultation recording in progress".
[0021] To reduce the power consumption of the image acquisition component and improve its resource utilization, the image acquisition component may optionally include an image sensor with adjusted operating parameters. These parameters are adjusted when the image sensor detects the user's hand contour. The operating parameters may include resolution, frame rate, encoding method, and / or supplementary lighting strategy. Specifically, when the service program redirects to the consultation page, the image sensor with the first operating parameter can be invoked to detect the user's hand contour. If the user's hand contour is detected, the image sensor with the second operating parameter can be invoked to acquire a sequence of gesture images. Here, the first operating parameter can be lower than the second operating parameter, thereby reducing the power consumption of the image sensor.
[0022] In specific implementation, the service program can perform a diagnostic test based on the user's gesture image sequence. After the diagnostic test is passed, the service program's diagnostic mode is activated. That is, the diagnostic mode can be activated after the diagnostic test based on the user's gesture image sequence is passed, thereby flexibly entering the diagnostic mode through gestures and improving the convenience of activating the diagnostic mode. In an optional implementation method provided in this embodiment, the diagnostic test is passed in the following way: Gesture recognition results are obtained by performing gesture recognition based on gesture image sequences; If the gesture recognition result indicates that the user's gesture is a consultation gesture, the consultation test is considered successful.
[0023] The user's gesture image sequence can be obtained by capturing gesture images through the front-facing camera and / or rear-facing camera of the user terminal. Specifically, it can be obtained by capturing gesture images through the front-facing camera and / or rear-facing camera based on a preset frame rate, such as capturing gesture images through the front-facing camera at 30fps to obtain the gesture image sequence.
[0024] Specifically, the user's gesture image sequence can be obtained by performing illumination compensation and / or noise reduction on each user gesture image in the initial gesture image sequence acquired by the image acquisition component.
[0025] Based on this, in order to improve the accuracy of gesture recognition, this embodiment provides an optional implementation method in which, during the process of obtaining gesture recognition results based on gesture image sequences, gesture recognition is performed based on the coordinates of key hand points and / or the duration of gestures in each user gesture image in the gesture image sequence. Specifically, the following operations can be performed: Hand keypoints are detected in each user gesture image in the gesture image sequence to obtain the coordinates of the hand keypoints; Gesture recognition results are obtained by using the coordinates of key hand points and the duration of the gesture.
[0026] Among them, the hand key point coordinates can be 21 hand key point coordinates, or palm key point coordinates and / or wrist key point coordinates; gesture duration refers to the duration for which the user makes a gesture.
[0027] Specifically, in the process of detecting hand keypoints and obtaining hand keypoint coordinates for each user gesture image in the gesture image sequence, each user gesture image in the gesture image sequence can be input into a keypoint detection model to detect hand keypoints and obtain hand keypoint coordinates; the keypoint detection model here can be a hand keypoint detection model, such as the MediaPipe Hands model (hand keypoint detection model).
[0028] To further refine gesture recognition and obtain more accurate results, thus avoiding erroneous activation of the consultation mode due to incorrect recognition, this embodiment provides an optional implementation method in which the following operations are performed during the process of obtaining gesture recognition results based on the coordinates of key hand points and the duration of the gesture: Calculate the direction vectors of the palm and wrist based on the coordinates of key points of the hand, and calculate the negative projection offset of the direction vectors on the image coordinate axes. If the negative projection offset is greater than the preset offset and the gesture duration is within the preset duration range, the gesture recognition result is determined to be a user's gesture for medical consultation. Among them, the image coordinate axis can be the Y-axis of the image; the negative projection offset refers to the projection offset of the direction vector in the negative direction of the image coordinate axis, specifically the projection length of the direction vector in the negative direction of the Y-axis of the image; the preset offset refers to the preset negative projection offset, for example, the preset offset is 0.2; the preset duration interval refers to the preset duration interval, for example, the preset duration interval is 0.3-0.5 seconds.
[0029] Specifically, a direction vector from the wrist to the center of the palm can be constructed based on the key points of the palm and wrist. The direction vector is normalized to obtain a normalized vector. The projection offset of the normalized vector in the negative direction of the image coordinate axis is calculated. If the projection offset is greater than the preset offset and the duration of the gesture is within the preset duration range, the gesture recognition result is determined to be a consultation gesture. The consultation gesture can be a consultation initiation gesture, that is, a gesture representing the start of consultation. The consultation gesture can be a gesture with the palm facing upward.
[0030] It should be noted that the operation of obtaining the user's and medical personnel's consultation voice stream in the consultation mode of the above-mentioned service program can be replaced by obtaining the user's and medical personnel's consultation voice stream, and forming a new implementation method with other steps provided in this embodiment.
[0031] Step S204: Perform speech separation and speech recognition on the consultation speech stream to obtain the user text and the medical personnel's diagnosis text.
[0032] The above-mentioned service acquisition program collects the user's and medical personnel's consultation voice stream in consultation mode. In this step, the consultation voice stream is separated into user text and medical personnel's diagnosis and consultation by performing voice separation and voice recognition. In this way, the consultation voice stream is separated into user voice track and medical personnel voice track, and voice recognition is performed on each track.
[0033] In this embodiment, the user text refers to the text obtained by performing speech recognition on the user's speech separated from the consultation speech stream; the medical personnel's diagnosis and treatment text refers to the diagnosis and treatment text obtained by performing speech recognition on the medical personnel's diagnosis and treatment speech separated from the consultation speech stream.
[0034] In practical applications, during offline consultations, users and medical personnel often naturally alternate in their conversations, resulting in a mixed audio stream. Therefore, to more clearly identify the user's consultation needs and the medical personnel's treatment of the user, audio separation can be performed on the consultation audio stream. Specifically, in one optional implementation of this embodiment, during the audio separation process, an inverse frequency domain transform is performed on the user's audio spectrum and the medical personnel's audio spectrum to obtain the user's text and the medical personnel's treatment text. Specifically, the following operations can be performed: The frequency domain transformation of the consultation voice stream is performed to obtain the voice spectrogram, and the voice spectrogram is input into the spectrum separation model for voiceprint recognition and spectrum separation to obtain the user's voice spectrum and the medical staff's voice spectrum; Perform inverse frequency domain transformation on the user's speech spectrum and the speech spectrum to obtain the user's speech and the medical staff's diagnostic speech.
[0035] Among them, the model architecture of the spectrum separation model can be U-Net (Convolutional Neural Network) architecture; the frequency domain transformation can be Short-Time Fourier Transform (STFT), the inverse frequency domain transformation can be Inverse Short-Time Fourier Transform; and the speech spectrogram can be Mel spectrogram.
[0036] Specifically, in the process of obtaining a speech spectrogram by performing frequency domain transformation on the consultation speech stream, the consultation speech stream can be transformed from the time domain to the frequency domain to obtain a speech spectrogram, or a short-time Fourier transform can be performed on the consultation speech stream to obtain a mixed Mel spectrogram.
[0037] The mask features of the medical personnel mentioned above can be the probability of belonging to the medical personnel in the mixed Mel spectrogram, and the mask features of the users can be the probability of belonging to the users in the mixed Mel spectrogram; among them, the voiceprint features of the medical personnel can be the voiceprint vector of the medical personnel, such as the 256-dimensional voiceprint vector of the medical personnel, and the voiceprint features of the users can be the voiceprint vector of the users, such as the 256-dimensional voiceprint vector of the users.
[0038] To improve the convenience and efficiency of voiceprint recognition, the voiceprint features of medical personnel can be pre-registered. Specifically, at least one voice of a medical personnel can be collected in advance, and the voiceprint vector can be extracted from the at least one voice and stored in the encrypted database of the medical personnel's terminal device. The extraction of the voiceprint vector can be achieved by the ECAPA-TDNN (a deep neural network model applied in speaker verification tasks) model, and the duration of each voice in the at least one voice can be a preset duration.
[0039] In addition, during the above-mentioned process of speech separation of the consultation speech stream, the consultation speech stream can also be denoised to obtain a denoised consultation speech stream. Specifically, the consultation speech stream can be denoised by spectral subtraction and / or filtering methods such as Wiener filtering. Then, the denoised consultation speech stream can be separated into user speech and medical personnel's diagnosis speech by deep clustering algorithm.
[0040] In the process of voiceprint recognition and spectrum separation described above, in an optional implementation of this embodiment, the user's speech spectrum and the medical personnel's speech spectrum are calculated based on the speech spectrogram, the medical personnel's voiceprint features, and the user's voiceprint features. Specifically, the following operations can be performed: The mask features of medical personnel are calculated based on speech spectrograms and voiceprint features of medical personnel, and the mask features of users are calculated based on speech spectrograms and voiceprint features of users. The speech spectrum is calculated based on the speech spectrogram and the masking characteristics of medical personnel, and the user's speech spectrum is calculated based on the speech spectrogram and the user's masking characteristics.
[0041] Among them, the user's voiceprint features can be voiceprint features extracted from the user's voice during the user's first consultation, or voiceprint features obtained by extracting the user's voice during the initial text consultation.
[0042] Specifically, mask features of medical personnel can be calculated based on the hybrid Mel spectrogram and the voiceprint features of medical personnel; user mask features can be calculated based on the hybrid Mel spectrogram and the user's voiceprint features; the Mel spectrogram of medical personnel can be calculated based on the mask features of medical personnel and the hybrid Mel spectrogram; and the Mel spectrogram of user can be calculated based on the mask features of user and the hybrid Mel spectrogram. In the process of performing inverse frequency domain transform on the user's speech spectrum and the speech spectrum to obtain the user's speech and the medical personnel's diagnostic speech, inverse short-time Fourier transforms can be performed on the medical personnel's Mel spectrogram and the user's Mel spectrogram to obtain the medical personnel's diagnostic speech and the user's speech, or to obtain the medical personnel's speech track and the user's speech track. Subsequently, user tags can be applied to the user's speech and / or medical personnel tags can be applied to the medical personnel's diagnostic speech, and the tagged user speech and medical personnel's diagnostic speech can be allocated to independent speech channels for storage.
[0043] The above-mentioned process of separating the consultation voice stream into user voice and medical personnel's diagnostic voice can be used to obtain user text and medical personnel's diagnostic text. Specifically, user voice and / or medical personnel's diagnostic voice can be input into a medical speech recognition engine for speech recognition to obtain user text and / or medical personnel's diagnostic text. For example, the speech recognition engine could be Whisper (a general speech recognition model). User text and / or diagnostic text can carry timestamps. For example, if the user text is "I've been feeling unwell in my stomach and have indigestion lately," the medical personnel's diagnostic text could be "Amlodipine 5mg once daily, fasting CT scan required."
[0044] In practical applications, during consultations between users and medical personnel, users may need to highlight certain medical orders given by the medical personnel. To address this, and to meet the diverse needs of users for marking medical orders, while also improving the convenience of this method, this embodiment provides an optional implementation that, based on the time interval of the timestamp carried in the diagnosis and treatment annotation request sent by the service program, marks key text blocks in the diagnosis and treatment text and marks key segments of the medical personnel's diagnosis and treatment voice stream in the subsequent consultation voice stream. Specifically, the following operations can be performed: Read the timestamp carried in the diagnostic annotation request sent by the service program; Based on the time interval of the timestamp, key text blocks are marked in the diagnosis and treatment text, and key segments are marked in the diagnosis and treatment voice segments of medical personnel in the subsequent consultation voice stream.
[0045] Among them, a diagnostic annotation request can be a request to annotate diagnostic texts.
[0046] For example Figure 5The consultation recording page shown features a treatment annotation reminder: "Waving your thumb up can record key medical orders." Users can annotate treatments based on these reminders.
[0047] Based on this, to improve the convenience of diagnostic annotation and enhance user engagement, in one optional implementation of this embodiment, the diagnostic annotation request is sent after the service program has passed the diagnostic annotation detection of the collected user gesture image sequence. The passing of the diagnostic annotation detection can be determined based on the coordinates of the key finger points in each user gesture image within the sequence. Specifically, the passing of the diagnostic annotation detection can be achieved in the following way: Finger key point recognition is performed on each user gesture image in the user gesture image sequence to obtain the coordinates of the finger key points; The direction of finger movement and the duration of finger movement are obtained by recognizing the coordinates of key points of the finger and calculating the movement duration. If the direction of finger movement and the duration of finger movement trigger the diagnostic annotation conditions, the diagnostic annotation detection is confirmed to be successful.
[0048] Among them, the coordinates of the key points of the fingers can be the coordinates of the thumb tip; the diagnostic annotation conditions can include the finger movement direction being upward and / or the finger movement duration being within a preset duration range.
[0049] It should be noted that the above-mentioned operation of performing speech separation and speech recognition on the consultation voice stream to obtain user text and medical personnel's diagnosis text can be replaced by performing speech separation and speech recognition on the consultation voice stream to obtain user text and / or medical personnel's diagnosis text, or it can be replaced by performing speech separation and / or speech recognition on the consultation voice stream to obtain user text and / or medical personnel's diagnosis text, and together with other processing steps provided in this embodiment, it forms a new implementation method.
[0050] In practical applications, medical personnel's medical records may or may not contain medical keywords. In the case where medical keywords are absent, to improve the success rate of constructing structured consultation information, a medical insurance engine can be invoked based on the medical personnel's input of medical insurance keywords to calculate medical expenses. Based on the user text, medical records, and medical expenses, a structured consultation report is then constructed. Specifically, in one optional implementation of this embodiment, after performing speech separation and speech recognition on the consultation voice stream to obtain the user text and the medical personnel's medical records, the following operations are also performed: If the medical keyword recognition result obtained from the diagnosis text is empty, the medical insurance engine is called to calculate the cost based on the medical insurance code associated with the medical keywords entered by the medical personnel, and the medical insurance payment cost and the user payment cost are obtained. A structured consultation report is constructed based on user text, medical treatment text, medical insurance payment fees, and user payment fees.
[0051] Among them, medical keywords refer to medical-related keywords in the diagnosis and treatment text. Medical keywords can be medical entities, such as disease keywords, drug keywords, and / or examination item keywords; the medical insurance engine can be the medical insurance interface of the medical insurance platform; the medical keywords entered by medical personnel can be the medical keywords selected by medical personnel from the candidate medical keywords; medical personnel can enter medical keywords through the medical personnel service program and send them to the server through the medical personnel service program, or they can enter medical keywords through the medical institution's medical program and send the medical keywords to the medical institution's system through the medical institution's system and then send the medical keywords to the server through the medical institution's system.
[0052] Specifically, in the process of calculating costs based on the medical insurance codes associated with medical keywords entered by medical personnel, and obtaining the medical insurance payment and user payment costs, the corresponding reimbursement rules can be queried based on the medical insurance codes, and the costs can be calculated based on the reimbursement rules to obtain the medical insurance payment and user payment costs. The reimbursement rules here may include whether the diseases, drugs and / or examination items corresponding to the medical keywords are reimbursed by medical insurance and / or the medical insurance reimbursement ratio. The reimbursement rules can specifically be medical insurance reimbursement rules.
[0053] In practical applications, medical keywords in medical personnel's diagnostic texts may be relatively colloquial, such as the drug "Norvasc" in the diagnostic text. To improve the success rate of subsequent cost calculations, these medical keywords can be standardized to obtain standard medical keywords. By standardizing the medical keywords to the standard medical keywords in the medical insurance catalog, subsequent cost calculations become easier. The standardized medical keywords may be single or multiple. To provide corresponding cost calculation methods for different standard medical keywords and improve the flexibility of cost calculation, in one optional implementation of this embodiment, after performing speech separation and speech recognition on the consultation voice stream to obtain the user text and the medical personnel's diagnostic text, the following operations are also performed: Search for standard medical keywords that match medical keywords in medical texts within the medical knowledge graph, and determine whether there are multiple standard medical keywords. If not, proceed to step S206 below, call the medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis and treatment text to obtain medical expenses.
[0054] Among them, the medical knowledge graph may include a knowledge graph in the medical field constructed based on medical keywords and standard medical keywords that match the medical insurance catalog.
[0055] Specifically, after obtaining the medical personnel's diagnosis text through speech recognition, medical entities can be obtained by medical entity recognition in the medical personnel's diagnosis text. The medical entities can then be standardized to obtain standard medical entities. Specifically, the medical personnel's diagnosis text can be input into the entity recognition model to obtain medical entities. The standard medical entities that match the medical entities can then be queried in the medical knowledge graph. In this way, the medical entities spoken by the medical personnel can be converted into standard medical entities in the medical insurance catalog, thereby improving the accuracy of subsequent cost calculation. Among them, medical entities can be medical keywords, standard medical entities can be standard medical keywords, and entity recognition models can be keyword recognition models; entity recognition models can be pre-trained language models, such as BioBERT (Bidirectional Encoder Representations from Transformers for Biomedical Text, a deep learning language model based on the BERT architecture in the biomedical field).
[0056] For example, a medical professional's diagnosis text is "Neurovir 5mg once daily, examination item: fasting CT". The medical entities identified from the diagnosis text are the drug name "Neurovir" and the examination item "fasting CT". The standard medical entity that matches the medical entity "Neurovir" in the medical knowledge graph is "amlodipine besylate tablets", and the standard medical entity that matches the medical entity "fasting CT" is "abdominal CT plain scan".
[0057] It should be noted that the above-mentioned operation of performing speech separation and speech recognition on the consultation speech stream to obtain user text and medical personnel's diagnosis text can be replaced by performing speech recognition on the consultation speech stream to obtain user text and / or medical personnel's diagnosis text, and forming a new implementation method with other processing steps provided in this embodiment.
[0058] Step S206: Call the medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis and treatment text.
[0059] The above-mentioned process of separating and recognizing the voice stream during a consultation to obtain user text and medical personnel's diagnosis text is followed by a step where the medical insurance interface is called to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text. This enables real-time linkage with the medical insurance platform, providing users with medical insurance settlement solutions, helping them to understand medical insurance reimbursement information in a timely manner, and improving their medical experience.
[0060] The medical insurance interface mentioned in this embodiment refers to the interface provided by the medical insurance platform. The medical insurance interface may include the settlement interface and / or payment interface of the medical insurance platform. In addition, the medical insurance interface may also include other types of interfaces of the medical insurance platform, such as the medical insurance consultation interface of the medical insurance platform. The user's registration information in the service program refers to the user's registration information in the service program. The registration information may include user attribute information, such as the user's name, whether they are male or female, and / or the user's occupation. The registration information may be user information that represents the uniqueness of the user.
[0061] The medical keywords in the diagnosis and treatment text refer to the key or important words in the text, such as disease keywords, drug keywords, and / or examination item keywords. Medical keywords can be medical entities in the diagnosis and treatment text. Medical expenses refer to the cost information of medical payments made by the user. Medical expenses can include medical insurance payments and / or user payments. Specifically, medical expenses can be the medical insurance payments and / or user payments under each medical keyword.
[0062] In practice, the medical insurance interface can be called to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and the standard medical keywords in the medical text. Specifically, in the process of querying medical insurance data and calculating medical expenses based on the user's registration information in the service program and the standard medical keywords in the medical text, the user's identity as a medical insurance participant can be verified based on the registration information. After the identity verification is passed, the reimbursement parameters can be queried based on the standard medical keywords to obtain the reimbursement parameters, and the medical expenses can be calculated based on the reimbursement parameters. Among them, the reimbursement parameter can be the reimbursement ratio; in the process of obtaining medical expenses by querying medical insurance data and calculating medical expenses based on the user's registration information in the service program and the standard medical keywords in the medical text, the user's insurance type can also be queried based on the registration information, the reimbursement ratio corresponding to the standard medical keywords can be queried based on the insurance type, and the medical insurance payment and / or user payment can be calculated as medical expenses based on the reimbursement ratio.
[0063] In practical applications, there may be more than one standard medical keyword matching. In this case, in order to still provide users with an accurate medical insurance settlement plan, this embodiment provides an optional implementation method. In the process of querying medical insurance data and calculating medical expenses based on medical keywords in the user's registration information and medical records in the service program, the reimbursement parameters are queried based on multiple standard medical keywords. The medical expenses under the corresponding standard medical keywords are then calculated based on the multiple reimbursement parameters obtained from the query. Specifically, the following operations can be performed: Verify the user's identity as a participant in medical insurance based on the registration information; After identity verification is passed, multiple reimbursement parameters are obtained by querying multiple standard medical keywords based on medical keywords, and the medical expenses under the corresponding standard medical keywords are calculated based on each reimbursement parameter.
[0064] Among them, verifying a user's identity as a participant in medical insurance can be either verifying whether the user is a participant in medical insurance or verifying whether the user has participated in medical insurance.
[0065] Specifically, reimbursement parameters for multiple standard medical keywords can be obtained by querying reimbursement parameters based on multiple standard medical keywords and / or the user's insurance type. The medical insurance payment and / or user payment under the corresponding standard medical keyword can be calculated based on each reimbursement parameter.
[0066] For example, after identity verification is passed, the medical keyword "vitamin B" has multiple standard medical keywords such as "compound vitamin B tablets", "vitamin B6 tablets", and "vitamin B6 injection". Based on these multiple standard medical keywords, the reimbursement ratio can be queried to obtain the reimbursement ratio corresponding to each of the standard medical keywords "compound vitamin B tablets", "vitamin B6 tablets", and "vitamin B6 injection". Based on each reimbursement ratio, the medical expenses corresponding to "compound vitamin B tablets", "vitamin B6 tablets", and "vitamin B6 injection" can be calculated.
[0067] In addition, users can be queried based on their registration information to determine their insurance type. Based on the insurance type and multiple standard medical keywords, reimbursement parameters can be queried to obtain multiple reimbursement parameters. Medical expenses under the corresponding standard medical keywords can then be calculated based on each reimbursement parameter.
[0068] Based on this, in practical application scenarios, some standard medical keywords may correspond to only one type of drug, while others may correspond to multiple types of drugs, such as drugs produced by multiple pharmaceutical companies. To address this, and to make cost calculation more accurate, this embodiment provides an optional implementation method in which the following operation is performed during the process of calculating medical costs under the corresponding standard medical keywords based on various reimbursement parameters: Based on the drug costs of each drug under each standard medical keyword and the reimbursement ratio corresponding to each standard medical keyword, calculate the medical insurance payment and user payment for each drug.
[0069] Furthermore, since the same examination items may be performed in fast and slow versions, resulting in different medical costs, the above operation of calculating the medical insurance payment and user payment for each drug based on the drug cost and reimbursement ratio of each standard medical keyword can be replaced by calculating the user's medical insurance payment and / or user payment for each medical task based on the task cost and reimbursement ratio of each medical task under each standard medical keyword; whereby medical tasks may include drugs, examination items, and / or surgical items.
[0070] As described above, the medical knowledge graph queries for standard medical keywords that match medical keywords in the diagnosis and treatment text. These standard medical keywords can be single or multiple. In one optional implementation of this embodiment, if there are multiple standard medical keywords, the medical insurance platform's payment interface is called to query medical insurance payment information based on registration information. A structured consultation report is then constructed based on the medical insurance payment information, user text, and / or diagnosis and treatment text. Specifically, when there are multiple standard medical keywords, the following operations can be performed: The payment interface of the medical insurance platform is called to query medical insurance payment information based on the registration information and obtain medical insurance payment information; The system retrieves medical insurance payment fees and user payment fees from medical insurance payment information, and constructs a structured consultation report based on user text, diagnosis and treatment text, medical insurance payment fees, and user payment fees.
[0071] Among them, medical insurance payment information refers to the payment information of users who pay medical expenses through medical insurance.
[0072] Specifically, the payment interface of the medical insurance platform can be called to query the medical insurance payment information of users who have paid medical expenses through medical insurance within a preset time period based on the registration information. The medical insurance payment fee and / or user payment fee for each medical task can be read from the medical insurance payment information. A structured consultation report can be constructed based on the user text, diagnosis and treatment text, and medical insurance payment fee and / or user payment fee. The process of constructing a structured consultation report based on the user text, diagnosis and treatment text, and medical insurance payment fee and / or user payment fee will be explained in detail below and will not be repeated here.
[0073] It should be noted that the above-mentioned operation of calling the medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text can be replaced by calling the medical insurance interface to query medical insurance data based on the user's registration information in the service program, or it can be replaced by calling the medical insurance interface to query medical insurance data and / or calculate medical expenses based on the user's registration information in the service program and / or medical keywords in the diagnosis text, or it can be replaced by predicting medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text. Here, the prediction of medical expenses includes the prediction of medical insurance payment expenses and / or user payment expenses, and together with other processing steps provided in this embodiment, they form a new implementation method.
[0074] Step S208: Construct consultation information based on user text, diagnosis text, and medical expenses, and return the constructed structured consultation information to the service program.
[0075] The aforementioned call to the medical insurance interface is based on the user's registration information in the service program and medical keywords in the diagnosis text to query medical insurance data and calculate medical expenses. In this step, consultation information is constructed based on user text, diagnosis text, and medical expenses, and the constructed structured consultation information is returned to the service program. The structured consultation information mentioned in this embodiment refers to consultation information displayed in a structured form, such as a structured consultation report.
[0076] In practice, user text, medical records, and medical expenses can be filled into the corresponding consultation fields in the electronic consultation template to obtain structured consultation information, which is then returned to the service program. The service program can then render and display the structured consultation information returned by the server, based on user text, medical records, and medical expenses. Specifically, the service program can call the template app to render and display the structured consultation information based on user text, medical records, and medical expenses; for example... Figure 6 The report generation reminder page shown is used to notify users that a structured consultation report is being generated; for example... Figure 7The consultation report page shown includes fields such as the patient's description of their condition, the doctor's preliminary diagnosis, examination items, prescription, and payment items. Based on the user's text, the field content for the "patient's description of their condition" is generated as "nasal congestion accompanied by allergy symptoms, short duration, and recent exposure to a new environment." Based on the user's text and the medical record, the field content for the "doctor's preliminary diagnosis" is generated as "nasal congestion accompanied by allergy symptoms, short duration, and recent exposure to a new environment, possibly caused by allergic rhinitis; further diagnosis requires allergen testing to identify the specific allergen." The fields corresponding to the "Examination Items" field, such as "nasal endoscopy, sinus CT / nasal CT no abnormalities," and the fields corresponding to the "Prescription" field, such as "xx1, xx2," are used to generate the fields corresponding to the "Paid Items" field based on the medical expenses, such as "actual payment for xx1 (user payment is ¥xx3) and medical insurance payment (medical insurance payment is ¥xx4)." These fields are then filled into the electronic consultation template to obtain a structured consultation report, which is sent to the service program. The service program then calls the template mini-program to generate the consultation report page based on the structured consultation report.
[0077] Based on the above calculation of the medical insurance payment and user payment costs for each drug according to the drug cost and reimbursement ratio corresponding to each standard medical keyword, since each standard medical keyword corresponds to each drug, in order to display more specific and effective drug medical costs to users and avoid displaying invalid medical costs, in the first optional implementation of this embodiment, during the process of constructing consultation information based on user text, diagnosis text, and medical costs, the consultation report is constructed based on the medical insurance payment and / or user payment costs of the target drugs selected by medical personnel in each drug category, along with the user text and / or diagnosis text. Specifically, the following operations can be performed: The medical insurance reimbursement and user payment fees for each drug will be distributed to medical personnel; The consultation report is constructed based on the medical insurance payment and user payment costs of the target drugs selected by medical personnel in various drug categories, as well as user text and diagnosis and treatment text.
[0078] Furthermore, as mentioned above, based on the task cost of each medical task under each standard medical keyword and the corresponding reimbursement ratio, the medical insurance payment and / or user payment for each medical task are calculated. On this basis, during the process of constructing consultation information based on user text, diagnosis and treatment text, and medical expenses, the medical insurance payment and / or user payment for each medical task can be distributed to medical personnel. Consultation reports are constructed based on the medical insurance payment and / or user payment for the target medical task selected by the medical personnel in each medical task, along with the user text and / or diagnosis and treatment text.
[0079] In addition, as mentioned above, key text blocks can be annotated in the medical records, and key segments can be annotated in the medical personnel's medical records in the subsequent consultation audio stream. Based on this, in the second optional implementation provided in this embodiment, the following operations are performed during the process of constructing consultation information based on user text, medical records, and medical expenses: The user text, the diagnosis text annotated with key text blocks, the diagnosis text fragments of the diagnosis audio fragments annotated with key segments, and the medical expenses are filled into the corresponding diagnosis fields in the electronic consultation template to obtain a structured consultation report.
[0080] By highlighting key medical orders in the diagnosis and treatment text, the structured consultation report can be improved to enhance its intuitiveness.
[0081] In this context, the medical staff's diagnosis and treatment voice segments in the subsequent consultation voice stream can be voice segments within a preset time interval; marking key text blocks in the diagnosis and treatment text can be marking key text blocks in the diagnosis and treatment text; the time interval of the timestamp can be a preset time interval before and / or after the timestamp.
[0082] For example, if the time interval of the timestamp is x seconds before and after the timestamp, key text blocks for the first x seconds are marked in the diagnosis text, and key segments of the medical staff's diagnosis voice segments in the next consultation voice stream are marked according to the last x seconds after the timestamp.
[0083] In practical implementation, the medical insurance platform and medical institutions can settle expenses according to disease type or illness, that is, a fixed fee can be set for a disease type or illness, which improves the convenience of expense settlement and helps optimize medical resources; in an optional implementation method provided in this embodiment, the following operations are also performed: The registration information and standard medical keywords are sent to the medical insurance platform to match the registration information and standard medical keywords with the user information and medical keyword information sent by the medical institution. If the match is successful, the matching medical insurance fund amount is queried in the medical insurance settlement package based on the disease keywords in the registration information and standard medical keywords, and the medical insurance fund is paid to the medical institution according to the medical insurance fund amount.
[0084] Among them, the medical insurance settlement package refers to the strategy for medical insurance platforms to settle expenses with medical institutions. Specifically, it can be a expense settlement strategy based on disease type or disease.
[0085] It should be added that after obtaining the user's text, the user's text can also be input into the medical entity recognition model to obtain medical entities. Based on the medical entities, a consultation knowledge graph of the user can be constructed in real time. Then, based on the consultation knowledge graph, disease prediction can be performed to obtain the user's predicted disease. Based on the user's predicted disease, treatment suggestions can be generated and sent to medical personnel. For example, if the medical entities identified from the user's text include "headache", "fever" and "throbbing pain", the predicted disease obtained from the disease prediction is "cold" and "migraine", and the treatment suggestion is "suggest screening for meningitis".
[0086] It should be noted that in this embodiment, the data transmission between the server and the service program can be encrypted. Encryption algorithms such as SM4 (SM4 Block Cipher Algorithm) can be used to encrypt the data before sending it.
[0087] This embodiment can continuously collect the voice streams of user and medical personnel during consultation mode in the service program. The service program sends the voice streams to the server in real time. The server can continuously perform speech separation and speech recognition on the voice streams to obtain user text and medical personnel's diagnosis text. It calls the medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text. When it detects that the service program has exited consultation mode, the server can construct consultation information based on user text, diagnosis text, and medical expenses, and return the constructed structured consultation information to the service program.
[0088] When the service program is in consultation mode, it can detect the user's hand contour. Upon detection, it can call the image acquisition component to capture a sequence of the user's second gesture image. Consultation exit detection is then performed based on this second gesture image sequence, and the consultation mode exits upon successful detection. For example... Figure 8 The consultation recording page shown in the image displays a consultation exit reminder: "Wave your palm down to end the consultation record." Users use gestures to end the consultation based on the consultation exit reminder. The image acquisition component is called to capture the user's second gesture image sequence. Consultation exit detection is performed based on the second gesture image sequence, and the consultation mode is exited after the detection is passed.
[0089] During the consultation exit detection process based on the second gesture image sequence, hand key point detection can be performed on each user gesture image in the second gesture image sequence to obtain the coordinates of the hand key points. Gesture recognition is then performed based on the hand key point coordinates and the duration of the gesture to obtain the gesture recognition result. Specifically, the direction vectors of the palm and wrist can be calculated based on the hand key point coordinates, and the forward projection offset of the direction vectors on the image coordinate axes can be calculated. If the forward projection offset is greater than a preset offset and the gesture duration is within a preset duration range, the gesture recognition result is determined to be a consultation exit gesture. If the user gesture is a consultation gesture, the service program can be in consultation mode. If the user gesture is a consultation exit gesture, the service program can exit consultation mode, that is, it can enter the general medical mode.
[0090] Among them, the image coordinate axis can be the Y-axis of the image; the forward projection offset refers to the projection offset of the direction vector in the positive direction of the image coordinate axis, specifically the projection length of the direction vector in the positive direction of the Y-axis of the image; the preset offset refers to the pre-set forward projection offset, for example, the preset offset is 0.2; the preset duration interval refers to the pre-set duration interval, for example, the preset duration interval is 0.3-0.5 seconds.
[0091] Specifically, a direction vector from the wrist to the center of the palm can be constructed based on the key points of the palm and wrist. The direction vector is normalized to obtain a normalized vector. The projection offset of the normalized vector in the positive direction of the image coordinate axis is calculated. If the projection offset is greater than the preset offset and the duration of the gesture is within the preset duration range, the gesture recognition result is determined to be the user's gesture as a consultation exit gesture. The consultation exit gesture can be a gesture of waving the palm downwards.
[0092] The above-described method of constructing consultation information based on user text, medical records, and medical expenses, and returning the structured consultation information obtained to the service program, can be replaced by constructing consultation information based on user text, medical records, and / or medical expenses, and returning the structured consultation information obtained to the service program. Alternatively, it can be replaced by constructing consultation information based on user text, medical records, and / or medical expenses to obtain structured consultation information, and combined with other processing steps provided in this embodiment to form a new implementation method.
[0093] It should be noted that the user data obtained in this manual, such as consultation voice streams and user voiceprint features, is authorized by the user and does not involve user privacy. The data obtained from medical personnel, such as medical personnel voiceprint features and consultation voice streams, is authorized by the medical personnel and does not involve medical personnel privacy.
[0094] It should be added that each optional implementation method and each feasible execution method in steps S202 to S208 provided in this embodiment can be executed independently as needed, or they can be combined and referenced with each other. At the same time, each specific execution step in each optional implementation method or each feasible execution method can also be executed independently or combined as needed. The execution conditions of "if" or "under what circumstances" involved in each step or operation can be directly deleted, and subsequent operations can be executed. This embodiment does not make specific limitations on this.
[0095] It should also be added that, depending on the actual application scenario, step S202 and any of the subsequent steps S204 to S208 can be deleted, or any feature in any step can be deleted. For example, the service program being in consultation mode in step S202 can be deleted, and the execution order of steps S202 to S208 can also be arbitrary.
[0096] The implementation process of the above-described voice consultation processing method can be executed by the server of the service program. The implementation process of the other voice consultation processing method provided in the following method embodiment can also be executed by the service program. The two can cooperate with each other during execution. Therefore, when reading the above implementation process, you can refer to the corresponding content of the following other voice consultation processing method embodiment. Correspondingly, when reading the following other voice consultation processing method embodiment, you can also refer to the corresponding content of the above method embodiment.
[0097] The following description uses the application of the voice consultation processing method provided in this embodiment in an offline consultation scenario as an example to further illustrate the voice consultation processing method provided in this embodiment. (See also...) Figure 9 The voice consultation processing method, which is applied to offline consultation scenarios, can be applied to the server side of medical service programs and includes the following steps.
[0098] Step S902: Obtain the voice stream of the consultation between the user and the doctor in the medical institution, which is collected when the medical service program is in consultation mode.
[0099] Optionally, the consultation mode is activated after a consultation test based on the user's gesture image sequence is passed.
[0100] Step S904: Perform frequency domain transformation on the consultation voice stream to obtain a voice spectrogram, and input the voice spectrogram into the spectrum separation model for voiceprint recognition and spectrum separation to obtain the user's voice spectrum and the doctor's voice spectrum.
[0101] Step S906: Perform inverse frequency domain transformation on the user's speech spectrum and the speech spectrum to obtain the user's speech and the doctor's diagnosis speech.
[0102] Step S908: Perform speech recognition on the user's voice and the doctor's diagnosis voice respectively to obtain the user's text and the doctor's diagnosis text.
[0103] Step S910: Call the medical insurance interface to verify the user's medical insurance participation identity based on the user registration information in the medical service program. After the identity verification is passed, query the reimbursement parameters based on multiple standard medical keywords in the diagnosis and treatment text to obtain multiple reimbursement parameters. Calculate the medical insurance payment and user payment for each medical task based on the cost of each medical task under each standard medical keyword and the reimbursement ratio corresponding to each medical task.
[0104] Step S912: Distribute the medical insurance payment and user payment fees for each medical task to the doctor.
[0105] Step S914: Based on the medical insurance payment and user payment costs of the target medical task selected by the doctor in each medical task, as well as the user text and diagnosis and treatment text, a consultation report is constructed, and the structured consultation report obtained is returned to the medical service program.
[0106] Optionally, the consultation report is built after the consultation exit detection is performed based on the user's second gesture image sequence and the detection passes.
[0107] It should be noted that any one or more of steps S902 to S914 can be replaced by the corresponding technical means provided in steps S202 to S208 as needed for implementation and deployment. Any one or more of steps S902 to S914 can also be combined into a new implementation method as needed for implementation and deployment. Furthermore, any one or more of steps S902 to S914 can also be combined with one or more of the steps provided in steps S202 to S208 to form a new implementation method, or combined with one or more optional implementation methods provided in steps S202 to S208 to form a new implementation method, as needed for actual deployment. These will not be elaborated on here.
[0108] One or more embodiments of another voice consultation processing method provided in this specification are as follows: Reference Figure 10 The voice consultation processing method provided in this embodiment can be applied to service programs, specifically including steps S1002 to S1006.
[0109] Step S1002: Perform a diagnostic test based on the user's gesture image sequence, and start the diagnostic mode of the service program after the diagnostic test is passed.
[0110] The service program described in this embodiment can be an application, subprogram (app), or web application that provides medical services; that is, the service program can be a medical service program, a medical application, a medical subprogram, or a medical web application. The user can be any user, such as a patient, a patient's family member, the patient's previous attending physician, and / or the patient's first patient. The medical personnel refer to personnel engaged in medical, preventive, and health care professional and technical work in medical institutions, such as staff, doctors, nurses, pharmacists, and / or laboratory personnel. The medical institution refers to an institution that provides medical services, such as a hospital, clinic, pharmacy, or school medical room.
[0111] The consultation voice stream refers to the voice communication between medical personnel and users during the consultation process. When the service program is in consultation mode, it means that the service program has started recording the consultation voice stream between the user and the medical personnel. The consultation voice stream can be collected using a preset sampling rate, such as 16kHz. Specifically, if the service program has not performed a consultation test or the consultation test has failed, it can be in general medical mode. When the service program is in general medical mode, users can access health records and have conversations with medical intelligent agents within the service program.
[0112] In practice, the service program can perform a diagnostic test based on the user's gesture image sequence, and start the diagnostic mode after the diagnostic test is passed.
[0113] It should be noted that the specific implementation process of the diagnosis detection based on the user's gesture image sequence is similar to the implementation process of the diagnosis detection passing in the voice diagnosis processing method executed by the server mentioned above. You can refer to it for reference, and it will not be repeated here.
[0114] Step S1004: Call the voice acquisition component to acquire the voice stream of the user and medical personnel in the consultation mode of the service program, and send the voice stream to the server.
[0115] In practice, the voice acquisition component can be invoked to collect the voice stream of the user and medical personnel in the consultation mode of the service program and send it to the server.
[0116] It should be noted that the above-mentioned operation of calling the voice acquisition component to collect the consultation voice stream between the user and the medical personnel in the consultation mode of the service program and sending the consultation voice stream to the server can be replaced by collecting the consultation voice stream between the user and the medical personnel in the consultation mode of the service program and sending it to the server; or it can be replaced by calling the voice acquisition component to collect the consultation voice stream of the user in the consultation mode of the service program and sending the consultation voice stream to the server, and forming a new implementation method with other processing steps provided in this embodiment.
[0117] Step S1006: Render and display the structured consultation information returned by the server, which is based on user text, medical personnel's diagnosis text, and medical expenses.
[0118] The above-mentioned voice acquisition component collects the voice stream of the user and medical personnel in the consultation mode of the service program and sends the voice stream to the server. In this step, the structured consultation information returned by the server is rendered and displayed based on the user text, the medical personnel's diagnosis text and medical expenses.
[0119] Optionally, medical expenses are obtained after calling the medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text. Specifically, the server can perform speech separation and speech recognition on the consultation voice stream to obtain the user's text and the medical personnel's diagnosis text. It then calls the medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text to obtain the medical expenses. Based on the user's text, diagnosis text, and medical expenses, the consultation information is constructed, and the structured consultation information is returned to the service program. The specific process executed by the server here has been described in detail in the above-mentioned server-side voice consultation processing method, and will not be repeated here.
[0120] The following describes the application of the voice consultation processing method provided in this embodiment in an offline consultation scenario as an example. The voice consultation processing method provided in this embodiment can be applied to medical service procedures and specifically includes the following steps.
[0121] Step 1: Perform a consultation test based on the user's gesture image sequence. Once the consultation test is passed, start the consultation mode of the medical service program.
[0122] Step 2: Call the voice acquisition component to collect the voice stream of the consultation between the user and the doctor in the medical service program in consultation mode within the medical institution, and send the consultation voice stream to the server.
[0123] Step 3: Render and display the structured consultation report returned by the server, which is based on user text, doctor's diagnosis text, medical insurance payment for the target medical task, and user payment.
[0124] Optionally, the medical insurance payment and user payment fees for the target medical task include the medical insurance payment and user payment fees for the target medical task selected by the doctor in each medical task. The medical insurance payment and user payment fees for each medical task are obtained by calling the medical insurance interface to verify the user's medical insurance participation identity based on the user's registration information in the medical service program. After the identity verification is passed, multiple reimbursement parameters are obtained by querying multiple standard medical keywords based on multiple medical keywords in the diagnosis and treatment text. The reimbursement parameters are calculated based on the cost of each medical task under each standard medical keyword and the reimbursement ratio corresponding to each medical task.
[0125] The embodiment of the voice consultation processing method applicable to medical service programs and offline consultation scenarios described here can be executed in conjunction with the above-described embodiment of the voice consultation processing method applicable to the server side of medical service programs and offline consultation scenarios. When reading this embodiment, please refer to the above embodiment, and when reading the above embodiment, please refer to this embodiment.
[0126] This manual provides an embodiment of a voice-based medical consultation processing device as follows: In the above embodiments, a voice consultation processing method is provided, and correspondingly, a voice consultation processing device is also provided, which will be described below with reference to the accompanying drawings.
[0127] Reference Figure 11 The diagram shows an embodiment of a voice consultation processing device provided in this embodiment.
[0128] Since the apparatus embodiments correspond to the method embodiments, the descriptions are relatively simple. For relevant parts, please refer to the corresponding descriptions of the method embodiments provided above. The apparatus embodiments described below are merely illustrative.
[0129] This embodiment provides a voice consultation processing device, including: The voice acquisition module 1102 is configured to acquire the voice stream of the user and medical personnel during the consultation mode collected by the service program when the consultation mode is in consultation mode; the consultation mode is activated after the consultation detection based on the user's gesture image sequence passes. The speech recognition module 1104 is configured to perform speech separation and speech recognition on the consultation speech stream to obtain user text and the medical personnel's diagnosis text; The interface call module 1106 is configured to call the medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text to obtain medical expenses. The information construction module 1108 is configured to construct consultation information based on the user text, the diagnosis text, and the medical expenses, and return the constructed structured consultation information to the service program.
[0130] Another embodiment of the voice consultation processing device provided in this manual is as follows: In the above embodiments, another voice consultation processing method is provided, and correspondingly, another voice consultation processing device is also provided, which will be described below with reference to the accompanying drawings.
[0131] Reference Figure 12 The diagram shows an embodiment of a voice consultation processing device provided in this embodiment.
[0132] Since the apparatus embodiments correspond to the method embodiments, the descriptions are relatively simple. For relevant parts, please refer to the corresponding descriptions of the method embodiments provided above. The apparatus embodiments described below are merely illustrative.
[0133] This embodiment provides a voice consultation processing device, including: The consultation detection module 1202 is configured to perform consultation detection based on the user's gesture image sequence, and to start the consultation mode of the service program after the consultation detection is passed; The voice sending module 1204 is configured to call the voice acquisition component to acquire the consultation voice stream between the user and the medical personnel when the service program is in the consultation mode, and send the consultation voice stream to the server. The rendering module 1206 is configured to render and display the structured consultation information returned by the server, which is based on user text, medical personnel's diagnosis text, and medical expenses. The medical expenses are obtained by calling the medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text.
[0134] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0135] This manual provides an example of a voice-based medical consultation processing device as follows: Corresponding to the voice consultation processing method described above, based on the same technical concept, one or more embodiments of this specification also provide a voice consultation processing device, which is used to execute the voice consultation processing method provided above. Figure 13 This is a schematic diagram of the structure of a voice consultation processing device provided for one or more embodiments of this specification.
[0136] This embodiment provides a voice consultation processing device, including: like Figure 13 As shown, device 1300 mainly consists of a communication interface 1302, a user interface 1304, a processor 1306, and a data storage 1308. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 1310. The communication interface 1302 enables device 1300 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 1302 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 1302 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 1302 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 1302 may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide-area wireless interfaces. The user interface 1304 includes receiving user input and providing output to the user. Therefore, user interface 1304 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 1304 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 1304 may include software, circuitry, or other forms of logic capable of transmitting data to and receiving data from external user input / output devices. Additionally or alternatively, device 1300 may support remote access from other devices via communication interface 1302 or another physical interface (not shown). User interface 1304 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 1304 may also be configured as a display device for rendering or displaying text fragments.
[0137] Processor 1306 may include one or more general-purpose processors and / or special-purpose processors. Data storage 1308 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 1306. Data storage 1308 may include removable and non-removable components.
[0138] Processor 1306 is capable of executing program instructions 1318 (e.g., compiled or uncompiled program logic and / or machine code) stored in data store 1308 to perform the various functions described herein. Data store 1308 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 1300, enable device 1300 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 1318 by processor 1306 may result in processor 1306 using data 1312. For example, program instructions 1318 may include an operating system 1322 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 1300 and one or more application programs 1320 (e.g., a browser, social application, or game application). Similarly, data 1312 may include operating system data 1316 and application data 1314. Operating system data 1316 is primarily accessible to operating system 1322, while application data 1314 is primarily accessible to one or more application programs 1320. Application data 1314 may reside in a file system visible or hidden to the user of device 1300. Application 1320 may communicate with operating system 1322 via one or more application programming interfaces (APIs). These APIs facilitate application 1320 reading and / or writing application data 1314, transmitting or receiving information via communication interface 1302, receiving or displaying information on user interface 1304, etc. In some terms, application 1320 may be simply referred to as an "app". Furthermore, application 1320 may be downloaded to device 1300 through one or more online app stores or app markets. However, applications may also be installed on device 1300 in other ways, such as through a web browser or a physical interface on device 1300 (e.g., a USB port).
[0139] In one specific embodiment, the voice consultation processing device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the voice consultation processing device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: The system acquires the voice stream of a user's consultation with a medical professional, collected while the service program is in consultation mode; the consultation mode is activated after a consultation detection based on the user's gesture image sequence is passed. The consultation voice stream is subjected to voice separation and voice recognition to obtain the user's text and the medical personnel's diagnosis text; The medical insurance interface is invoked to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis and treatment text. Based on the user text, the diagnosis text, and the medical expenses, consultation information is constructed, and the structured consultation information obtained is returned to the service program.
[0140] Another embodiment of the voice-based medical consultation processing device provided in this manual is as follows: Corresponding to the other voice consultation processing method described above, based on the same technical concept, one or more embodiments of this specification also provide another voice consultation processing device, which is used to execute the other voice consultation processing method provided above. Figure 14 This is a schematic diagram of the structure of a voice consultation processing device provided for one or more embodiments of this specification.
[0141] This embodiment provides a voice consultation processing device, including: like Figure 14As shown, device 1400 mainly consists of a communication interface 1402, a user interface 1404, a processor 1406, and a data storage 1408. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 1410. The communication interface 1402 enables device 1400 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 1402 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 1402 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 1402 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 1402 may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide-area wireless interfaces. The user interface 1404 includes receiving user input and providing output to the user. Therefore, user interface 1404 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 1404 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 1404 may include software, circuitry, or other forms of logic capable of transmitting data to and receiving data from external user input / output devices. Additionally or alternatively, device 1400 may support remote access from other devices via communication interface 1402 or another physical interface (not shown). User interface 1404 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 1404 may also be configured as a display device for rendering or displaying text fragments.
[0142] Processor 1406 may include one or more general-purpose processors and / or special-purpose processors. Data storage 1408 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 1406. Data storage 1408 may include removable and non-removable components.
[0143] Processor 1406 is capable of executing program instructions 1418 (e.g., compiled or uncompiled program logic and / or machine code) stored in data store 1408 to perform the various functions described herein. Data store 1408 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 1400, enable device 1400 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 1418 by processor 1406 may result in processor 1406 using data 1412. For example, program instructions 1418 may include an operating system 1422 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 1400 and one or more application programs 1420 (e.g., a browser, social application, or game application). Similarly, data 1412 may include operating system data 1416 and application data 1414. Operating system data 1416 is primarily accessible to operating system 1422, while application data 1414 is primarily accessible to one or more application programs 1420. Application data 1414 may reside in a file system visible or hidden from the user of device 1400. Application 1420 may communicate with operating system 1422 via one or more application programming interfaces (APIs). These APIs facilitate application 1420 reading and / or writing application data 1414, transmitting or receiving information via communication interface 1402, receiving or displaying information on user interface 1404, etc. In some terms, application 1420 may be simply referred to as an "app". Furthermore, application 1420 may be downloaded to device 1400 through one or more online app stores or app markets. However, applications may also be installed on device 1400 in other ways, such as through a web browser or a physical interface on device 1400 (e.g., a USB port).
[0144] In one specific embodiment, the voice consultation processing device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the voice consultation processing device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: The system performs a diagnostic test based on the user's gesture image sequence, and initiates the diagnostic mode of the service program after the diagnostic test is passed. The voice acquisition component is invoked to collect the voice stream of the user and medical personnel during the consultation mode of the service program, and the voice stream is sent to the server. Render and display the structured consultation information returned by the server, which is based on the user's text, the medical personnel's diagnosis text, and medical expenses; The medical expenses are obtained by calling the medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text.
[0145] This specification provides an embodiment of a computer-readable storage medium as follows: Corresponding to the voice consultation processing method described above, based on the same technical concept, one or more embodiments of this specification also provide a computer-readable storage medium.
[0146] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, which, when executed, perform the following steps: The system acquires the voice stream of a user's consultation with a medical professional, collected while the service program is in consultation mode; the consultation mode is activated after a consultation detection based on the user's gesture image sequence is passed. The consultation voice stream is subjected to voice separation and voice recognition to obtain the user's text and the medical personnel's diagnosis text; The medical insurance interface is invoked to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis and treatment text. Based on the user text, the diagnosis text, and the medical expenses, consultation information is constructed, and the structured consultation information obtained is returned to the service program.
[0147] It should be noted that the embodiments of a computer-readable storage medium described in this specification and the embodiments of a voice diagnosis processing method described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0148] Another embodiment of a computer-readable storage medium provided in this specification is as follows: In response to the other voice consultation processing method described above, based on the same technical concept, one or more embodiments of this specification also provide another computer-readable storage medium.
[0149] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, which, when executed, perform the following steps: The system performs a diagnostic test based on the user's gesture image sequence, and initiates the diagnostic mode of the service program after the diagnostic test is passed. The voice acquisition component is invoked to collect the voice stream of the user and medical personnel during the consultation mode of the service program, and the voice stream is sent to the server. Render and display the structured consultation information returned by the server, which is based on the user's text, the medical personnel's diagnosis text, and medical expenses; The medical expenses are obtained by calling the medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text.
[0150] It should be noted that the embodiments of another computer-readable storage medium described in this specification and the embodiments of another voice diagnosis processing method described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0151] This specification provides an example of a computer program product as follows: Corresponding to the voice consultation processing method described above, based on the same technical concept, one or more embodiments of this specification also provide a computer program product.
[0152] A computer program product includes a computer program / instructions that, when executed by a processor, perform the following steps: The system acquires the voice stream of a user's consultation with a medical professional, collected while the service program is in consultation mode; the consultation mode is activated after a consultation detection based on the user's gesture image sequence is passed. The consultation voice stream is subjected to voice separation and voice recognition to obtain the user's text and the medical personnel's diagnosis text; The medical insurance interface is invoked to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis and treatment text. Based on the user text, the diagnosis text, and the medical expenses, consultation information is constructed, and the structured consultation information obtained is returned to the service program.
[0153] It should be noted that the embodiments of a computer program product described in this specification and the embodiments of a voice diagnosis processing method described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0154] Another example of a computer program product provided in this specification is as follows: Corresponding to the other voice consultation processing method described above, based on the same technical concept, one or more embodiments of this specification also provide another computer program product.
[0155] A computer program product includes a computer program / instructions that, when executed by a processor, perform the following steps: The system performs a diagnostic test based on the user's gesture image sequence, and initiates the diagnostic mode of the service program after the diagnostic test is passed. The voice acquisition component is invoked to collect the voice stream of the user and medical personnel during the consultation mode of the service program, and the voice stream is sent to the server. Render and display the structured consultation information returned by the server, which is based on the user's text, the medical personnel's diagnosis text, and medical expenses; The medical expenses are obtained by calling the medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text.
[0156] It should be noted that the embodiments of another computer program product described in this specification and the embodiments of another voice diagnosis processing method described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0157] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments. For example, the device embodiment, equipment embodiment and computer-readable storage medium embodiment are all similar to the method embodiment, so the description is relatively simple. When reading the relevant content of the device embodiment, equipment embodiment and computer-readable storage medium embodiment, please refer to the description of the method embodiment.
[0158] Although one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is only one of many possible execution orders and does not represent the only execution order. Therefore, when the claims involve method steps, any changes or adjustments to the order of such steps, or the parallelism between steps, are also within the scope of protection of the claims.
[0159] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.
[0160] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0161] In the 1930s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many improvements to the methodology today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that an improvement to the methodology cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0162] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0163] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0164] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in one or more software and / or hardware.
[0165] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0166] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable test processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable test processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0167] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable test processing equipment to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0168] These computer program instructions can also be loaded onto a computer or other programmable test processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0169] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0170] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0171] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0172] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of features includes not only those features but also other features not expressly listed, or features inherent to such process, method, article, or apparatus. Without further limitations, a feature defined by the phrase "comprising one..." does not exclude the presence of other identical features in the process, method, article, or apparatus that includes said feature.
[0173] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0174] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0175] The above description is merely an embodiment of this document and is not intended to limit the scope of this document. Various modifications and variations can be made to this document by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this document should be included within the scope of the claims of this document.
Claims
1. A voice-based medical consultation processing method, comprising: The service program collects audio streams of consultations between users and medical personnel when the consultation mode is active. The consultation mode is activated after the consultation detection based on the user's gesture image sequence is passed. The consultation voice stream is subjected to voice separation and voice recognition to obtain the user's text and the medical personnel's diagnosis text; The medical insurance interface is invoked to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis and treatment text. Based on the user text, the diagnosis text, and the medical expenses, consultation information is constructed, and the structured consultation information obtained is returned to the service program.
2. The voice consultation processing method according to claim 1, wherein the service program, when the user is in a medical institution, redirects to the consultation page based on the user's access instruction; After being redirected to the consultation page, the service program calls the image acquisition component to collect the gesture image sequence.
3. The voice consultation processing method according to claim 2, wherein the image acquisition component includes an image sensor with adjusted operating parameters; the adjustment of operating parameters is performed when the image sensor detects the outline of the user's hand.
4. The voice consultation processing method according to claim 1, wherein the step of querying medical insurance data and calculating medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text to obtain medical expenses includes: Based on the registration information, the user's identity is verified to confirm their participation in medical insurance. After identity verification is passed, multiple reimbursement parameters are obtained by querying multiple standard medical keywords based on the aforementioned medical keywords, and the medical expenses under the corresponding standard medical keywords are calculated according to each reimbursement parameter.
5. The voice consultation processing method according to claim 4, wherein calculating the medical expenses under the corresponding standard medical keywords based on each reimbursement parameter includes: Based on the drug cost of each drug under each standard medical keyword and the reimbursement ratio corresponding to each standard medical keyword, calculate the medical insurance payment and user payment for each drug.
6. The voice consultation processing method according to claim 5, wherein constructing consultation information based on the user text, the diagnosis text, and the medical expenses includes: The medical insurance reimbursement and user payment fees for each of the aforementioned drugs will be distributed to the medical personnel. A consultation report is constructed based on the medical insurance payment and user payment costs of the target drugs selected by the medical personnel from the various drugs, the user text, and the diagnosis and treatment text.
7. The voice consultation processing method according to claim 1, wherein the voice separation of the consultation voice stream includes: The frequency domain transformation of the consultation voice stream is performed to obtain a voice spectrogram, and the voice spectrogram is input into a spectrum separation model for voiceprint recognition and spectrum separation to obtain the user's voice spectrum and the medical personnel's voice spectrum; Perform inverse frequency domain transformation on the user's speech spectrum and the speech spectrum to obtain the user's speech and the medical personnel's diagnostic speech.
8. The voice consultation processing method according to claim 7, wherein the voiceprint recognition and spectrum separation include: The mask features of the medical personnel are calculated based on the speech spectrogram and the voiceprint features of the medical personnel, and the mask features of the user are calculated based on the speech spectrogram and the user's voiceprint features. The speech spectrum is calculated based on the speech spectrogram and the masking features of the medical personnel, and the user's speech spectrum is calculated based on the speech spectrogram and the user's masking features.
9. The voice consultation processing method according to claim 1, wherein the consultation detection is passed, is implemented in the following manner: Gesture recognition results are obtained by performing gesture recognition based on the gesture image sequence; If the gesture recognition result indicates that the user's gesture is a consultation gesture, the consultation test is considered successful.
10. The voice consultation processing method according to claim 9, wherein obtaining the gesture recognition result by performing gesture recognition based on the gesture image sequence includes: Hand key point detection is performed on each user gesture image in the gesture image sequence to obtain the coordinates of the hand key points; The gesture recognition result is obtained by performing gesture recognition based on the coordinates of the key points of the hand and the duration of the gesture.
11. The voice consultation processing method according to claim 10, wherein obtaining the gesture recognition result by performing gesture recognition based on the coordinates of the hand key points and the duration of the gesture includes: Calculate the direction vectors of the palm and wrist based on the coordinates of the key points of the hand, and calculate the negative projection offset of the direction vectors on the image coordinate axes; If the negative projection offset is greater than a preset offset and the duration of the gesture is within a preset duration range, the gesture recognition result is determined to be a user gesture for medical consultation.
12. The voice consultation processing method according to claim 1, after the step of performing voice separation and voice recognition on the consultation voice stream to obtain user text and the medical personnel's diagnosis text, further includes: Read the timestamp carried in the diagnostic labeling request sent by the service program; Based on the time interval of the timestamp, key text blocks are marked in the diagnosis text, and key segments are marked in the diagnosis voice segments of the medical personnel in the subsequent consultation voice stream.
13. The voice consultation processing method according to claim 12, wherein the diagnostic annotation request is sent after the service program passes the diagnostic annotation detection of the collected user gesture image sequence; Accordingly, the diagnostic labeling test is passed in the following manner: Finger key point recognition is performed on each user gesture image in the user gesture image sequence to obtain the coordinates of the finger key points; The movement direction and duration of the finger are obtained by recognizing the movement direction and calculating the movement duration based on the coordinates of the key points of the finger. If the movement direction and duration of the finger trigger the diagnostic annotation conditions, the diagnostic annotation detection is determined to be successful.
14. The voice consultation processing method according to claim 12, wherein constructing consultation information based on the user text, the diagnosis text, and the medical expenses includes: The user text, the diagnosis text annotated with key text blocks, the diagnosis text fragments of the diagnosis audio fragments annotated with key segments, and the medical expenses are filled into the corresponding consultation fields in the electronic consultation template to obtain a structured consultation report.
15. The voice consultation processing method according to claim 1, after the step of performing voice separation and voice recognition on the consultation voice stream to obtain user text and the medical personnel's diagnosis text, it further includes: If the medical keyword recognition result obtained from the diagnosis text is empty, the medical insurance engine is called to calculate the cost based on the medical insurance code associated with the medical keywords entered by the medical personnel, and the medical insurance payment cost and the user payment cost are obtained. A structured consultation report is constructed based on the user text, the medical treatment text, the medical insurance payment, and the user payment.
16. The voice consultation processing method according to claim 1, after the step of performing voice separation and voice recognition on the consultation voice stream to obtain user text and the medical personnel's diagnosis text is executed, and before the step of calling the medical insurance interface to perform medical insurance data query and medical expense calculation based on the user's registration information in the service program and medical keywords in the diagnosis text to obtain medical expenses, further includes: The medical knowledge graph is used to search for standard medical keywords that match the medical keywords in the diagnosis and treatment text, and it is determined whether there are multiple standard medical keywords. If not, the steps of calling the medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text are executed to obtain medical expenses.
17. The voice consultation processing method according to claim 16, if the determination of whether the standard medical keyword is the result of multiple operations is yes, the following operation is performed: The payment interface of the medical insurance platform is called to query medical insurance payment information based on the registration information; The medical insurance payment amount and user payment amount are read from the medical insurance payment information, and a structured consultation report is constructed based on the user text, the diagnosis and treatment text, the medical insurance payment amount, and the user payment amount.
18. The voice consultation processing method according to claim 1, further comprising: The registration information and the standard medical keywords of the medical keywords are sent to the medical insurance platform to match the registration information and the standard medical keywords with the user information and medical keyword information sent by the medical institution. If the match is successful, the matching medical insurance fund amount is queried in the medical insurance settlement package based on the disease keywords in the registration information and the standard medical keywords, and the medical insurance fund is paid to the medical institution according to the medical insurance fund amount.
19. A voice consultation processing method, comprising: The system performs a diagnostic test based on the user's gesture image sequence, and initiates the diagnostic mode of the service program after the diagnostic test is passed. The voice acquisition component is invoked to collect the voice stream of the user and medical personnel during the consultation mode of the service program, and the voice stream is sent to the server. Render and display the structured consultation information returned by the server, which is based on the user's text, the medical personnel's diagnosis text, and medical expenses; The medical expenses are obtained by calling the medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text.
20. A voice consultation processing device, comprising: The voice acquisition module is configured to acquire the voice stream of the user and medical personnel during the consultation mode of the service program. The consultation mode is activated after the consultation detection based on the user's gesture image sequence is passed. The speech recognition module is configured to perform speech separation and speech recognition on the consultation speech stream to obtain user text and the medical personnel's diagnosis text; The interface call module is configured to call the medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text to obtain medical expenses; The information construction module is configured to construct consultation information based on the user text, the diagnosis text, and the medical expenses, and return the constructed structured consultation information to the service program.
21. A voice consultation processing device, comprising: The consultation detection module is configured to perform consultation detection based on the user's gesture image sequence, and start the consultation mode of the service program after the consultation detection is passed; The voice sending module is configured to call the voice acquisition component to collect the consultation voice stream between the user and the medical personnel when the service program is in the consultation mode, and send the consultation voice stream to the server. The rendering module is configured to render and display the structured consultation information returned by the server, which is based on user text, medical personnel's diagnosis text, and medical expenses. The medical expenses are obtained by calling the medical insurance interface to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis text.
22. A voice consultation processing device, comprising: processor; And, a memory configured to store computer-executable instructions, which, when executed, cause the processor to: The system acquires the voice stream of a user's consultation with a medical professional, collected while the service program is in consultation mode; the consultation mode is activated after a consultation detection based on the user's gesture image sequence is passed. The consultation voice stream is subjected to voice separation and voice recognition to obtain the user's text and the medical personnel's diagnosis text; The medical insurance interface is invoked to query medical insurance data and calculate medical expenses based on the user's registration information in the service program and medical keywords in the diagnosis and treatment text. Based on the user text, the diagnosis text, and the medical expenses, consultation information is constructed, and the structured consultation information obtained is returned to the service program.
23. A computer-readable storage medium for storing computer-executable instructions that, when executed, implement the steps of the method of claim 1 or 19.