Two-person dominated multi-round inquiry type voice dialogue intelligent identification method

By using the intelligent identification method of voice dialogue during the doctor-patient consultation process, the problems of large workload and inaccurate records in traditional consultations are solved, and efficient and accurate medical record recording and sorting are achieved.

CN120089143APending Publication Date: 2025-06-03HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510246479.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The traditional doctor-patient consultation process has a large workload and inaccurate records, which leads to heavy burden on the doctor and may cause logical inconsistent problems.

Method used

The intelligent identification method of two-person-led multi-round inquiry voice dialogue is adopted, and through sound collection, voiceprint feature recognition, contextual voice translation and templated medical record generation, efficient recording and sorting of doctor-patient dialogues are achieved.

Benefits of technology

It improves the efficiency and accuracy of the doctor-patient consultation process, reduces the workload of doctors, and ensures the logic and integrity of medical records.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089143A_ABST
    Figure CN120089143A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voiceprint recognition, in particular to a two-person dominated multi-round inquiry type voice dialogue intelligent recognition method. The method comprises the following steps: S1, erecting a sound collector to obtain the sound of a speaker, and carrying out sound source orientation discrimination; s2, performing speaker identity initial recognition and gender recognition according to the voiceprint features of the speaker voice, and performing sound source orientation correction and fusion to obtain a final speaker identity recognition result; s3, scene language translation is carried out on the speaking content through a speech translation algorithm in combination with the inquiry scene, and a translation result is output; s4, performing rolling type arrangement and integration on the multi-round dialogue content corresponding to the identity; and S5, generating a template medical record according to the sorted and integrated dialogue content. According to the invention, the inquiry process has the remarkable advantages of high efficiency and intelligence, the working efficiency of doctors can be greatly improved, and the efficiency advantage is particularly obvious under the condition of emergency treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voiceprint recognition, and particularly to an intelligent identification method for two-person-led multi-round inquiry-based voice conversations. Background Art

[0002] The doctor-patient inquiry-based voice conversation is an important stage for doctors to analyze the patient's condition. Based on the conversation content at this stage, it is necessary to accurately sort out and record the patient's medical history and description of the main discomfort symptoms. The traditional method of doctors talking to patients while sorting out and recording not only significantly increases the workload of doctors, but also may have problems such as inaccurate recording and inconsistent logic. Summary of the Invention

[0003] The present invention provides an intelligent identification method for two-person-led multi-round inquiry-based voice conversations, aiming to solve the technical problems such as large workload and inaccurate recording existing in traditional doctor-patient consultations.

[0004] The present invention provides an intelligent identification method for two-person-led multi-round inquiry-based voice conversations, including the following steps:

[0005] S1. Set up a sound collector to obtain the voice of the speaker and perform sound source azimuth discrimination;

[0006] S2. Perform initial identification of the speaker's identity and gender recognition based on the voiceprint characteristics of the speaker's voice, and perform sound source azimuth correction and fusion to obtain the final speaker identity recognition result;

[0007] S3. Through a voice translation algorithm combined with the inquiry scenario, perform scenario language translation on the spoken content and output the translation result;

[0008] S4. Scroll and synthesize the multi-round conversation content corresponding to the identity;

[0009] S5. Generate a templated medical record based on the sorted and synthesized conversation content.

[0010] As a further improvement of the present invention, the specific content of step S1 includes:

[0011] The collection surface of the sound collector faces the inquirer and is flush with the inquirer's head. When performing sound source azimuth discrimination, if the sound source azimuth is directly opposite the sound collector, the voice comes from the inquirer; if the sound source azimuth is in other directions except directly opposite the sound collector, the voice comes from the interviewee. At the same time, set a confidence threshold, and give the confidence of the discrimination result according to the sound source azimuth. Only when the confidence is higher than the threshold, combine the identity of the sound source azimuth discrimination with the initial identity recognition to perform identity recognition based on fusion to obtain the final speaker identity recognition result.

[0012] As a further improvement of the present invention, in the step S1, the sound collector adopts an array composed of a plurality of acoustic sensors. The specific process of the sound collector for discriminating the sound source direction includes:

[0013] S11. Initialization and parameter setting: Set the number M of acoustic sensors in the array and their relative positions. Assume that the spatial coordinates of the M acoustic sensors are respectively (x 1 , y 1 , z 1 ), (x 2 , y 2 , z 2 ),..., (x M , y M , z M );

[0014] S12. Processing of received signals: Assume that the signal received by the array is x i (t), where the subscript i represents the i-th acoustic sensor and t is the time index. Perform Fourier transform on the signals received by each acoustic sensor to obtain the frequency-domain signals X i (f);

[0015] S13. Calculate the autocovariance matrix: Use the autocovariance matrix R to describe the correlation between each acoustic sensor. Assume that the signal X i (f) is the frequency-domain signal of the i-th acoustic sensor in the array. The covariance matrix can be expressed as:

[0016]

[0017] Among them, Ε[·] represents the expectation operation, is the complex conjugate of X i (f);

[0018] S14. Calculate the beam spectrum of the target direction: Calculate the spatial spectrogram, that is, for each possible sound source direction, calculate the response of the array. Assume that scanning is to be performed at a certain azimuth angle θ and elevation angle φ. The beam response of the target direction is calculated by the following formula:

[0019]

[0020] Among them, k is the wave vector, r i is the position vector of the i-th acoustic sensor, represents the phase delay to the acoustic sensor, where the calculation formula of the wave vector is:

[0021]

[0022] Among them, f is the signal frequency and θ is the azimuth angle of the target direction;

[0023] k·r i is the phase delay from the acoustic sensor i to the sound source, and the calculation formula is:

[0024]

[0025] The beam response value is obtained by solving the optimization problem:

[0026]

[0027] where P(θ,φ) is the beam response at the direction angle θ and the elevation angle φ, and a(θ,φ) H is the conjugate transpose of a(θ,φ), and R -1 is the inverse of the autocovariance matrix;

[0028] S15. Calculation of the scanning direction and the beam spectrum: Calculate the beam response P(θ,φ) at each preset direction θ and φ. The scanning direction is within a specific range, that is, from θ min to θ max , from φ min to φ max to obtain a complete spatial spectrum diagram. Take θ ∈ [0,π] and For each direction, calculate and store its beam response;

[0029] S16. Solve the sound source direction: The beam response map P(θ) obtained through the scanning process will show the response intensity in different directions. The sound source will form a peak in the beam pattern, and this peak corresponds to the direction of the sound source. The direction θ source and φ source of the sound source will correspond to the maximum value of P(θ):

[0030]

[0031] The direction θ and φ of the sound source obtained by sound source localization;

[0032] S17. Calculate the distance from the sound source to the array using the time difference of arrival: The time difference of arrival is used to estimate the position of the sound source by measuring the time difference of signal propagation from the sound source to different sensors; Assume that the propagation times from the sound source to each sensor are t 1 , t 2 , t 3 , t 4 , t 5 , t 6 , respectively. The time differences of the signals received by each sensor can be known, and based on these time differences, the relationship of the distance differences is established.

[0033] The distance from the sound source to the sensor is obtained through the speed of sound v and the time t i :

[0034] d i = v·t i

[0035] Assume the distance from the sound source to sensor M 1 is d 1 , and the distance to sensor M 2 is d 2 , then the distance difference between the two is:

[0036] d 2 -d 1 = v·Δt 12

[0037] where Δt 12 = t 2 -t 1 is the arrival time difference from the sound source to sensor M 1 and M 2 . For the time differences between other sensors, equations are established similarly. By using the distance formula in three-dimensional space, the distance differences between the sound source and each sensor are established, and the least squares method is used to solve for the distance d of the sound source source ;

[0038] S18. Locate the three-dimensional coordinates of the sound source: Based on the direction (θ source , φ source ) and the distance d from the sound source to the array source , deduce the three-dimensional coordinates of the sound source:

[0039] x source = d source sin(θ source )cos(φ source )

[0040] y source = d source sin(θ source )cos(φ source ).

[0041] z source = d source cos(θ source )

[0042] As a further improvement of the present invention, in the step S2, the process of initially identifying the speaker's identity based on the voiceprint characteristics of the speaker's voice includes:

[0043] After the interrogation process starts, only the voiceprint characteristics, identity, and ID number of the interrogator are known. For all the subsequent detected speech of the speaker, the following recognition operations are performed:

[0044] Set a threshold for the similarity of voiceprint features. If the similarity between the voiceprint features of the current speaker and the voiceprint features of a certain interviewee with a determined ID number is greater than the threshold, match it to the identity with the determined ID number. If the similarity between the voiceprint features of the current speaker and the voiceprint features of any identity with a determined ID number is less than the threshold, identify it as a new identity and assign a new ID number to it.

[0045] As a further improvement of the present invention, in the step S2, the process of extracting the voiceprint features of the interviewee based on the voice of the interviewee includes:

[0046] The voiceprint features extracted from the voice of the interviewee reading the specified content before starting the interview every day are respectively recorded as F 1 ,…,F h , where F 1 ,…,F h respectively represent the first day to the hth day; after collecting the voiceprint features from the current speaking sentence, select the two voiceprint features with the highest similarity to the current voiceprint features from F 1 ,…,F h , which are called the most suitable voiceprint features and are recorded as E 1 、E 2 . Let the similarity between the current voiceprint feature and E 1 、E 2 be S 1 、S 2 respectively. The similarity calculation formula of the current voiceprint feature and the doctor is S = 0.9×S 1 +0.1×S 2 .

[0047] As a further improvement of the present invention, in the step S2, the gender recognition of the speaker is carried out based on the binary classification algorithm of training examples, specifically including:

[0048] Collect a batch of Chinese voice data recording the gender of the speaker in a question-and-answer scenario. After taking the method of sentence-by-sentence annotation of the data, use it as a training example; extract the voiceprint feature examples of men and women respectively, and assume that the number of training examples of men and women is r and s respectively. The class labels of men and women are respectively defined as symmetric row vectors (1 -1) and (-1 1), and the examples of men and women are respectively expressed as row vectors p 1 ,...,p r , q 1 ,...,q s ; Let Then let And L is a matrix composed of the class labels of each corresponding training example; assuming that there approximately exists a linear mapping relationship GY = L, and this mapping relationship is the model for speaker gender recognition, then Y is solved according to the following scheme: If cond(G) < 1.0e 6 , then Y = (G T G) -1 G T L, otherwise Y = (G T G + 0.0001) -1 G T L; cond(G) represents the condition number of matrix G; for the example of the gender to be discriminated, let the row vector representing its voiceprint feature be z, and let v = zY. Calculate the Euclidean distances between v and the row vectors (1 -1) and (-1 1) respectively, and denote them as d 1 、d 2 , if d 1 < d 2 , then the current example of the gender to be discriminated is determined to be male, otherwise it is determined to be female.

[0049] As a further improvement of the present invention, in the step S2, the modeling process of voiceprint feature acquisition includes:

[0050] Collect a batch of voice data from the inquiry dialogue scenario, mark all the voices in sentences, and use different ID numbers to represent different speakers. Use a Gaussian mixture model containing five Gaussian distributions to model the Mel frequency cepstral coefficients of the collected voice data: The specific formula of the Gaussian mixture model is In the formula, w j is the weight of the j-th Gaussian model, d is the dimension of the voiceprint feature x represented by the vector, μ j is the expectation of the j-th Gaussian model, Φ j is the covariance matrix of the j-th Gaussian model, and det(Φ j ) represents the value of the determinant of the covariance matrix Φ j .

[0051] As a further improvement of the present invention, in the step S2, the judgment process of the fusion of sound source azimuth, initial identity recognition, and gender recognition correction includes:

[0052] When the speaker's identity is initially identified as the inquirer, if the confidence level of the speaker's gender recognition result is higher than the threshold and the gender recognition result does not match the actual gender of the inquirer, it indicates that the initial identification of the speaker's identity is an incorrect result, and its final identity should be identified as the interviewee; if the orientation discrimination result of the current speech is directly in front of the sound collector and the confidence level is higher than the threshold, but the initial identification of the speaker's identity is the interviewee, it indicates an incorrect identification, and its final identity should be identified as the inquirer; among them, the confidence level threshold of the sound source orientation discrimination result is higher than the confidence level thresholds of the initial identity recognition and gender recognition discrimination results.

[0053] As a further improvement of the present invention, in step S3, after the inquiry scenario is adapted, the construction of a homophone and near-homophone vocabulary list for the inquiry scenario is included in the scenario language translation:

[0054] The homophone and near-homophone vocabulary list includes common noun phrases with wide coverage in the inquiry scenario, their word frequencies, and their Mandarin pronunciations and pinyin; for the speech translation results obtained by the speech translation algorithm, the post-processing algorithm first identifies all the noun phrases therein, and then queries each noun phrase in the constructed homophone and near-homophone vocabulary list. If the noun phrase cannot be found in the list, it is considered that the speech translation result is incorrect, and the following rules are adopted to replace the noun phrase: calculate the similarity between the pronunciation of the noun phrase and the pronunciations of all phrases in the homophone and near-homophone vocabulary list, and then replace the noun phrase in the original speech translation result with the phrase with the highest pronunciation similarity; if there are more than one noun phrase with the same and highest pronunciation similarity, the noun phrase with the highest word frequency is adopted.

[0055] As a further improvement of the present invention, in step S5, the medical records generated by the templatized medical record generation step include: a medical record manuscript, a preliminary medical record report, and a confirmed official medical record report; the medical record manuscript is a record of the patient's conversation in the actual order of the conversation; the editable preliminary medical record report is a medical record report that synthesizes the actual conversation content and conforms to the conventional format; the official medical record report is a medical record report supplemented, edited, and modified by the inquirer for the preliminary medical record report and finally submitted after confirmation;

[0056] The rule for subsequent medical record modification is: when the content of a certain medical record is modified, the corresponding content in all subsequent medical records is automatically modified, and the previous medical records are not modified.

[0057] The beneficial effects of the present invention are: for multi-round voice conversations in hospital scenarios where doctors and patients are the main parties and there may be participation of other people, an intelligent recognition and processing method is designed, which includes voice collection, instant recognition of the speaker's identity, situational voice translation, and templated medical record generation. This makes the consultation process have the significant advantages of high efficiency and intelligence, and can greatly improve the work efficiency of doctors. The efficiency advantage is particularly obvious in emergency treatment situations. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 This is the overall flow chart of the intelligent recognition method of two-person multi-round inquiry voice dialogue of the present invention;

[0059] Figure 2 It is the main flow chart of speaker identification in the present invention;

[0060] Figure 3 This is a flow chart of the initial recognition and matching of doctor-patient identities based on speech of the present invention;

[0061] Figure 4 is a flow chart of the situational speech translation of the present invention;

[0062] Figure 5 is a flow chart of the query scene adaptability post-processing algorithm of the present invention;

[0063] Figure 6 is a flow chart of the steps of generating a templated medical record of the present invention;

[0064] Figure 7 It is a flow chart of the subsequent penetration rules of medical record modification of the present invention. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0066] like Figure 1 As shown, the present invention provides a method for intelligently identifying a two-person multi-round inquiry voice dialogue, characterized in that it includes the following steps:

[0067] S1. Set up a sound collector to obtain the speaker's voice and determine the direction of the sound source;

[0068] S2. Perform initial speaker identification and gender recognition based on the voiceprint features of the speaker's voice, and perform sound source orientation correction and fusion to obtain the final speaker identification result;

[0069] S3. Use the voice translation algorithm to combine the inquiry scenario, translate the spoken content into situational language, and output the translation result;

[0070] S4. Scroll and synthesize the multi-round conversation content corresponding to the identity;

[0071] S5. Generate a templatized medical record based on the sorted and synthesized conversation content.

[0072] This embodiment is directed to multi-round voice conversations in a hospital scenario, mainly between a doctor and a patient and possibly with other bystanders involved. The inquirer is the doctor, and the person being inquired is the patient or other bystanders. A microphone can be used as the sound collector.

[0073] When conducting a medical interview in a hospital scenario, first, before the doctor officially starts the outpatient service for the day, according to the prompt, read a sentence of the specified content (the specified content is: "I am doctor ***, male / female, and now start today's outpatient service work"), and accordingly extract the voiceprint features reflecting the doctor's voice characteristics and record their gender. Secondly, during the next time of the day, while performing on-site scenario voice translation, perform initial speaker identity recognition and gender recognition based on the voiceprint features; if the speaker is determined to be the doctor, the initial identity of the corresponding speaker's speech content is marked as the doctor, otherwise the initial identity of the corresponding speaker's speech content is marked as the patient. The voiceprint features in the present invention are all Mel-frequency cepstral coefficients with a dimension of 12 for the original one-dimensional voice data.

[0074] As Figure 2 shown, in order to improve the accuracy of doctor-patient identity recognition, sound source identification (judgment of the direction of the speaking voice) is also performed, and it is combined with the gender recognition result based on the voice, as well as the initial identity recognition result based on the voiceprint. For each sentence, based on the comprehensive result, the corresponding doctor-patient identity recognition conclusion and the final matching of the speaker's identity and content are obtained. The microphone is installed in a way that faces the doctor directly and is basically at the same height as the doctor's head. Therefore, when performing sound source direction judgment, if the speaking voice comes from the doctor, the sound direction is in the direct front of the microphone. On the contrary, the voices of the patient and their family members come from other directions except the direct front of the microphone. At the same time, a confidence threshold is set. When performing sound source direction judgment and speaking voice gender judgment, the confidence of the judgment result is given at the same time. Only when the confidence is higher than the threshold, the identity determined by the sound source direction judgment is combined with the initial identity recognition to perform identity recognition based on fusion to obtain the final speaker identity recognition result.

[0075] The microphone for sound collection adopts a small cross-shaped structure composed of 6 sound sensors. The main steps of the sound source identification (judgment of the direction of the speaking voice) scheme are as follows:

[0076] S11. Initialization and parameter setting: Set the number M of sound sensors in the array and their relative positions. Assume the spatial coordinates of the M sound sensors are (x 1 , y 1 , z 1 ), (x2 , y 2 , z 2 )),...,(x M , y M , z M ); Determine the speed of sound v to be 343 m / s and the sampling frequency of the signal to be 48000 Hz.

[0077] S12. Processing of the received signal: Assume that the signal received by the array is x i (t), where the subscript i represents the i-th acoustic sensor and t is the time index. Perform Fourier transform on the signal received by each acoustic sensor to obtain the frequency-domain signal X i (f).

[0078] S13. Calculate the autocovariance matrix: Use the autocovariance matrix R to describe the correlation between each acoustic sensor. Assume that the signal X i (f) is the frequency-domain signal of the i-th acoustic sensor in the array. The covariance matrix can be expressed as:

[0079]

[0080] where Ε[·] represents the expectation operation, is the complex conjugate of X i (f).

[0081] S14. Calculate the beam spectrum of the target direction: Calculate the spatial spectrogram (beam pattern), that is, for each possible sound source azimuth, calculate the response of the array. Assume that scanning is to be performed at a certain direction angle θ and elevation angle φ. The beam response of the target direction is calculated by the following formula:

[0082]

[0083] where k is the wave vector, r i is the position vector of the i-th acoustic sensor, represents the phase delay to the acoustic sensor, where the calculation formula of the wave vector is:

[0084]

[0085] where f is the signal frequency and θ is the azimuth angle of the target direction;

[0086] k·r i is the phase delay from the acoustic sensor i to the sound source, and the calculation formula is:

[0087]

[0088] Obtain the beam response value by solving the optimization problem:

[0089]

[0090] Among them, P(θ, φ) is the beam response of the azimuth angle θ and the elevation angle φ, and a(θ, φ) H is the conjugate transpose of a(θ, φ), and R -1 is the inverse of the autocovariance matrix.

[0091] S15. Calculation of the scanning direction and the beam spectrum: Calculate the beam response P(θ, φ) at each preset direction θ and φ. The scanning direction is carried out within a specific range, that is, from θ min to θ max , from φ min to φ max to obtain a complete spatial spectrum diagram. Take θ ∈ [0, π] and For each direction, calculate and store its beam response.

[0092] S16. Solve the sound source direction: The beam response map P(θ) obtained through the scanning process will show the response intensity in different directions. The sound source will form a peak in the beam pattern, and this peak corresponds to the direction of the sound source. The direction θ source and φ source of the sound source will correspond to the maximum value of P(θ):

[0093]

[0094] The direction θ and φ of the sound source obtained by sound source localization.

[0095] S17. Calculate the distance from the sound source to the array using the time difference of arrival (TDOA): The time difference of arrival is used to estimate the position of the sound source by measuring the time difference of the signal propagation from the sound source to different sensors; assume that the propagation times from the sound source to each sensor are t 1 , t 2 , t 3 , t 4 , t 5 , t 6 , and the time differences of the signals received by each sensor can be known. Based on these time differences, the relationship of the distance differences is established.

[0096] The distance from the sound source to the sensor can be obtained through the sound speed v and the time t i :

[0097] d i = v · t i

[0098] Assume that the distance from the sound source to sensor M 1 is d 1 , and the distance to sensor M 2 is d 2 , then the distance difference between the two is:

[0099] d 2 -d 1 =v·Δt 12

[0100] where Δt 12 =t 2 -t 1 The time difference of arrival from the sound source to the sensor to M 1 and M 2 For the time differences between other sensors, equations are similarly established. By using the distance formula in three-dimensional space, the distance differences between the sound source and each sensor are established, and the least squares method is used to solve for the distance d of the sound source source .

[0101] S18. Locate the three-dimensional coordinates of the sound source: Based on the direction (θ source , φ source ) and the distance d from the sound source to the array source , deduce the three-dimensional coordinates of the sound source:

[0102] x source =d source sin(θ source )cos(φ source )

[0103] y source =d source sin(θ source )cos(φ source ).

[0104] z source =d source cos(θ source )

[0105] In step S2, the process of initial identification and matching of the speaker's identity based on the voiceprint characteristics of the speaker's voice includes: setting a threshold for the similarity of voiceprint characteristics. As shown in Figure 3 , the dotted line with an arrow on the right side of the figure indicates that when the similarity between the voiceprint characteristics of the current speaker and the voiceprint characteristics of a certain identity with a determined ID number is greater than the threshold, it is matched to the identity with the determined ID number. On the contrary, the solid line with an arrow on the right side of the figure indicates that when the similarity between the voiceprint characteristics of the current speaker and the voiceprint characteristics of any identity with a determined ID number is less than the threshold, it is identified as a new identity and a new ID number is assigned to it. After the diagnosis and treatment process of a new patient begins, only the voiceprint characteristics and identity of the doctor and their ID number are known. Then, for all the spoken voices detected subsequently, the initial identification and matching steps of the identity in this Figure 3 are performed, the spoken voice is automatically segmented, and each sentence of the voice is processed according to the steps in Figure 3 .

[0106] In the present invention, voice-based doctor identity matching plays a very important role. Given that voice is collected and voiceprint features are extracted for each doctor's outpatient service every day, the present invention designs an optimal voiceprint feature automatic screening scheme for doctor identity matching, which can effectively adapt to the possible voice changes of doctors due to physical reasons (for example, a normal voice may become somewhat hoarse after long-term speaking). Specifically, the voiceprint features extracted from the voice of the doctor reading the specified content before starting the outpatient service every day are respectively recorded as F 1 , …, F h (F 1 , …, F h respectively represent the first day to the h-th day). After collecting the voiceprint features from the current spoken sentence, in order to determine whether it is the speech content of the doctor, two voiceprint features with the highest similarity to the current voiceprint feature are selected from F 1 , …, F h , which are called the optimal voiceprint features and denoted as E 1 , E 2 . Let the similarities between the current voiceprint feature and E 1 , E 2 be S 1 , S 2 respectively. The similarity calculation formula between the current voiceprint feature and the doctor obtained through the optimal experiment is S = 0.9×S 1 +0.1×S 2 .

[0107] The gender recognition scheme of the speaker adopts a binary classification algorithm based on training examples. Specifically, first, a batch of Chinese voice data recording the genders of speakers in a question-and-answer scenario is collected. After the data is marked sentence by sentence, it is used as training examples. Then, the voiceprint feature examples of men and women are respectively extracted. Suppose the numbers of training examples of men and women are r and s respectively, and the class labels of men and women are defined as symmetric row vectors (1 -1) and (-1 1) respectively. The examples of men and women are respectively represented as row vectors p 1 , ·, p r , q 1 ,..., q s ; Let Then let and L be the matrix composed of the class labels of the corresponding training examples; assume that there is approximately a linear mapping relationship GY = L (where Y is unknown), and this mapping relationship is the model for speaker gender recognition. Then, Y is solved according to the following scheme: If cond(G) < 1.0e 6 , then Y = (G T G) -1 G TL, otherwise Y = (G T G + 0.0001) -1 G T L. cond(G) represents the condition number of matrix G; for the sample whose gender is to be discriminated, let the row vector representing its voiceprint feature be z, and let v = zY. Calculate the Euclidean distances between v and the row vectors (1 -1) and (-1 1) respectively, and denote them as d 1 and d 2 . If d 1 < d 2 , then determine the sample whose gender is to be discriminated as male, otherwise determine it as female.

[0108] Voiceprint feature extraction is an important basis of the present invention. To enhance adaptability, first collect a batch of voice data from the doctor-patient dialogue scenario and label the speaker identities of all the voices in sentences (using different ID numbers to represent different speakers). Use a Gaussian mixture model containing five Gaussian distributions to model the Mel frequency cepstral coefficients of the collected voice data: The specific formula of the Gaussian mixture model is In the formula, w j is the weight of the j-th Gaussian model, d is the dimension of the voiceprint feature x represented by the vector, μ j is the expectation of the j-th Gaussian model, Φ j is the covariance matrix of the j-th Gaussian model, and det(Φ j ) represents the value of the determinant of the covariance matrix Φ j .

[0109] In step S2, the judgment process of the fusion of sound source orientation, initial identity recognition, and gender recognition correction includes: when the speaker's identity is initially identified as a doctor, if the confidence level of the gender recognition result is higher than the threshold and the gender recognition result does not match the actual gender of the doctor, it indicates that the initial identification of the speaker's identity is an incorrect result, and its final identification should be the identity of the patient (patient's family member). Similarly, if the orientation discrimination result of the current speech is directly in front of the microphone and the confidence level is higher than the threshold, but the initial identification of the speaker's identity is the identity of the patient (patient's family member), it indicates an incorrect identification, and its final identification should be the identity of the doctor. The confidence level threshold of the orientation discrimination result of the speech is set to a relatively high value, and the confidence level threshold of the sound source orientation discrimination result is higher than the confidence level thresholds of the initial identity recognition and gender recognition discrimination results. All the thresholds in the present invention are optimized based on experiments using training data. Specifically, first, an equally spaced threshold range for selection is given based on experience, and then the experimental accuracy rates obtained with different thresholds are compared, and finally a set of thresholds corresponding to the highest experimental accuracy rate is selected. In addition, the confidence levels in the present invention all refer to the probability values or approximate probability values between 0 and 1 corresponding to the classification decisions. Specifically, for a C-class classification problem, the model calculates the probability values or approximate probability values of a specific instance belonging to different classes, and then discriminates the instance as the class with the highest probability value or approximate probability value, and the corresponding (approximate) probability value is the confidence level associated with the classification decision.

[0110] As Figure 4 and Figure 5 shown, in step S3, the scenario speech translation function is implemented through the scenario speech translation software module. The core point of the scenario speech translation software module is the speech translation method adapted to the hospital multi-department consultation scenario. This method consists of an open-source speech translation algorithm plus a post-processing algorithm designed for the hospital multi-department consultation scenario adaptability. The key point of the post-processing algorithm for the hospital multi-department consultation scenario adaptability is the construction of a homophone and near-homophone vocabulary list for the hospital multi-department consultation scenario. Specifically, the homophone and near-homophone vocabulary list includes common noun phrases with wide coverage in the hospital multi-department consultation scenario (including basic vocabulary, combined vocabulary, and abbreviated vocabulary), word frequencies, and their Mandarin pronunciations (pinyin). For the speech translation result obtained by the open-source speech translation algorithm, the post-processing algorithm first identifies all the noun phrases in it, and then queries each noun phrase in the constructed homophone and near-homophone vocabulary list. If the noun phrase cannot be found in the table, it is considered that the speech translation result is incorrect, and the following rules are adopted to replace the noun phrase: calculate the similarity between the pronunciation of the noun phrase and the pronunciations of all the phrases in the homophone and near-homophone vocabulary list, and then replace the noun phrase in the original speech translation result with the phrase with the highest pronunciation similarity; if there are more than one phrase with the same and highest pronunciation similarity, the phrase with the highest word frequency is adopted.

[0111] AsFigure 6 As shown, in step S5, the medical records generated by the templated medical record generation step include: a medical record manuscript, a preliminary draft of a medical report, and a formal medical report after being confirmed by a doctor. The doctor and the software system can conveniently compare and verify based on the three reports. Both the preliminary draft of the medical report and the formal medical report are templated reports that comply with relevant specifications, and the department and other known information are automatically imported into the medical record by the software system. The medical record manuscript is a record of the patient's conversation in the actual order of the conversation. The editable preliminary draft of the medical report is a medical report that synthesizes the actual conversation content and complies with the conventional format. The preliminary draft of the medical report and the formal medical report adopt a templated design and both include parts such as the patient's chief complaint, medical history, past history, preliminary diagnosis, treatment suggestions, etc.; in specific applications, the doctor can select the parts that need to be included in the case through the check and delete functions provided by the software system; at the same time, the software system allows the doctor to add new parts to the case template. The chief complaint records the main symptoms and the duration. The medical history refers to the current medical history, which mainly records the onset date of the current illness, the main symptoms, the diagnosis and treatment conditions in other hospitals and the curative effect. The past history records the past history related to the current disease, personal history and family history. For the preliminary draft of the medical report, the doctor makes necessary supplements (including diagnosis conclusions, etc.), as well as edits and modifies, and finally submits it after confirmation to form a formal medical report.

[0112] As Figure 7 shown, the doctor is allowed to modify and supplement the medical record manuscript, the preliminary draft of the medical report, and the formal medical report. At the same time, in order to ensure the consistency of the content among the medical record manuscript, the preliminary draft of the medical report, and the formal medical report, a subsequent penetration scheme for the doctor's modified and supplemented content is designed. Specifically, the rules of the subsequent penetration scheme are: when the doctor modifies a certain medical record, the corresponding content in all subsequent medical records is automatically modified, and the previous medical records are not modified. For example, when the doctor modifies the medical record manuscript, the corresponding preliminary draft of the medical report and the formal medical report are modified in sequence. When the doctor modifies the preliminary draft of the medical report, the system modifies the corresponding formal medical report, but does not modify the medical record manuscript in the previous step. The doctor is allowed to make multiple modifications, and after each modification, the subsequent penetration scheme is automatically executed.

[0113] Finally, a data storage scheme to ensure data integrity is adopted, and at the same time, the original doctor-patient conversation audio, the extracted voiceprint features, the medical record manuscript, the preliminary draft of the medical report, the formal medical report, and the log of the doctor's modification of the case content are stored.

[0114] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for intelligent recognition of multi-round inquiry voice dialogues led by two people, characterized in that: The following steps are involved: S1. Set up a sound collector to obtain the speaker's voice and determine the direction of the sound source; S2. Perform initial speaker identification and gender recognition based on the voiceprint features of the speaker's voice, and perform sound source orientation correction and fusion to obtain the final speaker identification result; S3. Use the voice translation algorithm to combine the inquiry scenario, translate the spoken content into situational language, and output the translation result; S4. rolling sorting and integration of the contents of multiple rounds of conversations corresponding to the identities; S5. Produce templated medical records based on the collated and integrated conversation content.

2. The method for intelligent recognition of multi-round inquiry voice dialogues led by two persons according to claim 1 is characterized in that: The step S1 specifically includes: The collection surface of the sound collector is facing the inquirer and is at the same height as the inquirer's head. When the sound source direction is determined, if the sound source direction is located directly opposite the sound collector, the speaking voice comes from the inquirer. If the sound source direction is located in a direction other than directly opposite the sound collector, the speaking voice comes from the person being inquired. At the same time, a confidence threshold is set, and the confidence of the determination result is given according to the sound source direction. Only when the confidence is higher than the threshold, the identity determined by the sound source direction is combined with the initial identity recognition to perform fusion-based identity recognition to obtain the final speaker identity recognition result.

3. The method for intelligent recognition of multi-round inquiry voice dialogues led by two persons according to claim 1 is characterized in that: In step S1, the sound collector uses an array composed of multiple sound sensors, and the specific process of the sound collector performing sound source orientation determination includes: S11. Initialization and parameter setting: Set the number of acoustic sensors in the array M and their relative positions. Assume that the spatial coordinates of the M acoustic sensors are (x1, y1, z1), (x2, y2, z2), ..., (x M ,y M ,z M ); S12. Processing of received signals: Assume that the signal received by the array is x i (t), where the subscript i represents the i-th acoustic sensor and t is the time index. The signal received by each acoustic sensor is Fourier transformed to obtain the frequency domain signal X of each acoustic sensor. i (f); S13. Calculate the autocovariance matrix: Use the autocovariance matrix R to describe the correlation between each acoustic sensor. Assume that the signal X i (f) is the frequency domain signal of the i-th acoustic sensor in the array, and the covariance matrix can be expressed as: Where, Ε[·] represents the expected operation, For X i the complex conjugate of (f); S14. Calculate the beam spectrum in the target direction: Calculate the spatial spectrum, that is, for each possible sound source orientation, calculate the response of the array. Assuming that scanning is to be performed at a certain direction angle θ and pitch angle φ, the beam response in the target direction is calculated by the following formula: Where k is the wave vector, r i is the position vector of the ith acoustic sensor, represents the phase delay to the acoustic sensor, where the wave vector is calculated as: Where f is the signal frequency, θ is the azimuth of the target direction; k·r i is the phase delay from acoustic sensor i to the sound source, calculated as: The beam response value is obtained by solving the optimization problem: where P(θ,φ) is the beam response at azimuth angle θ and elevation angle φ, and a(θ,φ) H is the conjugate transpose of a(θ,φ), R -1 is the inverse of the autocovariance matrix; S15. Calculation of scanning direction and beam spectrum: Calculate the beam response P(θ,φ) at each preset direction θ and φ. The scanning direction is performed within a specific range, that is, from θ min to θ max , from φ min to φ max , to obtain the complete spatial spectrum, taking θ∈[0,π] and For each direction, calculate and store its beam response; S16. Solve the direction of the sound source: The beam response diagram P(θ) obtained through the scanning process will show the response intensity in different directions. The sound source will form a peak in the beam diagram, which corresponds to the direction of the sound source. The direction of the sound source θ source and φ source will correspond to the maximum value of P(θ): The sound source directions θ and φ obtained by sound source localization; S17. Use the arrival time difference to calculate the distance from the sound source to the array: The arrival time difference is to infer the position of the sound source by measuring the signal propagation time difference from the sound source to different sensors; assuming that the propagation time from the sound source to each sensor is t1, t2, t3, t4, t5, and t6 respectively, the time difference of each sensor receiving the signal can be known, and based on these time differences, the relationship between the distance differences is established. The distance from the sound source to the sensor is determined by the speed of sound v and time t i get: d i =v·t i Assuming that the distance from the sound source to sensor M1 is d1, and the distance to sensor M2 is d2, the distance difference between the two is: d2-d1=v·Δt 12 Where Δt 12 = t2-t1 The arrival time difference from the sound source to the sensor to M1 and M2. For the time difference between other sensors, similar equations are established. Through the distance formula in three-dimensional space, the distance difference between the sound source and each sensor is established. The least squares method is used to solve the distance d of the sound source. source ; S18. Locate the three-dimensional coordinates of the sound source: based on the direction (θ source ,φ source ) and the distance d from the sound source to the array source , calculate the three-dimensional coordinates of the sound source:

4. The method for intelligent recognition of multi-round inquiry voice dialogues led by two persons according to claim 1 is characterized in that: In step S2, the process of initially identifying the speaker's identity based on the voiceprint features of the speaker's voice includes: After the inquiry process begins, only the voiceprint features, identity, and ID number of the inquirer are known. For all subsequent detected speech, the following recognition operations are performed: A threshold for the similarity of voiceprint features is set. If the similarity between the current speaker's voiceprint features and the voiceprint features of a questioned person with a determined ID number is greater than the threshold, it is matched as the identity of the determined ID number. If the similarity between the current speaker's voiceprint features and the voiceprint features of any identity with a determined ID number is less than the threshold, it is identified as a new identity and a new ID number is assigned to it.

5. The method for intelligent recognition of multi-round inquiry voice dialogues led by two persons according to claim 4 is characterized in that: In step S2, the process of extracting the voiceprint feature of the inquirer based on the voice of the inquirer includes: The voiceprint features extracted from the voice of the inquirer reading the prescribed content before starting the inquiry every day are recorded as F1,…,F h , where F1,…,F h represent the first day to the hth day respectively; after collecting the voiceprint features from the current spoken sentence, from F1,…,F h Select the two voiceprint features with the greatest similarity to the current voiceprint features, called the most suitable voiceprint features, and record them as E1 and E2. Let the similarities between the current voiceprint features and E1 and E2 be S1 and S2 respectively. The similarity calculation formula between the current voiceprint features and the doctor is S=0.9×S1+0.1×S2.

6. The method for intelligent recognition of multi-round inquiry voice dialogues led by two persons according to claim 1 is characterized in that: In step S2, the speaker's gender recognition is performed based on a binary classification algorithm of training samples, specifically including: A batch of Chinese speech data with the gender of the speaker recorded in a question-and-answer scenario is collected. The data is annotated sentence by sentence and used as training samples. The voiceprint feature samples of men and women are extracted respectively. It is assumed that the number of training samples for men and women is r and s respectively. The category labels of men and women are defined as symmetric row vectors (1-1) and (-1 1) respectively. The samples of men and women are represented as row vectors p1,...,p r ,q1,...,q s ;make Re-order And L is the matrix composed of the category labels of the corresponding training examples; assuming that there is an approximate linear mapping relationship GY = L, this mapping relationship is the model of speaker gender recognition, then solve Y according to the following scheme: if cond(G) < 1.0e 6 , then Y=(G T G) -1 G T L, otherwise Y=(G T G+0.0001) -1 G T L; cond(G) represents the condition number of the matrix G; for the sample whose gender is to be determined, let the row vector representing its voiceprint feature be z, and let v=zY, calculate the Euclidean distance between v and the row vector (1-1) and (-1 1) respectively, and express them as d1 and d2. If d1<d2, the sample whose gender is to be determined is determined to be male, otherwise it is determined to be female.

7. The method for intelligent recognition of multi-round inquiry voice dialogues led by two persons according to claim 1 is characterized in that: In step S2, the modeling process of voiceprint feature collection includes: A batch of speech data is collected from the inquiry dialogue scene, and all the speech in sentences is marked. Different ID numbers are used to represent different speakers. The Mel-frequency coefficients of the collected speech data are modeled using a Gaussian mixture model containing five Gaussian distributions: The specific formula of the Gaussian mixture model is: w in the formula j is the weight of the jth Gaussian model, d is the dimension of the voiceprint feature x represented by the vector, μ j is the expectation of the jth Gaussian model, Φ j is the covariance matrix of the j-th Gaussian model, det(Φ j ) represents the covariance matrix Φ j The value of the determinant of .

8. The method for intelligent recognition of multi-round inquiry voice dialogues led by two persons according to claim 1 is characterized in that: In step S2, the judgment process of the sound source orientation, initial identity recognition, and gender recognition correction fusion includes: When the speaker's identity is initially identified as the inquirer, if the confidence of the speaker's gender recognition result is higher than the threshold, and the gender recognition result does not match the actual gender of the inquirer, it means that the initial identification of the speaker's identity is an incorrect result, and it should be finally identified as the inquired person; if the direction determination result of the current speaking voice is directly opposite the sound collector and the confidence is higher than the threshold, but the speaker's identity is initially identified as the inquired person, it means that it is an incorrect identification, and it should be finally identified as the inquirer; the confidence threshold of the sound source direction determination result is higher than the confidence threshold of the initial identity identification and gender recognition results.

9. The method for intelligent recognition of multi-round inquiry voice dialogues led by two persons according to claim 1 is characterized in that: In step S3, after the inquiry scenario is adapted, the situational language translation includes the construction of a homophone and near-homophone vocabulary for the inquiry scenario: The homophone and near-phonetic vocabulary list includes commonly used noun phrases with wide coverage in inquiry scenarios, word frequencies, Mandarin pronunciations, and pinyins; for the speech translation results obtained by the speech translation algorithm, the post-processing algorithm first identifies all the noun phrases therein, and then searches for each noun phrase in the constructed homophone and near-phonetic vocabulary list. If the noun phrase cannot be found in the table, the speech translation result is considered to be incorrect, and the noun phrase is replaced using the following rules: the pronunciation of the noun phrase is calculated similarly to the pronunciation of all phrases in the homophone and near-phonetic vocabulary list, and then the noun phrase in the original speech translation result is replaced with the phrase with the greatest pronunciation similarity; if the pronunciation similarity of more than one noun phrase is the greatest and equal, the noun phrase with the highest frequency is used.

10. The method for intelligent recognition of multi-round inquiry voice dialogues led by two persons according to claim 1 is characterized in that: In step S5, the medical record generated by the templated medical record generation step includes: a medical record manuscript, a medical record report draft, and a confirmed formal medical record report; the medical record manuscript is a patient conversation recorded in the actual order of the conversation; the editable medical record report draft is a medical record report that is a summary of the actual conversation content and conforms to the conventional format; the formal medical record report is a medical record report that the inquirer supplements, edits, and modifies the medical record report draft and finally submits after confirmation; The rule for subsequent modification of medical records is: when the content of a medical record is modified, the corresponding content in all subsequent medical records will be automatically modified, and the previous medical records will not be modified.

Citation Information

Cited By

  • Auxiliary turning-over equipment for nursing old people and voice control method thereof

    CN120808768A