Intelligent reminding system and method for medicine use in elderly department based on voice interaction

By extracting phoneme characteristics and pronunciation characteristics of the elderly's voice, combining multi-intention recognition and semantic integrity, the problem of fuzzy speech distortion recognition in the elderly is solved, and the accurate response of the elderly's medication reminder system is achieved.

CN120340217AInactive Publication Date: 2025-07-18CHONGQING NO 3 PEOPLES HOSPITAL
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510536913.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The speech content of the elderly is fuzzy and distorted during the speech interaction due to physiological changes such as hearing loss, unclear pronunciation, changes in accent, and cognitive decline. It is difficult for traditional speech recognition technology to accurately identify and respond to the elderly’s medication needs.

Method used

By extracting phoneme characteristics and pronunciation characteristics of the elderly under different historical dialogue rounds, pronunciation loss and semantic constraints are determined, and the speech content is corrected by combining multi-intention recognition and semantic integrity.

Benefits of technology

It improves the adaptability and accuracy of speech recognition, ensures that the medication reminder system for the elderly can accurately understand and respond to the real needs of the elderly, and reduces the fault tolerance of identifying errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340217A_ABST
    Figure CN120340217A_ABST
Patent Text Reader

Abstract

The invention provides a voice interaction-based intelligent medicine use reminding system and method for the elderly department. The voice interaction-based intelligent medicine use reminding method comprises the following steps of: extracting phoneme characteristics of a target elderly in each historical dialogue round from each interaction voice; determining the pronunciation loss of the target old person in each historical dialogue round according to the phoneme similarity between the interactive voices in each historical dialogue round and the rhythm feature of each interactive voice, and further determining the semantic constraint condition of the target old person in each historical dialogue round; performing multi-intention recognition on the interaction voice of the target old person in the current dialogue round to obtain different semantic intentions of the target old person in the current dialogue round; further determining the semantic integrity of the target old person in the current dialogue round; and correcting the voice content of the target old person in the current dialogue round through the semantic integrity and the intention confidence of each semantic intention. By adopting the scheme of the invention, the fuzzy and distorted voice content of the elderly can be corrected in the voice interaction process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice interaction technology. More specifically, this application relates to an intelligent reminder system and method for geriatric internal medicine medication based on voice interaction. Background Art

[0002] Voice interaction realizes voice communication between humans and machines through speech recognition and natural language processing. The core of voice interaction technology lies in speech recognition, semantic understanding, and response generation. Among them, speech recognition involves the extraction, feature analysis, and matching of audio signals, semantic understanding is the in-depth intention parsing of the input speech, and response generation is to provide accurate feedback according to user needs. With the development of technologies such as machine learning and deep learning, the recognition accuracy, response speed, and intelligence level of voice interaction systems have been significantly improved, and they can adapt to changes in different languages, accents, and environmental noises. Voice interaction technology is widely used in many fields such as smart homes, in-vehicle systems, and voice assistants, providing users with a more convenient and efficient operation experience.

[0003] With the advent of an aging society, the health management of the elderly has gradually attracted wide attention. Especially in the management of the elderly's medication, speech recognition technology is used to help the elderly take medicine on time and avoid health problems caused by forgetting or taking the wrong medicine. However, with the physiological changes of the elderly, such as hearing loss, unclear pronunciation, accent changes, cognitive decline, etc., their voice content becomes blurred and distorted. This blurred and distorted phenomenon is particularly obvious in the process of voice interaction, especially the deviations in phoneme pronunciation, intonation, speech rate, etc. This change poses a great challenge to traditional speech recognition technology. Therefore, how to correct the blurred and distorted voice content of the elderly in the process of voice interaction has become a difficult problem faced by the industry. Summary of the Invention

[0004] This application provides an intelligent reminder system and method for geriatric internal medicine medication based on voice interaction, which can correct the blurred and distorted voice content of the elderly in the process of voice interaction.

[0005] In a first aspect, this application provides a speech recognition correction method for use in a geriatric internal medicine medication intelligent reminder system to perform speech recognition correction during voice interaction, including the following steps: Obtain the interaction voice of the target elderly during medication reminders in different historical conversation turns, and extract the phoneme features of the target elderly in each historical conversation turn from each interaction voice; Determine the pronunciation loss of the target elderly in each historical dialogue turn according to the phoneme similarity between interactive voices in each historical dialogue turn and the prosodic features of each interactive voice, and determine the semantic constraint conditions of the target elderly in each historical dialogue turn through all phoneme features and the pronunciation loss of each historical dialogue turn; Perform multi-intent recognition on the interactive voice of the target elderly in the current dialogue turn to obtain different semantic intents of the target elderly in the current dialogue turn; Determine the semantic integrity of the target elderly in the current dialogue turn according to all semantic intents and the semantic constraint conditions in each historical dialogue turn; Correct the speech content of the target elderly in the current dialogue turn through the semantic integrity and the intent confidence of each semantic intent.

[0006] In some embodiments, extracting the phoneme features of the target elderly in each historical dialogue turn from each interactive voice specifically includes: Perform phoneme recognition on each interactive voice to obtain the phoneme labels of each audio frame within each interactive voice; Determine the phoneme features of the target elderly in each historical dialogue turn according to the phoneme labels of each audio frame within each interactive voice.

[0007] In some embodiments, determining the pronunciation loss of the target elderly in each historical dialogue turn according to the phoneme similarity between interactive voices in each historical dialogue turn and the prosodic features of each interactive voice specifically includes: Determine the phoneme similarity between interactive voices in each historical dialogue turn; Determine the intonation curve of the target elderly in each historical dialogue turn according to the prosodic features of each interactive voice; Determine the pronunciation loss of the target elderly in each historical dialogue turn through the phoneme similarity and the intonation curve of the target elderly in each historical dialogue turn.

[0008] In some embodiments, determining the semantic constraint conditions of the target elderly in each historical dialogue turn through all phoneme features and the pronunciation loss of each historical dialogue turn specifically includes: Determine the semantic deviation coefficient of the target elderly during medication reminder in each historical dialogue turn according to all phoneme features; Perform constraint analysis on the semantic deviation coefficient of the target elderly during medication reminder in each historical dialogue turn through all pronunciation losses to obtain the semantic constraint conditions of the target elderly in each historical dialogue turn.

[0009] In some embodiments, performing multi-intent recognition on the interactive voice of the target elderly in the current dialogue turn to obtain different semantic intents of the target elderly in the current dialogue turn specifically includes: Convert the interactive speech of the target elderly person in the current conversation turn into text to obtain the speech content; Extract different semantic intents of the target elderly person in the current conversation turn from the speech content through a pre-trained multi-intent recognition model.

[0010] In some embodiments, determining the semantic integrity of the target elderly person in the current conversation turn according to all semantic intents and the semantic constraint conditions in each historical conversation turn specifically includes: Determine the semantic graph of the target elderly person in the current conversation turn according to all semantic intents and the semantic constraint conditions in each historical conversation turn; Determine the semantic change amount of the target elderly person in the current conversation turn through the intent confidence of each semantic intent; Determine the semantic integrity of the target elderly person in the current conversation turn according to the semantic graph and the semantic change amount.

[0011] In some embodiments, correcting the speech content of the target elderly person in the current conversation turn through the semantic integrity and the intent confidence of each semantic intent specifically includes: Obtain the speech content of the interactive speech of the target elderly person in the current conversation turn; Structurally complete the speech content according to the semantic integrity to obtain semantically completed content; Semantically correct the semantically completed content through the intent confidence of each semantic intent to obtain corrected speech content.

[0012] In a second aspect, the present application provides an intelligent reminder system for geriatric internal medicine medication based on voice interaction. The system includes a voice recognition and correction unit, and the voice recognition and correction unit includes: An acquisition module, configured to acquire the interactive speech of the target elderly person during medication reminder in different historical conversation turns, and extract the phoneme features of the target elderly person in each historical conversation turn from each interactive speech; A processing module, configured to determine the pronunciation loss of the target elderly person in each historical conversation turn according to the phoneme similarity between the interactive speeches in each historical conversation turn and the prosodic features of each interactive speech, and determine the semantic constraint conditions of the target elderly person in each historical conversation turn through all phoneme features and the pronunciation loss of each historical conversation turn; The processing module is further configured to perform multi-intent recognition on the interactive speech of the target elderly person in the current conversation turn to obtain different semantic intents of the target elderly person in the current conversation turn; The processing module is further configured to determine the semantic integrity of the target elderly person in the current conversation turn according to all semantic intents and the semantic constraint conditions in each historical conversation turn; An execution module, configured to correct the speech content of the target elderly person in the current conversation turn based on the semantic integrity and the intention confidence of each semantic intention.

[0013] In a third aspect, the present application provides a computer device, which includes a memory and a processor. The memory stores code, and the processor is configured to obtain the code and execute the above-mentioned speech recognition correction method.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned speech recognition correction method is implemented.

[0015] The technical solutions provided by the disclosed embodiments of the present application have the following beneficial effects: In the intelligent reminder system and method for geriatric internal medicine medication based on voice interaction provided by the present application, the interactive speech during medication reminder of the target elderly person in different historical conversation turns is obtained, and the phoneme features of the target elderly person in each historical conversation turn are extracted from each interactive speech; the pronunciation loss of the target elderly person in each historical conversation turn is determined according to the phoneme similarity between the interactive speeches in each historical conversation turn and the prosody features of each interactive speech, and the semantic constraint conditions of the target elderly person in each historical conversation turn are determined through all the phoneme features and the pronunciation loss of each historical conversation turn; multi-intention recognition is performed on the interactive speech of the target elderly person in the current conversation turn to obtain different semantic intentions of the target elderly person in the current conversation turn; the semantic integrity of the target elderly person in the current conversation turn is determined according to all the semantic intentions and the semantic constraint conditions in each historical conversation turn; the speech content of the target elderly person in the current conversation turn is corrected based on the semantic integrity and the intention confidence of each semantic intention.

[0016] It can be seen that in this application, the semantic integrity of the target elderly person in the current conversation turn can be determined according to all semantic intents and the semantic constraint conditions in each historical conversation turn. First, the phoneme features of the target elderly person in different historical conversation turns are extracted. The phoneme features can quantify the personalized pronunciation features of the target elderly person in different historical conversation turns, so as to capture the pronunciation deviation of the target elderly person, enable the system to adapt to the speech pattern of the target elderly person, and improve the understanding ability of their speech. Second, by calculating the phoneme similarity of the interactive speech in each historical conversation turn, the system can identify which phonemes of the target elderly person are missing, weakened or deformed, and further analyze the rules of these pronunciation changes in combination with prosodic features. For example, if some consonants of the elderly person are pronounced weakly, or some phonemes are elongated or swallowed due to the decrease in speech speed, the system can detect and record these deviations to form a personalized pronunciation loss model, so that the system can automatically compensate for the possibly lost phoneme information when recognizing the speech of the elderly person, and improve the adaptability and accuracy of speech recognition. Furthermore, using the phoneme features and pronunciation loss, the semantic constraint conditions of the target elderly person are constructed to accurately correct the recognition errors caused by fuzzy and distorted pronunciation. Through the semantic constraint conditions, the system can combine the previous pronunciation rules and historical conversation content of the elderly person, limit the reasonable semantic range, and automatically correct the possible recognition errors. Then, through multi-intent recognition, the system can analyze the speech content from different angles and generate multiple possible semantic candidates. Even if some phonemes cause recognition deviations due to fuzzy distortion, the system can still perform intelligent correction based on multiple possible semantic paths to ensure that the medication reminder system accurately understands the real needs of the elderly. Further, through semantic integrity calculation, the system can evaluate the currently recognized semantic information, judge whether it lacks key content, and supplement it based on historical semantic constraint conditions. This process effectively solves the recognition problem caused by the change of the elderly person's speech features, improves the fault tolerance and response accuracy of the system. Finally, the speech content of the target elderly person in the current conversation turn is corrected through the semantic integrity and the intent confidence of each semantic intent. In summary, the solution of this application can correct the fuzzy and distorted speech content of the elderly person during the speech interaction process. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is an exemplary flowchart of a speech recognition correction method according to some embodiments of the present application; Figure 2 is a schematic flowchart of determining pronunciation loss according to some embodiments of the present application; Figure 3 is a schematic flowchart of determining semantic integrity according to some embodiments of the present application; Figure 4 is a schematic structural diagram of a speech recognition correction unit according to some embodiments of the present application; Figure 5 It is a schematic structural diagram of a computer device for implementing a voice recognition correction method according to some embodiments of the present application. Detailed implementation manners

[0018] To better understand the technical solutions of the present application, the technical solutions of the present application will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0019] Refer to Figure 1 , this figure is an exemplary flowchart of a voice recognition correction method according to some embodiments of the present application. The voice recognition correction method 100 mainly includes the following steps: In step 101, obtain the interactive voice of the target elderly person during medication reminders in different historical conversation turns, and extract the phoneme features of the target elderly person in each historical conversation turn from each interactive voice.

[0020] Specifically, the interactive voice of the target elderly person during medication reminders in different historical conversation turns is obtained through an intelligent medication interaction device (such as an intelligent medicine box).

[0021] It should be noted that the interactive voice described in the present application refers to the dialogue voice record during the voice interaction process, and each interactive voice corresponds to a historical conversation turn.

[0022] In some embodiments, the extraction of the phoneme features of the target elderly person in each historical conversation turn from each interactive voice can be implemented by the following steps: Perform phoneme recognition on each interactive voice to obtain the phoneme labels of each audio frame in each interactive voice; Determine the phoneme features of the target elderly person in each historical conversation turn according to the phoneme labels of each audio frame in each interactive voice.

[0023] In specific implementation, for each interactive voice, phoneme recognition is performed to obtain the phoneme labels of each audio frame within each interactive voice, which can be implemented in the following manner: for each interactive voice, an audio processing library (such as the torchaudio library) is used to read the interactive voice, preprocess the interactive voice, and use the data obtained after preprocessing as audio preprocessing data. Among them, the preprocessing steps include adjusting the sampling rate of the interactive voice (such as 16 kHz), converting the interactive voice to mono, and normalizing the interactive voice. Then, the audio preprocessing data is input into a pre-trained acoustic model (such as the Wav2Vec 2.0 model), and the output result of the speech model is used as an audio tensor. The audio tensor consists of the phoneme probability distributions of multiple audio frames. Furthermore, the connectionist temporal classification (CTC) algorithm is used to perform sequence labeling on the audio tensor, so as to obtain the pronunciation marks of each audio frame within the interactive voice, and all the obtained pronunciation marks are used as phoneme labels. In other embodiments, other methods can also be used for implementation, which will not be elaborated here.

[0024] It should be noted that the phoneme labels described in this application represent the marks of the smallest pronunciation units in the speech signal.

[0025] In specific implementation, to determine the phoneme features of the target elderly person in each historical conversation turn according to the phoneme labels of each audio frame within each interactive voice, it can be implemented in the following manner: select an interactive voice as the selected interactive voice, calculate the frequency of each phoneme label appearing in the selected interactive voice, and use all the obtained frequencies as the phoneme distribution frequencies. Then, sort all the phoneme distribution frequencies in descending order, and use the obtained sequence as the phoneme features of the target elderly person in the historical conversation turn corresponding to the selected interactive voice. Continue to determine the phoneme features of the target elderly person in the historical conversation turns corresponding to the remaining interactive voices. In other embodiments, other methods can also be used for implementation, which will not be elaborated here.

[0026] It should be noted that the phoneme features described in this application represent the distribution features of different pronunciation phonemes of the target elderly person in the interactive voice. Through the phoneme features, the pronunciation habits of the target elderly person for different phonemes can be reflected. In addition, each phoneme feature corresponds to a historical conversation turn.

[0027] In addition, it should be noted that by counting the occurrence frequencies of phoneme tags in the target elderly's interactive speech, the usage of different phonemes by the target elderly is quantified, thereby constructing their personalized phoneme distribution pattern. Calculating the phoneme distribution frequency can reflect the pronunciation preferences of the target elderly. Sorting all phoneme frequencies in descending order helps to remove the interference of speech content on phoneme statistics, making the phoneme features more stable and comparable. Traversing all historical interactive speeches ensures that the phoneme features cover the long-term pronunciation habits of the target elderly, making them representative and complete.

[0028] In step 102, according to the phoneme similarity between interactive speeches in each historical dialogue turn and the prosodic features of each interactive speech, the pronunciation loss of the target elderly in each historical dialogue turn is determined. Based on all the phoneme features and the pronunciation loss in each historical dialogue turn, the semantic constraint conditions of the target elderly in each historical dialogue turn are determined.

[0029] In some embodiments, as Figure 2 shown, this figure is a schematic flowchart of determining pronunciation loss in some embodiments of the present application. In this embodiment, determining the pronunciation loss of the target elderly in each historical dialogue turn according to the phoneme similarity between interactive speeches in each historical dialogue turn and the prosodic features of each interactive speech can be implemented by the following steps: Determine the phoneme similarity between interactive speeches in each historical dialogue turn; According to the prosodic features of each interactive speech, determine the intonation curve of the target elderly in each historical dialogue turn; Determine the pronunciation loss of the target elderly in each historical dialogue turn through the phoneme similarity and the intonation curve of the target elderly in each historical dialogue turn.

[0030] Specifically, when implementing, the phoneme similarity between interactive speeches in each historical dialogue turn can be implemented in the following way, that is: First, obtain the phoneme tags of each audio frame in each interactive speech, then use one-hot encoding to vectorize the phoneme tags of each audio frame in each interactive speech, and take the obtained vectors as the phoneme tag vectors of each interactive speech respectively. Then, calculate the cosine similarity between every two phoneme tag vectors, and take the average value of all cosine similarities as the phoneme similarity between interactive speeches in each historical dialogue turn. In other embodiments, other methods can also be used for implementation, which is not limited here.

[0031] It should be noted that the phoneme similarity described in the present application represents the similarity degree of pronunciation phonemes of the target elderly between each historical dialogue turn.

[0032] In addition, it should be noted that the prosodic features described in this application represent the features describing the pitch distribution of the voices in the interactive speech. First, the interactive speech can be converted into a frequency spectrum using the Fourier transform, and then the pitch of each time frame in the frequency spectrum can be extracted using a pitch extraction algorithm (such as the YIN algorithm), and the set of all pitches is used as the prosodic features of the interactive speech.

[0033] When specifically implemented, determining the intonation curve of the target elderly person in each historical conversation turn according to the prosodic features of each interactive speech can be achieved in the following way, that is: for each interactive speech, arrange the pitches of each time frame in the prosodic features of the interactive speech in chronological order, then plot them into a curve, and use the obtained curve as the intonation curve of the target elderly person in the corresponding historical conversation turn of the interactive speech, so as to obtain the intonation curve of the target elderly person in each historical conversation turn. In other embodiments, other methods can also be used to implement this, which will not be elaborated here.

[0034] It should be noted that the intonation curve described in this application represents the trajectory curve of the pitch changing with time. In addition, each intonation curve corresponds to a historical conversation turn.

[0035] When specifically implemented, determining the pronunciation loss of the target elderly person in each historical conversation turn through the phoneme similarity and the intonation curve of the target elderly person in each historical conversation turn can be achieved in the following way, that is: first, for each intonation curve, use a spline fitting algorithm (such as the B-spline fitting algorithm) to fit the intonation curve to obtain a fitted curve, then use the difference method to calculate the change rate at each position on the fitted curve, and use the obtained change rates as the pitch change rates. Further, standardize all the pitch change rates, then take the difference of all the standardized values, and then add the mean of all the values obtained from the difference to the phoneme similarity, and use the added value as the pronunciation loss of the target elderly person in the historical conversation turn corresponding to the intonation curve, so as to obtain the pronunciation loss of the target elderly person in each historical conversation turn. In other embodiments, other methods can also be used to implement this, which is not limited here.

[0036] It should be noted that the pronunciation loss described in this application represents the parameter value for evaluating the pronunciation stability of the target elderly person during the speech interaction process.

[0037] In addition, it should be noted that in this application, by fitting the intonation curve, calculating the pitch change rate and the volatility of the change, the instability or difficulty of the speech can be evaluated. Combining with the phoneme similarity, the calculation of the pronunciation loss takes into account the accuracy and fluency of the speech, quantifies the pronunciation loss of the target elderly person in the interaction, and helps to evaluate the decline of their speech clarity and expression ability.

[0038] In some embodiments, the semantic constraint conditions of the target elderly person in each historical conversation turn can be determined by all phoneme features and the pronunciation loss of each historical conversation turn through the following steps: Determine the semantic deviation coefficient of the target elderly person during medication reminder in each historical conversation turn according to all phoneme features; Conduct constraint analysis on the semantic deviation coefficient of the target elderly person during medication reminder in each historical conversation turn through all pronunciation losses to obtain the semantic constraint conditions of the target elderly person in each historical conversation turn.

[0039] Specifically, when implemented, determining the semantic deviation coefficient of the target elderly person during medication reminder in each historical conversation turn according to all phoneme features is achieved in the following manner, that is: First, obtain the reference phoneme distribution from the intelligent reminder system for geriatric internal medicine medication based on voice interaction. Among them, the reference phoneme distribution represents the statistical representation of all phonemes in the acoustic feature space and temporal distribution under standard pronunciation conditions for the healthy elderly group. The reference phoneme distribution consists of the mean vector and covariance matrix of the phoneme probability distribution. Then, for each phoneme feature, use the Mahalanobis distance to quantify the deviation degree between the phoneme feature and the reference phoneme distribution, and use the quantified value as the semantic deviation coefficient of the target elderly person during medication reminder in the historical conversation turn corresponding to the phoneme feature, so as to obtain the semantic deviation coefficient of the target elderly person during medication reminder in each historical conversation turn. In other embodiments, other methods can also be used for implementation, which is not limited here.

[0040] It should be noted that the semantic deviation coefficient described in this application is a parameter representing the degree of semantic deviation of the target elderly person during medication reminder in the voice interaction process. Among them, each semantic deviation coefficient corresponds to a historical conversation turn.

[0041] In addition, it should also be noted that the Mahalanobis distance can accurately measure the degree of deviation in the multi-dimensional phoneme feature space. The phoneme features of the target elderly person are the phoneme distribution formed in their actual conversations, which may deviate due to factors such as individual pronunciation habits and health conditions. The Mahalanobis distance takes into account the correlation between the dimensions of the phoneme features and is more suitable for quantifying the differences in high-dimensional feature data compared to the Euclidean distance. When the Mahalanobis distance is large, it indicates that the phoneme features of the target elderly person deviate significantly from the reference phoneme distribution, that is, the pronunciation pattern has a large difference from the standard pronunciation, which in turn affects the accuracy and comprehensibility of semantic expression. Therefore, this distance can be used as a quantization standard for the semantic deviation coefficient to characterize the degree of deviation of the target elderly person's voice expression in the medication reminder scenario.

[0042] In specific implementation, by using all pronunciation losses to perform constraint analysis on the semantic deviation coefficient during medication reminder for the target elderly in each historical conversation turn, the semantic constraint conditions for the target elderly in each historical conversation turn can be implemented in the following manner, that is: First, perform standardization processing (such as Min-Max standardization) on all semantic deviation coefficients and all pronunciation losses, so as to obtain multiple standardized semantic deviation coefficients and multiple standardized pronunciation losses. Calculate the exponential function with the natural logarithm e as the base for each standardized pronunciation loss, and then use the obtained values after calculating the exponential function with the natural logarithm e as the base as constraint weights. Among them, each historical conversation turn corresponds to a constraint weight. Then, for each historical conversation turn, multiply the semantic deviation coefficient during medication reminder for the target elderly in the historical conversation turn by the constraint weight corresponding to the historical conversation turn, and use the multiplied value as the semantic constraint condition for the target elderly in the historical conversation turn, thereby obtaining the semantic constraint conditions for the target elderly in each historical conversation turn. In other embodiments, other methods can also be used for implementation, which will not be elaborated here.

[0043] It should be noted that the semantic constraint conditions described in this application represent the condition parameters used to correct the semantic ambiguity and distortion of the target elderly during voice interaction.

[0044] In addition, it should also be noted that since semantic communication in this application depends on clear and accurate speech, the increase in pronunciation loss will exacerbate semantic deviation, thereby affecting the interaction effect. Therefore, by standardizing pronunciation loss and semantic deviation coefficients, enabling them to be compared on the same scale, and introducing a negative exponential function to calculate the constraint weight, the stronger the correction of semantic deviation in the turn with a larger pronunciation loss, so as to obtain dynamically adjusted semantic constraint conditions. This process ensures that in voice interaction, even if there is a pronunciation deviation, through reasonable semantic adjustment, the intelligibility and interaction accuracy of the voice content can be improved.

[0045] In step 103, perform multi-intent recognition on the interaction voice of the target elderly in the current conversation turn to obtain different semantic intents of the target elderly in the current conversation turn.

[0046] In some embodiments, performing multi-intent recognition on the interaction voice of the target elderly in the current conversation turn to obtain different semantic intents of the target elderly in the current conversation turn can be implemented by the following steps: Perform text conversion on the interaction voice of the target elderly in the current conversation turn to obtain the voice content; Extract different semantic intents of the target elderly in the current conversation turn from the voice content through a pre-trained multi-intent recognition model.

[0047] In specific implementation, the interactive speech of the target elderly person in the current conversation turn is converted into text. The following method can be used to obtain the interactive speech text, that is: the intelligent reminder system for geriatric internal medicine medication based on voice interaction calls the iFLYTEK API of iFlytek to convert the interactive speech of the target elderly person in the current conversation turn into text, and uses the obtained text content as the speech content. In other embodiments, other methods can also be used to implement this, which will not be elaborated here.

[0048] It should be noted that the pre-trained multi-intent recognition model in this application can adopt the BERT model. The BERT model is a pre-trained language model based on the Transformer architecture. The full name of BERT is Bidirectional Encoder Representations from Transformers. It captures complex semantic relationships in the text through its bidirectional context understanding ability. In the multi-intent recognition task, BERT performs excellently and can extract multiple intent labels from the input text at the same time. Its core advantage lies in BERT's bidirectional self-attention mechanism, which can fully understand the context relationship of each word in the context and capture multiple intents in the text. Through the fine-tuning process, BERT can be optimized in the multi-label classification task to identify multiple possible intent labels in the text. Each intent label is predicted through an independent binary classification task. The output of the model is the probability value of whether each label exists. Such a mechanism enables BERT to handle multiple co-occurring intents, thereby obtaining accurate recognition results in multi-intent recognition.

[0049] In specific implementation, the following method can be used to extract different semantic intents of the target elderly person in the current conversation turn from the speech content through the pre-trained multi-intent recognition model, that is: input the interactive speech text into the multi-intent recognition model, and then the multi-intent recognition model outputs multiple labels, and all the obtained labels are used as the semantic intents of the target elderly person in the current conversation turn. Among them, the semantic intent represents the specific semantic label of the target elderly person regarding medication information in the current conversation turn. In other embodiments, other methods can also be used to implement this, which is not limited here.

[0050] In step 104, the semantic integrity of the target elderly person in the current conversation turn is determined according to all the semantic intents and the semantic constraint conditions in each historical conversation turn.

[0051] In some embodiments, with reference to Figure 3 As shown, this figure is a schematic flowchart of determining semantic integrity in some embodiments of this application. In this embodiment, the following steps can be used to determine the semantic integrity of the target elderly person in the current conversation turn according to all the semantic intents and the semantic constraint conditions in each historical conversation turn: First, in step 1041, determine the semantic graph of the target elderly person in the current conversation turn according to all semantic intents and the semantic constraint conditions in each historical conversation turn. Secondly, in step 1042, determine the semantic change amount of the target elderly person in the current conversation turn through the intent confidence of each semantic intent. Then, in step 1043, determine the semantic integrity of the target elderly person in the current conversation turn according to the semantic graph and the semantic change amount.

[0052] In specific implementation, determining the semantic graph of the target elderly person in the current conversation turn according to all semantic intents and the semantic constraint conditions in each historical conversation turn can be implemented in the following way: First, for each pair of semantic intents, the cosine similarity can be used to measure the correlation relationship between the two semantic intents, where the weight of the correlation relationship is the value of the cosine similarity. Then, calculate the average value of all semantic constraint conditions and use the obtained average value as the constraint parameter for each correlation relationship, so as to update the weight of each correlation relationship. Furthermore, take each semantic intent as a node, construct a graph according to the correlation relationship between every two semantic intents, and use the obtained graph as the semantic graph of the target elderly person in the current conversation turn. In other embodiments, other methods can also be used for implementation, which will not be elaborated here.

[0053] It should be noted that the semantic graph in this application represents a graph composed of different semantic intents and their semantic association structures of the target elderly person in the current conversation turn, where nodes represent semantic intents and edges represent the logical connections between semantic intents.

[0054] It should be noted that the intent confidence in this application represents the credibility of the recognition result of the semantic intent, and the intent confidence of each semantic intent can be output through the Softmax function of the multi-intent recognition model.

[0055] In specific implementation, determining the semantic change amount of the target elderly person in the current conversation turn through the intent confidence of each semantic intent can be implemented in the following way: Calculate the standard deviation of all intent confidences and use the obtained standard deviation as the semantic change amount of the target elderly person in the current conversation turn. In other embodiments, other methods can also be used for implementation, which will not be elaborated here.

[0056] It should be noted that the semantic change amount in this application is an index of the semantic change degree of the speech content of the target elderly person in the current conversation turn.

[0057] When specifically implemented, determining the semantic integrity of the target elderly person in the current conversation turn according to the semantic graph and the semantic change amount can be implemented in the following manner, that is: multiply the weight of each edge in the semantic graph by the semantic change amount, then sum all the multiplied values, and use the summed value as the semantic integrity of the interactive speech in the current conversation turn. In other embodiments, other methods can also be used for implementation, which is not limited here.

[0058] It should be noted that the semantic integrity described in this application represents the completeness of the semantic expression content of the target elderly person in the current conversation turn.

[0059] In step 105, the speech content of the target elderly person in the current conversation turn is corrected according to the semantic integrity and the intention confidence of each semantic intention.

[0060] In some embodiments, correcting the speech content of the target elderly person in the current conversation turn according to the semantic integrity and the intention confidence of each semantic intention can be implemented by the following steps: Obtain the speech content of the interactive speech of the target elderly person in the current conversation turn; Structurally complete the speech content according to the semantic integrity to obtain semantic completion content; Semantically correct the semantic completion content through the intention confidence of each semantic intention to obtain corrected speech content.

[0061] When specifically implemented, structurally completing the speech content according to the semantic integrity to obtain semantic completion content can be implemented in the following manner, that is: First, the tokenizer of the BERT model converts the speech content into a Token sequence, then expands the semantic integrity into a weight vector of the same length as the Token sequence, further multiplies the obtained weight vector by the Token sequence, and uses the multiplied result as the input embedding of the BERT model. Furthermore, calculate the area to be completed through the Masked Language Model task of BERT, insert MASK tokens at these positions, encode the word embeddings, word embeddings, and paragraph embeddings of the BERT model. Then, BERT calculates the relationship between each Token in the Token sequence through multi-head self-attention to make the BERT model focus on the information missing area. Finally, BERT calculates the most likely word to complete at the MASK position in the Softmax classification layer, fills the predicted word into the original speech content, thereby generating completion content, and uses the obtained completed speech content as the semantic completion content. In other embodiments, other methods can also be used for implementation, which is not limited here.

[0062] It should be noted that the semantic completion content described in this application represents the speech content obtained after structurally completing the speech content.

[0063] In specific implementation, semantic correction is performed on the semantic completion content through the intent confidence of each semantic intent. The corrected speech content can be implemented in the following manner, that is: First, use the tokenizer of the BERT model to convert the semantic completion content into a Token sequence. Then, create an initial vector with the same length as the Token sequence, and the values of all elements in the initial vector can be set to a relatively small default value, such as 0.1. Further, traverse each Token in the Token sequence, and use the named entity recognition model to identify the relevance of each Token to each semantic intent. If a Token is relevant to a certain semantic intent, assign the intent confidence of this semantic intent to the element corresponding to the position in the initial vector. Then, multiply the updated initial vector by the word embedding of the Token sequence to update the word embedding. Next, input the updated word embedding into the BERT model. Through the multi-layer self-attention mechanism of its Transformer structure, the BERT model performs in-depth reasoning and context correction on the semantic completion content. Then, use the output result of the BERT model as the corrected speech content. In other embodiments, other methods can also be used for implementation, which will not be elaborated here.

[0064] It should be noted that when the present application performs semantic correction, the intent confidence reflects the degree of certainty and importance of a specific semantic intent, indicating which intents are key or most need to be considered preferentially in the current conversation. The self-attention mechanism of BERT captures the relationships between various words in the text and determines which words should receive more attention during the reasoning process. When multiplying the mean of the intent confidence by the attention weights of BERT, it is equivalent to strengthening the influence of the semantic intent by adjusting the attention distribution. This process can improve the model's sensitivity to key information, enhance the accuracy and context consistency of semantic completion, and thus correct irrelevant or incorrect content.

[0065] In addition, on the other hand of the present application, in some embodiments, the present application provides an intelligent reminder system for geriatric internal medicine medication based on voice interaction. The system includes a voice recognition and correction unit. Refer to Figure 4 , which is a schematic structural diagram of the voice recognition and correction unit according to some embodiments of the present application. The voice recognition and correction unit 400 includes: an acquisition module 401, a processing module 402, and an execution module 403, which are described as follows: The acquisition module 401. In the present application, the acquisition module 401 is mainly used to acquire the interactive voice of the target elderly person during medication reminders in different historical conversation turns, and extract the phoneme features of the target elderly person in each historical conversation turn from each interactive voice. The processing module 402. In this application, the processing module 402 is used to determine the pronunciation loss of the target elderly person in each historical conversation turn based on the phoneme similarity between the interactive voices in each historical conversation turn and the prosodic features of each interactive voice, and determine the semantic constraint conditions of the target elderly person in each historical conversation turn through all the phoneme features and the pronunciation loss in each historical conversation turn; It should be noted that the processing module 402 in this application is also used to perform multi-intent recognition on the interactive voice of the target elderly person in the current conversation turn to obtain different semantic intents of the target elderly person in the current conversation turn; In addition, the processing module 402 in this application is also used to determine the semantic integrity of the target elderly person in the current conversation turn according to all the semantic intents and the semantic constraint conditions in each historical conversation turn; The execution module 403. In this application, the execution module 403 is mainly used to correct the speech content of the target elderly person in the current conversation turn through the semantic integrity and the intent confidence of each semantic intent.

[0066] In addition, this application also provides a computer device, which includes a memory and a processor. The memory stores code, and the processor is configured to obtain the code and execute the above speech recognition correction method.

[0067] In some embodiments, refer to Figure 5 , this figure is a schematic structural diagram of a computer device for implementing the speech recognition correction method according to some embodiments of this application. The speech recognition correction method in the above embodiments can be implemented by Figure 5 the computer device shown. The computer device 500 includes at least one processor 501, a communication bus 502, a memory 503, and at least one communication interface 504.

[0068] The processor 501 can be a general-purpose central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more for controlling the execution of the speech recognition correction method in this application.

[0069] The communication bus 502 can be used to transmit information between the above components.

[0070] The memory 503 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disks, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 503 can exist independently and be connected to the processor 501 through the communication bus 502. The memory 503 can also be integrated with the processor 501.

[0071] Among them, the memory 503 is used to store the program code for executing the solution of this application and is controlled by the processor 501 to execute. The processor 501 is used to execute the program code stored in the memory 503. The program code can include one or more software modules. The method described in the above method embodiments can be implemented by one or more software modules in the program code of the processor 501 and the memory 503.

[0072] The communication interface 504 uses any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.

[0073] In a specific implementation, as an embodiment, the computer device can include multiple processors, and each of these processors can be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0074] The computer device described above may be a general-purpose computer device or a special-purpose computer device. In a specific implementation, the computer device may be a desktop computer, a laptop computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. The embodiments of the present application do not limit the type of the computer device.

[0075] In addition, the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned speech recognition and correction method is implemented.

[0076] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0077] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A voice recognition correction method for correcting voice recognition during the voice interaction process of an intelligent reminder system for geriatric internal medicine medication, characterized in that, The steps are as follows: Obtain the interactive speech of the target elderly person during medication reminders in different historical conversation turns, and extract the phoneme features of the target elderly person in each historical conversation turn from each interactive speech; Determine the pronunciation loss of the target elderly person in each historical conversation turn according to the phoneme similarity between the interactive speeches in each historical conversation turn and the prosody features of each interactive speech, and determine the semantic constraint conditions of the target elderly person in each historical conversation turn through all the phoneme features and the pronunciation loss of each historical conversation turn; Perform multi-intent recognition on the interactive speech of the target elderly person in the current conversation turn to obtain different semantic intents of the target elderly person in the current conversation turn; Determine the semantic integrity of the target elderly person in the current conversation turn according to all the semantic intents and the semantic constraint conditions in each historical conversation turn; Correct the speech content of the target elderly person in the current conversation turn through the semantic integrity and the intent confidence of each semantic intent.

2. The method according to claim 1, wherein Specifically, extracting the phoneme features of the target elderly person in each historical conversation turn from each interactive speech includes: Perform phoneme recognition on each interactive speech to obtain the phoneme labels of each audio frame in each interactive speech; Determine the phoneme features of the target elderly person in each historical conversation turn according to the phoneme labels of each audio frame in each interactive speech.

3. The method according to claim 1, wherein Specifically, determining the pronunciation loss of the target elderly person in each historical conversation turn according to the phoneme similarity between the interactive speeches in each historical conversation turn and the prosody features of each interactive speech includes: Determine the phoneme similarity between the interactive speeches in each historical conversation turn; Determine the intonation curve of the target elderly person in each historical conversation turn according to the prosody features of each interactive speech; Determine the pronunciation loss of the target elderly person in each historical conversation turn through the phoneme similarity and the intonation curve of the target elderly person in each historical conversation turn.

4. The method according to claim 1, wherein Specifically, determining the semantic constraint conditions of the target elderly person in each historical conversation turn through all the phoneme features and the pronunciation loss of each historical conversation turn includes: Determine the semantic deviation coefficient of the target elderly person during medication reminder in each historical conversation turn according to all the phoneme features; Perform constraint analysis on the semantic deviation coefficient of the target elderly person during medication reminder in each historical conversation turn through all the pronunciation losses to obtain the semantic constraint conditions of the target elderly person in each historical conversation turn.

5. The method according to claim 1, wherein Specifically, performing multi-intent recognition on the interactive speech of the target elderly person in the current conversation turn to obtain different semantic intents of the target elderly person in the current conversation turn includes: Perform text conversion on the interactive speech of the target elderly person in the current conversation turn to obtain the speech content; Extract different semantic intents of the target elderly person in the current conversation turn from the speech content through a pre-trained multi-intent recognition model.

6. The method according to claim 1, wherein Specifically, determining the semantic integrity of the target elderly person in the current conversation turn according to all the semantic intents and the semantic constraint conditions in each historical conversation turn includes: Determine the semantic graph of the target elderly person in the current conversation turn according to all the semantic intents and the semantic constraint conditions in each historical conversation turn; Determine the semantic change amount of the target elderly person in the current conversation turn through the intention confidence of each semantic intention; Determine the semantic integrity of the target elderly person in the current conversation turn according to the semantic map and the semantic change amount.

7. The method according to claim 1, wherein Correcting the speech content of the target elderly person in the current conversation turn through the semantic integrity and the intention confidence of each semantic intention specifically includes: Obtain the speech content of the interactive speech of the target elderly person in the current conversation turn; Structurally complete the speech content according to the semantic integrity to obtain semantically completed content; Semantically correct the semantically completed content through the intention confidence of each semantic intention to obtain corrected speech content.

8. An intelligent reminder system for geriatric internal medicine medications based on voice interaction, the system includes a voice recognition and correction unit, characterized in that, The speech recognition correction unit includes: An acquisition module, configured to acquire the interactive speech of the target elderly person during medication reminder in different historical conversation turns, and extract the phoneme features of the target elderly person in each historical conversation turn from each interactive speech; A processing module, configured to determine the pronunciation loss of the target elderly person in each historical conversation turn according to the phoneme similarity between the interactive speeches in each historical conversation turn and the prosodic features of each interactive speech, and determine the semantic constraint conditions of the target elderly person in each historical conversation turn through all the phoneme features and the pronunciation loss of each historical conversation turn; The processing module is further configured to perform multi-intention recognition on the interactive speech of the target elderly person in the current conversation turn to obtain different semantic intentions of the target elderly person in the current conversation turn; The processing module is further configured to determine the semantic integrity of the target elderly person in the current conversation turn according to all the semantic intentions and the semantic constraint conditions in each historical conversation turn; An execution module, configured to correct the speech content of the target elderly person in the current conversation turn through the semantic integrity and the intention confidence of each semantic intention.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the speech recognition correction method described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the speech recognition correction method described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Session processing method and device

    CN121166852A

  • Elderly semantic recognition system constructed based on corpus

    CN121438829A