AI voice interview intelligent generation method and system for chronic disease management
By using an AI-powered voice-based follow-up questionnaire intelligent generation method, the problem of traditional chronic disease voice follow-up systems being unable to dynamically adjust has been solved. This enables standardized information collection, structured extraction, and personalized interaction, thereby improving the precision of chronic disease management and patient participation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG YISHAN SMART MEDICAL RES CO LTD
- Filing Date
- 2026-06-24
- Publication Date
- 2026-07-24
AI Technical Summary
Traditional voice follow-up systems for chronic diseases cannot dynamically adjust their questioning strategies and expressions based on individual patient feedback, resulting in a mechanical and rigid interaction process and incomplete information collection, which fails to meet the needs of refined management of chronic diseases.
The AI-powered voice-based randomized questionnaire generation method is adopted. Through voice recognition, terminology standardization, natural language understanding, compliance scoring, and adaptive questionnaire path generation, the questionnaire nodes and wording style are dynamically adjusted to achieve personalized interaction.
It has achieved standardized collection of chronic disease follow-up information, structured extraction of key information, quantitative assessment of compliance status, and real-time adaptive adjustment of interactive content, thereby improving the completeness of information collection and patients' sense of participation.
Smart Images

Figure CN122455201A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent voice follow-up technology, and in particular to an AI voice follow-up questionnaire intelligent generation method and system for chronic disease management. Background Technology
[0002] In the long-term management of chronic diseases, regular follow-up is a core means of monitoring changes in a patient's condition, assessing treatment effectiveness, and adjusting intervention plans. With the widespread adoption of voice interaction technology, automated voice follow-up is gradually becoming an efficient and low-cost form of follow-up. This involves calling patients via telephone or smart devices, asking questions according to a pre-set questionnaire, and collecting information such as the patient's self-reported symptoms, medications, and lifestyle habits. This information is then used to assist healthcare professionals in conducting large-scale chronic disease monitoring. Traditional voice follow-up for chronic diseases typically uses a fixed script-driven approach. The system reads out pre-arranged standard questionnaire questions sequentially, regardless of what information the patient has already provided or what state they exhibit during the current follow-up. While this method enables the bulk collection of patient information, the follow-up process itself lacks the ability to adapt to individual differences.
[0003] Traditional, fixed-script-based voice follow-up for chronic diseases has significant shortcomings. Because the content of the follow-up interaction cannot be dynamically adjusted based on the actual amount of information, the quality of the patient's expression, and their current state, previously obtained information may be repeatedly asked, key risk points lack in-depth exploration, and ambiguous or contradictory statements cannot be clarified in a timely manner. This results in a mechanical and rigid follow-up dialogue, decreased patient participation, and difficulty in ensuring the completeness and effectiveness of the collected data. This limitation makes traditional voice follow-up unable to meet the requirements of refined chronic disease management for interactive flexibility and information accuracy, and also fails to provide patients with targeted long-term management support. Summary of the Invention
[0004] This application provides an AI-powered voice-based intelligent generation method and system for chronic disease management, which improves upon the technical problems in existing technologies where voice interaction during chronic disease follow-up often relies on fixed scripts, making it impossible to adjust the questioning strategy and expression in real time based on individual patient feedback. This results in a mechanical interview process, poor patient experience, and incomplete information collection, hindering the personalization and continuity of long-term chronic disease management.
[0005] This application discloses the following technical solution: In a first aspect, this application provides an AI-powered method for intelligently generating voice-based follow-up visit questionnaires for chronic disease management, the method comprising: Acquire user voice response data, perform speech recognition and terminology standardization processing on the voice response data to obtain standardized text data; Natural language understanding processing is performed on the standardized text data to extract key medical information and generate structured follow-up records; Based on the structured follow-up records, calculate the user's overall compliance score; Based on a pre-defined directed decision graph of follow-up questionnaires, and with the comprehensive compliance score and the information coverage status in the structured follow-up records as inputs, the questionnaire nodes in the decision graph are dynamically adjusted to generate an adaptive questionnaire path for the current follow-up round. Following the adaptive questionnaire path, the next round of voice follow-up interaction will be conducted.
[0006] Secondly, this application provides an AI-powered voice-based intelligent generation system for chronic disease management, the system comprising: The voice acquisition module is used to acquire the user's voice response data, and to perform voice recognition and terminology standardization processing on the voice response data to obtain standardized text data; The semantic parsing module is used to perform natural language understanding processing on the standardized text data, extract key medical information, and generate structured follow-up records; The compliance assessment module is used to calculate the user's overall compliance score based on the structured follow-up records; The questionnaire generation module is used to dynamically adjust the questionnaire nodes in the decision graph based on a preset follow-up questionnaire directed decision graph, and with the comprehensive compliance score and the information coverage status in the structured follow-up record as input, to generate an adaptive questionnaire path for the current follow-up round. The interactive execution module is used to execute the next round of voice follow-up interaction according to the adaptive questionnaire path.
[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages: The technical solution of this application provides an AI-powered voice follow-up questionnaire intelligent generation method for chronic disease management. First, by recognizing and standardizing the terminology of the user's voice response data, the patient's colloquial expressions are transformed into standardized medical terminology text, which realizes the standardized connection of the chronic disease follow-up information collection entry point and solves the problem that existing voice follow-up systems are unable to accurately capture the core semantics of chronic diseases due to the diversity of colloquial expressions and the prevalence of non-standard terms.
[0008] Furthermore, by performing semantic understanding on standardized text, key information related to chronic diseases is automatically extracted and structured follow-up records are generated. This achieves automated semantic assembly and formatting of follow-up information, solving the problems of low efficiency and insufficient generalization ability for complex expressions when relying on manual review or keyword matching in traditional methods.
[0009] Furthermore, by conducting multi-dimensional analysis and quantification of patients' protocol implementation based on structured records, a unified cooperation score is generated, realizing the transformation of patient follow-up status from qualitative description to quantitative assessment, and solving the problem that single behavioral indicators or experience judgments cannot fully reflect patients' true willingness to cooperate.
[0010] Furthermore, by using cooperation level scores and current information coverage status as decision-making criteria, preset questionnaire nodes are trimmed, expanded, and rearranged to dynamically generate personalized question paths that match the patient's current status. This enables real-time adaptive adjustment of follow-up interaction content and solves the problems of interaction redundancy and risk omission caused by the inability of fixed questionnaires to flexibly change according to individual differences.
[0011] Finally, by sensing patients' emotional feedback during the questionnaire execution phase and dynamically switching the communication style according to their emotional state and level of cooperation, we achieved emotional interactive intervention during the follow-up process, which solved the problem that fixed communication styles are difficult to match with patients' psychological state and affect long-term compliance and willingness to participate.
[0012] In summary, the technical solution of this application achieves standardized response collection, structured key information, quantitative assessment of compliance status, dynamic adjustment of questionnaire path, and emotionally adaptive interaction in the voice interaction of chronic disease follow-up. Through speech recognition and terminology mapping in the chronic disease field, NLU (Natural Language Understanding) semantic parsing and template filling, multi-dimensional compliance fusion scoring, adaptive pruning of directed decision graph nodes, and emotionally driven dialogue mode switching, it effectively improves the technical problems in existing technologies, such as fixed and rigid follow-up dialogues, inability to adjust the content and method of inquiries in real time according to individual patient feedback, resulting in information redundancy and omission, distorted compliance assessment, insufficient patient participation, and difficulty in supporting the continuous optimization of refined chronic disease management. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 A schematic diagram of the overall process of the AI-based voice-based intelligent generation method for chronic disease management provided in the embodiments of this application; Figure 2 A flowchart illustrating the dynamic adjustment of questionnaire nodes in the AI-based voice-based intelligent generation method for chronic disease management provided in this application embodiment; Figure 3This is a schematic diagram of the structure of an AI voice-based intelligent generation system for chronic disease management, provided in an embodiment of this application.
[0015] In the attached diagram, Figure 3 The components represented by each number are explained as follows: Voice acquisition module 11, semantic parsing module 12, compliance assessment module 13, questionnaire generation module 14, and interaction execution module 15. Detailed Implementation
[0016] This application provides an AI-powered voice-based follow-up questionnaire intelligent generation method and system for chronic disease management. It addresses the technical problem that existing chronic disease follow-up voice systems often use standardized questionnaires, which cannot dynamically adjust subsequent questions and interaction methods based on individual patient feedback. This results in a lack of targeted follow-up interaction, insufficient user participation, and difficulty in accurately capturing disease details to support long-term management.
[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. It should be noted that the numerical values in the embodiments are for illustrative purposes only and do not constitute a limitation on this application.
[0018] Example 1, as shown in the appendix Figure 1 As shown, this application provides an AI-powered voice-based intelligent generation method for chronic disease management, the method comprising the following steps: S100: Acquire the user's voice response data, perform speech recognition and terminology standardization processing on the voice response data, and obtain standardized text data; In this embodiment of the application, in the scenario of voice interaction during chronic disease follow-up, patients often freely describe their condition through speech. However, their spoken expressions are often mixed with non-standard terms, dialects, or habitual omissions. General speech recognition methods struggle to accurately capture the core semantics of chronic diseases, resulting in a lack of reliable and standardized information input for subsequent questionnaire generation. Therefore, a speech recognition and terminology standardization method for the chronic disease field is needed to automatically convert patients' spoken language into standard medical terminology, thereby solving the problems of information extraction bias and interaction strategy failure caused by colloquial expressions.
[0019] Step S100 in the method provided in this application embodiment includes: An ASR model pre-trained with a chronic disease medical terminology vocabulary is used to convert speech response data into initial text. A medical terminology mapping dictionary is then used to map colloquial expressions in the initial text to standard medical terms or pharmacopoeia codes. A detailed explanation follows: In this embodiment, the ASR (Automatic Speech Recognition) model refers to an automatic speech recognition model that uses deep learning technology to convert speech signals into text sequences and is pre-trained on a corpus containing chronic disease terminology to improve the accuracy of recognizing chronic disease-related expressions. The chronic disease medical terminology vocabulary refers to a vocabulary containing standard terms and common colloquial variations of symptoms, medications, and examinations related to chronic diseases such as hypertension and diabetes, used to enhance the ASR model's ability to recognize domain-specific terms. The medical terminology mapping dictionary refers to a table mapping colloquial expressions to standard medical terms or pharmacopoeia codes, for example, mapping "tang da le" to "blood sugar elevation." Colloquial expressions refer to non-standard, omitted, or habitual terms used by patients in their voice responses, such as vague descriptions like "a little high" or "that medicine." Standard medical terms or pharmacopoeia codes refer to standardized disease, symptom, and drug names or corresponding national drug codes, such as nifedipine controlled-release tablets or hyperglycemia.
[0020] In this step, firstly, in order to accurately convert the patient's chronic disease-related speech into initial text, an ASR model with a pre-trained chronic disease medical terminology vocabulary is used to process the speech response data. This will produce initial text that reflects the patient's original semantics but retains the colloquial expression. For example, if the patient says, "My blood sugar has been a bit high lately, and I keep forgetting to take my medicine," the ASR model will output the corresponding text content.
[0021] Furthermore, in order to eliminate colloquial ambiguity and standardize expressions, it is necessary to map colloquial expressions in the initial text to standard medical terms or pharmacopoeia codes through a medical terminology mapping dictionary. This will result in standardized text, allowing subsequent information extraction to be directly based on unified terminology. For example, "slightly high blood sugar" can be mapped to "elevated blood sugar," and "frequently forgets to take medication" can be mapped to "irregular medication use."
[0022] For example, taking a patient with hypertension as the follow-up subject, their voice response was obtained through a telephone voice channel. The patient stated: "I've been feeling a bit dizzy these past few days, and my blood pressure seems to be high again. I sometimes forget to take the nifedipine prescribed by the doctor, and I take one when I remember." The initial text obtained by the ASR model was a colloquial expression. After standardization using a medical terminology mapping dictionary, the standardized text "dizziness, high blood pressure, nifedipine, irregular medication" was output. The entire process took approximately 1.2 seconds, with a speech recognition error rate of less than 5%.
[0023] In summary, this step transforms patients' unstructured spoken responses into standardized medical text through speech recognition and terminology standardization processing adapted to the chronic disease domain. Compared with existing technologies, this step has the following advantages: First, the use of a chronic disease medical terminology vocabulary significantly improves the recognition accuracy of chronic disease-related spoken expressions and reduces the false recognition rate of general ASR; second, through medical terminology mapping, it achieves automatic standardization from spoken language to standard terminology, eliminating ambiguity in subsequent structured information extraction and providing a high-quality data foundation for dynamically generating personalized follow-up questionnaires.
[0024] S200: Perform natural language understanding processing on the standardized text data, extract key medical information, and generate structured follow-up records; In this embodiment, even with standardized text data, the text, though organized into medical terminology, remains natural language fragments and cannot be directly used for compliance assessment and questionnaire logic operations. Traditional methods often rely on manual review or simple keyword matching to extract follow-up information. When faced with diverse expressions and complex semantics, key information is easily missed or misjudged, and standardized structured records cannot be generated to support subsequent automated decision-making. Therefore, a method is needed that can automatically understand semantics from standardized text, accurately extract elements according to chronic disease follow-up needs, and organize them into structured data to solve the problems of incomplete follow-up information caused by low efficiency of manual processing and poor generalization of rule matching.
[0025] Step S200 in the method provided in this application embodiment includes: An NLU model, fine-tuned based on a chronic disease-annotated corpus, is used to perform intent recognition and slot filling on the standardized text data. Extracted key slot information is automatically mapped to preset public health follow-up template fields to generate structured follow-up records. Detailed explanation follows: In this embodiment, the NLU model refers to a Natural Language Understanding model, which is pre-trained on general text and fine-tuned on annotated corpora such as chronic disease consultation dialogues and follow-up records to enable it to extract medical-related semantic elements. Intent recognition refers to determining the main purpose category expressed by the user's input statement, such as reporting symptoms, providing medication feedback, or asking questions. Slot filling refers to extracting predefined key information fragments from the text and filling them into corresponding slots, such as extracting the value of "elevated" blood pressure and filling it into the blood pressure status slot. Public health follow-up template fields refer to structured information items preset according to the chronic disease public health follow-up specifications, including symptoms, drug names, medication frequency, blood pressure values, blood sugar values, and descriptions of discomfort. Structured follow-up records refer to follow-up information sets organized using key-value pairs or tables, which can be directly read by a computer; for example, the symptom value is dizziness, and the blood pressure status value is elevated.
[0026] In this step, in order to identify the patient's expressed medical intent and extract specific medical elements from the standardized text, an NLU model finely tuned based on chronic disease annotated corpus is used to perform intent recognition and slot filling, obtaining key slot value pairs with intent labels. For example, the intent is identified as reporting symptoms and providing feedback on medication, and the extracted symptom values are dizziness, blood pressure status is elevated, medication name is nifedipine, and medication pattern is irregular.
[0027] Furthermore, in order to integrate the extracted information into a standardized format for use in subsequent steps, key slot information needs to be automatically mapped to public health follow-up template fields to generate structured follow-up records. For example, after filling the slot values into the corresponding fields, the symptom description is dizziness, the blood pressure assessment is elevated, the drug name is nifedipine, and the medication adherence is irregular.
[0028] For example, taking a patient with hypertension as the follow-up subject, the standardized text obtained from the aforementioned steps—dizziness, elevated blood pressure, nifedipine, and irregular medication use—is processed. An NLU model identifies the intent, outputting intent categories as stating symptoms and stating medication use. Slot filling is performed, extracting the symptom slot for dizziness, the blood pressure slot for elevated blood pressure, the medication slot for nifedipine, and the medication adherence slot for irregularity. These slots are then mapped to a public health follow-up template to generate a structured follow-up record: symptom description: dizziness; blood pressure assessment: elevated; medication name: nifedipine; medication adherence: irregularity. The processing time is approximately 0.8 seconds, and the slot extraction accuracy is greater than 92%.
[0029] In summary, this step utilizes a finely tuned NLU model for chronic disease domains to perform semantic understanding on standardized text and generate structured follow-up records. Its advantages include: the NLU model can identify medical intentions and slots in complex contexts, avoiding the rigidity and omissions of rule matching; and automatic mapping to public health follow-up templates ensures a unified and complete output record structure, providing an immediate and quantifiable data foundation for subsequent comprehensive adherence scoring and dynamic questionnaire adjustments.
[0030] S300: Calculate the user's overall compliance score based on the structured follow-up records; In this embodiment, under the scenario where structured follow-up records are available, the records contain factual information such as the patient's self-reported medication behavior and monitoring behavior. However, these behavioral records alone are insufficient to fully reflect the patient's true level of cooperation. Traditional compliance assessments often rely solely on single behavioral indicators or human experience judgment, ignoring emotional signals such as resistance, perfunctoriness, or anxiety naturally revealed in patient communication. This leads to a discrepancy between the assessment results and the actual willingness to cooperate, failing to provide a reliable reference for subsequent interactive decisions. Therefore, a method is needed that can integrate behavioral facts and emotional attitudes to comprehensively quantify the patient's level of cooperation, in order to solve the problem of one-dimensional assessment being biased and unable to accurately represent the patient's compliance status.
[0031] Step S300 in the method provided in this application embodiment includes: The user's self-reported behavior extracted from the structured follow-up records is compared with a pre-set standard intervention plan to calculate a content consistency score; voice emotion recognition is performed on the user's voice response data, and text emotion classification is performed on the standardized text data to calculate an emotion tone score; the content consistency score and the emotion tone score are weighted and fused to obtain a comprehensive compliance score. A detailed explanation follows: In this embodiment, the standard intervention plan refers to a pre-set treatment plan based on the patient's chronic disease type, including standardized medical orders such as drug name, dosage, monitoring frequency, and lifestyle recommendations, serving as a baseline for measuring the degree of patient compliance. The content consistency dimension score refers to the behavioral compliance score obtained by quantitatively calculating the degree of matching between the patient's self-reported medication use, monitoring, and other behaviors and the standard intervention plan. A higher score indicates that the behavior is closer to the medical order requirements.
[0032] Voice emotion recognition analyzes the acoustic features of voice responses to identify the underlying emotional states, such as positive, neutral, anxious, or resistant. Text emotion classification analyzes standardized text data to determine the positive, negative, or neutral emotions expressed in a patient's written expression. Emotional tone scoring combines voice emotion recognition and text emotion classification results, converting them according to preset rules to obtain a score reflecting the patient's emotional compliance. Weighted fusion combines the content consistency score and the emotional tone score with preset weighting coefficients to obtain a comprehensive compliance score. The weights can be adjusted according to the specific application scenario; for example, a weight of 0.6 for the behavioral dimension and 0.4 for the emotional dimension ensures the final score reflects both behavioral compliance and emotional compliance.
[0033] In this step, firstly, in order to objectively measure the degree of consistency between the patient's self-reported behavior and the doctor's orders, the user's self-reported behavior in the structured follow-up record needs to be compared with the standard intervention plan item by item, and the content consistency dimension score is calculated. For example, the medication name, medication frequency, and blood pressure monitoring behavior are compared to see if they are consistent with the plan. The scores of each matching item are added together and normalized to a percentage score.
[0034] Furthermore, in order to capture the emotional attitudes revealed by patients during communication, it is necessary to perform voice emotion recognition on the raw voice response data and text emotion classification on the standardized text data. The combined analysis results of the two are used to generate an emotional tone dimension score. For example, when the voice recognition detects hesitation, sighing, or when words such as helplessness or boredom appear in the text, the emotional score will decrease accordingly.
[0035] Finally, in order to obtain a unified score that can comprehensively represent the patient's level of cooperation, the content consistency dimension score and the affective tone dimension score need to be weighted and integrated to obtain a comprehensive compliance score, which can be used for dynamic decision-making in subsequent questionnaire paths.
[0036] For example, taking a patient with hypertension as the follow-up subject, the structured follow-up record obtained from the aforementioned steps shows that the symptom description is dizziness, the blood pressure assessment is elevated, the medication name is nifedipine, and the medication adherence is irregular. The medication adherence field is compared with the standard intervention protocol's requirement of once-daily nifedipine, and the content consistency dimension score is calculated to be 40 points. Simultaneously, sentiment analysis is performed on the original speech and standardized text. The speech sentiment recognition result is slightly helpless, and the text sentiment classification is negative, resulting in a combined sentiment tone dimension score of 35 points. A weighted fusion is performed using a behavioral weight of 0.6 and a sentiment weight of 0.4, calculating a comprehensive adherence score of 40. 0.6+35 0.4 38 points. This score is below the preset low compliance threshold of 40 points, indicating that targeted follow-up questions are needed.
[0037] In summary, this step assesses and weights compliance from two dimensions—content consistency and emotional tone—to obtain a comprehensive adherence score. Its advantages are: it quantifies the gap between the patient's actual behavior and the doctor's orders, while also introducing an emotional attitude dimension to compensate for the limitations of a single behavioral assessment; the unified quantitative score allows the subsequent decision-making module to dynamically adjust questionnaire nodes based on clear numerical thresholds, improving the relevance and personalization of follow-up interactions.
[0038] S400: Based on the preset directed decision graph of the follow-up questionnaire, and with the comprehensive compliance score and the information coverage status in the structured follow-up record as input, dynamically adjust the questionnaire nodes in the decision graph to generate an adaptive questionnaire path for the current follow-up round. In this embodiment, under the scenario where structured follow-up records and comprehensive compliance scores have been obtained, traditional follow-up interactions often employ questionnaire scripts in a fixed order. Regardless of what information has been collected in this round, whether the patient cooperates, or whether the answers are clear and complete, the same process is mechanically advanced. This leads to repeated questioning of already covered content, failure to delve into low-risk cooperation points, and a lack of clarification for ambiguous or contradictory statements, making the follow-up a mere formality and resulting in inconsistent data quality. Therefore, a method is needed that can plan the content and depth of subsequent questions in real time based on the completeness of the information, willingness to cooperate, and quality of expression provided by the patient, in order to solve the problems of interaction redundancy, risk omission, and information distortion caused by fixed questionnaires failing to match individual differences.
[0039] Step S400 in the method provided in this application embodiment includes: When the structured follow-up record already contains the target information to be asked at a certain questionnaire node, that node is automatically skipped; when the overall compliance score is lower than a first preset threshold, follow-up questions associated with low compliance items are added or activated; when the overall compliance score is higher than a second preset threshold and the relevant indicators meet the stability condition, the questioning frequency of the corresponding questionnaire node is skipped or reduced, wherein the first preset threshold is less than the second preset threshold; when ambiguous, contradictory, or missing information is detected in the user's answer, a preset clarification follow-up question is automatically inserted. Detailed explanation follows: In this embodiment, the directed decision graph of the follow-up questionnaire refers to a questionnaire structure pre-constructed by experts and organized in the form of a directed graph. Nodes in the graph represent a question topic or information collection item, and directed edges represent the sequential jump relationship between nodes. Information coverage status refers to the matching status of each field in the structured follow-up record with the target information of the decision graph nodes. If all necessary slots corresponding to a node have been filled with valid values, then the node information is determined to be covered. Adaptive questionnaire path refers to a personalized node sequence generated after dynamic adjustment, containing only the question nodes required for the current round. Follow-up node refers to a pre-set in-depth inquiry branch in the decision graph targeting specific low-compliance-risk items, such as the specific reasons and frequency of missed medication doses. Stable condition refers to a state where key physiological indicators are within the normal range and without significant fluctuations or a continuous deterioration trend in multiple consecutive follow-ups, for example, blood pressure values in the last three follow-ups are all below 140 / 90 mmHg and without an increasing trend. The first preset threshold refers to the scoring threshold used to determine low-compliance-risk, typically 40 points; below this value, the activation of low-compliance follow-up nodes is triggered. The second preset threshold is a scoring threshold used to determine a high compliance and stable state. A typical value is 80 points. When the score is higher than this value and the indicator is stable, the query frequency of the corresponding node is reduced.
[0040] The overall process of dynamically adjusting questionnaire nodes in the decision-making diagram in this step is as follows: Figure 2 As shown, the specific steps include the following.
[0041] First, to reduce unnecessary duplicate queries, when a node's target information is detected to already be present in a structured record, that node is automatically skipped. For example, if a blood pressure-related field is already filled, blood pressure will not be asked again.
[0042] Furthermore, in order to address situations of insufficient compliance, when the overall compliance score is lower than the first preset threshold, follow-up question nodes associated with low compliance items are added or activated. For example, if medication is taken irregularly, follow-up questions about the reasons for missed doses are inserted.
[0043] Furthermore, to avoid disturbing stable patients, when the overall compliance score is higher than the second preset threshold and the key indicators meet the stable condition, the frequency of the corresponding questionnaire node is skipped or reduced. For example, if blood pressure is normal for a long time, the questioning frequency is changed from every follow-up visit to once every 3 visits.
[0044] Finally, when the answer contains ambiguous, contradictory, or missing information, a clarification questioning node is automatically inserted to improve the quality of information. The specific judgment and processing methods for this sub-step will be further elaborated in subsequent steps.
[0045] Step S400 in the method provided in this application embodiment further includes: When a key information field in a user's answer is found to be empty or its value exceeds a preset normal range, it is determined to be missing information, and a supplementary inquiry node for that field is inserted. When a discrepancy is found between the user's current answer and the same indicator in historical follow-up records, it is determined to be contradictory information, and a double inquiry node for verification is inserted, marking the indicator as pending confirmation. When a user's answer contains vague quantifiers or uncertain expressions, it is determined to be ambiguous information, and a follow-up inquiry node for specific quantification is inserted to guide the user to provide precise values or a clear status. Detailed explanations are as follows: In this embodiment, key information fields refer to the core slots that must be collected with the access volume, such as medication name, frequency of medication, and blood glucose measurement values. Supplementary query nodes refer to query nodes designed for missing fields, used to obtain the necessary information that was originally missing. Dual query nodes refer to nodes that re-query the same contradictory indicator in different ways to verify it, while marking its status as pending confirmation to prevent it from being considered covered. A pending confirmation status indicates that an indicator is temporarily marked due to contradiction or suspicion, and must be verified again before it can be considered valid. Vague quantifiers refer to imprecise words used in user statements, such as sometimes, probably, not too much, etc., lacking quantitative basis.
[0046] In this step, to handle missing information, when a key information field is detected to be empty or its value exceeds the normal range, a supplementary inquiry node is inserted. For example, if the blood glucose measurement value for this week is not mentioned, a follow-up inquiry for the specific reading is added. Furthermore, to handle contradictory information, when the current answer is found to be inconsistent with the same indicator in the historical record, a double inquiry node is inserted for verification and marked as pending confirmation. For example, if the patient says their blood pressure is normal this time, but the previous record showed it was high, a follow-up inquiry is made to confirm again and the result is temporarily marked. Finally, to handle ambiguous information, when the answer contains vague quantifiers or uncertain expressions, a quantitative follow-up inquiry node is inserted. For example, if the patient says they don't exercise much, a follow-up inquiry is made to ask how many times they exercise per week and for how many minutes each time.
[0047] For example, taking a patient with hypertension as the follow-up subject, the structured record obtained from the aforementioned steps includes symptoms of dizziness, blood pressure assessment of elevated, medication name of nifedipine, medication adherence of irregularity, and a comprehensive adherence score of 38. The preset decision graph includes blood pressure, medication, symptom, and exercise nodes. Since the necessary information for the blood pressure, medication, and symptom nodes has been covered, this step automatically skips these three nodes. If the comprehensive adherence score is lower than the first preset threshold of 40, the low medication adherence follow-up node is activated, adding inquiries about the frequency and reasons for missed doses. At the same time, the current answer uses the vague quantifier "irregularity," triggering a clarifying follow-up inquiry, continuing to ask how many doses were missed per week. The generated adaptive path for the current round is: low medication adherence follow-up node to missed dose frequency quantification follow-up node to exercise node. The entire process takes approximately 0.3 seconds, with a decision jump accuracy rate greater than 95%.
[0048] In summary, this step uses comprehensive adherence scores and information coverage status as inputs to dynamically adjust decision graph nodes and generate personalized questionnaire paths. Its advantages include: automatically skipping covered information to reduce redundant questions; automatically adding follow-up questions when adherence is low and reducing the frequency when adherence is high and stable, thus focusing the interaction on risk items; identifying and clarifying ambiguous, contradictory, and missing information, significantly improving the completeness and accuracy of collected data; and enabling flexible follow-up interactions to adapt to individual patient differences, supporting the continuous optimization of refined chronic disease management.
[0049] S500: Execute the next round of voice follow-up interaction according to the adaptive questionnaire path.
[0050] In this embodiment, under the scenario where a personalized questionnaire path has been generated, traditional voice interaction uses a fixed script to read each question one by one. Regardless of the patient's current mood or level of cooperation, the same tone and wording are used throughout. This can easily exacerbate negative feelings when the patient is anxious or resistant, and lack positive feedback when cooperation is good, weakening the sustainability of follow-up and the patient's enthusiasm for participation. Therefore, a method is needed that can dynamically adjust the script based on the patient's overall compliance status and real-time emotional tone to solve the problem of low interaction acceptance and unsustainable compliance improvement caused by the fixed script failing to match the patient's psychological state.
[0051] Step S500 in the method provided in this application embodiment includes: Obtain the comprehensive compliance score and affective tone dimension analysis results; when resistance or anxiety is detected in the affective tone dimension analysis results, switch the dialogue mode to empathetic or encouraging expression; when the user is identified as actively cooperating and the comprehensive compliance score is higher than a preset threshold, add positive incentive dialogue. Detailed explanation follows: In this embodiment, the emotional tone dimension analysis result refers to the patient's emotional state comprehensively identified from voice and text responses, including categories such as positive, neutral, resistant, and anxious, and follows the emotional analysis output from the aforementioned steps. The discourse pattern refers to the expression style and language strategy adopted by the system in voice interaction, such as neutral questioning, empathetic reassurance, encouraging guidance, or positive motivation. Empathic expression refers to incorporating statements of understanding and emotional recognition of the patient's situation into the discourse, such as "I understand your recent troubles; let's adjust together slowly." Encouraging expression refers to statements conveying confidence and companionship in a gentle, supportive tone, such as "You've already done very well; keep going, and you'll do even better." Positive motivational discourse refers to statements that clearly affirm and praise the good behavior already demonstrated by the patient, such as "You've done a fantastic job consistently measuring your blood pressure this week." The preset threshold can be the same as the aforementioned second preset threshold; in this example, it is set to 80 points, meaning that positive motivational discourse is triggered when the overall compliance score is higher than 80 points and the patient's emotions are positive.
[0052] In this step, to perceive the patient's current emotions and respond appropriately, it is necessary to obtain the comprehensive compliance score and emotional tone dimension analysis results, which will serve as trigger signals for switching the script. Furthermore, when resistance or anxiety is detected, the script mode is immediately switched to empathetic or encouraging expressions to reduce psychological defenses and negative emotions; for example, if impatience or tension is detected, a reassuring script is used. Further, when the patient is found to be actively cooperating and the comprehensive compliance score is above a preset threshold, positive reinforcement is added to strengthen good behavior; for example, praise is added when the score is above 80 and the patient's emotions are positive.
[0053] For example, taking a patient with hypertension as the follow-up subject, the comprehensive compliance score obtained from the aforementioned steps was 38 points. The emotional tone dimension analysis result showed a slightly helpless tone and negative text sentiment, triggering the identification of resistance / anxiety. When performing questioning according to the adaptive path, the system switched the wording from a standard question to an encouraging expression, such as "Don't worry, we'll take it one step at a time." At the same time, since the score was only 38 points, it did not reach the high compliance threshold required to trigger positive incentives, and the sentiment analysis did not show positive cooperation, so positive incentive words were not triggered. If the score improved to 82 points in subsequent follow-ups and the sentiment turned positive, positive incentive statements such as "You've been doing very well recently, please keep it up" were added.
[0054] In summary, this step dynamically switches the communication style based on compliance and emotion recognition results during the questionnaire execution phase. Its advantages include: timely reassurance of negative emotions can prevent patients from giving up or resisting; reinforcement of positive behaviors can increase willingness to cooperate; and follow-up interactions are made more humane, thereby improving compliance and effectiveness in long-term chronic disease management.
[0055] This application embodiment also includes the following steps: Based on the overall compliance score and its trend in the current round, the duration of the next follow-up cycle is automatically adjusted. When the overall compliance score is below a third preset threshold for several consecutive cycles, the follow-up cycle duration is shortened; when the overall compliance score is above a fourth preset threshold for several consecutive cycles, the follow-up cycle duration is extended. The third preset threshold is less than or equal to a first preset threshold, and the fourth preset threshold is greater than or equal to a second preset threshold. Detailed explanation follows: In this embodiment, the follow-up period refers to the time interval between two adjacent voice follow-up interactions. The third preset threshold is a scoring threshold used to determine the risk of long-term low compliance. If the score is below this threshold for several cycles, the follow-up interval is shortened. A typical value is 40 points. The fourth preset threshold is a scoring threshold used to determine long-term high and stable compliance. If the score is above this threshold for several cycles, the follow-up interval is extended. A typical value is 85 points.
[0056] In this step, to avoid insufficient attention to patients with low adherence or excessive disturbance to stable patients, it is necessary to track the continuous trend of changes in the comprehensive adherence score. For example, if the score is below 40 for three consecutive cycles, the follow-up period is shortened from 30 days to 15 days; if the score is above 85 for three consecutive cycles, the follow-up period is extended from 30 days to 45 days. Through dynamic adjustment of the cycle, the follow-up frequency is matched with the patient's risk level, improving the efficiency of management resource utilization.
[0057] This application embodiment also includes the following steps: The structured follow-up records, comprehensive compliance scores, and adaptive questionnaire path execution logs generated in each round of follow-up are stored as historical data. This historical data is used to iteratively update the NLU model, the weight parameters of each node in the directed decision graph, and the fusion strategy for the comprehensive compliance score. A detailed explanation follows: In this embodiment, historical data refers to the collection of medical information records, scoring results, and decision execution processes generated during each follow-up visit. NLU model update refers to incrementally training the NLU model using newly labeled or pseudo-labeled historical data to improve the accuracy of intent recognition and slot filling. Decision graph weight parameters refer to the weight values of transition edges or node priorities between nodes in the directed decision graph, which are optimized by statistically analyzing the effectiveness of historical paths. Fusion strategy refers to the weighting coefficients or fusion methods of content consistency dimension scores and sentiment tone dimension scores in the comprehensive compliance score, which can be dynamically adjusted through correlation analysis between compliance scores and actual compliance outcomes in historical data.
[0058] In this step, after each round of interaction, the structured records, scores, and path logs are persistently stored, and offline iterations are performed periodically using the accumulated historical data. For example, after every 500 follow-up records, the NLU model is retrained and the decision graph weights and fusion coefficients are updated to make the semantic understanding of subsequent follow-ups more accurate, the path pruning more reasonable, and the compliance assessment closer to the real situation.
[0059] For example, taking a patient with hypertension as the follow-up subject, after the first round of follow-up, their comprehensive adherence score was 38 points. After three consecutive rounds of follow-up, the score remained below 40 points. The system automatically shortened the original 30-day follow-up period to 15 days and added follow-up questions on medication adherence. After two cycles of intensive intervention, the patient's score rose to 62 points. After three consecutive rounds of follow-up, the score stabilized between 55 and 70 points, but still did not meet the high adherence standard, so the system reverted to the default 30-day cycle. When the scores in the 6th, 7th, and 8th rounds were 87, 90, and 92 points respectively, and the patient's emotional state was positive, the system extended the cycle to 45 days. Meanwhile, after accumulating complete follow-up data for the patient for 8 rounds and data from 1,000 other patients during the same period in the background, the system started offline iteration: using 1,200 newly labeled chronic disease follow-up responses to incrementally fine-tune the NLU model, improving slot extraction accuracy from 92% to 94%; recalculating the jump weights of each node in the decision graph, reducing the adaptive path pruning idle rate by 12%; adjusting the weight ratio of behavior and emotion in the fusion strategy to 0.65 to 0.35, improving the consistency between the overall compliance score and the actual clinical assessment results by 8%.
[0060] In summary, the dynamic adjustment of follow-up cycles and the data storage and iterative optimization steps enable the follow-up frequency to adaptively scale with patient compliance trends, providing more intensive attention to high-risk patients and reducing unnecessary disturbances to stable patients. At the same time, through closed-loop iteration and continuous optimization of semantic understanding, decision-making paths, and scoring strategies, the personalization and effectiveness of chronic disease follow-up management are continuously improved with the accumulation of data.
[0061] The embodiments of this application, through the specific implementation methods described above, achieve the following technical effects: This application proposes an AI-powered intelligent generation method for voice-based follow-up questionnaires for chronic disease management. First, voice response data is collected and standardized to obtain standardized text data. Then, key medical information is extracted from the standardized text using natural language understanding processing to generate structured follow-up records. Next, a comprehensive adherence score is calculated based on the structured follow-up records. Using the comprehensive adherence score and the information coverage status of the structured follow-up records as input, the questionnaire nodes in the decision graph are dynamically adjusted to generate an adaptive questionnaire path for the current follow-up round. When the adherence score triggers a preset threshold or when information is ambiguous, contradictory, or missing, follow-up questions are automatically inserted or the questioning frequency is adjusted. Finally, the next round of voice-based follow-up interaction is executed according to the adaptive path, and the follow-up cycle is adjusted based on adherence trends. Simultaneously, the model and decision parameters are iteratively optimized using accumulated historical data. This method achieves real-time adaptive generation and continuous optimization of chronic disease follow-up interaction content through a closed-loop linkage of voice understanding, adherence quantification, and dynamic adjustment of the decision graph.
[0062] Example 2, as shown in the appendix Figure 3 As shown, based on the inventive concept of the AI voice-based intelligent generation method for chronic disease management provided in Embodiment 1, this application also provides an AI voice-based intelligent generation system for chronic disease management, specifically including: The voice acquisition module 11 is used to acquire the user's voice response data, and to perform voice recognition and terminology standardization processing on the voice response data to obtain standardized text data. Semantic parsing module 12 is used to perform natural language understanding processing on the standardized text data, extract key medical information, and generate structured follow-up records; The compliance assessment module 13 is used to calculate the user's overall compliance score based on the structured follow-up records; The questionnaire generation module 14 is used to dynamically adjust the questionnaire nodes in the decision graph based on a preset follow-up questionnaire directed decision graph, and with the comprehensive compliance score and the information coverage status in the structured follow-up record as input, to generate an adaptive questionnaire path for the current follow-up round. The interaction execution module 15 is used to execute the next round of voice follow-up interaction according to the adaptive questionnaire path.
[0063] In one embodiment, the voice acquisition module 11 is further configured to: use an ASR model pre-trained with a chronic disease medical terminology vocabulary to convert voice response data into initial text; and map colloquial expressions in the initial text to standard medical terms or pharmacopoeia codes through a medical terminology mapping dictionary.
[0064] In one embodiment, the semantic parsing module 12 is further configured to: use an NLU model finely tuned based on chronic disease annotated corpus to perform intent recognition and slot filling on the standardized text data; automatically map the extracted key slot information to preset public health follow-up template fields to generate structured follow-up records.
[0065] In one embodiment, the compliance assessment module 13 is further configured to: compare the user's self-reported behavior extracted from the structured follow-up records with a preset standard intervention plan, and calculate a content consistency dimension score; perform voice emotion recognition on the user's voice response data, and perform text emotion classification on the standardized text data, and calculate an emotion tone dimension score; and perform weighted fusion of the content consistency dimension score and the emotion tone dimension score to obtain a comprehensive compliance score.
[0066] In one embodiment, the questionnaire generation module 14 is further configured to: automatically skip a questionnaire node when the structured follow-up record already contains the target information to be asked at that questionnaire node; increase or activate follow-up question nodes associated with low compliance items when the overall compliance score is lower than a first preset threshold; skip or reduce the questioning frequency of the corresponding questionnaire node when the overall compliance score is higher than a second preset threshold and the relevant indicators meet the stability condition, wherein the first preset threshold is less than the second preset threshold; and automatically insert preset clarification follow-up question nodes when ambiguous, contradictory, or missing information is detected in the user's answer.
[0067] Furthermore, the questionnaire generation module 14 is also used to: determine that information is missing when a key information field in a user's answer is empty or the value exceeds the preset normal range, and insert a supplementary inquiry node for that field; determine that information is contradictory when a user's current answer is inconsistent with the same indicator in historical follow-up records, insert a double inquiry node for verification, and mark the indicator as pending confirmation; and determine that information is ambiguous when a user's answer contains vague quantifiers or uncertain expressions, insert a follow-up question node for specific quantification, and guide the user to provide precise values or clear status.
[0068] In one embodiment, the interactive execution module 15 is further configured to: obtain the comprehensive compliance score and the emotional tone dimension analysis results; when resistance or anxiety is detected in the emotional tone dimension analysis results, switch the speech mode to empathic or encouraging expression; when the user is identified as actively cooperating and the comprehensive compliance score is higher than a preset threshold, add positive incentive speech.
[0069] Furthermore, the interactive execution module 15 is also used to: automatically adjust the duration interval of the next follow-up cycle based on the comprehensive compliance score of the current round and its changing trend; shorten the follow-up cycle duration when the comprehensive compliance score is lower than the third preset threshold for multiple consecutive cycles; and extend the follow-up cycle duration when the comprehensive compliance score is higher than the fourth preset threshold for multiple consecutive cycles, wherein the third preset threshold is less than or equal to the first preset threshold, and the fourth preset threshold is greater than or equal to the second preset threshold.
[0070] Furthermore, the interactive execution module 15 is also used to: store the structured follow-up records, comprehensive compliance scores, and adaptive questionnaire path execution logs generated in each round of follow-up as historical data; and use the historical data to iteratively update the NLU model, the weight parameters of each node in the directed decision graph, and the fusion strategy of the comprehensive compliance score.
[0071] The AI-powered voice-based follow-up questionnaire intelligent generation system for chronic disease management provided in this application enables adaptive voice interaction closed-loop management in scenarios such as chronic disease follow-up management and remote health monitoring. This management encompasses everything from voice response data collection and terminology standardization to key medical information extraction and compliance quantification assessment, and then to dynamic decision graph trimming and personalized questionnaire path generation. It can be integrated into intelligent voice follow-up platforms or chronic disease management cloud systems, effectively improving the personalization of follow-up content and patients' long-term cooperation willingness. Simultaneously, it provides structured data support, including compliance scores, structured follow-up records, and interaction execution logs, for refined chronic disease management and intervention strategy optimization. For specific interaction processes and decision details of this system, please refer to Example 1.
[0072] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
Claims
1. An AI-powered intelligent generation method for voice-based follow-up questionnaires for chronic disease management, characterized in that: The method includes: Acquire user voice response data, perform speech recognition and terminology standardization processing on the voice response data to obtain standardized text data; Natural language understanding processing is performed on the standardized text data to extract key medical information and generate structured follow-up records; Based on the structured follow-up records, the user's overall compliance score is calculated; Based on a pre-defined directed decision graph of follow-up questionnaires, and with the comprehensive compliance score and the information coverage status in the structured follow-up records as inputs, the questionnaire nodes in the decision graph are dynamically adjusted to generate an adaptive questionnaire path for the current follow-up round. Following the adaptive questionnaire path, the next round of voice follow-up interaction will be conducted.
2. The AI-powered voice-based intelligent generation method for chronic disease management according to claim 1, characterized in that, The voice response data is subjected to speech recognition and terminology standardization processing, including: An ASR model with a pre-trained vocabulary of chronic disease medical terms was used to convert speech response data into initial text. The colloquial expressions in the initial text are mapped to standard medical terms or pharmacopoeia codes using a medical terminology mapping dictionary.
3. The AI-powered voice-based intelligent generation method for chronic disease management according to claim 1, characterized in that, Natural language understanding processing of the standardized text data includes: An NLU model, finely tuned based on chronic disease annotated corpus, is used to perform intent recognition and slot filling on the standardized text data. The extracted key slot information is automatically mapped to the preset public health follow-up template fields to generate structured follow-up records.
4. The AI-powered voice-based intelligent generation method for chronic disease management according to claim 1, characterized in that, Based on the structured follow-up records, the user's overall compliance score is calculated, including: The user self-reported behaviors extracted from the structured follow-up records are compared with the preset standard intervention plan, and the content consistency dimension score is calculated. Voice emotion recognition is performed on the user's voice response data, and text emotion classification is performed on the standardized text data to calculate the emotion tone dimension score; The content consistency dimension score and the sentiment tone dimension score are weighted and fused to obtain a comprehensive compliance score.
5. The AI-powered voice-based intelligent generation method for chronic disease management according to claim 1, characterized in that, Dynamically adjust the questionnaire nodes in the decision graph, including: When the structured follow-up record already contains the target information to be asked in a certain questionnaire node, that node is automatically skipped; When the overall compliance score is lower than a first preset threshold, add or activate follow-up question nodes associated with low compliance items; When the overall compliance score is higher than the second preset threshold and the relevant indicators meet the stability condition, the frequency of the corresponding questionnaire node is skipped or reduced, wherein the first preset threshold is less than the second preset threshold. When the system detects that a user's answer contains ambiguous, contradictory, or missing information, it automatically inserts a preset clarification question.
6. The AI-powered voice-based intelligent generation method for chronic disease management according to claim 5, characterized in that, When ambiguous, contradictory, or missing information is detected in a user's answer, a preset clarification follow-up question is automatically inserted, including: When a key information field in a user's answer is found to be empty or its value exceeds the preset normal range, it is determined that information is missing, and a supplementary query node for that field is inserted. When a discrepancy is detected between the user's current answer and the same indicator in the historical follow-up records, it is determined to be an information contradiction. A double query node is inserted for verification, and the indicator is marked as pending confirmation. When a user's answer is found to contain vague quantifiers or uncertain expressions, it is determined to be ambiguous information. A follow-up question node for specific quantification is inserted to guide the user to provide precise values or clear status.
7. The AI-powered voice-based intelligent generation method for chronic disease management according to claim 1, characterized in that, Following the adaptive questionnaire path, the next round of voice follow-up interaction will be performed, including: Obtain the results of the comprehensive compliance score and affective tone dimension analysis; When resistance or anxiety is detected in the emotional tone dimension analysis results, the speech pattern is switched to empathetic or encouraging expression; When a user is identified as actively cooperating and their overall compliance score is higher than a preset threshold, positive incentive messages are added.
8. The AI-powered voice-based intelligent generation method for chronic disease management according to claim 1, characterized in that, Also includes: The duration interval of the next follow-up cycle is automatically adjusted based on the overall compliance score and its trend in the current round. When the comprehensive compliance score is below the third preset threshold for several consecutive periods, the follow-up period is shortened. When the comprehensive compliance score is higher than the fourth preset threshold for multiple consecutive periods, the follow-up period is extended, wherein the third preset threshold is less than or equal to the first preset threshold, and the fourth preset threshold is greater than or equal to the second preset threshold.
9. The AI-powered voice-based intelligent generation method for chronic disease management according to claim 1, characterized in that, Also includes: The structured follow-up records, comprehensive compliance scores, and adaptive questionnaire path execution logs generated in each round of follow-up will be stored as historical data. Using the historical data, the NLU model, the weight parameters of each node in the directed decision graph, and the fusion strategy for the comprehensive compliance score are iteratively updated.
10. An AI-powered voice-based intelligent generation system for chronic disease management, characterized in that: The system is used to execute the AI-powered voice-based intelligent generation method for chronic disease management as described in any one of claims 1-9, the system comprising: The voice acquisition module is used to acquire the user's voice response data, and to perform voice recognition and terminology standardization processing on the voice response data to obtain standardized text data; The semantic parsing module is used to perform natural language understanding processing on the standardized text data, extract key medical information, and generate structured follow-up records; The compliance assessment module is used to calculate the user's overall compliance score based on the structured follow-up records; The questionnaire generation module is used to dynamically adjust the questionnaire nodes in the decision graph based on a preset follow-up questionnaire directed decision graph, and with the comprehensive compliance score and the information coverage status in the structured follow-up record as input, to generate an adaptive questionnaire path for the current follow-up round. The interactive execution module is used to execute the next round of voice follow-up interaction according to the adaptive questionnaire path.