Deep learning-based breast cancer rehabilitation question and answer method, device and equipment and medium

By using deep learning technology to analyze and fuse the features of multimodal data input by users, and combining graph attention neural networks and rule engines, conflict-free personalized breast cancer rehabilitation guidance is generated. This solves the problem that existing systems cannot combine professional medical knowledge with common sense, and provides safe and reliable personalized advice.

CN120878294APending Publication Date: 2025-10-31GANSU DAR HEALTH REHABILITATION HOSPITAL CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511292466.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing medical question-and-answer systems cannot effectively combine professional medical knowledge with common sense during the recovery process of breast cancer patients, resulting in the inability to provide rehabilitation guidance that is both professional and practical. This may affect postoperative recovery, especially when there is a potential conflict between medical advice and life experience.

Method used

Using a deep learning-based approach, multimodal feature analysis is performed on user-input text, voice, and lesion image data to generate a knowledge demand vector. Then, a graph attention neural network and a pre-built rule engine are used to perform weighted fusion and conflict detection of medical knowledge subgraphs and general knowledge fragments to generate conflict-free personalized rehabilitation guidance.

Benefits of technology

It achieves seamless integration of professional medical knowledge and common sense, providing personalized rehabilitation guidance that is both professional and accurate, as well as practical and life-oriented, ensuring safety and reliability, and avoiding potential conflicts and inapplicable advice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120878294A_ABST
    Figure CN120878294A_ABST
Patent Text Reader

Abstract

The invention relates to a deep learning-based breast cancer rehabilitation question and answer method, apparatus and device, and a medium. According to the method, patient demands are accurately captured through multi-modal feature analysis to generate knowledge demand vectors, and medical knowledge sub-graphs and general knowledge fragments are retrieved in parallel to form a knowledge basis; based on demand vector parameter dynamic weighting fusion of the two types of knowledge, a collaborative knowledge unit is constructed, and then importance distribution is performed on the medical knowledge by using a graph attention neural network to generate attention distribution; semantic integration is performed in combination with attention distribution and general knowledge to form fusion representation, and conflicts are arbitrated through a rule engine to ensure safety and reliability; finally, the conflict-free reasoning path and the patient demand vector are combined to generate professional, accurate and living personalized rehabilitation guidance, and seamless cooperation of medical preciseness and life practicability is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical rehabilitation technology, and in particular relates to a breast cancer rehabilitation question-and-answer method, device, equipment and medium based on deep learning. Background Technology

[0002] During the recovery process of breast cancer patients, they often face complex, interdisciplinary questions. Current medical question-and-answer systems tend to overemphasize specialized medical fields. When patients inquire about specific matters related to daily life management, these systems either provide mechanical, standard medical answers or refuse to answer altogether. While general-purpose dialogue systems can handle everyday questions, their lack of an authoritative medical knowledge framework poses safety risks when answering questions involving drug interactions or contraindications for recovery. This disconnect between specialized medical knowledge and common sense prevents patients from receiving both professional, reliable, and practical recovery guidance. This is especially true when medical advice potentially conflicts with everyday experience, such as when certain daily activities might affect postoperative recovery. Summary of the Invention

[0003] Therefore, it is necessary to provide a breast cancer rehabilitation question-and-answer method, device, equipment, and medium based on deep learning to address the above-mentioned technical problems.

[0004] Firstly, this application provides a deep learning-based question-answering method for breast cancer rehabilitation, including:

[0005] S1. Perform multimodal feature analysis on the user-input text data, voice data, and lesion image data to obtain the knowledge demand vector;

[0006] S2. Based on the knowledge demand vector, perform subgraph retrieval on the medical knowledge base to obtain medical knowledge subgraphs; based on the knowledge demand vector, perform fragment retrieval on the general knowledge base to obtain general knowledge fragments.

[0007] S3. Based on the parameter components in the knowledge demand vector, the medical knowledge subgraph and general knowledge fragments are weighted and fused to obtain a weighted fused knowledge unit.

[0008] S4. Based on weighted fusion knowledge units, a graph attention neural network is used to assign node importance weights to the medical knowledge subgraph, generating a node attention distribution.

[0009] S5. Based on node attention distribution and common knowledge fragments, knowledge is integrated through semantic fusion to generate a fused knowledge representation;

[0010] S6. Based on the fusion knowledge representation, conflict detection and arbitration are performed through a pre-built rule engine to obtain a conflict-free reasoning path;

[0011] S7. Based on the conflict-free reasoning path and knowledge requirement vector, generate response text through a text generator.

[0012] Secondly, this application also provides a deep learning-based breast cancer rehabilitation question-answering device for implementing the method described in the first aspect, the device comprising:

[0013] The multimodal feature parsing module is used to perform multimodal feature parsing on user-input text data, voice data, and lesion image data to obtain a knowledge requirement vector;

[0014] The knowledge retrieval module is used to perform subgraph retrieval on the medical knowledge base based on the knowledge demand vector to obtain medical knowledge subgraphs; and to perform fragment retrieval on the general knowledge base based on the knowledge demand vector to obtain general knowledge fragments.

[0015] The weighted fusion processing module is used to perform weighted fusion of medical knowledge subgraphs and general knowledge fragments based on the parameter components in the knowledge demand vector to obtain weighted fusion knowledge units.

[0016] The graph attention allocation module is used to allocate node importance weights to the medical knowledge subgraph based on weighted fusion knowledge units and through a graph attention neural network to generate node attention distribution;

[0017] The semantic fusion and integration module is used to integrate knowledge based on node attention distribution and common knowledge fragments through semantic fusion, and generate fused knowledge representation.

[0018] The conflict detection and arbitration module is used to detect and arbitrate conflicts based on fused knowledge representation and through a pre-built rule engine to obtain conflict-free reasoning paths.

[0019] The text generation module is used to generate response text based on conflict-free reasoning paths and knowledge requirement vectors through a text generator.

[0020] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement a deep learning-based breast cancer rehabilitation question-and-answer method as described in the first aspect.

[0021] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a deep learning-based breast cancer rehabilitation question-and-answer method as described in the first aspect.

[0022] The aforementioned deep learning-based breast cancer rehabilitation question-answering method, device, equipment, and medium accurately captures patient needs through multimodal feature analysis to generate a knowledge demand vector. Based on this vector, medical knowledge subgraphs and general knowledge fragments are retrieved in parallel to form a knowledge foundation. Collaborative knowledge units are constructed by dynamically weighting and fusing the two types of knowledge based on the demand vector parameters. Then, a graph attention neural network is used to allocate the importance of medical knowledge to generate an attention distribution. The attention distribution and general knowledge are semantically integrated to form a fusion representation. A rule engine arbitrates conflicts to ensure safety and reliability. Finally, a conflict-free reasoning path is combined with the patient demand vector to generate personalized rehabilitation guidance that is both professional and accurate, yet also practical and relevant to daily life, achieving a seamless synergy between medical rigor and practical application. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A flowchart illustrating a deep learning-based question-and-answer method for breast cancer rehabilitation provided by this invention;

[0025] Figure 2 This is a schematic diagram of the process of generating a conflict-free inference path in an optional embodiment of the present invention.

[0026] Figure 3 This is a schematic diagram of the structure of a breast cancer rehabilitation question-and-answer device based on deep learning provided by the present invention. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0028] refer to Figure 1 The document presents a flowchart illustrating a deep learning-based question-and-answer method for breast cancer rehabilitation provided in this application. The method includes the following steps:

[0029] S1. Perform multimodal feature analysis on the user-input text data, voice data, and lesion image data to obtain the knowledge demand vector.

[0030] Specifically, for text data, deep learning-based natural language processing techniques can be used for parsing. Word embedding techniques, such as Word2Vec or GloVe, map words in the text to a high-dimensional vector space to capture semantic relationships between words. Then, recurrent neural networks (RNNs) or their variants, Long Short-Term Memory (LSTM) and Gated Recurrent Units (GRUs), are used to encode the text sequence and obtain semantic feature representations. Simultaneously, convolutional neural networks (CNNs) can be used to extract local features of the text, such as n-gram grammatical features, and these text features are fused to form a feature vector for the text modality.

[0031] For speech data, the first step is speech signal preprocessing, including speech framing and feature extraction. Then, the extracted speech features are input into a deep neural network (DNN) or convolutional neural network (CNN) for speech recognition, converting the speech into corresponding text content. Afterward, the same text processing methods are used to extract semantic features from the recognized text, obtaining the semantic feature vector corresponding to the speech modality.

[0032] For lesion image data, convolutional neural networks (CNNs) can be used for image feature extraction. For example, CNN architectures such as VGG and ResNet can be used to perform multi-level convolution and pooling operations on lesion images to progressively extract low-level features (such as edges and textures) and high-level semantic features (such as the shape, size, and density of lesions). The final image feature vector is then used as the feature representation of the image modality.

[0033] After obtaining feature vectors from text, speech, and image modalities, they are fused to construct a knowledge demand vector. Considering the differences between modal features, multimodal fusion techniques can be employed, such as early fusion (directly concatenating different modal features at a low level before processing), mid-stage fusion (fusion occurring in the middle of feature extraction), or late-stage fusion (processing each modal feature separately before making a fusion decision). During the fusion process, attention mechanisms can be used to dynamically adjust the weights of different modal features, enabling the model to automatically learn the importance of different modalities to the current user's question, thus more accurately reflecting the user's knowledge needs. For example, by constructing an attention network, using each modal feature as input, calculating the attention weight of each modal feature, and then weighted summing the modal features according to their weights, a comprehensive knowledge demand vector is obtained. This vector can comprehensively represent the user's integrated knowledge needs across text, speech, and image modalities, providing a foundation for subsequent knowledge retrieval and fusion.

[0034] S2. Based on the knowledge demand vector, perform subgraph retrieval on the medical knowledge base to obtain medical knowledge subgraphs; based on the knowledge demand vector, perform fragment retrieval on the general knowledge base to obtain general knowledge fragments.

[0035] Specifically, based on the knowledge demand vector, searches are performed on both the medical knowledge base and the general knowledge base. Regarding subgraph retrieval within the medical knowledge base, which is organized as a graph structure, it contains numerous medical concepts (such as symptoms, drugs, and treatment methods) as nodes, and the relationships between these concepts (such as etiology, treatment methods, and drug interactions) as edges. First, the knowledge demand vector is matched for similarity with each node in the medical knowledge base. Methods such as cosine similarity and Euclidean distance can be used to calculate the similarity between the knowledge demand vector and the feature vectors of each node in the medical knowledge base, filtering out the medical concept nodes most relevant to the user's knowledge needs. Then, based on graph database query techniques (such as using the Cypher query language for the Neo4j graph database), a certain number of neighboring nodes and their connecting edges are expanded outwards from these relevant medical concept nodes, thereby constructing a subgraph of medical knowledge closely related to the user's question. For example, if a user's question involves dietary precautions for post-operative recovery of breast cancer, then by finding medical concept nodes related to "post-operative breast cancer" and "diet" through similarity matching, the query is expanded to obtain related nodes and relationship edges such as "suitable foods," "forbidden foods," and "nutritional balance" connected to these nodes, forming a subgraph around knowledge about post-operative dietary recovery of breast cancer.

[0036] For fragment retrieval in a general knowledge base, the general knowledge base can be a large-scale text corpus covering various aspects of daily life, encyclopedic knowledge, etc. The knowledge demand vector can be converted into a corresponding text query statement, and retrieval algorithms based on vector space models or semantic retrieval models based on deep learning (such as BERT-Siamese Network) can be used to retrieve the general knowledge base. The vector space model vectorizes the knowledge demand vector and the text fragments in the general knowledge base, ranks them by calculating the similarity between the two, and returns the set of general knowledge fragments that best match the knowledge demand vector. Semantic retrieval models based on deep learning can better understand the semantic information of the text and uncover general knowledge fragments that are semantically relevant to the user's question. For example, in the context of breast cancer rehabilitation, if a user asks for exercise advice during rehabilitation, the general knowledge base retrieval returns general knowledge fragments containing information such as the benefits of moderate exercise for physical recovery and simple rehabilitation exercises suitable for home use. These fragments can provide users with specific, practical advice.

[0037] S3. Based on the parameter components in the knowledge demand vector, the medical knowledge subgraph and general knowledge fragments are weighted and fused to obtain a weighted fused knowledge unit.

[0038] Specifically, the parameter components in the knowledge demand vector reflect the intensity of users' needs across different knowledge dimensions (such as medical expertise and relevance to daily life). Based on the values ​​of these parameter components, the weighting coefficients for the medical knowledge subgraph and general knowledge fragments are determined during the fusion process. For example, if the medical expertise parameter component in the knowledge demand vector is large, it indicates a high degree of reliance on specialized medical knowledge for the user's current problem; therefore, the medical knowledge subgraph is given greater weight during fusion. Conversely, if the relevance to daily life parameter component is dominant, the weight of the general knowledge fragments is increased.

[0039] A specific weighted fusion method can employ a linear weighted fusion strategy. First, the medical knowledge subgraph and general knowledge fragments are vectorized separately. For the medical knowledge subgraph, its structural information (such as the connections between nodes and edges) can be encoded using a graph neural network (such as a graph convolutional network GCN or a graph attention network GAT) to obtain the subgraph's structural feature vector. Simultaneously, the semantic feature vectors of each medical concept node in the subgraph are extracted (obtained through natural language processing methods based on text descriptions and other information from the medical knowledge base). The structural and semantic feature vectors are then fused to form a comprehensive feature vector for the medical knowledge subgraph. For general knowledge fragments, natural language processing techniques (such as using pre-trained language models like BERT to extract text semantic features) are used to convert each fragment into a semantic feature vector. Weights can be adjusted based on factors such as the fragment's source reliability and its relevance to the user's question to obtain a weighted semantic feature vector for the general knowledge fragment.

[0040] Then, according to the weight coefficients determined by the parameter components in the knowledge demand vector, the comprehensive feature vector of the medical knowledge subgraph and the weighted semantic feature vector of the general knowledge fragment are linearly weighted and summed. That is, the feature vector of the weighted fused knowledge unit is expressed as: V_fused = α × V_medical + β × V_general, where V_fused is the feature vector of the weighted fused knowledge unit, V_medical is the comprehensive feature vector of the medical knowledge subgraph, V_general is the weighted semantic feature vector of the general knowledge fragment, α and β are the weight coefficients of medical knowledge and general knowledge determined according to the parameter components of the knowledge demand vector, and α + β = 1. In this way, medical expertise and general life knowledge are fused according to the proportion of user needs, so that the fused knowledge unit can simultaneously take into account professionalism and practicality, providing a more comprehensive knowledge foundation for subsequent reasoning and answer generation.

[0041] S4. Based on weighted fusion knowledge units, a graph attention neural network is used to assign node importance weights to the medical knowledge subgraph, generating a node attention distribution.

[0042] Specifically, after obtaining the weighted fusion knowledge units, they are used as input to a graph attention neural network (GAT) to assign importance weights to each node in the medical knowledge subgraph. The core idea of ​​the graph attention neural network is to automatically learn the correlation between nodes through an attention mechanism, thereby determining the importance of each node in answering the user's question within the current knowledge requirement and the context of fused knowledge.

[0043] First, the node feature vectors of the medical knowledge subgraph (which can be previously extracted semantic feature vectors or comprehensive feature vectors, etc.) are input into the graph attention layer. Each node's feature vector is projected through a learnable linear transformation matrix to obtain the node's feature representation vector. Then, for each node in the subgraph, the attention coefficient between it and all other nodes is calculated. The formula for calculating the attention coefficient is: e_ij = a(Wh_i,Wh_j), where W is a learnable weight matrix, h_i and h_j are the feature vectors of node i and node j, respectively, and a is the attention calculation function, such as a single-layer neural network or a hyperbolic tangent function. This attention coefficient reflects the importance of node j to node i, that is, how much influence the information of node j has on determining the importance of node i.

[0044] Next, the calculated attention coefficients are normalized using the softmax function, ensuring that the sum of the attention coefficients between each node and its neighbors is 1. The normalized attention coefficients represent the attention weights between nodes, indicating the importance of information transmission between them in the graph structure. By employing a multi-head attention mechanism (i.e., running multiple independent attention calculation processes simultaneously and then concatenating or averaging the results), we can capture the correlations between nodes in different aspects, enriching the feature representation of the nodes.

[0045] After multi-layer processing by a graph attention neural network (which may include stacking multiple graph attention layers), the final result is the importance and weight allocation of each node, i.e., the node attention distribution. Each element in this node attention distribution vector corresponds to the importance weight of a node in the medical knowledge subgraph. The larger the weight value, the more important the node is in the current user question and the context of integrated knowledge, and the more crucial the knowledge content it contains is for generating accurate and reasonable answers. For example, in the knowledge subgraph of breast cancer rehabilitation, if the user question focuses on the side effects of a certain drug and how to deal with them, then after processing by the graph attention neural network, medical concept nodes related to the drug and its side effects (such as drug name nodes, side effect symptom nodes, treatment measures nodes, etc.) will receive higher attention weights, while other nodes with weaker associations (such as unrelated disease nodes) will have relatively lower weights. This highlights key medical knowledge nodes and provides accurate node weights for subsequent knowledge integration and reasoning.

[0046] S5. Based on node attention distribution and common knowledge fragments, knowledge is integrated through semantic fusion to generate a fused knowledge representation.

[0047] Specifically, after obtaining the node attention distribution and general knowledge fragments, the two are semantically fused to generate a fused knowledge representation. First, the nodes in the medical knowledge subgraph are weighted according to the node attention distribution. The importance weight of each node (i.e., the corresponding value in the node attention distribution) is multiplied by the semantic feature vector of the node to obtain a weighted set of medical knowledge node feature vectors. This step allows more important nodes in the medical knowledge subgraph to have a greater influence in the subsequent fusion process, highlighting their knowledge value.

[0048] Then, the semantic feature vectors of the general knowledge fragments (which have been previously extracted and weighted) are fused with the weighted set of medical knowledge node feature vectors. The fusion can be achieved using an attention-based semantic fusion mechanism. A fusion attention model is constructed, taking the medical knowledge node feature vectors and the general knowledge fragment feature vectors as input, and calculating the semantic relevance weights between them. For example, mechanisms such as dot product attention or additive attention can be used to measure the semantic similarity and correlation between each general knowledge fragment and the medical knowledge node. Based on the calculated relevance weights, the feature vectors of the general knowledge fragments are weighted and summed to obtain the general knowledge fusion vector semantically related to the medical knowledge subgraph.

[0049] Next, the weighted set of medical knowledge node feature vectors is aggregated using operations such as averaging, summing, or max pooling to obtain a comprehensive medical knowledge representation vector. This medical knowledge representation vector is then concatenated with or element-wise summed with a general knowledge fusion vector to generate the final fused knowledge representation vector. This fused knowledge representation vector contains both attention-weighted professional medical knowledge and semantically related general knowledge, comprehensively reflecting the cross-domain knowledge required for the user's problem and providing a unified and rich knowledge representation foundation for subsequent conflict detection and reasoning.

[0050] S6. Based on the fusion knowledge representation, conflict detection and arbitration are performed through a pre-built rule engine to obtain a conflict-free reasoning path.

[0051] Specifically, a fused knowledge representation vector is used and input into a pre-built rule engine for conflict detection and arbitration. The pre-built rule engine contains a series of reasoning rules based on medical knowledge and common sense. These rules can be summarized by domain experts (including medical experts, nutritionists, rehabilitation therapists, etc.) based on clinical experience and authoritative knowledge, and are used to determine whether there are potential conflicts or contradictions between different pieces of knowledge.

[0052] First, various conflict detection rule patterns are defined in the rule engine. For example, regarding drug interactions, if the fused knowledge representation contains both drugs A and B, and a rule exists in the medical knowledge rule base stating that "simultaneous use of drug A and drug B may lead to serious side effects X," then conflict detection is triggered. The rule engine will identify this potential drug interaction conflict based on the semantic information in the fused knowledge representation. Similarly, for situations where medical advice conflicts with life experience, such as when the fused knowledge representation contains "medical advice that breast cancer patients should avoid a certain type of exercise" and "general life knowledge suggests that a certain type of exercise is beneficial for physical recovery," the conflict rules in the rule engine will identify and mark these conflicts.

[0053] Upon detecting a conflict, the rules engine initiates an arbitration mechanism. The arbitration process follows pre-defined priority rules and conflict resolution strategies. Priority rules can be set based on factors such as the reliability of the knowledge source (e.g., medical knowledge takes precedence over general knowledge) and the degree of medical risk (e.g., medical advice that may affect life safety is prioritized). For example, in cases of drug interaction conflicts, authoritative drug interaction guidelines from the medical knowledge base are prioritized, excluding or correcting advice from conflicting general knowledge fragments. In cases of conflicts between medical rehabilitation advice and life experience, if the medical advice is supported by sufficient clinical evidence, it takes precedence; if some flexibility is possible, the rules engine will generate a compromise suggestion based on the patient's specific rehabilitation stage, physical condition, and other factors, combined with reasonable aspects of common sense, such as moderately referencing helpful advice from life experience within medically permissible limits. Through a series of rule matching, conflict detection, and arbitration operations, a conflict-free reasoning path is ultimately obtained. The knowledge content within this path is coordinated and logically consistent, providing users with a safe and reliable framework for rehabilitation guidance.

[0054] S7. Based on the conflict-free reasoning path and knowledge requirement vector, generate response text through a text generator.

[0055] Specifically, the text generator can adopt a deep learning-based sequence-to-sequence (Seq2Seq) model, such as a neural network with an encoder-decoder structure. The encoder encodes the knowledge content (including the fusion information of medical knowledge and general knowledge) and the knowledge requirement vector in the conflict-free reasoning path as input to generate a context representation vector. The decoder then generates a sequence of words for the response text step by step based on this context representation vector.

[0056] To improve the quality and accuracy of text generation, the encoder can use a bidirectional recurrent neural network (such as a bidirectional LSTM or bidirectional GRU) to encode each knowledge node and common knowledge fragment in the conflict-free reasoning path, fully capturing the dependencies and semantic information of the knowledge content. Simultaneously, the knowledge demand vector is used as an additional input feature and fused with the knowledge content for encoding, enabling the decoder to better understand the user's key needs and focus. In the decoder, an attention mechanism (such as Bahdanau attention or Luong attention) is employed to dynamically focus on the most relevant information in the conflict-free reasoning path for each word generated in the response text, improving the coherence and accuracy of text generation.

[0057] Furthermore, a template-guided mechanism can be introduced during the text generation process. Based on different types of breast cancer rehabilitation questions (such as medication consultation, dietary advice, exercise guidance, etc.), corresponding text templates are preset. These templates contain the basic framework and key information points for answering the questions. When generating response text, the text generator can refer to these templates, filling in the knowledge content from the conflict-free reasoning path into the corresponding positions of the templates, while supplementing and refining it with freely generated text content to generate response text that is both standardized and personalized. For example, for consultation questions about drug side effects, the template might include information points such as "drug name, side effect manifestations, coping measures, and when to seek medical attention." The text generator fills in these information points based on the knowledge content in the reasoning path and organizes them into a complete response text in fluent and natural language, such as, "The medication XX you are using may cause side effects such as XX and XX. In daily life, if mild XX symptoms occur, they can be relieved by XX methods, but if the symptoms continue to worsen or serious conditions such as XX occur, please seek medical attention immediately."

[0058] In this way, the generated response text can accurately reflect the knowledge content in the conflict-free reasoning path, while meeting the personalized needs expressed by users in the knowledge demand vector, providing professional, considerate, and practical rehabilitation guidance for breast cancer patients, and helping them better cope with various problems in the rehabilitation process.

[0059] The aforementioned deep learning-based question-answering method for breast cancer rehabilitation accurately captures patient needs through multimodal feature parsing to generate a knowledge demand vector. Based on this vector, it retrieves medical knowledge subgraphs and general knowledge fragments in parallel to form a knowledge foundation. It then constructs collaborative knowledge units by dynamically weighting and fusing the two types of knowledge based on the demand vector parameters. Furthermore, it uses a graph attention neural network to allocate the importance of medical knowledge to generate an attention distribution. Finally, it integrates the attention distribution with general knowledge semantically to form a fusion representation, and uses a rule engine to arbitrate conflicts to ensure safety and reliability. Ultimately, it combines a conflict-free reasoning path with the patient demand vector to generate personalized rehabilitation guidance that is both professional and accurate, yet also practical and relevant to daily life, achieving a seamless synergy between medical rigor and everyday practicality.

[0060] In one optional embodiment, multimodal feature parsing is performed on user-input text data, voice data, and lesion image data to obtain a knowledge demand vector, including the following steps:

[0061] S11. Using a language model trained on breast cancer rehabilitation corpus, medical terminology is identified in the text data to obtain the medical terminology identification results; the medical terminology density index is calculated based on the medical terminology identification results.

[0062] Specifically, a large amount of corpus related to breast cancer rehabilitation was collected, including medical literature, rehabilitation guidelines, medical records, and expert consultation dialogues. This corpus underwent preprocessing, such as word segmentation, stop word removal, and part-of-speech tagging. Then, a language model based on a recurrent neural network (RNN) or transformer architecture was built using deep learning frameworks (such as TensorFlow or PyTorch). The preprocessed corpus was input into the model for training. By adjusting the model parameters, it was made able to accurately predict the next word in the text, while simultaneously learning professional terminology and semantic features in the field of breast cancer rehabilitation.

[0063] The input text data is processed using a pre-trained language model. The text is input word by word into the model, which outputs a context-dependent feature vector for each word. By analyzing these feature vectors, medical terms are identified. A threshold for medical term recognition can be set; when the similarity between the feature vector of a word and the feature vector of a known medical term exceeds the threshold, it is determined to be a medical term. All identified medical terms and their positions within the text are recorded.

[0064] Next, the medical terminology density index is calculated. This involves counting the number of medical terms in the text, as well as the total number of words. The medical terminology density index can be calculated using the following formula: Medical Terminology Density Index = (Number of Medical Terms / Total Number of Words) × 100%. This index reflects the proportion of medical content in the text; a higher value indicates a greater emphasis on medical expertise, suggesting a higher potential demand for professional medical answers from users.

[0065] S12. Extract the acoustic features of the speech data, input the acoustic features into a classifier trained on a clinical anxiety speech dataset, and generate an anxiety index that represents the user's anxiety level.

[0066] Specifically, professional audio processing tools (such as the librosa library) are used to preprocess the speech data, including speech framing and background noise removal. Then, various acoustic feature parameters are extracted, such as pitch (fundamental frequency), timbre (Mel-frequency cepstral coefficients, MFCC), speech rate (estimated by analyzing the rate of energy change in the speech signal), and loudness (sound pressure level). These feature parameters are combined into an acoustic feature vector to characterize the acoustic properties of the speech data.

[0067] A large dataset of clinical anxiety speech samples was collected, containing voice samples from patients clinically diagnosed with anxiety symptoms, along with corresponding anxiety level labels (e.g., mild, moderate, severe anxiety). Acoustic features were extracted from these speech samples to obtain feature vectors and corresponding labels for training. A classifier was constructed using machine learning algorithms (e.g., Support Vector Machine (SVM), Random Forest, or Convolutional Neural Network (CNN) in deep learning). The extracted acoustic feature vectors and anxiety level labels were input into the classifier for training. By adjusting the classifier's parameters, it was made capable of accurately predicting anxiety levels based on acoustic features.

[0068] The extracted acoustic feature vectors of the speech data are input into a trained classifier, which outputs an anxiety index that characterizes the user's anxiety level. This index can be a continuous value (e.g., between 0 and 1, representing anxiety levels from none to extremely severe) or a classification label (e.g., mild, moderate, severe anxiety). The anxiety index reflects the user's underlying anxious psychological state in their speech expression, which helps in generating more reassuring and patient explanations for anxious users in subsequent response generation, or in prioritizing the provision of key, directly relevant rehabilitation information to alleviate the user's anxiety.

[0069] S13. For lesion image data, the rehabilitation stage is identified through a medical image recognition model to obtain the rehabilitation stage code.

[0070] Specifically, a large amount of breast cancer lesion image data was collected. These images should cover the characteristics of lesions at different recovery stages and be labeled by professional doctors to clearly define the recovery stage corresponding to each image (such as early postoperative recovery, intermediate recovery, and near-cure stages). A medical image recognition model based on a convolutional neural network (CNN) was constructed using a deep learning framework. The labeled lesion image data was input into the model for training. By adjusting the parameters of the model, such as convolutional layers, pooling layers, and fully connected layers, it was made able to automatically learn the feature representations of lesion images at different recovery stages, such as the changing patterns of lesion size, shape, boundary clarity, and internal tissue density.

[0071] The input lesion image data undergoes preprocessing operations, such as image resizing and pixel value normalization, to meet the model's input requirements. The preprocessed image is then input into a trained medical image recognition model. The model automatically extracts image features and performs feature transformation and abstraction through its internal multi-layer neural network structure. Finally, the model outputs a rehabilitation stage code, a predefined label or numerical value used to uniquely identify the rehabilitation stage of the lesion. For example, rehabilitation stage codes can be defined as numbers such as 0, 1, and 2, each corresponding to a different rehabilitation stage. This rehabilitation stage code provides crucial information for subsequent integration of user medical status information, helping to filter out medical and general knowledge closely related to the user's current rehabilitation stage during knowledge retrieval and fusion, thus improving the relevance and practicality of the responses.

[0072] S14. Construct a knowledge demand vector based on medical terminology density index, anxiety index, and rehabilitation stage coding.

[0073] Specifically, through analysis of extensive user feedback data and actual rehabilitation Q&A scenarios, the influence of different features (medical terminology density index, anxiety index, and rehabilitation stage coding) on ​​users' knowledge needs was statistically analyzed. For example, the analysis found that the medical terminology density index was highly correlated with users' need for professional medical knowledge, the anxiety index was strongly correlated with users' need for emotional support and reassurance in responses, and the rehabilitation stage coding was closely related to users' need for specific medical treatment and rehabilitation guidance plans. Based on these correlation analysis results, a weight coefficient was assigned to each parameter. The weight coefficients can be determined using regression analysis methods in machine learning, or adjusted and optimized through expert scoring and trial and error. For example, the weight of the medical terminology density index was set to 0.4, the weight of the anxiety index to 0.3, and the weight of the rehabilitation stage coding to 0.3. These weight coefficients reflect the relative importance of each parameter in constructing the knowledge need vector.

[0074] Then, the parameters are normalized. Since the medical terminology density index is a percentage value (0% to 100%), the anxiety index may be a continuous value (e.g., 0 to 1), and the rehabilitation stage code is a discrete integer label, their numerical ranges and types differ. To enable comprehensive calculation, each parameter is normalized, mapping it to the same numerical range (e.g., 0 to 1). Normalization methods can include min-max normalization or Z-score normalization. For example, the medical terminology density index can be converted to a value between 0 and 1 by dividing by 100; for the anxiety index, if its original range is 0 to 1, it can remain unchanged or be adjusted according to the actual distribution; for the rehabilitation stage code, it can be mapped to equally spaced values ​​between 0 and 1 according to the order of the rehabilitation stages, such as rehabilitation stage code 0 mapping to 0, code 1 mapping to 0.5, code 2 mapping to 1, and so on.

[0075] Finally, the normalized parameter values ​​are multiplied by their respective weighting coefficients, treating each parameter as a separate dimension to preserve its independence. This allows the knowledge demand vector to be represented as a multi-dimensional vector, such as V = [Normalized medical terminology density index × 0.4, Normalized anxiety index × 0.3, Normalized rehabilitation stage code × 0.3]. This representation more clearly reflects the contribution of each parameter to the knowledge demand, and allows for independent weight adjustments and analysis of each dimension during subsequent knowledge retrieval and fusion processes.

[0076] In one optional embodiment, based on the parameter components in the knowledge demand vector, a weighted fusion of the medical knowledge subgraph and general knowledge fragments is performed to obtain a weighted fused knowledge unit, including the following steps:

[0077] S21. Linearly combine the medical terminology density index and rehabilitation stage encoding in the knowledge demand vector to obtain a linear combination representation; input the linear combination representation into the S-type activation function for normalization to obtain the medical knowledge weight coefficient with a value between 0 and 1, and subtract the medical knowledge weight coefficient from 1 to obtain the general knowledge weight coefficient.

[0078] Specifically, the weights of these two parameters in the linear combination can be determined based on preliminary analysis of large amounts of user data and expert opinions. For example, suppose analysis shows that the medical terminology density index is more indicative of medical knowledge demand, so its weight is set to 0.6, while the weight of the rehabilitation stage code is set to 0.4. The linear combination formula can be expressed as: Linear combination = (Medical terminology density index × 0.6) + (Rehabilitation stage code × 0.4). Here, the medical terminology density index is a normalized value between 0 and 1, and the rehabilitation stage code is also preprocessed to a value between 0 and 1 (e.g., 0.3 for early rehabilitation stage, 0.6 for mid-stage, and 0.9 for late-stage, etc.).

[0079] The resulting linear combination representation is then input into a sigmoid activation function (such as the Logistic function) for normalization. The formula for the sigmoid activation function is: Here, k is the steepness parameter of the function, and x0 is the midpoint parameter of the function. Appropriate values ​​for k and x0 are set according to actual needs, for example, k = 10 and x0 = 0.5, so that the function has a significant gradient change near x0, compressing the linear combination representation into the interval between 0 and 1, thus obtaining the medical knowledge weight coefficient. Then, this medical knowledge weight coefficient is subtracted from 1 to obtain the general knowledge weight coefficient. For example, if the normalized medical knowledge weight coefficient is 0.7, then the general knowledge weight coefficient is 0.3. These two weight coefficients reflect the relative importance of medical knowledge and general knowledge under the current user needs, respectively.

[0080] S22. Combining the weight coefficients of medical knowledge, perform a linear transformation on the medical knowledge subgraph to obtain a weighted medical knowledge representation.

[0081] Specifically, the medical knowledge subgraph is stored in a graph structure, containing multiple nodes and edges, each with a corresponding feature vector. First, the feature vector of each node in the medical knowledge subgraph is represented as a matrix. Assuming the dimension of each node's feature vector is d, and the subgraph has n nodes, the node feature matrix of the entire subgraph can be represented as an n×d matrix. Similarly, the edge features can also be represented as a similar matrix.

[0082] Then, the node and edge feature matrices are weighted using a linear transformation. This linear transformation can be achieved through matrix multiplication, i.e., multiplying the corresponding feature matrix by the medical knowledge weight coefficient. For example, for the node feature matrix, the weighted node feature matrix is ​​the original node feature matrix multiplied by the medical knowledge weight coefficient; the edge feature matrix is ​​processed similarly. In this way, the feature vector of each node and edge is assigned a corresponding weight, so that in subsequent processing, the overall representation of the medical knowledge subgraph can reflect its importance in the user's current needs. For example, if the medical knowledge weight coefficient is high, it indicates that the user has a strong demand for medical expertise, and the weighted medical knowledge subgraph features will be more prominent, which is beneficial for highlighting the influence of medical knowledge in the subsequent fusion process.

[0083] S23. Combining the general knowledge weight coefficients, embedding spatial projection is performed on the general knowledge fragments to obtain a weighted general knowledge representation.

[0084] Specifically, general knowledge fragments are text fragments, each with its own semantic content. These text fragments can be converted into fixed-dimensional semantic vectors using pre-trained semantic embedding models (such as BERT, Word2Vec, etc.).

[0085] Next, these embedding vectors are weighted using a general knowledge weight coefficient. Specifically, the semantic embedding vector of each general knowledge fragment is multiplied by the general knowledge weight coefficient. For example, for a semantic embedding vector *v* of a general knowledge fragment, the weighted vector is *v* × the general knowledge weight coefficient. In this way, each general knowledge fragment is assigned a corresponding weight, reflecting its relative importance in the context of the user's current needs. If the general knowledge weight coefficient is low, it means that the user's need for general knowledge is relatively weak; the magnitude of the weighted semantic vector of the general knowledge fragment will be smaller, and its influence in subsequent fusion processes will also be correspondingly reduced.

[0086] S24. Based on weighted medical knowledge representation and weighted general knowledge representation, a weighted fused knowledge unit is obtained by fusion processing through vector addition.

[0087] Specifically, for the medical knowledge subgraph, the weighted node and edge feature matrices can be recombine into a new graph structure representation, or their feature vectors can be aggregated (e.g., by averaging, summing, etc.) to obtain a comprehensive medical knowledge vector representation. For example, averaging all weighted node feature vectors yields a d-dimensional comprehensive medical knowledge vector. For general knowledge fragments, the same aggregation operation is performed on all weighted semantic embedding vectors to obtain a comprehensive general knowledge vector.

[0088] To merge the two, they can be projected onto the same semantic space. For example, a linear transformation matrix can be designed to map the medical knowledge vector from its original dimension d to the 768-dimensional space of the general knowledge vector. Assuming the medical knowledge vector is mapped to the 768-dimensional space, this transformation is achieved through matrix multiplication. That is, a d×768 linear transformation matrix W is constructed, and the comprehensive medical knowledge vector V_medical (d-dimensional) is multiplied by matrix W to obtain the transformed medical knowledge vector V_medical' = V_medical×W. Here, the dimension of V_medical' is 768, consistent with the dimension of the general knowledge vector.

[0089] Then, the transformed weighted medical knowledge representation vector V_medical' and the weighted general knowledge representation vector V_general are added element-wise to obtain the fused vector V_fused = V_medical' + V_general. This fused vector V_fused is the weighted fused knowledge unit, which integrates medical and general knowledge, with each part assigned a corresponding weight according to user needs. This vector will be used in subsequent reasoning steps, such as as input to a graph attention neural network, or to generate the final response content.

[0090] In practical applications, the linear transformation matrix W can be determined as follows: During the training phase, prepare some known sample data containing medical knowledge and general knowledge, along with corresponding expected results after fusion. Extract medical knowledge vectors and general knowledge vectors from these sample data respectively, and then train the matrix W using optimization algorithms such as gradient descent to minimize the error between the fused vector V_fused and the expected result. During the inference phase, directly use the trained matrix W to transform the new medical knowledge vector, and then add it to the general knowledge vector to obtain the fused vector.

[0091] In one optional embodiment, based on weighted fusion knowledge units, a graph attention neural network is used to assign node importance weights to the medical knowledge subgraph to generate a node attention distribution, including the following steps:

[0092] S31. Based on the weighted medical knowledge representation in the weighted fusion knowledge unit, the node feature vectors of the medical knowledge subgraph are extracted through the graph embedding layer to obtain the initial feature matrix of the nodes.

[0093] Specifically, the weighted medical knowledge representation is a vector that integrates the weighted information of the user's medical knowledge needs and the features of the medical knowledge subgraph. The graph embedding layer can be a neural network structure used to map each node in the graph to a low-dimensional vector space while preserving the local and global structural information of the nodes. In practice, the embedding layers of Graph Convolutional Networks (GCNs) or Graph Attention Networks (GATs) can be used as the graph embedding layer.

[0094] The structural information of the medical knowledge subgraph (such as the adjacency matrix) and the original features of the nodes (such as word vectors for medical terms and rehabilitation stages) are input into the graph embedding layer. The graph embedding layer calculates the initial feature vector for each node through a series of linear transformations and activation functions. These initial feature vectors constitute the node's initial feature matrix. For example, for a medical knowledge subgraph containing n nodes, if the dimension of the initial feature vector of each node is d, then the shape of the node's initial feature matrix is ​​n×d.

[0095] S32. Based on the initial feature matrix of the nodes, the correlation between nodes is calculated through a multi-head attention mechanism, and the attention coefficient matrix is ​​obtained according to the correlation between nodes.

[0096] Specifically, the multi-head attention mechanism captures different aspects of the correlation between nodes through multiple parallel attention heads. Each attention head independently calculates the attention coefficients between nodes, and then the outputs of multiple attention heads are concatenated or averaged to obtain the final attention coefficient matrix. The specific calculation steps are as follows:

[0097] For each attention head, the initial feature matrix H (of shape n×d) of the node is transformed by two independent linear transformations to obtain the query matrix Q and the key matrix K (both of shape n×d). k ), where d k This is the dimension of each attention head. A linear transformation can be achieved through the weight matrix W. Q and W K Achieve, i.e.: Q = HW Q K = HW K .

[0098] Calculate the dot product between the query matrix and the key matrix to obtain the original attention score matrix A0 (n×n), i.e.: A0 = QK T .

[0099] Divide the original attention score matrix by d k To perform scaling, preventing gradient vanishing or exploding, i.e.:

[0100] The scaled attention score matrix is ​​passed through a non-linear activation function (such as LeakyReLU) to introduce non-linear properties, i.e.: A2 = LeakyReLU(A1).

[0101] The attention score matrix A2 of each attention head is concatenated or averaged to obtain the final attention coefficient matrix.

[0102] S33. Based on the attention coefficient matrix, normalized attention weights are obtained by row-wise softmax normalization.

[0103] Specifically, row-wise softmax normalization involves performing a softmax operation on each row of the attention coefficient matrix, ensuring that the sum of the elements in each row is 1. This step ensures that the attention weights of each node can be represented as a probability distribution, reflecting the relative importance of that node to other nodes. The specific calculation is as follows:

[0104] For each row a of the attention coefficient matrix A i Calculate softmax:

[0105]

[0106] Normalized attention weight matrix A norm each line This represents the attention weight of node i to all other nodes.

[0107] S34. Based on the normalized attention weights and the initial feature matrix of the nodes, feature fusion processing is performed through weighted aggregation operation to obtain the node attention distribution.

[0108] Specifically, the weighted aggregation operation combines the feature vector of each node with its attention weights to other nodes to obtain a new feature representation. Specifically, for each node i, its initial feature vector is weighted and summed with its attention weights to other nodes to obtain a new feature vector. The specific calculation is as follows:

[0109] The normalized attention weight matrix A norm Performing matrix multiplication with the initial feature matrix H of the nodes yields a new feature matrix H', i.e., H' = A. norm H.

[0110] Each row vector h' in the new feature matrix H' i Let i represent the new feature vector of node i, which is obtained by weighting the initial feature vector of node i with its attention weights to other nodes.

[0111] The node attention distribution can be represented as a new feature matrix H', where the eigenvector of each node reflects its relative importance in the graph and its correlation with other nodes.

[0112] refer to Figure 2 In one optional embodiment, based on fused knowledge representation, conflict detection and arbitration are performed through a pre-built rule engine to obtain a conflict-free reasoning path, including the following steps:

[0113] S41. Based on medical entity nodes and life concept nodes in the fusion knowledge representation, semantic conflict detection is performed through cosine similarity calculation to obtain a set of conflict node pairs.

[0114] Specifically, the fused knowledge representation is a vector representation that integrates medical knowledge and general knowledge, containing feature vectors of medical entity nodes (such as drugs, symptoms, and treatments) and life concept nodes (such as daily activities and dietary habits). To detect semantic conflicts between these nodes, the cosine similarity between each pair of medical entity nodes and life concept nodes is calculated. The specific calculation is as follows:

[0115] For medical entity node i and life concept node j, their feature vectors are v respectively. i and v j The formula for calculating the cosine similarity sim(i,j) is:

[0116]

[0117] Among them, v i ·v j It is the dot product of vectors, ||v i ||||v j || represents the L2 norm of the vector.

[0118] Set a similarity threshold θ (e.g., θ = 0.7). If the cosine similarity of a pair of nodes is greater than θ, then the pair of nodes is considered to have a possible semantic conflict and is added to the set of conflicting node pairs C.

[0119] S42. Based on the conflict node pair set and the rehabilitation stage encoding, rule retrieval is performed through a conditional expression matcher to obtain the matching taboo rule identifier.

[0120] Specifically, the pre-built rule engine stores a large number of contraindication rules related to breast cancer rehabilitation. Each rule has a unique identifier and contains a conditional expression and corresponding weight information. The conditional expression describes under what circumstances the rule applies; for example, a specific combination of medical entities and life concepts might conflict at a certain stage of rehabilitation. The specific steps for rule retrieval can be as follows:

[0121] Traverse the set C of conflicting node pairs. For each pair of conflicting nodes (i,j), extract its medical entity type t. i Life concept type t j And rehabilitation phase codes.

[0122] Searching for conditional expressions and (t) in the rule engine i ,t j The rules for matching ,s). For example, the conditional expression for a rule could be: IF Medical Entity Type = t i AND Life Concept Type = t j AND recovery phase = s, THEN taboo rule ID = XXX. Collect all matching rule identifiers into a set R.

[0123] S43. Based on the taboo rule identifier, the arbitration weight is allocated through the rule priority decision-maker to obtain the medical constraint weight and the life advice weight.

[0124] Specifically, the rule prioritization decision-maker assigns arbitration weights based on the rule's priority and type. Rule priorities can be pre-defined by domain experts, for example, based on the degree of medical risk, the strength of clinical evidence, etc. Rule types are divided into medical constraint rules and lifestyle recommendation rules. The specific allocation steps can be as follows:

[0125] For each matching rule r∈R, obtain its priority p. r and type r (Medical restrictions or lifestyle advice).

[0126] Calculate the average priority of medical constraint rules:

[0127]

[0128] Among them, R medical It is a subset of medical constraint rules, |R medical | Represents a subset R of medical constraint rules medical The number of rules contained therein.

[0129] Calculate the average priority of life advice rules:

[0130]

[0131] Among them, R lifestyle It is a subset of life advice rules, |R lifestyle | Represents a subset R of life advice rules lifestyle The number of rules contained therein.

[0132] Convert priorities into weights, for example, through normalization.

[0133] S44. Use the attention weights in the node attention distribution as the weight values ​​of the corresponding nodes in the medical knowledge subgraph, delete the nodes in the medical knowledge subgraph whose weight values ​​are lower than the medical constraint weights and their corresponding associated edges, and obtain the filtered knowledge subgraph.

[0134] Specifically, the node attention distribution is a vector, where each element represents the importance weight of the corresponding node in the medical knowledge subgraph. The specific selection steps are as follows: For each weight value 'a' in the node attention distribution... i With medical constraint weight w medical Compare them. For each node i, if a... i <w medical If the node is not found, then remove the node and its associated edges from the medical knowledge subgraph. Retain the remaining nodes and their associated edges to form the filtered knowledge subgraph.

[0135] S45. Based on the node attention distribution, assign weights to the knowledge points associated with each node in the medical knowledge subgraph in the general knowledge fragment; delete the knowledge points in the general knowledge fragment whose weight is greater than the weight of life advice and which are associated with the node pairs in the set of conflicting node pairs, and obtain the filtered general knowledge fragment.

[0136] Specifically, for each knowledge point in the general knowledge fragment, determine its association with nodes in the medical knowledge subgraph. This can be achieved through semantic similarity calculation (such as cosine similarity) or predefined association rules. Based on the node attention distribution, assign the attention weights of the associated nodes to the corresponding knowledge points in the general knowledge fragment. Iterate through each knowledge point in the general knowledge fragment, checking whether it is associated with node pairs in the set of conflicting node pairs C, and whether its weight is greater than the weight w of the life advice. lifestyle If a knowledge point meets the above conditions, it is removed from the general knowledge fragment. The remaining knowledge points are retained to form the filtered general knowledge fragment.

[0137] S46. Generate a conflict-free reasoning path based on the filtered knowledge subgraph and the filtered general knowledge fragment.

[0138] Specifically, the filtered knowledge subgraphs and general knowledge fragments are integrated to form a unified knowledge graph. Graph traversal algorithms (such as depth-first search or breadth-first search) are used to extract reasoning paths relevant to the user's question from the knowledge graph. It is ensured that the reasoning paths do not contain any conflicting node pairs from the set C, thus guaranteeing the path's non-conflictability. The extracted reasoning paths are then sorted and optimized, and the most relevant and reasonable path is selected as the final conflict-free reasoning path.

[0139] In one optional embodiment, a response text is generated by a text generator based on a conflict-free reasoning path and a knowledge requirement vector, including the following steps:

[0140] S51. Based on the conflict-free reasoning path, the text skeleton is constructed through a sequence-to-sequence generation model to obtain the initial response draft.

[0141] Specifically, a large-scale medical question-answering dataset was used to train a sequence-to-sequence (Seq2Seq) generative model. This model is based on an encoder-decoder architecture, where the encoder encodes the sequence of knowledge nodes in a conflict-free reasoning path into a fixed-length context vector, and the decoder generates a text sequence based on this vector. During training, an attention mechanism is employed to ensure the decoder focuses on the most relevant part of the reasoning path when generating each word. For the breast cancer rehabilitation domain, the dataset contains numerous question-answer pairs relevant to this field. After multiple rounds of iterative training, the model parameters were optimized to accurately capture domain-specific expressions and logical connections. Faced with a new conflict-free reasoning path input, the model generates a logically clear and coherent initial draft response based on its learned knowledge.

[0142] Given the specialized and rigorous nature of the medical field, a domain-specific optimization was performed on top of the general Seq2Seq model. A professional terminology dictionary related to breast cancer rehabilitation was introduced, and the model's vocabulary was expanded and updated to ensure accurate generation of professional terms such as "adjuvant chemotherapy after breast cancer surgery" and "radiotherapy dosage." Simultaneously, by incorporating domain expert knowledge, rules were applied to the text generated by the model to ensure the standardization of drug usage and dosage descriptions and the rationality of the order of rehabilitation recommendations, thereby improving the accuracy and credibility of the initial draft response in terms of professional knowledge.

[0143] S52. Based on the anxiety index in the knowledge demand vector, a reassuring statement is inserted into the initial response draft using an emotion template matcher to obtain an emotion-enhanced response text.

[0144] Specifically, an emotional template library is built, encompassing various emotional comfort and encouragement templates suitable for users with different levels of anxiety. The templates are designed based on psychological principles of reassurance, such as expressing understanding, offering hope, and emphasizing professional support. Upon receiving the anxiety index from the knowledge needs vector, the emotional template matcher activates, accurately matching the appropriate emotional template based on the anxiety index's numerical range. For example, when the anxiety index is in the moderate range, a template like "I understand you're a little worried right now, which is normal. In fact, many patients have achieved good results through standardized rehabilitation and positive mindset adjustment." is matched. Natural language processing technology is used to smoothly insert emotional statements into the initial response draft, ensuring logical coherence and natural emotional transitions, enhancing the emotional warmth and humanistic care of the response.

[0145] In addition to preset templates, a personalized emotional statement generation module based on sentiment analysis and language generation can be introduced. This module analyzes the emotional tone of the user's historical dialogue records and current questions to uncover their underlying emotional needs. Combining anxiety levels, it uses deep learning language models to generate emotional statements tailored to the individual user. For example, if a user repeatedly mentions their fear of relapse, a personalized statement like, "I completely understand your concerns; relapse is indeed a major concern for many patients in the recovery period. But please know that regular checkups and current rehabilitation monitoring methods are very comprehensive, and we are gradually reducing the risks," can be generated. This enhances the targeted nature of emotional support and makes responses more aligned with the user's psychological state.

[0146] S53. Based on the medical terminology density index in the knowledge demand vector, the emotionally enhanced response text is processed by replacing professional terms through a terminology conversion rule base to obtain the final response text with appropriate terms.

[0147] Specifically, a terminology conversion rule base for breast cancer rehabilitation was constructed through collaboration between medical experts and linguists. This base includes detailed professional terms and their simplified explanations and expressions, such as converting "adjuvant endocrine therapy" to "adjuvant endocrine therapy, which uses medication to regulate hormone levels in the body to prevent cancer recurrence." Based on the medical terminology density index in the knowledge demand vector, a terminology conversion strategy is intelligently selected. If the index indicates a user preference for simplified expressions, the text generator replaces professional terms in batches in the emotionally enhanced response text according to the rule base; conversely, it appropriately retains professional terms, only providing simplified explanations for particularly complex terms, ensuring a dynamic balance between professional knowledge and comprehensibility in the response text, allowing users with different medical knowledge backgrounds to accurately understand rehabilitation advice.

[0148] To ensure the quality of terminology conversion, an evaluation process can be implemented. This involves collecting user feedback on the difficulty of understanding the terms, and using professional medical text readability analysis tools to assess the medical accuracy and accessibility of the converted text. Based on the evaluation results, the terminology conversion rule base is continuously optimized, terminology explanations are updated, and conversion matching rules are refined to continuously improve terminology adaptation, ensuring that the response text accurately conveys rehabilitation knowledge to different user groups.

[0149] The aforementioned deep learning-based question-answering method for breast cancer rehabilitation accurately captures patient needs through multimodal feature parsing to generate a knowledge demand vector. Based on this vector, it retrieves medical knowledge subgraphs and general knowledge fragments in parallel to form a knowledge foundation. It then constructs collaborative knowledge units by dynamically weighting and fusing the two types of knowledge based on the demand vector parameters. Furthermore, it uses a graph attention neural network to allocate the importance of medical knowledge to generate an attention distribution. Finally, it integrates the attention distribution with general knowledge semantically to form a fusion representation, and uses a rule engine to arbitrate conflicts to ensure safety and reliability. Ultimately, it combines a conflict-free reasoning path with the patient demand vector to generate personalized rehabilitation guidance that is both professional and accurate, yet also practical and relevant to daily life, achieving a seamless synergy between medical rigor and everyday practicality.

[0150] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0151] Based on the same inventive concept, this application also provides an apparatus for implementing the deep learning-based breast cancer rehabilitation question-answering method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations of one or more deep learning-based breast cancer rehabilitation question-answering apparatus embodiments provided below can be found in the limitations of the deep learning-based breast cancer rehabilitation question-answering method described above, and will not be repeated here.

[0152] In one exemplary embodiment, such as Figure 3 As shown, a deep learning-based breast cancer rehabilitation question-answering device 30 is provided to implement the methods in the above-described method embodiments. The device includes:

[0153] The multimodal feature parsing module 31 is used to perform multimodal feature parsing on the user-input text data, voice data, and lesion image data to obtain the knowledge requirement vector.

[0154] The knowledge retrieval module 32 is used to perform subgraph retrieval on the medical knowledge base according to the knowledge demand vector to obtain a medical knowledge subgraph; and to perform fragment retrieval on the general knowledge base according to the knowledge demand vector to obtain a general knowledge fragment.

[0155] The weighted fusion processing module 33 is used to perform weighted fusion of the medical knowledge subgraph and the general knowledge fragment based on the parameter components in the knowledge demand vector to obtain a weighted fusion knowledge unit.

[0156] The graph attention allocation module 34 is used to allocate node importance weights to the medical knowledge subgraph based on the weighted fusion knowledge unit through a graph attention neural network, thereby generating a node attention distribution.

[0157] The semantic fusion and integration module 35 is used to integrate knowledge through semantic fusion based on the node attention distribution and the general knowledge fragments to generate a fused knowledge representation.

[0158] The conflict detection and arbitration module 36 is used to perform conflict detection and arbitration based on the fused knowledge representation through a pre-set rule engine to obtain a conflict-free reasoning path.

[0159] The text generation module 37 is used to generate response text through a text generator based on the conflict-free reasoning path and the knowledge requirement vector.

[0160] Embodiments of this application also provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the aforementioned method embodiments.

[0161] Embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method embodiments.

[0162] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0163] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.

Claims

1. A deep learning-based question-and-answer method for breast cancer rehabilitation, characterized in that, The method includes: S1. Perform multimodal feature analysis on the user-input text data, voice data, and lesion image data to obtain the knowledge demand vector; S2. Based on the knowledge demand vector, perform subgraph retrieval on the medical knowledge base to obtain a medical knowledge subgraph; based on the knowledge demand vector, perform fragment retrieval on the general knowledge base to obtain a general knowledge fragment. S3. Based on the parameter components in the knowledge demand vector, the medical knowledge subgraph and the general knowledge fragment are weighted and fused to obtain a weighted fused knowledge unit. S4. Based on the weighted fusion knowledge unit, the medical knowledge subgraph is weighted by a graph attention neural network to generate a node attention distribution; S5. Based on the node attention distribution and the general knowledge fragments, knowledge is integrated through semantic fusion to generate a fused knowledge representation; S6. Based on the fused knowledge representation, conflict detection and arbitration are performed through a pre-set rule engine to obtain a conflict-free reasoning path; S7. Based on the conflict-free reasoning path and the knowledge requirement vector, generate the response text using a text generator.

2. The method according to claim 1, characterized in that, The process of performing multimodal feature analysis on user-input text data, voice data, and lesion image data to obtain a knowledge requirement vector includes: S11. Using a language model trained on breast cancer rehabilitation corpus, medical terminology is identified in the text data to obtain medical terminology identification results; a medical terminology density index is calculated based on the medical terminology identification results. S12. Extract the acoustic features of the speech data, input the acoustic features into a classifier trained on a clinical anxiety speech dataset, and generate an anxiety index that represents the user's anxiety level. S13. The lesion image data is used to identify the rehabilitation stage through a medical image recognition model to obtain the rehabilitation stage code. S14. Construct the knowledge demand vector based on the medical terminology density index, the anxiety index, and the rehabilitation stage coding.

3. The method according to claim 2, characterized in that, The step involves weighted fusion of the medical knowledge subgraph and the general knowledge fragment based on the parameter components in the knowledge demand vector to obtain a weighted fused knowledge unit, including: S21. Linearly combine the medical terminology density index and the rehabilitation stage code in the knowledge demand vector to obtain a linear combination representation; input the linear combination representation into the S-type activation function for normalization to obtain the medical knowledge weight coefficient with a value between 0 and 1, and subtract the medical knowledge weight coefficient from 1 to obtain the general knowledge weight coefficient. S22. Combining the medical knowledge weight coefficients, perform a linear transformation on the medical knowledge subgraph to obtain a weighted medical knowledge representation; S23. Combine the general knowledge weight coefficients to perform embedded spatial projection on the general knowledge fragments to obtain a weighted general knowledge representation; S24. Based on the weighted medical knowledge representation and the weighted general knowledge representation, the weighted fused knowledge unit is obtained by fusion processing through vector addition.

4. The method according to claim 3, characterized in that, The step of assigning node importance weights to the medical knowledge subgraph based on the weighted fusion knowledge unit and generating a node attention distribution by using a graph attention neural network includes: S31. Based on the weighted medical knowledge representation in the weighted fusion knowledge unit, the node feature vectors of the medical knowledge subgraph are extracted through the graph embedding layer to obtain the initial feature matrix of the nodes; S32. Based on the initial feature matrix of the nodes, calculate the correlation between nodes through a multi-head attention mechanism, and obtain the attention coefficient matrix according to the correlation between nodes; S33. Based on the attention coefficient matrix, normalized attention weights are obtained by row-wise softmax normalization. S34. Based on the normalized attention weights and the initial feature matrix of the nodes, feature fusion processing is performed through weighted aggregation operation to obtain the node attention distribution.

5. The method according to claim 2, characterized in that, The process of obtaining a conflict-free reasoning path based on the fused knowledge representation and through a pre-built rule engine for conflict detection and arbitration includes: S41. Based on the medical entity nodes and life concept nodes in the fused knowledge representation, semantic conflict detection is performed by cosine similarity calculation to obtain a set of conflict node pairs; S42. Based on the set of conflicting nodes and the rehabilitation stage encoding, a rule retrieval is performed using a conditional expression matcher to obtain a matching taboo rule identifier. S43. Based on the taboo rule identifier, arbitration weight allocation is performed through the rule priority decision-maker to obtain medical constraint weight and lifestyle advice weight; S44. Use the attention weight in the node attention distribution as the weight value of the corresponding node in the medical knowledge subgraph, delete the nodes in the medical knowledge subgraph whose weight value is lower than the medical constraint weight and their corresponding associated edges, and obtain the filtered knowledge subgraph. S45. Based on the node attention distribution, assign weights to the knowledge points associated with each node in the medical knowledge subgraph in the general knowledge fragment; delete the knowledge points in the general knowledge fragment whose weights are greater than the weights of the life advice and which are associated with the node pairs in the conflict node pair set, to obtain the filtered general knowledge fragment. S46. Based on the filtered knowledge subgraph and the filtered general knowledge fragment, generate a conflict-free reasoning path.

6. The method according to any one of claims 2 to 5, characterized in that, The step of generating response text using a text generator based on the conflict-free reasoning path and the knowledge requirement vector includes: S51. Based on the conflict-free reasoning path, the text skeleton is constructed using a sequence-to-sequence generation model to obtain an initial response draft. S52. Based on the anxiety index in the knowledge demand vector, a reassuring statement is inserted into the initial response draft using an emotion template matcher to obtain an emotion-enhanced response text. S53. Based on the medical terminology density index in the knowledge demand vector, the emotionally enhanced response text is processed by replacing professional terms through a terminology conversion rule base to obtain the final response text with appropriate terms.

7. A deep learning-based breast cancer rehabilitation question-and-answer device, used to implement the method of any one of claims 1 to 6, characterized in that, The device includes: The multimodal feature parsing module is used to perform multimodal feature parsing on user-input text data, voice data, and lesion image data to obtain a knowledge requirement vector; The knowledge retrieval module is used to perform subgraph retrieval on the medical knowledge base according to the knowledge demand vector to obtain medical knowledge subgraphs; and to perform fragment retrieval on the general knowledge base according to the knowledge demand vector to obtain general knowledge fragments. The weighted fusion processing module is used to perform weighted fusion of the medical knowledge subgraph and the general knowledge fragment based on the parameter components in the knowledge demand vector to obtain a weighted fusion knowledge unit; The graph attention allocation module is used to allocate node importance weights to the medical knowledge subgraph based on the weighted fusion knowledge unit through a graph attention neural network, thereby generating a node attention distribution. The semantic fusion and integration module is used to integrate knowledge through semantic fusion based on the node attention distribution and the general knowledge fragments, and generate a fused knowledge representation. The conflict detection and arbitration module is used to perform conflict detection and arbitration based on the fused knowledge representation and through a pre-built rule engine to obtain a conflict-free reasoning path. The text generation module is used to generate response text through a text generator based on the conflict-free reasoning path and the knowledge requirement vector.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.