Intelligent modeling system for journey of breast cancer patient based on multi-modal data fusion
By using multimodal data fusion and temporal attention neural networks, a full-cycle health status model for breast cancer patients is established, which solves the problem of lack of temporal modeling and real-time updates in existing technologies. It enables automatic identification and visualization of key nodes in the journey, improving the accuracy and practicality of breast cancer patient management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-31
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies lack the ability to model the entire lifecycle of breast cancer patients, cannot update in real time and provide cross-modal semantic consistency, and lack automatic identification and labeling of key nodes in the journey, peaks of emotional fluctuations and pain points, making it difficult to provide accurate intervention timing.
By integrating multimodal data from structured medical records, semi-structured interview texts, unstructured emotional speech, and physiological signals from wearable devices, and employing cross-modal semantic alignment technology and temporal attention neural networks, a full-cycle health status evolution model for patients is established. This model automatically identifies key nodes in the journey and generates a visual map with real-time update capabilities.
It enables dynamic monitoring and prediction of the health status of breast cancer patients throughout their entire life cycle, provides a panoramic understanding of the patient experience, improves the accuracy and timeliness of clinical decision-making, and reduces the incidence of adverse experiences.
Smart Images

Figure CN121768637A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical data processing and intelligent modeling technology, specifically involving an intelligent modeling system for the journey of breast cancer patients based on multimodal data fusion, which is applicable to the whole-cycle health management, patient experience optimization and precision medical intervention for breast cancer patients. Background Technology
[0002] Patient journey mapping is an important tool in the field of healthcare service design, visually representing a patient's experiences, emotional fluctuations, and key touchpoints throughout their journey from initial diagnosis to recovery. Traditional patient journey mapping primarily relies on subjective qualitative research methods such as questionnaires and focus group interviews, which suffer from technical limitations including limited data collection dimensions, difficulty in real-time updates, and a lack of objective physiological indicators. With the development of artificial intelligence and multimodal data fusion technologies, the automatic construction of dynamically updated patient journey models using heterogeneous data sources such as structured medical records, semi-structured interview texts, unstructured voice and emotional data, and physiological signals from wearable devices has become a research hotspot.
[0003] Existing technologies include a statistical learning-based multimodal analysis and prediction system for patient behavior (see Chinese Patent CN111920420A). This system includes a multimodal acquisition module, a posture-based patient behavior recognition module, a physiological signal-based patient behavior recognition module, an emotional signal-based patient behavior recognition module, a speech signal-based patient behavior recognition module, and a multi-kernel learning-based fusion module. This system trains and combines each kernel function separately using a multi-kernel classifier to achieve the recognition and prediction of heterogeneous patient behavior data. However, this existing technology has the following shortcomings: The system only focuses on the identification and classification of patients' immediate behaviors, lacking the ability to model the entire treatment cycle temporally, and cannot capture the evolution of patients' health status at different stages such as screening and diagnosis, perioperative period, continued treatment, and rehabilitation and return to work; while the multi-core learning fusion method used in this system can combine multimodal features, it does not solve the semantic alignment problem of different modal data, resulting in a lack of cross-modal semantic consistency in the fused feature vectors; the system lacks a visual representation mechanism for the patient journey, lacking automatic identification and annotation functions for task nodes, emotional peaks, and pain points, making it difficult for medical staff to intuitively grasp the panoramic information of the patient experience; the system lacks dynamic update capabilities, unable to adjust the journey model in real time based on newly generated data during the patient's treatment process, and unable to predict future pain points and optimal intervention times. These technical deficiencies limit the system's application value in disease scenarios requiring long-term, precise management, such as breast cancer.
[0004] Currently, there is no systematic technical solution for intelligent modeling of the entire life cycle of breast cancer patients. Therefore, how to integrate multimodal data such as structured medical record data, semi-structured interview text, unstructured emotional speech, and physiological signals from wearable devices, and establish a dynamic evolution model of the patient's health status through temporal neural networks, automatically identify key nodes in the journey, and generate a real-time updated visual journey map, so as to provide medical staff with a panoramic understanding of the patient experience and accurate prediction of intervention timing, has become an urgent technical problem to be solved in this field. Summary of the Invention
[0005] The purpose of this invention is to overcome the above-mentioned shortcomings of the prior art and provide an intelligent modeling system for breast cancer patient journey based on multimodal data fusion. By integrating multi-source heterogeneous data such as structured medical records, semi-structured interview texts, unstructured emotional speech, and physiological signals from wearable devices, the system utilizes cross-modal semantic alignment technology and temporal attention neural networks to establish a full-cycle health status evolution model for patients. It automatically identifies task nodes, emotional fluctuation peaks, and pain point triggering conditions, and generates a real-time updatable visual journey map, solving the technical problems of single data collection dimensions, strong subjectivity, and difficulty in dynamic updating in the traditional patient journey map construction process.
[0006] This invention provides an intelligent modeling system for the journey of breast cancer patients based on multimodal data fusion, comprising a multimodal data acquisition module, a semantic feature extraction module, a multimodal semantic alignment module, a temporal journey modeling module, a journey node identification module, and a visualization output module. The multimodal data acquisition module collects structured medical record data, semi-structured interview text data, unstructured emotional speech data, and wearable device physiological signal data from breast cancer patients during the screening and diagnosis phase, perioperative phase, continuation treatment phase, and rehabilitation and return-to-work phase. The semantic feature extraction module performs natural language processing on the interview text data to obtain text semantic features, performs acoustic feature extraction and emotion recognition on the emotional speech data to obtain speech emotion features, performs medical entity recognition on the medical record data to obtain structured medical features, and performs temporal feature extraction on the physiological signal data to obtain physiological state features. The multimodal semantic alignment module maps different modal features to a unified semantic space based on a cross-modal attention mechanism, generating an aligned multimodal fusion feature vector. The temporal journey modeling module uses a temporal attention neural network to learn temporal dependencies between multimodal fusion feature vectors and timestamp information, identifying patterns in the patient's health status transitions at different treatment stages and generating a temporal state evolution sequence. The journey node identification module identifies key task nodes, peak emotional fluctuation nodes, and pain trigger nodes in the patient's journey. The visualization output module generates a visualized map of the patient's journey, including a timeline, health status curves, and intervention timing annotations, and updates it in real time based on newly acquired data.
[0007] Compared with the prior art, the present invention has the following beneficial effects:
[0008] This invention integrates four types of multimodal data—structured medical records, semi-structured interview texts, unstructured emotional speech, and physiological signals from wearable devices—to construct a comprehensive data acquisition system covering medical records, patient narratives, emotional states, and physiological indicators. Compared to existing technologies that only collect real-time behavioral data such as posture, physiology, emotion, and speech, the data source of this invention is closer to the full-cycle management scenario of breast cancer patients, with richer data dimensions, laying a solid data foundation for subsequent accurate modeling.
[0009] This invention innovatively introduces a multimodal semantic alignment module, which uses a bidirectional cross-modal attention mechanism to map features from different modalities to a unified semantic space. This solves the problem in existing multi-core learning methods that simply combine features at the feature level without semantic alignment, thus enabling the fused feature vectors to have cross-modal semantic consistency and improving the effectiveness and accuracy of multimodal data fusion.
[0010] This invention uses a temporal attention neural network to establish a patient's full-cycle health status evolution model. Compared with the existing static multi-core classifier, this invention can capture the health status transition patterns and temporal dependencies of patients in different stages such as screening and diagnosis, perioperative period, continued treatment and rehabilitation and return to work, realizing a technological leap from static behavior recognition to dynamic journey modeling.
[0011] This invention is the first to achieve automatic identification and classification of key nodes in the patient journey. Through anomaly detection algorithms and threshold judgment mechanisms, it automatically identifies key task nodes, peak emotional fluctuation nodes, and pain trigger nodes. Compared with existing technologies that cannot provide journey node information, this invention provides medical staff with a panoramic visualization of the patient experience, significantly improving the accuracy and timeliness of clinical decision-making.
[0012] The visual journey map constructed by this invention has the ability to update in real time. When the system acquires new patient data, it can automatically trigger the entire process of feature extraction, semantic alignment, temporal modeling and node recognition, and dynamically update the time axis, health status curve and node annotation in the journey map. Compared with the traditional static patient journey map, which requires periodic manual redrawing, this invention realizes the automated continuous evolution of the journey model, which greatly improves the practical value of the system.
[0013] This invention uses an intervention timing prediction module to predict the time interval and probability of future pain points based on the temporal state evolution sequence and historical patient data, providing medical staff with a window of opportunity for early intervention. Compared with existing technologies that can only identify the current state without predictive ability, this invention realizes a paradigm shift from passive response to proactive prevention, effectively reducing the incidence of adverse patient experiences. Attached Figure Description
[0014] Figure 1 This is a schematic diagram of the overall architecture of the intelligent modeling system for breast cancer patient journey based on multimodal data fusion, as proposed in this invention.
[0015] Figure 2 This is a detailed structural diagram of the semantic feature extraction module of the present invention.
[0016] Figure 3 This is a schematic diagram of the workflow of the multimodal semantic alignment module of the present invention.
[0017] Figure 4 This is a schematic diagram of the node recognition algorithm flow of the journey node recognition module of the present invention.
[0018] Figure 5 This is a schematic diagram of the prediction algorithm flow of the intervention timing prediction module of the present invention. Detailed Implementation
[0019] Please refer to the attached document. Figures 1-5 The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.
[0020] Reference Figure 1 The intelligent modeling system for breast cancer patient journey based on multimodal data fusion provided by the present invention includes a multimodal data acquisition module 1, a semantic feature extraction module 2, a multimodal semantic alignment module 3, a temporal journey modeling module 4, a journey node identification module 5, a visualization output module 6, and an intervention timing prediction module 7.
[0021] The multimodal data acquisition module 1 is used to collect structured medical record data, semi-structured interview text data, unstructured emotional voice data, and wearable device physiological signal data of breast cancer patients during the screening and diagnosis stage, perioperative period, continued treatment stage, and rehabilitation and return to work stage.
[0022] In specific implementations, structured medical record data originates from the Hospital Information System (HIS) and Electronic Medical Record System (EMR), including basic patient information (age, gender, medical history, etc.), diagnostic information (ICD-10 code, TNM stage, pathological type, etc.), treatment information (surgical procedure, chemotherapy regimen, radiotherapy dose, etc.), and examination indicators (tumor markers, imaging results, etc.). Semi-structured interview text data is obtained through in-depth interviews with patients, covering aspects such as the patient's feelings about the diagnostic process, physical reactions during treatment, changes in psychological state, family support, and evaluation of medical services. Interview recordings are converted to text format using speech recognition technology and then stored. Unstructured emotional speech data is obtained by synchronously recording the patient's voice signals during the interview, preserving paralinguistic features such as tone, speech rate, pauses, and vibrato. Wearable device physiological signal data is collected through devices such as smart bracelets or smartwatches, including physiological and behavioral indicators such as heart rate, heart rate variability (HRV), skin conductance level (SCL), sleep duration, sleep quality score, and daily activity steps. All data collection complies with medical data privacy protection guidelines, obtains informed consent from patients, and anonymizes sensitive information. The data collection frequency is flexibly set according to the patient's treatment stage. Data collection is more frequent during the screening and diagnosis stage and the perioperative stage, usually daily or weekly, while the data collection frequency is moderately reduced during the continuing treatment stage and the rehabilitation and return-to-work stage, usually weekly or monthly.
[0023] The semantic feature extraction module 2 is connected to the multimodal data acquisition module 1. It is used to perform natural language processing on interview text data to obtain text semantic features, perform acoustic feature extraction and emotion recognition on emotional speech data to obtain speech emotion features, perform medical entity recognition on medical record data to obtain structured medical features, and perform temporal feature extraction on physiological signal data to obtain physiological state features.
[0024] Reference Figure 2 The semantic feature extraction module 2 includes a text processing unit, an emotion recognition unit, a medical record parsing unit, and a physiological signal processing unit.
[0025] The text processing unit performs natural language processing on the interview text data. First, it uses the jieba word segmentation tool for Chinese word segmentation. Then, it identifies parts of speech (nouns, verbs, adjectives, etc.) through part-of-speech tagging. Next, it uses a BERT-based named entity recognition model to extract symptom descriptions (e.g., "nausea," "fatigue," "pain"), treatment experiences (e.g., "anxiety," "fear," "hope"), and psychological state keywords (e.g., "worried about relapse," "affecting work") from the patient's narrative. The text semantic features are encoded using a pre-trained Chinese-BERT-wwm model, mapping the text sequence to a 768-dimensional semantic vector. Preferably, using a pre-trained MC-BERT model in the medical field can further improve the accuracy of medical semantic understanding.
[0026] The emotion recognition unit performs acoustic feature extraction and emotion recognition on emotional speech data. Acoustic feature extraction includes Mel-frequency cepstral coefficient (MFCC) extraction, pitch feature (fundamental frequency F0) extraction, and speech rate feature (syllable rate) extraction. MFCC features are obtained using a standard speech signal processing procedure: first, the speech signal is pre-emphasized, then windowed in frames, followed by Fast Fourier Transform (FFT), filtering with a Mel filter bank, taking the logarithm, and then performing Discrete Cosine Transform (DCT) to finally obtain 13-dimensional MFCC coefficients. Pitch features are extracted using autocorrelation or cepstral methods to extract the fundamental frequency F0, reflecting the pitch changes during the patient's speech. Speech rate features are obtained by counting the number of syllables per unit time. Emotion recognition uses a deep convolutional neural network (CNN) model, with a network structure consisting of 3 convolutional layers, 2 pooling layers, and 2 fully connected layers. The convolutional kernel size is 3×3, the pooling layers use max pooling, and the fully connected layers have 256 and 128 neurons, respectively. The network output layer uses the softmax activation function to classify speech signals into six emotion types: joy, calmness, sadness, anger, fear, and anxiety, and outputs the emotion intensity value (range 0 to 1).
[0027] The medical record parsing unit standardizes medical terminology and identifies entities in structured medical record data. For diagnostic codes, they are uniformly converted to ICD-10 encoding format; for treatment plans, key information such as surgery name, chemotherapy drug name, and radiotherapy dose is extracted; for drug dosages, they are standardized to standard units (e.g., mg, mL); and for examination indicator values, the values and reference ranges of tumor markers (e.g., CA15-3, CEA) are extracted. The medical record parsing unit uses a combination of rule-based and machine learning methods. First, it extracts structured fields through regular expression matching, and then uses a CRF model to identify medical entities in unstructured medical record text. Finally, the structured medical features are represented as a composite feature vector containing diagnostic codes, treatment stage identifiers, drug type vectors, and examination indicator vectors, with a dimension of 512.
[0028] The physiological signal processing unit preprocesses and extracts features from the physiological signal data collected by the wearable device. For heart rate signals, noise reduction filtering is first performed using a bandpass filter ranging from 0.5Hz to 40Hz to remove baseline drift and high-frequency noise. Then, R-wave detection is performed to calculate instantaneous heart rate and heart rate variability indices, including time-domain indices (SDNN, RMSSD) and frequency-domain indices (LF, HF, LF / HF ratio). For skin conductance signals, after baseline correction, skin conductance level (SCL) and skin conductance response (SCR) features are extracted. For sleep data, parameters such as total sleep duration, deep sleep duration, light sleep duration, and sleep efficiency are extracted. The physiological state features are ultimately represented as a 128-dimensional feature vector containing heart rate statistics, HRV indices, SCL values, sleep quality parameters, and activity levels.
[0029] The multimodal semantic alignment module 3 is connected to the semantic feature extraction module 2. It is used to map text semantic features, speech emotion features, structured medical features and physiological state features to a unified semantic space based on the cross-modal attention mechanism, and generate an aligned multimodal fusion feature vector.
[0030] Reference Figure 3 The multimodal semantic alignment module 3 employs a bidirectional cross-modal attention mechanism for semantic alignment. Specifically, text semantic features, speech emotion features, structured medical features, and physiological state features are denoted as follows: , , and The dimensions are respectively , , and First, a linear transformation is used to project each modal feature onto the same dimension. ,get , , and ,in , , and This is a learnable weight matrix.
[0031] The computational process of the cross-modal attention mechanism is as follows: for any two modalities... and modality Features As a query vector, modality Features As a vector of keys and values, calculate the attention weights:
[0032] ,
[0033] in, For modality For modes Attention weight matrix, for transpose, For the hidden layer dimension, This is a normalized exponential function. This formula calculates the modal... Each feature and mode The similarity of all features is considered, and features with higher similarity have greater weight in the combination.
[0034] Based on attention weights for modality The features are weighted and fused to obtain the modality. Cross-modal features from different perspectives:
[0035] ,
[0036] in, For modality From modality Semantic features extracted from [the data].
[0037] Bidirectional cross-modal attention was calculated for all modal pairs, resulting in 12 sets of cross-modal features. , , , , , , , , , , , ).
[0038] The original features of each modality are concatenated with the cross-modal features extracted from other modalities to obtain the enhanced features:
[0039] ,
[0040] in, This represents a vector concatenation operation. , , To remove The other three modes besides.
[0041] Finally, the enhanced features of the four modalities are concatenated and fused through a two-layer fully connected network to obtain the aligned multimodal fused feature vector. :
[0042] ,
[0043] in, and The weight matrix is a learnable matrix. and For bias vectors, To correct the linear unit activation function. Multimodal fusion feature vector. The dimension is 256, containing semantic information from four modalities and cross-modal semantic association information. In a preferred embodiment, The value is 512, and the hidden layer dimension of the fully connected layer is 1024, which can control the model parameter scale while maintaining the feature representation capability.
[0044] The temporal journey modeling module 4 is connected to the multimodal semantic alignment module 3 and the multimodal data acquisition module 1. It is used to learn the temporal dependency relationship of multimodal fusion feature vector and corresponding timestamp information based on the temporal attention neural network, identify the health status transformation pattern of patients in different treatment stages, and generate a temporal state evolution sequence.
[0045] The temporal journey modeling module 4 employs a bidirectional long short-term memory network (Bi-LSTM) combined with a self-attention mechanism for temporal modeling. Specifically, it arranges all multimodal fusion feature vectors from the patient's initial diagnosis to the current moment in chronological order to form a temporal input sequence. ,in This represents the number of time steps, corresponding to the number of times patient data was collected. Each time step... It also includes corresponding timestamp information. This includes absolute time (date) and relative time (number of days since the initial consultation, number of days since the surgery, etc.).
[0046] Bidirectional Long Short-Term Memory (LSTM) networks consist of two parts: a forward LSTM and a backward LSTM. The forward LSTM operates from time step 1 to time step 2. The input sequence is processed sequentially to capture the forward dependencies of the patient's health status; the inverse LSTM starts from the time step. At time step 1, the input sequence is processed in reverse to capture the backward dependencies of the patient's health status. The core computational units of LSTM include forget gates, input gates, output gates, and cell states, selectively retaining or forgetting historical information through gating mechanisms. For time steps... The calculation process of LSTM is as follows:
[0047] Forgotten Gate: ;
[0048] Input Gate: ;
[0049] Candidate cell status: ;
[0050] Cell status update: ;
[0051] Output gate: ;
[0052] Hidden output: ;
[0053] in, It is the sigmoid activation function. The hyperbolic tangent activation function is used. For element-wise multiplication, , , , This is the weight matrix. , , , For bias vectors, This is the hidden state from the previous time step. This represents the cell state at the previous time step.
[0054] Forward LSTM outputs the forward hidden state sequence The reverse LSTM outputs the backward hidden state sequence. Concatenating the two yields a bidirectional hidden state sequence. .
[0055] To further enhance the historical moment information that is highly correlated with the current state, a self-attention mechanism is introduced to assign weights to the bidirectional hidden state sequence. The calculation process of self-attention is as follows:
[0056] Query vector: ;
[0057] Key vector: ;
[0058] Value vector: ;
[0059] in, For a matrix that stacks the hidden states of all time steps row by row, , , This is a learnable weight matrix.
[0060] Attention weight calculation:
[0061] ,
[0062] in, Let be the dimension of the key vector. Normalize each row so that the sum of the attention weights of each time step to all historical moments is 1.
[0063] The weighted hidden state sequence output by the self-attention mechanism is denoted as follows: That is, the temporal state evolution sequence, each Includes patient time steps The comprehensive health status information. In a preferred embodiment, the hidden layer dimension of the Bi-LSTM is set to 256, and the key vector dimension of the self-attention mechanism is... Setting it to 128 allows for the capture of long-range temporal dependencies while maintaining computational efficiency.
[0064] Timestamp information Position encoding is incorporated into the time-series modeling process. Specifically, sine and cosine functions are used to encode relative time, resulting in a time position vector. Then With multimodal fusion feature vector The summation is then fed into the LSTM network, enabling the model to perceive the absolute position and relative distance at different points in time.
[0065] The journey node identification module 5 is connected to the temporal journey modeling module 4. It is used to identify key task nodes, peak emotional fluctuation nodes, and pain trigger nodes in the patient's journey based on the temporal state evolution sequence, and to determine the time position, duration, and impact intensity of each node.
[0066] Reference Figure 4 The journey node identification module 5 identifies three types of key nodes through anomaly detection algorithms and threshold judgment.
[0067] Key task node identification: analysis of time-series state evolution sequences Calculate the rate of change of state within the sliding window. Define the window size as... (Preferred) ), for time step Calculation window Rate of change of state within:
[0068] ,
[0069] in, It is the L2 norm. It reflects the magnitude of change in the patient's health status within a unit of time step. When Exceeding the preset first threshold Mark time step This is a critical node in the task. First threshold. The value ranges from 0.3 to 0.5, with a preferred value of 0.4. Key nodes in the task typically correspond to important events in the patient's treatment process, such as diagnosis, surgery, the start of a chemotherapy cycle, and the end of radiotherapy.
[0070] Emotional fluctuation peak node identification: Extracting emotional intensity value sequences from speech emotional features Peak detection is performed on the sequence. The peak detection employs the local maximum method, where the peak value is determined at time step [missing information]. Emotional intensity value satisfy and ,Right now It is a local maximum, and at the same time Exceeding the preset second threshold Mark time step This represents the peak point of emotional fluctuations. The second threshold. The value ranges from 0.6 to 0.8, with a preferred value of 0.7. The peak value of emotional fluctuation reflects the moments of intense fluctuation in the patient's emotional state, such as upon first receiving the diagnosis, the peak of chemotherapy side effects, and when the condition improves.
[0071] Pain trigger node identification: Extract key physiological indicators (such as heart rate variability, skin conductance level, and sleep quality score) from physiological state characteristics, and calculate the statistical distribution characteristics (mean) of these indicators. and standard deviation When continuous Preferred time points or The physiological signal data deviated from the normal range by more than the preset third threshold. At that time, the starting point of the time interval is marked as the pain point trigger node. The normal range is defined as follows: The third threshold The value ranges from 1.5 to 2.0, with a preferred value of 1.8. Pain trigger points correspond to periods when the patient's physical or psychological discomfort significantly increases, such as severe insomnia, persistent anxiety, or intense pain.
[0072] For each identified node, the system records its time and location. Duration cycle (Calculated by the time point when the detected state returns to normal) and the intensity of the impact. (Quantified by the magnitude of state changes or the degree of deviation of indicators). Node information is stored in a node list. Among them Indicate the node type (critical task node, peak emotional fluctuation node, or pain point trigger node). This represents the total number of nodes identified.
[0073] The visualization output module 6 is connected to the journey node recognition module 5. It is used to generate a patient journey visualization map with time axis, health status curve and intervention timing annotation based on key task nodes, peak emotional fluctuation nodes and pain trigger nodes, and update the patient journey visualization map in real time based on newly collected data.
[0074] The patient journey visualization map generated by Visualization Output Module 6 includes the following components:
[0075] Horizontal Timeline: The horizontal axis represents time, marking the start and end times of the screening and diagnosis phase, perioperative phase, continuation treatment phase, and rehabilitation and return-to-work phase. The timeline uses dates or relative times (such as "30 days after diagnosis") to facilitate medical staff in quickly locating the patient's treatment stage.
[0076] Health status curve: The vertical axis represents the patient's overall health score, and the horizontal axis represents time. The overall health score is calculated by analyzing the evolution sequence of the patient's health status over time. A linear mapping is performed, resulting in a score range of 0 to 100, with higher scores indicating better patient health. The health status curve visually reflects the trend of changes in the patient's health status throughout the treatment cycle, with peaks and troughs corresponding to periods of improvement and deterioration.
[0077] Node Label Layer: Marks key task nodes, peak emotional fluctuation nodes, and pain point trigger nodes on the timeline. These three types of nodes are distinguished by different colors and icons: key task nodes are marked with blue triangles, peak emotional fluctuation nodes with yellow circles, and pain point trigger nodes with red stars. Hovering the mouse over a node icon displays detailed information about the node, including its time location, duration, impact intensity, and a description of the triggering reason.
[0078] Intervention Recommendation Area: Near the identified pain point triggers, the system recommends corresponding medical interventions and psychological support programs based on its knowledge base and historical cases. For example, for pain points caused by chemotherapy side effects, medical interventions such as adjusting chemotherapy dosage, using antiemetics, and strengthening nutritional support are recommended; for pain points caused by anxiety, psychological support programs such as psychological counseling, relaxation training, and family companionship are recommended. Intervention recommendations are displayed on the map in the form of text labels or pop-ups.
[0079] The visualized map features an interactive design, supporting zooming, panning, and filtering operations. Users can filter and display travel information for specific time periods, choose to hide or show certain types of nodes, and click on nodes to view associated raw data (such as interview recordings, medical records, and physiological signal waveforms).
[0080] The visualization output module 6 updates the patient journey visualization map in real time as follows: When the multimodal data acquisition module 1 acquires new patient data, it triggers the semantic feature extraction module 2, the multimodal semantic alignment module 3, and the temporal journey modeling module 4 to process the new patient data sequentially, generating new multimodal fusion feature vectors and temporal states. The processed new state information is then appended to the temporal state evolution sequence. At the end, the sequence length starts from Growth to Re-execute the node identification process of journey node identification module 5 and update the node list. The system dynamically updates the timeline (extending to new time points), health status curves (plotting new data points and smoothly connecting them), and node annotation layers (adding newly identified node annotations or updating the duration of existing nodes) on the patient journey visualization map. The update process uses incremental calculation, recalculating only new data and affected historical data to avoid the computational overhead of a full recalculation. Preferably, the system is configured with a timed update mechanism, automatically triggering data collection and map update processes every 24 hours or after each patient visit to ensure the timeliness of map information.
[0081] The intervention timing prediction module 7 is connected to the journey node identification module 5 and the temporal journey modeling module 4. It is used to predict the time interval and probability of future pain point triggering nodes based on the temporal state evolution sequence and historical patient data, providing medical staff with a time window for early intervention.
[0082] Reference Figure 5 The intervention timing prediction module 7 uses a sequence-to-sequence (Seq2Seq) model for prediction. The Seq2Seq model consists of an encoder and a decoder. The encoder converts the historical portion of the temporal state evolution sequence... As input, the LSTM network learns the evolutionary pattern of the patient's health status, and outputs an encoding vector. Encoding vector This is the hidden state of the encoder at the last time step, containing all the information of the historical sequence. The decoder uses the encoded vector... Using the initial state, a state prediction sequence for future time periods is generated through an LSTM network. ,in For the number of time steps to be predicted, the preferred method is... This corresponds to predicting the status over the next 7 data collection periods (such as the next 7 days or the next 7 weeks).
[0083] Each time step of the decoder The calculation process is as follows:
[0084] ,
[0085] in, For the decoder LSTM network, This represents the hidden state of the decoder at the previous time step.
[0086] State prediction sequence The pain point trigger node identification algorithm of the journey node identification module 5 is applied to mark the time steps in the predicted sequence where pain points may occur. Specifically, key physiological indicators are extracted from the predicted physiological state features, and their deviation from the normal range is calculated. When the indicators of consecutive time steps deviate beyond a third threshold... When the time comes, mark it as the predicted pain point trigger node.
[0087] For each predicted pain point trigger node, a confidence probability is calculated. The confidence probability is quantified by the model's prediction uncertainty using the Monte Carlo Dropout method. A Dropout layer is introduced into the decoder, performing multiple forward propagations (ideally 20 times) to obtain multiple prediction results. The variance of these prediction results is calculated as a measure of uncertainty; a smaller variance indicates a more reliable prediction. The confidence probability is defined as:
[0088] ,
[0089] in, The standard deviation of the prediction results The maximum possible value of the standard deviation. The range is from 0 to 1, with the closer to 1 indicating a higher confidence level.
[0090] Predictive nodes with a confidence level higher than 0.7 are selected as valid early warning signals. The system pushes early warning notifications to medical staff, including the predicted trigger time of the pain point (e.g., "The patient is expected to experience severe insomnia between day 3 and day 5 in the future"), the predicted trigger cause (based on correlation pattern analysis in historical data), and suggested early intervention measures (e.g., "Adjust the medication regimen in advance and increase sleep support therapy"). Early warning notifications are sent to responsible medical staff via the hospital information system's push notification function or SMS.
[0091] The Seq2Seq model is trained using historical patient data. The system collects a large amount of complete journey data from breast cancer patients, dividing each patient's temporal state evolution sequence into historical and future parts. The historical part is used as the encoder input, and the future part as the target output of the decoder. The model parameters are trained using supervised learning. The loss function is the mean squared error (MSE), which minimizes the difference between the predicted and actual states. The model is trained using the Adam optimizer with a learning rate of 0.001, a batch size of 32, and 100 training epochs.
[0092] In a preferred embodiment, the system integrates knowledge distillation technology. First, a complex, large-scale Transformer model is trained as the teacher model, and then the knowledge is transferred to a Seq2Seq student model with fewer parameters. This significantly reduces inference time while maintaining prediction accuracy, meeting the needs of real-time clinical prediction.
[0093] In a preferred embodiment, the system also includes a feedback learning mechanism. After medical staff intervene in advance based on the warning signal, the system records the type of intervention, the intervention time, and the intervention effect (assessing whether the patient's condition has improved through subsequently collected data). This feedback data is used to continuously optimize the prediction model of the intervention timing prediction module 7 and the threshold parameters of the journey node identification module 5, forming a closed-loop learning process of "prediction → intervention → feedback → optimization," so that the system's predictive ability continuously improves with the increase of usage time.
[0094] In a preferred embodiment, the system employs federated learning technology to train the model using data from multiple hospitals while protecting patient privacy. Each hospital trains the model locally, uploading only the model parameter gradients to the central server for aggregation. This avoids the cross-institutional transmission of raw patient data, complying with regulations protecting medical data privacy. Federated learning significantly expands the scale of training data, improving the model's generalization ability and prediction accuracy.
[0095] In a preferred embodiment, the system provides a personalized customization function, allowing medical staff to adjust the threshold parameters for node recognition based on the characteristics of different patients. For example, for emotionally sensitive patients, the second threshold for peak emotion fluctuation nodes can be lowered. This allows the system to capture more subtle emotional changes; for patients with poor physical tolerance, it can lower the third threshold of pain trigger points. This allows for the early identification of potential discomfort signals. Personalized parameter settings are stored in the patient's file and automatically loaded during subsequent data processing.
[0096] In a preferred embodiment, the system integrates natural language generation technology to convert a patient journey visualization map into a journey report in natural language format. The report includes an overview of the patient's treatment stages, a timeline of key events, analysis of emotional change trends, a summary of pain points, and a summary of intervention recommendations. The natural language report facilitates quick browsing and understanding of the patient journey by healthcare professionals and helps patients and their families understand treatment progress and key management priorities. Report generation employs a template-based approach combined with the BERT language model to ensure the professionalism and readability of the report's language.
[0097] In a preferred embodiment, the system supports multi-patient comparative analysis. Healthcare professionals can select multiple patients' journey maps for parallel display, comparing the health status and node distribution of different patients at the same treatment stage to identify common problems and individual differences. Multi-patient comparative analysis helps medical teams summarize treatment experience, optimize treatment plans, and improve the overall quality of medical services.
[0098] This invention integrates multimodal data such as structured medical records, semi-structured interview texts, unstructured emotional speech, and physiological signals from wearable devices. It utilizes cross-modal semantic alignment technology and temporal attention neural networks to establish a full-cycle health status evolution model for patients. This model automatically identifies task nodes, peak emotional fluctuations, and pain trigger conditions, generates a real-time updatable visual journey map, and predicts future intervention opportunities. It provides breast cancer patients with panoramic, precise, and dynamic intelligent journey modeling support, significantly improving the personalization of medical services and the quality of patient experience.
[0099] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A breast cancer patient journey intelligent modeling system based on multimodal data fusion, characterized in that, include: The multimodal data acquisition module is used to collect structured medical record data, semi-structured interview text data, unstructured emotional voice data, and wearable device physiological signal data of breast cancer patients during the screening and diagnosis stage, perioperative period, continued treatment stage, and rehabilitation and return to work stage. The semantic feature extraction module, connected to the multimodal data acquisition module, is used to perform natural language processing on the interview text data to obtain text semantic features, perform acoustic feature extraction and emotion recognition on the emotional speech data to obtain speech emotion features, perform medical entity recognition on the medical record data to obtain structured medical features, and perform temporal feature extraction on the physiological signal data to obtain physiological state features. The multimodal semantic alignment module, connected to the semantic feature extraction module, is used to map the text semantic features, the speech emotion features, the structured medical features and the physiological state features to a unified semantic space based on a cross-modal attention mechanism, and generate an aligned multimodal fusion feature vector. The temporal journey modeling module is connected to the multimodal semantic alignment module and the multimodal data acquisition module. It is used to learn the temporal dependency relationship of the multimodal fusion feature vector and the corresponding timestamp information based on the temporal attention neural network, identify the health status transformation pattern of the patient in different treatment stages, and generate a temporal state evolution sequence. The journey node identification module, connected to the temporal journey modeling module, is used to identify key task nodes, peak emotional fluctuation nodes, and pain triggering nodes in the patient's journey based on the temporal state evolution sequence, and to determine the time location, duration, and intensity of influence of each node. The visualization output module, connected to the journey node identification module, is used to generate a patient journey visualization map containing a timeline, health status curve, and intervention timing annotations based on the key task nodes, the peak emotional fluctuation nodes, and the pain point trigger nodes, and to update the patient journey visualization map in real time based on newly collected data.
2. The intelligent modeling system for breast cancer patient journey based on multimodal data fusion according to claim 1, characterized in that, The semantic feature extraction module includes: The text processing unit is used to perform word segmentation, part-of-speech tagging, and named entity recognition on the interview text data, and to extract keywords related to symptom descriptions, treatment experiences, and psychological states from the patient's narrative. The emotion recognition unit is used to extract Mel frequency cepstral coefficients, pitch features, and speech rate features from the emotional speech data, and to identify the patient's emotion type and intensity through a deep convolutional neural network. The medical record parsing unit is used to standardize medical terminology in the structured medical record data and extract diagnostic codes, treatment plans, drug dosages, and examination index values. The physiological signal processing unit is used to perform noise reduction filtering, baseline drift correction, and feature waveform extraction on the physiological signal data to obtain parameters such as heart rate variability, skin conductance level, and sleep quality.
3. The intelligent modeling system for breast cancer patient journey based on multimodal data fusion according to claim 1, characterized in that, The multimodal semantic alignment module employs a bidirectional cross-modal attention mechanism for semantic alignment, specifically including: The text semantic features, the voice emotion features, the structured medical features, and the physiological state features are respectively used as the query vector and the key-value vector; The cross-modal attention mechanism is used to calculate the similarity weights between features of different modalities and to identify the feature combinations with the strongest semantic relevance. Based on the similarity weights, the features of each modality are weighted and fused to generate an aligned feature vector containing multimodal semantic information; The aligned feature vector is mapped to the fixed-dimensional multimodal fusion feature vector through a fully connected layer.
4. The intelligent modeling system for breast cancer patient journey based on multimodal data fusion according to claim 1, characterized in that, The temporal journey modeling module employs a bidirectional long short-term memory network combined with a self-attention mechanism for temporal modeling, specifically including: The multimodal fusion feature vectors are arranged in chronological order to form a temporal input sequence; The bidirectional long short-term memory network is used to encode the temporal input sequence in both forward and reverse directions to capture the sequential dependencies of the patient's health status. The self-attention mechanism is introduced to assign weights to the encoded hidden state vector, thereby strengthening the historical moment information that is highly correlated with the current state. Output the time-series state evolution sequence containing time-series context information.
5. The intelligent modeling system for breast cancer patient journey based on multimodal data fusion according to claim 1, characterized in that, The journey node identification module identifies key nodes through anomaly detection algorithms and threshold judgments, specifically including: Calculate the rate of state change within a sliding window for the time-series state evolution sequence, and mark the task critical node when the rate of state change exceeds a preset first threshold. Peak detection is performed on the emotion intensity value in the voice emotion feature. When the emotion intensity value reaches a local maximum value and exceeds a preset second threshold, it is marked as the emotion fluctuation peak node. Abnormal patterns in the physiological state characteristics are identified, and when the physiological signal data at multiple consecutive time points deviates from the normal range by more than a preset third threshold, it is marked as the pain point trigger node.
6. The intelligent modeling system for breast cancer patient journey based on multimodal data fusion according to claim 5, characterized in that, The first threshold ranges from 0.3 to 0.5, the second threshold ranges from 0.6 to 0.8, and the third threshold ranges from 1.5 to 2.0 times the standard deviation.
7. The intelligent modeling system for breast cancer patient journey based on multimodal data fusion according to claim 1, characterized in that, The patient journey visualization map generated by the visualization output module includes: A horizontal timeline is used to mark the start and end times of the screening and diagnosis phase, the perioperative phase, the continuing treatment phase, and the rehabilitation and return-to-work phase. The health status curve reflects the trend of a patient's comprehensive health score at different time points; A node labeling layer marks the key nodes of the task, the peak nodes of emotional fluctuations, and the pain point trigger nodes on the timeline, and distinguishes them with different colors and icons; The intervention suggestion area displays recommended medical interventions and psychological support plans near the identified pain point trigger nodes.
8. The intelligent modeling system for breast cancer patient journey based on multimodal data fusion according to claim 1, characterized in that, The visualization output module updates the patient journey visualization map in real time in the following way: When the multimodal data acquisition module acquires new patient data, it triggers the semantic feature extraction module, the multimodal semantic alignment module, and the temporal journey modeling module to process the new patient data in sequence. The processed new state information is appended to the end of the temporal state evolution sequence; Re-execute the node recognition process of the journey node recognition module to update the position and number of the key task nodes, the peak emotion fluctuation nodes, and the pain point trigger nodes; The timeline, health status curve, and node label layer are dynamically updated on the patient journey visualization map.
9. The intelligent modeling system for breast cancer patient journey based on multimodal data fusion according to claim 1, characterized in that, The system also includes: The intervention timing prediction module, connected to the journey node identification module and the temporal journey modeling module, is used to predict the time interval and probability of the pain point triggering node that may occur in the future based on the temporal state evolution sequence and historical patient data, so as to provide medical staff with a time window for early intervention.
10. The intelligent modeling system for breast cancer patient journey based on multimodal data fusion according to claim 9, characterized in that, The intervention timing prediction module uses a sequence-to-sequence model for prediction, specifically including: The historical portion of the temporal state evolution sequence is used as the encoder input to learn the evolutionary pattern of the patient's health status; Generate state prediction sequences for future time periods using a decoder; The identification algorithm of the journey node identification module is applied to the state prediction sequence to mark the predicted pain point triggering nodes; Calculate the confidence probability of each prediction node, and select nodes with a confidence level higher than 0.7 as valid early warning signals.
Citation Information
Patent Citations
Patient behavior multimodal analysis and prediction system based on statistical learning
CN111920420A