A method, device, equipment and medium for predicting user symptoms

CN122575767APending Publication Date: 2026-08-14CHINA MOBILE GROUP SHANDONG +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

现有技术多依赖于单一模态数据或有限的多模态融合,未能充分整合病症发展的复杂病理生理机制;同时,现有方案在时序数据处理上往往采用静态或简单分段方法,无法精准对齐多源异构数据在疾病进展过程中的动态变化规律,导致预测结果缺乏时间敏感性和临床相关性

Benefits of technology

[0019]本申请实施例的技术方案,从多个信息系统中获取同一用户的多模态数据;其中,多模态数据包括统计学信息、临床记录、实验室化验结果、医学影像原始数据以及中医四诊数据中的多项数据;将多模态数据进行数据整合和时序对齐,得到西医指标特征向量与中医证素特征向量;根据所述西医指标特征向量与所述中医证素特征向量确定融合向量,并基于所述融合向量对用户病症进行预测。上述方案能够获取多模态数据并进行整合和时序对齐得到西医指标特征向量与中医证素特征向量,实现中西医特征的兼顾,并将西医指标特征向量与中医证素特征向量进行语义层面的深度融合,模拟了中医辨证论治的思维过程,从而更全面地捕捉病症发生发展的多因素影响,提升了预测的全面性和病理机制相关性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575767A_ABST
    Figure CN122575767A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and medium for predicting user symptoms. The method includes: acquiring multimodal data of the same user from multiple information systems; wherein the multimodal data includes multiple data from statistical information, clinical records, laboratory test results, raw medical imaging data, and traditional Chinese medicine (TCM) diagnostic data; integrating and temporally aligning the multimodal data to obtain Western medicine indicator feature vectors and TCM syndrome element feature vectors; determining a fusion vector based on the Western medicine indicator feature vectors and the TCM syndrome element feature vectors; and predicting the user's symptoms based on the fusion vector. This solution deeply integrates multimodal medical data, achieves accurate temporal alignment and dynamic evolution modeling, and possesses interpretable decision support capabilities, thereby comprehensively improving the accuracy, robustness, and clinical applicability of symptom prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart healthcare technology, and in particular to a method, device, equipment, and medium for predicting user symptoms. Background Technology

[0002] Currently, some diseases are widespread, affecting hundreds of millions of people, and the age of onset is trending younger. Faced with this health challenge, treatment after onset is far from sufficient. Therefore, early prediction and risk assessment are crucial. By identifying potential high-risk groups and intervening effectively in advance, the risk of developing these diseases can be significantly reduced, saving countless lives. Currently, various technologies are attempting to improve prediction accuracy from different perspectives, but certain limitations still exist.

[0003] Several key limitations exist in existing disease incidence risk prediction technologies, which restrict the accuracy, interpretability, and clinical applicability of prediction models. Existing technologies largely rely on single-modal data or limited multimodal fusion, failing to fully integrate the complex pathophysiological mechanisms of disease development. Furthermore, existing solutions often employ static or simple segmentation methods in processing time-series data, failing to accurately align the dynamic changes of multi-source heterogeneous data during disease progression, resulting in prediction results lacking time sensitivity and clinical relevance. Summary of the Invention

[0004] This application provides a method, apparatus, device, and medium for predicting user symptoms, which deeply fuses multimodal data to improve the accuracy of symptom prediction.

[0005] According to one aspect of this application, a method for predicting user symptoms is provided, the method comprising:

[0006] Multimodal data of the same user is obtained from multiple information systems; the multimodal data includes statistical information, clinical records, laboratory test results, raw medical imaging data, and multiple data from the four diagnostic methods of traditional Chinese medicine.

[0007] Multimodal data is integrated and time-series aligned to obtain the feature vectors of Western medicine indicators and the feature vectors of traditional Chinese medicine syndrome elements;

[0008] A fusion vector is determined based on the feature vectors of Western medicine indicators and the feature vectors of traditional Chinese medicine syndrome elements, and the user's symptoms are predicted based on the fusion vector.

[0009] According to one aspect of this application, a user symptom prediction device is provided, the device comprising:

[0010] The multimodal data acquisition module is used to acquire multimodal data of the same user from multiple information systems; the multimodal data includes multiple data from statistical information, clinical records, laboratory test results, raw medical imaging data, and traditional Chinese medicine diagnostic data.

[0011] The data processing module is used to integrate and time-series align multimodal data to obtain feature vectors of Western medicine indicators and feature vectors of traditional Chinese medicine syndrome elements.

[0012] The symptom prediction module is used to determine a fusion vector based on the Western medicine indicator feature vector and the traditional Chinese medicine syndrome element feature vector, and to predict the user's symptom based on the fusion vector.

[0013] According to another aspect of this application, an electronic device is provided, the electronic device comprising:

[0014] At least one processor; and,

[0015] A memory that is communicatively connected to at least one processor; wherein,

[0016] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the user symptom prediction method of any embodiment of the present application.

[0017] According to another aspect of this application, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement the user symptom prediction method of any embodiment of this application.

[0018] According to another aspect of this application, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the user symptom prediction method of any embodiment of this application.

[0019] The technical solution of this application embodiment acquires multimodal data of the same user from multiple information systems. This multimodal data includes statistical information, clinical records, laboratory test results, raw medical imaging data, and multiple data points from the four diagnostic methods of Traditional Chinese Medicine (TCM). The multimodal data is integrated and time-aligned to obtain Western medicine indicator feature vectors and TCM syndrome element feature vectors. A fusion vector is determined based on these two vectors, and the user's symptoms are predicted based on the fusion vector. This solution acquires multimodal data, integrates and time-aligns it to obtain Western medicine indicator feature vectors and TCM syndrome element feature vectors, achieving a balance between Western and TCM features. Furthermore, it performs deep semantic fusion of the Western medicine indicator feature vectors and TCM syndrome element feature vectors, simulating the thought process of TCM syndrome differentiation and treatment. This allows for a more comprehensive capture of the multi-factor influences on the occurrence and development of symptoms, improving the comprehensiveness of prediction and its relevance to pathological mechanisms.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A flowchart illustrating a user symptom prediction method provided in this application embodiment;

[0023] Figure 2 A flowchart of a user symptom prediction method provided in another embodiment of this application;

[0024] Figure 3 A flowchart of a user symptom prediction method provided in another embodiment of this application;

[0025] Figure 4 This is a schematic diagram of the structure of a user symptom prediction device provided in an embodiment of this application;

[0026] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0028] It should be noted that the terms "first," "second," "third," "fourth," "actual," "preset," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] The acquisition, storage, use, and processing of data in this application comply with relevant national laws and regulations. The acquired data is obtained with authorization and will not be disclosed without permission, used for illegal purposes, purposes detrimental to the interests of others, or for personalized analysis or product promotion. It should be noted that certain software, components, models, and other existing industry solutions may be mentioned in the embodiments of this application. These should be considered exemplary and intended only to illustrate the feasibility of implementing the technical solution of this application, but do not imply that the applicant has already used or necessarily used the relevant content of such solutions.

[0030] Figure 1 This is a flowchart illustrating a user symptom prediction method provided in an embodiment of this application. This embodiment is applicable to situations where the likelihood of a symptom recurrence needs to be predicted. The method can be executed by a user symptom prediction device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0031] S110. Obtain multimodal data of the same user from multiple information systems; wherein, the multimodal data includes multiple data from statistical information, clinical records, laboratory test results, raw medical imaging data, and traditional Chinese medicine diagnostic data.

[0032] The multiple information systems can represent various methods, locations, departments, and other dimensions related to symptom detection for users. Examples include hospital information systems, laboratory information systems, image archiving systems, and traditional Chinese medicine (TCM) diagnostic systems. Each information system establishes a standard application programming interface (API), and the user's symptom prediction device connects to these APIs to acquire multimodal data. This multimodal data includes statistical information, clinical records, laboratory test results, raw medical imaging data, and various data from the four diagnostic methods of TCM. Laboratory test results include parameters such as blood lipids, blood pressure, blood glucose, and myocardial enzyme levels. Raw hospital imaging data includes parameters such as electrocardiograms, ultrasound data, X-ray data, MRI data, and coronary angiography data. TCM diagnostic data includes parameters such as tongue appearance, pulse diagnosis, and patient history information.

[0033] In this embodiment of the application, for users who need to predict their symptoms, the user's multimodal data is obtained through multiple information systems, and then the user's symptoms are predicted to recur based on the multimodal data.

[0034] Before further use, the acquired multimodal data can be cleaned and transformed. Data fields from different sources are mapped to a unified medical data model (OMOP CDM) to eliminate semantic ambiguity. Secondly, a unified medical language system (UMLS) is used to perform entity recognition and terminology encoding on unstructured information in clinical texts; for example, mapping "chest tightness" to a unique concept identifier in UMLS. For numerical indicators, Z-score normalization is performed to make them conform to the model input requirements. The Z-score normalization formula is as follows:

[0035] ;

[0036] Where x is the original value, μ is the mean of the data, and σ is the standard deviation of the data.

[0037] S120. Integrate and align the multimodal data over time to obtain the feature vectors of Western medicine indicators and the feature vectors of traditional Chinese medicine syndrome elements.

[0038] For example, in multimodal data, each dimension of the data corresponds to information such as the data collection time and location. Before using multimodal data for disease prediction, it is necessary to integrate and align the multimodal data temporally to facilitate subsequent use in disease prediction.

[0039] For example, data integration and temporal alignment of multimodal data can involve integrating the data's numerical information, the time of occurrence, the clinical stage, and local dynamic change models to achieve deep semantic modeling of medical time-series data. Additionally, multimodal data can be aligned along the time dimension, that is, multimodal data corresponding to the same timestamp can be aligned. This allows for temporal analysis and prediction of multimodal data at the same time point during subsequent data processing, making the analysis results more reliable.

[0040] S130. Determine a fusion vector based on the Western medicine indicator feature vector and the traditional Chinese medicine syndrome feature vector, and predict the user's symptoms based on the fusion vector.

[0041] Among them, the Western medicine indicator feature vector and the traditional Chinese medicine syndrome element feature vector reflect the user's physical health characteristics from the perspectives of Western medicine and traditional Chinese medicine, respectively. However, the Western medicine indicator feature vector and the traditional Chinese medicine syndrome element feature vector are two independent feature vectors. A fusion vector can be determined based on the Western medicine indicator feature vector and the traditional Chinese medicine syndrome element feature vector to achieve deep fusion of the two feature vectors.

[0042] In the process of determining the fusion vector based on the characteristic vectors of Western medicine indicators and the characteristic vectors of TCM syndrome elements, the characteristic vectors of Western medicine indicators and the characteristic vectors of TCM syndrome elements can be concatenated, weighted, or fused using an attention mechanism.

[0043] After determining the fusion vector, it can be input into the prediction model to predict the user's symptoms and determine whether the symptoms will recur.

[0044] The technical solution of this application embodiment acquires multimodal data of the same user from multiple information systems. This multimodal data includes statistical information, clinical records, laboratory test results, raw medical imaging data, and multiple data points from the four diagnostic methods of Traditional Chinese Medicine (TCM). The multimodal data is integrated and time-aligned to obtain Western medicine indicator feature vectors and TCM syndrome element feature vectors. A fusion vector is determined based on these two vectors, and the user's symptoms are predicted based on the fusion vector. This solution acquires multimodal data, integrates and time-aligns it to obtain Western medicine indicator feature vectors and TCM syndrome element feature vectors, achieving a balance between Western and TCM features. Furthermore, it performs deep semantic fusion of the Western medicine indicator feature vectors and TCM syndrome element feature vectors, simulating the thought process of TCM syndrome differentiation and treatment. This allows for a more comprehensive capture of the multi-factor influences on the occurrence and development of symptoms, improving the comprehensiveness of prediction and its relevance to pathological mechanisms.

[0045] In this embodiment of the application, after obtaining multimodal data of the same user from multiple information systems, the method further includes:

[0046] The encrypted multimodal data is decrypted using the decryption key to obtain the original multimodal data. The encrypted multimodal data is obtained by each information system encrypting the original data using a symmetric encryption key generated by a random bit generator before sending the multimodal data.

[0047] In this embodiment, before sending user data to the user's symptom prediction device, the information system needs to encrypt the user data to prevent data leakage. Specifically, a 256-bit symmetric encryption key can be generated using a random bit generator. The key manager is responsible for the key's lifecycle management, including distribution, rotation, and attribute-based access control. The user data to be sent is encrypted using the AES-256 algorithm in GCM mode, which provides both confidentiality and data integrity authentication. The encrypted data is transmitted through a secure transmission channel established based on the SSL / TLS 1.3 protocol, which provides forward security to prevent the decryption of historical communications due to session key leakage. After receiving the encrypted multimodal data, the user's symptom prediction device decrypts it using the corresponding decryption key to restore the original multimodal data that can be processed.

[0048] To protect patient privacy, privacy enhancement techniques are implemented before data leaves various information systems. The specific implementation involves: applying a differential privacy algorithm to add calibrated noise to the query results; simultaneously, k-anonymizing the quasi-identifiers (k≥5) ensures that any record in the dataset is indistinguishable from fewer than k individuals. The noise addition formula for differential privacy is as follows:

[0049] ;

[0050] Where f(D) is the query function, and Noise is noise sampled from a Gaussian distribution.

[0051] Figure 2 This is a flowchart illustrating a user symptom prediction method according to another embodiment of this application. This embodiment is an optimization based on the above embodiment; solutions not described in detail in this embodiment are found in the above embodiment. Figure 2 As shown, the method in this embodiment of the application specifically includes the following steps:

[0052] S210. Obtain multimodal data of the same user from multiple information systems; wherein, the multimodal data includes multiple data from statistical information, clinical records, laboratory test results, raw medical imaging data, and traditional Chinese medicine diagnostic data.

[0053] S220. Adaptive time binning and multi-scale temporal position coding are performed on the Western medicine indicator data to obtain Western medicine time series data.

[0054] For example, in processing Western medicine indicator data along the time dimension, the data can be adaptively binned according to its stage characteristics, forming multiple time intervals corresponding to those stages. Furthermore, a vector representation containing deep semantics is generated for each time bucket, and multi-scale temporal position encoding is performed.

[0055] In this embodiment of the application, adaptive time binning and multi-scale temporal position encoding are performed on the Western medicine indicator data to obtain Western medicine time-series data, including:

[0056] Using the date of diagnosis or treatment as the origin, the timeline is divided into different disease stages based on the clinical path of disease development, resulting in multiple time intervals.

[0057] For Western medicine indicator data in each time interval, the absolute and relative position codes are calculated based on sine and cosine functions to determine the stage embedding vector of the Western medicine indicator data, and the codes of data statistical features within the time interval are determined.

[0058] Based on the absolute and relative position encoding, stage embedding vector, and data statistical feature encoding, the Western medicine time series data is determined.

[0059] For example, the timeline can be divided into multiple time intervals based on the user's diagnosis date or treatment date, according to the clinical pathway of disease development, with different disease stages as the origin. The clinical pathway of disease development refers to a standardized diagnosis and treatment process for a specific disease, rather than simply the trajectory of the disease itself. For instance, for heart disease, based on the typical clinical pathway of heart disease development, the timeline can be divided into multiple time intervals according to several medically significant stages such as "acute phase," "recovery phase," and "long-term follow-up phase." Each time interval uses a binning granularity that matches its physiological rate of change: in the highly dynamic acute phase, a "hour" or "day" level granularity is used to capture rapid fluctuations; in the long-term follow-up phase where changes tend to level off, a "month" level granularity is used to represent long-term trends. The boundaries of this binning strategy can also be optimized by key clinical events such as abnormal electrocardiograms and elevated myocardial enzyme levels, ensuring that the model's focus aligns with clinical decision points.

[0060] For Western medicine indicator data across different time intervals, absolute and relative position encoding is calculated based on sine and cosine functions to determine the stage embedding vector of the Western medicine indicator data and the encoding of statistical features within the time interval. Specifically, multi-scale temporal coding can be fused from three components with clear physical meaning: absolute relative position encoding calculated based on sine and cosine functions, learnable clinical stage embedding vectors, and encoding of statistical features within the time interval. The absolute and relative position encoding calculated based on sine and cosine functions provides basic temporal sequence information, much like labeling each time bucket in chronological order, allowing the model to distinguish the order in which events occurred. Here, it is generated using the classic sine and cosine method, simultaneously representing absolute position and relative distances between different positions. The learnable clinical stage embedding vector represents the clinical stage to which the time interval belongs. Each time bucket interval is labeled with its corresponding clinical stage. Through model training, the vectors automatically learn the common patterns of different stages, enabling the model to implicitly learn common patterns across different stages. Without manually writing rules, the model can automatically distinguish the feature differences between different stages. The encoding of statistical features of data within a time interval uses a lightweight neural network to map the dynamic changes of the data within the bucket into a high-dimensional vector. It specifically incorporates actual data information from within the time interval: first, it extracts statistical features of patient data (such as changes in blood pressure and heart rate) within the time interval, and then uses a lightweight neural network to convert this dynamic change information into a high-dimensional vector, ensuring that the final encoding includes information about actual changes in the patient's condition.

[0061] Based on the encoding of absolute and relative positions, stage embedding vectors, and data statistical features, Western medicine time-series data is determined. For example, the weighted sum of the three component vectors can be used, with the weights dynamically calculated based on the overall characteristics of the current patient through a simple attention mechanism. This encoding method ensures that each data point, when input into the model, not only carries its own numerical information but also integrates its time point, clinical stage, and local dynamic data change patterns, thereby achieving deep semantic modeling of medical time-series data.

[0062] S230. Normalize the structured Western medicine indicator data and calculate the Western medicine correlation data between the corresponding Western medicine indicator data sequences.

[0063] For example, there may be significant differences between the two levels of structured Western medicine indicator data. The structured Western medicine indicator data can be normalized to unify the data volume. For the normalized structured Western medicine indicator data, the correlation data between the corresponding Western medicine indicator data sequences can be calculated to reflect the magnitude of the correlation between the Western medicine indicator data.

[0064] In this embodiment, the Spearman correlation coefficient between Western medicine indicator data sequences can be calculated to reflect the correlation data of Western medicine. The formula for the Spearman correlation coefficient is as follows: ; where d i The difference between the i-th Western medicine indicator data in two Western medicine indicator data sequences, where n is the total number of indicator data in the Western medicine indicator sequence.

[0065] S240. Using TCM syndrome elements as nodes, construct a graph structure by determining the weights of edges between nodes based on preset prior weights and the strength of associations.

[0066] For example, a graph structure is constructed using each TCM syndrome element as a node, with edges connecting the nodes. The edge weights can be determined based on preset prior weights and the discovered correlation strength. The preset prior weights can be based on classical TCM theories, and the discovered correlation strength can be the statistical correlation strength between TCM syndrome elements obtained through mining and learning from massive clinical data. The initial feature of each node is its standardized weight.

[0067] S250. Based on the graph attention mechanism, the interaction between TCM syndrome elements, Western medicine time series data and Western medicine correlation data is carried out to determine the TCM syndrome element feature vector, which reflects the intensity of TCM syndrome elements, the topological relationship with other syndrome elements and the correlation with Western medicine indicators.

[0068] For example, a graph attention mechanism can be used to interact with TCM syndrome elements, Western medicine time-series data, and Western medicine correlation data to determine TCM syndrome element feature vectors, reflecting the strength of TCM syndrome elements, their topological relationships with other syndrome elements, and their correlation with Western medicine indicators. For instance, based on the "Syndrome Element Differentiation" standard, the collected TCM four diagnostic methods information is converted into initial quantitative weights for each syndrome element, and these initial weights are represented by feature vectors. Simultaneously, Western medicine time-series data and Western medicine correlation data are also represented by feature vectors. The graph attention mechanism interacts with these feature vectors to determine the TCM syndrome element feature vectors, reflecting the strength of the TCM syndrome elements, their topological relationships with other syndrome elements, and their correlation with Western medicine indicators. For example, the "blood stasis in the heart" syndrome element node will tend to focus on the dynamic trajectory of ST segment changes on electrocardiograms and myocardial enzyme levels at various stages of the disease. Through multiple rounds of graph convolution and feature propagation, the final embedding vector obtained for each syndrome element node is a high-dimensional composite representation that integrates its own strength, its topological relationships with other syndrome elements, and the context of relevant objective Western medicine evidence. This representation is mapped to a uniform 512-dimensional dense vector through an embedding layer, providing a highly information-rich query basis for subsequent deep fusion.

[0069] S260. Determine a fusion vector based on the Western medicine indicator feature vector and the traditional Chinese medicine syndrome feature vector, and predict the user's symptoms based on the fusion vector.

[0070] The scheme in this application embodiment performs adaptive time binning and multi-scale temporal position encoding on Western medicine indicator data to obtain Western medicine time-series data; it normalizes the structured Western medicine indicator data and calculates the Western medicine correlation data between the corresponding Western medicine indicator data sequences; and it determines the Western medicine indicator feature vector based on the Western medicine time-series data and the Western medicine correlation data. The above scheme proposes a dynamic adaptive time binning mechanism, which adaptively adjusts the time granularity according to the clinical stage of the disease, and combines multi-scale fusion temporal position encoding to transform absolute time into a clinically meaningful temporal representation. This enables the model to accurately align the temporal dynamics of multi-source heterogeneous data, better capture the evolutionary patterns of disease risk, and enhance the temporal sensitivity and clinical applicability of predictions.

[0071] Figure 3 This is a flowchart illustrating a user symptom prediction method as another embodiment of this application. This embodiment is an optimization based on the above embodiments; solutions not described in detail in this embodiment are found in the above embodiments. Figure 3 As shown, the method in this embodiment of the application specifically includes the following steps:

[0072] S310. Obtain multimodal data of the same user from multiple information systems; wherein, the multimodal data includes multiple data from statistical information, clinical records, laboratory test results, raw medical imaging data, and traditional Chinese medicine diagnostic data.

[0073] S320. Integrate and align the multimodal data over time to obtain the feature vectors of Western medicine indicators and the feature vectors of traditional Chinese medicine syndrome elements.

[0074] S330. Based on the hierarchical cross-attention mechanism, determine the correlation feature vector between the Western medicine indicator feature vector and the traditional Chinese medicine syndrome feature vector according to the TCM syndrome feature vector, the Western medicine time series data and the Western medicine correlation data in the Western medicine indicator feature vector.

[0075] For example, a hierarchical cross-attention mechanism can be used to determine the correlation feature vector between the Western medicine indicator feature vector and the TCM syndrome feature vector based on the TCM syndrome feature vector, the Western medicine time-series data in the Western medicine indicator feature vector, and the Western medicine correlation data. Specifically, the query vector, key vector, and value vector can be determined using the TCM syndrome feature vector, TCM time-series data, and Western medicine correlation data, and then attention calculation can be performed.

[0076] In this embodiment of the application, based on the hierarchical cross-attention mechanism, the correlation feature vector between the Western medicine indicator feature vector and the traditional Chinese medicine syndrome feature vector is determined according to the Western medicine time-series data and Western medicine correlation data in the traditional Chinese medicine syndrome feature vector, including:

[0077] The TCM syndrome feature vector is input into a gated recurrent unit network to obtain a syndrome state vector containing event dynamic information.

[0078] The state vector of the syndrome element is used as the query vector, the time-series features of Western medicine are used as the key vector, and the relevant data of Western medicine are used as the value vector. Cross-attention calculation is performed to obtain the associated feature vector.

[0079] In this embodiment, the TCM syndrome feature vectors undergo in-depth processing. Instead of treating them as static query vectors, a gated recurrent unit (GRU) network is introduced to encode the syndrome vector sequence. This GRU network can capture the dynamic evolution of syndromes such as "blood stasis" at different postoperative stages, thereby generating syndrome state vectors containing temporal dynamic information. This vector better reflects the prognosis trend of the syndrome, providing query conditions with more temporal contextual information for fusion. The update and reset gate formulas of the GRU are as follows:

[0080] ;

[0081] ;

[0082] ;

[0083] ;

[0084] in, It's an update door. It's a door reset. It is in a hidden state. It is input. It is the sigmoid function. It is element-wise multiplication. t These are input features.

[0085] Deep fusion is achieved through a hierarchical cross-attention mechanism. The evolved syndrome element state vector is used as the query, and cross-attention calculation is performed with Western medicine time-series data and related Western medicine data (as Key and Value) to achieve syndrome element-indicator association perception. The purpose of this layer is to enable each syndrome element state to actively and dynamically retrieve microscopic evidence related to its pathophysiological level from massive Western medicine time-series data, forming "syndrome-disease" association features.

[0086] The formula for calculating the query matrix is ​​as follows:

[0087] Query Matrix ;

[0088] Key matrix ;

[0089] Value matrix ;

[0090] in , , These are learnable parameters.

[0091] The formula for calculating attention is as follows:

[0092] ;

[0093] ;

[0094] in, This is the feature matrix of the evidence element associations output from the first layer. Each row vector represents the feature of an evidence element after association with all indicators. d is the dimension of the key matrix.

[0095] S340. Based on the multi-head self-attention mechanism, determine the fused feature vector according to the associated feature vector.

[0096] For example, the feature vectors associated with each syndrome element are considered as a set of "syndrome element-evidence" reflecting the patient's overall condition. This set is then input into a multi-head self-attention layer. In this layer, the feature vectors associated with each syndrome element perform attention calculations, realizing interaction and global representation between syndrome elements, thereby simulating the integration process of "pathogenesis" in traditional Chinese medicine theory, where syndrome elements are combined and influence each other. The output of this layer is a global, weighted feature vector after interaction between syndrome elements, comprehensively reflecting the patient's overall and dynamic disease state.

[0097] The specific calculation formula is as follows:

[0098] Multi-head self-attention: For each attention head i (where i = 1, ..., h)

[0099] ;

[0100] in, These are learnable parameters.

[0101] Attention calculation for each head:

[0102] ;

[0103] Merge multiple outputs:

[0104] ;

[0105] Linear transformation:

[0106] ;

[0107] in These are learnable parameters.

[0108] Global fusion: for Average pooling is performed to obtain the global fused feature vector.

[0109] ;

[0110] in yes The Row vectors.

[0111] S350. The fused feature vector is residually concatenated with the original features to obtain the fused vector; wherein, the original features include the Western medicine indicator feature vector and the traditional Chinese medicine syndrome element feature vector.

[0112] For example, to stabilize training and preserve original information, a residual connection is made between the fused feature vector and the original features. Subsequently, a gating mechanism is used to regulate the contribution ratio of the fused feature vector and the original features to the final prediction, achieving adaptive information flow control. The gating fusion formula is as follows:

[0113] ;

[0114] ;

[0115] in, It is a gating value. It is a feature of fusion. It is a primitive feature.

[0116] S360. Predict the user's symptoms based on the fusion vector.

[0117] For example, the fused feature representation is input into a fully connected classification layer. To address the potential imbalance between relapse and non-relapse classes in clinical data, a class-weighted cross-entropy loss function is employed. The entire network is fine-tuned end-to-end using the AdamW optimizer (learning rate set to 2e-5, weight decay 0.01), ultimately outputting a relapse probability value between 0 and 1.

[0118] The solution in this application embodiment, based on a hierarchical cross-attention mechanism, determines the correlation feature vector between the Western medicine indicator feature vector and the TCM syndrome feature vector according to the TCM syndrome feature vector, the Western medicine temporal data in the Western medicine indicator feature vector, and the Western medicine correlation data; based on a multi-head self-attention mechanism, it determines the fusion feature vector according to the correlation feature vector; and it performs a residual connection between the fusion feature vector and the original features to obtain the fusion vector; wherein, the original features include the Western medicine indicator feature vector and the TCM syndrome feature vector. The above solution simulates the dynamic thinking process of "holistic review and comprehensive consideration of the four diagnostic methods" in TCM syndrome differentiation and treatment. By introducing temporal dynamism and the correlation between syndrome elements, it elevates TCM syndrome elements from static features to dynamic and intelligent query vectors that guide information fusion. It innovatively uses TCM syndrome elements as dynamic intelligent query vectors to guide the fusion of multimodal information. By using a convolutional network of syndrome element relationship graphs and a cross-modal attention mechanism, deep semantic interaction between TCM syndrome elements and Western medicine indicators was achieved, simulating the thought process of TCM syndrome differentiation and treatment. This allows for a more comprehensive capture of the multi-factor influence on the occurrence and development of heart disease, improving the comprehensiveness of prediction and the relevance to pathological mechanisms.

[0119] In this embodiment of the application, the method further includes:

[0120] The feature contribution of the TCM syndrome element feature vector and the Western medicine indicator feature vector to disease prediction is calculated based on model interpretability technology.

[0121] Based on the prediction error of the user's symptom prediction method and the feature contribution of the TCM syndrome feature vector and the Western medicine indicator feature vector to the symptom prediction, the user's symptom prediction method is optimized; wherein, in the optimization process, the weights of the TCM syndrome feature vector and the Western medicine indicator feature vector are positively correlated with the feature contribution.

[0122] For example, based on model interpretability techniques such as SHAP or LIME, the feature contribution of each feature (TCM syndrome elements and Western medicine indicators) to the individual prediction result is calculated. This is then visually displayed using SHAP force waterfall plots, radar charts, heatmaps, and other visualization methods. Simultaneously, selected significant biomarkers are mapped to pathway databases such as KEGG and Reactome for enrichment analysis, explaining the potential mechanisms of relapse risk from a biological pathway perspective. By comprehensively considering the predicted risk level, feature contribution, and pathway analysis results, a knowledge base built based on clinical guidelines and TCM expert knowledge is queried to generate a personalized plan including Western medicine recommendations, TCM prescription suggestions, lifestyle interventions, and follow-up plans. Model parameters are dynamically adjusted based on the model's prediction error on new data and the calculated feature contribution. A weighted moving average algorithm is used to prioritize real-time, small-amplitude optimization of the mapping weights of TCM syndrome element embedding vectors with high feature contributions, enabling the model to adaptively learn new knowledge without requiring complete retraining.

[0123] This application provides a specific implementation method, including:

[0124] The overall process includes the following six main steps:

[0125] S100: Multimodal medical data acquisition and standardized preprocessing: Collect and clean multimodal data from heterogeneous source systems to provide high-quality input for subsequent analysis.

[0126] S200: Data Security Encryption and Transmission: Establishing an end-to-end secure channel to ensure the confidentiality and integrity of sensitive medical data during transmission.

[0127] S300: Multi-source heterogeneous data integration and precise temporal alignment: Vectorizes and aligns Western medicine indicators, imaging features, omics data and traditional Chinese medicine syndrome elements within a unified temporal framework.

[0128] S400: Cross-modal fusion prediction guided by TCM syndrome elements: Utilizing a pre-trained large model, multimodal features are fused through an innovative attention mechanism, and the probability of relapse risk is calculated.

[0129] S500: Explainable clinical decision support and dynamic model feedback optimization: Transforms prediction results into clinically applicable decision-making basis and uses feedback signals to optimize model parameters in real time.

[0130] S600: Blockchain-based secure evidence storage and multi-center model verification: It realizes the immutable evidence storage of data and model parameters, and ensures the robustness and generalization ability of the model through internal and external verification.

[0131] In the above steps, the feedback signal generated in step S500 will act on the model parameters in steps S300 and S400, forming a continuously optimized closed-loop learning system.

[0132] (II) Detailed explanation of the method and steps

[0133] S100. Multimodal Medical Data Acquisition and Standardized Preprocessing

[0134] This step forms the basis of the method's data input, aiming to address the challenges of integrating and initially standardizing multi-source heterogeneous medical data. Its sub-steps are as follows:

[0135] S101. Data Connection and Extraction: By establishing standard application programming interface connections with hospital information systems, laboratory information systems, image archiving systems, and traditional Chinese medicine diagnostic equipment, data is extracted according to preset time intervals or event triggering mechanisms. This includes patient demographic information, clinical records, laboratory test results (such as blood lipids, blood pressure, blood glucose, myocardial enzyme spectrum, etc.), raw medical image data (such as electrocardiogram, echocardiography, coronary CTA), and information from the four diagnostic methods of traditional Chinese medicine (such as tongue appearance, pulse appearance, and consultation information).

[0136] S102. Data Standardization and Terminology Mapping: The extracted raw data is cleaned and transformed. First, data fields from different sources are mapped to a unified medical data model (OMOP CDM) to eliminate semantic ambiguity. Second, a unified medical language system (UMLS) is used to perform entity recognition and terminology encoding on unstructured information in clinical texts; for example, mapping "chest tightness" to a unique concept identifier in UMLS. For numerical indicators, Z-score normalization is performed to make them conform to the model input requirements. The Z-score normalization formula is as follows:

[0137] ;

[0138] in, It is the original value. It is the mean of the features. It is the standard deviation of the feature.

[0139] S103. Pre-processing for Privacy Computation: To protect patient privacy, privacy enhancement techniques are implemented before data leaves the original data source. Specifically, a differential privacy algorithm is applied to add calibrated noise to the query results; simultaneously, k-anonymization (k≥5) is performed on quasi-identifiers, ensuring that any record in the dataset is indistinguishable from fewer than k individuals. The noise addition formula for differential privacy is as follows:

[0140] ;

[0141] in, It is a query function. It is noise sampled from a Gaussian distribution.

[0142] S200. Data Security Encryption and Transmission

[0143] This step is responsible for building a secure bridge from the data source to the central processing node.

[0144] S201. Key Generation and Management: A 256-bit symmetric encryption key is generated using a random bit generator. The key manager is responsible for the key's lifecycle management, including distribution, rotation, and attribute-based access control.

[0145] S202. Data Encryption and Secure Transmission: Using the key generated in step S201, the standardized data output from step S100 is encrypted in GCM mode using the AES-256 algorithm. This mode provides both confidentiality and data integrity authentication. The encrypted data is transmitted through a secure transmission channel established based on the SSL / TLS 1.3 protocol. This protocol provides forward security to prevent the decryption of historical communications due to session key leakage.

[0146] S203. Data Decryption: After receiving the encrypted data at the central processing node, the corresponding decryption key is used to decrypt the data and restore the original data that can be processed by step S300.

[0147] S300. Multi-source heterogeneous data integration and precise time-series alignment

[0148] The core of this step lies in addressing the heterogeneity of multimodal medical data across time scales. Its key innovation is embodied in the dynamic adaptive temporal binning and multi-scale temporal position encoding method proposed in sub-step S301. This method aims to transform discrete medical events into continuous temporal representations with rich clinical semantics, providing high-quality input for subsequent deep fusion.

[0149] S301. Dynamic Adaptive Temporal Bucketing and Multi-Scale Temporal Position Encoding: This step proposes an adaptive bucketing mechanism based on clinical pathological stages. First, the date of diagnosis of acute myocardial infarction or the date of cardiac surgery is used as the absolute time origin T0. The bucketing process follows the typical clinical pathway of cardiac disease development, dividing the time axis into several medically significant stages such as "acute phase," "recovery phase," and "long-term follow-up phase." Within each stage, a bucketing granularity matching its physiological rate of change is adopted: in the highly dynamic acute phase, an "hourly" or "daily" granularity is used to capture rapid fluctuations; in the long-term follow-up phase where changes tend to level off, a "monthly" granularity is used to characterize long-term trends. The boundaries of this bucketing strategy can also be optimized by key clinical events such as abnormal electrocardiograms and elevated myocardial enzyme levels, ensuring that the model's focus aligns with clinical decision points.

[0150] Building upon the bucketing approach, this method employs a multi-scale fusion temporal position encoding to generate a vector representation with deep semantic meaning for each time bucket. This encoding is composed of three physically meaningful components: first, an absolute / relative position encoding based on sine and cosine functions, providing fundamental temporal sequence information; second, a learnable clinical stage embedding vector to identify the clinical stage to which the time bucket belongs, enabling the model to implicitly learn common patterns across different stages; and third, an encoding based on the statistical features of the data within the time bucket, mapping the dynamic changes in the data within the bucket into a high-dimensional vector using a lightweight neural network. The final temporal position encoding is a weighted sum of these three vector components, with the weights dynamically calculated based on the overall characteristics of the current patient using a simple attention mechanism. This encoding method ensures that each data point, when input into the model, not only carries its own numerical information but also integrates its time point, clinical stage, and local dynamic data change patterns, thereby achieving deep semantic modeling of medical time-series data.

[0151] The sine and cosine position coding formulas are as follows:

[0152] ;

[0153] ;

[0154] in, Where i is the time position and 'i' is the dimension index. It is the model dimension.

[0155] S302. Standardization and Correlation Analysis of Western Medicine Indicators: Structured Western medicine clinical indicators (such as blood pressure, blood lipids, and electrocardiogram features) are subjected to Min-Max normalization, scaling them to the [0,1] interval. Simultaneously, the Spearman correlation coefficient matrix among these indicators is calculated, strongly correlated feature pairs are selected, and a feature association network is constructed for subsequent interpretability analysis. The Min-Max normalization formula is as follows:

[0156] ;

[0157] Simultaneously, the Spearman correlation coefficient matrix between indicators is calculated to identify strongly correlated feature pairs and construct a feature association network for subsequent interpretability analysis. The Spearman correlation coefficient formula is as follows:

[0158] ;

[0159] in, It's a difference in rank. That is the number of samples.

[0160] S303. Quantification and Dynamic Calibration of TCM Syndrome Elements: The core innovation of this sub-step lies in constructing a quantitative representation system that can dynamically reflect the complex interactions between syndrome elements and their correlation with Western medicine indicators. First, strictly adhering to the standards of *Syndrome Element Differentiation*, the collected TCM four diagnostic methods information is transformed into initial quantitative weights for each syndrome element. To overcome the limitations of static weight representation, this invention introduces a syndrome element relationship graph convolutional network. In this network, syndrome elements act as nodes, and the edge weights between nodes consist of two parts: one is the prior weight set based on classical TCM theory; the other is the statistical correlation strength between syndrome elements obtained through mining and learning from massive clinical data. The initial feature of each node is its standardized weight.

[0161] The graph structure then interacts cross-modal with the temporalized Western medicine features processed in steps S301 and S302. To this end, a graph-temporal feature interaction module is designed. This module utilizes a graph attention mechanism, enabling each syndrome element node to adaptively focus on the temporal patterns of the Western medicine indicators most relevant to it, based on its semantic connotation. For example, the "blood stasis" syndrome element node will tend to focus on the dynamic trajectories of ST segment changes on electrocardiograms and myocardial enzyme levels at various stages of the disease course. Through multiple rounds of graph convolution and feature propagation, the final embedding vector obtained for each syndrome element node is a high-dimensional composite representation that integrates its own strength, topological relationships with other syndrome elements, and the context of associated objective Western medicine evidence. This representation is mapped to a unified 512-dimensional dense vector through an embedding layer, providing a highly information-rich query foundation for subsequent deep fusion.

[0162] S304. Missing Data Imputation: For data such as genomics and proteomics that may contain missing values, conditional generative adversarial networks (GANs) are used for imputation. The generator uses observed features as conditions to generate missing values ​​consistent with the distribution of the real data, while the discriminator distinguishes between the generated data and the real data. Adversarial training improves the reliability of the imputed data. The loss function of the conditional generative adversarial network is as follows:

[0163] ;

[0164] Where G is the generator, D is the discriminator, and y is the condition information.

[0165] S400. Cross-modal fusion prediction guided by TCM syndrome elements

[0166] This step is the core computational unit for achieving deep integration of traditional Chinese medicine and Western medicine information and risk prediction. Its innovation is mainly reflected in the multi-level, dynamically evolving cross-modal attention fusion mechanism proposed in sub-step S402. The core of this mechanism lies in simulating the dynamic thinking process of "holistic review and comprehensive consideration of the four diagnostic methods" in traditional Chinese medicine syndrome differentiation and treatment. By introducing temporal dynamism and the correlation between syndrome elements, it elevates traditional Chinese medicine syndrome elements from static features to dynamic and intelligent query vectors that guide information fusion.

[0167] S401. Basic Model Transfer and Domain Adaptation: A Transformer model pre-trained on large public medical datasets (MIMIC-III, PubMed) is used as the foundation. The model's vocabulary is expanded by adding approximately 200 new TCM syndrome element terms and their corresponding word vectors, enabling the model to understand TCM domain knowledge.

[0168] S402. Cross-modal attention fusion: This sub-step is the core innovation of this invention. It constructs a dynamic, multi-layered fusion architecture that incorporates the logic of traditional Chinese medicine diagnosis.

[0169] Syndrome element state evolution encoding: First, the TCM syndrome element embedding vector output from step S303 is further processed. Instead of treating it as a static query vector, a gated recurrent unit (GRU) network is introduced to encode the syndrome element vector sequence in postoperative temporal order. This GRU network can capture the dynamic evolution of syndrome elements such as "blood stasis" at different postoperative stages, thereby generating syndrome element state vectors containing temporal dynamic information. This vector better reflects the prognosis trend of the syndrome, providing query conditions with more temporal contextual information for fusion. The formulas for the GRU update gate and reset gate are as follows:

[0170] ;

[0171] ;

[0172] ;

[0173] ;

[0174] in, It's an update door. It's a door reset. It is in a hidden state. It is input. It is the sigmoid function. It is element-wise multiplication.

[0175] Layered Cross-Attention Fusion: Deep fusion is achieved through a layered cross-attention mechanism. The first layer uses the evolved syndrome element state vector as the query and performs cross-attention calculation with the Western medicine indicator feature vectors (as Key and Value) processed by S301 and S302 at all time points to achieve syndrome element-indicator association perception. The purpose of this layer is to enable each syndrome element state to actively and dynamically retrieve microscopic evidence related to it at the pathophysiological level from massive Western medicine time-series data, forming "syndrome-disease" association features.

[0176] The formula for calculating the query matrix is ​​as follows:

[0177] ;

[0178] ;

[0179] ;

[0180] in , , These are learnable parameters.

[0181] The formula for calculating attention is as follows:

[0182] ;

[0183] ;

[0184] in, It is the evidence element association feature matrix output by the first layer. Each row vector represents the feature of an evidence element after association with all indicators.

[0185] The second layer treats the feature vectors associated with each syndrome element output from the first layer as a set of "syndrome element-evidence" reflecting the patient's overall condition. This set is then input into a multi-head self-attention layer. In this layer, attention calculations are performed on the associated features of each syndrome element, achieving interaction and global representation between syndrome elements, thereby simulating the integration process of "pathogenesis" in traditional Chinese medicine theory, where syndrome elements are combined and influence each other. The output of this layer is a global, weighted feature vector after interaction between syndrome elements, which comprehensively reflects the patient's overall and dynamic disease state.

[0186] The specific calculation formula is as follows:

[0187] Multi-head self-attention: For each attention head i (where i = 1, ..., h)

[0188] ;

[0189] in, These are learnable parameters.

[0190] Attention calculation for each head:

[0191] ;

[0192] Merge multiple outputs:

[0193] ;

[0194] Linear transformation:

[0195] ;

[0196] in Learnable parameters

[0197] Global fusion: for Average pooling is performed to obtain the global fused feature vector.

[0198] ;

[0199] in yes The Row vectors.

[0200] Residual Connection and Gated Fusion: To stabilize training and preserve original information, a residual connection is made between the global fusion features output from the second layer and the input from the first layer (i.e., the aggregated representation of the basic Western medicine indicator feature sequence). Subsequently, a gating mechanism is used to regulate the contribution ratio of the fusion features and the original features to the final prediction, achieving adaptive information flow control. The gated fusion formula is as follows:

[0201] ;

[0202] ;

[0203] in, It is a gating value. It is a feature of fusion. It is a primitive feature.

[0204] S403. Risk Prediction Fine-tuning: The fused feature representation output from step S402 is input into a fully connected classification layer. To address the potential class imbalance between relapsed and non-relapsed data in clinical data, a class-weighted cross-entropy loss function is employed. The AdamW optimizer is used (learning rate set to 2e). -5 The entire network is fine-tuned end-to-end with a weight decay of 0.01, and the final output is a recurrence probability value between 0 and 1.

[0205] S500. Explainable clinical decision support and dynamic feedback optimization of models.

[0206] This step translates model predictions into clinical action and enables the system to evolve on its own.

[0207] S501. Risk Visualization and Mechanism Analysis: Based on model interpretability techniques such as SHAP or LIME, the feature contribution of each characteristic (TCM syndrome elements and Western medicine indicators) to the individual prediction result is calculated. This is visually presented using SHAP force waterfall plots, radar charts, heatmaps, and other visualization methods. Simultaneously, significant biomarkers selected are mapped to pathway databases such as KEGG and Reactome for enrichment analysis, explaining the potential mechanisms of relapse risk from a biological pathway perspective.

[0208] S502. Personalized Intervention Recommendation Generation: Based on the comprehensive prediction of risk level, feature contribution, and pathway analysis results, the system queries a knowledge base built on clinical guidelines and TCM expert knowledge to generate personalized solutions that include Western medicine recommendations, TCM prescription recommendations, lifestyle interventions, and follow-up plans.

[0209] S503. Dynamic Feedback Optimization of Model Parameters: Based on the model's prediction error on new data and the feature contribution calculated in step S501, the model parameters are dynamically adjusted. A weighted moving average algorithm is used to prioritize real-time, small-amplitude optimization of the mapping weights of TCM syndrome element embedding vectors with high feature contributions, enabling the model to adaptively learn new knowledge without requiring complete retraining.

[0210] S600. Secure Evidence Storage and Multi-Center Model Verification Based on Blockchain

[0211] This step provides security and reliability assurance for the entire method and evaluates the model's generalization ability.

[0212] S601. Data Security Evidence Storage: SHA-256 hash values ​​are calculated for key data such as feature vectors, model parameters, and prediction results throughout the entire process. These hash values, along with encrypted data index pointers, timestamps, and operator identities, are then written into the blockchain network. Leveraging the immutability of the blockchain, a complete and reliable data audit trail is created. Large-scale raw data can be stored in distributed storage systems such as IPFS; the blockchain only stores its content-addressing hash.

[0213] S602. Internal and External Validation: Internal Validation: Five-fold cross-validation is used to assess the model's stability, and bootstrap sampling is used to calculate the 95% confidence intervals for metrics such as AUC, accuracy, and F1 score. External Validation: Through smart contracts deployed on the blockchain, multiple independent medical centers are coordinated to run the model on local encrypted data, and the validation results from each center are aggregated to evaluate the model's generalization performance on unseen data.

[0214] S603. Continuous Optimization of Feature System: Attribution analysis is performed on the prediction error samples identified in S602 verification to identify the features that caused the errors. For features with consistently low contribution or that introduce noise, their weights are gradually reduced until they are removed from the feature system, thereby achieving iterative optimization of feature selection.

[0215] Figure 4 This is a schematic diagram of a user symptom prediction device provided in an embodiment of this application. This device can execute the user symptom prediction method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing the method. For example... Figure 4 As shown, the device includes:

[0216] The multimodal data acquisition module 410 is used to acquire multimodal data of the same user from multiple information systems; wherein, the multimodal data includes multiple data from statistical information, clinical records, laboratory test results, raw medical imaging data, and traditional Chinese medicine four diagnostic methods data;

[0217] Data processing module 420 is used to integrate and time-series align multimodal data to obtain feature vectors of Western medicine indicators and feature vectors of traditional Chinese medicine syndrome elements.

[0218] The symptom prediction module 430 is used to determine a fusion vector based on the Western medicine indicator feature vector and the traditional Chinese medicine syndrome feature vector, and to predict the user's symptom based on the fusion vector.

[0219] In this embodiment, the data processing module 420 integrates and aligns the multimodal data over time to obtain a feature vector of Western medicine indicators, including:

[0220] Adaptive time binning and multi-scale temporal location coding were performed on the Western medicine indicator data to obtain Western medicine time-series data.

[0221] The structured Western medicine indicator data were normalized, and the correlation data between the Western medicine indicator data sequences corresponding to each Western medicine indicator were calculated.

[0222] The feature vector of Western medicine indicators is determined based on the Western medicine time series data and the Western medicine correlation data.

[0223] In this embodiment, the data processing module 420 performs adaptive time binning and multi-scale temporal position encoding on the Western medicine indicator data to obtain Western medicine time-series data, including:

[0224] Using the date of diagnosis or treatment as the origin, the timeline is divided into different disease stages based on the clinical path of disease development, resulting in multiple time intervals.

[0225] For Western medicine indicator data in each time interval, the absolute and relative position codes are calculated based on sine and cosine functions to determine the stage embedding vector of the Western medicine indicator data, and the codes of data statistical features within the time interval are determined.

[0226] Based on the absolute and relative position encoding, stage embedding vector, and data statistical feature encoding, the Western medicine time series data is determined.

[0227] In this embodiment, the data processing module 420 integrates and aligns the multimodal data according to time series to obtain a TCM syndrome feature vector, including:

[0228] Using TCM syndrome elements as nodes, and determining the weights of edges between nodes based on preset prior weights and the strength of associations, a graph structure is constructed.

[0229] Based on the graph attention mechanism, the interaction between TCM syndrome elements, Western medicine time series data and Western medicine correlation data is carried out to determine the feature vector of TCM syndrome elements, which reflects the intensity of TCM syndrome elements, the topological relationship with other syndrome elements and the correlation with Western medicine indicators.

[0230] In this embodiment, the symptom prediction module 430 determines a fusion vector based on the Western medicine indicator feature vector and the traditional Chinese medicine syndrome element feature vector, including:

[0231] Based on the hierarchical cross-attention mechanism, the correlation feature vector between the TCM syndrome feature vector and the TCM syndrome feature vector is determined according to the TCM syndrome feature vector, the Western medicine time-series data and Western medicine correlation data in the Western medicine indicator feature vector;

[0232] Based on the multi-head self-attention mechanism, a fused feature vector is determined according to the associated feature vector;

[0233] The fused feature vector is obtained by performing a residual concatenation with the original features; wherein, the original features include the Western medicine indicator feature vector and the traditional Chinese medicine syndrome element feature vector.

[0234] In this embodiment, the symptom prediction module 430, based on a hierarchical cross-attention mechanism, determines the correlation feature vector between the Western medicine indicator feature vector and the traditional Chinese medicine syndrome feature vector according to the TCM syndrome feature vector, the Western medicine time-series data in the Western medicine indicator feature vector, and the Western medicine correlation data, including:

[0235] The TCM syndrome feature vector is input into a gated recurrent unit network to obtain a syndrome state vector containing event dynamic information.

[0236] The state vector of the syndrome element is used as the query vector, the time-series features of Western medicine are used as the key vector, and the relevant data of Western medicine are used as the value vector. Cross-attention calculation is performed to obtain the associated feature vector.

[0237] In this embodiment of the application, the device further includes an optimization module, used for:

[0238] The feature contribution of the TCM syndrome element feature vector and the Western medicine indicator feature vector to disease prediction is calculated based on model interpretability technology.

[0239] Based on the prediction error of the user's symptom prediction method and the feature contribution of the TCM syndrome feature vector and the Western medicine indicator feature vector to the symptom prediction, the user's symptom prediction method is optimized; wherein, in the optimization process, the weights of the TCM syndrome feature vector and the Western medicine indicator feature vector are positively correlated with the feature contribution.

[0240] In this embodiment of the application, the device further includes a decryption module, used for:

[0241] The encrypted multimodal data is decrypted using the decryption key to obtain the original multimodal data. The encrypted multimodal data is obtained by each information system encrypting the original data using a symmetric encryption key generated by a random bit generator before sending the multimodal data.

[0242] The user symptom prediction device provided in this application embodiment can execute a user symptom prediction method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of executing the method.

[0243] Figure 5 A schematic diagram of the structure of an electronic device 10 that can be used to implement embodiments of this application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital processors, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of this application as described herein and / or the first preset requirements.

[0244] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0245] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0246] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as user symptom prediction methods.

[0247] In some embodiments, the user symptom prediction method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the user symptom prediction method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the user symptom prediction method by any other suitable means (e.g., by means of firmware).

[0248] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0249] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable user symptom prediction device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0250] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0251] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0252] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0253] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0254] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the user symptom prediction method as provided in any embodiment of this application.

[0255] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0256] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired information of the technical solution of this application can be achieved, and this is not limited herein.

[0257] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made based on the first design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for predicting user symptoms, characterized in that, The method includes: Multimodal data of the same user is obtained from multiple information systems; the multimodal data includes statistical information, clinical records, laboratory test results, raw medical imaging data, and multiple data from the four diagnostic methods of traditional Chinese medicine. Multimodal data is integrated and time-series aligned to obtain the feature vectors of Western medicine indicators and the feature vectors of traditional Chinese medicine syndrome elements; A fusion vector is determined based on the feature vectors of Western medicine indicators and the feature vectors of traditional Chinese medicine syndrome elements, and the user's symptoms are predicted based on the fusion vector.

2. The method according to claim 1, characterized in that, Multimodal data is integrated and time-series aligned to obtain the feature vector of Western medicine indicators, including: Adaptive time binning and multi-scale temporal location coding were performed on the Western medicine indicator data to obtain Western medicine time-series data. The structured Western medicine indicator data were normalized, and the correlation data between the Western medicine indicator data sequences corresponding to each Western medicine indicator were calculated. The feature vector of Western medicine indicators is determined based on the Western medicine time series data and the Western medicine correlation data.

3. The method according to claim 2, characterized in that, Adaptive time binning and multi-scale temporal position encoding were applied to the Western medicine indicator data to obtain Western medicine time-series data, including: Using the date of diagnosis or treatment as the origin, the timeline is divided into different disease stages based on the clinical path of disease development, resulting in multiple time intervals. For Western medicine indicator data in each time interval, the absolute and relative position codes are calculated based on sine and cosine functions to determine the stage embedding vector of the Western medicine indicator data, and the codes of data statistical features within the time interval are determined. Based on the absolute and relative position encoding, stage embedding vector, and data statistical feature encoding, the Western medicine time series data is determined.

4. The method according to claim 1, characterized in that, Multimodal data is integrated and time-series aligned to obtain TCM syndrome feature vectors, including: Using TCM syndrome elements as nodes, and determining the weights of edges between nodes based on preset prior weights and the strength of associations, a graph structure is constructed. Based on the graph attention mechanism, the interaction between TCM syndrome elements, Western medicine time series data and Western medicine correlation data is carried out to determine the feature vector of TCM syndrome elements, which reflects the intensity of TCM syndrome elements, the topological relationship with other syndrome elements and the correlation with Western medicine indicators.

5. The method according to claim 1, characterized in that, The fusion vector is determined based on the Western medicine indicator feature vector and the traditional Chinese medicine syndrome element feature vector, including: Based on the hierarchical cross-attention mechanism, the correlation feature vector between the TCM syndrome feature vector and the TCM syndrome feature vector is determined according to the TCM syndrome feature vector, the Western medicine time-series data and Western medicine correlation data in the Western medicine indicator feature vector; Based on the multi-head self-attention mechanism, a fused feature vector is determined according to the associated feature vector; The fused feature vector is obtained by performing a residual concatenation with the original features; wherein, the original features include the Western medicine indicator feature vector and the traditional Chinese medicine syndrome element feature vector.

6. The method according to claim 5, characterized in that, Based on a hierarchical cross-attention mechanism, the correlation feature vector between the TCM syndrome feature vector and the TCM syndrome feature vector is determined according to the TCM syndrome feature vector, the Western medicine time-series data in the Western medicine indicator feature vector, and the Western medicine correlation data, including: The TCM syndrome feature vector is input into a gated recurrent unit network to obtain a syndrome state vector containing event dynamic information. The state vector of the syndrome element is used as the query vector, the time-series features of Western medicine are used as the key vector, and the relevant data of Western medicine are used as the value vector. Cross-attention calculation is performed to obtain the associated feature vector.

7. The method according to claim 1, characterized in that, The method further includes: The feature contribution of the TCM syndrome element feature vector and the Western medicine indicator feature vector to disease prediction is calculated based on model interpretability technology. Based on the prediction error of the user's symptom prediction method and the feature contribution of the TCM syndrome feature vector and the Western medicine indicator feature vector to the symptom prediction, the user's symptom prediction method is optimized; wherein, in the optimization process, the weights of the TCM syndrome feature vector and the Western medicine indicator feature vector are positively correlated with the feature contribution.

8. The method according to claim 1, characterized in that, After obtaining multimodal data of the same user from multiple information systems, the method further includes: The encrypted multimodal data is decrypted using the decryption key to obtain the original multimodal data. The encrypted multimodal data is obtained by each information system encrypting the original data using a symmetric encryption key generated by a random bit generator before sending the multimodal data.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the user symptom prediction method according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the user symptom prediction method according to any one of claims 1-8.