Intelligent diagnosis auxiliary analysis system based on medical data driving

By designing an intelligent diagnostic auxiliary analysis system, the problem of multi-source medical data fusion was solved, achieving efficient and accurate medical diagnosis. The accuracy and efficiency of diagnosis were improved by utilizing graph neural networks and Transformer models.

CN121034597APending Publication Date: 2025-11-28张美华
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511146588.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Traditional medical diagnosis relies on doctors' experience and judgment, and cannot effectively integrate multi-source medical data, resulting in low diagnostic efficiency and accuracy.

Method used

Design an intelligent diagnostic auxiliary analysis system based on medical data, including a multi-source medical data acquisition module, a dynamic feature engineering analysis module, a medical feature screening module, and a diagnostic auxiliary analysis decision module. Diagnostic reasoning is performed by combining graph neural networks and Transformer models through missing difference completion, temporal alignment, dynamic feature extraction, and feature screening.

Benefits of technology

It enables efficient integration and analysis of multi-source medical data, improves the accuracy and efficiency of diagnosis, provides a comprehensive understanding of the potential relationships between different data sources, and offers scientific diagnostic recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034597A_ABST
    Figure CN121034597A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical information, in particular to an intelligent diagnosis auxiliary analysis system based on medical data driving. The system comprises a multi-source medical data acquisition module, a dynamic feature engineering analysis module, a medical feature screening module and a diagnosis auxiliary analysis decision module, and can acquire medical image data, medical laboratory examination data and medical gene reaction data and perform missing difference complementation and time sequence alignment. Obtaining a multi-source medical time sequence driving data set; performing dynamic feature engineering analysis and feature offset contribution screening on the multi-source medical time sequence driving data set to obtain a multi-source medical important contribution feature set; and constructing a corresponding medical drive diagnosis analysis model based on the graph neural network and Transform in a mixed manner to carry out diagnosis reasoning prediction so as to output a diagnosis probability corresponding to the medical data, and carrying out decision analysis according to the diagnosis probability to obtain a corresponding medical diagnosis suggestion scheme. The accuracy of medical diagnosis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of medical information technology, and particularly relates to an intelligent diagnosis auxiliary analysis system based on medical data driving. BACKGROUND

[0002] With the continuous progress of medical technology, the collection and application of medical data are becoming more and more extensive, especially in the process of clinical diagnosis and treatment, the diversity and complexity of medical data make the diagnosis process face great challenges. In recent years, with the rapid development of artificial intelligence technology, especially machine learning and deep learning, the intelligent diagnosis auxiliary analysis method driven by medical data has attracted widespread attention. Using machine learning and deep learning technology, a large amount of medical data (such as image data, laboratory test data, medical record data, etc.) can be automatically analyzed to mine potential rules and information, assist doctors in disease prediction, diagnosis, decision-making, etc., especially in the fields of medical imaging, genomics, electronic health records (EHR), intelligent data analysis methods have achieved remarkable results. However, traditional medical diagnosis mainly depends on the experience and judgment of doctors, although doctors have rich knowledge and clinical experience in some aspects, but when facing a large amount of medical data, due to the heterogeneity of different types of medical data, it is impossible to comprehensively integrate multi-source medical data, resulting in low diagnosis efficiency and low accuracy. SUMMARY

[0003] Therefore, it is necessary to provide an intelligent diagnosis auxiliary analysis system based on medical data driving to solve at least one of the above technical problems.

[0004] To achieve the above purpose, an intelligent diagnosis auxiliary analysis system based on medical data driving includes the following modules:

[0005] A multi-source medical data acquisition module is used to acquire medical image data, medical laboratory examination data and medical genetic reaction data, and to complete the missing difference and time sequence alignment of the medical image data, medical laboratory examination data and medical genetic reaction data to obtain a multi-source medical time sequence driving data set;

[0006] A dynamic feature engineering analysis module is used to perform dynamic feature engineering analysis on the corresponding medical image, laboratory data and genetic reaction data in the multi-source medical time sequence driving data set to obtain a multi-source medical data feature full set composed of medical image features, laboratory index fluctuation features and genetic pathological reaction scores;

[0007] A medical feature screening module is used to perform feature offset contribution screening on each medical sub-feature in the multi-source medical data feature full set to obtain a multi-source medical important contribution feature set;

[0008] The diagnostic auxiliary analysis decision module is configured to construct a medical-driven diagnosis analysis model based on a graph neural network and a Transformer hybrid, input a multi-source medical important contribution feature set into the medical-driven diagnosis analysis model for diagnosis reasoning and prediction, output a diagnosis probability corresponding to the medical data, and analyze a corresponding medical diagnosis suggestion scheme based on the diagnosis probability.

[0009] Further, the multi-source medical data acquisition module includes the following functions:

[0010] Obtaining medical image data;

[0011] Obtaining medical laboratory examination data;

[0012] Obtaining medical genetic reaction data;

[0013] Performing pixel blur calculation on the medical image data to obtain medical image pixel blur; and performing blur denoising processing on the medical image data based on the medical image pixel blur to generate medical denoised image data;

[0014] Performing missing difference completion analysis on the medical laboratory examination data and the medical genetic reaction data to obtain medical laboratory difference completion data and medical genetic reaction difference completion data;

[0015] Performing time sequence alignment processing on the medical denoised image data, the medical laboratory difference completion data, and the medical genetic reaction difference completion data by a dynamic time warping algorithm to obtain a multi-source medical time sequence driven data set.

[0016] Further, the missing difference completion analysis on the medical laboratory examination data and the medical genetic reaction data includes:

[0017] Determining data missing points of the medical laboratory examination data and the medical genetic reaction data to obtain corresponding data missing points in the laboratory data and the genetic data;

[0018] Setting a neighborhood range interval corresponding to a neighborhood radius of 1, and filling the corresponding data missing points with the data mean in the neighborhood radius based on the neighborhood range interval to obtain corresponding medical missing completion data;

[0019] Obtaining a completion data distribution from the medical missing completion data, and calculating a data distribution difference value between the completion data distribution and an original data distribution before completion based on the difference between the completion data distribution and the original data distribution to obtain a data distribution difference value between the completion data and the original data;

[0020] The data distribution difference value between the completed data and the original data is compared and judged according to a preset distribution difference threshold 0.05, if the data distribution difference value between the completed data and the original data is less than or equal to the preset distribution difference threshold 0.05, the neighborhood radius corresponding to the neighborhood range interval is increased by 1 and the corresponding data missing point is refilled, until the data distribution difference value between the completed data and the original data is greater than the preset distribution difference threshold 0.05, to obtain the medical laboratory difference completed data and the medical gene reaction difference completed data.

[0021] Further, the dynamic feature engineering analysis module includes the following functions:

[0022] The corresponding medical images in the multi-source medical time series driven data set are subjected to grayscale processing to generate a medical grayscale image sequence;

[0023] The corresponding gray level co-occurrence matrix is obtained through the medical grayscale image sequence, and the medical grayscale image sequence is subjected to image feature statistical analysis based on the gray level co-occurrence matrix to obtain medical image features, including contrast, uniformity, correlation, energy and entropy corresponding to image texture features;

[0024] The laboratory data corresponding to the multi-source medical time series driven data set is subjected to experimental index fluctuation analysis to obtain laboratory index fluctuation features;

[0025] The pathological reaction of the gene reaction data corresponding to the multi-source medical time series driven data set is evaluated to obtain a gene pathological reaction score;

[0026] The corresponding medical image features, laboratory index fluctuation features and gene pathological reaction scores are combined in a data set to obtain a corresponding multi-source medical data feature set.

[0027] Further, the laboratory data corresponding to the multi-source medical time series driven data set is subjected to experimental index fluctuation analysis, including:

[0028] The laboratory data corresponding to the multi-source medical time series driven data set is subjected to experimental index change amplitude analysis to obtain laboratory index change amplitude;

[0029] The laboratory data corresponding to the multi-source medical time series driven data set is subjected to experimental index change frequency statistical analysis based on the laboratory index change amplitude to obtain laboratory index change frequency;

[0030] The laboratory data corresponding to the multi-source medical time series driven data set is subjected to experimental index change amplitude analysis to obtain laboratory index change amplitude;

[0031] The index fluctuation slope is calculated based on the change abnormal fluctuation segment corresponding to each laboratory index, and the laboratory index fluctuation slope is obtained.

[0032] The laboratory index fluctuation characteristics are obtained by combining the laboratory index change amplitude, the laboratory index change frequency and the laboratory index fluctuation slope.

[0033] Further, the index fluctuation slope calculation based on the change abnormal fluctuation segment corresponding to each laboratory index comprises:

[0034] The corresponding abnormal fluctuation time start point and abnormal fluctuation time end point are obtained through the change abnormal fluctuation segment corresponding to each laboratory index.

[0035] The corresponding abnormal fluctuation duration is calculated according to the abnormal fluctuation time start point and the abnormal fluctuation time end point.

[0036] The corresponding laboratory index change amplitude is obtained through the change abnormal fluctuation segment corresponding to each laboratory index, and the index fluctuation slope calculation is performed on the laboratory index change amplitude based on the abnormal fluctuation duration, and the laboratory index fluctuation slope is obtained.

[0037] Further, the pathological reaction evaluation of the corresponding gene reaction data in the multi-source medical time series driven data set comprises:

[0038] The gene reaction feature matrix analysis is performed on the corresponding gene reaction data in the multi-source medical time series driven data set, and the medical gene reaction feature matrix is obtained, which includes the change rate of gene reaction expression, the corresponding clinical treatment reaction time point and the influence factors of gene reaction mutation;

[0039] The disease pathological reaction stage is obtained, and the pathological reaction path simulation is performed on the disease pathological reaction stage based on the medical gene reaction feature matrix and combined with the Monte Carlo simulation method, to generate the individualized gene pathological process reaction path;

[0040] Based on the individualized gene pathological process reaction path, the gene pathological reaction comprehensive evaluation calculation is performed on the medical gene reaction feature matrix by using the weighted average method and combining the actual medical clinical reaction corresponding to the patient, and the gene pathological reaction score is obtained.

[0041] Further, the medical feature screening module comprises the following functions:

[0042] The feature dimension distribution of each medical sub-feature in the multi-source medical data feature set is statistically obtained, and the data feature dimension distribution corresponding to each medical sub-feature is obtained.

[0043] Based on the data feature dimension distribution corresponding to each medical sub-feature, KL divergence is calculated for the corresponding medical sub-features in the full set of multi-source medical data features to obtain the feature distribution offset KL divergence corresponding to each medical sub-feature.

[0044] Based on the feature distribution offset KL divergence corresponding to each medical sub-feature, medical sub-features with a value greater than 0.1 are selected, and the corresponding feature contribution selection process is automatically triggered. The contribution of each medical sub-feature is quantified by the Shapley value. At the same time, the medical sub-features are sorted from largest to smallest and the top 80% are retained to obtain a set of important contribution features from multiple sources of medical care.

[0045] Furthermore, the diagnostic auxiliary analysis and decision-making module includes the following functions:

[0046] A medical-driven diagnostic analysis model is constructed based on a hybrid graph neural network and Transformer.

[0047] The medical sub-features within the multi-source medical important contribution feature set are divided into three branches: imaging branch, experimental indicator branch, and gene branch. The medical sub-features corresponding to the three branches are then input into the medical-driven diagnostic analysis model for diagnostic inference and prediction, so as to output the diagnostic probability corresponding to the medical data.

[0048] Obtain a medical clinical knowledge base, and based on the diagnostic probability and the decision analysis of the medical clinical knowledge base, generate corresponding medical diagnostic recommendations.

[0049] Furthermore, the step of inputting the medical sub-features corresponding to the three branches into the medical-driven diagnostic analysis model for diagnostic inference and prediction includes:

[0050] The learning rate of the medical-driven diagnostic analysis model was set to 0.5.

[0051] By inputting the medical sub-features corresponding to the three branches into the medical-driven diagnostic analysis model, and by calculating the network AUC improvement rate corresponding to each epoch update time on the corresponding branches of the medical-driven diagnostic analysis model;

[0052] The weights on each branch are dynamically adjusted based on the learning rate and the network AUC improvement rate corresponding to each epoch update time. The corresponding diagnosis probability is calculated by weighted inference prediction based on the weights and the corresponding medical sub-features on each branch, so as to output the diagnosis probability corresponding to the medical data.

[0053] The beneficial effects of this invention are:

[0054] The intelligent diagnostic auxiliary analysis system based on medical data proposed in this invention consists of a multi-source medical data acquisition module, a dynamic feature engineering analysis module, a medical feature screening module, and a diagnostic auxiliary analysis and decision-making module. Compared with the prior art, the beneficial effect of this application lies in acquiring and integrating multi-source data. By completing missing values ​​and aligning time series, it provides an accurate and coherent data foundation for subsequent analysis. In actual medical data processing, due to differences in the acquisition time, frequency, and quality of different data sources, data loss or time series inconsistencies often occur. By completing missing data, it can be ensured that the model will not be biased due to missing values ​​during processing. Through time series alignment, it can be ensured that the time nodes of each data source can be accurately matched, so that the interrelationships between data can be effectively reflected at the same time. Through data fusion and processing, a structured and time-consistent dataset can be obtained, which is convenient for further analysis and mining of potential patterns, thereby better integrating heterogeneous medical data. Secondly, through in-depth analysis of medical imaging, laboratory data, and gene response data, dynamic features are extracted. For example, important imaging features can be extracted from imaging data using convolutional neural networks, while the fluctuation patterns of various indicators can be extracted from laboratory data based on time series characteristics. Gene response data can capture the dynamic changes in pathological responses through specific scoring systems. Through these multi-dimensional feature extractions, the potential relationships and interactions between different data sources can be comprehensively understood. Then, to improve the accuracy and efficiency of the model, these analyzed features need to be screened to identify those that contribute most to the analysis task. Feature shift contribution screening is mainly based on the influence of features on the target task, thus retaining the most representative features. This step aims to improve model performance and interpretability by reducing the interference of redundant features, avoiding overfitting, and improving generalization ability. Through shift analysis of multi-source data features, it is possible to reveal which features have significant commonalities and correlations across different data sources, providing accurate input to the model and thus improving the accuracy of diagnostic inference and prediction.Finally, a hybrid model combining Graph Neural Networks (GNNs) and Transformers will be used to process multi-source medical data features. GNNs excel at handling data with graph structures, capturing complex relationships between nodes (such as different medical data sources and medical indicators), while Transformers, with their powerful sequence modeling capabilities, can efficiently process time-series data. This hybrid model leverages the strengths of both GNNs and Transformers to accurately capture structural and temporal features in the data, thereby improving the accuracy and efficiency of medical data analysis. By inputting the selected set of important multi-source medical features into the model, the model can perform deep learning based on these features, deduce relevant patterns in the medical data, and output corresponding analysis results. The model's diagnostic inference predictions can not only provide more accurate analysis but also provide scientific basis for decision-makers based on the prediction results. This approach efficiently processes multi-source heterogeneous data and utilizes advanced deep learning technology to uncover potential complex relationships, thereby improving the efficiency and accuracy of diagnosis. Attached Figure Description

[0055] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0056] Figure 1 This is a schematic diagram of the modules of the intelligent diagnostic auxiliary analysis system based on medical data driven by the present invention;

[0057] Figure 2 for Figure 1 Functional flowchart of the multi-source medical data acquisition module;

[0058] Figure 3 for Figure 1 A functional flowchart of the dynamic feature engineering analysis module. Detailed Implementation

[0059] The technical system of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0060] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor systems and / or microcontroller systems.

[0061] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0062] To achieve the above objectives, please refer to Figures 1 to 3 This invention provides an intelligent diagnostic auxiliary analysis system based on medical data, the system comprising the following modules:

[0063] The multi-source medical data acquisition module is used to acquire medical imaging data, medical laboratory test data, and medical gene response data, and to perform missing difference completion and time series alignment on the medical imaging data, medical laboratory test data, and medical gene response data to obtain a multi-source medical time series driven dataset.

[0064] The dynamic feature engineering analysis module is used to perform dynamic feature engineering analysis on the corresponding medical images, laboratory data and gene response data in the multi-source medical time-series driven dataset, and obtain a complete set of multi-source medical data features composed of medical image features, laboratory indicator fluctuation features and gene pathological response scores.

[0065] The medical feature filtering module is used to filter the feature offset contribution of each medical sub-feature in the full set of multi-source medical data features to obtain the set of important contribution features of multi-source medical data.

[0066] The diagnostic auxiliary analysis and decision-making module is used to construct a corresponding medical-driven diagnostic analysis model based on a hybrid graph neural network and Transformer. It inputs a set of important contribution features from multiple sources of medical data into the medical-driven diagnostic analysis model to perform diagnostic inference and prediction, output the diagnostic probability corresponding to the medical data, and make corresponding medical diagnostic suggestions based on the diagnostic probability decision analysis.

[0067] In the embodiments of this invention, please refer toFigure 1 The diagram shown is a schematic representation of the modules of the intelligent diagnostic auxiliary analysis system based on medical data driven by the present invention. In this example, the intelligent diagnostic auxiliary analysis system based on medical data driven by the present invention includes the following modules:

[0068] S1: Multi-source medical data acquisition module, used to acquire medical imaging data, medical laboratory test data and medical gene response data, and to perform missing difference completion and time series alignment on medical imaging data, medical laboratory test data and medical gene response data to obtain multi-source medical time series driven dataset;

[0069] In this embodiment of the invention, medical image data is acquired from the hospital's Picture Archiving and Communication System (PACS). Taking lung CT images as an example, image data of 200 patients is acquired, each image with a resolution of 512×512 pixels, stored in DICOM format. Medical laboratory test data is acquired from the Laboratory Information System (LIS), covering 20 test items such as blood routine and biochemical indicators for these 200 patients. Medical gene response data is acquired from the gene detection database, containing expression level data of 1000 gene loci for these patients. For the medical image data, pixel blurring is used to remove corresponding blurry noise points in the images. For the medical laboratory test data and gene response data, an iterative completion method based on neighborhood mean is used, with a neighborhood radius of 1. If a patient's white blood cell count data is missing, the mean of the white blood cell counts of the patients before and after the missing patient is calculated (e.g., the white blood cell counts of the patients before and after the missing patient are 6.0×10⁻⁶). 9 / L and 7.0×10 9 / L, then the mean is 6.5×10 9 The data is filled using the / L) method. After filling, the difference in distribution between the completed data and the original data is calculated. If the difference is greater than a preset threshold of 0.05, the process stops; otherwise, the neighborhood radius is increased and the data is filled again. Then, the time series is aligned using the Dynamic Time Warping (DTW) algorithm. Assuming that the medical imaging data has 10 time points, the laboratory examination data has 8 time points, and the gene response data has 12 time points, a time series distance matrix is ​​constructed. The optimal alignment path is found through dynamic programming, and the data is interpolated or sampled to make the time points of the three data points consistent, ultimately resulting in a multi-source medical time series driven dataset.

[0070] S2: Dynamic Feature Engineering Analysis Module, used to perform dynamic feature engineering analysis on the corresponding medical images, laboratory data and gene response data in the multi-source medical time-series driven dataset, to obtain a complete set of multi-source medical data features composed of medical image features, laboratory indicator fluctuation features and gene pathological response scores.

[0071] In this embodiment of the invention, medical images within a multi-source medical time-series driven dataset are converted to grayscale using a weighted average method, with the formula Gray = 0.299R + 0.587G + 0.114B. The color images are then converted to grayscale images, and a gray-level co-occurrence matrix is ​​calculated. The grayscale level is set to 16, the calculation directions are 0°, 45°, 90°, and 135°, and the distance is 1 pixel. A 16×16 matrix is ​​constructed to count the number of pixel pairs, and the contrast ratio is calculated using a formula. Uniformity We obtain medical imaging features, and for laboratory data, taking white blood cell count as an example, we calculate the variation range between adjacent time points. For instance, if a patient's white blood cell count at adjacent time points is 6.0 × 10⁻⁶... 9 / L and 7.0×10 9 / L, with a variation range of 1.0×10 9 / L, and by setting the change range threshold to 1.5×10 9 / L, statistically analyze the frequency of changes, determine the abnormal fluctuation range, calculate the duration, amplitude, and slope of abnormal fluctuations, and obtain the fluctuation characteristics of laboratory indicators; for gene response data, select 50 key gene loci. If the gene locus expression level is within the normal reference range of 90%-110%, it gets 0 points; if it is within the range of 80%-90% or 110%-120%, it gets 1 point; if it is below 80% or above 120%, it gets 2 points. Add up the scores of the 50 loci to obtain the gene pathological response score. Integrate the above characteristics to form a complete set of multi-source medical data characteristics.

[0072] S3: Medical Feature Filtering Module, used to filter the feature offset contribution of each medical sub-feature in the full set of multi-source medical data features to obtain the set of important contribution features of multi-source medical data;

[0073] In this embodiment of the invention, by analyzing the various medical sub-features within the complete set of multi-source medical data features, such as contrast in medical imaging features, white blood cell count variation in laboratory indicator fluctuation features, and gene pathological response scores, the feature dimension distribution is first statistically analyzed. The contrast value range [0, 100] is divided into 10 intervals, and the frequency of feature values ​​in each interval is statistically analyzed. Based on the distribution of each sub-feature dimension, the KL divergence is calculated to measure the distribution difference. Let the contrast feature distribution be P(x), and the white blood cell count variation feature distribution be Q(x), according to the formula... Calculate the divergence between the two features, select sub-features with a KL divergence greater than 0.1, and quantify their contribution using the Shapley value. Assuming there are three sub-features A, B, and C, calculate v(A), v(A,B), etc., by combination, and use the formula... Calculate the Shapley value for each sub-feature, sort them from largest to smallest, retain the top 80% of sub-features, and finally obtain the set of important contribution features of multi-source medical care.

[0074] S4: Diagnostic Assisted Analysis and Decision Module, which is used to construct a corresponding medical-driven diagnostic analysis model based on a hybrid graph neural network and Transformer, and input the multi-source medical important contribution feature set into the medical-driven diagnostic analysis model for diagnostic inference and prediction, so as to output the diagnostic probability corresponding to the medical data, and make corresponding medical diagnostic suggestions based on the diagnostic probability decision analysis.

[0075] In this embodiment of the invention, a medical-driven diagnostic analysis model is constructed based on a hybrid graph neural network (GNN) and a Transformer. The GNN employs a graph attention network (GAT). For the medical knowledge graph, nodes represent diseases, symptoms, etc., and edges represent relationships. The node feature update formula is as follows: Attention coefficient The Transformer employs a multi-head self-attention mechanism: MultiHead(Q,K,V) = Concat(head1,...,head2,...). h W O ,That The model inputs a set of key medical contribution features from multiple sources. Image features are processed by 3D convolution, laboratory features by temporal convolution, and gene features by fully connected layers. The output of each branch is calculated using the formula P = σ(w1 × MLP). image (o image )+w2×MLP lab (o lab )+w3×MLP gene (o gene )), where σ is the Sigmoid function, mapping the output to the interval [0, 1], and w1, w2, and w3 are dynamically adjusted weights, MLP image MLP lab MLP gene It consists of multilayer perceptrons in various branches, acquiring a medical clinical knowledge base. When the diagnostic probability is greater than a threshold (e.g., 0.6), it identifies a possible disease, queries the knowledge base for disease diagnostic criteria and treatment guidelines, and combines this with specific patient data, such as through formulas. Where w i f is the weight of the i-th medical sub-feature. i (x) is the support function of the feature for the treatment plan, which generates medical diagnostic recommendations by ranking the recommendations.

[0076] Furthermore, as an embodiment of the present invention, reference is made to... Figure 2 As shown, Figure 1 A functional flowchart of the multi-source medical data acquisition module is shown in this embodiment. The multi-source medical data acquisition module includes the following functions:

[0077] S11: Acquire medical imaging data;

[0078] In this embodiment of the invention, medical image data is acquired from a hospital's Picture Archiving and Communication System (PACS). Taking a patient's lung CT image as an example, the image data is stored in DICOM format with a resolution of 512×512 pixels. The grayscale value of each pixel ranges from 0 to 4095, and a total of 200 tomographic images are included. Using medical image reading software, the DICOM format image data is parsed into a computer-processable digital matrix form. Each tomographic image corresponds to a two-dimensional matrix, and the element values ​​in the matrix are the grayscale values ​​of the pixels. The acquired medical image data covers the structural information of different layers of the patient's lungs, providing the raw data foundation for subsequent image analysis.

[0079] S12: Obtain medical laboratory test data;

[0080] In this embodiment of the invention, medical laboratory test data is extracted from the hospital's Laboratory Information System (LIS). Taking a set of blood routine and biochemical test data containing 50 patients as an example, each patient's data record includes 30 indicators such as red blood cell count, white blood cell count, hemoglobin concentration, blood glucose, liver function, and kidney function. The data is stored in tabular form, with each row representing a patient's test record and each column corresponding to a test indicator. For example, the white blood cell count of the 10th patient is 6.5 × 10⁻⁶. 9 The blood glucose level was 5.2 mmol / L. The acquired data underwent format conversion, and non-numerical data (such as the test item name) was encoded and converted into a computer-recognizable numerical form for subsequent analysis.

[0081] S13: Obtain medical gene response data;

[0082] In this embodiment of the invention, medical gene response data is obtained from the database of a gene testing laboratory. Assuming the data contains gene expression profile information of 50 patients, the expression levels of 1000 gene loci were tested for each patient. The gene response data is stored in the form of a text file, with each line recording the gene expression data of one patient and each column corresponding to one gene locus. The value represents the expression level of that gene locus. For example, the expression level of gene locus A of the 5th patient is 2.3 and the expression level of gene locus B is 1.8. After obtaining the data, outliers are preliminarily processed. Extreme values ​​that obviously do not conform to biological laws (such as negative expression levels) are set as missing values ​​to prepare for subsequent data completion and analysis.

[0083] S14: Calculate the pixel blur of the medical image data to obtain the pixel blur of the medical image; perform blur denoising processing on the medical image data based on the pixel blur of the medical image to generate medical denoised image data;

[0084] In this embodiment of the invention, pixel blurring is calculated on medical image data. Taking a tomographic image of a lung CT scan as an example, the gradient magnitude method is used to calculate the pixel blurring. For each pixel (x, y) in the image, the Sobel operator is used to calculate its gradient in the x and y directions. The template of the Sobel operator in the x direction is... The template in the y-direction is The gradient G in the x-direction is obtained through convolution. x (x,y) and gradient G in the y direction y (x,y), then according to the formula The gradient magnitude is calculated; the smaller the gradient magnitude, the higher the pixel blur. Then, based on the calculated pixel blur, the medical image data is subjected to blur denoising processing using a Gaussian filtering method. The standard deviation of the Gaussian filter is set to σ = 1.5, and the filter size is 3σ. The Gaussian filtering formula is as follows: Where (x0, y0) is the filter center and (i, j) are the coordinates within the filter. For each pixel, its neighborhood is convolved with the Gaussian filter to obtain the denoised pixel value, thus generating medical denoised image data. For example, for the pixel with coordinates (100, 100) in the original image, after Gaussian filtering, its grayscale value is adjusted from the original 200 to 195.

[0085] S15: Perform missing data and medical gene response data completion analysis on medical laboratory test data and medical gene response data to obtain medical laboratory data completion data and medical gene response data completion data.

[0086] In this embodiment of the invention, by performing missing data and medical gene response data completion analysis, firstly, missing data points are determined through a traversal approach. For medical laboratory test data, if a patient's data for a certain indicator is empty or marked as "NA", then that position is recorded as a missing point. Assume a total of 20 missing points are found. For medical gene response data, 15 missing points are also marked. An iterative completion method based on neighborhood mean is used. The neighborhood radius is initially set to 1. For a missing point in the medical laboratory test data, such as the missing "hemoglobin concentration" indicator for the 15th patient, the mean value of the corresponding data within a neighborhood radius of 1 is calculated and used to fill the missing point. If the hemoglobin concentrations of the two patients before and after this patient are 130 g / L and 140 g / L respectively, then the filling value is (130 + 140) / 2 = 135 g / L. After filling all missing points, the distribution characteristics of the completed data, including the mean, are calculated. variance Then perform a chi-square test on the distribution of the data and the original data distribution, and calculate the chi-square statistic. Among them O j E is the expected frequency of the completed data in the j-th interval. j is the expected frequency of the original data in the j-th interval, and k is the number of intervals. If the chi-square statistic is less than the preset distribution difference threshold of 0.05, such as when the calculated chi-square statistic is 0.03, the neighborhood radius is increased by 1, and the missing points are filled and tested again until the chi-square statistic is greater than the threshold. After multiple iterations, the medical laboratory difference completion data and medical gene response difference completion data are finally obtained.

[0087] S16: The dynamic time warping algorithm is used to perform time-series alignment processing on medical denoised image data, medical laboratory differential completion data, and medical gene response differential completion data to obtain a multi-source medical time-series driven dataset.

[0088] In this embodiment of the invention, the Dynamic Time Warping (DTW) algorithm is used to perform time-series alignment processing on medical denoised image data, medical laboratory differential completion data, and medical gene response differential completion data. Assuming the medical image data contains 200 time points (tomographic images), the medical laboratory examination data has 50 time points (different examination times), and the medical gene response data has 30 time points (gene testing batch times), the DTW algorithm calculates the optimal alignment path between two time series to minimize their distance. The distance metric is defined as Euclidean distance. For the medical denoised image data sequence X = [x1, x2, ..., x...],... 200 ] and medical laboratory difference completion data sequence Y=[y1,y2,…,y 50 Construct a 200×50 distance matrix D, where D(i,j)=||x i -y j || 2A dynamic programming approach is employed, starting from the top left corner of the matrix, to calculate the optimal path using the recursive formula DTW(i,j) = D(i,j) + min{DTW(i-1,j), DTW(i,j-1), DTW(i-1,j-1)}. Here, DTW(i,j) represents the optimal alignment distance between the first i elements of sequence X and the first j elements of sequence Y. After finding the optimal path, interpolation or sampling operations are performed on the data according to the path to align the two sequences. The same operation is performed on the medical gene response data sequences. Finally, the medical denoised image data, medical laboratory differential completion data, and medical gene response differential completion data are aligned in the time dimension, resulting in a multi-source medical time-series driven dataset. For example, the 100th time point of the medical image data is aligned with the 25th time point of the medical laboratory examination data and the 15th time point of the medical gene response data, enabling joint analysis of different types of data on the same time scale and providing a unified data foundation for subsequent intelligent diagnostic auxiliary analysis.

[0089] Furthermore, the missing data and differential expression analysis of medical laboratory test data and medical gene response data includes:

[0090] Data missing points were identified in both medical laboratory test data and medical gene response data to obtain the corresponding data missing points in the laboratory data and gene data.

[0091] In this embodiment of the invention, using medical laboratory test data and medical gene response data as examples, missing data points are determined. The medical laboratory test data includes 100 samples of patients' blood routine and biochemical indicators, with 20 indicators per sample. The medical gene response data records the gene expression levels of 50 samples, with 30 gene loci per sample. The data is examined by traversal. For the medical laboratory test data, if a certain indicator of a sample is marked with the special symbol "NA" or a null value, then that position is determined to be a missing data point. For example, if the "white blood cell count" indicator data of the 10th sample is null, then that position is a missing data point. Similarly, the medical gene response data is examined. If a certain gene locus of a sample is missing, its position is recorded. Finally, it is found that there are 15 missing data points in the medical laboratory test data and 10 missing data points in the medical gene response data, thus clarifying the specific positions that need to be filled in.

[0092] Preferably, by setting a neighborhood range interval corresponding to a neighborhood radius of 1, and filling the corresponding missing data points by calculating the mean of the corresponding data within the neighborhood radius at the corresponding missing data point, the corresponding medical missing data is obtained.

[0093] In this embodiment of the invention, by setting a neighborhood range interval corresponding to a neighborhood radius of 1, taking a missing data point in medical laboratory test data as an example, assuming that the missing point is located at the 8th indicator of the 5th sample, the mean value of the corresponding data within a neighborhood radius of 1 is calculated and used to fill the missing data point. For one-dimensional data, the neighborhood range refers to one data point before and after the missing point (if they exist), namely the 7th and 9th indicator data of the 5th sample. If these two data points are 10 and 12 respectively, the mean value is (10+12) / 2=11. The value of 11 is then filled into the missing point of the 8th indicator of the 5th sample. Similarly, for medical gene response data, the same method is used. With each missing point as the center, the mean value of adjacent data within a neighborhood radius of 1 is calculated and used to fill the missing data. By filling all 15 and 10 missing data points in medical laboratory test data and medical gene response data, the corresponding medical missing data is finally obtained after completion.

[0094] Preferably, the corresponding distribution of completed data is obtained through medical missing data completion, and the difference between the distribution of completed data and the original data distribution before completion is calculated to obtain the data distribution difference value between the completed data and the original data.

[0095] In this embodiment of the invention, the distribution of the corresponding completed data is obtained through medical missing data completion. Taking medical laboratory test data as an example, the mean of the completed data is calculated. Where n is the total number of medical missing data points, and x i To complete the i-th value in the missing medical data, the variance is... The distribution frequency of the data in different intervals was statistically analyzed, and a histogram was plotted to represent the distribution of the completed data. For the original data, its mean, variance, and distribution frequency were also calculated. The difference between the distribution of the completed data and the original data distribution was calculated using a chi-square test. The formula for the chi-square statistic is: Among them O j E is the expected frequency of the completed data in the j-th interval. j It is the expected frequency of the original data in the j-th interval, where k is the number of intervals. The specific process of calculating the expected frequency is based on the mean of the corresponding interval. and variance σ 2 Calculate the corresponding normal distribution probability density function value. Select several points x within interval j. j1 x j2 , ..., x jm The probability P of interval j is approximated by numerically integrating the probability density function values ​​within the interval (e.g., using the trapezoidal integral method). j For example, using the trapezoidal integral method, the interval j = [aj ,b j Divide into m-1 smaller intervals, each with a width of [missing value]. Then the corresponding probability Therefore, the expected frequency is specifically the product of the sample size and the corresponding probability. Combining this with the above formula, we obtain the chi-square statistic, which represents the difference in data distribution between the completed and original data. Similarly, we calculate the difference in data distribution between the completed and original medical gene response data.

[0096] Preferably, the data distribution difference between the completed data and the original data is compared and judged according to a preset distribution difference threshold of 0.05. If the data distribution difference between the completed data and the original data is less than or equal to the preset distribution difference threshold of 0.05, the neighborhood radius corresponding to the neighborhood range interval is increased by 1 and the corresponding missing data points are refilled until the data distribution difference between the completed data and the original data is greater than the preset distribution difference threshold of 0.05, so as to obtain medical laboratory difference completion data and medical gene response difference completion data.

[0097] In this embodiment of the invention, the data distribution difference between the completed data and the original data is compared and judged according to a preset distribution difference threshold of 0.05. Taking medical laboratory test data as an example, if the calculated data distribution difference value is 0.03, which is less than the preset distribution difference threshold of 0.05, the neighborhood radius corresponding to the neighborhood range interval is increased by 1, that is, the neighborhood radius becomes 2, and the missing data points are filled again. For the missing point of the 8th indicator of the 5th sample, the neighborhood range includes the 6th, 7th, 9th and 10th indicator data of the 5th sample. The mean of these 4 data is calculated and filled. After filling, the data distribution difference value between the completed data and the original data is calculated again. Assuming that the difference value calculated this time is 0.06, which is greater than the preset distribution difference threshold of 0.05, the operation is stopped, and medical laboratory difference completed data is obtained. The same operation process is performed on medical gene response data until the data distribution difference value between the completed data and the original data is greater than the preset distribution difference threshold of 0.05, and finally medical gene response difference completed data is obtained.

[0098] Furthermore, as an embodiment of the present invention, reference is made to... Figure 3 As shown, Figure 1 A functional flowchart of the dynamic feature engineering analysis module is shown in this embodiment. The dynamic feature engineering analysis module includes the following functions:

[0099] S21: Perform grayscale processing on the corresponding medical images in the multi-source medical time-series driven dataset to generate medical grayscale image sequences.

[0100] In this embodiment of the invention, taking the lung CT image sequence of 100 patients contained in a multi-source medical time-series driven dataset as an example, each patient's image sequence consists of 200 tomographic images, and the images are color images with a resolution of 512×512 pixels. For each color medical image, a weighted average method is used for grayscale processing, with the formula Gray = 0.299R + 0.587G + 0.114B, where R, G, and B are the pixel values ​​of the red, green, and blue channels in the color image, respectively. For each image, all pixels are examined. For a pixel with coordinates (x, y), if its R value is 200, G value is 150, and B value is 100, then the gray value Gray = 59.8 + 88.05 + 11.4 = 159.25. This value is used as the gray value of the pixel. This operation is performed on all 20,000 color medical images (100 × 200) for 100 patients to generate a corresponding medical grayscale image sequence, converting each image from color to grayscale for subsequent image feature analysis.

[0101] S22: Obtain the corresponding gray-level co-occurrence matrix through the medical gray-level image sequence, and perform image feature statistical analysis on the medical gray-level image sequence based on the gray-level co-occurrence matrix to obtain medical image features, including image texture features corresponding to contrast, uniformity, correlation, energy and entropy.

[0102] In this embodiment of the invention, taking one of 200 medical grayscale images of a patient from a previously generated medical grayscale image sequence as an example, the grayscale level is set to 16 (the grayscale range of 0-255 is divided into 16 intervals). The calculation directions are selected as 0°, 45°, 90°, and 135°, and the distance d is set to 1 pixel to construct a 16×16 grayscale co-occurrence matrix GLCM(i,j), where i and j represent grayscale levels. For each pixel (x, y) in the image, the number of pixel pairs from grayscale level i to grayscale level j is counted at the specified direction and distance, and filled into the GLCM(i,j) matrix. For example, in the 0° direction, starting from a pixel of grayscale level 3, there are 10 pixel pairs reaching grayscale level 5 at a distance of 1 pixel. Then, the value of GLCM(3,5) in the 0° direction is 10. Based on the grayscale co-occurrence matrix, image feature statistical analysis is performed, and the contrast is: Reflects the drasticness of grayscale changes in an image, and its uniformity: It measures the uniformity of grayscale distribution and its correlation. Where μ x μ y These are the mean values ​​of GLCM in the row and column directions, σ x σ yThese are the standard deviations of GLCM in the row and column directions, reflecting the linear correlation of gray-level distribution in the image, while energy: Entropy represents the degree of concentration of elements in the gray-level co-occurrence matrix. Reflecting the randomness of grayscale distribution in an image, the average value of each feature value calculated in four directions is taken to obtain the image texture features corresponding to contrast, uniformity, correlation, energy, and entropy of the image. This operation is performed on all 20,000 images in the medical grayscale image sequence to finally obtain complete medical image features.

[0103] S23: Perform experimental index fluctuation analysis on the corresponding laboratory data within the multi-source medical time-series driven dataset to obtain the fluctuation characteristics of laboratory indicators;

[0104] In this embodiment of the invention, taking the laboratory data such as blood routine and biochemical indicators of 100 patients in a multi-source medical time-series driven dataset as an example, each patient contains 30 laboratory indicators, and data are recorded at 5 different time points. For each laboratory indicator, such as "white blood cell count", the absolute difference between the indicator values ​​at adjacent time points is first calculated to obtain the amplitude of change; a threshold for the amplitude of change is set, and the frequency of change is statistically analyzed; based on the amplitude and frequency of change, the abnormal fluctuation range is determined, the start and end points of the abnormal fluctuation time are obtained, and the duration of the abnormal fluctuation is calculated; the amplitude of the indicator change within the abnormal fluctuation range is calculated, and thus the slope of the laboratory indicator fluctuation is obtained. For example, for the "white blood cell count" indicator, the 10th patient has an abnormal fluctuation range at time points 2-4, and the indicator value at time point 2 is 6.0 × 10⁻⁶. 9 / L, the index value at time point 4 is 8.0×10 9 / L, the abnormal fluctuation duration is 4-2+1=3 time units, and the index change amplitude is 8.0×10 9 / L-6.0×10 9 / L=2.0×10 9 If / L, then the slope of the laboratory index fluctuation is 2.0×10 9 / L / 3≈0.67×10 9 / L, the above operation was performed on 30 laboratory indicators of all patients, and the change range, change frequency and fluctuation slope of each indicator were finally obtained. These three data were combined into a feature vector, and finally the complete fluctuation characteristics of the laboratory indicators were obtained.

[0105] S24: Evaluate the pathological response of the corresponding gene response data in the multi-source medical time-series driven dataset to obtain a gene pathological response score;

[0106] In this embodiment of the invention, using gene expression profile data from 50 patients within a multi-source medical time-series driven dataset as an example, the expression levels of 1000 gene loci were detected for each patient. A gene pathological response assessment model was established, selecting 100 key gene loci related to the disease. Scoring was based on the degree of difference between the expression level of each gene locus and the normal reference value. The scoring rules were as follows: if the expression level of a gene locus is within 90%-110% of the normal reference value, 0 points are awarded; if it is within 80%-90% or 110%-120%, 1 point is awarded; if it is below 80% or above 120%, a score of 0 is awarded. For example, for gene locus A in patient number 5, the normal reference value is 2.0, and its expression level is 2.2, which is in the range of 110%-120%, so this locus gets 1 point; gene locus B has a normal reference value of 1.5, and its expression level is 1.3, which is in the range of 80%-90%, so this locus also gets 1 point. The scores of the patient's 100 key gene loci are added together to obtain the gene pathology response score. If the total score of the 100 loci is 30 points, then the patient's gene pathology response score is 30 points. This operation is performed on all 50 patients to obtain their respective gene pathology response scores.

[0107] S25: Combine the corresponding medical imaging features, laboratory indicator fluctuation features, and gene pathological response scores into one dataset to obtain the complete set of corresponding multi-source medical data features.

[0108] In this embodiment of the invention, the previously obtained medical imaging features, laboratory indicator fluctuation features, and gene pathology response scores are merged. Taking the first patient as an example, their medical imaging features include five texture feature values ​​such as contrast and uniformity, assumed to be [0.8, 0.7, 0.6, 0.5, 0.4]; the laboratory indicator fluctuation features include fluctuation feature vectors of 30 indicators, assumed to be [1.2, 0.8, 0.6] for the "white blood cell count" indicator; and the gene pathology response score is 25 points. These data are integrated into a dataset to form the complete set of multi-source medical data features for this patient, such as [[0.8, 0.7, 0.6, 0.5, 0.4], [[1.2, 0.8, 0.6], [...], 25]. This merging operation is performed on 100 patients to obtain a complete set of multi-source medical data features. This dataset integrates key features from multiple sources such as medical imaging, laboratory tests, and gene responses, providing comprehensive data support for intelligent diagnostic auxiliary analysis based on medical data.

[0109] Furthermore, the analysis of experimental index fluctuations in the corresponding laboratory data within the multi-source medical time-series driven dataset includes:

[0110] We analyzed the magnitude of changes in experimental indicators in the corresponding laboratory data within the multi-source medical time-series driven dataset to obtain the magnitude of changes in laboratory indicators.

[0111] In this embodiment of the invention, taking the blood routine and biochemical index test data of 100 patients contained in a multi-source medical time-series driven dataset as an example, the variation range of laboratory indicators is analyzed. Assuming the dataset records 30 laboratory indicator data for each patient at 5 different time points, for each laboratory indicator, such as "white blood cell count," the absolute difference between indicator values ​​at adjacent time points is calculated on a patient-by-patient basis. For a certain patient, the white blood cell count at time point 1 is 6.0 × 10⁻⁶. 9 / L, the white blood cell count at time point 2 was 7.0 × 10⁹ / L. 9 If the white blood cell count is / L, then the change in white blood cell count between these two time points is |7.0×10 9 -6.0×10 9 |=1.0×10 9 The process iterates through all patients for every laboratory indicator at each time point, calculating the magnitude of change between adjacent time points. Then, for each laboratory indicator, the average magnitude of change across all patients is calculated using the following formula: Where n is the number of patients, t is the number of time points, and x is the number of time points. ij This represents the indicator value of the i-th patient at time j. For example, the mean change in the "white blood cell count" indicator for all patients is calculated to be 0.8 × 10⁻⁶. 9 / L, and so on, to obtain the average change of each laboratory indicator, that is, the change of the laboratory indicator, which provides basic data for subsequent analysis.

[0112] Preferably, the laboratory index change frequency is statistically analyzed for the corresponding laboratory data in the multi-source medical time-series driven dataset based on the change magnitude of laboratory indexes, so as to obtain the change frequency of laboratory indexes.

[0113] In this embodiment of the invention, based on the previously obtained changes in laboratory indicators, the frequency of changes in corresponding laboratory indicators within the multi-source medical time-series driven dataset is statistically analyzed. Taking the "white blood cell count" indicator as an example, a threshold for the change range is set, assuming it is 1.5 × 10⁻⁶. 9 / L, for each patient's time series data of the "white blood cell count" indicator, count the number of times the change between adjacent time points exceeds this threshold. For example, if a patient's white blood cell count changes by 0.5 × 10 at 5 time points... 9 / L, 2.0×10 9 / L, 0.8×10 9 / L, 1.6×10 9 / L and 0.3×109 / L, where the variation is greater than 1.5×10 9 The / L value was measured twice. For each laboratory indicator across all patients, the number of changes was recorded. Then, the average number of changes for each laboratory indicator across all patients was calculated using the following formula: Where n is the number of patients, c i This represents the number of times a certain indicator changes for the i-th patient. For example, the frequency of change for the "white blood cell count" indicator is calculated to be 1.2 times per patient. The frequency of change for each laboratory indicator is obtained in this way, which reflects the degree of frequency of indicator changes.

[0114] Preferably, based on the magnitude and frequency of changes in laboratory indicators, abnormal fluctuation range analysis is performed on each corresponding laboratory indicator within the laboratory data to obtain the abnormal fluctuation range corresponding to each laboratory indicator.

[0115] In this embodiment of the invention, abnormal fluctuation ranges are analyzed for each laboratory indicator within the laboratory data based on the magnitude and frequency of changes in laboratory indicators. Taking "blood glucose" as an example, the normal fluctuation range is first determined. Assuming that, based on a large amount of historical data, the average normal fluctuation range of "blood glucose" is 0.5 mmol / L and the normal frequency is 0.8 times per patient, for each patient's time series data of "blood glucose," if the average fluctuation range of a certain consecutive time point is greater than 1.5 times the average normal fluctuation range, and the frequency of change is greater than 1.5 times the normal frequency of change, then that time period is determined to be an abnormal fluctuation range. For example, a patient's "blood glucose" at time points 3-5... "The average change in the indicator was 0.9 mmol / L, with a frequency of 1.5 changes, both exceeding the corresponding threshold. Therefore, time points 3-5 represent the abnormal fluctuation range of the patient's 'blood glucose' indicator. By iterating through each laboratory indicator for all patients, the abnormal fluctuation range for each indicator was determined. Then, for each laboratory indicator, the abnormal fluctuation ranges of all patients were integrated, and duplicates were removed to obtain the corresponding abnormal fluctuation segment. For example, after integrating the abnormal fluctuation ranges of the 'blood glucose' indicator for all patients, the overall abnormal fluctuation segments were obtained as time points 3-5 and 8-10. In this way, the abnormal fluctuation segments corresponding to each laboratory indicator were obtained, clarifying the specific time period of abnormal fluctuation of the indicator."

[0116] Preferably, the slope of the fluctuation of the laboratory indicators is calculated based on the abnormal fluctuation range corresponding to each laboratory indicator.

[0117] In this embodiment of the invention, the slope of index fluctuation is calculated based on the abnormal fluctuation range corresponding to each laboratory indicator. Taking a certain abnormal fluctuation range of the "hemoglobin concentration" indicator (assumed to be time points 2-4) as an example, for the data of each patient within this abnormal fluctuation range, the slope of fluctuation is calculated using a linear regression method. Let the time point be x and the indicator value be y. According to the principle of least squares, the formula for calculating the slope k is as follows: Where n is the number of time points within the abnormal fluctuation range (here n=3). For example, a patient's "hemoglobin concentration" values ​​at time points 2, 3, and 4 are 120g / L, 125g / L, and 130g / L, respectively, and the corresponding time points x are 2, 3, and 4. Substituting these values ​​into the formula, we get k=5. By iterating through the data of all patients within the abnormal fluctuation range of each laboratory indicator, we can calculate the fluctuation slope of each indicator. Then, we can calculate the average fluctuation slope of each laboratory indicator among all patients, and finally obtain the fluctuation slope of the laboratory indicator. This method quantifies the trend of the indicator within the abnormal fluctuation range.

[0118] Preferably, the amplitude of change of laboratory indicators, the frequency of change of laboratory indicators, and the slope of fluctuation of laboratory indicators are combined as fluctuation characteristics to obtain the fluctuation characteristics of laboratory indicators.

[0119] In this embodiment of the invention, the fluctuation characteristics are obtained by combining the magnitude of laboratory indicator changes, the frequency of laboratory indicator changes, and the slope of laboratory indicator fluctuations. Taking the "liver function" indicator as an example, its magnitude of laboratory indicator change is 1.2 (assuming it is the average magnitude of the change of a specific liver function indicator, with the unit depending on the indicator), the frequency of laboratory indicator change is 1.0 times / patient, and the slope of laboratory indicator fluctuation is 2.5 (assuming the unit is the unit of the change rate of the corresponding indicator). These three data are combined into a feature vector [1.2, 1.0, 2.5]. This vector is the fluctuation characteristic of the "liver function" indicator. For each laboratory indicator in the multi-source medical time-series driven dataset, this combination is performed to obtain the corresponding fluctuation characteristics of the laboratory indicator. These fluctuation characteristics comprehensively reflect the information of the laboratory indicator in terms of magnitude, frequency, and trend of change, providing key feature data for intelligent diagnostic auxiliary analysis based on medical data, which helps doctors or intelligent diagnostic systems to judge the patient's health status and disease development trend.

[0120] Furthermore, the calculation of the slope of index fluctuation based on the abnormal fluctuation range corresponding to each laboratory index includes:

[0121] By identifying the abnormal fluctuation ranges corresponding to the changes in various laboratory indicators, we can obtain the corresponding time start and time end points of abnormal fluctuations.

[0122] In this embodiment of the invention, taking the "blood glucose" indicator as an example, in a multi-source medical time-series driven dataset, the abnormal fluctuation segments corresponding to the "blood glucose" indicator are previously determined to be time points 3-5 and 8-10. For each abnormal fluctuation segment, its start and end time points are directly extracted as the abnormal fluctuation time start and end points. For example, in the time point 3-5 segment, the abnormal fluctuation time start point is time point 3, and the abnormal fluctuation time end point is time point 5; in the time point 8-10 segment, the abnormal fluctuation time start point is time point 8, and the abnormal fluctuation time end point is time point 9. Point 10: For each laboratory indicator of all patients, the corresponding abnormal fluctuation range is traversed in this manner, and the abnormal fluctuation time start and abnormal fluctuation time end of each range are extracted. Assuming there are 100 patients, for the "hemoglobin concentration" indicator, the abnormal fluctuation range of the 10th patient is time point 6-8. Then the abnormal fluctuation time start of the "hemoglobin concentration" indicator of this patient is time point 6, and the abnormal fluctuation time end is time point 8. And so on, to complete the extraction of the abnormal fluctuation time start and end points of each laboratory indicator of all patients, providing basic data for subsequent calculations.

[0123] Preferably, the duration of the abnormal fluctuation is calculated based on the start and end points of the abnormal fluctuation time.

[0124] In this embodiment of the invention, the duration of abnormal fluctuation is calculated based on the previously obtained start and end points of the abnormal fluctuation time. Taking the "white blood cell count" indicator in a "complete blood count" as an example, for an abnormal fluctuation segment of a patient, assuming the start point of the abnormal fluctuation time is time point 2 and the end point of the abnormal fluctuation time is time point 4, since the time points are arranged sequentially and adjacent time points are spaced the same, the duration of the abnormal fluctuation is calculated as follows: Abnormal fluctuation duration = Abnormal fluctuation end point - Abnormal fluctuation start point + 1 (considering the inclusion of the start and end time points), i.e., 4 - 2 + 1 = 3. For each laboratory indicator of all patients, the duration of abnormal fluctuation is calculated according to the above formula for each abnormal fluctuation segment. For example, for the "liver function" indicator, the 20th patient has two abnormal fluctuation segments, namely time points 1-3 and 5-7. The duration of abnormal fluctuation in the first segment is 3-1+1=3, and the duration of abnormal fluctuation in the second segment is 7-5+1=3. Through this calculation, the duration of abnormal fluctuation of each laboratory indicator of all patients under different abnormal fluctuation segments is obtained, which prepares for further calculation of the indicator fluctuation slope.

[0125] Preferably, the amplitude of the change of each laboratory indicator is obtained through the abnormal fluctuation range corresponding to each laboratory indicator, and the slope of the index fluctuation is calculated based on the duration of the abnormal fluctuation to obtain the slope of the laboratory indicator fluctuation.

[0126] In this embodiment of the invention, the amplitude of change of the corresponding laboratory indicator is obtained through the abnormal fluctuation segment corresponding to each laboratory indicator. Taking the "kidney function" indicator as an example, for a patient's abnormal fluctuation segment at time point 2-4, assuming that the patient's "kidney function" indicator value at time point 2 is 80 (the unit is set according to the actual indicator) and the indicator value at time point 4 is 100, then the amplitude of change of the laboratory indicator = the indicator value corresponding to the end of the abnormal fluctuation time - the indicator value corresponding to the start of the abnormal fluctuation time, that is, 100-80=20. Based on the duration of the abnormal fluctuation, the slope of the laboratory indicator fluctuation amplitude is calculated. The calculation formula is: laboratory indicator fluctuation slope = laboratory indicator fluctuation amplitude / duration of abnormal fluctuation. For the above example of the "kidney function" indicator, the duration of abnormal fluctuation is 4-2+1=3, then the slope of the laboratory indicator fluctuation = 20 / 3≈6.67. For each laboratory indicator of all patients, each abnormal fluctuation segment is traversed, and the slope of the laboratory indicator fluctuation is calculated in the above manner. For example, regarding the "blood lipid" indicator, the 30th patient exhibited an abnormal fluctuation range between time points 3 and 6. The indicator value at time point 3 was 5.0, and the indicator value at time point 6 was 7.5. The duration of the abnormal fluctuation was 6-3+1=4, and the amplitude of the laboratory indicator change was 7.5-5.0=2.5. Therefore, the slope of the laboratory indicator fluctuation for this patient's "blood lipid" indicator in this range was 2.5 / 4=0.625. By calculating the various indicators for all patients, the complete laboratory indicator fluctuation slope can be obtained. These slopes can intuitively reflect the changing trend of laboratory indicators within the abnormal fluctuation range, providing an important basis for intelligent diagnostic auxiliary analysis based on medical data.

[0127] Furthermore, the pathological response assessment of the corresponding gene response data within the multi-source medical time-series driven dataset includes:

[0128] Gene response feature matrix analysis was performed on the gene response data in the multi-source medical time-series driven dataset to obtain the medical gene response feature matrix, which includes the change rate of gene response expression, the corresponding clinical treatment response time point, and the influencing factors of gene response mutation.

[0129] In this embodiment of the invention, using gene expression profile data from 50 patients within a multi-source medical time-series driven dataset as an example, the expression levels of 1000 gene loci were detected for each patient, and data from 5 different time points were recorded. For each gene locus, the rate of change in expression level between adjacent time points was calculated, using the following formula: Where E tThis represents the gene expression level at time point t. For example, if a patient's gene locus A has an expression level of 2.0 at time point 2 and 2.4 at time point 3, then the rate of change of this gene locus from time point 2 to 3 is 2.4 - 2.0 / 2.0 × 100% = 20%. The clinical treatment response time point is determined by comparing the change in gene expression level with the implementation time of the clinical treatment. When the gene expression level shows a significant change after the implementation of the treatment (the rate of change exceeds a preset threshold of 15%), this time point is recorded as the clinical treatment response time point. For example, if a patient received specific treatment at time point 3, and the rate of change of gene locus B from time point 3 to 4 is 25%, then time point 4 is recorded as the clinical treatment response time point. The clinical treatment response time points of this gene locus were analyzed, and the influencing factors corresponding to gene response mutations were analyzed. An evaluation system containing 10 influencing factors was established, such as the mutation type, mutation location, and mutation frequency of the gene locus. The mutation of each gene locus was quantified according to a preset scoring standard. For example, a missense mutation was assigned 2 points, and a nonsense mutation was assigned 3 points. The change rate of each gene locus at each time point, the clinical treatment response time point, and the mutation influencing factor scores were integrated to construct a 1000×5×12 medical gene response feature matrix (1000 gene loci, 5 time points, 12 columns including change rate, time point, and 10 influencing factor scores), and the gene response feature matrix analysis was completed.

[0130] Preferably, the disease pathological response stage is obtained, and the pathological response path of the disease pathological response stage is simulated based on the medical gene response feature matrix and combined with the Monte Carlo simulation method to generate a personalized gene pathological process response path.

[0131] In this embodiment of the invention, by acquiring the pathological response stage of the disease, and according to clinical diagnostic criteria, the disease development is divided into three stages: early, middle, and late. For each patient's medical gene response feature matrix, a Monte Carlo simulation method is used to simulate the pathological response path. The simulation is set to 1000 times, and sample paths conforming to the probability distribution of gene expression changes are randomly generated. For a given patient, starting from the early stage of the disease, based on the rate of change and influencing factor scores in their gene response feature matrix, the possible changes in the expression level of each gene locus in the next stage are calculated. For example, if the rate of change of gene locus C from the early to the middle stage is 18%, and the influencing factor score is 7, its expression level in the middle stage is predicted using a preset regression model. The same operation is performed on all gene loci to obtain... In the intermediate stage of the gene expression profile, changes in the gene expression profile are analyzed in each simulation, combined with the clinical treatment response time point, to determine whether the disease has progressed to the next stage. For example, when the expression level of more than 50% of key genes changes by more than 20%, the disease is considered to have progressed to the intermediate stage. This process is repeated to simulate the disease's progression from the early stage to the intermediate stage and then to the late stage. After 1000 simulations, the frequency of each path is counted, and the path with the highest frequency is selected as the patient's personalized gene pathological process response path. For example, if the simulation results for a patient show that the path "early stage → intermediate stage → late stage" appears 600 times, this path is selected as the patient's personalized gene pathological process response path. Finally, a complete path is generated that includes changes in gene expression, treatment response time points, and disease stage transitions.

[0132] Preferably, a weighted average method is used based on the personalized gene pathology process response path, and the actual medical clinical response of the patient is combined to perform a comprehensive evaluation calculation of the gene pathology response feature matrix to obtain a gene pathology response score.

[0133] In this embodiment of the invention, a comprehensive evaluation of the gene pathology response is performed on the medical gene response feature matrix based on a personalized gene pathology process response path, using a weighted average method combined with the patient's actual medical clinical response. The weights of each evaluation indicator are determined: the rate of change is weighted at 0.4, the clinical treatment response time point at 0.3, and the mutation influencing factor score at 0.3. For a patient's gene locus D, its rate of change score in the personalized gene pathology process response path is 8 points (mapped to a 0-10 score range based on the magnitude of the rate of change), and the clinical treatment response time score is 7 points (based on the early response time). The late-stage and effectiveness score), with a mutation influencing factor score of 6, results in a comprehensive score of (8×0.4+7×0.3+6×0.3=3.2+2.1+1.8=7.1 points). This calculation is performed on all gene loci to obtain the comprehensive score for each locus. The comprehensive scores of all gene loci are then weighted and summed to obtain the patient's gene pathology response score. Assuming a total of 1000 gene loci are involved in the assessment, the weight of each gene locus is determined based on its correlation with the disease; for example, a key gene locus has a weight of 0.002, and a non-key gene locus has a weight of 0.0008. Finally... Where S i The comprehensive score of the i-th gene locus, W i These are the corresponding weights. For example, after calculation, a patient's gene pathology response score is 75 points. This completes the comprehensive assessment calculation of gene pathology response. The above embodiment details the calculation process of gene pathology response score.

[0134] Furthermore, the medical feature screening module includes the following functions:

[0135] Statistical analysis of the feature dimension distribution of each medical sub-feature within the full set of multi-source medical data features is performed to obtain the data feature dimension distribution corresponding to each medical sub-feature.

[0136] In this embodiment of the invention, taking a complete set of multi-source medical data features as an example, this complete set includes medical imaging features, laboratory indicator fluctuation features, and medical sub-features such as gene pathology response scores. For the contrast feature in medical imaging features, assuming its value range is [0, 100], it is divided into 10 intervals, namely [0, 10), [10, 20), ..., [90, 100]. The frequency of the contrast feature value in each interval is counted. For example, it appears 20 times in the interval [10, 20), 30 times in the interval [20, 30), etc., thus obtaining the data feature dimension distribution of the contrast feature. The white blood cell count change amplitude feature in the laboratory indicator fluctuation features is processed similarly, assuming its value range is [-5×10]. 9 / L, 5×10 9 / L] is divided into 10 intervals, such as [-5×10 9 / L, -4×10 9 / L), [-4×10 9 / L, -3×10 9 By analyzing the frequency of occurrence of feature values ​​within each interval (e.g., / L), the data feature dimension distribution of the white blood cell count variation amplitude is obtained. Furthermore, by dividing the gene pathological response score into 5 intervals (e.g., [0, 20), [20, 40)) based on its value range of [0, 100], the frequency of occurrence of the score within each interval is analyzed to obtain the data feature dimension distribution of the gene pathological response score. By performing this processing on each medical sub-feature within the complete set of multi-source medical data features, the data feature dimension distribution corresponding to each medical sub-feature is finally obtained.

[0137] Preferably, KL divergence is calculated for the corresponding medical sub-features within the full set of multi-source medical data features based on the data feature dimension distribution corresponding to each medical sub-feature, to obtain the feature distribution offset KL divergence corresponding to each medical sub-feature.

[0138] In this embodiment of the invention, based on the previously obtained data feature dimension distributions corresponding to each medical sub-feature, taking the contrast feature in medical image features and the white blood cell count variation amplitude feature in laboratory indicator fluctuation features as examples, the KL divergence between them is calculated. Let the data feature dimension distribution of the contrast feature be P(x), and the data feature dimension distribution of the white blood cell count variation amplitude feature be Q(x). The KL divergence formula is: Suppose that P(x) has a probability of 0.1 in the interval [0, 10) and a probability of 0.2 in [10, 20), etc.; and Q(x) has a probability of 0.15 in the interval [0, 10) and a probability of 0.18 in [10, 20), etc., then By calculating the feature distribution offset KL divergence between the contrast feature and the white blood cell count variation feature, this KL divergence calculation is performed on every two medical sub-features in the full set of multi-source medical data features, thereby obtaining the feature distribution offset KL divergence corresponding to each medical sub-feature. For example, the KL divergence between the contrast feature and the gene pathology response score, as well as the KL divergence between the white blood cell count variation feature and the gene pathology response score, are then calculated to comprehensively assess the distribution differences between each medical sub-feature.

[0139] Preferably, based on the feature distribution offset KL divergence corresponding to each medical sub-feature, medical sub-features with a value greater than 0.1 are selected, and the corresponding feature contribution selection process is automatically triggered to quantify the contribution of each medical sub-feature using Shapley value. At the same time, the medical sub-features are sorted from largest to smallest and the top 80% are retained to obtain a set of important contribution features from multiple sources of medical care.

[0140] In this embodiment of the invention, based on the previously obtained feature distribution offset KL divergence corresponding to each medical sub-feature, medical sub-features with a value greater than 0.1 are selected. Assuming that contrast and uniformity in medical imaging features, white blood cell count variation amplitude and blood glucose variation frequency in laboratory indicator fluctuation features, and gene pathological response scores are selected, the corresponding feature contribution screening process is automatically triggered. The Shapley value is used to quantify the contribution of each selected medical sub-feature. For example, assuming there are three medical sub-features A, B, and C, their combinations are {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, and {A, B, C}. For the combination {A}, its contribution value v(A) to the final diagnostic result is calculated; for the combination {A, B}, its contribution value v(A, B) is calculated. The formula for calculating the Shapley value is... Where N is the set of all medical sub-features, i is a specific medical sub-feature, and S is a subset of N. The Shapley value of each medical sub-feature is calculated, such as 0.3 for medical sub-feature A, 0.25 for B, 0.2 for C, etc. They are sorted from largest to smallest, assuming the sorted values ​​are A, B, C, D, and E. The top 80% of the medical sub-features are retained, assuming the top 80% includes A, B, C, and D, resulting in the multi-source medical important contribution feature set {A, B, C, D}. These features have high importance and contribution in medical data-driven intelligent diagnostic auxiliary analysis.

[0141] Furthermore, the diagnostic auxiliary analysis and decision-making module includes the following functions:

[0142] A medical-driven diagnostic analysis model is constructed based on a hybrid graph neural network and Transformer.

[0143] In this embodiment of the invention, a medical-driven diagnostic analysis model is constructed based on a hybrid graph neural network (GNN) and a transformer. The model comprises three main components: a GNN encoder, a transformer encoder, and a decoder. The GNN encoder employs a graph attention network (GAT) architecture to process relational structures in multi-source medical data. For a medical knowledge graph, nodes represent entities such as diseases, symptoms, and examination items, and edges represent relationships between them (such as causal relationships and accompaniment relationships). The feature vector of each node is updated through a GAT layer, using the following formula: in, W is the feature vector of node i in the l-th layer. (l) It is a learnable weight matrix, α ijThe attention coefficient is calculated using the following formula: Where 'a' is a learnable attention vector, || denotes the concatenation operation, and the Transformer encoder is used to process sequence features, employing a standard multi-head self-attention mechanism: MultiHead(Q,K,V) = Concat(head1,...,head2,...). h W O ,That and W O It is a learnable weight matrix. The decoder uses a multilayer perceptron (MLP) structure, fusing the outputs of the GNN encoder and the Transformer encoder to generate diagnostic predictions: Prediction = MLP(Concat(h GNN h Transformer h GNN and h Transformer These are the output feature vectors of the GNN encoder and the Transformer encoder, respectively, which are used to construct the corresponding medical-driven diagnostic analysis model.

[0144] Preferably, the medical sub-features within the multi-source medical important contribution feature set are divided into three branches, including an imaging branch, an experimental indicator branch, and a gene branch. The medical sub-features corresponding to the three branches are then input into a medical-driven diagnostic analysis model for diagnostic reasoning and prediction, so as to output the diagnostic probability corresponding to the medical data.

[0145] In this embodiment of the invention, the various medical sub-features within the multi-source medical important contribution feature set are divided into three branches: an imaging branch, an experimental index branch, and a gene branch. The imaging branch includes texture features such as contrast and uniformity in medical imaging features; the experimental index branch includes fluctuation features such as the amplitude of white blood cell count changes and the frequency of blood glucose changes in laboratory index fluctuation features; and the gene branch includes gene-related features such as gene pathology response scores. The medical sub-features corresponding to the three branches are respectively input into different modules of the medical-driven diagnostic analysis model. The features of the imaging branch are processed by a 3D convolutional neural network (CNN) to extract spatial features; the features of the experimental index branch are processed by a temporal convolutional network (TCN) to capture temporal changes; and the features of the gene branch are encoded by a fully connected layer. The outputs of each branch are o. image o lab and o gene The final diagnosis probability is calculated using the following formula: P = σ(w1 × MLP) image (o image )+w2×MLP lab (o lab )+w3×MLP gene (o gene)), where σ is the Sigmoid function, mapping the output to the interval [0, 1], and w1, w2, and w3 are dynamically adjusted weights, MLP image MLP lab MLP gene It is a multilayer perceptron with branches. For example, when a patient's multi-source medical data is input, the output of the imaging branch is 0.75, the output of the experimental index branch is 0.82, and the output of the gene branch is 0.69. The dynamically adjusted weights are 0.35, 0.40, and 0.25, respectively. Then the diagnosis probability is P = 0.683, and the final output is the diagnosis probability corresponding to this set of medical data.

[0146] Preferably, a medical clinical knowledge base is acquired, and a corresponding medical diagnosis recommendation is generated based on the diagnosis probability and the decision analysis of the medical clinical knowledge base.

[0147] In this embodiment of the invention, a medical clinical knowledge base is acquired. This knowledge base contains structured knowledge such as disease diagnostic criteria, treatment guidelines, and drug information. Based on the diagnostic probability and the decision analysis of the medical clinical knowledge base, corresponding medical diagnostic recommendations are generated. First, a list of possible diseases is determined based on the diagnostic probability, with a threshold of 0.6. When the diagnostic probability is greater than 0.6, the disease is considered to be possible. For example, if the diagnostic probability is 0.683, the corresponding disease is diabetes, so diabetes is added to the list of possible diseases. Then, the diagnostic criteria and treatment guidelines for the disease are queried from the medical clinical knowledge base. For example, for diabetes, the knowledge base records that the diagnostic criteria are fasting blood glucose ≥7.0 mmol / L or random blood glucose ≥11.1 mmol / L, and the treatment guidelines include diet control, exercise therapy, oral hypoglycemic agents, and insulin therapy. Next, combined with the patient's specific medical data, such as the frequency of blood glucose changes in laboratory indicators and gene pathology response scores, a personalized diagnostic recommendation is generated. If the patient's blood glucose change frequency is high and the gene pathology response score shows sensitivity to insulin, insulin therapy is recommended as the first choice. Finally, the recommendation degree of each treatment plan is calculated using the following formula: Where w i f is the weight of the i-th medical sub-feature. i (x) is the support function of this feature for the treatment plan. For example, for an insulin treatment plan, the weight of the frequency of blood glucose changes is 0.4, and the support function is f. glucose (x) = 0.8; the weight of the gene pathological response score is 0.6, and the support function is f. geneIf (x) = 0.9, then the recommendation rate of the insulin treatment regimen is: Recommendation = 0.4 × 0.8 + 0.6 × 0.9 = 0.86. Based on the recommendation rate, the regimens are sorted from high to low to generate the final medical diagnosis recommendation plan, including the preferred treatment plan and its detailed description.

[0148] Furthermore, the step of inputting the medical sub-features corresponding to the three branches into the medical-driven diagnostic analysis model for diagnostic inference and prediction includes:

[0149] The learning rate of the medical-driven diagnostic analysis model was set to 0.5.

[0150] In this embodiment of the invention, a medical-driven diagnostic analysis model is constructed. This model contains three branches, corresponding to medical imaging features, laboratory indicator fluctuation features, and gene pathological response scores, respectively. The learning rate is set to 0.5 to control the step size of the model parameter updates, thereby ensuring the stability and convergence of the model.

[0151] Preferably, the medical sub-features corresponding to the three branches are input into the medical-driven diagnostic analysis model, and the network AUC improvement rate corresponding to each epoch update time is calculated on the corresponding branches in the medical-driven diagnostic analysis model.

[0152] In this embodiment of the invention, the medical sub-features corresponding to the three branches are input into the medical-driven diagnostic analysis model. For the medical image feature branch, the input includes texture features such as contrast and uniformity; for the laboratory indicator fluctuation feature branch, the input includes fluctuation features such as the amplitude of white blood cell count changes and the frequency of blood glucose changes; and for the gene pathology response score branch, the input is the gene pathology response score. At each epoch update time of the model training, the corresponding network AUC improvement rate is calculated. Assuming the AUC value of the t-th epoch is AUC... t The AUC value of the (t-1)th epoch is AUC. t-1 The formula for calculating the AUC improvement rate is: For example, if the AUC value of the 5th epoch is 0.85 and the AUC value of the 4th epoch is 0.82, then the AUC improvement rate of the 5th epoch is 0.85 - 0.82 / 0.82 × 100 = 3.66%. Calculate the AUC improvement rate for each branch separately. For example, the AUC improvement rate of the medical imaging features branch is 4.2%, the AUC improvement rate of the laboratory indicator fluctuation features branch is 2.8%, and the AUC improvement rate of the gene pathology response score branch is 3.1%.

[0153] Preferably, the weights on each branch are dynamically adjusted based on the learning rate and the network AUC improvement rate corresponding to each epoch update time. The corresponding diagnostic probability is calculated by weighted inference prediction based on the weights and the corresponding medical sub-features on each branch, so as to output the diagnostic probability corresponding to the medical data.

[0154] In this embodiment of the invention, the weights of each branch are dynamically adjusted based on the learning rate and the network AUC improvement rate corresponding to each epoch update time. Let the initial weights of the three branches be w1 = 0.3, w2 = 0.4, and w3 = 0.3, respectively. In the t-th epoch, the weights are adjusted according to the AUC improvement rates r1, r2, and r3 of each branch. The calculation formula is as follows: Where α = 0.5 is the adjustment coefficient. For example, in a certain epoch, the AUC improvement rate r1 of the medical imaging feature branch is 4.2%, r2 of the laboratory indicator fluctuation feature branch is 2.8%, and r3 of the gene pathology response score branch is 3.1%. The new weights of the medical imaging feature branches are as follows: Similarly, the new weights for the laboratory indicator fluctuation feature branch are approximately 0.39, and the new weights for the gene pathology response score branch are approximately 0.30. Based on the adjusted weights and the weighted inference prediction of the corresponding medical sub-features on each branch, the corresponding diagnostic probabilities are calculated. Let the outputs of the three branches be o1, o2, and o3, respectively, then the diagnostic probabilities are... For example, if the outputs of the three branches are 0.7, 0.8, and 0.6 respectively, and the adjusted weights are 0.31, 0.39, and 0.30 respectively, then the diagnosis probability P = 0.709, and the diagnosis probability corresponding to the medical data is output in this way.

[0155] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A medical data-driven intelligent diagnostic auxiliary analysis system, characterized in that, Includes the following modules: The multi-source medical data acquisition module is used to acquire medical imaging data, medical laboratory test data, and medical gene response data, and to perform missing difference completion and time series alignment on the medical imaging data, medical laboratory test data, and medical gene response data to obtain a multi-source medical time series driven dataset. The dynamic feature engineering analysis module is used to perform dynamic feature engineering analysis on the corresponding medical images, laboratory data and gene response data in the multi-source medical time-series driven dataset, and obtain a complete set of multi-source medical data features composed of medical image features, laboratory indicator fluctuation features and gene pathological response scores. The medical feature filtering module is used to filter the feature offset contribution of each medical sub-feature in the full set of multi-source medical data features to obtain the set of important contribution features of multi-source medical data. The diagnostic auxiliary analysis and decision-making module is used to construct a corresponding medical-driven diagnostic analysis model based on a hybrid graph neural network and Transformer. It inputs a set of important contribution features from multiple sources of medical data into the medical-driven diagnostic analysis model to perform diagnostic inference and prediction, output the diagnostic probability corresponding to the medical data, and make corresponding medical diagnostic suggestions based on the diagnostic probability decision analysis.

2. The intelligent diagnostic auxiliary analysis system based on medical data as described in claim 1, characterized in that, The multi-source medical data acquisition module includes the following functions: Acquire medical imaging data; Obtain medical laboratory test data; Acquiring medical gene response data; The pixel blur of the medical image data is calculated to obtain the pixel blur of the medical image; based on the pixel blur of the medical image, the medical image data is subjected to blur denoising processing to generate medical denoised image data; Missing data and medical gene response data were analyzed to obtain medical laboratory test data and medical gene response data. The dynamic time warping algorithm is used to perform time-series alignment processing on medical denoised image data, medical laboratory differential completion data, and medical gene response differential completion data to obtain a multi-source medical time-series driven dataset.

3. The intelligent diagnostic auxiliary analysis system based on medical data as described in claim 2, characterized in that, The missing difference completion analysis of medical laboratory test data and medical gene response data includes: Data missing points were identified in both medical laboratory test data and medical gene response data to obtain the corresponding data missing points in the laboratory data and gene data. By setting a neighborhood range interval corresponding to a neighborhood radius of 1, and filling the corresponding missing data points by calculating the mean of the corresponding data within the neighborhood radius at the corresponding missing data points, the corresponding medical missing data can be obtained. The distribution of the completed data is obtained by completing medical missing data, and the difference between the distribution of the completed data and the original data distribution before completion is calculated to obtain the data distribution difference value between the completed data and the original data. The data distribution difference between the completed data and the original data is compared and judged based on a preset distribution difference threshold of 0.

05. If the data distribution difference between the completed data and the original data is less than or equal to the preset distribution difference threshold of 0.05, the neighborhood radius corresponding to the neighborhood range interval is increased by 1 and the corresponding missing data points are refilled until the data distribution difference between the completed data and the original data is greater than the preset distribution difference threshold of 0.05, so as to obtain medical laboratory difference completion data and medical gene response difference completion data.

4. The intelligent diagnostic auxiliary analysis system based on medical data as described in claim 1, characterized in that, The dynamic feature engineering analysis module includes the following functions: The medical images in the multi-source medical time-series driven dataset are converted to grayscale to generate medical grayscale image sequences. The corresponding gray-level co-occurrence matrix is ​​obtained by medical gray-level image sequences, and image feature statistical analysis is performed on the medical gray-level image sequences based on the gray-level co-occurrence matrix to obtain medical image features, including image texture features corresponding to contrast, uniformity, correlation, energy and entropy. Experimental index fluctuation analysis was performed on the corresponding laboratory data within the multi-source medical time-series driven dataset to obtain the fluctuation characteristics of laboratory indicators. Pathological response assessment is performed on the corresponding gene response data within the multi-source medical time-series driven dataset to obtain gene pathological response scores; The corresponding medical imaging features, laboratory indicator fluctuation features, and gene pathology response scores are merged into a single dataset to obtain the complete set of corresponding multi-source medical data features.

5. The intelligent diagnostic auxiliary analysis system based on medical data as described in claim 4, characterized in that, The analysis of experimental index fluctuations in the laboratory data within the multi-source medical time-series driven dataset includes: We analyzed the magnitude of changes in experimental indicators in the corresponding laboratory data within the multi-source medical time-series driven dataset to obtain the magnitude of changes in laboratory indicators. Based on the magnitude of changes in laboratory indicators, the frequency of changes in laboratory indicators is statistically analyzed for corresponding laboratory data within the multi-source medical time-series driven dataset to obtain the frequency of changes in laboratory indicators. Based on the magnitude and frequency of changes in laboratory indicators, abnormal fluctuation ranges are analyzed for each laboratory indicator within the laboratory data to obtain the abnormal fluctuation ranges for each laboratory indicator. The slope of the fluctuation of the laboratory indicators is calculated based on the abnormal fluctuation range of each laboratory indicator. The fluctuation characteristics of laboratory indicators are obtained by combining the magnitude of changes in laboratory indicators, the frequency of changes in laboratory indicators, and the slope of fluctuations in laboratory indicators.

6. The intelligent diagnostic auxiliary analysis system based on medical data as described in claim 5, characterized in that, The calculation of the slope of index fluctuation based on the abnormal fluctuation range corresponding to each laboratory index includes: By identifying the abnormal fluctuation ranges corresponding to the changes in various laboratory indicators, we can obtain the corresponding time start and time end points of abnormal fluctuations. The duration of the abnormal fluctuation is calculated based on the start and end points of the abnormal fluctuation time. By obtaining the corresponding laboratory indicator change amplitude through the abnormal fluctuation range of each laboratory indicator, and calculating the indicator fluctuation slope based on the duration of abnormal fluctuation, the laboratory indicator fluctuation slope is obtained.

7. The intelligent diagnostic auxiliary analysis system based on medical data as described in claim 4, characterized in that, The pathological response assessment of the corresponding gene response data within the multi-source medical time-series driven dataset includes: Gene response feature matrix analysis was performed on the gene response data in the multi-source medical time-series driven dataset to obtain the medical gene response feature matrix, which includes the change rate of gene response expression, the corresponding clinical treatment response time point, and the influencing factors of gene response mutation. The disease pathological response stages are obtained, and the pathological response paths of the disease pathological response stages are simulated based on the medical gene response feature matrix and combined with the Monte Carlo simulation method to generate personalized gene pathological process response paths. Based on the personalized gene pathology process response path, a weighted average method is used, combined with the patient's actual medical clinical response, to comprehensively evaluate and calculate the gene pathology response score by analyzing the medical gene response feature matrix.

8. The intelligent diagnostic auxiliary analysis system based on medical data as described in claim 1, characterized in that, The medical feature screening module includes the following functions: Statistical analysis of the feature dimension distribution of each medical sub-feature within the full set of multi-source medical data features is performed to obtain the data feature dimension distribution corresponding to each medical sub-feature. Based on the data feature dimension distribution corresponding to each medical sub-feature, KL divergence is calculated for the corresponding medical sub-features in the full set of multi-source medical data features to obtain the feature distribution offset KL divergence corresponding to each medical sub-feature. Based on the feature distribution offset KL divergence corresponding to each medical sub-feature, medical sub-features with a value greater than 0.1 are selected, and the corresponding feature contribution selection process is automatically triggered. The contribution of each medical sub-feature is quantified by the Shapley value. At the same time, the medical sub-features are sorted from largest to smallest and the top 80% are retained to obtain a set of important contribution features from multiple sources of medical care.

9. The intelligent diagnostic auxiliary analysis system based on medical data as described in claim 1, characterized in that, The diagnostic auxiliary analysis and decision-making module includes the following functions: A medical-driven diagnostic analysis model is constructed based on a hybrid graph neural network and Transformer. The medical sub-features within the multi-source medical important contribution feature set are divided into three branches: imaging branch, experimental indicator branch, and gene branch. The medical sub-features corresponding to the three branches are then input into the medical-driven diagnostic analysis model for diagnostic inference and prediction, so as to output the diagnostic probability corresponding to the medical data. Obtain a medical clinical knowledge base, and based on the diagnostic probability and the decision analysis of the medical clinical knowledge base, generate corresponding medical diagnostic recommendations.

10. The intelligent diagnostic auxiliary analysis system based on medical data as described in claim 9, characterized in that, The step of inputting the medical sub-features corresponding to the three branches into the medical-driven diagnostic analysis model for diagnostic reasoning and prediction includes: The learning rate of the medical-driven diagnostic analysis model was set to 0.

5. By inputting the medical sub-features corresponding to the three branches into the medical-driven diagnostic analysis model, and by calculating the network AUC improvement rate corresponding to each epoch update time on the corresponding branches of the medical-driven diagnostic analysis model; The weights on each branch are dynamically adjusted based on the learning rate and the network AUC improvement rate corresponding to each epoch update time. The corresponding diagnosis probability is calculated by weighted inference prediction based on the weights and the corresponding medical sub-features on each branch, so as to output the diagnosis probability corresponding to the medical data.