Lung cancer drug resistance risk assessment system based on multi-modal data and light-weight network
The lung cancer drug resistance risk assessment system, which utilizes multimodal data and lightweight networks, addresses the issues of limited data dimensions and insufficient adaptability in traditional assessments. It enables personalized drug resistance risk assessment, improving assessment accuracy and clinical applicability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIAN MEDICAL UNIV
- Filing Date
- 2026-02-13
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional lung cancer drug resistance risk assessment relies on unimodal data, which has a single information dimension and is difficult to adapt to individual differences among different patients, resulting in poor assessment accuracy and affecting clinical medication decisions.
An evaluation system based on multimodal data and lightweight networks is adopted. Multiple response subgroups are obtained through a data clustering module to construct a specific evaluation network library. Network matching and risk assessment are performed through a pattern matching retrieval module.
It enables accurate assessment of lung cancer drug resistance risk, improves the accuracy and specificity of the assessment, avoids assessment bias caused by incomplete feature coverage or poor model generalization ability, and enhances clinical applicability.
Smart Images

Figure CN121709271B_ABST
Abstract
Description
A lung cancer drug resistance risk assessment system based on multimodal data and lightweight networks Technical Field
[0001] This application relates to the field of lung cancer drug resistance risk assessment, and in particular to a lung cancer drug resistance risk assessment system based on multimodal data and lightweight networks. Background Technology
[0002] With the development of lung cancer treatment technology, drug resistance risk assessment has become a key link in guiding clinical medication and optimizing treatment plans. Accurate and personalized assessment results are of great significance for improving patient treatment outcomes and reducing the risk of ineffective treatment.
[0003] Currently, traditional lung cancer drug resistance risk assessments rely heavily on single-modality data. The limited information dimensions lead to poor assessment accuracy and make it difficult to adapt to individual differences among patients. This not only fails to provide reliable decision-making support for clinical practice but may also affect treatment strategy formulation due to assessment bias, thereby reducing treatment effectiveness. Summary of the Invention
[0004] This application provides a lung cancer drug resistance risk assessment system based on multimodal data and lightweight networks, which improves the current situation of traditional lung cancer drug resistance risk assessment relying on single-modal data and limited information, and solves the problems of poor assessment accuracy and difficulty in adapting to individual patient differences.
[0005] The embodiments of this application disclose the following technical solutions:
[0006] This application provides a lung cancer drug resistance risk assessment system based on multimodal data and lightweight networks, the system comprising:
[0007] The data clustering and partitioning module is used to acquire multimodal data of historical patients' samples and corresponding drug resistance data, and extract response distribution patterns for clustering and partitioning to obtain multiple response subgroups;
[0008] The network library construction module is used to train multiple specific evaluation networks based on lightweight networks for each of the aforementioned response subgroups, thereby forming a specific evaluation network library.
[0009] The pattern matching retrieval module is used to obtain the target response distribution pattern of the target patient and perform network matching by traversing the specific assessment network library according to the target response distribution pattern.
[0010] The drug resistance risk assessment module is used to acquire target multimodal data of target patients, and combine the network matching results to perform risk assessment on the target multimodal data to generate drug resistance risk assessment results.
[0011] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0012] This application proposes a lung cancer drug resistance risk assessment system based on multimodal data and lightweight networks. By collecting historical patient data and target patient data step-by-step, constructing a specific assessment network library, matching and adapting assessment models, and conducting differentiated risk assessments, it achieves accurate assessment of drug resistance risk in lung cancer patients. First, multimodal data and corresponding drug resistance data of historical patients are collected. After data diversity verification, response distribution patterns are extracted, and multiple response subgroups are divided using adaptive clustering analysis. Then, a lightweight specific assessment network is constructed and trained for each response subgroup, and the specific response distribution patterns corresponding to each network are extracted and integrated to form a specific assessment network library. Next, medication records of target patients are obtained, target multimodal data and target drug resistance data are extracted, target response distribution patterns are constructed, and cross-entropy is calculated to obtain network matching degree. Single or multiple adapted specific assessment networks are selected according to the matching degree threshold, and fusion weights are calculated simultaneously in multi-model scenarios. Finally, the target multimodal data is input into the network corresponding to the matching result. A single model directly outputs the assessment result, while multiple models are weighted and fused based on the fusion weights to obtain the final drug resistance risk assessment result.
[0013] The technical solution of this application solves the problems of single data dimension, insufficient model adaptability and low assessment efficiency in traditional lung cancer drug resistance assessment. It avoids assessment bias caused by incomplete feature coverage or poor model generalization ability, and improves the accuracy, pertinence and clinical applicability of lung cancer drug resistance risk assessment. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 is a schematic diagram of the structure of the lung cancer drug resistance risk assessment system based on multimodal data and lightweight network provided in the embodiment of this application;
[0016] Figure 2 is a schematic diagram of the process for constructing a specific evaluation network library provided in an embodiment of this application.
[0017] The components represented by each number in the attached diagram are explained below:
[0018] Data clustering and partitioning module 01, network library construction module 02, pattern matching and retrieval module 03, and drug resistance risk assessment module 04. Detailed Implementation
[0019] This application provides a lung cancer drug resistance risk assessment system based on multimodal data and lightweight networks to address the technical problems in existing technologies where traditional lung cancer drug resistance risk assessment relies on single-modal data, provides limited information, has poor assessment accuracy, and is difficult to adapt to individual patient differences, resulting in the inability to provide personalized assessment results and affecting clinical medication decisions.
[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0022] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0023] Example 1, as shown in Figure 1, this application provides a lung cancer drug resistance risk assessment system based on multimodal data and lightweight networks. The system includes the following steps:
[0024] The data clustering and partitioning module 01 is used to acquire multimodal data of historical patients' samples and corresponding drug resistance data, and extract response distribution patterns for clustering and partitioning to obtain multiple response subgroups.
[0025] In this embodiment of the application, in the scenario where lung cancer drug resistance risk assessment relies on data to support model training, in order to ensure that the assessment model can adapt to the individual characteristics of different patients, it is necessary to first collect relevant historical patient data and classify it according to response characteristics to improve the pertinence and reliability of the assessment results.
[0026] Specifically, the process begins by combining the target patients' medication usage information with pre-defined sample window constraints to collect raw sample data and form a raw sample dataset. The sample window constraints must be set in accordance with the temporal characteristics of the clinical data to ensure that the collected data effectively reflects the correlation between drug use and drug resistance.
[0027] Furthermore, a first preset number of sample extractions are performed on the original sample dataset, followed by data diversity verification using diversity entropy. If the diversity entropy of the extraction results reaches a preset first threshold, the result is directly output; otherwise, a second preset number of sample extractions are performed and the results are randomly updated until the first threshold requirement is met.
[0028] The first preset number of times is N times the second preset number of times, and N is greater than or equal to 5. The diversity and representativeness of the sample data are ensured by extracting and verifying multiple times.
[0029] After acquiring the sample data, the multimodal data of the samples are traversed to extract and form multiple multimodal sequence data groups, with each data group corresponding to a historical patient.
[0030] Furthermore, by combining the sample multimodal sequence data set with the corresponding sample drug resistance data, the correlation coefficient between each mode and the sample drug resistance data is calculated using the correlation analysis method. Then, the correlation coefficients of all modes are normalized and concatenated to form a response distribution vector, which is then output as the response distribution pattern.
[0031] Finally, adaptive clustering analysis was performed on the response distribution patterns of all historical patients. Based on the analysis results, each cluster was defined as a response pattern. Then, the multimodal data of the samples and the corresponding drug resistance data of the samples were divided according to multiple response patterns to obtain multiple response subgroups.
[0032] In the system provided in this application embodiment, the data clustering and partitioning module 01 includes:
[0033] Raw sample data is collected based on the target patients' medication usage information and the preset sample window constraints to obtain the raw sample dataset;
[0034] A first preset number of sample extractions are performed on the original sample dataset, and the sample extraction results are verified for data diversity based on diversity entropy.
[0035] If the diversity entropy of the sample extraction result meets the preset first threshold, the sample extraction result is output; otherwise, a second preset number of sample extractions are performed and the sample extraction result is randomly updated until the preset first threshold is met.
[0036] The sample extraction results include the sample multimodal data and the corresponding sample drug resistance data, wherein the first preset number of times is N times the second preset number of times, and N is greater than or equal to 5.
[0037] The sample multimodal data is traversed to extract multiple sample multimodal sequence data sets, wherein each sample multimodal sequence data set corresponds to a historical patient.
[0038] By combining the sample multimodal sequence data set with the corresponding sample drug resistance data, the correlation coefficient between each modality and the sample drug resistance data is calculated using a correlation analysis method.
[0039] The correlation coefficients of multiple modes are normalized and then concatenated to form a response distribution vector, which is then output as the response distribution pattern.
[0040] Adaptive clustering analysis was performed on multiple response distribution patterns corresponding to multiple historical patients;
[0041] Based on the results of adaptive clustering analysis, each cluster is defined as a response mode, and the multimodal data of the sample and the corresponding drug resistance data of the sample are divided according to multiple response modes to obtain multiple response subgroups.
[0042] In this embodiment of the application, in order to ensure that the drug resistance risk assessment can be adapted to the individual characteristics of different patients, it is necessary to collect relevant historical patient data, extract response distribution patterns and perform reasonable clustering through a standardized process, so as to improve the training efficiency of the subsequent assessment model and the accuracy of the assessment results.
[0043] Specifically, the process begins by collecting raw sample data based on the target patients' medication usage information and pre-defined sample window constraints, forming a raw sample dataset. The sample window constraints must be set in conjunction with the cyclical characteristics of clinical treatment; for example, a timeframe of 3-6 months after medication use may be selected as the sample window to avoid insufficient data correlation due to a timeframe that is too short or too long.
[0044] Meanwhile, during the data collection process, it is necessary to ensure the integrity and accuracy of the original sample data, covering various modal data related to lung cancer treatment and corresponding drug resistance results records, to ensure that the data can support subsequent analysis.
[0045] Furthermore, after obtaining the original sample dataset, a first preset number of sample extractions are performed, and the extraction results are validated for data diversity using diversity entropy to avoid data bias caused by samples being concentrated on a single feature, and to ensure that the extracted samples can cover patient characteristics of different medication situations and disease stages.
[0046] Specifically, if the diversity entropy of the extracted results reaches a preset first threshold, the result is directly output as the final set of sample multimodal data and corresponding sample drug resistance data. Conversely, if the diversity entropy of the extracted results does not reach the preset first threshold, a second preset number of sample extractions are performed and the results are randomly updated until the first threshold requirement is met.
[0047] The first preset number of times is N times the second preset number of times, and N is greater than or equal to 5. For example, the first preset number of times is set to 50 times and the second preset number of times is set to 10 times. Through multiple extractions and verifications, the sample data is made more diverse and the samples can cover patients with different medication regimens and different stages of the disease.
[0048] Furthermore, after sample extraction is completed, the multimodal data of the samples are traversed to extract the corresponding multimodal sequence data set for each historical patient to ensure the correspondence between the data and the individual patient.
[0049] Meanwhile, by combining the multimodal sequence data set of the samples with the corresponding drug resistance data of the samples, the correlation coefficient between each modality and the drug resistance data of the samples was calculated through correlation analysis to quantify the degree of influence of different modal data on drug resistance results and screen out key modal information that is strongly correlated with drug resistance risk.
[0050] Specifically, for each modality in the multimodal sequence data set, the linear or nonlinear correlation strength between each modality and the corresponding drug resistance data is calculated, and redundant modalities with no or weak correlation are eliminated to ensure that the retained modal data has practical reference value.
[0051] Furthermore, the correlation coefficients of all modes are normalized and then concatenated to form a response distribution vector that reflects the association between each mode and drug resistance, which is then output as the response distribution pattern.
[0052] For example, if a sample contains three modalities: imaging features, clinical laboratory indicators, and medication cycle, and the correlation coefficients between these three modalities and drug resistance data are calculated to be 0.72, 0.65, and 0.58, respectively, and after normalization they are 0.38, 0.34, and 0.28, respectively, and the concatenation of these modalities forms a response distribution vector of [0.38, 0.34, 0.28], which is the response distribution pattern of the sample.
[0053] Furthermore, adaptive clustering analysis was performed on the response distribution patterns of all historical patients to identify patient groups with similar drug resistance response characteristics. This provides a clear data grouping basis for subsequent targeted training of the specific evaluation network and improves the network's adaptability to different patient groups.
[0054] The system provided in this application embodiment performs adaptive clustering analysis on multiple response distribution patterns corresponding to multiple historical patients, including:
[0055] The similarity between multiple response distribution patterns is calculated using radial basis function kernels to construct a similarity matrix;
[0056] Perform normalized spectral clustering on the similarity matrix, including:
[0057] Calculate the first M eigenvalues of the Laplacian matrix corresponding to the similarity matrix, where M is a positive integer;
[0058] The optimal number of clusters K is determined based on the numerical gaps between the first M feature values;
[0059] Based on the optimal number of clusters K, K-means clustering is performed on multiple response distribution patterns after dimensionality reduction by spectral clustering to obtain adaptive clustering analysis results.
[0060] Specifically, the similarity between multiple response distribution patterns is first calculated using radial basis function kernels to construct a similarity matrix. Radial basis function kernels, as a commonly used nonlinear kernel function, can effectively capture the complex nonlinear relationships between response distribution patterns in high-dimensional space, and are particularly suitable for processing unstructured data such as response distribution patterns formed after multimodal data fusion.
[0061] For example, for two response distribution patterns that contain multimodal association information such as imaging features, clinical test indicators, and medication history, the radial basis function kernel can quantify the similarity between the two in drug resistance association features, avoiding the shortcomings of traditional linear similarity calculation methods in characterizing complex associations.
[0062] During the calculation process, the bandwidth parameter of the radial basis function kernel needs to be set appropriately. For example, based on the feature dimension and numerical range of the response distribution pattern, the bandwidth parameter can be set to a reasonable value between 0.1 and 1.0 to ensure the stability of the similarity calculation. This ensures that response distribution patterns with similar features exhibit high similarity values, while patterns with significant feature differences maintain low similarity values, thus providing a reliable similarity measurement basis for subsequent clustering.
[0063] Furthermore, after constructing the similarity matrix, normalized spectral clustering is performed. This process is a key step in achieving adaptive clustering. By extracting eigenvalues from the Laplacian matrix transformed from the similarity matrix, determining the optimal number of clusters, and completing the clustering, the intrinsic correlation of the response distribution pattern can be accurately analyzed.
[0064] Specifically, the first M eigenvalues of the Laplacian matrix corresponding to the similarity matrix are calculated, where M is a positive integer. The Laplacian matrix effectively reflects the structural information of the similarity matrix. By performing eigenvalue decomposition on it and extracting the first M eigenvalues, dimensionality reduction of the high-dimensional response distribution pattern can be achieved while preserving the key structural features of the data.
[0065] For example, when the dimension of the response distribution pattern is high, setting M to a value between 20 and 50 can significantly reduce the data dimension and subsequent computational load, while preserving the core information related to clustering to the greatest extent.
[0066] Meanwhile, high-precision numerical calculation methods must be employed during eigenvalue calculation to avoid inaccurate eigenvalue extraction due to numerical errors. For example, the Jacobi iteration method can be used for eigenvalue decomposition of the Laplace matrix, and the numerical accuracy can be optimized through multiple iterations to ensure that the first M eigenvalues can truly reflect the structural correlation characteristics of the response distribution pattern.
[0067] Furthermore, the optimal number of clusters K is determined based on the numerical gaps between the first M feature values. The numerical gaps between the feature values directly reflect the natural clustering trend of the data. The logic is as follows: if there is a significant numerical jump between the Kth feature value and the (K+1)th feature value among the first M feature values, it indicates that there is a clear group division boundary in this dimension. In this case, K is the number of clusters that best fits the data distribution characteristics.
[0068] For example, if among the first 30 feature values, the fourth feature value is 2.8 and the fifth feature value drops sharply to 1.2, forming a significant numerical gap, then the optimal number of clusters K can be determined to be 4. This method of determining the number of clusters based on the gap between feature values relies entirely on the distribution pattern of the data itself, avoiding the problem of unreasonable division that may be caused by manually pre-setting the number of clusters, and making the clustering results more scientific.
[0069] Finally, based on the determined optimal number of clusters K, K-means clustering is performed on the multiple response distribution patterns after dimensionality reduction by spectral clustering to obtain the final adaptive clustering analysis results. The K-means clustering algorithm can quickly divide the dimensionality-reduced response distribution patterns into K independent clusters with similar internal characteristics.
[0070] For example, when the optimal number of clusters K is determined to be 5, the K-means clustering algorithm can accurately divide all response distribution patterns into 5 clusters. The response distribution patterns within each cluster are highly similar in terms of the characteristics of each modality and drug resistance association, representing patient groups with similar drug resistance response characteristics. For example, some clusters correspond to patients who are sensitive to a certain type of targeted drug and have a low risk of drug resistance, while some clusters correspond to patients with a moderate risk of drug resistance and are related to specific clinical indicators.
[0071] Furthermore, after obtaining the adaptive clustering analysis results, each cluster is defined as a response mode based on the results, and the multimodal data of the samples and the corresponding drug resistance data of the samples are divided according to multiple response modes to obtain multiple response subgroups.
[0072] Specifically, each cluster is first assigned a unique response pattern identifier, which directly corresponds to the core feature of all response distribution patterns within the cluster. For example, clusters focusing on the characteristic of targeted drug sensitivity are defined as response pattern A, and clusters with the characteristic of high risk of chemotherapy drug resistance are defined as response pattern B, etc., so that each response pattern has a clear characteristic orientation.
[0073] Furthermore, the multimodal data of the samples and the corresponding drug resistance data of the samples are associated and classified based on multiple response patterns. Since each response distribution pattern has a one-to-one correspondence with the multimodal data and drug resistance data of specific historical patients, the corresponding patient data can be classified into the same response subgroup by the cluster to which the response distribution pattern belongs (i.e., the response pattern).
[0074] For example, for all patients whose response distribution pattern belongs to response pattern A, their corresponding multimodal data and drug resistance data will be classified into response subgroup A; the data of patients whose response pattern belongs to response pattern B will be classified into response subgroup B, and so on.
[0075] During the partitioning process, data integrity verification is also required for each response subgroup after partitioning to ensure that each subgroup contains a sufficient number of sample data and covers the corresponding multimodal data and drug resistance data, so as to avoid the impact of missing data on the subsequent model training effect.
[0076] For example, if clustering yields 5 response patterns, corresponding to 5 characteristics: targeted drug sensitivity, chemotherapy resistance, moderate immunotherapy response, high risk of multidrug resistance, and initial treatment sensitivity followed by drug resistance, then 5 response subgroups can be obtained by division.
[0077] The response subgroup 1 includes imaging data, clinical indicators, medication records, and corresponding drug resistance results for all patients with targeted drug sensitivity characteristics; the response subgroup 4 gathers various relevant data for patients at high risk of multidrug resistance. The patient data within each response subgroup are highly homogeneous in terms of drug resistance response characteristics, providing a highly targeted data foundation for training specific assessment networks for different subgroups and improving the accuracy of subsequent risk assessments.
[0078] Network library construction module 02 is used to train multiple specific evaluation networks based on lightweight networks for each of the aforementioned response subgroups, thereby forming a specific evaluation network library;
[0079] In this embodiment of the application, in order to have a dedicated evaluation model for each response subgroup, a lightweight specific evaluation network needs to be constructed and trained, and then integrated into the specific evaluation network library after associating features, so as to support subsequent rapid network matching and risk assessment.
[0080] Specifically, the first step is to initialize and construct multiple specific evaluation networks based on lightweight neural networks. These lightweight networks feature simplified parameters and high computational efficiency, enabling them to reduce computational resource consumption while ensuring evaluation effectiveness, thus meeting the needs of rapid clinical evaluation.
[0081] Furthermore, the segmented multimodal data and drug resistance data of the samples were used as the training basis, with each response subgroup serving as an independent training data set to train multiple specific evaluation networks. This allowed each network to focus on learning the characteristic association patterns of its corresponding subgroup, improving the model's adaptability to specific populations.
[0082] After training, the subgroup response distribution pattern of each network's corresponding response subgroup is extracted, the central pattern is identified and output as the specific response distribution pattern, which embodies the core adaptation features of the corresponding network.
[0083] Finally, multiple trained specific assessment networks are linked and integrated with their corresponding specific response distribution patterns to form a complete specific assessment network library. This module, through targeted training and system integration, provides reliable support for subsequent rapid matching of appropriate networks based on target patient characteristics.
[0084] As shown in Figure 2, in the system provided by this application embodiment, the network library construction module 02 includes:
[0085] Initialize and construct multiple specific evaluation networks based on lightweight neural networks;
[0086] The multimodal data of the sample after division and the drug resistance data of the sample are defined as the response subgroups, and multiple specific evaluation networks are trained using each response subgroup as a set of training data.
[0087] For the trained specific evaluation network, extract the corresponding subgroup response distribution patterns, identify the center pattern, and output the specific response distribution patterns.
[0088] The specific evaluation network library is obtained by associating multiple trained specific evaluation networks with their corresponding specific response distribution patterns.
[0089] In this embodiment of the application, in order to avoid the problem of low evaluation accuracy or long computation time due to insufficient adaptability of general models, it is necessary to construct a training-specific evaluation network and integrate it into a structured network library to achieve efficient calling of the evaluation model, thereby supporting the need for rapid and reliable drug resistance risk assessment in clinical scenarios.
[0090] Specifically, the first step is to initialize and construct multiple specific evaluation networks based on lightweight neural networks. Among them, lightweight neural networks, with their advantages of simplified parameter size and low computational complexity, can significantly reduce the computation time in the inference stage while ensuring the model's learning ability, thus meeting the urgent clinical demand for evaluation efficiency.
[0091] During the construction process, it is necessary to combine the task characteristics of lung cancer drug resistance assessment and rationally design the input layer dimension, hidden layer structure and output layer form of the network to ensure that the network can effectively receive multimodal data, capture the correlation between data and drug resistance risk, and avoid resource waste caused by redundant network structure.
[0092] For example, to meet the fusion requirements of multimodal data, a multi-channel data access interface is designed in the network input layer, such as image feature channel, clinical indicator channel, and medication time series channel, so that different types of modal data can be accurately input and participate in model training, laying a structural foundation for subsequent targeted learning.
[0093] Furthermore, the segmented multimodal data and drug resistance data of the samples are defined as response subgroups, and each response subgroup is used as an independent training data set to specifically train multiple specific evaluation networks. The patient data within each response subgroup exhibits high homogeneity in drug resistance response characteristics; using this as dedicated training data avoids the decline in model generalization ability caused by the mixing of characteristics from different groups.
[0094] During training, an appropriate loss function is used to measure the deviation between the model's predictions and actual drug resistance data. Gradient descent-type optimization algorithms are used to continuously adjust network parameters until the model's prediction accuracy on the training data reaches a preset standard. For example, for a response subgroup whose core feature is sensitivity to targeted drugs, the model will focus on learning the correlation between imaging features, clinical indicators, and low drug resistance risk within this subgroup during training, so that the trained model can accurately identify the drug resistance characteristics of similar patients.
[0095] Furthermore, after all the specific evaluation networks have completed training, it is necessary to extract the subgroup response distribution pattern of the response subgroup to which the training data of each network belongs, identify the center pattern from it, and output it as the specific response distribution pattern.
[0096] Among them, the subgroup response distribution pattern is a concentrated manifestation of the response distribution characteristics of all patients within the response subgroup, while the central pattern is the most representative distribution pattern that can reflect the core characteristics of the group and can be used as the core identifier of the specific assessment network's suitability.
[0097] During the extraction process, statistical analysis of all response distribution patterns within a subgroup is required. Central patterns are identified by calculating feature mean and cluster centers to ensure they accurately represent the characteristics of the corresponding network's target population. For example, after a specific assessment network is trained on a subgroup exhibiting a "high risk of chemotherapy resistance," its corresponding specific response distribution patterns will centrally reflect the core characteristics of this subgroup in the correlation between various modalities and drug resistance data, providing a feature-based basis for subsequent network matching.
[0098] Finally, multiple trained specific evaluation networks are associated and stored with their corresponding specific response distribution patterns to form a complete specific evaluation network library. During the association and storage process, a clear indexing mechanism needs to be established so that each specific evaluation network can be quickly retrieved through its corresponding specific response distribution pattern.
[0099] For example, a feature vector index is established for each specific response distribution pattern. Once the target response distribution pattern of the target patient is subsequently obtained, the appropriate specific assessment network can be quickly located by calculating feature similarity. Simultaneously, the network library needs to be structurally managed to ensure an accurate one-to-one correspondence between networks and feature patterns, avoiding indexing errors during the matching process.
[0100] Through the above steps, the constructed specific assessment network library contains both dedicated models adapted to different patient groups and structured features for efficient retrieval, providing reliable model support for subsequent network matching and risk assessment of target patients.
[0101] The pattern matching retrieval module 03 is used to obtain the target response distribution pattern of the target patient and perform network matching by traversing the specific evaluation network library according to the target response distribution pattern.
[0102] In this embodiment of the application, in order to avoid evaluation bias due to improper model adaptation, it is necessary to first obtain the target response distribution pattern that can fully reflect the patient characteristics, and then use a scientific matching algorithm to traverse the network library to screen the optimal or better model, so as to ensure the reliability of subsequent drug resistance risk assessment.
[0103] Specifically, the first step is to obtain the medication records of the target patients. These records contain key information about the patients' treatment process, and the target multimodal data and target drug resistance data can be extracted based on these records.
[0104] Furthermore, based on the correlation analysis method, each modality of the target multimodal data is traversed, and the correlation coefficient between each modality and the target drug resistance data is calculated one by one to quantify the degree of influence of different modal data on drug resistance results and screen out key information that is strongly correlated with drug resistance risk.
[0105] Subsequently, all calculated correlation coefficients are normalized to eliminate the influence of differences in the dimensions of data from different modalities. The normalized correlation coefficients are then concatenated in a preset order to form a target response distribution vector, which is finally output as the target response distribution pattern.
[0106] Furthermore, using the target response distribution pattern as the core of the retrieval, the specific assessment network library is traversed for network matching. During the traversal, the similarity between the target response distribution pattern and the specific response distribution pattern corresponding to each specific assessment network is quantified by calculating the cross-entropy. The smaller the cross-entropy value, the higher the fit between the two patterns. Therefore, the reciprocal of the cross-entropy is defined as the network matching degree to intuitively reflect the degree of fit between the target patient and each specific assessment network.
[0107] Furthermore, the matching scores of multiple networks are iterated and compared with a preset first matching score threshold. If at least one network matching score satisfies the first matching score threshold, the specific evaluation network corresponding to the largest one is output as the network matching result.
[0108] Conversely, if no network matching degree satisfies the first matching degree threshold, the specific evaluation network corresponding to the top Z network matching degrees that satisfy the preset second matching degree threshold is extracted to form the network matching result.
[0109] In the system provided in this application embodiment, the pattern matching retrieval module 03 includes:
[0110] Obtain the medication records of the target patient, and extract target multimodal data and target drug resistance data based on the medication records, wherein the target multimodal data is time series data;
[0111] Based on the correlation analysis method, each mode of the target multimodal data is traversed, the correlation coefficient with the target drug resistance data is calculated, and the correlation coefficient is normalized and concatenated into a target response distribution vector;
[0112] The target response distribution vector is output as the target response distribution pattern.
[0113] Traverse the specific evaluation network library, calculate the cross-entropy between the target response distribution pattern and the specific response distribution pattern corresponding to each specific evaluation network, and output the reciprocal of the cross-entropy as the network matching degree.
[0114] The network matching degree of multiple networks is compared with a preset first matching degree threshold. If at least one network matching degree satisfies the first matching degree threshold, the specific evaluation network corresponding to the largest of the multiple network matching degrees is output as the network matching result.
[0115] If no network matching degree satisfies the first matching degree threshold, then the specific evaluation network corresponding to the top Z network matching degrees that satisfy the preset second matching degree threshold and have the largest value is extracted as the network matching result.
[0116] In this embodiment of the application, in order to enable the target patient to quickly match the most suitable specific assessment network, it is necessary to first construct a target response distribution pattern that reflects the individual characteristics of the patient, and then traverse the specific assessment network library through matching and quantification indicators for screening. At the same time, a weight foundation is reserved for multi-model fusion to ensure the reliability of subsequent assessment results.
[0117] Specifically, the medication records of target patients are first obtained through clinical data carriers such as electronic medical record systems and hospital information management platforms. These records comprehensively document key information such as the patient's treatment plan, medication cycle, and dosage adjustments, serving as a crucial basis for extracting core data.
[0118] Furthermore, target multimodal data and target drug resistance data are extracted based on medication records. The target multimodal data is time-series data, which can dynamically present various characteristic changes of patients at different time points throughout the entire medication process, such as the evolution of imaging characteristics, fluctuations in clinical test indicators, and changes in gene expression levels in the early, middle, and late stages of medication.
[0119] Furthermore, based on correlation analysis, each modality of the target multimodal data is traversed, and the correlation coefficient between each modality and the target drug resistance data is calculated one by one. This process aims to quantify the degree of influence of different modalities on drug resistance results, screen out key information that is strongly correlated with drug resistance risk, and eliminate irrelevant or weakly correlated redundant data.
[0120] Furthermore, all calculated correlation coefficients are normalized to eliminate interference caused by differences in dimensions and numerical ranges of data from different modalities, making the correlation coefficients of each modality comparable. The normalized correlation coefficients are then concatenated in a preset order to form a target response distribution vector, which is finally output as the target response distribution pattern.
[0121] For example, if the target patient's medication record covers a 6-month treatment cycle of a certain targeted drug, the target multimodal data extracted based on the record includes time-series data such as chest CT imaging features, serum tumor marker levels, and gene mutation status before treatment, 2 months after treatment, 4 months after treatment, and 6 months after treatment. The target drug resistance data is the drug resistance determination result after 6 months of treatment.
[0122] Based on the correlation analysis method, the correlation coefficients between imaging feature modality, tumor marker modality, gene mutation modality and drug resistance results were calculated, yielding 0.75, 0.68, and 0.82, respectively. After normalization, these correlation coefficients were converted to 0.34, 0.31, and 0.35. These coefficients were then concatenated in a preset order of imaging features, tumor markers, and gene mutations to form a target response distribution vector of [0.34, 0.31, 0.35]. This vector represents the target response distribution pattern of the patient, which centrally reflects the correlation between its multimodal characteristics and drug resistance risk.
[0123] Furthermore, after constructing the target response distribution pattern, this pattern is used as the core search condition to traverse the specific evaluation network library for network matching. During the traversal, the similarity between the target response distribution pattern and the specific response distribution pattern corresponding to each specific evaluation network is quantified by calculating the cross-entropy between them.
[0124] Cross-entropy is a commonly used indicator to measure the difference between two probability distributions. The smaller the value, the higher the fit between the two patterns. Therefore, the reciprocal of cross-entropy is defined as the network matching degree. The larger the network matching degree value, the more suitable the corresponding specific assessment network is for the target patient.
[0125] For example, if the cross-entropy between the target response distribution pattern and a certain specific response distribution pattern is 0.2, the corresponding network matching degree is 5; if the cross-entropy is 0.5, the corresponding network matching degree is 2, which intuitively reflects that the former's adaptability is much higher than the latter.
[0126] Furthermore, all calculated network matching scores are compared with a preset first matching score threshold. The first matching score threshold is a key criterion for determining whether an optimally matched model exists. Its setting needs to be combined with the accuracy requirements of clinical assessment and the performance of the network library model. For example, setting the first matching score threshold to 4 means that only models with a network matching score of not less than 4 can meet the accuracy requirements of optimal matching.
[0127] If at least one network matching degree meets the first matching degree threshold, it indicates that there is a specific evaluation network that is highly consistent with the characteristics of the target patient. In this case, the specific evaluation network corresponding to the largest network matching degree among multiple network matching degrees that meet the conditions is output as the network matching result, ensuring that the matched model can fit the patient characteristics to the greatest extent.
[0128] Conversely, if there is no network matching degree that meets the first matching degree threshold, it means that no single model can perfectly match the target patient. In this case, the specific evaluation network corresponding to the top Z network matching degrees that meet the preset second matching degree threshold and have the largest value should be extracted as the network matching result.
[0129] The second matching threshold is set lower than the first matching threshold, for example, 2. This ensures basic fit between the selected model and patient characteristics while avoiding situations where no matching results are obtained due to excessively high standards. Meanwhile, Z is a positive integer, and its value needs to balance assessment accuracy and computational efficiency. For example, setting it to 3 selects the top three models in terms of fit, using multi-model fusion to compensate for the shortcomings of a single model's insufficient fit.
[0130] In the system provided in this application embodiment, the top Z network matching scores that satisfy a preset second matching score threshold and have the highest specific evaluation scores are extracted as network matching results. Afterwards, the system further includes:
[0131] Obtain the Z network matching degrees corresponding to the network matching results;
[0132] The fusion weights are calculated by combining a preset nonlinear function and Z network matching degrees.
[0133] The fusion weights are associated with and stored in relation to the network matching results.
[0134] Specifically, the Z network matching degrees corresponding to the network matching results are first obtained. These Z network matching degrees are quantitative indicators of the degree of fit between the top Z optimal fit-specific evaluation networks and the target response distribution pattern, and their values reflect the evaluation adaptation potential of the corresponding model for the target patient.
[0135] For example, when Z=3, the obtained three network matching scores are 3.9, 3.3, and 2.8, clearly showing the adaptation differences of each model. During the acquisition process, it is necessary to ensure the accuracy of the association between the network matching score and the corresponding specific evaluation network, ensuring that each network matching score corresponds to its respective specific evaluation network.
[0136] Furthermore, the fusion weights are calculated by combining a preset nonlinear function with the matching degrees of Z networks. The selection of the nonlinear function should highlight the dominant role of the high-matching-degree model while also considering the auxiliary value of the low-matching-degree model, avoiding overly even or extreme weight distribution.
[0137] For example, an exponential function or a sigmoid function can be used. These functions can amplify the differences in network matching, so that the model with the higher matching degree gets a larger weight ratio, thus playing a core role in the fusion evaluation.
[0138] Taking the exponential function as an example, if the network matching degrees are 3.9, 3.3, and 2.8, the numerical growth corresponding to higher matching degrees is more significant after calculation using the exponential function. After normalization, a reasonable fusion weight allocation can be obtained. During the calculation process, the matching degrees of the Z networks need to be standardized first to eliminate the influence of differences in numerical ranges. Then, they are substituted into the nonlinear function for calculation. Finally, the calculation results are normalized to the 0-1 interval to ensure that the sum of all fusion weights is 1.
[0139] For example, let Z=3, corresponding to network matching degrees of 3.9, 3.3, and 2.8, using the exponential function f(x)=e x As a non-linear function, the exponent values of each matching degree are first calculated, namely e-1. 3.9 ≈49.4, e 3.3 ≈27.1, e 2.8 The sum is approximately 16.4. Then, the total is calculated to be 49.4 + 27.1 + 16.4 ≈ 92.9. The fusion weights are obtained by dividing each index value by the sum, which are 49.4 / 92.9 ≈ 0.53, 27.1 / 92.9 ≈ 0.29, and 16.4 / 92.9 ≈ 0.18.
[0140] Finally, the calculated fusion weights are associated and stored with the corresponding network matching results. During this association and storage process, a clear data index structure needs to be established so that each specific evaluation network can be associated one-to-one with its corresponding fusion weight. For example, a structured data table can be constructed using the mapping relationship between model IDs and weight values.
[0141] The storage method described above ensures that the weights corresponding to each model can be quickly retrieved during subsequent multi-model fusion evaluations, eliminating the need for recalculation and improving evaluation efficiency. Simultaneously, the stored data must be validated for integrity to ensure that all Z specific evaluation networks have corresponding fusion weights, and that the weight values conform to a preset reasonable range, preventing errors in subsequent fusion evaluations due to missing or abnormal data.
[0142] Ultimately, the above process provides data support for subsequent multi-model weighted fusion evaluation, ensuring that even without an absolutely optimal model, the advantages of each model can be complemented through scientific weight allocation, thereby improving the accuracy and reliability of the overall drug resistance risk assessment.
[0143] The drug resistance risk assessment module 04 is used to acquire target multimodal data of the target patient, and combine the network matching results to perform risk assessment on the target multimodal data to generate drug resistance risk assessment results.
[0144] In this embodiment of the application, in order to make the evaluation results fit the individual characteristics of the patient and the model fit, it is necessary to first obtain complete target multimodal data, and then adopt the corresponding evaluation method according to the number of models in the network matching results, so as to ensure that the evaluation results are accurate and reliable and provide a scientific basis for clinical treatment decisions.
[0145] Specifically, the first step is to acquire target multimodal data based on the medication records of the target patients to comprehensively capture the dynamic changes in key characteristics during the treatment process. Simultaneously, the acquisition process must ensure data integrity and validity, removing outliers to guarantee data quality meets assessment requirements.
[0146] Furthermore, the target multimodal data is input into the specific evaluation network corresponding to the network matching result to obtain the drug resistance risk assessment result. If the network matching result contains only a single specific evaluation network, it indicates that the model is highly adapted to the patient characteristics, and its output is directly output as the assessment result, ensuring assessment efficiency and accuracy.
[0147] Conversely, if the network matching results contain multiple specific evaluation networks, the output results of multiple models are weighted and fused according to the pre-constructed fusion weights to integrate the adaptation advantages of each model, make up for the limitations of a single model, and ultimately generate more robust and reliable drug resistance risk assessment results, providing scientific support for the formulation and adjustment of clinical treatment plans.
[0148] In the system provided in this application embodiment, the drug resistance risk assessment module 04 includes:
[0149] Based on the medication records of the target patients, obtain the target multimodal data of the target patients;
[0150] The target multimodal data is input into the specific evaluation network corresponding to the network matching result to obtain the drug resistance risk assessment result, including:
[0151] If the network matching result contains only a single specific evaluation network, then the output result of the specific evaluation network is directly output as the drug resistance risk assessment result;
[0152] If the network matching result contains multiple specific evaluation networks, then the output results of the multiple specific evaluation networks are weighted and fused according to the pre-constructed fusion weights in the network matching result to obtain the drug resistance risk assessment result.
[0153] In this embodiment of the application, in order to ensure that the drug resistance risk assessment results of the target patients are both consistent with individual characteristics and can give full play to the assessment advantages of the fitting model, it is necessary to first obtain high-quality target multimodal data, and then adopt a differentiated assessment strategy based on the number of models in the network matching results to ensure the scientificity and reliability of the assessment results.
[0154] Specifically, the first step is to acquire target multimodal data for the target patients based on their medication records. These medication records detail key information such as the patient's treatment plan, medication cycle, and dosage adjustments, serving as the core basis for extracting target multimodal data and ensuring a high correlation between the extracted data and the patient's treatment progress.
[0155] In addition, the target multimodal data covers various key features at different time points in the patient's treatment process, including imaging features, clinical test indicators, and gene expression data. Moreover, these data are presented in time-series form, which can dynamically reflect the physiological and pathological changes of patients during the treatment process, providing rich and individualized data support for a comprehensive assessment of drug resistance risk.
[0156] During the data acquisition process, it is necessary to control the quality of the data, remove abnormal data caused by detection errors, data entry errors, etc., and supplement missing key modal data to ensure the integrity and consistency of the data, thereby avoiding the impact of poor-quality data on the evaluation results.
[0157] Furthermore, the quality-verified target multimodal data is input into the specific evaluation network corresponding to the network matching result to initiate the drug resistance risk assessment process. This process employs differentiated assessment strategies based on the number of specific evaluation networks included in the network matching result to maximize the assessment effectiveness of the adaptation model.
[0158] Specifically, if the network matching result contains only a single specific assessment network, it indicates that this specific assessment network is the model best suited to the characteristics of the target patient. The feature association patterns learned during its training process highly match the multimodal data characteristics of the target patient, and it has the ability to accurately assess the patient's drug resistance risk. In this case, directly outputting the output of this specific assessment network as the drug resistance risk assessment result can ensure the accuracy of the assessment result to the greatest extent.
[0159] For example, after being trained on a "targeted therapy-sensitive population of EGFR gene-mutant lung cancer", this specific assessment network can accurately identify features related to this population in the multimodal data of target patients, and quickly output the corresponding drug resistance risk level, risk probability and key risk association factors, providing clinicians with clear and specific assessment references.
[0160] Conversely, if the network matching results contain multiple specific evaluation networks, it indicates that no single model can fully adapt to the complex characteristics of the target patient. In this case, it is necessary to perform weighted fusion of the output results of multiple specific evaluation networks based on the pre-constructed fusion weights in the network matching results, so as to integrate the evaluation advantages of each model and make up for the limitations of a single model.
[0161] Among them, the fusion weight is calculated by nonlinear function based on the network matching degree of each model. It can objectively reflect the degree of fit of each model to the target patient. The higher the fit of the model, the greater the fusion weight, and the higher the proportion of its output in the final evaluation result.
[0162] Specifically, in the weighted fusion process, the output of each specific assessment network is first standardized, converting the risk assessment values from different models to the same numerical range (0-1 interval) to ensure the comparability of the outputs. Then, the standardized outputs are weighted and summed according to the fusion weights to obtain a comprehensive drug resistance risk assessment result.
[0163] For example, if the network matching result contains three specific evaluation networks with corresponding fusion weights of 0.55, 0.3, and 0.15, and the standardized output risk probabilities are 0.85, 0.72, and 0.68, the final drug resistance risk probability can be calculated by weighted fusion as 0.85×0.55+0.72×0.3+0.68×0.15=0.7955, which is approximately 0.80, corresponding to a high drug resistance risk level.
[0164] Ultimately, the fusion results fully absorbed the evaluation advantages of each model, highlighting the leading role of the highly adapted model while also taking into account the auxiliary reference value of other models. The results are more scientific and reliable than the output of a single model, and can effectively address situations where the characteristics of the target patients are complex and a single model cannot fully cover them.
[0165] Through the aforementioned differentiated assessment strategies, reliable drug resistance risk assessment results can be generated regardless of whether a single optimal model exists or multiple models are required for synergistic assessment. These results not only reflect the drug resistance risk level of the target patient but also provide clinicians with relevant risk-related information, helping them to adjust treatment strategies in a timely manner and thereby improve the treatment outcomes for lung cancer.
[0166] The embodiments of this application, through the specific implementation methods described above, achieve the following technical effects:
[0167] This application proposes a lung cancer drug resistance risk assessment system based on multimodal data and a lightweight network. First, raw sample data is collected by combining target patient drug usage information with sample window constraints. Valid datasets containing multimodal data and corresponding drug resistance data are selected through diversity entropy verification. Next, sequence data sets are extracted from the multimodal data sets, and correlation coefficients for each modality are calculated and normalized to form a response distribution pattern. A similarity matrix is constructed using radial basis function kernels, and multiple response subgroups are divided using normalized spectral clustering and K-means clustering. Subsequently, a lightweight specific assessment network is initialized and constructed for each response subgroup. Targeted training is conducted to extract specific response distribution patterns for each network, which are then integrated to form a specific assessment network library. Next, medication records of target patients are acquired, and time-series target multimodal data and target drug resistance data are extracted. Correlation coefficients are calculated and normalized to generate target response distribution patterns. The network library is traversed to calculate network matching degree using cross-entropy, and one or more suitable networks are selected according to a dual threshold rule. In multi-network scenarios, fusion weights are calculated synchronously and stored in association. Finally, the target multimodal data is input into the network corresponding to the matching result. A single network directly outputs the assessment result, while multiple networks are weighted and fused based on fusion weights to generate the final drug resistance risk assessment result.
[0168] The system provided in this application, through the technical solution of "data acquisition and verification - response pattern extraction and clustering - specific network training and library construction - target pattern matching - differentiated risk assessment", solves the problems of incomplete data feature coverage, poor model adaptability and low assessment efficiency in traditional lung cancer drug resistance assessment. It avoids assessment bias caused by data homogeneity and insufficient model generalization ability, and provides scientific and technical support for the formulation and adjustment of clinical personalized treatment plans.
[0169] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0170] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0171] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application intends to include such modifications and variations.
Claims
1. A lung cancer drug resistance risk assessment system based on multimodal data and lightweight networks, characterized in that, include: The data clustering and partitioning module is used to acquire multimodal data of historical patients' samples and corresponding drug resistance data, and extract response distribution patterns for clustering and partitioning to obtain multiple response subgroups; A network library construction module is used to train multiple specific evaluation networks based on lightweight networks for each of the aforementioned response subgroups, forming a specific evaluation network library. This includes: initializing and constructing multiple specific evaluation networks based on lightweight neural networks; defining the segmented multimodal data and drug resistance data of the samples as the response subgroups, and training multiple specific evaluation networks using each response subgroup as a set of training data; extracting the corresponding subgroup response distribution patterns from the trained specific evaluation networks, identifying the center pattern, and outputting it as a specific response distribution pattern; associating the multiple trained specific evaluation networks with the corresponding specific response distribution patterns to obtain the specific evaluation network library; a pattern matching retrieval module is used to obtain the target response distribution pattern of the target patient and perform network matching by traversing the specific evaluation network library according to the target response distribution pattern; and a drug resistance risk assessment module is used to obtain the target response distribution pattern of the target patient. The method involves obtaining target multimodal data of a target patient and performing a risk assessment on the target multimodal data in conjunction with network matching results to generate a drug resistance risk assessment result. This includes: acquiring target multimodal data of the target patient based on their medication records; inputting the target multimodal data into a specific evaluation network corresponding to the network matching result to obtain the drug resistance risk assessment result; if the network matching result contains only a single specific evaluation network, directly outputting the output of the specific evaluation network as the drug resistance risk assessment result; if the network matching result contains multiple specific evaluation networks, weighted fusion of the outputs of the multiple specific evaluation networks based on pre-constructed fusion weights in the network matching result to obtain the drug resistance risk assessment result.
2. The lung cancer drug resistance risk assessment system based on multimodal data and lightweight networks as described in claim 1, characterized in that, Acquiring historical patient sample multimodal data and corresponding sample drug resistance data includes: collecting raw sample data based on the target patient's drug usage information and a preset sample window constraint to obtain a raw sample dataset; performing a first preset number of sample extractions on the raw sample dataset and performing data diversity verification based on diversity entropy on the sample extraction results; if the diversity entropy of the sample extraction results meets a preset first threshold, then outputting the sample extraction results; otherwise, performing a second preset number of sample extractions and randomly updating the sample extraction results until the preset first threshold is met; wherein, the sample extraction results include the sample multimodal data and the corresponding sample drug resistance data, the first preset number of extractions is N times the second preset number of extractions, and N is greater than or equal to 5.
3. The lung cancer drug resistance risk assessment system based on multimodal data and lightweight networks as described in claim 1, characterized in that, The process involves extracting response distribution patterns and performing clustering to obtain multiple response subgroups. This includes: traversing the multimodal data of the samples to extract multiple multimodal sequence data sets, where each multimodal sequence data set corresponds to a historical patient; combining the multimodal sequence data sets with the corresponding drug resistance data of the samples, calculating the correlation coefficient between each modality and the drug resistance data of the samples based on association analysis; normalizing and concatenating the correlation coefficients of the multiple modalities to form a response distribution vector, which is then output as the response distribution pattern; performing adaptive clustering analysis on the multiple response distribution patterns corresponding to multiple historical patients; defining each cluster as a response pattern based on the adaptive clustering analysis results, and dividing the multimodal data of the samples and the corresponding drug resistance data of the samples according to the multiple response patterns to obtain multiple response subgroups.
4. The lung cancer drug resistance risk assessment system based on multimodal data and lightweight networks as described in claim 1, characterized in that, Obtaining the target response distribution pattern of a target patient includes: obtaining the target patient's medication records, and extracting target multimodal data and target drug resistance data based on the medication records, wherein the target multimodal data is time-series data; based on an association analysis method, traversing each mode of the target multimodal data, calculating the correlation coefficient with the target drug resistance data, and normalizing and concatenating them into a target response distribution vector; and outputting the target response distribution vector as the target response distribution pattern.
5. The lung cancer drug resistance risk assessment system based on multimodal data and lightweight networks as described in claim 1, characterized in that, The process of performing network matching based on the target response distribution pattern involves traversing the specific evaluation network library, calculating the cross-entropy between the target response distribution pattern and the specific response distribution pattern corresponding to each specific evaluation network, and outputting the reciprocal of the cross-entropy as the network matching degree; comparing multiple network matching degrees with a preset first matching degree threshold; if at least one network matching degree satisfies the first matching degree threshold, outputting the specific evaluation network corresponding to the largest of the multiple network matching degrees as the network matching result; if no network matching degree satisfies the first matching degree threshold, extracting the specific evaluation networks corresponding to the top Z network matching degrees that satisfy the preset second matching degree threshold as the network matching result.
6. The lung cancer drug resistance risk assessment system based on multimodal data and lightweight networks as described in claim 5, characterized in that, The specific evaluation network corresponding to the top Z network matching degrees that meet the preset second matching degree threshold is extracted as the network matching result. Then, the method further includes: obtaining the Z network matching degrees corresponding to the network matching results; calculating and obtaining the fusion weight by combining the preset nonlinear function and the Z network matching degrees; and storing the fusion weight in association with the network matching results.
7. The lung cancer drug resistance risk assessment system based on multimodal data and lightweight networks as described in claim 3, characterized in that, Adaptive clustering analysis is performed on multiple response distribution patterns corresponding to multiple historical patients, including: calculating the similarity between multiple response distribution patterns using radial basis function kernels to construct a similarity matrix; performing normalized spectral clustering on the similarity matrix, including: calculating the first M eigenvalues of the Laplacian matrix corresponding to the similarity matrix, where M is a positive integer; determining the optimal number of clusters K based on the numerical gaps of the first M eigenvalues; and performing K-means clustering on the multiple response distribution patterns after dimensionality reduction by spectral clustering based on the optimal number of clusters K to obtain the adaptive clustering analysis results.
Citation Information
Patent Citations
Gastric cancer treatment effect evaluation method and system based on multi-modal data and medium
CN120257111A