Large model optimization method based on multi-round diagnosis and treatment experience learning

By employing a large-scale model optimization method based on multi-round clinical experience learning, and utilizing a heterogeneous temporal multi-graph representation model and a dynamic replay strategy, the problem of insufficient feature representation in multimodal data is solved, thereby improving the accuracy of disease progression prediction and the model's generalization ability.

CN120998516APending Publication Date: 2025-11-21SECOND AFFILIATED HOSPITAL OF COLLEGE OF MEDICINEOF XIAN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510152359.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies cannot effectively discover and encode the internal relationships between different modalities, cannot learn higher quality feature representations from multimodal input data, and suffer from data gaps that affect the accuracy of disease progression models.

Method used

We employ a large model optimization method based on multi-round clinical experience learning. We use a heterogeneous temporal multi-graph representation model for multimodal temporal modeling and unified embedding representation learning. We combine recurrent neural networks and dynamic replay strategies, and utilize low-rank adaptive expert selection technology for model optimization.

Benefits of technology

It enables the learning of higher-quality feature representations from multimodal input data, improving the accuracy of disease progression prediction and the model's generalization ability, and solving the problems of missing data and data imbalance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998516A_ABST
    Figure CN120998516A_ABST
Patent Text Reader

Abstract

The invention discloses a large model optimization method based on multi-round diagnosis and treatment experience learning. The method comprises the steps that variable-length multi-modal time sequence modeling and unified embedded representation learning are carried out on dynamic domain experience through a heterogeneous tense multi-graph expression model; the disease course evolution is constructed into a heterogeneous, multi-modal, multi-scale and uncertainty weighted signal graph network cluster with a time sequence concept; carrying out strategic random walk resampling with a time concept on the network cluster; and performing large model optimization according to the sampling and multi-round diagnosis and treatment experience. According to the method, through a multi-modal domain experience unified representation modeling technology based on heterogeneous multi-graph progressive embedding learning, higher-quality feature expression is learned from multi-modal input data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of large models, and particularly relates to a large model optimization method based on multi-round diagnosis and treatment experience learning. BACKGROUND

[0002] In clinical practice, predicting the progression of a disease will help improve diagnosis and enable patients to take active care measures. However, most existing methods are only directed at individuals with a fixed number of historical visits, and can only predict the progression within a fixed time range in the future, which cannot meet actual requirements.

[0003] Some disease progression modeling models are developed based on the assumption of "data completeness", while in actual situations, data missing is a common and serious problem that always exists in the process of patient disease recovery. For example, patients may not appear at the time points previously agreed upon or even drop out of the study, which is referred to as "missing visits". In addition, not all patients have undergone comprehensive clinical examinations or provided complete data sets. For example, some patients only complete part of the necessary detection items due to economic conditions or other reasons, resulting in missing data in some modalities (such as laboratory indicators, imaging examination results, etc.), which is referred to as partial modality missing. In addition, during the follow-up of disease management or research, some patients may not complete follow-up at one or more time points, resulting in the inability to obtain data at key time nodes of the disease process, which becomes clinical follow-up missing.

[0004] These problems will affect researchers in building and verifying disease progression models, because continuous, complete and high-quality data are usually needed in model development to depict the evolution law of the disease. In order to solve such problems, statisticians and researchers usually use various data processing techniques, such as multiple imputation, pattern filling, Bayesian methods and other missing value analysis methods to estimate and process missing data as reasonably as possible, so as to reduce the potential bias caused by data missing.

[0005] The prior art proposes a medical image fusion method based on boundary measurement of pulse coupled neural network modulation, and estimates the future course of cancer patients by fusing features learned independently from clinical data, mRNA expression data, microRNA expression data and whole slide images of histopathology. However, these methods cannot discover and encode internal relations between different modalities, and cannot learn higher quality feature expressions from multi-modality input data. SUMMARY

[0006] The embodiment of the present application provides a large model optimization method based on multi-round diagnosis and treatment experience learning, to solve the problem that the prior art cannot discover and encode internal relations between different modalities, and cannot learn higher quality feature expressions from multi-modality input data.

[0007] In an aspect, the embodiments of the present application provide a large model optimization method based on multi-round diagnosis and treatment experience learning, comprising:

[0008] Variable-length multi-modal time-series modeling and unified embedding representation learning of dynamic domain experience are performed by a heterogeneous time-varying multi-graph expression model;

[0009] The course of disease evolution is constructed as a heterogeneous, multi-modal, multi-scale, and uncertainty-weighted signal graph network cluster with a time concept;

[0010] Strategic random walk resampling with a time concept is performed on the network cluster;

[0011] The large model is optimized according to the sampling and multi-round diagnosis and treatment experience.

[0012] In a possible implementation, the representation learning comprises:

[0013] Degenerate network is used for multi-modal domain experience fusion;

[0014] A recurrent neural network is used for multi-round diagnosis and treatment experience modeling.

[0015] In a possible implementation, the sampling comprises:

[0016] Historical data is utilized by a dynamic replay strategy;

[0017] Layered sampling is performed by using historical data.

[0018] In a possible implementation, the sampling further comprises limiting the number of sampling batches to control the training complexity.

[0019] In a possible implementation, the large model optimization according to the sampling and multi-round diagnosis and treatment experience further comprises fine-tuning the large model by a low-rank adaptive expert selection technique.

[0020] In a possible implementation, the low-rank adaptive expert selection technique further comprises constructing a specific task low-rank adaptive expert predictor technique.

[0021] In a possible implementation, the specific task is a diagnosis and treatment task for a specific disease.

[0022] In a possible implementation, before the large model optimization according to the sampling and multi-round diagnosis and treatment experience, an auxiliary cache system is constructed.

[0023] In a possible implementation, the auxiliary cache system is used to store logs generated by a tuned low-rank adaptive expert selection technique.

[0024] The large model optimization method based on multi-round diagnosis and treatment experience learning has the following advantages:

[0025] The multi-modal field experience unified representation modeling technology based on heterogeneous multi-graph progressive embedding learning realizes learning of higher quality feature expression from multi-modal input data. BRIEF DESCRIPTION OF DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0027] Figure 1 The heterogeneous time-state multi-graph perception and unified embedding representation learning framework for dynamic multi-modal field experience modeling of the large model optimization method based on multi-round diagnosis and treatment experience learning provided by the embodiments of the present application. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0029] Figure 1 The heterogeneous time-state multi-graph perception and unified embedding representation learning framework for dynamic multi-modal field experience modeling of the large model optimization method based on multi-round diagnosis and treatment experience learning provided by the embodiments of the present application. The embodiments of the present application provide a large model optimization method based on multi-round diagnosis and treatment experience learning, which comprises:

[0030] The dynamic field experience is variably long multi-modal time sequence modeling and unified embedding representation learning through the heterogeneous time-state multi-graph expression model;

[0031] The disease course evolution is constructed as a heterogeneous, multi-modal, multi-scale, uncertainty weighted signal graph network cluster with time sequence concept;

[0032] The network cluster is subjected to strategic random walk resampling with time concept;

[0033] The large model is optimized according to the sampling and multi-round diagnosis and treatment experience.

[0034] Exemplarily, as Figure 1The figure is annotated as follows: Prediction; Sequence Learning; addition operator; dot-product operator; concatenating operator; reconstruct; Degradation; In the optimization research of large models based on multi-round diagnosis and treatment experience learning, first, based on the heterogeneous time-varying multi-graph expression model, the variable-length multi-modal time sequence modeling and unified embedding representation learning of dynamic domain experience are carried out; secondly, the course evolution of a specific patient for a specific disease is constructed as a heterogeneous, multi-modal, multi-scale, and uncertain weight signal graph network cluster with time concept; finally, the above time sequence network cluster is carried out with strategic random walk resampling with time concept, and the large model online learning sample stream about the specific patient population and disease diagnosis and treatment experience is dynamically generated, and the large model optimization for multi-round diagnosis and treatment experience is carried out accordingly.

[0035] In one possible embodiment, the representation learning includes:

[0036] Multi-modal domain experience fusion is carried out through the degradation network.

[0037] Multi-round diagnosis and treatment experience modeling is carried out through the recurrent neural network.

[0038] Exemplarily, each modality degradation layer contains multiple dense layers. Each time access, the latent representation will pass through the degradation layer to reconstruct the original modality. For the case of partial modality missing, the learned representation is needed to reconstruct the available modality to ensure full use of multi-modal data. In order to use the latent representation to reconstruct the corresponding modality through the vth degradation network, the loss function for the vth degradation network is designed as

[0039]

[0040] where f v (·; Θ v ) is the degradation layer of the vth modality, the parameter is Θ v , including multiple dense layers, represents the vth modality data of the ith individual at the tth time point, represents the missing data; if the ith person has the vth modality data at the tth access, otherwise Note that the multi-modal fusion module includes V degradation networks, and the reconstruction loss is developed as follows:

[0041]

[0042] Based on the learned longitudinal representations, a sequence learning module based on a recurrent neural network is used to encode and capture the temporal relevance of these representations. Taking the i-th individual as an example, assume that the longitudinal latent representation {h1, h2, ..., h} has been obtained from the multimodal fusion module. T}.from Figure 1 As can be seen from this, for each access t, the hidden representation h is... t and expert rating vector y t Connect them together as the vertical input s t In order to solve s t To address the missing data issue, the "model imputation" method is used, based on the hidden state z. t-1 The predicted input s from the dense layer t ,Right now Specifically, firstly z t-1 Input to dense layer for prediction Among them W d This represents the weight parameters of the dense layer. Note the estimated value. Depends on the hidden state z t-1 This is unavailable at the first time point. If data from the first time point is missing, imputation will be performed using the average of all time points for all training individuals. Based on the estimated values... In the interpolation layer, for s t Perform element-wise interpolation: in s t A masking vector is used to mask the missing values ​​in the data, and its definition is as follows:

[0043]

[0044] Where, δ t,d and s t,d δ t and s t The d-th component in the data. Finally, the interpolated data... Unit state c t-1 and hidden state z t-1 The input is fed into the sequence learning unit to update the current state c. t and z t In the longitudinal data processing stage, minimize the fitting loss. as follows:

[0045]

[0046] Where |·| represents the absolute value of the entity, Let p represent the input vector of individual i at time point i, and g(.;Ω) represent the network including an interpolation layer, a sequence learning module, and a dense layer with parameter Ω.i,t is an indicator vector, p (i,j),t = 1 if the t-th feature of the i-th individual is available at the t-th time point, otherwise p (i,j),t = 0.

[0047] In one possible embodiment, the sampling comprises:

[0048] exploiting historical data through a dynamic replay strategy;

[0049] layered sampling through historical data.

[0050] The sampling further comprises controlling training complexity by limiting the number of sampled batches.

[0051] Exemplarily, the continuous fine-tuning of the large model aims to teach the large model to handle a series of tasks in a specific application scenario while using the data that is constantly emerging. The set of data patterns handled in the k-th task is denoted as and the corresponding training set and the evaluation set At the k-th step, given the visible training data the model is required to obtain satisfactory results on the evaluation set of the new task and the evaluation sets of all k-1 historical tasks, i.e. the pattern set Each example in the data set is composed of a set of multi-modal time-series domain experience data and a true value label. Then, the large model needs to predict the knowledge type label along the task sequence with stable satisfactory accuracy.

[0052] Based on the large model framework of low-rank autonomous adaptive learning, compared with the traditional method of relying on fixed memory sample replay, a more advanced dynamic replay strategy is adopted. In mathematical expression, it can be described as follows:

[0053] Suppose there is a full historical data set D, which contains a large number of samples x i ,y i , i = 1, 2,..., N, where N represents the total number of samples, x i represents the input feature, and y i is the corresponding label.

[0054] Through a dynamic sampling strategy, in each iteration t, a mini-batch data subset D t of size B t is selected from D, satisfying:

[0055]

[0056] where B tlimited by a pre-set maximum batch size to control the training complexity, and B t ≤B max .

[0057] To achieve more comprehensive data coverage and prevent overfitting and forgetting, a hierarchical robust sampling method is used to ensure the diversity and representativeness of the sampling. In this process, the model dynamically adjusts the sampling strategy according to the current learning state and historical gradient information, effectively mining key information from the overall historical data, and improving the model's generalization ability and learning efficiency.

[0058] In summary, the model optimization technology based on low-rank autonomous adaptive learning not only breaks the storage space constraint on data sampling, but also realizes the full and effective use of the full historical data under controllable training complexity through the innovative dynamic replay mechanism.

[0059] The limitation of the number of sampling batches is also introduced to control the training complexity, specifically, the storage space constraint is eliminated to improve the data coverage. On the contrary, the example of replay is dynamically selected from the full volume data under the limitation of the maximum batch size at each step, which ensures the complexity independent of the number of tasks.

[0060] For example, GPT-3175B, its training corpus contains about 300B tokens. Since a single English token contains an average of about 4 characters1, the total storage of its training corpus is about 1-2 TB. For most modern servers used in artificial intelligence research, such memory size is completely acceptable. Although the training data of recent large models is considered to occupy more space, due to the progress of storage hardware technology, the storage price has become moderate, generally no more than a few hundred yuan per TB. In addition, since most of the data sets for downstream tasks need to be filtered and annotated to ensure high quality and provide supervision for training, their size is difficult to reach the TB level, and usually requires much less storage space than the training corpus of large models.

[0061] During the process of full data storage and continuous learning, especially when facing a growing sequence of tasks, there are indeed problems of statistical inefficiency and data imbalance. To solve these problems, an effective method is to implement a hierarchical sampling strategy when revisiting historical data. The following is a detailed description of this strategy and its mathematical expression form:

[0062] Suppose there is a data set containing multiple historical tasks, each task T i has its own sub-data set D i , where i = 1, 2,..., N represents different task numbers, and N is the total number of tasks.

[0063] The core idea of ​​stratified sampling is to ensure that each historical task has balanced representativeness during training, rather than allowing random sampling to lead to some tasks being overrepresented while others are neglected. The specific steps are as follows:

[0064] 1. Task ID hierarchical selection:

[0065] First, according to the preset probability distribution P(T) i To independently extract the ID of an old task, for example, using a uniform probability distribution, each task has an equal probability of being selected.

[0066] 2. Sample selection within a subset of data:

[0067] Select task T i Then, from its subset D i A sample x is randomly selected from the sample. j , where j = 1, 2, ..., |D i | indicates that the j-th sample is in task T i The data set.

[0068] 3. Maintain the ratio of new to old samples:

[0069] Each training batch B contains a fixed ratio of new and old samples. Let r be the proportion of new samples. Then the batch size |B| contains r|B| samples from the current new task and (1-r)|B| stratified samples from the old task.

[0070] In the formula, it is assumed that there are M training batches in one iteration, and the number of new samples in each batch is m. r =r|B|, where the number of old samples is m o = (1-r)∣B∣. When selecting old samples, task T i The expected number of selections can be represented as E[Selections of T]. i ] = M·m o ·P(T i ).

[0071] Through this hierarchical sampling mechanism, the model can continuously review and balance the absorption of information from all historical tasks while adapting to new learning tasks, thereby improving learning efficiency and generalization ability when facing data imbalance problems. Especially in continuous learning and transfer learning scenarios, this strategy helps to mitigate catastrophic forgetting and ensures the retention of skills from past tasks.

[0072] Furthermore, a fixed ratio of new to old examples is maintained in each training batch. In this way, the dynamic selection of low-rank adaptation is more robust to the problem of data imbalance and statistically more efficient in terms of equal coverage of each historical task.

[0073] In one possible embodiment, the large model optimization based on the sampling and multi-round diagnostic experience further includes fine-tuning the large model using low-rank adaptive expert selection technology.

[0074] The low-rank adaptive expert selection technique also includes the technique of constructing task-specific low-rank adaptive expert predictors.

[0075] The specific task refers to a diagnostic or treatment task targeting a specific disease.

[0076] For example, to avoid the uncontrollable linear growth of training costs in previous dynamic architecture-based methods, a low-rank adaptation expert selection pre-selection is constructed to select a fixed number of the most important low-rank adaptation modules. Specifically, in step k, a low-rank adaptation expert selector is trained to select from... A fixed number t of task pattern sets are identified for each example. Then, the low-rank adaptation modules for the selected t pattern sets are retained as active modules and participate in subsequent inference.

[0077] Formally, given an input example x, the large model integrated with the low-rank adaptation module M first encodes x as its example representation, and then projects it onto the corresponding log vector through the linear head contained in M. This process can be represented as: s = f(M, x), where f and s represent the encoding function and the log vector, respectively. Specifically, the low-rank adaptation expert selector transforms each example x into a selection score vector s of size k. sel The j-th element represents x belonging to The confidence level. sel The first t elements are indexed as {i1,i2,...,i...} t} determines the low-rank adaptation module for t activities. The selection process involves a teacher-mandated strategy, whereby the correct low-rank adaptive module corresponding to the example is always selected for training. Through pre-selection, the forward propagation cost is unaffected by the number of low-rank adaptive modules.

[0078] To obtain the final prediction results, the activity module was integrated. The information learned in the process. Formally, each active module receives a logarithmic vector of its corresponding pattern set, while the logarithmic vectors of inactive modules (i.e., those unselected modules) are assigned the same vector with a sufficiently low value α. This process can be represented by the following formula:

[0079] s j =f(M j ,x),j∈{i1,i2,...,i t},

[0080]

[0081] pred = arg max[s1, s2,..., s k ].

[0082] where [·] denotes concatenation, s j is the logit vector produced by the jth low-rank adaptation, and pred denotes the predicted label.

[0083] In one possible embodiment, before the large model optimization according to the sampling and multi-round diagnosis and treatment experience, an auxiliary cache system is constructed, which is used to store the logs generated by the tuned low-rank adaptation expert selection technique.

[0084] To solve the problem of repeated log calculation, an auxiliary cache system is constructed to store the logs generated by the tuned low-rank adaptation module. Specifically, each cache entry includes an example ID, a low-rank adaptation module ID (the index of a specific task module or selector), and a log vector. When the adjusted low-rank adaptation module encounters an input example, the auxiliary cache system checks the database using the example ID and low-rank adaptation module ID. If there is a match, the time-consuming calculation of the equation s = f(M, x) can be avoided. The cache system can only be used after the specific module completes the tuning process, and will remain fixed thereafter. For example, the low-rank adaptation selector can only access the cache after the training process of the pre-selection stage is completed.

[0085] Working principle of auxiliary cache system and related formulas:

[0086] Suppose there is a set of low-rank adaptation modules M = M1, M2,..., M m , each module M i is responsible for processing a specific type of input and generating a corresponding log probability vector s i = f(M i , x), where x is the input example and f is the calculation function. To avoid recalculating the log probability every time the same input is encountered, an auxiliary cache system C is established:

[0087] C = (E j , I k , s ik ): i ∈ [1, m], j ∈ ExampleSet

[0088] Each cache entry here contains three parts:

[0089] Example ID E j : is a unique identifier that uniquely identifies the input example.

[0090] Low-rank adaptation module ID I k: is the index or identification of the module M k that generated this entry.

[0091] : is the log-probability vector s ik resulting from the processing of the example E k by the module M j .

[0092] When a tuned low-rank adaptive module M l needs to process a new or repeated input x t , the system performs the following steps:

[0093] : construct the query key Q t = (E t , I l ), where E t is the ID of the current example x t and I l is the ID of the module M l .

[0094] : retrieve the cache system C and check if there is a matching entry Q t :

[0095]

[0096] : if there is a matching entry, directly extract the log-probability vector s {lk} from the cache:

[0097] s lk = C[(E t , I l )] if (E t , I l ) ∈ C.

[0098] This avoids the time-consuming computation s lk = f(M l , x t ).

[0099] : if there is no matching entry, perform the computation and store the result in the cache:

[0100] S lk = f(M l , x t ),

[0101] and add the new entry to the cache system:

[0102] C ← C ∪ {(E t , I l , s lk}.

[0103] In addition, only when a certain low-rank adaptation module M l After the tuning training process is completed, its calculation results are allowed to be written into the cache system. This means that after the pre-selection phase ends and the module parameters are determined, the log probabilities of any examples processed by the module will be stored and retrieved according to the above strategy.

[0104] The low-rank autonomous self-adaptive learning-based model self-optimization paradigm also addresses the low scalability problem through the dynamic structure of task-specific low-rank adaptation modules and selectors. Unlike previous hybrid expert architectures, the low-rank autonomous self-adaptive learning-based model self-optimization paradigm is more suitable for large model adjustment, as each expert is a lightweight and highly pluggable low-rank adaptation module that can be adjusted without changing the original large model structure or parameters. Compared to traditional hybrid expert architectures, this method has the following advantages:

[0105] Lightweight and pluggable: Each low-rank adaptation module is an independent and lightweight component that can achieve local optimization and adjustment without significantly affecting the overall structure of the large model.

[0106] Dynamic structure: The low-rank adaptation selector can dynamically select and combine appropriate low-rank adaptation modules according to task requirements, improving the flexibility and efficiency of the model when processing diverse tasks.

[0107] In summary, this low-rank autonomous self-adaptive learning-based model self-optimization scheme can effectively cope with the challenges of data distribution changes in complex large model environments, while reducing computational cost and improving training efficiency.

[0108] Although preferred embodiments of the application have been described, those skilled in the art on obtaining the basic inventive concept can make additional changes and modifications to these embodiments. Therefore, the appended claims are intended to include the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0109] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A large model optimization method based on multi-round diagnosis and treatment experience learning, characterized in that, The method comprises the following steps: Performing variable-length multi-modal time series modeling and unified embedding representation learning on dynamic domain experience through a heterogeneous time-state multi-graph expression model; Constructing a disease course evolution as a heterogeneous, multi-modal, multi-scale, and uncertainty-weighted signal graph network cluster with a time concept; Performing a strategic random walk resampling with a time concept on the network cluster; Optimizing a large model according to the sampling and multi-round diagnosis and treatment experience.

2. The method of claim 1, wherein the method is based on multi-round diagnosis and treatment experience learning. The representation learning comprises the following steps: Fusing multi-modal domain experience through a degenerative network; Modeling multi-round diagnosis and treatment experience through a recurrent neural network.

3. The method of claim 1, wherein the method is based on multi-round diagnosis and treatment experience learning. The sampling comprises the following steps: Using historical data through a dynamic replay strategy; Performing hierarchical sampling through historical data.

4. The method of claim 3, wherein the method further comprises: The sampling further comprises limiting the number of sampling batches to control the training complexity.

5. The method of claim 1, wherein the method is based on multi-round diagnosis and treatment experience learning. The large model optimization according to the sampling and multi-round diagnosis and treatment experience further comprises fine-tuning the large model through a low-rank adaptive expert selection technology.

6. The method of claim 5, wherein the method further comprises: The low-rank adaptive expert selection technology further comprises constructing a specific task low-rank adaptive expert predictor technology.

7. The method of claim 6, wherein the method further comprises: The specific task is a diagnosis and treatment task for a specific disease.

8. The method of claim 6, wherein the method further comprises: Before the large model optimization according to the sampling and multi-round diagnosis and treatment experience, an auxiliary cache system is constructed.

9. The method of claim 8, wherein the method further comprises: The auxiliary cache system is used to store the logs generated by the tuned low-rank adaptive expert selection technology.