Large model training system and method based on multi-modal medical data fusion
By extracting features and optimizing attention scores on multimodal medical data, the problem of inaccurate attention resource allocation in the existing technology is solved, more efficient medical data fusion and model training are achieved, and the learning and reasoning performance of the model is improved.
Patent Information
- Application Number
- CN202510849286.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-24
AI Technical Summary
The existing multimodal feature fusion method based on attention mechanism cannot effectively distinguish the reliability and criticality of information in medical data processing, resulting in inaccurate allocation of attention resources, affecting model learning and reasoning performance.
By collecting multimodal medical data, feature extraction and pre-calculation are performed, the information distinction value of sample features is evaluated and the label correlation intensity is determined, modal consistency and significance analysis are performed, attention scores are optimized, and the final multimodal data feature fusion representation is generated.
The accuracy of multimodal fusion feature representation is improved, the efficiency and accuracy of the model to learn medical knowledge is enhanced, the interference of low-value or redundant information is suppressed, and the performance of the model on target tasks is improved.
Smart Images

Figure CN120372296A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a large model training system and method based on multi-modal medical data fusion. Background Art
[0002] When artificial intelligence is applied to the medical field, especially when dealing with tasks involving multiple information sources (such as medical images, text reports, electronic medical record records, etc.), multi-modal feature fusion is a fundamental and key technology. This technology aims to integrate the feature information extracted from each independent modal data through specific algorithms or strategies to generate a unified feature representation that can comprehensively reflect the sample information. This unified representation is the direct input of machine learning models (such as classifiers, segmentation models, generative models, etc.) for subsequent complex analysis, prediction, or decision-making. Currently, the common technical paths for realizing multi-modal feature fusion include: First, directly concatenating the feature vectors extracted from each modality in the dimension to form a longer combined vector; Second, performing element-level mathematical operations on the feature vectors of each modality, such as element-wise addition, multiplication, or taking the average; Third, adopting a method based on the attention mechanism, which can learn the importance weights of different modal features according to the data content and perform weighted summation on the modal features based on this weight to obtain a dynamically adjusted fusion representation. Among them, the fusion method based on the attention mechanism is considered a relatively advanced and effective technical approach because it can adaptively focus on more relevant modal information.
[0003] However, existing multi-modal feature fusion technologies, especially the applied attention mechanism-based methods, have a technical limitation when calculating the contribution degree of each modality to the final fusion representation. The attention mechanism usually generates the original attention scores by calculating the similarity between a query vector and the key vectors derived from the features of each modality, and this score is subsequently normalized to the final attention weights. This core mechanism that completely relies on the numerical similarity between vectors to determine the information contribution degree fails to fully consider the complex characteristics inherent in medical multi-modal data, namely the inherent reliability differences of data signals and the relative criticality differences of information content. On the one hand, the features of a modality may generate numerically strong feature signals due to noise, artifacts, or low-quality expressions in the original data (such as blurred regions in images, templatized or non-specific descriptions in reports). Such signals may accidentally generate high similarity scores with the query vector or the key vectors of other modalities, but they do not reflect real and valuable information. The standard attention score calculation process lacks an evaluation link for the reliability of such signal sources, resulting in its susceptibility to being misled by such "false" signals and assigning inappropriate high attention to low-quality or noise sources. On the other hand, even if the signal source is reliable, the information values carried by different modalities in a specific sample are not equal. A high similarity score may correspond to a common description of a general background information or a normal state (constituting information redundancy), or it may correspond to an accurate expression of a key and diagnostically discriminative pathological feature (constituting key information). The existing similarity-based attention score calculation mechanism itself cannot effectively distinguish the information value differences behind these two "high similarity" scenarios, thus possibly evenly distributing or wrongly focusing on redundant information for attention resources, rather than fully amplifying the truly key and decision-making significant modality information.
[0004] In summary, the core link in the existing attention fusion mechanism that only calculates attention scores based on vector similarity, due to its inability to simultaneously evaluate the reliability of signals and distinguish the criticality and redundancy of information, leads to insufficient accuracy and effectiveness of attention allocation. The finally generated fusion feature representation fails to optimally focus on real and key medical information, which directly limits the performance potential of downstream artificial intelligence systems that rely on this representation for learning and reasoning. Summary of the Invention
[0005] To solve the above technical problem of insufficient accuracy and effectiveness of attention allocation, the present invention aims to propose a large model training system and method based on multi-modal medical data fusion to improve the accuracy of weight evaluation of vectors in the attention mechanism.
[0006] To achieve the above object, the technical solution of the present invention is realized as follows: First aspect, the present application provides a large model training method based on multi-modal medical data fusion, and the method includes the following steps: Step S1: Collect multi-modal medical data, and perform feature extraction and pre-computation on the multi-modal medical data; Step S2: Through the evaluation and fusion analysis of the discrimination value of the sample feature information of the medical data and the association strength with the medical data label, obtain the context adaptability adjustment factor of the sample data; Step S3: Through the modal consistency analysis and modal significance analysis of the multi-modal medical data samples, obtain the information value deepening adjustment factor of the sample data; Step S4: Optimize and adjust the original attention score through the context adaptability adjustment factor and the information value deepening adjustment factor to obtain the final attention score; Step S5: Perform multi-modal data feature fusion through the final attention score and perform model training.
[0007] Further, according to the collection of multi-modal medical data and the feature extraction and pre-computation of the multi-modal medical data, it specifically includes: Collect multi-modal medical data samples in the medical information system and perform desensitization processing to construct a multi-modal medical data set; for any sample in the multi-modal medical data set, obtain its medical image data, text diagnosis report data, and target label data; perform standardized pre-processing on the medical image data and input it into a pre-trained deep convolutional neural network to extract image feature vectors; perform text pre-processing on the text diagnosis report data and input it into a pre-trained language model to extract text feature vectors; classify all samples according to the target label, and for any target category and any modality, calculate and store the modal feature mean vector and covariance matrix of the samples within the category; for any modality, calculate and store the average L2 norm of the feature vectors of the modality on all samples in the data set.
[0008] Further, according to the evaluation and fusion analysis of the discrimination value of the sample feature information of the medical data and the association strength with the medical data label to obtain the context adaptability adjustment factor of the sample data, it includes: Obtain the feature vector of the sample data in the multi-modal medical dataset and the category to which the feature vector of the sample data belongs; through intra-class difference analysis of the feature vector of the sample data, obtain the first typicality evaluation of the sample data; through inter-class center difference analysis of the feature vector of the sample data, obtain the first outlier separation evaluation of the sample data; through discrimination value analysis of the first typicality evaluation and the first outlier separation evaluation of the sample data, obtain the first discrimination degree of the sample data; through target label correlation analysis of the feature vector of the sample data, obtain the label correlation strength of the sample data; through fusion analysis of the first discrimination degree and the label correlation strength of the sample data, obtain the context adaptability adjustment factor of the sample data.
[0009] Further, according to the above-mentioned obtaining the first typicality evaluation of the sample data through intra-class difference analysis of the feature vector of the sample data; obtaining the first outlier separation evaluation of the sample data through inter-class center difference analysis of the feature vector of the sample data; obtaining the first discrimination degree of the sample data through discrimination value analysis of the first typicality evaluation and the first outlier separation evaluation of the sample data, specifically including: Obtain the feature vector of the target sample, and call the mean vector and covariance matrix of its target category, evaluate the multi-dimensional distance among the three, standardize the evaluation result according to the feature dimension to obtain the first typicality evaluation; compare the feature vector of the target sample with the mean vectors of all other target categories one by one, select the minimum distance and take its square value to obtain the first outlier separation evaluation; obtain the atypicality threshold of the target category, use the first outlier separation evaluation as the numerator and the sum of a constant one and the first typicality evaluation as the denominator to form a fraction as the first discrimination degree evaluation factor; map the difference between the first typicality evaluation and the threshold through the Sigmoid function to obtain the second discrimination degree evaluation factor; multiply the first discrimination degree evaluation factor and the second discrimination degree evaluation factor to obtain the first discrimination degree of the sample data.
[0010] Further, the above-mentioned obtaining the label correlation strength of the sample data through target label correlation analysis of the feature vector of the sample data, and obtaining the context adaptability adjustment factor of the sample data through fusion analysis of the first discrimination degree and the label correlation strength of the sample data, specifically including: Obtain the distances from the sample data to the mean vectors of the nearest heterogeneous categories and the distances to the mean vectors of its own category respectively, and determine the label association strength of the sample data through the ratio of the two; set the non - negative fusion weight coefficient of the first discrimination degree and the non - negative fusion weight coefficient of the label association strength, perform logarithmic mapping on the first discrimination degree and the label association strength respectively, and perform weighted summation according to the weight coefficients to obtain the first fusion evaluation; map the first fusion evaluation through the hyperbolic tangent function and output it as the context adaptability adjustment factor of the sample data.
[0011] Further, by performing modal consistency analysis and modal significance analysis on the multi - modal medical data samples, an information value deepening adjustment factor of the sample data is obtained, specifically including: Obtain the feature vectors of the sample data in the multi - modal medical data set, and use the average cosine similarity between the feature vectors of any dimension of the sample data and the feature vectors of all other dimensions of the sample data as the inter - modal similarity degree of the sample data; obtain the relative significance degree of the sample data by performing significance evaluation on the feature vectors of the sample data; obtain the information value deepening adjustment factor of the sample data by performing fusion evaluation on the inter - modal similarity degree and the relative significance degree of the sample data.
[0012] Further, the obtaining of the relative significance degree of the sample data by performing significance evaluation on the feature vectors of the sample data includes: Obtain the feature vectors of the sample data and the average L2 norm of the feature vectors of any modality of the sample data on all samples in the multi - modal medical data set; use the calculation result of dividing the L2 norm of the feature vectors of any modality of the sample data by the average L2 norm of the feature vectors of this modality on all samples in the multi - modal medical data set as the relative significance degree of the sample data.
[0013] Further, according to the obtaining of the information value deepening adjustment factor of the sample data by performing fusion evaluation on the inter - modal similarity degree and the relative significance degree of the sample data, it specifically includes: Set the adjustment weight of the relative significance influence; use the calculation result of adding the constant 1 and the inter - modal similarity degree of the sample data as the first adjustment factor; use the calculation result of multiplying the adjustment weight of the relative significance influence, the relative significance degree of the sample data and the inter - modal similarity degree of the sample data as the second adjustment factor; use the mapping result of performing hyperbolic tangent function mapping on the calculation result of adding the first adjustment factor and the second adjustment factor as the first mapping evaluation; use the calculation result of adding the constant 1 and the first mapping evaluation as the information value deepening adjustment factor of the sample data.
[0014] Further, the original attention score is optimized and adjusted according to the context adaptability adjustment factor and the information value deepening adjustment factor to obtain the final attention score, which specifically includes: Obtain the original attention score, context adaptability adjustment factor, and information value deepening adjustment factor of the sample data; use the calculation result of multiplying the context adaptability adjustment factor and the information value deepening adjustment factor as the optimization weight, and use the calculation result of multiplying the optimization weight and the original attention score as the final attention score.
[0015] In a second aspect, the present application provides a large model training system based on multi-modal medical data fusion, including: a processor and a memory, where the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a large model training method based on multi-modal medical data fusion is implemented.
[0016] Compared with the prior art, the present invention has the following advantages: The large model training system and method based on multi-modal medical data fusion according to the present invention collect multi-modal medical data, extract features and perform pre-calculation on the multi-modal medical data; through the fusion analysis of the information discrimination value evaluation and label association strength of the sample features, obtain the context adaptability adjustment factor of the sample data; through the modal consistency analysis and modal significance analysis of the multi-modal medical data samples, obtain the information value deepening adjustment factor of the sample data; optimize and adjust the original attention score through the context adaptability adjustment factor and the information value deepening adjustment factor to obtain the final attention score; perform multi-modal data feature fusion through the final attention score and perform model training, where the original attention score is optimized and adjusted through the context adaptability adjustment factor and the information value deepening adjustment factor and normalized through the standard Softmax function during the model training process to obtain the final attention weight that can accurately reflect the comprehensive contribution value of each modality, and the value vectors corresponding to each modality are weighted and combined according to the final attention weight to generate an optimized multi-modal fusion feature representation, so that the proportion of key non-redundant information is intelligently enhanced during the fusion process of the optimized multi-modal fusion feature representation, while suppressing the interference of low-value or redundant information. Therefore, compared with the feature representation generated by the prior art, it can more effectively encapsulate the core information of the sample, thereby improving the efficiency and accuracy of the model in learning medical knowledge and improving the final performance of the model trained by the training system on the target task. Description of the Drawings
[0017] In the drawings: Figure 1 It is a method flow chart of the large model training system and method based on multi-modal medical data fusion according to the embodiment of the present invention. Detailed Embodiments
[0018] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0019] See Figure 1 , which is a flowchart of a method for a large model training system and method based on multi-modal medical data fusion provided in the first embodiment of the present invention. As Figure 1 shown, the large model training system and method based on multi-modal medical data fusion may include: S1, collect multi-modal medical data, and perform feature extraction and pre-computation on the multi-modal medical data.
[0020] It should be noted that during the data collection process, relevant privacy protection regulations are strictly observed to thoroughly anonymize or pseudonymize all information related to patient identities.
[0021] First, collect medical data samples containing multiple information sources (multiple modalities) from the medical information system. For the th sample in the collected multi-source medical data, ensure the acquisition of the following core data: Medical image data : Digital image files such as CT (Computed Tomography), MRI (Magnetic Resonance Imaging), X-ray, ultrasound, etc.
[0022] Text diagnosis report : A text-format diagnosis report written by a radiologist or other clinician corresponding to the image data , containing descriptions of the image findings, measurement values, diagnostic opinions, etc.
[0023] Target label : The final analysis target or reference fact label corresponding to each sample i. This label is crucial for subsequent calculations and is used to define the category C ( ) to which the sample belongs. For example it can be the name of the disease diagnosed by pathology (such as "lung adenocarcinoma", "benign nodule"), the clinical diagnosis stage, the risk level.
[0024] After that, perform basic feature extraction and data space definition: Preprocess the collected raw data and extract the basic feature vectors of each modality. These feature vectors form the basis for subsequent calculations. All feature vectors exist in the feature spaces of their respective modalities.
[0025] Image feature extraction: For the image data Perform standard preprocessing and then input it into a pre-trained deep convolutional neural network (CNN) model (e.g., ResNet, EfficientNet, etc. pre-trained on large general image datasets or medical image datasets) to extract its deep visual features and obtain an image feature vector of a fixed dimension. All the image feature vectors constitute the feature space of the image modality.
[0026] Text feature extraction: Perform text preprocessing (word segmentation, stop word removal, conversion to lowercase) on the text diagnosis report and then input it into a pre-trained language model to extract its semantic representation and obtain a text feature vector of a fixed dimension. All the text feature vectors constitute the feature space of the text modality.
[0027] After that, perform pre-computation and storage of global statistical information: To support the calculation of adjustment factors in subsequent steps, the following global statistical information needs to be pre-computed and stored on the entire training dataset (including samples). These calculations are based on the feature spaces of each modality extracted in the basic feature extraction and data space definition: Category-related statistics: For each target category C (defined by all possible values) and each modality ( ), calculate and store the mean vector of the modality features of all samples in category C ; calculate and store the covariance matrix of the modality features of all samples in category C .
[0028] Modal global statistics: For each modality , calculate and store the average L2 norm of the feature vectors of this modality on the entire training dataset samples (for subsequent calculation of relative significance).
[0029] These pre-computed global statistics (category mean vector, category covariance matrix, modal average norm) will be used as known constants or lookup tables and called in subsequent steps.
[0030] Finally, perform data alignment confirmation: Confirm that for each sample , its data of different modalities ( , ), the corresponding target label and the association between the extracted basic feature vectors ( , ) is accurate.
[0031] After completing the above steps, the obtained dataset and pre-computed information can be used for subsequent optimization and fusion processing.
[0032] S2. By evaluating and fusing the discrimination value of the sample feature information of medical data and the association strength with medical data labels, a context adaptability adjustment factor for sample data is obtained.
[0033] Existing attention-based multimodal fusion methods focus on determining attention allocation by calculating the numerical similarity between query vectors and key vectors of each modality. Although this mechanism realizes data-driven weight adjustment, its limitation is that similarity itself is a relative and context-unaware metric, which cannot independently reflect the true information value and contribution potential of a modality feature in the specific analysis context of a current sample. Even when the data has been cleaned to exclude obvious errors and noise, this scoring method based on single vector similarity still faces problems in terms of the ambiguity of feature space representation and the non-linear relationship between similarity and information contribution.
[0034] Among them, in terms of the ambiguity of feature space representation, the feature vectors extracted from complex raw data (high-dimensional images, long reports) are a compressed expression of the original information. In this process, it is possible that different original signal patterns are mapped to similar positions in the feature space, or the importance differences of the same original signal in different contexts cannot be fully reflected in a single feature vector. This means that a high similarity score , the original information pattern corresponding to it and its actual meaning in the current sample analysis task are uncertain. For example, in the analysis of lung CT images, a feature pattern representing a small round shadow is common. When the analysis task is to identify early lung cancer, this pattern is a key signal that needs to be highly concerned, so its feature will have a high similarity with the query vector for finding suspicious lesions. However, in samples with routine screening or a consistent benign medical history, the same small round shadow feature pattern most likely represents common benign calcifications or scars, and its diagnostic importance is significantly reduced. Although the visual appearance and feature pattern are similar, resulting in a high numerical similarity between this pattern and the query vector, according to its working principle, the standard attention mechanism cannot and is not designed to be able to distinguish whether this shadow is a key suspected signal or common benign background information in the current specific context.
[0035] On the other hand, regarding the nonlinear relationship between similarity and information contribution, in medical diagnosis analysis, the role of information is often not linear. A feature that is highly similar to the query, if the information it represents is ubiquitous or known, such as a description of a normal anatomical structure, has a very low marginal contribution to the final decision. On the contrary, a feature that is not necessarily the most similar to the query but represents some differential information, such as an imaging sign or a suggestive differential diagnosis in a report, will have a higher marginal contribution. The standard attention score calculation is proportional to the similarity and cannot capture the nonlinear or even inversely correlated complex relationship between contribution and similarity. It tends to amplify highly similar information regardless of its true contribution value.
[0036] Therefore, relying solely on vector similarity to allocate attention will inevitably lead to deviations in the true value assessment of each modal information in the current specific sample analysis context. In response to the problem of true value assessment deviation, the present invention obtains a first typicality assessment of the sample data by performing intra-class difference analysis on the feature vectors of the sample data; obtains a first heterogeneous separation assessment of the sample data by performing heterogeneous center difference analysis on the feature vectors of the sample data; obtains a first degree of discrimination of the sample data by performing a differentiation value analysis on the first typicality assessment of the sample data and the first heterogeneous separation assessment; obtains the label association strength of the sample data by performing a target label association analysis on the feature vectors of the sample data; and obtains the contextual adaptability adjustment factor of the sample data by performing a fusion analysis on the first degree of discrimination of the sample data and the label association strength.
[0037] The present invention constructs context-adaptive adjustment factors for sample data, which no longer rely on judging the absolute "signal quality" or whether it conforms to the "universal prior" (to avoid accidental injury in rare cases), but focuses on evaluating the relative value of feature information in the current context: Evaluate the discrimination and specificity of information: Valuable information should help distinguish the current sample from other possible disease states. This requires evaluating the extent to which a modality feature is a unique representative of its category (rather than a common shared feature) and the degree of difference between it and other different categories. The present invention achieves this evaluation through the calculation of specific statistical indicators, aiming to identify those feature sources that are more likely to point to key diagnostic clues or rare manifestations.
[0038] The direct relevance of the evaluation information to the sample analysis goal: The ultimate value of a feature is reflected in its contribution to the achievement of the analysis goal. This invention quantifies the feature by directly comparing the geometric relationship between the feature and its category and the closest different category. and sample target labels The association strength. The distance information in the feature space is used to determine to what extent a feature clearly points to its true category, so as to evaluate its actual utility for the current task objective.
[0039] Among them, through the within-class difference analysis of the feature vectors of the sample data, the first typicality evaluation of the sample data is obtained; through the difference analysis of the heterogeneous centers of the feature vectors of the sample data, the first heterogeneous separation evaluation of the sample data is obtained; through the discrimination value analysis of the first typicality evaluation and the first heterogeneous separation evaluation of the sample data, the first discrimination degree of the sample data is obtained, specifically including: Obtain the mean vector and covariance matrix of the feature vector of the sample data corresponding to the target category of the feature vector; through the feature vector of the sample data, the mean vector of the target category corresponding to the feature vector of the sample data and the covariance matrix of the target category corresponding to the feature vector of the sample data, conduct the Mahalanobis distance evaluation between the feature vector and the category center to obtain the first Mahalanobis distance of the feature vector of the sample data; take the calculation result of dividing the square of the first Mahalanobis distance by the dimension of the feature vector of the sample data as the first typicality evaluation of the sample data; In one embodiment, assume the th sample and the th modality feature vector is , the covariance matrix of modality in category pre-calculated based on the training data set is , the mean vector of modality in category pre-calculated based on the training data set is , then the calculation formula for the first typicality evaluation of the th sample and the th modality is: Among them, represents the first typicality evaluation of the th sample and the th modality; represents the th sample and the th modality feature vector; represents the mean vector of modality in category pre-calculated based on the training data set; represents the covariance matrix of modality in category pre-calculated based on the training data set; represents the dimension of the feature vector , represents the transpose of a vector.
[0040] After that, by evaluating the distance between the eigenvector of the sample data and the mean vectors corresponding to all other target categories, the closest outlier mean vector of the sample data is obtained; the square of the Euclidean distance between the eigenvector of the sample data and the closest outlier mean vector is used as the first outlier separation evaluation of the sample data. Obtain the atypicality threshold within the target category of the sample data; use the first outlier separation evaluation of the sample data as the numerator, and the calculation result of adding the constant 1 and the first typicality evaluation of the sample data as the denominator to form a fraction as the first discrimination degree evaluation factor; map the calculation result of subtracting the first typicality evaluation of the sample data from the atypicality threshold within the target category of the sample data through the Sigmoid function, and the mapping result is used as the second discrimination degree evaluation factor; the calculation result of multiplying the first discrimination degree evaluation factor and the second discrimination degree evaluation factor is used as the first discrimination degree of the sample data.
[0041] In one embodiment, assume the th sample, and the first outlier separation evaluation of the th modality is ; the atypicality threshold within the target category of the sample data is , then the calculation formula for the first discrimination degree of the th sample in the th modality is: where represents the first discrimination degree of the th sample in the th modality; represents the first outlier separation evaluation of the th sample in the th modality; represents the first typicality evaluation of the th sample in the th modality; represents the atypicality threshold within the target category of the sample data; represents the activation function.
[0042] It should be noted that considering that in medical applications, only those features that can effectively distance themselves from different classes and show a certain uniqueness (not completely typical) within their own class are the most distinguishable. The above formula ensures that only when the features simultaneously meet the requirements of distancing from different classes and having sufficient uniqueness within their own class, and exceed the threshold, will their distinguishability increase significantly. In this embodiment, the initial value of the atypia threshold within the target class of the sample data is set to 1, and in actual training, the seventy-fifth percentile of the squared Mahalanobis distance distribution of the modal features of the corresponding class can be used for setting.
[0043] After evaluating the distinguishing value of the features, it is also necessary to measure their direct association strength with the current sample analysis task objective. The value of medical information ultimately lies in its contribution to achieving specific clinical objectives (such as accurate diagnosis). Therefore, the present invention calculates the association strength index between the features and the target label, and this index aims to quantify how much the features clearly point to their true class rather than other different classes. This is achieved by comparing the distance from the feature to the nearest different-class center with the distance to its own class center. The larger the distance ratio, the more clearly it belongs to its own class and the stronger its association with the target label.
[0044] Specifically, after obtaining the first degree of distinguishability of the sample data, the label association strength of the sample data can be obtained by performing target label association analysis on the feature vector of the sample data; by performing fusion analysis on the first degree of distinguishability of the sample data and the label association strength, the context adaptability adjustment factor of the sample data is obtained, including: Obtaining the mean vector of the nearest other target class corresponding to the feature vector of the sample data and the mean vector of the target class corresponding to the feature vector of the sample data; using the mean vector of the nearest other target class corresponding to the feature vector of the sample data as the mean vector of the nearest different-class cluster of the sample data; using the mean vector of the target class corresponding to the feature vector of the sample data as the class mean vector of the sample data; Taking the Euclidean distance between the feature vector of the sample data and the mean vector of the nearest different-class cluster of the sample data as the numerator, and taking the calculation result of the fraction formed by adding a minimum positive number to the Euclidean distance between the feature vector of the sample data and the class mean vector of the sample data as the denominator as the label association strength of the sample data; In one embodiment, assuming the th sample data, the mean vector of the nearest other target class corresponding to the feature vector of the th modality is then the formula for calculating the label association strength of the th sample data for the th modality is: Among them, represents the label association strength of the -th modality in the -th sample data; represents the feature vector of the -th modality in the -th sample; represents the mean vector of the nearest other target classes corresponding to the feature vector of the -th modality in the -th sample data; represents the mean vector of modality in class pre-computed based on the training data set; represents the Euclidean distance; represents a very small positive number used to prevent the denominator from being zero.
[0045] It should be noted that this formula directly calculates the ratio of the minimum inter-class distance to the intra-class distance. The higher the value, the closer the feature of modality is to the target for completing the current sample analysis task.
[0046] Obtain the set non-negative fusion weight coefficient of the first discrimination degree and the set non-negative fusion weight coefficient of the label association strength; use the result of logarithmic mapping of the calculation result of adding the constant 1 to the first discrimination degree of the sample data as the first logarithmic mapping result; use the result of logarithmic mapping of the calculation result of adding the constant 1 to the label association strength of the sample data as the second logarithmic mapping result; perform weighted summation calculations on the first logarithmic mapping result and the second logarithmic mapping result through the non-negative fusion weight coefficient of the first discrimination degree and the non-negative fusion weight coefficient of the label association strength respectively, and use the calculation result of the weighted summation as the first fusion evaluation; use the mapping result of the hyperbolic tangent function mapping of the first fusion evaluation as the context adaptability adjustment factor of the sample data.
[0047] In one embodiment, assume that the non-negative fusion weight coefficient of the first discrimination degree is ; the non-negative fusion weight coefficient of the label association strength is Then the calculation formula for the context adaptability adjustment factor of the -th modality in the -th sample data is: Among them, represents the context adaptability adjustment factor of the -th modality in the -th sample data; represents the first degree of discrimination non - negative fusion weight coefficient. In the embodiments of the present invention, the initial value of this weight coefficient is set to ; represents the th sample, the first degree of discrimination of the th modality; represents the label - association strength non - negative fusion weight coefficient. In the embodiments of the present invention, the initial value of this weight coefficient is set to ; represents the th sample data, the label - association strength of the th modality; represents the hyperbolic tangent function; represents the logarithmic function with base
[0048] It should be noted that the context - adaptability adjustment factor of the sample data designed in the present invention, its core objective is to make up for the defect that the existing attention mechanism only relies on vector similarity to calculate the attention score and cannot accurately measure the true contribution potential of information in the context of a specific sample. To achieve this goal, is constructed by closely combining the characteristics of medical data and the requirements of the analysis task, reflecting clear causal logic and problem - solving ideas.
[0049] Aiming at the problem that the existing attention mechanism is difficult to distinguish the information value behind high similarity, the present invention first evaluates the discrimination value of information by calculating the within - class relative density and between - class separation index of features. When calculating the within - class distance, the Mahalanobis distance is used, which can take into account the complex correlations between dimensions in medical feature data, making the measurement of feature typicality more accurate. Combining the Euclidean distance with the nearest outlier center directly associates the feature with the ability to distinguish different class states.
[0050] A regulatory term based on the Sigmoid function and threshold is introduced into the first degree of discrimination of the sample data. This design mimics the emphasis on information specificity in clinical judgment: only those information that can effectively distinguish from outliers and show a certain uniqueness (not completely typical) within their own category are considered to have high discrimination value. This refined quantification of the discrimination value effectively overcomes the limitation that information cannot be judged as discriminative only by similarity.
[0051] Aiming at the problem that the existing attention mechanism lacks task - orientation consideration, the present invention calculates the association strength index of the feature and the target label. Using the geometric relationship in the feature space to evaluate the feature and its true class label The degree of association tightness. By directly calculating the ratio of the Euclidean distance from a feature to the nearest outlier center to the Euclidean distance from the feature to its own class center, intuitively quantifies the extent to which the feature clearly points to its true class rather than other classes. A feature that is close to its own class center and far from the outlier center has a naturally higher value, indicating a strong association with the target label.
[0052] By adopting logarithmic transformation smoothing and combining weighted summation with an activation function in this way, we obtain , ensuring that the adjustment factor can comprehensively and balancedly reflect the discrimination value and task contribution of modal features in the current context.
[0053] So far, through the modal consistency analysis and modal significance analysis of samples, an information value deepening adjustment factor for sample modalities is obtained.
[0054] S3. By conducting modal consistency analysis and modal significance analysis on multi-modal medical data samples, an information value deepening adjustment factor for the sample data is obtained.
[0055] After calculating the context adaptability adjustment factor of the sample data through step S2 , a preliminary assessment of the expected contribution value of the sample data in a specific scenario (based on its discrimination ability and task relevance) is obtained. This step improves the understanding of the value of a single modality in attention allocation, but the fundamental defect of the existing attention mechanism has not been completely solved, that is, the mutual relationship between different information sources (modalities) has not been fully considered, especially the degree of consistency in their content and the contrast between the strengths of the signals themselves, which is crucial for finally determining the optimal attention weights.
[0056] Specifically, even if the feature vector of the sample data of a certain modality is rated as having a relatively high expected contribution value (that is, has a high value), the information content it conveys may be highly similar to that of other modalities with relatively high value. In medical practice, images, reports, and medical records often describe the same disease condition or physiological state from different aspects, and the information not only corroborates each other but also has repetitions. If we only increase its attention score based on the independent expected contribution value of each modality, it is easy to give excessive cumulative attention to this repetitive information, which instead dilutes the attention allocated to other modalities that may contain unique perspectives or key supplementary information. On the other hand, although the initially evaluated contribution value of a modality is It is not the highest, but its information content is significantly different from other modalities, or its signal strength is much higher than other modalities although the content is similar, which indicates that this modality may play a particularly important role in the current sample.
[0057] Therefore, in order to compensate for After the adjustment, the problem of inaccurate attention allocation due to failure to fully consider the consistency and relative strength between modalities still exists. The present invention introduces an information value deepening adjustment factor. The evaluation is further based on the modality The attention contribution of the modality can be more finely adjusted based on the consistency of the information content with other modalities and the relative strength of its own signal.
[0058] The idea of constructing the information value deepening modifier is to measure these two key interactive characteristics: One is to evaluate the consistency of modal content: by calculating the modal The average similarity between the features of the modality and the features of all other modalities directly understands the extent to which the modality information overlaps or coordinates with other source information. High consistency means that information is repeated; low consistency implies that there are unique differences.
[0059] The second is to evaluate the relative strength of the sample data: by comparing the modes The signal strength of the feature and its average strength level in the entire data set determine whether the modality is particularly prominent or abnormally weak in the current sample. Modalities with high signal strength often convey clearer information, regardless of whether their content is consistent with other modalities, and their strength itself provides additional judgment clues.
[0060] Although the sample data context adaptation adjustment factor The problem that the existing attention mechanism cannot evaluate the information distinguishing value and task association strength has been initially solved. Correcting the raw attention score The fundamental limitation of calculating scores based on similarity alone has not been completely overcome. The information of the corresponding modality after the improvement of attention score may be comparable to other modal information with high attention score. The ratings have the same modality, resulting in information redundancy; or Modes with moderate values provide key complementary or conflicting information with other modes. To solve the suboptimal attention allocation problem caused by ignoring the interaction between modalities after adjustment, the present invention further designs an information value deepening adjustment factor, which aims to The consistency level with other modal information and its own relative significance have a great impact on the The attention contribution potential after preliminary adjustment is recalibrated to ultimately achieve the purpose of preferentially highlighting key non-redundant information.
[0061] The specific steps for obtaining the information value deepening adjustment factor of sample data include obtaining the feature vectors of sample data in a multi-modal medical dataset, obtaining the inter-modal similarity degree of the sample data by performing inter-modal consistency analysis on the feature vectors of the sample data; obtaining the relative significance degree of the sample data by performing significance evaluation on the feature vectors of the sample data; and obtaining the information value deepening adjustment factor of the sample data by performing fusion evaluation on the inter-modal similarity degree and the relative significance degree of the sample data.
[0062] First, by performing inter-modal consistency analysis on the feature vectors of the sample data, the inter-modal similarity degree of the sample data is obtained. Specifically, the feature vectors of the sample data are obtained, and the average cosine similarity between the feature vectors of any dimension of the sample data and the feature vectors of all other dimensions of the sample data is used as the inter-modal similarity degree of the sample data.
[0063] After that, by performing significance evaluation on the feature vectors of the sample data, the relative significance degree of the sample data is obtained. Specifically, the feature vectors of the sample data are obtained, and the average L2 norm of the feature vectors of any modality of the sample data on all samples in the multi-modal medical dataset; the calculation result of dividing the L2 norm of the feature vectors of any modality of the sample data by the average L2 norm of the feature vectors of this modality on all samples in the multi-modal medical dataset is used as the relative significance degree of the sample data.
[0064] Finally, by performing fusion evaluation on the inter-modal similarity degree and the relative significance degree of the sample data, the information value deepening adjustment factor of the sample data is obtained. Specifically, the set adjustment weight of the relative significance impact is obtained; the calculation result of adding the constant 1 to the inter-modal similarity degree of the sample data is used as the first adjustment factor; the calculation result of multiplying the adjustment weight of the relative significance impact, the relative significance degree of the sample data, and the inter-modal similarity degree of the sample data is used as the second adjustment factor; the mapping result of performing hyperbolic tangent function mapping on the calculation result of adding the first adjustment factor and the second adjustment factor is used as the first mapping evaluation; the calculation result of adding the constant 1 to the first mapping evaluation is used as the information value deepening adjustment factor of the sample data.
[0065] In one embodiment, assume the th sample data, the th modality, the inter-modal similarity degree is ; the th sample data, the th modality, the relative significance degree is ; the adjustment weight for the relative significance impact is , then for the th sample data, the calculation expression for the information value deepening adjustment factor of the th modality is: Wherein, represents the information value deepening adjustment factor of the th sample data for the th modality; represents the similarity degree between modalities of the th sample data for the th modality; represents the relative significance degree of the th sample data for the th modality; represents the adjustment weight for the relative significance impact; represents the hyperbolic tangent function; represents the constant 1.
[0066] It should be noted that for the problem of information redundancy that may still exist after preliminary adjustment, the present invention quantifies the consistency between modalities by calculating the average cosine similarity . This index provides a direct clue for identifying modalities with high overlap with information from other sources. To distinguish the signal strength, a relative significance measure is introduced, and using the normalized L2 norm can effectively capture the prominence of the current modality signal strength relative to the average level. By the ( ) term, the weight of information that is inconsistent ( low) with other modalities is explicitly increased, solving the problem that the key supplementary or conflicting information cannot be fully highlighted after adjustment. By the term, the weight of information that is both consistent ( high value) and significant ( high value) is retained and enhanced, which avoids over-suppressing all consistent information and solves the bias caused by only considering inconsistency.
[0067] Thus far, through the modality consistency analysis and modality significance analysis of the multi-modal medical data samples, the information value deepening adjustment factor of the sample data is obtained.
[0068] S4. Optimally adjust the original attention score through the context adaptability adjustment factor and the information value deepening adjustment factor to obtain the final attention score.
[0069] After obtaining the adjustment factor reflecting the context adaptability of the sample dataand a regulation factor reflecting the deepening of information value After that, these two factors need to be further applied to the existing attention mechanism-based feature fusion method. Specifically, the calculation link of the original attention score is corrected by these two factors. This correction aims to integrate the evaluation results of the potential contribution of the present invention to the modality into the final determination process of the attention weight, so as to overcome the limitation that the original mechanism only depends on vector similarity.
[0070] This step is executed within the standard attention fusion framework. First, according to the standard process of the existing attention mechanism, for the current sample and the query vector , the key vector corresponding to the basic feature of each modality and the value vector are calculated, and the original attention score between the query vector and the key vector is calculated (through dot product similarity ). Subsequently, the present invention corrects the original attention score and are sequentially applied to the original score in a two-step product regulation manner to obtain the finally regulated attention score , and the specific calculation is shown in the following formula: where represents the final attention score of the th modality in the th sample; represents the original attention score of the th modality in the th sample; represents the context adaptability regulation factor of the th modality in the th sample data; represents the information value deepening regulation factor of the th modality in the th sample data.
[0071] So far, the original attention score is optimized and regulated by the context adaptability regulation factor and the information value deepening regulation factor to obtain the final attention score.
[0072] S5. Perform multi-modal data feature fusion through the final attention score and perform model training.
[0073] After calculating the attention scores adjusted for each modality, this step completes the final feature fusion and applies the result to the training of a machine learning model. The adjusted attention scores are normalized through the standard Softmax function to obtain the final attention weights that can accurately reflect the comprehensive contribution value of each modality. Based on these optimized weights, the value vectors corresponding to each modality are weighted and combined to generate an optimized multi-modal fusion feature representation. This optimized representation can more effectively encapsulate the core information of the sample compared to the representations generated by existing technologies because it has intelligently enhanced the proportion of key and non-redundant information during the fusion process while suppressing the interference of low-value or redundant information. Finally, in the training system of the present invention, the optimized fusion feature representation generated in this way is used as input data to train a machine learning model that needs to process multi-modal medical data. Training with the optimized feature representation obtained by adopting the method of the present invention aims to improve the efficiency and accuracy of the model in learning medical knowledge, thereby improving the final performance of the model trained by this training system on the target task.
[0074] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A large model training method based on multimodal medical data fusion, characterized in that, The method includes the following steps: Step S1: Collect multi-modal medical data, and perform feature extraction and pre-computation on the multi-modal medical data; Step S2: Evaluate and fuse the association strength between the value of the sample feature information of the medical data and the medical data label to obtain the context adaptability adjustment factor of the sample data; Step S3: Perform modal consistency analysis and modal significance analysis on the multi-modal medical data samples to obtain the information value deepening adjustment factor of the sample data; Step S4: Optimize and adjust the original attention score through the context adaptability adjustment factor and the information value deepening adjustment factor to obtain the final attention score; Step S5: Perform multi-modal data feature fusion through the final attention score and conduct model training.
2. The large model training method based on multimodal medical data fusion according to claim 1, wherein, According to the collection of multi-modal medical data and the feature extraction and pre-computation of the multi-modal medical data, it specifically includes: Collect multi-modal medical data samples in the medical information system and perform desensitization processing to construct a multi-modal medical data set; for any sample in the multi-modal medical data set, obtain its medical image data, text diagnosis report data, and target label data; perform standardized pre-processing on the medical image data and input it into a pre-trained deep convolutional neural network to extract image feature vectors; perform text pre-processing on the text diagnosis report data and input it into a pre-trained language model to extract text feature vectors; classify all samples according to the target label, and for any target category and any modality, calculate and store the modal feature mean vector and covariance matrix of the samples within the category; for any modality, calculate and store the average L2 norm of the feature vectors of the modality on all samples in the data set.
3. The large model training method based on multimodal medical data fusion according to claim 1, wherein, According to the evaluation and fusion analysis of the association strength between the value of the sample feature information of the medical data and the medical data label to obtain the context adaptability adjustment factor of the sample data, it includes: Obtain the feature vector of the sample data in the multi-modal medical data set and the category to which the sample data feature vector belongs; perform intra-class difference analysis on the feature vector of the sample data to obtain the first typicality evaluation of the sample data; perform different-class center difference analysis on the feature vector of the sample data to obtain the first different-class separation evaluation of the sample data; perform discrimination value analysis on the first typicality evaluation and the first different-class separation evaluation of the sample data to obtain the first discrimination degree of the sample data; perform target label association analysis on the feature vector of the sample data to obtain the label association strength of the sample data; perform fusion analysis on the first discrimination degree and the label association strength of the sample data to obtain the context adaptability adjustment factor of the sample data.
4. The large model training method based on multi-modal medical data fusion according to claim 3, characterized in that, According to the intra-class difference analysis of the feature vector of the sample data to obtain the first typicality evaluation of the sample data; perform different-class center difference analysis on the feature vector of the sample data to obtain the first different-class separation evaluation of the sample data; Perform discrimination value analysis on the first typicality evaluation and the first different-class separation evaluation of the sample data to obtain the first discrimination degree of the sample data, specifically including: Obtain the feature vector of the target sample, call the mean vector and covariance matrix of its target category, evaluate the multi-dimensional distance among the three, standardize the evaluation result according to the feature dimension to obtain the first typicality evaluation; compare the feature vector of the target sample with the mean vectors of all other target categories one by one, select the minimum distance and take its square value to obtain the first outlier separation evaluation; obtain the atypicality threshold of the target category, and form a fraction with the first outlier separation evaluation as the numerator and the sum of a constant one and the first typicality evaluation as the denominator as the first discrimination degree evaluation factor; map the difference between the first typicality evaluation and the threshold through the Sigmoid function to obtain the second discrimination degree evaluation factor; take the product of the first discrimination degree evaluation factor and the second discrimination degree evaluation factor as the first discrimination degree of the sample data.
5. The large model training method based on multi-modal medical data fusion according to claim 3, wherein, By performing target label correlation analysis on the feature vector of the sample data to obtain the label correlation strength of the sample data, and by performing fusion analysis on the first discrimination degree and the label correlation strength of the sample data to obtain the context adaptability adjustment factor of the sample data, specifically including: Respectively obtain the distance from the feature vector of the sample data to the mean vector of the nearest outlier category and the distance to the mean vector of its own category, and determine the label correlation strength of the sample data through the ratio of the two; set the non-negative fusion weight coefficient of the first discrimination degree and the non-negative fusion weight coefficient of the label correlation strength, perform logarithmic mapping on the first discrimination degree and the label correlation strength respectively, and perform weighted summation according to the weight coefficient to obtain the first fusion evaluation; take the result of the hyperbolic tangent mapping of the first fusion evaluation as the context adaptability adjustment factor of the sample data.
6. The large model training method based on multimodal medical data fusion according to claim 1, wherein By performing modal consistency analysis and modal significance analysis on the multi-modal medical data sample to obtain the information value deepening adjustment factor of the sample data, specifically including: Obtain the feature vector of the sample data in the multi-modal medical data set, and use the average cosine similarity between the feature vector of any dimension of the sample data and the feature vectors of all other dimensions of the sample data as the inter-modal similarity degree of the sample data; perform significance evaluation on the feature vector of the sample data to obtain the relative significance degree of the sample data; perform fusion evaluation on the inter-modal similarity degree and the relative significance degree of the sample data to obtain the information value deepening adjustment factor of the sample data.
7. The large model training method based on multimodal medical data fusion according to claim 6, wherein By performing significance evaluation on the feature vector of the sample data to obtain the relative significance degree of the sample data, including: Obtain the average L2 norm of the feature vector of the sample data and the feature vectors of all samples in any modality of the sample data in the multi-modal medical data set; take the ratio of the L2 norm of the feature vector of any modality of the sample data to the average L2 norm of the feature vectors of all samples in this modality in the multi-modal medical data set as the relative significance degree of the sample data.
8. The large model training method based on multi-modal medical data fusion according to claim 6, wherein According to the above, by performing fusion evaluation on the inter-modal similarity degree and the relative significance degree of the sample data to obtain the information value deepening adjustment factor of the sample data, specifically including: Set the adjustment weight for the relative salience impact; use the calculation result of adding the constant 1 to the inter-modal similarity degree of the sample data as the first adjustment factor; use the calculation result of multiplying the adjustment weight of the relative salience impact, the relative salience degree of the sample data, and the inter-modal similarity degree of the present data as the second adjustment factor; use the mapping result of applying the hyperbolic tangent function mapping to the calculation result of adding the first adjustment factor and the second adjustment factor as the first mapping evaluation; use the calculation result of adding the constant 1 to the first mapping evaluation as the information value deepening adjustment factor for the sample data.
9. The large model training method based on multimodal medical data fusion according to claim 1, wherein Optimize and adjust the original attention score according to the context adaptability adjustment factor and the information value deepening adjustment factor to obtain the final attention score, specifically including: Obtain the original attention score, the context adaptability adjustment factor, and the information value deepening adjustment factor of the sample data; use the product of the context adaptability adjustment factor and the information value deepening adjustment factor as the optimization weight, and use the calculation result of multiplying the optimization weight by the original attention score as the final attention score.
10. A large model training system based on multimodal medical data fusion, characterized in that, Including: A processor and a memory, where the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the large model training method based on multi-modal medical data fusion according to any one of claims 1-9 is implemented.
Citation Information
Patent Citations
Image scene classification method and device based on target semantics and attention mechanism
CN111104898A
Text classification method, system and equipment based on multi-label association and medium
CN118227790A
Radar moving target detection method based on small sample transfer learning and attention mechanism
CN119250118A
Machine language large model construction method and system
CN119721118A
Cited By
Multi-modal data quality evaluation and optimization method and system of Internet of Vehicles
CN120804608A
Method and device for establishing diagnosis and treatment system of digestive system disease multi-modal information
CN120998466A
Internet-based intelligent household electrical appliance fault monitoring, collecting and processing system
CN121785295A
An internet-based intelligent household appliance fault monitoring, collecting and processing system
CN121785295B