Medical data processing method and system based on model distillation and medium

By combining cross-attention and symmetric attention mechanisms with multiple loss functions, the problems of multimodal feature reconstruction and cross-modal correlation knowledge transfer are solved, which improves the diagnostic accuracy and robustness of small models in medical data processing and is suitable for a variety of medical data combinations.

CN120806051APending Publication Date: 2025-10-17NORTH CHINA DIGITAL HEALTH TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510666631.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-10-17

Smart Images

  • Figure CN120806051A_ABST
    Figure CN120806051A_ABST
Patent Text Reader

Abstract

The invention discloses a medical data processing method and system based on model distillation and a medium, mainly relates to the technical field of medical data processing, and is used for solving the problems that an existing distillation model is difficult to effectively balance the relationship among multi-modal feature reconstruction, attention distribution matching and classification task optimization, and in addition, distillation is carried out only for a single modal, so that the efficiency is low. And cross-modal association knowledge cannot be inherited. Comprising the following steps: acquiring hidden data of a middle layer of a trained first model, and mapping the hidden data to a middle layer of a second model according to a preset layer number matching relationship; migrating the first attention weight matrix and the second attention weight matrix to a second model; configuring a classification cross entropy hyper-parameter, an intermediate layer MSE hyper-parameter and a weight matrix KL divergence hyper-parameter of the second model to obtain a preliminary second model; and training the preliminary second model by using the visual feature sequence and the word vector sequence to obtain a trained second model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of medical data processing technology, and in particular to a medical data processing method, system and medium based on model distillation. Background Art

[0002] In the field of artificial intelligence, especially healthcare, deep learning-based models have been widely used in medical imaging, disease prediction, clinical diagnosis, and other fields. Large models, often due to their large number of parameters and high computational resource requirements, are difficult to deploy and apply in low-resource environments. Therefore, compressing large models into smaller ones (through techniques such as model distillation) has become a common optimization method. Model distillation is a method that transfers knowledge from a large model to a smaller one, significantly improving the performance of the smaller model.

[0003] While distilled models demonstrate good performance on some tasks, existing techniques generally employ a single loss function (such as categorical cross-entropy) or a simple linear combination of loss weights, making it difficult to effectively balance multimodal feature reconstruction, attention distribution matching, and classification task optimization. Furthermore, medical data is inherently multimodal (images + text), but existing data distillation models focus solely on a single modality, preventing small models from inheriting cross-modal knowledge.

[0004] Therefore, there is an urgent need for a medical data processing method, system and medium based on model distillation to solve the above technical problems. Summary of the Invention

[0005] The present application provides a medical data processing method, system and medium based on model distillation to solve the problem that existing distillation models are difficult to effectively balance the relationship between multimodal feature reconstruction, attention distribution matching and classification task optimization. In addition, the existing data model distillation only distills a single modality, resulting in the inability of small models to inherit cross-modal correlation knowledge.

[0006] In a first aspect, the present application provides a medical data processing method based on model distillation, the method comprising: Encode the imaging data into a sequence of visual features and the clinical report text into a sequence of word vectors; Using cross attention, we obtain the first attention weight matrix from the visual feature sequence to the word vector sequence; using symmetric attention, we obtain the second attention weight matrix from the word vector sequence to the visual feature sequence; Obtain a trained first model using the visual feature sequence, the word vector sequence, the first attention weight matrix and the second attention weight matrix; obtaining hidden data of an intermediate layer of the trained first model, mapping the hidden data to an intermediate layer of a second model according to a preset number-of-layers matching relationship; migrating the first attention weight matrix and the second attention weight matrix to the second model; configuring the second model with respect to a classification cross-entropy hyperparameter, an intermediate layer MSE hyperparameter, and a weight matrix KL divergence hyperparameter, to obtain a preliminary second model; training the preliminary second model by using the visual feature sequence and the word vector sequence, to obtain a trained second model.

[0007] In an implementation manner of the present application, the image data is encoded into a visual feature sequence, specifically including: The image data is down-sampled from 512*512 to 32*32 in stages according to resolutions, to obtain four-stage sampling data; and each stage of the sampling data is input into a corresponding encoder to output data features of different granularities; The visual feature sequence of the current stage is calculated by the formula: represents the i-th stage, and i [1, 3], represents the data feature of the i-th stage, The visual feature sequence of the current stage is calculated by the formula:

[0008] In an implementation manner of the present application, cross-attention is used to obtain a first attention weight matrix from the visual feature sequence to the word vector sequence; and symmetric attention is used to obtain a second attention weight matrix from the word vector sequence to the visual feature sequence, specifically including: obtaining a visual feature sequence V and a word vector sequence T; mapping the visual feature sequence as Query1 and the word vector sequence as Key1 and Value1; and further obtaining: represents the dimension of the visual feature sequence, represents the dimension of the word vector sequence, and d represents a preset alignment dimension; represents linear projection of Query1, represents linear projection of Key1, represents linear projection of Value1; ​​​​​​​​​​​​​, the first attention weight matrix is obtained by calculation; mapping the word vector sequence into Query2, mapping the word vector sequence into Key2 and Value2; further obtaining: , , ; wherein, , , , linear projection of Query2, linear projection of Key2, linear projection of Value2; by the formula: the second attention weight matrix is obtained by calculation.

[0009] In an implementation form of the present application, the hidden data of the intermediate layer of the trained first model is obtained, and the hidden data is mapped to the intermediate layer of the second model according to the preset number of layers matching relationship, specifically comprising: intercepting the output of the preset specified intermediate layer when the first model is forward propagated; According to the preset number of layers matching relationship, the output of the preset specified intermediate layer is projected to the corresponding intermediate layer of the second model, and the alignment loss is calculated; By minimizing the alignment loss, the intermediate layer MSE of the second model is obtained.

[0010] In an implementation form of the present application, the first attention weight matrix and the second attention weight matrix are migrated to the second model; the second model is configured with classification cross entropy super parameter, intermediate layer MSE super parameter and weight matrix KL divergence super parameter, to obtain a preliminary second model, specifically comprising: After migrating the first attention weight matrix and the second attention weight matrix to the second model, the attention distribution of the first model and the second model is obtained, and then the minimum weight matrix KL divergence super parameter corresponding to the minimum weight matrix KL divergence loss function is obtained by the weight matrix KL divergence loss function composed of the attention distribution of the first model and the second model, and the super parameter.

[0011] In a second aspect, the present application provides a medical data processing system based on model distillation, comprising: The encoding module is configured to encode the image data into a visual feature sequence and encode the clinical report text into a word vector sequence; The weight module is configured to obtain a first attention weight matrix from the visual feature sequence to the word vector sequence by using cross attention, and obtain a second attention weight matrix from the word vector sequence to the visual feature sequence by using symmetric attention; The training module is configured to obtain a trained first model by using the visual feature sequence, the word vector sequence, the first attention weight matrix, and the second attention weight matrix; obtain hidden data of an intermediate layer of the trained first model, map the hidden data to an intermediate layer of a second model according to a preset layer number matching relationship; migrate the first attention weight matrix and the second attention weight matrix to the second model; configure the second model with a classification cross-entropy hyperparameter, an intermediate layer MSE hyperparameter, and a weight matrix KL divergence hyperparameter to obtain a preliminary second model; and train the preliminary second model by using the visual feature sequence and the word vector sequence to obtain a trained second model.

[0012] In an implementation manner of the present application, the encoding module comprises an encoding unit configured to downsample the image data from 512x512 to 32x32 in stages to obtain four levels of sampling data; and input each level of sampling data into a corresponding encoder to output data features of different granularities. The visual feature sequence of the current level is calculated by the formula: , wherein i represents the i-th level, and i [1, 3], represents the data feature of the i-th level, the visual feature sequence of the i-th level is itself.

[0013] In an implementation manner of the present application, the training module comprises an intermediate layer unit configured to intercept the output of a preset designated intermediate layer when the first model is forward propagated; project the output of the preset designated intermediate layer to an intermediate layer of a corresponding second model according to a preset layer number matching relationship, and calculate an alignment loss; adjust the intermediate layer MSE of the second model by minimizing the alignment loss.

[0014] In an implementation manner of the present application, the training module comprises a matrix parameter unit configured to obtain the attention distribution of the first model and the second model after migrating the first attention weight matrix and the second attention weight matrix to the second model, and then obtain the minimum weight matrix KL divergence hyperparameter corresponding to the minimum weight matrix KL divergence loss function by using the weight matrix KL divergence loss function composed of the attention distribution of the first model and the second model and the hyperparameters.

[0015] In a third aspect, the present application provides a non-volatile computer storage medium having computer instructions stored thereon, the computer instructions being executed to implement a medical data processing method based on model distillation according to any one of the above aspects. ​​​

[0016] From the above technical solutions, the present application has the following advantages: 1. Realize the collaborative optimization of multi-modal features and attention mechanism: ‌Balance multi-task relationship: By introducing the hyperparameter configuration of three loss functions of classification cross-entropy, intermediate layer MSE (mean square error) and weight matrix KL divergence, the problem of difficult balance of multi-modal feature reconstruction, attention distribution matching and classification task optimization in traditional distillation is solved. This multi-objective optimization mechanism ensures that the model considers feature alignment, knowledge transfer and final classification performance during training, avoiding performance deviation caused by single task dominance.

[0017] ‌Cross-modal related knowledge transfer: Use cross-attention (visual to text) and symmetric attention (text to visual) mechanisms to generate bidirectional attention weight matrices and transfer them to the student model (second model), so that the student model can inherit the cross-modal related knowledge in the first model (first model). This design breaks through the limitations of single modal distillation and effectively solves the problem of performance degradation of small models due to lack of cross-modal interaction.

[0018] 2. Improve the hierarchical knowledge transfer to improve model efficiency: ‌Intermediate layer hidden data mapping: By mapping the intermediate layer hidden data of the first model to the student model according to the pre-set number of layers matching relationship, the transfer of deep feature representation is realized. Compared with the traditional method of only transferring the output layer, this strategy makes the student model more accurately mimic the internal representation of the first model, especially suitable for complex multi-modal features in medical data (such as the implicit association between images and text).

[0019] ‌Structured transfer of attention weights: Directly transfer the first and second attention weight matrices, which preserves the key attention patterns of the first model on multi-modal data (such as the correspondence between the visual features of the lesion area and the text description), improving the student model's ability to analyze medical data.

[0020] 3. Improve the practical value in medical scenarios: ‌Improve the diagnostic reliability of small models: In resource-constrained medical scenarios, the student model (lightweight version) can more accurately combine image and text information for joint reasoning by inheriting cross-modal knowledge, reducing the risk of misdiagnosis. For example, in the diagnosis task of combining pathological images and clinical reports, the student model can use visual features and text semantics simultaneously to simulate the comprehensive judgment process of expert doctors.

[0021] Enhanced model generalization: Multi-modal joint training and attention matching mechanism reduce the overfitting tendency of the model to a single data modality, especially suitable for the common modality missing or imbalance problem in medical data (such as some cases only have image or text data), which improves the robustness of the model in practical application.

[0022] 4. Improved scalability of technical solutions: Flexible adaptation to different modalities: This application is not limited to image and text, and can be extended to other multi-modal medical data (such as gene sequence and electronic health record), which can be adapted to new modalities by adjusting the encoder and attention mechanism, providing a general framework for future multi-modal fusion of medical AI.

[0023] Hyperparameter configurability: By dynamically adjusting the loss weights of classification, feature reconstruction and attention transfer, the model focus can be optimized for different medical tasks (such as disease classification, prognosis prediction), enhancing the scene adaptability of the method. BRIEF DESCRIPTION OF DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed to be used in the description, obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0025] Figure 1 is a medical data processing method flowchart based on model distillation provided by an embodiment of the present application.

[0026] Figure 2 is a medical data processing system internal structure schematic diagram based on model distillation provided by an embodiment of the present application. DETAILED DESCRIPTION

[0027] The technical solutions in the embodiments of the present application will be described in detail below with reference to the drawings in the embodiments of the present application, obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0028] Those skilled in the art should understand that the embodiments described below are only preferred embodiments of the present disclosure, and do not represent the only way the present disclosure can be implemented. The preferred embodiments are only used to explain the technical principles of the present disclosure, and are not intended to limit the protection scope of the present disclosure. Based on the preferred embodiments provided in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative labor should fall within the protection scope of the present disclosure.

[0029] It should also be noted that the terms "comprising", "containing" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus that includes a list of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.

[0030] The technical solutions of the embodiments of the present application will be described in detail below with reference to the drawings.

[0031] The embodiments provide a medical data processing method based on model distillation, as shown in Figure 1 The method provided by the embodiments of the present application mainly includes the following steps: Step 110, encode the image data into a visual feature sequence and encode the clinical report text into a word vector sequence.

[0032] In some embodiments, the image data is encoded into a visual feature sequence, specifically including: The image data is down-sampled from 512x512 to 32x32 in stages according to the resolution, to obtain four levels of sampling data; the sampling data of each level is input into the corresponding encoder to output data features containing different granularities; The visual feature sequence of the current level is calculated by the formula: , represents the i-th level, and i∈[1,3], represents the data feature of the i-th level, The visual feature sequence of the i-th level is It is itself.

[0033] ​​​It should be noted that by gradually downsampling the image data from 512x512 to 32x32 (four levels of sampling), the model can simultaneously extract features of different granularities. For example, high-resolution levels (such as 512x512) retain local details (such as the texture of small lesions), while low-resolution levels (such as 32x32) focus on global structures (such as organ morphology). This multi-scale design makes the model adapt to the detection needs of different size targets in medical images (such as small calcification points and overall organ lesions). By fusing the data features of the i-th level (i∈[1, 3]), the features of different levels are dynamically weighted or spliced to form the final visual feature sequence. For example, high-level features may enhance semantic information (such as lesion type), and low-level features supplement spatial details (such as lesion boundary). This fusion mechanism avoids information loss at a single resolution and improves the comprehensive representation ability of complex medical images.

[0034] Step 120, obtaining a first attention weight matrix of the visual feature sequence to the word vector sequence by cross attention, and obtaining a second attention weight matrix of the word vector sequence to the visual feature sequence by symmetric attention.

[0035] It should be noted that this step can be specifically: obtaining a visual feature sequence V and a word vector sequence T; Mapping the visual feature sequence to Query1, and mapping the word vector sequence to Key1 and Value1; Further obtaining: , , ; Wherein, , , , and represents the dimension of the visual feature sequence, represents the dimension of the word vector sequence, and d represents a preset alignment dimension; represents the linear projection of Query1, represents the linear projection of Key1, represents the linear projection of Value1; The first attention weight matrix is calculated by the formula: Mapping the word vector sequence to Query2, and mapping the word vector sequence to Key2 and Value2; Further obtaining: , , ; Wherein, , , , ​linear projection of Query2, linear projection of Key2, linear projection of Value2. By formula: The second attention weight matrix is calculated.

[0036] It should be noted that the cross-attention of vision→text: by mapping the visual feature sequence to Query1, the text sequence to Key1 and Value1, the model can learn the dependency of the image features on the text description. For example, the lesion area in the CT image (visual feature) will focus on the keywords such as “tumor size” or “edge blur” in the clinical report, realizing fine-grained alignment of vision to text.

[0037] Symmetric attention of text→vision: mapping the text sequence to Query2 and the visual feature sequence to Key2 and Value2 (if the user description is ambiguous, it needs to be corrected to the visual feature) in reverse, the model can capture the guiding effect of the text description on the image features. For example, the description of “pulmonary nodule” in the text will strengthen the model’s attention to the corresponding area in the image, forming a two-way complement.

[0038] Two-way attention coordination: the two attention mechanisms form a closed-loop interaction, simulating the two-way reasoning process of the doctor “looking at the picture and writing the report” and “checking the picture according to the text”, improving the accuracy of multi-modal joint diagnosis.

[0039] Step 130, using the visual feature sequence, the word vector sequence, the first attention weight matrix and the second attention weight matrix, obtaining a trained first model.

[0040] Step 140, obtaining hidden data of an intermediate layer of the trained first model, mapping the hidden data to an intermediate layer of a second model according to a preset number of layers matching relationship; migrating the first attention weight matrix and the second attention weight matrix to the second model; configuring the second model about classification cross-entropy hyperparameters, intermediate layer MSE hyperparameters and weight matrix KL divergence hyperparameters, obtaining a preliminary second model.

[0041] In some embodiments, obtaining hidden data of an intermediate layer of the trained first model, mapping the hidden data to an intermediate layer of a second model according to a preset number of layers matching relationship, specifically includes: When the first model is forward propagated, intercepting the output of the preset designated intermediate layer; According to the preset number of layers matching relationship, projecting the output of the preset designated intermediate layer to the corresponding intermediate layer of the second model, calculating the alignment loss; By minimizing the alignment loss, adjusting the intermediate layer MSE of the second model.

[0042] In some embodiments, the first attention weight matrix and the second attention weight matrix are migrated to the second model; the second model is configured with a classification cross-entropy super parameter, an intermediate layer MSE super parameter and a weight matrix KL divergence super parameter to obtain a preliminary second model, specifically comprising: After migrating the first attention weight matrix and the second attention weight matrix to the second model, the attention distribution of the first model and the second model is obtained, and then the minimum weight matrix KL divergence super parameter corresponding to the minimum weight matrix KL divergence loss function is obtained through the weight matrix KL divergence loss function composed of the attention distribution of the first model and the second model and the super parameter.

[0043] It should be noted that by intercepting the hidden data of the preset intermediate layer of the teacher model (the first model) and projecting and mapping to the corresponding layer of the student model (the second model), the migration of deep feature representation is realized. For example, the lesion edge features (intermediate layer output) encoded in the teacher model can directly guide the learning of similar features in the student model, avoiding shallow knowledge migration caused by relying only on the final output. The preset number of layers matching relationship allows flexible adaptation to the structural differences between the teacher and student models (such as the teacher model being ResNet-50 and the student model being MobileNet). For example, the 10th layer features of the teacher model are mapped to the 5th layer of the student model, and the dimension mismatch problem is solved through a projection matrix (such as linear transformation), ensuring the effectiveness of cross-architecture knowledge transfer. By minimizing the intermediate layer MSE loss, the student model is forced to approximate the teacher model in the intermediate layer output, mimicking its internal feature distribution and improving the student model's analysis accuracy for medical multi-modal data. The bidirectional attention weight matrix (first and second matrices) of the teacher model is directly migrated to the student model, preserving the cross-modal association pattern. For example, the attention weight of the teacher model "pulmonary nodule image features ->'spiculation' text description" is reused by the student model, and the cross-modal association can be established without retraining, significantly reducing the training cost. By comparing the attention distribution difference between the teacher and student models through the KL divergence loss function (such as ), the weight matrix KL divergence super parameter is dynamically optimized. For example, when the attention distribution of the student model deviates from the key area (such as the tumor location) of the teacher model, the KL divergence loss will adjust the super parameter to strengthen the alignment, ensuring the fidelity of cross-modal knowledge transfer. By configuring the classification cross-entropy super parameter (optimizing the diagnostic accuracy), the intermediate layer MSE super parameter (feature alignment) and the weight matrix KL divergence super parameter (attention distribution matching), the balanced optimization of the classification task and knowledge transfer is realized. For example, in the early stage of training, the MSE loss is focused on to stabilize the feature alignment, and in the later stage, the classification cross-entropy weight is increased to improve the diagnostic performance, avoiding model bias caused by single loss dominance.

[0044] Step 150, training a preliminary second model by using the visual feature sequence and the word vector sequence, and obtaining a trained second model.

[0045] In addition, the present application Figure 2 A medical data processing system based on model distillation is provided for an embodiment of the present application. Figure 2 As shown in the system provided by the embodiment of the present application, the system mainly comprises: The encoding module 210 is configured to encode the image data into a visual feature sequence and encode the clinical report text into a word vector sequence.

[0046] The encoding module 210 comprises an encoding unit configured to downsample the image data from 512x512 to 32x32 in stages according to the resolution, to obtain four levels of sampling data; and input each level of sampling data into a corresponding encoder to output data features of different granularities. The visual feature sequence of the current level is calculated by the formula: , wherein i represents the i-th level, and i is an element in the interval [1, 3], , wherein represents the data feature of the i-th level, , and the visual feature sequence of the i-th level is itself.

[0047] The weight module 220 is configured to obtain a first attention weight matrix from the visual feature sequence to the word vector sequence by using cross attention, and obtain a second attention weight matrix from the word vector sequence to the visual feature sequence by using symmetric attention.

[0048] The training module 230 is configured to obtain a trained first model by using the visual feature sequence, the word vector sequence, the first attention weight matrix, and the second attention weight matrix; map hidden data of an intermediate layer of the trained first model to an intermediate layer of a second model according to a preset number-of-layers matching relationship; migrate the first attention weight matrix and the second attention weight matrix to the second model; configure the second model with respect to a classification cross-entropy hyperparameter, an intermediate layer MSE hyperparameter, and a weight matrix KL divergence hyperparameter to obtain a preliminary second model; train the preliminary second model by using the visual feature sequence and the word vector sequence to obtain a trained second model.

[0049] The training module 230 comprises an intermediate layer unit configured to intercept the output of a preset designated intermediate layer when the first model is forward propagated; project the output of the preset designated intermediate layer to an intermediate layer of the corresponding second model according to the preset number-of-layers matching relationship, and calculate an alignment loss; ​​​​​The intermediate layer MSE of the second model is adjusted by minimizing the alignment loss.

[0050] The training module 230 comprises a matrix parameter unit configured to obtain attention distributions of the first model and the second model after the first attention weight matrix and the second attention weight matrix are migrated to the second model, and further obtain a minimum weight matrix KL divergence hyperparameter corresponding to a minimum weight matrix KL divergence loss function by a weight matrix KL divergence loss function composed of the attention distributions of the first model and the second model and the hyperparameter.

[0051] In addition, the embodiment of the present application further provides a nonvolatile computer storage medium, which has executable instructions stored thereon, and the executable instructions, when executed, realize the medical data processing method based on model distillation.

[0052] The above description of disclosed embodiments enables those skilled in the art to carry out or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A medical data processing method based on model distillation, characterized in that: The method comprises: Encode the imaging data into a sequence of visual features and the clinical report text into a sequence of word vectors; Using cross attention, we obtain the first attention weight matrix from the visual feature sequence to the word vector sequence; using symmetric attention, we obtain the second attention weight matrix from the word vector sequence to the visual feature sequence; Obtain a trained first model using the visual feature sequence, the word vector sequence, the first attention weight matrix and the second attention weight matrix; Obtain the hidden data of the middle layer of the trained first model, and map the hidden data to the middle layer of the second model according to the preset layer matching relationship; migrate the first attention weight matrix and the second attention weight matrix to the second model; configure the second model with respect to the classification cross entropy hyperparameter, the middle layer MSE hyperparameter, and the weight matrix KL divergence hyperparameter to obtain a preliminary second model; The visual feature sequence and the word vector sequence are used to train a preliminary second model to obtain a trained second model.

2. The medical data processing method based on model distillation according to claim 1, characterized in that: Encode image data into a sequence of visual features, including: The image data is downsampled from 512×512 to 32×32 in order to obtain four levels of sampling data. Each level of sampling data is input into the corresponding encoder to output data features with different granularities. By formula: , calculate the visual feature sequence of the current level ; in, , represents the i-th level, and i∈[1,3], represents the data features of level i, Visual feature sequence For itself.

3. The medical data processing method based on model distillation according to claim 1, characterized in that: Using cross attention, we obtain the first attention weight matrix from the visual feature sequence to the word vector sequence; using symmetric attention, we obtain the second attention weight matrix from the word vector sequence to the visual feature sequence: specifically, Get the visual feature sequence V and word vector sequence T; Map the visual feature sequence to Query1 and the word vector sequence to Key1 and Value1; And then get: , , ; in, , , ,and represents the dimension of the visual feature sequence, Represents the dimension of the word vector sequence, and d represents the preset alignment dimension; represents the linear projection of Query1, represents the linear projection of Key1, Represents the linear projection of Value1; By formula: , calculate and obtain the first attention weight matrix; Map the word vector sequence to Query2, and the word vector sequence to Key2 and Value2; And then get: , , ; in, , , , represents the linear projection of Query2, represents the linear projection of Key2, Represents the linear projection of Value2; By formula: , calculate and obtain the second attention weight matrix.

4. The medical data processing method based on model distillation according to claim 1, characterized in that: Obtain the hidden data of the middle layer of the trained first model and map the hidden data to the middle layer of the second model according to the preset layer matching relationship, specifically including: During the forward propagation of the first model, the output of the preset intermediate layer is intercepted; According to the preset layer matching relationship, the output of the preset specified intermediate layer is projected to the corresponding intermediate layer of the second model, and the alignment loss is calculated; The middle layer MSE of the second model is obtained by adjusting the minimization of the alignment loss.

5. The medical data processing method based on model distillation according to claim 4, characterized in that: Migrating the first attention weight matrix and the second attention weight matrix to the second model; Configure the second model's classification cross entropy hyperparameters, intermediate layer MSE hyperparameters, and weight matrix KL divergence hyperparameters to obtain a preliminary second model, specifically including: After migrating the first attention weight matrix and the second attention weight matrix to the second model, the attention distribution of the first model and the second model is obtained, and then the minimum weight matrix KL divergence hyperparameters corresponding to the minimum weight matrix KL divergence loss function are obtained through the weight matrix KL divergence loss function composed of the attention distribution and hyperparameters of the first model and the second model.

6. A medical data processing system based on model distillation, characterized in that: The system comprises: The encoding module is used to encode the image data into a sequence of visual features and the clinical report text into a sequence of word vectors; A weight module is used to obtain a first attention weight matrix from a visual feature sequence to a word vector sequence using cross attention; and a second attention weight matrix from a word vector sequence to a visual feature sequence using symmetric attention; The training module is used to obtain a trained first model using a visual feature sequence, a word vector sequence, a first attention weight matrix and a second attention weight matrix; obtain the hidden data of the middle layer of the trained first model, and map the hidden data to the middle layer of the second model according to a preset layer matching relationship; migrate the first attention weight matrix and the second attention weight matrix to the second model; configure the second model with respect to the classification cross entropy hyperparameter, the middle layer MSE hyperparameter and the weight matrix KL divergence hyperparameter to obtain a preliminary second model; use the visual feature sequence and the word vector sequence to train the preliminary second model to obtain a trained second model.

7. The medical data processing system based on model distillation according to claim 6, characterized in that The encoding module includes an encoding unit for downsampling the image data from 512×512 to 32×32 in accordance with the resolution to obtain four levels of sampling data; inputting the sampling data at each level into the corresponding encoder to output data features containing different granularities; By formula: , calculate the visual feature sequence of the current level ; in, , represents the i-th level, and i∈[1,3], represents the data features of level i, Visual feature sequence For itself.

8. The medical data processing system based on model distillation according to claim 6, characterized in that: The training module includes an intermediate layer unit, which is used to intercept the output of a preset intermediate layer during the forward propagation of the first model; According to the preset layer matching relationship, the output of the preset specified intermediate layer is projected to the corresponding intermediate layer of the second model, and the alignment loss is calculated; The middle layer MSE of the second model is obtained by adjusting the minimization of the alignment loss.

9. The medical data processing system based on model distillation according to claim 8, characterized in that: The training module includes a matrix parameter unit, which is used to obtain the attention distribution of the first model and the second model after migrating the first attention weight matrix and the second attention weight matrix to the second model, and then obtain the minimum weight matrix KL divergence hyperparameters corresponding to the minimum weight matrix KL divergence loss function through the weight matrix KL divergence loss function composed of the attention distribution and hyperparameters of the first model and the second model.

10. A non-volatile computer storage medium, characterized in that Computer instructions are stored thereon, and when the computer instructions are executed, they implement a medical data processing method based on model distillation as described in any one of claims 1 to 5.