Medical image lesion recognition method and device based on pre-trained large model
By compressing the pre-trained large model, the problems of high computational complexity and poor real-time performance caused by large model parameters are solved, and efficient adaptation and high-precision lesion recognition are achieved in the medical field.
Patent Information
- Application Number
- CN202510230379.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
AI Technical Summary
The pre-trained large model has large parameters and cannot be adapted to the medical field, resulting in high complexity in model calculation and poor real-time performance.
Generate a lightweight compressed model by model compression of pre-trained large models, including updating the output layer, fine-tuning training, structural compression and structural adjustment.
While maintaining high diagnostic accuracy, it significantly reduces the computing volume and storage needs of model, and adapts to the resource limitations of medical robot equipment.
Smart Images

Figure CN120163781A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical image recognition technology. Specifically, it relates to a method and device for medical image lesion recognition based on a pre-trained large model. Background Art
[0002] Large models have an extremely large number of parameters and can capture complex patterns and features, thus enabling the generation of high-quality content and accurate medical image lesion recognition.
[0003] How to adapt the pre-trained Transformer model with a large number of parameters to the field of medical robots to solve the problems of high model computational complexity and poor real-time performance in medical image diagnosis tasks. Summary of the Invention
[0004] The problem solved by this application is that the pre-trained large model has a large number of parameters and cannot be adapted to the medical field.
[0005] To solve the above problems, the first aspect of this application provides a method for medical image lesion recognition based on a pre-trained large model, including:
[0006] Preprocess the obtained medical image to obtain a corresponding input sequence;
[0007] Obtain a pre-trained large model;
[0008] Based on the input sequence, compress the pre-trained large model to obtain a compressed model;
[0009] Input the medical image to be recognized into the compressed model to obtain the lesion recognition result of the medical image.
[0010] The second aspect of this application provides a device for medical image lesion recognition based on a pre-trained large model, which includes:
[0011] An image preprocessing module, which is used to preprocess the obtained medical image to obtain a corresponding input sequence;
[0012] A model acquisition module, which is used to obtain a pre-trained large model;
[0013] A model compression module, which is used to compress the pre-trained large model based on the input sequence to obtain a compressed model;
[0014] A lesion recognition module, which is used to input the medical image to be recognized into the compressed model to obtain the lesion recognition result of the medical image.
[0015] The third aspect of this application provides an electronic device, which includes: a memory and a processor;
[0016] The memory is used to store programs;
[0017] The processor, coupled to the memory, is configured to execute the programs for:
[0018] Preprocess the acquired medical images to obtain corresponding input sequences;
[0019] Obtain a pre-trained large model;
[0020] Based on the input sequences, compress the pre-trained large model to obtain a compressed model;
[0021] Input the medical images to be recognized into the compressed model to obtain the lesion recognition results of the medical images.
[0022] A fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the above-mentioned method for recognizing medical image lesions based on a pre-trained large model.
[0023] In the present application, through compression technology, while maintaining high diagnostic accuracy, the model computation amount and storage requirements are significantly reduced, adapting to the resource limitations of medical robot devices. Description of the Drawings
[0024] Figure 1 It is a flowchart of the method for recognizing medical image lesions based on a pre-trained large model according to an embodiment of the present application;
[0025] Figure 2 It is a structural block diagram of the device for recognizing medical image lesions based on a pre-trained large model according to an embodiment of the present application;
[0026] Figure 3 It is a structural block diagram of an electronic device according to an embodiment of the present application. Detailed Embodiments
[0027] To make the above objects, features, and advantages of the present application more obvious and understandable, the following detailed description of the specific embodiments of the present application is provided in conjunction with the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0028] It should be noted that unless otherwise specified, the technical terms or scientific terms used in the present application should have the ordinary meaning understood by those skilled in the art to which the present application belongs.
[0029] In view of the above problems, the present application provides a new medical image lesion recognition solution based on a pre-trained large model, which can compress the model and solve the problem that the pre-trained large model has a large number of parameters and cannot be adapted to the medical field.
[0030] An embodiment of the present application provides a medical image lesion recognition method based on a pre-trained large model. The specific solution of this method is Figure 1 as shown. This method can be executed by a medical image lesion recognition device based on a pre-trained large model. The medical image lesion recognition device based on a pre-trained large model can be integrated into electronic devices such as computers, servers, computers, server clusters, and data centers. Combining Figure 1 as shown, it is a flowchart of a medical image lesion recognition method based on a pre-trained large model according to an embodiment of the present application; wherein, the medical image lesion recognition method based on a pre-trained large model includes:
[0031] S101, preprocess the obtained medical image to obtain a corresponding input sequence;
[0032] In the present application, the medical image (such as CT, MRI or X-ray) is adjusted to a fixed size (such as 224×224 pixels) to ensure input consistency; the pixel values are normalized, and the original pixel values (such as 0~255) are mapped to the range of [0,1] to facilitate model processing.
[0033] In the present application, the adjusted image is segmented into a sequence of squares with a preset pixel size (such as 16×16 pixels); each square is converted into an embedding vector through linear mapping to generate an input sequence.
[0034] In the present application, data augmentation is also performed on the training set images, including random horizontal flipping (probability 50%), rotation (range of ±15°), and brightness adjustment (range of ±20%).
[0035] S102, obtain a pre-trained large model;
[0036] In the present application, the pre-trained large model is a 24-layer Transformer architecture, and the pre-training task is image classification or natural language processing; the model weights are loaded from a public pre-trained model library (such as Hugging Face or PyTorch Hub).
[0037] S103, based on the input sequence, compress the pre-trained large model to obtain a compressed model;
[0038] S104, input the medical image to be recognized into the compressed model to obtain the lesion recognition result of the medical image.
[0039] In this application, the medical image to be recognized is preprocessed into an input sequence and input into the compression model; the model outputs the lesion recognition results (such as lesion type and lesion area).
[0040] In this application, through compression technology, while maintaining high diagnostic accuracy, the model calculation amount and storage requirements are significantly reduced to adapt to the resource limitations of medical robot devices.
[0041] In one implementation, S101 preprocesses the obtained medical image to obtain the corresponding input sequence, including:
[0042] Adjust the medical image to a fixed size and normalize the pixel values;
[0043] Segment the image into a sequence of blocks of preset pixels and generate an embedding vector through linear mapping, and the embedding vector is the input sequence.
[0044] In this application, the bilinear interpolation method is used to adjust the medical image to 224×224 pixels; the pixel values are normalized.
[0045] In this application, the image is segmented into a sequence of 16×16 pixel blocks, each block is flattened into a 256-dimensional vector; the 256-dimensional vector is converted into a 768-dimensional embedding vector through linear mapping.
[0046] In one implementation, the pre-trained large model consists of 24 layers of Transformer architecture.
[0047] In this application, the pre-trained large model is a 24-layer Transformer architecture, each layer contains a multi-head self-attention mechanism and a feed-forward neural network; the model embedding dimension is 768, and the number of attention heads is 12.
[0048] In one implementation, S103 compresses the pre-trained large model based on the input sequence to obtain a compressed model, including:
[0049] Update the output layer of the pre-trained large model;
[0050] Based on the input sequence, perform fine-tuning training on the updated pre-trained large model;
[0051] Perform structural compression and structural adjustment on the fine-tuned pre-trained large model;
[0052] Perform secondary training on the adjusted pre-trained large model to obtain the compressed model.
[0053] In this application, the output layer of the pre-trained large model is replaced with a fully connected layer that is consistent with the number of categories of the medical lesion recognition task (such as 2 categories: benign / malignant); based on the medical image input sequence, the updated pre-trained large model is fine-tuned, using the AdamW optimizer, with the learning rate set to 0.0001, and trained for 10 epochs.
[0054] In this application, pruning and structural adjustment are performed on the fine-tuned model. The mean absolute value of the gradients output by each layer of the Transformer is calculated, and the 6 layers with the lowest contribution are removed; the L1 norm of the weight matrix of each attention head is statistically calculated, and the 2 heads with the smallest norm in each layer are removed; an intermediate module (which can be composed of two fully connected layers and reduced to 192 dimensions) is inserted between the remaining adjacent layers.
[0055] In this application, the adjusted model is retrained. The structural parameters of the Transformer are frozen, and only the intermediate module is trained, with the learning rate set to 0.00001 to optimize the parameters of the intermediate module.
[0056] In one implementation manner, the fine-tuning training of the updated pre-trained large model based on the input sequence includes:
[0057] Freeze the parameters of the preset layers in front of the pre-trained large model;
[0058] Based on the input sequence, fine-tuning training is performed on the pre-trained large model after freezing; the learning rate of the optimizer in the fine-tuning training is set to 0.0001.
[0059] In this application, the parameters are frozen: the parameters of the first 18 layers of the Transformer are frozen, and only the last 6 layers and the output layer are trained; optimizer settings: use the AdamW optimizer, with the learning rate set to 0.0001; loss function: adopt the cross-entropy loss function to calculate the difference between the model output and the true label.
[0060] In this application, the purpose of freezing the parameters is to retain the knowledge of the pre-trained model in general tasks and avoid overfitting.
[0061] In this application, the parameters of the last 6 layers of the Transformer and the output layer parameters are unfrozen so that they can be adjusted according to the medical image data;
[0062] In this application, the medical image dataset is divided into a training set and a validation set (such as 80% for training and 20% for validation); each training batch contains 32 images, and it is trained for 10 epochs; after each epoch ends, the validation set is used to evaluate the model performance, and the model weights with the lowest validation loss are saved.
[0063] In this application, the knowledge of the pre-trained model on general tasks is retained by freezing the parameters of the first 18 layers; the model is adapted to the medical image lesion recognition task by fine-tuning the last 6 layers and the output layer.
[0064] In this application, the AdamW optimizer and the cross-entropy loss function are used to achieve fast convergence and high-precision recognition.
[0065] In this application, by freezing the parameters of the first 18 layers, fine-tuning the last 6 layers and the output layer, and combining the AdamW optimizer and the cross-entropy loss function, efficient fine-tuning training is achieved, and a high-precision model adapted to the medical image lesion recognition task is generated.
[0066] In one implementation, the structural compression and structural adjustment of the fine-tuned pre-trained large model include:
[0067] Calculating the mean of the absolute values of the gradients output by each layer of the Transformer structure;
[0068] Statistically calculating the norm of the weight matrix of each attention head in the Transformer structure;
[0069] Removing several layers of the Transformer structure with the smallest mean of the absolute values of the gradients, and removing several attention heads with the smallest norm of the weight matrix in the remaining Transformer structure;
[0070] Inserting intermediate modules between adjacent remaining layers of the Transformer structure.
[0071] In this application, calculating the mean of the absolute values of the gradients: calculating the mean of the absolute values of the gradients output by each layer of the Transformer to evaluate the contribution degree of each layer.
[0072] In this application, statistically calculating the weight norm of the attention heads: statistically calculating the L1 norm of the weight matrix of each attention head to evaluate the importance of each head.
[0073] In this application, removing redundant structures and attention heads: removing 6 layers of the Transformer structure with the smallest mean of the absolute values of the gradients; removing 2 attention heads with the smallest norm of the weight matrix in each layer.
[0074] In this application, sorting according to the mean of the absolute values of the gradients and removing several layers (such as 6 layers) with the lowest contribution degree; for example, if the original model has 24 layers, 18 layers are retained after removal.
[0075] In this application, in each remaining layer of the Transformer, sorting according to the norm of the weight matrix and removing several attention heads with the lowest importance (such as removing 2 heads in each layer); for example, if each layer originally has 8 heads, 6 heads are retained after removal.
[0076] In this application, by removing redundant layers and attention heads, the number of model parameters and the amount of computation are significantly reduced;
[0077] In this application, an intermediate module is inserted: an intermediate module is inserted between adjacent layers that are retained.
[0078] In this application, an intermediate module is inserted between the retained adjacent layer Transformer structures to enhance the model's expressive ability; for example, 17 intermediate modules are inserted in an 18-layer Transformer.
[0079] In this application, inserting an intermediate module enhances the model's expressive ability and improves the accuracy of lesion recognition;
[0080] In one implementation, the intermediate module can be composed of two fully connected layers and is dimension-reduced to 192 dimensions.
[0081] In another implementation, the structure and processing process of the intermediate module include:
[0082] The input feature map is divided into blocks to obtain independent blocks;
[0083] For each independent block, first neighborhood blocks and second neighborhood blocks with different spacings are obtained;
[0084] Based on the independent block and the first neighborhood blocks, a first feature block is generated;
[0085] Based on the independent block and the second neighborhood blocks, a second feature block is generated;
[0086] Feature compression is performed on the first feature block and the second feature block to obtain a compressed block;
[0087] All independent blocks are traversed, and an output feature map is generated based on the obtained compressed blocks.
[0088] In this application, dividing the input feature map into blocks means dividing the input feature map into corresponding feature map blocks through a checkerboard; among them, the feature map block can be at the pixel level (that is, each pixel is a feature map block), or it can be at other levels, and the specific division depends on the actual processing situation.
[0089] In this application, a sliding window or a fixed step size is used to divide the feature map into blocks of the same size.
[0090] It should be noted here that if the input feature map is a two-dimensional feature map, it is directly divided into a checkerboard, and each grid is a feature map block; if the input feature map is a three-dimensional feature map, a surface is selected for checkerboard division, and each grid is a strip-shaped grid with a lot of depth (the depth is the depth of the three-dimensional feature map), and this strip-shaped grid is a feature map block.
[0091] Preferably, in the present application, each feature patch is 100 - 1000 pixels, so as to perform more feature calculations between local regions on the basis of ensuring the generation accuracy and reducing the computational amount.
[0092] In the present application, a feature patch is selected as an independent block, and the adjacent feature patches on the upper side, lower side, left side, and right side of the independent block are the first neighborhood blocks; the feature patches separated by one grid on the upper side, lower side, left side, and right side of the independent block are the second neighborhood blocks. The distances between the first neighborhood blocks and the second neighborhood blocks and the independent block are different.
[0093] In the present application, neighborhood information is extracted for each independent block to capture local structures.
[0094] In the present application, to generate the first feature block, a local feature representation is generated using the independent block and its first neighborhood blocks. Specifically, it can be: the independent block and the first neighborhood blocks are processed through a convolutional layer and an attention layer to obtain the first feature block.
[0095] In the present application, the specific structures and specific parameters of the convolutional layer and the attention layer can be obtained according to the training data or determined according to the actual situation.
[0096] It should be noted that in the present application, there are four first neighborhood blocks and multiple first feature blocks.
[0097] In the present application, the independent block and the first neighborhood blocks are processed through a convolutional layer and an attention layer to obtain the first feature block. The specific process is: the independent block and the four neighborhood blocks are concatenated together to form a multi-channel input, and the convolutional layer is used to extract features from the concatenated block; the self-attention mechanism or the channel attention mechanism is used to enhance important features, calculate the attention weights, and weight the output of the convolutional layer to enhance important features; the output of the attention layer is split into multiple feature blocks, and each feature block corresponds to the processing result of the independent block and at least one neighborhood block.
[0098] In the present application, to generate the second feature block, a more extensive local feature representation is generated using the independent block and its second neighborhood blocks. The specific generation process is the same as that of the first feature block, except that the parameters of the convolutional layer and the attention layer are different.
[0099] In the present application, the generated feature blocks are compressed into a more compact representation to reduce the computational amount and retain key information. Pooling operations (such as max pooling or average pooling) or fully connected layers are used for feature compression.
[0100] In this way, through compression, multiple first feature blocks and second feature blocks are compressed into a compressed block, which has the same size and position as the independent block and is used to replace the independent block. All feature patches are replaced by compressed blocks to obtain the output feature map.
[0101] In this application, by means of traversal, each feature map block of the input feature map is traversed to obtain the corresponding compressed block.
[0102] In this application, for the feature map blocks / independent blocks near the edge, their first neighborhood blocks and second neighborhood blocks are incomplete. At this time, the first neighborhood blocks and second neighborhood blocks in the relative positions are copied for completion. For example, if the first neighborhood block above the independent block does not exist, the first neighborhood block below is copied and used as the block above.
[0103] In this application, through completion, the processing accuracy of the edge feature map blocks is greatly improved.
[0104] In this application, the intermediate module is used to capture the similarity relationship between local regions, thereby enhancing the feature representation.
[0105] In this application, through preprocessing, fine-tuning, compression, and secondary training, a lightweight compression model is generated, significantly reducing the computational complexity and storage requirements while maintaining the high-precision lesion recognition ability.
[0106] In this application, block embedding and linear mapping are combined to adapt the Transformer to process medical images; through gradient mean pruning and attention head pruning, efficient structure compression is achieved; an intermediate module is inserted and secondary training is performed to further improve the model performance.
[0107] In one implementation, the secondary training of the adjusted pre-trained large model to obtain the compression model includes:
[0108] Freeze the Transformer structure parameters of the adjusted pre-trained large model;
[0109] Input the labeled sample data into the pre-trained large model to obtain the prediction result;
[0110] Calculate the overall loss according to the label and the prediction result;
[0111] Adjust the parameters of the intermediate module in the pre-trained large model based on the overall loss until the loss converges.
[0112] In this application, during the secondary training process, the parameters of all Transformer layers are frozen to ensure that only the parameters of the intermediate module are trained; the purpose of freezing the parameters is to avoid destroying the structural stability of the compressed model and reduce the training computational amount at the same time.
[0113] In this application, a medical image dataset (such as BraTS or CheXpert) is used as the training sample, and each sample contains an image and the corresponding label (such as the lesion type or region); the preprocessed image is input into the compression model, and the model outputs the prediction result (such as the lesion probability distribution or segmentation mask).
[0114] In this application, the model output is compared with the true label to calculate the loss value; the cross-entropy loss function and the Dice loss function are used, and the weighted sum is calculated using weight allocation to calculate the loss value.
[0115] In this application, the optimizer settings are as follows: the AdamW optimizer is used, and the learning rate is set to 0.00001; the optimizer only updates the parameters of the intermediate module.
[0116] In this application, the training process is as follows: the training data is divided into multiple batches (Batch), each batch contains 32 images; the loss is calculated for each batch and backpropagated to update the parameters of the intermediate module.
[0117] In this application, the convergence condition is: when the validation set loss does not decrease for 5 consecutive epochs, stop training; save the final model weights.
[0118] In this application, by freezing the Transformer parameters and only training the intermediate module, the training computational amount is significantly reduced; the intermediate module parameters are optimized by secondary training to further improve the performance of the model in the medical image lesion recognition task; since only a small number of parameters are trained, the model can converge in a short time.
[0119] An embodiment of this application provides a medical image lesion recognition device based on a pre-trained large model, which is used to execute a medical image lesion recognition method described above in this application. The following is a detailed description of the medical image lesion recognition device based on a pre-trained large model.
[0120] As Figure 2 shown, the medical image lesion recognition device based on a pre-trained large model includes:
[0121] An image preprocessing module 101, which is used to preprocess the acquired medical image to obtain a corresponding input sequence;
[0122] A model acquisition module 102, which is used to acquire a pre-trained large model;
[0123] A model compression module 103, which is used to compress the pre-trained large model based on the input sequence to obtain a compressed model;
[0124] A lesion recognition module 104, which is used to input the medical image to be recognized into the compressed model to obtain the lesion recognition result of the medical image.
[0125] In one implementation, the image preprocessing module 101 is further used for:
[0126] Adjust the medical image to a fixed size and normalize the pixel values; segment the image into a sequence of squares of preset pixels, and generate an embedding vector through linear mapping, and the embedding vector is the input sequence.
[0127] In one embodiment, the pre-trained large model consists of a 24-layer Transformer architecture.
[0128] In one embodiment, the model compression module 103 is further configured to:
[0129] Update the output layer of the pre-trained large model; based on the input sequence, perform fine-tuning training on the updated pre-trained large model; perform structure compression and structure adjustment on the fine-tuned pre-trained large model; perform secondary training on the adjusted pre-trained large model to obtain the compressed model.
[0130] In one embodiment, the model compression module 103 is further configured to:
[0131] Freeze the parameters of the preset layers in the front of the pre-trained large model; based on the input sequence, perform fine-tuning training on the frozen pre-trained large model; the learning rate of the optimizer in the fine-tuning training is set to 0.0001.
[0132] In one embodiment, the model compression module 103 is further configured to:
[0133] Calculate the mean of the absolute values of the gradients output by each layer of the Transformer structure; count the norm of the weight matrix of each attention head in the Transformer structure; remove several layers of the Transformer structure with the smallest mean of the absolute values of the gradients, and remove several attention heads with the smallest norm of the weight matrix in the remaining Transformer structures; insert intermediate modules between the remaining adjacent layers of the Transformer structure.
[0134] In one embodiment, the model compression module 103 is further configured to:
[0135] Freeze the parameters of the Transformer structure of the adjusted pre-trained large model; input the sample data with labels into the pre-trained large model to obtain a prediction result; calculate the overall loss according to the labels and the prediction result; adjust the parameters of the intermediate module in the pre-trained large model based on the overall loss until the loss converges.
[0136] A medical image lesion recognition device based on a pre-trained large model provided by the above embodiments of the present application and a medical image lesion recognition method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0137] The internal functions and structures of a medical image lesion recognition device based on a pre-trained large model are described above. For example, Figure 3 as shown, in practice, the medical image lesion recognition device based on a pre-trained large model can be implemented as an electronic device, including: a memory 301 and a processor 303.
[0138] The memory 301 can be configured to store programs.
[0139] In addition, the memory 301 can also be configured to store various other data to support operations on the electronic device. Examples of these data include instructions for any application program or method for operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.
[0140] The memory 301 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0141] The processor 303, coupled to the memory 301, is used to execute the programs in the memory 301 for:
[0142] Preprocessing the acquired medical image to obtain a corresponding input sequence;
[0143] Obtaining a pre-trained large model;
[0144] Based on the input sequence, performing model compression on the pre-trained large model to obtain a compressed model;
[0145] Inputting the medical image to be recognized into the compressed model to obtain the lesion recognition result of the medical image.
[0146] In one embodiment, the processor 303 is further used for:
[0147] Adjusting the medical image to a fixed size and normalizing the pixel values; segmenting the image into a sequence of squares of preset pixels and generating an embedding vector through linear mapping, and the embedding vector is the input sequence.
[0148] In one embodiment, the pre-trained large model consists of 24 layers of Transformer architecture.
[0149] In one embodiment, the processor 303 is further used for:
[0150] Update the output layer of the pre-trained large model; perform fine-tuning training on the updated pre-trained large model based on the input sequence; perform structure compression and structure adjustment on the fine-tuned pre-trained large model; perform secondary training on the adjusted pre-trained large model to obtain the compressed model.
[0151] In one implementation, the processor 303 is further configured to:
[0152] Freeze the parameters of the preset layers in the front of the pre-trained large model; perform fine-tuning training on the frozen pre-trained large model based on the input sequence; the learning rate of the optimizer in the fine-tuning training is set to 0.0001.
[0153] In one implementation, the processor 303 is further configured to:
[0154] Calculate the mean of the absolute values of the gradients output by each layer of the Transformer structure; count the norm of the weight matrix of each attention head in the Transformer structure; remove several layers of the Transformer structure with the smallest mean of the absolute values of the gradients, and remove several attention heads with the smallest norm of the weight matrix in the remaining Transformer structure; insert intermediate modules between the remaining adjacent layers of the Transformer structure.
[0155] In one implementation, the processor 303 is further configured to:
[0156] Freeze the parameters of the Transformer structure of the adjusted pre-trained large model; input the sample data with annotations into the pre-trained large model to obtain a prediction result; calculate the overall loss according to the annotations and the prediction result; adjust the parameters of the intermediate module in the pre-trained large model based on the overall loss until the loss converges.
[0157] In this application, Figure 3 only some components are schematically shown, which does not mean that the electronic device only includes Figure 3 the components shown.
[0158] The electronic device provided in this embodiment and a medical image lesion recognition method based on a pre-trained large model provided in an embodiment of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0159] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0160] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0161] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0162] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0163] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0164] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (Flash RAM). The memory is an example of a computer-readable medium.
[0165] The present application also provides a computer-readable storage medium corresponding to a medical image lesion recognition method based on a pre-trained large model provided in the foregoing embodiments. A computer program (i.e., a program product) is stored thereon. When the computer program is run by a processor, it will execute a medical image lesion recognition method based on a pre-trained large model provided in any of the foregoing embodiments.
[0166] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CDROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0167] The computer-readable storage medium provided in the above embodiments of the present application and a medical image lesion recognition method based on a pre-trained large model provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored thereon.
[0168] It should be noted that a large number of specific details are set forth in the specification provided herein. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known structures and technologies have not been shown in detail so as not to obscure the understanding of this specification.
[0169] It should also be noted that the term "comprises", "comprising", or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, commodity, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity, or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity, or device comprising the element.
[0170] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for identifying lesions in medical images based on a pre-trained large model, characterized in that: include: Preprocessing the acquired medical images to obtain corresponding input sequences; Get a pre-trained large model; Based on the input sequence, compressing the pre-trained large model to obtain a compressed model; The medical image to be identified is input into the compression model to obtain a lesion identification result of the medical image.
2. The method for medical image lesion recognition based on a pre-trained large model according to claim 1, characterized in that: The preprocessing of the acquired medical images to obtain a corresponding input sequence includes: Resize medical images to a fixed size and normalize pixel values; The image is divided into a block sequence of preset pixels, and an embedding vector is generated through linear mapping, where the embedding vector is the input sequence.
3. The method for medical image lesion recognition based on a pre-trained large model according to claim 1, characterized in that: The pre-trained large model consists of a 24-layer Transformer architecture.
4. The method for medical image lesion recognition based on a pre-trained large model according to claim 3, characterized in that: The method of compressing the pre-trained large model based on the input sequence to obtain a compressed model includes: Updating the output layer of the pre-trained large model; Based on the input sequence, fine-tune the updated pre-trained large model; Perform structural compression and structural adjustment on the fine-tuned pre-trained large model; The adjusted pre-trained large model is trained again to obtain the compressed model.
5. The method for medical image lesion recognition based on a pre-trained large model according to claim 4, characterized in that: The step of fine-tuning the updated pre-trained large model based on the input sequence comprises: Freeze the parameters of the front preset layer of the pre-trained large model; Based on the input sequence, the frozen pre-trained large model is fine-tuned; in the fine-tuning training, the learning rate of the optimizer is set to 0.0001.
6. The method for medical image lesion recognition based on a pre-trained large model according to claim 4, characterized in that: The structural compression and structural adjustment of the fine-tuned pre-trained large model includes: Calculate the mean absolute value of the gradient output of each layer of the Transformer structure; Count the norm of the weight matrix of each attention head in the Transformer structure; Remove several layers of Transformer structures with the smallest mean absolute value of gradient, and remove several attention heads with the smallest norm of weight matrix in the retained Transformer structure; Intermediate modules are inserted between the retained adjacent layer Transformer structures.
7. The method for medical image lesion recognition based on a pre-trained large model according to claim 4, characterized in that: The second training of the adjusted pre-trained large model to obtain the compressed model includes: Freeze the adjusted Transformer structure parameters of the pre-trained large model; Input the labeled sample data into the pre-trained large model to obtain the prediction results; Calculate the overall loss based on the annotations and the prediction results; Adjust the parameters of the intermediate modules in the pre-trained large model based on the overall loss until the loss converges.
8. A medical image lesion recognition device based on a pre-trained large model, characterized in that: include: An image preprocessing module, which is used to preprocess the acquired medical images to obtain a corresponding input sequence; Model acquisition module, which is used to obtain a pre-trained large model; A model compression module, which is used to compress the pre-trained large model based on the input sequence to obtain a compressed model; The lesion recognition module is used to input the medical image to be recognized into the compression model to obtain the lesion recognition result of the medical image.
9. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor, coupled to the memory, is configured to execute the program to: Preprocessing the acquired medical images to obtain corresponding input sequences; Get a pre-trained large model; Based on the input sequence, compressing the pre-trained large model to obtain a compressed model; The medical image to be identified is input into the compression model to obtain a lesion identification result of the medical image.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by the processor to implement a medical image lesion recognition method based on a pre-trained large model as described in any one of claim 17.