Training Method of Pre-trained Network Model, Medical Image Processing Method and Device

By training the initial pre-trained network model based on feature images, the problem of poor training effect of deep learning model caused by the small amount of medical image data is solved, and the effect of improving training efficiency and generalization ability is achieved.

CN114782768BActive Publication Date: 2025-05-30SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210239706.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-12
Publication Date
2025-05-30
Estimated Expiration
2042-03-12

AI Technical Summary

Technical Problem

Due to the small amount of medical image data and the need for manual annotation, the training effect of deep learning models under low data volume is poor, and the prior art is difficult to effectively solve this problem.

Method used

By obtaining the initial pretrained network model and multiple pretrained images, performing single-channel feature calculations, obtaining feature images, and training the initial pretrained network model based on these feature images to obtain the pretrained network model. The visual network parameters of the pre-trained network model are used as the initial parameters of the image processing model, simplifying the training process of the image processing model and improving the training efficiency.

Benefits of technology

Through this method, the generalization ability of the pre-trained model is enhanced, the accuracy of visual network parameters is improved, the training process of medical image processing models is simplified, and the training efficiency and effect are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114782768B_ABST
    Figure CN114782768B_ABST
Patent Text Reader

Abstract

The present application provides a training method for a pre-trained network model, a medical image processing method and device, relating to the technical field of image processing. In the present application, single-channel feature calculations are performed on multiple pre-trained images to obtain feature images corresponding to the respective pre-trained images. An initial pre-trained network model is trained based on the respective pre-trained images and their corresponding feature images to obtain a pre-trained network model. In this embodiment, using the feature images as the training targets makes it easier to learn the general features of the pre-trained images, enhances the generalization ability of the pre-trained model, and makes the visual network parameters of the obtained pre-trained network model more accurate. Furthermore, using the visual network parameters of the pre-trained network model as the initial parameters of the visual network to be trained in the training of the image processing model can simplify the training process of the image processing model and improve the training efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and in particular, to a training method for a pre-trained network model, a medical image processing method, and a device. Background Art

[0002] The wide application of deep learning in the field of medical images, especially the success of the new generation of vision-based deep learning models, has led to an increasing demand for resources such as computing power required for training models and medical image data for training.

[0003] Due to the relatively fixed data of patients with specific diseases and the limitation of patient data privacy, the amount of medical image data collected for training is small. Moreover, high-performance deep learning models often require manual annotation by professional doctors, which increases the workload of doctors and further squeezes medical resources. Therefore, high-quality data available for deep neural network training is still scarce. So the training effect of vision neural network training for medical image tasks is poor under the condition of low data volume, which remains an unsolved problem. Summary of the Invention

[0004] The purpose of the present application is to provide a training method for a pre-trained network model, a medical image processing method, and a device to solve the above technical problems.

[0005] In a first aspect, a training method for a pre-trained network model is provided, including:

[0006] Obtaining an initial pre-trained network model and multiple pre-training images, where the initial pre-trained network model includes a visual network to be trained and a first decoding network to be trained;

[0007] Performing single-channel feature calculation on each of the pre-training images to obtain a feature image corresponding to each of the pre-training images; training the initial pre-trained network model according to each of the pre-training images and its corresponding feature image to obtain the pre-trained network model;

[0008] Wherein, the visual network parameters of the pre-trained network model can be used as the initial parameters of the visual network to be trained for training an image processing model, and the image processing model is used for image processing of a to-be-processed two-dimensional medical image.

[0009] Based on the above solution, this solution calculates single-channel features for multiple pre-trained images to obtain feature images corresponding to each pre-trained image. The initial pre-trained network model is trained according to each pre-trained image and its corresponding feature image to obtain the pre-trained network model. In this embodiment, the feature image is used as the training target, which makes it easier to learn the general features of the pre-trained images, enhances the generalization ability of the pre-trained model, and the visual network parameters of the obtained pre-trained network model are more accurate. Furthermore, the visual network parameters of the pre-trained network model are used as the initial parameters of the visual network to be trained in the training of the image processing model, which can simplify the training process of the image processing model and improve the training efficiency.

[0010] In a possible implementation manner, the step of training the initial pre-trained network model according to each pre-trained image and its corresponding feature image to obtain the pre-trained network model includes:

[0011] Perform masking processing on each pre-trained image to generate a masking map corresponding to each pre-trained image;

[0012] Input each masking map into the initial pre-trained network model to output a comparison image corresponding to each masking map; calculate a loss value according to the comparison image corresponding to each masking map and its corresponding feature map by using the masked mean square error loss function;

[0013] Determine whether the initial pre-trained network model converges according to the loss value; if so, determine the current initial pre-trained network model as the pre-trained network model that has completed training; if not, perform iterative training until the initial pre-trained network model converges to obtain the pre-trained network model.

[0014] In a possible implementation manner, the step of calculating single-channel features for each pre-trained image to obtain a feature image corresponding to each pre-trained image includes:

[0015] Based on a preset grid size, divide each pre-trained image into grids to obtain multiple blocks;

[0016] Calculate the single-channel features of each block of each pre-trained image to obtain a feature map corresponding to each pre-trained image;

[0017] Correspondingly, the step of performing masking processing on each pre-trained image to generate a masking map corresponding to each pre-trained image includes:

[0018] Generate a mask for each pre-trained image based on a preset grid size;

[0019] Mask each of the pre-trained images according to the mask to generate a masked image corresponding to each of the pre-trained images.

[0020] In a possible implementation, the pre-trained image is a pre-trained two-dimensional medical image and / or a natural image; the pre-trained two-dimensional medical image includes any one or more of the following: CT image, MRI image.

[0021] In a possible implementation, the calculating single-channel features of each of the pre-trained images to obtain the single-channel features of the feature images corresponding to each of the pre-trained images includes:

[0022] Calculate target features for each of the pre-trained images to obtain a feature image corresponding to each of the pre-trained images, where the target features include any one of Haar features, Gabor features, and LBP features.

[0023] In a second aspect, a medical image processing method is provided, including:

[0024] Obtain an initial image processing model to be trained, where the initial image processing model includes a visual network to be trained and a second decoding network to be trained, and the initial parameters of the visual network to be trained of the initial image processing model are the parameters of the visual network of the pre-trained network model; the pre-trained network model is obtained based on multiple pre-trained images and their corresponding feature images;

[0025] Obtain multiple two-dimensional medical images and their corresponding annotation information;

[0026] Train the initial image processing model according to the multiple two-dimensional medical images and their corresponding annotation information to obtain the image processing model;

[0027] Obtain a two-dimensional medical image to be processed, and use the image processing model to perform image processing on the two-dimensional medical image to be processed.

[0028] In a possible implementation, the performing image processing on the two-dimensional medical image to be processed by using the image processing model includes:

[0029] Input the two-dimensional medical image to be processed into the visual network of the image processing model to obtain an encoded feature map;

[0030] Input the encoded feature map into the second decoder of the image processing model to obtain an image processing result.

[0031] In a possible implementation, the image processing includes any one of target recognition, image segmentation, and image classification.

[0032] In a third aspect, a training device for a pre-trained network model is provided, including:

[0033] A pre-training information acquisition module, configured to acquire an initial pre-trained network model and multiple pre-training images, where the initial pre-trained network model includes a visual network to be trained and a first decoding network to be trained;

[0034] A feature calculation module, configured to perform single-channel feature calculation on each of the pre-training images to obtain a feature image corresponding to each of the pre-training images;

[0035] A pre-training module, configured to train the initial pre-trained network model according to each of the pre-training images and the corresponding feature image thereof to obtain the pre-trained network model;

[0036] Wherein, the visual network parameters of the pre-trained network model can be used as the initial parameters of the visual network to be trained in the training of the image processing model, and the image processing model is used to perform image processing on the to-be-processed two-dimensional medical image.

[0037] In a possible implementation manner, the pre-training module includes:

[0038] A mask processing unit, configured to perform mask processing on each of the pre-training images to generate a masking map corresponding to each of the pre-training images;

[0039] An output unit, configured to input each of the masking maps into the initial pre-trained network model and output a to-be-compared image corresponding to each of the masking maps;

[0040] A loss value calculation unit, configured to calculate according to the to-be-compared image corresponding to each of the masking maps and the corresponding feature map by using a masking mean square error loss function to obtain a loss value;

[0041] A pre-trained network model acquisition unit, configured to determine whether the initial pre-trained network model converges according to the loss value; if so, determine the current initial pre-trained network model as the pre-trained network model that has been trained; if not, perform iterative training until the initial pre-trained network model converges to obtain the pre-trained network model.

[0042] In a possible implementation manner, the feature calculation module includes:

[0043] A grid division unit, configured to perform grid division on each of the pre-training images based on a preset grid size to obtain multiple blocks; a feature calculation unit, configured to perform single-channel feature calculation on each block of each of the pre-training images to obtain a feature map corresponding to each of the pre-training images;

[0044] Correspondingly, the mask processing unit includes:

[0045] A mask generation subunit, configured to generate a mask for each of the pre-training images based on a preset grid size;

[0046] A mask processing subunit, configured to perform mask processing on each of the pre-training images according to the mask to generate a masked image corresponding to each of the pre-training images.

[0047] In a possible implementation, the pre-training images are pre-training two-dimensional medical images and / or natural images; the pre-training two-dimensional medical images include any one or more of the following: CT images, MRI images.

[0048] In a possible implementation, the feature calculation module includes:

[0049] A feature calculation unit, configured to perform target feature calculation on each of the pre-training images to obtain a feature image corresponding to each of the pre-training images, where the target features include any one of Haar features, Gabor features, and LBP features.

[0050] In a fourth aspect, a medical image processing device is provided, including:

[0051] An initial image processing model acquisition module, configured to acquire an initial image processing model to be trained, where the initial image processing model includes a visual network to be trained and a second decoding network to be trained, and the initial parameters of the visual network to be trained of the initial image processing model are the parameters of the visual network of the pre-training network model; the pre-training network model is obtained according to multiple pre-training images and their corresponding feature images;

[0052] A training set acquisition module, configured to acquire multiple two-dimensional medical images and their corresponding annotation information;

[0053] An image processing model acquisition module, configured to train the initial image processing model according to multiple two-dimensional medical images and their corresponding annotation information to obtain the image processing model;

[0054] An image processing module, configured to acquire a two-dimensional medical image to be processed and perform image processing on the two-dimensional medical image to be processed by using the image processing model.

[0055] In a possible implementation, when the image processing module performs image processing on the two-dimensional medical image to be processed by using the image processing model, it is configured to:

[0056] Input the two-dimensional medical image to be processed into the visual network of the image processing model to obtain an encoded feature map;

[0057] Input the encoded feature map into the second decoder of the image processing model to obtain an image processing result.

[0058] In a possible implementation, the image processing includes any one of object recognition, image segmentation, and image classification.

[0059] In a fifth aspect, an electronic device is provided, including:

[0060] One or more processors;

[0061] A memory;

[0062] One or more applications, where the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to: execute the training method of the pre-trained network model shown in any possible implementation of the first aspect, or execute the medical image processing method shown in any possible implementation of the second aspect.

[0063] In a sixth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the training method of the pre-trained network model shown in any possible implementation of the first aspect, or implements the medical image processing method shown in any possible implementation of the second aspect when executed.

[0064] In summary, the present application includes at least one of the following beneficial technical effects:

[0065] The present application provides a training method for a pre-trained network model, a medical image processing method, and a device. In the present application, single-channel feature calculation is performed on multiple pre-trained images to obtain feature images corresponding to the respective pre-trained images, and the initial pre-trained network model is trained based on the respective pre-trained images and their corresponding feature images to obtain a pre-trained network model. In this embodiment, the feature image is used as the training target, making it easier to learn the general features of the pre-trained images, enhancing the generalization ability of the pre-trained model, and making the visual network parameters of the obtained pre-trained network model more accurate. Furthermore, using the visual network parameters of the pre-trained network model as the initial parameters of the visual network to be trained in the training of the image processing model can simplify the training process of the image processing model and improve the training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 It is a schematic flowchart of a training method for a pre-trained network model provided by an embodiment of the present application;

[0067] Figure 2 It is a schematic pre-training flowchart provided by an embodiment of the present application;

[0068] Figure 3 Schematic flowchart of a medical image processing method provided by an embodiment of this application;

[0069] Figure 4 Schematic flowchart of the training process of an image processing model provided by an embodiment of this application;

[0070] Figure 5 Schematic diagram of a pre-training provided by an embodiment of this application;

[0071] Figure 6 Block diagram of the structure of a training device for a pre-trained network model provided by an embodiment of this application;

[0072] Figure 7 Block diagram of the structure of a medical image processing device provided by an embodiment of this application;

[0073] Figure 8 Block diagram of the structure of an electronic device provided by an embodiment of this application. Detailed implementation manners

[0074] The following will make a further detailed description of this application in conjunction with the attached Figure 1 to the attached Figure 8 for a further detailed description of this application.

[0075] This specific embodiment is only an interpretation of this application and does not limit this application. After reading this specification, those skilled in the art can make modifications to this embodiment without creative contributions as needed, but as long as it is within the scope of this application, it is protected by the patent law.

[0076] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of this application.

[0077] In addition, the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, unless otherwise specified.

[0078] When training an image processing model using deep learning in the field of medical images, due to the relatively fixed data of patients with specific diseases and the restrictions on patient data privacy, the amount of medical image data collected for training is small. Moreover, high-performance deep learning models often require manual annotation by professional doctors, which increases the workload of doctors and further squeezes medical resources. Therefore, high-quality data available for deep neural network training is still scarce. So, the training effect of visual neural network training for medical image tasks is poor under low data volume, which remains an issue to be solved.

[0079] To solve the above technical problems, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings of the specification.

[0080] The embodiments of the present application provide a training method for a pre-trained network model, which is executed by an electronic device. The electronic device can be a server or a terminal device. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and the embodiments of the present application do not make any restrictions here, such as Figure 1 shown Figure 1 is a schematic flowchart of a training method for a pre-trained network model provided by the embodiments of the present application. The method includes:

[0081] S110, obtain an initial pre-trained network model and multiple pre-training images. The initial pre-trained network model includes a visual network to be trained and a first decoding network to be trained;

[0082] In the embodiments of the present application, a large number of pre-training images are required for training the pre-trained network model. The pre-training images can be two-dimensional medical images, specifically, any one or more of CT images, MRI (Magnetic Resonance Imaging) images, and ultrasonic images. However, since it is difficult to collect two-dimensional medical images, the pre-training images can also be natural images taken by natural light imaging. If the multiple pre-training images include two-dimensional medical images and natural images, the embodiments of the present application do not limit the ratio of the number of two-dimensional medical images to natural images, and users can set it according to actual needs. Among them, training the initial pre-trained network model with natural images can learn general information, and the general information includes but is not limited to: edge information, contrast information, shape information, texture information, and hue information; training the initial pre-trained network model with two-dimensional medical images can learn relevant medical information.

[0083] Preferably, the modality of the pre-trained image is consistent with the modality of the input image of the image processing model in the actual application scenario, so as to improve the accuracy of the parameters of the trained pre-trained network model. Furthermore, the visual network parameters of the pre-trained network model are used as the initial parameters of the visual network of the image processing model. The image processing model obtained by training according to multiple two-dimensional medical images and their corresponding annotation information can effectively improve the training efficiency. For example, when the image processing model is an MRI image artifact removal model, the corresponding pre-trained image is an MRI image; another example is that when the image processing model is a CT image disease recognition model, the determined pre-trained image is a CT image.

[0084] Further, the method for obtaining multiple pre-trained images may include: obtaining multiple pre-trained images from local storage, or crawling multiple pre-trained images from the network, or obtaining multiple pre-trained images input by the user.

[0085] Further, before step S110, it may further include: obtaining multiple initial pre-trained images; performing image size processing and image enhancement processing on the multiple initial pre-trained images to obtain corresponding pre-trained images, where all the pre-trained images have the same size. The image enhancement processing includes but is not limited to any one or more of contrast enhancement, denoising, filtering, edge sharpening, etc., to improve the visual effect of the image, highlight the features of the image, and facilitate further analysis of the image by the machine.

[0086] Further, before step S110, it may further include: obtaining multiple initial pre-trained images, where the initial pre-trained images include pre-trained two-dimensional medical images; performing affine transformation and / or elastic transformation on the multiple pre-trained two-dimensional medical images to obtain multiple augmented pre-trained two-dimensional medical images, where the multiple pre-trained images include all the pre-trained two-dimensional medical images and all the augmented pre-trained two-dimensional medical images.

[0087] Specifically, due to the particularity of medical images, it is difficult to obtain two-dimensional medical images. Therefore, the pre-trained two-dimensional medical images are limited. Therefore, when only using pre-trained two-dimensional medical images as pre-trained images, the robustness of the obtained pre-trained network model is poor. Correspondingly, the accuracy of the visual network parameters of the obtained pre-trained network model is insufficient. Therefore, in this embodiment, affine transformation and elastic transformation are used to augment the pre-trained images to obtain augmented pre-trained images. Among them, the affine transformation performs a translation, rotation, scaling, shear, and symmetry on the multiple initial pre-trained images. The elastic transformation performs an elastic transformation operation on the multiple initial pre-trained images, or on the initial pre-trained images after affine transformation to augment the samples. After affine transformation and elastic transformation, the final pre-trained images obtained can increase the diversity of the pre-trained images, so that the pre-trained network model can learn multiple features.

[0088] Specifically, in the embodiments of the present application, the initial pre-trained network model includes a visual network to be trained as the encoder of the initial pre-trained network model and a first decoding network to be trained connected to the visual network. The first decoding network to be trained can be a fully connected layer with linear activation. Specifically, the visual network can be ResNET18 without the output layer. Of course, it may also be other structures, which are not limited in this embodiment as long as they can achieve the purpose of this embodiment.

[0089] S120. Calculate single-channel features for each pre-trained image to obtain a feature image corresponding to each pre-trained image; medical images are generally single-channel grayscale images. Therefore, in this embodiment, single-channel feature calculation is performed on the pre-trained images to obtain feature images corresponding to each pre-trained image.

[0090] Specifically, it can be to calculate target features for each pre-trained image to obtain a feature image corresponding to each pre-trained image, where the target features include: Haar features, Gabor features, and LBP features. In this embodiment, local features are extracted through target feature calculation, and then the local similarity in the pre-trained images can be utilized to improve the utilization of data.

[0091] Generally, the process of model training using the transfer learning method includes: pre-training using a large dataset to enable the model to learn the general features of images, and then fine-tuning the parameters on the dataset of specific tasks to "transfer" the natural image knowledge learned by the model to the task images for classification and diagnosis of task images. However, the pre-trained data required by the method needs to be labeled, that is, supervised, and only the unique features of the labels can be learned. And, the feature map obtained in this step is used as the training target of the initial pre-trained network model, which can make the initial pre-trained network model more easily learn the general features of the pre-trained images, and can fully understand the spatial structure, spatial brightness, contrast, etc. of the pre-trained images, avoiding the unique features of the pre-trained images learned by the initial pre-trained network model by using the method of manually labeled data, and improving the generalization ability of the model.

[0092] S130. Train the initial pre-trained network model according to each pre-trained image and its corresponding feature image to obtain a pre-trained network model. The visual network parameters of the pre-trained network model can be used as the initial parameters of the visual network to be trained for image processing model training.

[0093] Specifically, a to-be-compared image can be obtained based on a pre-trained image and an initial pre-trained network model; the to-be-compared image and the feature image are calculated using a preset loss function to obtain a loss value; the initial pre-trained network model is iteratively trained based on the loss value. When the loss value reaches a preset loss threshold, or all pre-trained images have been trained, or the image training for a preset period has been completed, it is determined that the initial pre-trained network model converges, and the current initial pre-trained network model is determined as the trained pre-trained network model, where the preset loss threshold can be custom-set by the user or set according to experience.

[0094] The parameters of the pre-trained network model in this embodiment include: visual network parameters and first decoder parameters, and the visual network parameters are used as the initial parameters of the to-be-trained visual network for the training of the image processing model.

[0095] In summary, compared with most existing neural network pre-training methods, the pre-training data set in the embodiments of this application does not need to be labeled, is self-supervised, and because the main change in two-dimensional medical images is the contrast, compared with other self-supervised pre-training methods, the embodiments of this application use the feature map as the pre-training target, enabling the neural network to learn to predict the local contrast information of the image, and the generalization effect is better.

[0096] Through the above technical solution, it can be seen that in the embodiments of this application, single-channel feature calculation is performed on multiple pre-trained images to obtain the feature images corresponding to the respective pre-trained images, and the initial pre-trained network model is trained according to each pre-trained image and its corresponding feature image to obtain the pre-trained network model. In this embodiment, the feature image is used as the training target, which is easier to learn the general features of the pre-trained images, enhances the generalization ability of the pre-trained model, and the visual network parameters of the obtained pre-trained network model are more accurate. Furthermore, using the visual network parameters of the pre-trained network model as the initial parameters of the to-be-trained visual network for the training of the image processing model can simplify the training process of the image processing model, improve the training efficiency and training effect.

[0097] Furthermore, the pre-trained network model introduced in the above embodiment is obtained by training the initial pre-trained network model according to each pre-trained image and its corresponding feature image.

[0098] The embodiments of this application provide a pre-training method with simple operation. Specifically, S130 may include: inputting the pre-trained image into the initial pre-trained network model to obtain a to-be-compared image; calculating the to-be-compared image and the feature image using a preset loss function to obtain a loss value; and iteratively training the initial pre-trained network model based on the loss value to obtain the pre-trained network model. In this method, the pre-trained image is directly input into the initial pre-trained network model for model training, and the method is simple and highly operable.

[0099] The embodiment of the present application provides a pre-training method to improve the training efficiency. Please refer to Figure 2 , Figure 2 which is a schematic diagram of a pre-training process provided by the embodiment of the present application. Specifically, S130 may include: S131, S132, S133, S134, where:

[0100] S131. Perform masking processing on each pre-training image to generate a masking map corresponding to each pre-training image;

[0101] In order to further explore the internal connection of the pre-training image and make the feature extraction obtained by the initial pre-training network model more accurate, in this embodiment, the pre-training image is subjected to masking processing to obtain a masking map.

[0102] Among them, the masking processing is to block the local part of the pre-training image. The commonly used mask for blocking can be a multi-valued image or a binary matrix. In the embodiment of the present application, the mask can be used to block the local part of the pre-training image to obtain a masking map, and the masking map is used to train the initial preprocessing model, so as to increase the difficulty of local feature extraction of the initial preprocessing model, and enable the training process to learn and predict the local comparison information of the image. It can be clear that the mask of each pre-training image can be set according to actual needs as long as it can achieve the purpose of this embodiment.

[0103] S132. Input each masking map into the initial pre-training network model and output a to-be-compared image corresponding to each masking map; S133. According to the to-be-compared image corresponding to each masking map and its corresponding feature map, calculate using the masking mean square error loss function to obtain a loss value;

[0104] In this embodiment, the masking mean square error loss function is used to calculate the loss value to obtain the loss value of the to-be-compared image corresponding to each masking map and its corresponding feature map. It can be understood that the embodiment of the present application uses the masking mean square error loss function for calculation, and only calculates the mean square error of the output feature of the masked part and the corresponding part of the feature image as the loss value, and this loss value is used to represent the gap between the to-be-compared image corresponding to the masking map and the feature map.

[0105] S134. Determine whether the initial pre-training network model converges according to the loss value; if so, determine the current initial pre-training network model as the pre-training network model that has completed training; if not, perform iterative training until the initial pre-training network model converges to obtain the pre-training network model.

[0106] Specifically, when the loss value is within the first range, it is determined that the initial pre-trained network model converges. When the loss value is not within the first range, it is determined that the initial pre-trained network model does not converge. Based on the loss value, backpropagation is performed to update the weight parameters of the initial pre-trained network model, and then iterative training is carried out until the pre-trained network model is obtained. Among them, the first range can be set by the user according to actual needs or according to empirical values, as long as the purpose of this embodiment can be achieved.

[0107] It can be seen that in the embodiment of the present application, each pre-trained image is masked to obtain a corresponding masked image, and the masked image is used as the input data of the initial pre-trained model to output the images to be compared corresponding to each masked image. Using the images to be compared corresponding to the masked images as the images to be compared and the feature map as the target image, the masked mean square error loss function is used for calculation, and based on the obtained loss value, the convergence situation of the current initial pre-trained model is determined to obtain the finally trained pre-trained network model. It can make full use of the existing scale of images, increase the difficulty of local feature extraction of the initial preprocessing model by masking images, further explore the internal relationship of pre-trained images, perform feature prediction, improve the generalization ability, make the formal training effect of the model better, and the training speed faster.

[0108] Further, this embodiment provides a specific method for calculating single-channel features. Specifically, S120 includes: S121 (not shown in the drawings), S122 (not shown in the drawings), where:

[0109] S121: Divide each pre-trained image into grids based on a preset grid size to obtain a plurality of blocks;

[0110] The preset grid size is set according to the size of the pre-trained image. Specifically, it can be user-defined or set according to user experience. For example, when the size of the pre-trained image is h×w, h is the height value of the pre-trained image, and w is the width value of the pre-trained image. If the preset grid size is h g ×w g , then n h ×n w grids can be divided, where and the content in each grid is called a block, and the block size is h g ×w g .

[0111] S122: Calculate the single-channel features of each block of each pre-trained image to obtain the feature map corresponding to each pre-trained image;

[0112] In this embodiment, single-channel feature calculation is performed on each block. Specifically, the number of features calculated for each block can be custom-set by the user according to the actual situation. Correspondingly, when the number of features of each block is determined, the number of features of the feature map is also determined, specifically the product of the number of blocks and the number of features of each block. Correspondingly, the number of neurons of the second encoder is consistent with the number of features of the feature map.

[0113] Correspondingly, S131 includes: S131-1 (not shown in the drawings), S131-2 (not shown in the drawings), where:

[0114] S131-1 generates a mask for each pre-trained image based on a preset grid size;

[0115] This embodiment does not limit the method of generating the mask, and the user can set it customarily. Specifically, it can be for each pre-trained image to generate a mask where is a binary matrix, and each matrix element in the binary matrix corresponds to a block, 1 indicates covered, and 0 indicates not covered.

[0116] Specifically, a random mask can be generated for each pre-trained image based on a preset grid size. Specifically, when the training image includes: MRI image, in the actual sampling process, in order to accelerate the imaging speed of the nuclear magnetic resonance instrument, parallel imaging and undersampling techniques are often used. Among them, the undersampling technique mainly includes Cartesian sampling modes such as uniform sampling and random traversal; non-Cartesian sampling modes such as spiral and radial, and active sampling. The embodiment of the present application generates a mask in a random manner, which can be combined based on the data acquisition mode, and the production method of the mask can be set according to the acquisition mode.

[0117] S131-2 performs mask processing on each pre-trained image according to the mask to generate a masked image corresponding to each pre-trained image.

[0118] It can be seen that the embodiment of the present application performs grid division on each pre-trained image based on a preset grid size to obtain multiple blocks, and performs single-channel feature calculation based on each block to obtain a feature image corresponding to the pre-trained image, which can simplify the calculation and avoid causing the operating pressure of the electronic device due to the calculation amount of performing feature calculation on the entire pre-trained image.

[0119] Further, the pre-trained images are pre-trained two-dimensional medical images and / or natural images; the pre-trained two-dimensional medical images include any one or more of the following: CT images, MRI images. Since it is difficult to collect two-dimensional medical images, natural images captured by natural light can also be used as pre-trained images to obtain a large number of pre-trained images for training, so as to improve the model parameters of the trained pre-trained network model.

[0120] It can be understood that the model parameters of a pre-trained network model provided by an embodiment of the present application can be applied to two-dimensional medical images and can also be applied to natural images.

[0121] For the application of two-dimensional medical images, please refer to Figure 3 , Figure 3 which is a schematic flowchart of a medical image processing method provided by an embodiment of the present application, including:

[0122] S210. Obtain an initial image processing model to be trained. The initial image processing model includes a visual network to be trained and a second decoding network to be trained. The initial parameters of the visual network of the initial image processing model to be trained are the parameters of the visual network of the pre-trained network model. The pre-trained network model is obtained according to multiple pre-trained images and their corresponding feature images. Among them, the structure of the visual network of the initial image processing model is the same as that of the visual network of the initial pre-trained network model, but the second decoder of the initial image processing model is different from the first decoder of the initial pre-trained network model, and the second decoder of the initial image processing model can be selected according to specific image processing tasks.

[0123] S220. Obtain multiple two-dimensional medical images and their corresponding annotation information;

[0124] The types of two-dimensional medical images and the corresponding annotation information can be specifically determined according to the image processing task. For example, when the image processing is the target recognition of MIR images, the type of the corresponding two-dimensional medical image is MIR images, and the annotation information is specifically the target information of the two-dimensional medical image.

[0125] The ways to obtain multiple two-dimensional medical images can include: obtaining multiple two-dimensional medical images from local storage, or crawling multiple two-dimensional medical images from the network, or obtaining multiple two-dimensional medical images input by the user. Moreover, image enhancement, affine transformation, and / or elastic transformation of multiple two-dimensional medical images can also be performed, which are not limited in this embodiment, and users can set according to actual needs.

[0126] S230. Train the initial image processing model according to multiple two-dimensional medical images and their corresponding annotation information to obtain an image processing model;

[0127] Specifically, multiple two-dimensional medical images are input into an initial image processing model for image processing to obtain multiple two-dimensional medical images to be compared; based on the multiple two-dimensional medical images to be compared and their respective annotation information, the loss value of the initial image processing model is obtained, and the initial image processing model is iteratively trained based on the loss value of the initial image processing model until the loss value of the initial image processing model reaches a preset loss threshold or the preset number of iterations of training is completed, and then the current initial image processing model is determined as the trained image processing model.

[0128] Further, outputting the two-dimensional medical image to the initial image processing model for image processing may include: outputting the two-dimensional medical image to the visual network of the initial image processing model to output a training encoded feature map; inputting the training encoded feature map into a decoder for decoding processing to output a two-dimensional medical image to be compared.

[0129] S240. Obtain a two-dimensional medical image to be processed, and use an image processing model to perform image processing on the two-dimensional medical image to be processed.

[0130] The two-dimensional medical image to be processed is the medical image actually processed. In the embodiments of the present application, the obtained image processing model can be used to perform corresponding image processing on the two-dimensional medical image so as to output corresponding results, improving the processing efficiency.

[0131] Specifically, in the embodiments of the present application, using the image processing model to perform image processing on the two-dimensional medical image to be processed may include: inputting the two-dimensional medical image to be processed into the visual network of the image processing model to obtain an encoded feature map; inputting the encoded feature map into the second decoder of the image processing model to obtain an image processing result.

[0132] Further, before inputting the encoded feature map into the second decoder of the image processing model to obtain an image processing result, it may further include: inputting the encoded feature map into a feature pyramid module for processing to obtain a processed encoded feature map; correspondingly, inputting the encoded feature map into the second decoder of the image processing model to obtain an image processing result includes: inputting the processed encoded feature map into the decoder for decoding processing to obtain an image processing result.

[0133] It can be seen that in the embodiments of the present application, using the visual network parameters of the pre-trained network model as the initial parameters of the visual network to be trained in the training of the image processing model can simplify the training process of the image processing model and improve the training efficiency.

[0134] Further, the image processing includes any one of object recognition, image segmentation, and image classification.

[0135] For example, taking an MRI image as an example, when performing object recognition, the medical image recognition method may include:

[0136] S1A. The training process of the target recognition model. Specifically, obtain the initial image processing model to be trained; obtain multiple MRI images and their corresponding target annotation information; train the initial image processing model according to the multiple MRI images and their corresponding target annotation information to obtain the target recognition model.

[0137] S2A. The target recognition process. Specifically, obtain the MRI image to be recognized; input the MRI image to be recognized into the target recognition model for target recognition and output the recognition result. Among them, inputting the MRI image to be recognized into the target recognition model for target recognition includes: inputting the MRI image to be recognized into the visual network for encoding processing to obtain an encoded feature map; inputting the encoded feature map into the second decoder for decoding processing to obtain the recognition result.

[0138] For another example, taking MRI images as an example, when performing image segmentation, the medical image processing method may include:

[0139] S1B. The training process of the segmentation model. Specifically, obtain the initial image processing model to be trained; obtain multiple MRI images and their corresponding segmentation annotation information; train the initial image processing model according to the multiple MRI images and their corresponding segmentation annotation information to obtain the segmentation model.

[0140] S2B. The image segmentation process. Specifically, obtain the MRI image to be segmented; input the MRI image to be segmented into the segmentation model for segmentation and output the segmentation result. Among them, inputting the MRI image to be segmented into the segmentation model for segmentation includes: inputting the MRI image to be segmented into the visual network for encoding processing to obtain an encoded feature map; inputting the encoded feature map into the second decoder for decoding processing to obtain the segmentation result.

[0141] For another example, taking MRI images as an example, when performing image classification, the medical image processing method may include:

[0142] S1C. The training process of the classification model. Specifically, obtain the initial image processing model to be trained; obtain multiple MRI images and their corresponding class annotation information; train the initial image processing model according to the multiple MRI images and their corresponding class annotation information to obtain the classification model.

[0143] S2C. The classification process. Specifically, obtain the MRI image to be classified; input the MRI image to be classified into the classification model for target recognition and output the recognition result. Among them, inputting the MRI image to be recognized into the classification model for classification includes: inputting the MRI image to be classified into the visual network for encoding processing to obtain an encoded feature map; inputting the encoded feature map into the second decoder for decoding processing to obtain the classification result.

[0144] The following is described through a specific application scenario example for illustration. Please refer to Figure 4 , Figure 4 which is a schematic diagram of the training process of an image processing model provided by an embodiment of the present application, where:

[0145] I. Dataset collection

[0146] For the deep learning neural network training of medical images in the embodiment of the present application, the steps include pre-training and formal training, so it is necessary to collect data for the corresponding steps.

[0147] First, collect the pre-training dataset. The dataset can include a large number of non-repeating two-dimensional medical images, preferably with the same modality as the data for formal training, and no manual annotation is required. Since it is difficult to collect two-dimensional medical images, two-dimensional medical images and / or natural images of different modalities can also be collected. The pre-training dataset includes multiple pre-training images, specifically a total of n p images, denoted as For each the pre-training image needs to be scaled to a fixed h×w size and undergo simple preprocessing.

[0148] Secondly, collect the formal training dataset. The formal training dataset includes multiple two-dimensional medical images and accurate manual annotations for the corresponding tasks. Each two-dimensional medical image in the formal training dataset is scaled to a fixed h×w size and undergoes simple preprocessing. It can be understood that if it is a segmentation task, the annotation also needs to be scaled using nearest neighbor interpolation.

[0149] II. Pre-training.

[0150] Step 1, divide the grid. First, determine that the preset grid size of a single grid is h g ×w g , and each pre-training image can be divided into a total of n h ×n w grids, where in a pre-training image, the content in a grid is called a block, and the block size is h g ×w g .

[0151] Step 2, calculate Haar features. For one calculate n f Haar features for each of them, and a feature map Y f ×n h ×n w can be obtained. (i) .

[0152] Step 3: Pre-training preparation. Select a visual network, e.g., ResNet18 with its output layer removed, as the encoder of this model, denoted as E. Connect a fully-connected layer D with linear activation behind E 0 As the decoder output layer, D 0 The number of neurons in D out = n f × n h × n w .

[0153] Step 4: Conduct pre-training. During pre-training, for each pre-training image Generate a random mask where is a binary matrix, and each matrix element corresponds to a block. 1 indicates covering, and 0 indicates not covering. After generation, apply masking to the corresponding image blocks with a mask value of 1. The masking process specifically replaces the block content with matrix T, where T is a trainable parameter of size h g × w g . For details, please refer to Figure 5 "Mask" in a pre-training schematic diagram provided in an embodiment of this application. The image after applying the mask is denoted as J (i) . Then input J (i) into the initial pre-training network model to obtain an output vector Then reshape the output vector into an image of size n f × n h × n w . The target vector for pre-training is the result of Step 2, i.e., the feature map Y (i) . The loss function L (i) for pre-training is the masked mean squared error between the network output vector and the target vector Y (i) . The formula is as follows:

[0154]

[0155] where j is the channel index of the feature map, k is the pixel row number in the feature map and the mask, l is the pixel column number in the feature map and the mask, The symbol represents element-wise multiplication of the corresponding positions of the matrices.

[0156] Perform backpropagation operations according to this loss function, and loop Step 4 until the entire pre-training set is traversed Take this as one pre-training cycle until the pre-training network model is obtained.

[0157] Specifically in combination with Figure 5A description is given. Specifically, Haar features are calculated for the pre-trained image (the original image "Figure") to obtain the target vector (feature map); a random mask is applied to the pre-trained image, and the image after applying the mask is used as the input image and input into the initial pre-trained network model including a visual network and a linear fully-connected layer to obtain the output vector (the image to be compared), and training is performed based on the target vector, the output vector, and the loss function to obtain the pre-trained network model.

[0158] Step Five: Save the parameters for convenient calling during formal training.

[0159] Among them, the pre-trained network model includes: visual network parameters and first decoder parameters. The visual network parameters of the pre-trained network model can be used as the initial parameters of the visual network to be trained in the training of the image processing model.

[0160] III. Formal Training

[0161] During formal training, the pre-trained model parameters are called. According to the task of image processing, a suitable second decoder is selected and denoted as D 1 . Then the forward propagation calculation process of the neural network becomes D 1 (E(J (i) ))). Finally, formal training is performed based on the formal training data and the initial image processing model to obtain the image processing model.

[0162] Next, a training device for a pre-trained network model provided in an embodiment of the present application is introduced. The training device for the pre-trained network model described below can be correspondingly referred to the training method for the pre-trained network model described above. The training device for the pre-trained network model in this embodiment is set in an electronic device. Refer to Figure 6 , Figure 6 which is the structural block diagram of a training device for a pre-trained network model provided in an embodiment of the present application, including:

[0163] A pre-training information acquisition module 610, configured to acquire an initial pre-trained network model and multiple pre-trained images. The initial pre-trained network model includes a visual network to be trained and a first decoding network to be trained;

[0164] A feature calculation module 620, configured to perform single-channel feature calculation on each pre-trained image to obtain a feature image corresponding to each pre-trained image;

[0165] A pre-training module 630, configured to train the initial pre-trained network model according to each pre-trained image and its corresponding feature image to obtain the pre-trained network model;

[0166] Among them, the visual network parameters of the pre-trained network model can be used as the initial parameters of the visual network to be trained in the training of the image processing model, and the image processing model is used to perform image processing on the two-dimensional medical image to be processed.

[0167] In a possible implementation, the pre-training module includes:

[0168] A mask processing unit, configured to perform mask processing on each pre-training image to generate a masking map corresponding to each pre-training image; an output unit, configured to input each masking map into the initial pre-training network model and output a comparison image corresponding to each masking map; a loss value calculation unit, configured to calculate, according to the comparison image corresponding to each masking map and the respective corresponding feature map, using the masked mean square error loss function, to obtain a loss value;

[0169] A pre-training network model acquisition unit, configured to determine whether the initial pre-training network model converges according to the loss value; if so, determine the current initial pre-training network model as the pre-training network model that has been trained; if not, perform iterative training until the initial pre-training network model converges to obtain the pre-training network model.

[0170] In a possible implementation, the feature calculation module includes:

[0171] A grid division unit, configured to perform grid division on each pre-training image based on a preset grid size to obtain a plurality of blocks; a feature calculation unit, configured to calculate single-channel features for each block of each pre-training image to obtain a feature map corresponding to each pre-training image.

[0172] Correspondingly, the mask processing unit includes:

[0173] A mask generation subunit, configured to generate a mask for each pre-training image based on a preset grid size;

[0174] A mask processing subunit, configured to perform mask processing on each pre-training image according to the mask to generate a masking image corresponding to each pre-training image.

[0175] In a possible implementation, the pre-training image is a pre-training two-dimensional medical image and / or a natural image; the pre-training two-dimensional medical image includes any one or more of the following: CT image, MRI image.

[0176] In a possible implementation, the feature calculation module includes:

[0177] A feature calculation unit, configured to perform target feature calculation on each pre-training image to obtain a feature image corresponding to each pre-training image, where the target feature includes any one of Haar feature, Gabor feature, and LBP feature.

[0178] The following introduces a medical image processing device provided by an embodiment of the present application. The medical image processing device described below can be referred to in correspondence with the medical image processing method described above. The medical image processing device of this embodiment is set in an electronic device. Refer to Figure 7 , Figure 7 which is a structural block diagram of a medical image processing device provided by an embodiment of the present application, including:

[0179] An initial image processing model acquisition module 710, configured to acquire an initial image processing model to be trained. The initial image processing model includes a visual network to be trained and a second decoding network to be trained. The initial parameters of the visual network of the initial image processing model to be trained are the parameters of the visual network of a pre-trained network model. The pre-trained network model is obtained based on multiple pre-trained images and their respective corresponding feature images;

[0180] A training set acquisition module 720, configured to acquire multiple two-dimensional medical images and their respective corresponding annotation information;

[0181] An image processing model obtaining module 730, configured to train the initial image processing model according to multiple two-dimensional medical images and their respective corresponding annotation information to obtain an image processing model;

[0182] An image processing module 740, configured to acquire a two-dimensional medical image to be processed, and perform image processing on the two-dimensional medical image to be processed by using the image processing model.

[0183] In a possible implementation manner, when the image processing module 740 performs image processing on the two-dimensional medical image to be processed by using the image processing model, it is configured to:

[0184] Input the two-dimensional medical image to be processed into the visual network of the image processing model to obtain an encoded feature map;

[0185] Input the encoded feature map into the second decoder of the image processing model to obtain an image processing result.

[0186] In a possible implementation manner, the image processing includes any one of object recognition, image segmentation, and image classification.

[0187] The following introduces an electronic device provided by an embodiment of the present application. The electronic device described below can be referred to in correspondence with the method described above.

[0188] An embodiment of the present application provides an electronic device, as Figure 8 shown, Figure 8 which is a structural block diagram of an electronic device provided by an embodiment of the present application. Specifically, Figure 8The electronic device 800 shown includes: a processor 801 and a memory 803. Among them, the processor 801 and the memory 803 are connected, such as through a bus 802. Optionally, the electronic device 800 may further include a transceiver 804. It should be noted that in practical applications, the transceiver 804 is not limited to one, and the structure of the electronic device 800 does not constitute a limitation to the embodiments of the present application.

[0189] The processor 801 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 801 may also be a combination that implements computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0190] The bus 802 may include a path for transmitting information between the above components. The bus 802 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 802 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 8 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0191] The memory 803 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0192] The memory 803 is used to store the application program code for executing the solution of this application, and is controlled by the processor 801 for execution. The processor 801 is used to execute the application program code stored in the memory 803 to implement the training method of the pre-trained network model shown in the foregoing method embodiments, or to implement the medical image processing method shown in the foregoing method embodiments when executed.

[0193] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The electronic device shown is only an example and should not bring any limitations to the functions and usage scope of the embodiments of this application.

[0194] Next, a computer-readable storage medium provided by the embodiments of this application will be introduced. The computer-readable storage medium described below can be correspondingly referred to the method described above.

[0195] The embodiments of this application provide a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the training method of the pre-trained network model shown in the foregoing method embodiments, or implements the medical image processing method shown in the foregoing method embodiments.

[0196] Since the embodiments of the computer-readable storage medium part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the computer-readable storage medium part, and will not be elaborated here for the time being.

[0197] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this document, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0198] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A training method for a pre-trained network model, characterized in that, it includes: Obtain an initial pre-trained network model and multiple pre-training images, where the initial pre-trained network model includes a visual network to be trained and a first decoding network to be trained; Perform single-channel feature calculation on each of the pre-training images to obtain a feature image corresponding to each of the pre-training images; Train the initial pre-trained network model according to each of the pre-training images and their corresponding feature images to obtain the pre-trained network model; Among them, the visual network parameters of the pre-trained network model can be used as the initial parameters of the visual network to be trained in the training of the image processing model, and the image processing model is used to perform image processing on the two-dimensional medical image to be processed.

2. The training method for the pre-trained network model according to claim 1, characterized in that, The training the initial pre-trained network model according to each of the pre-training images and their corresponding feature images to obtain the pre-trained network model includes: Perform masking processing on each of the pre-training images to generate a masking map corresponding to each of the pre-training images; Input each of the masking maps into the initial pre-trained network model and output a comparison image corresponding to each of the masking maps; Calculate according to the comparison image corresponding to each of the masking maps and their corresponding feature maps using the masked mean square error loss function to obtain a loss value; Determine whether the initial pre-trained network model converges according to the loss value; if so, determine the current initial pre-trained network model as the pre-trained network model that has completed training; if not, perform iterative training until the initial pre-trained network model converges to obtain the pre-trained network model.

3. The training method for the pre-trained network model according to claim 2, characterized in that, The performing single-channel feature calculation on each of the pre-training images to obtain a feature image corresponding to each of the pre-training images includes: Perform grid division on each of the pre-training images based on a preset grid size to obtain a plurality of blocks; Perform single-channel feature calculation on each block of each of the pre-training images to obtain a feature map corresponding to each of the pre-training images; Correspondingly, the performing masking processing on each of the pre-training images to generate a masking map corresponding to each of the pre-training images includes: Generate a mask for each of the pre-training images based on a preset grid size; Perform masking processing on each of the pre-training images according to the mask to generate a masking image corresponding to each of the pre-training images.

4. The training method for the pre-trained network model according to claim 1, characterized in that, The pre-training images are pre-training two-dimensional medical images and / or natural images; The pre-training two-dimensional medical images include any one or more of the following: CT images, MRI images.

5. The training method for the pre-trained network model according to any one of claims 1 to 4, characterized in that, The performing single-channel feature calculation on each of the pre-training images to obtain a single-channel feature of the feature image corresponding to each of the pre-training images includes: Calculate the target features for each of the pre-trained images to obtain the feature images corresponding to the pre-trained images, where the target features include any one of Haar features, Gabor features, and LBP features.

6. A medical image processing method Characterized in that It includes: Obtain an initial image processing model to be trained, where the initial image processing model includes a visual network to be trained and a second decoding network to be trained, and the initial parameters of the visual network of the initial image processing model to be trained are the parameters of the visual network of the pre-trained network model; the pre-trained network model is obtained based on multiple pre-trained images and their corresponding feature images; Obtain multiple two-dimensional medical images and their corresponding annotation information; Train the initial image processing model according to the multiple two-dimensional medical images and their corresponding annotation information to obtain the image processing model; Obtain a two-dimensional medical image to be processed, and use the image processing model to perform image processing on the two-dimensional medical image to be processed.

7. The medical image processing method according to claim 6, Characterized in that The using the image processing model to perform image processing on the two-dimensional medical image to be processed includes: Input the two-dimensional medical image to be processed into the visual network of the image processing model to obtain an encoded feature map; Input the encoded feature map into the second decoder of the image processing model to obtain an image processing result.

8. The medical image processing method according to claim 6, Characterized in that The image processing includes any one of target recognition, image segmentation, and image classification.

9. An electronic device Characterized in that It includes: One or more processors; A memory; One or more applications, where the one or more applications are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs are configured to: execute the training method of the pre-trained network model according to any one of claims 1 to 5, or execute the medical image processing method according to any one of claims 6 to 8.

10. A computer-readable storage medium Characterized in that A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the training method of the pre-trained network model according to any one of claims 1 to 5, or executes the medical image processing method according to any one of claims 6 to 8.

Citation Information

Patent Citations

  • Model training and image processing method and device, medium and electronic equipment

    CN110163237A

  • Facial action recognition model training method and facial action recognition method

    CN110909595A