Training method, device, equipment and storage medium for medical image recognition model
Through the mask autoencoder pre-training and semi-supervised learning method, combined with the downstream medical task decoder for model stitching and fine-tuning training, the problems of long training time of existing medical image recognition models and dependence on annotation samples are solved, and the effect of structure simplification and real-time improvement is achieved.
Patent Information
- Application Number
- CN202210713767.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-06-22
AI Technical Summary
The existing medical image recognition models require a large number of labeled samples during the training process and the model structure is complex, resulting in a long training time and is not conducive to real-time.
The mask autoencoder pre-training method is used to train the visual transformation model through semi-supervised learning, extract the encoder weight parameters and migrate it to the local encoder, and combine it with the downstream medical task decoder for model stitching and fine-tuning training to reduce dependence on the annotated samples.
The structure of the medical image recognition model is simplified, the number of labeled samples required for training is reduced, and the real-time and training efficiency of the model are improved.
Smart Images

Figure CN115205225B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a training method, device, equipment and computer-readable storage medium for a medical image recognition model. Background Art
[0002] At present, downstream medical tasks include a variety of image recognition, such as X-ray scan images, CT images, and MRI brain segmentation images. In order to meet the needs of various image tasks, a huge amount of medical images of various types are required for the training process when building a medical image recognition model. In the training process, the convolutional neural network based on the attention mechanism is generally fully supervised, which leads to the need for a large number of labeled samples, and each labeled sample needs to be labeled by manpower, which greatly limits the construction of the medical image recognition model. In addition, the convolutional neural network based on the attention mechanism has a complex structure and a large amount of calculation, which makes the medical recognition model take a long time to process data, which is not conducive to the real-time performance of downstream medical tasks. Therefore, there is an urgent need for a training method with a simple structure that does not require a large number of labeled samples to train the medical image recognition model. Summary of the invention
[0003] The present invention provides a training method, device, equipment and storage medium for a medical image recognition model, the main purpose of which is to optimize the structure of the medical image recognition model and the required number of samples.
[0004] To achieve the above object, the present invention provides a training method for a medical image recognition model, comprising:
[0005] According to the mask autoencoder pre-training method, a pre-constructed visual conversion model is semi-supervised trained using a pre-constructed labeled sample set and an unlabeled sample set to obtain a trained visual conversion model;
[0006] Extracting encoder weight parameters in the visual conversion model, and migrating the encoder weight parameters to a pre-built local encoder to obtain a local update encoder;
[0007] Calling a preset downstream medical task decoder, connecting the local update encoder and the downstream medical task decoder to obtain a pre-trained medical image recognition model;
[0008] The pre-trained medical image recognition model is trained using the task sample set corresponding to the downstream medical task decoder to obtain a trained medical image recognition model.
[0009] Optionally, the pre-training method of the mask autoencoder is used to perform semi-supervised training on the pre-constructed visual conversion model using the pre-constructed labeled sample set and the unlabeled sample set to obtain the trained visual conversion model, including:
[0010] The visual transformation model is trained using pre-built labeled samples;
[0011] According to a preset masking ratio, each sample data in the pre-constructed unlabeled sample set is masked to obtain a masked data set;
[0012] Using the visual conversion model to identify the masked data set, and obtaining a recognition result and an accuracy probability score corresponding to each masked data;
[0013] Obtain recognition results corresponding to masked data whose accuracy probability scores are greater than or equal to a preset qualified threshold value as unlabeled samples, and use the recognition results as pseudo-labels of the unlabeled samples to obtain pseudo-labeled samples;
[0014] Calculating an average value of each of the accuracy probability scores, and determining the convergence of the average value;
[0015] When the average value does not converge, the pseudo-annotated samples are imported into the annotated samples, and the step of training the visual conversion model using the pre-constructed annotated samples is returned to update the visual conversion model.
[0016] When the average value converges, a trained visual conversion model is obtained.
[0017] Optionally, the visual conversion model obtained by training using pre-built labeled samples includes:
[0018] Splitting the labeled samples into images to obtain a set of blocks of the labeled samples;
[0019] Using a preset mapping matrix to perform one-dimensional mapping on each tile in the tile set to obtain a tile vector set;
[0020] Use the pre-built fully connected layer to connect and arrange the tile vector set to obtain a two-dimensional vector;
[0021] According to the two-dimensional vector, a column feature extraction operation and a row feature extraction operation are successively performed to obtain a position relationship feature vector and a channel relationship feature vector respectively;
[0022] The visual conversion model is trained using the Gaussian error linear loss function, the position relationship feature vector and the channel relationship feature vector.
[0023] Optionally, the using the task sample set corresponding to the downstream medical task decoder to train the pre-trained medical image recognition model to obtain a trained medical image recognition model includes:
[0024] Using a pre-trained medical image recognition model to perform model prediction on a task sample set corresponding to the downstream medical task decoder to obtain a prediction result set;
[0025] According to a preset loss function, the loss value of the predicted result set and the actual result of each task sample in the task sample set is calculated;
[0026] Minimize the loss value to obtain the network parameters when the loss value is the smallest, and use the network parameters to perform network back propagation to update the pre-trained medical image recognition model;
[0027] Determine whether the loss value is less than a preset threshold;
[0028] When the loss value is greater than or equal to the preset threshold, return to the step of using the pre-trained medical image recognition model to perform model prediction on the task sample set corresponding to the downstream medical task decoder to obtain a prediction result set, and iteratively update the pre-trained medical image recognition model;
[0029] When the loss value is less than the preset threshold, a trained medical image recognition model is obtained.
[0030] Optionally, extracting encoder weight parameters in the visual conversion model and migrating the encoder weight parameters to a pre-built local encoder to obtain a local update encoder includes:
[0031] Use the torch.save function package to obtain the storage address of the encoder weight parameters of the visual conversion model;
[0032] According to the storage address, the encoder weight parameters in the storage address are migrated to the pre-built local encoder using the torch.load function package to obtain a local update encoder.
[0033] In order to solve the above problems, the present invention also provides a training device for a medical image recognition model, the device comprising:
[0034] A semi-supervised learning module is used to perform semi-supervised training on a pre-constructed visual conversion model using a pre-constructed labeled sample set and an unlabeled sample set according to a mask autoencoder pre-training method to obtain a trained visual conversion model;
[0035] A knowledge transfer module, used for extracting encoder weight parameters in the visual conversion model, and migrating the encoder weight parameters to a pre-built local encoder to obtain a local update encoder;
[0036] A model splicing module, used to call a preset downstream medical task decoder, connect the local update encoder and the downstream medical task decoder, and obtain a pre-trained medical image recognition model;
[0037] The fine-tuning training module is used to train the pre-trained medical image recognition model using the task sample set corresponding to the downstream medical task decoder to obtain a trained medical image recognition model.
[0038] Optionally, the pre-training method of the mask autoencoder is used to perform semi-supervised training on the pre-constructed visual conversion model using the pre-constructed labeled sample set and the unlabeled sample set to obtain the trained visual conversion model, including:
[0039] The visual transformation model is trained using pre-built labeled samples;
[0040] According to a preset masking ratio, each sample data in the pre-constructed unlabeled sample set is masked to obtain a masked data set;
[0041] Using the visual conversion model to identify the masked data set, and obtaining a recognition result and an accuracy probability score corresponding to each masked data;
[0042] Obtain recognition results corresponding to masked data whose accuracy probability scores are greater than or equal to a preset qualified threshold value as unlabeled samples, and use the recognition results as pseudo-labels of the unlabeled samples to obtain pseudo-labeled samples;
[0043] Calculating an average value of each of the accuracy probability scores, and determining the convergence of the average value;
[0044] When the average value does not converge, the pseudo-annotated samples are imported into the annotated samples, and the step of training the visual conversion model using the pre-constructed annotated samples is returned to update the visual conversion model.
[0045] When the average value converges, a trained visual conversion model is obtained.
[0046] Optionally, the visual conversion model obtained by training using pre-built labeled samples includes:
[0047] Splitting the labeled samples into images to obtain a set of blocks of the labeled samples;
[0048] Using a preset mapping matrix to perform one-dimensional mapping on each tile in the tile set to obtain a tile vector set;
[0049] Use the pre-built fully connected layer to connect and arrange the tile vector set to obtain a two-dimensional vector;
[0050] According to the two-dimensional vector, a column feature extraction operation and a row feature extraction operation are successively performed to obtain a position relationship feature vector and a channel relationship feature vector respectively;
[0051] The visual conversion model is trained using the Gaussian error linear loss function, the position relationship feature vector and the channel relationship feature vector.
[0052] In order to solve the above problem, the present invention further provides an electronic device, the electronic device comprising:
[0053] at least one processor; and,
[0054] a memory communicatively connected to the at least one processor; wherein,
[0055] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the training method of the medical image recognition model described above.
[0056] In order to solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is executed by a processor in an electronic device to implement the above-mentioned training method of the medical image recognition model.
[0057] The embodiment of the present invention obtains a visual conversion model through semi-supervised learning through a mask autoencoder pre-training method, wherein the mask autoencoder adopts a multi-layer perception mechanism, can divide the image into blocks, and splice each block through a fully connected layer, and finally extracts the features of the columns and rows of the splicing results based on the image position relationship and channel domain relationship to obtain the medical image features. The present invention changes the previous method of extracting image features through the attention mechanism, simplifies the complexity of the medical image recognition model, and increases the real-time performance; in addition, the sample order of magnitude required for semi-supervised learning is lower, which is conducive to the training process of the medical image recognition model. Therefore, the training method, device, equipment and storage medium of a medical image recognition model provided by the embodiment of the present invention can construct a medical image recognition model through a method with a simpler structure and fewer training samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1A schematic diagram of a flow chart of a method for training a medical image recognition model provided by an embodiment of the present invention;
[0059] Figure 2 A detailed flowchart of a step in a method for training a medical image recognition model provided by an embodiment of the present invention;
[0060] Figure 3 A detailed flowchart of a step in a method for training a medical image recognition model provided by an embodiment of the present invention;
[0061] Figure 4 A detailed flowchart of a step in a method for training a medical image recognition model provided by an embodiment of the present invention;
[0062] Figure 5 A detailed flowchart of a step in a method for training a medical image recognition model provided by an embodiment of the present invention;
[0063] Figure 6 A functional module diagram of a training device for a medical image recognition model provided by an embodiment of the present invention;
[0064] Figure 7 A schematic diagram of the structure of an electronic device for implementing the training method of the medical image recognition model provided by one embodiment of the present invention.
[0065] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0066] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0067] The embodiment of the present application provides a training method for a medical image recognition model. In the embodiment of the present application, the execution subject of the training method of the medical image recognition model includes but is not limited to at least one of the electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the training method of the medical image recognition model can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.
[0068] Reference Figure 1 FIG. 1 is a flow chart of a method for training a medical image recognition model according to an embodiment of the present invention. In this embodiment, the method for training a medical image recognition model includes steps S1 to S4:
[0069] S1. According to the mask autoencoder pre-training method, a pre-constructed labeled sample set and an unlabeled sample set are used to perform semi-supervised training on the pre-constructed visual conversion model to obtain a trained visual conversion model.
[0070] In an embodiment of the present invention, the mask autoencoder pre-training is a computer vision self-supervised learning method, which is a method of randomly masking part of the pixels of the image and then reconstructing the masked part. Among them, the mask autoencoder in the embodiment of the present invention includes an encoder (Encoder) based on a multi-layer perception mechanism (MLP) and a standard lightweight transformer decoder (TransformerDecoder). Among them, the MLPEncoder includes a block layer (Per-patch), a fully connected layer (Fully-connected), and a mixing layer (Mixer Layer).
[0071] For more details, please refer to Figure 2 As shown, in the embodiment of the present invention, the pre-training method of the mask autoencoder uses a pre-constructed labeled sample set and an unlabeled sample set to perform semi-supervised training on the pre-constructed visual conversion model to obtain a trained visual conversion model, including steps S11-S17:
[0072] S11, using pre-built annotated samples to train the visual conversion model;
[0073] S12, masking each sample data in the pre-constructed unlabeled sample set according to a preset masking ratio to obtain a masked data set;
[0074] S13, using the visual conversion model to identify the masked data set, and obtaining a recognition result and an accuracy probability score corresponding to each masked data;
[0075] S14, obtaining the recognition result corresponding to the masked data whose accuracy probability score is greater than or equal to the preset qualified threshold value as the unlabeled sample, and using the recognition result as the pseudo-label of the unlabeled sample to obtain the pseudo-labeled sample;
[0076] S15, calculating the average value of each of the accuracy probability scores, and determining the convergence of the average value;
[0077] When the average value has not converged, S16, importing the pseudo-annotated samples into the annotated samples, and returning to the step of training the visual conversion model using the pre-constructed annotated samples to update the visual conversion model;
[0078] When the average value converges, S17, a trained visual conversion model is obtained.
[0079] The embodiment of the present invention first trains the Transformer model with labeled samples, and then uses the Transformer model to identify the unlabeled sample set to obtain multiple pseudo-recognition results corresponding to each unlabeled sample and the accuracy probability score corresponding to each pseudo-recognition result. In the embodiment of the present invention, the pseudo-recognition result with the highest accuracy probability score is used as the final recognition result of the unlabeled sample, and the accuracy probability score of the recognition result is counted.
[0080] When the accuracy probability score is very high, it indicates that the model can clearly identify the image. When the accuracy probability score is not high, it indicates that the model cannot clearly identify the image. Therefore, in an embodiment of the present invention, the unlabeled samples with an accuracy probability score greater than a preset qualified threshold, such as 85%, are used as pseudo-labeled samples, and the pseudo-labeled samples are used together with the labeled samples to train the Transformer model.
[0081] In the embodiment of the present invention, the prediction results of each unlabeled sample set are recorded, and the average value of each accuracy probability score is obtained as the result evaluation standard of this training. When the average value converges, it indicates that the recognition result of the Transformer model for the image is gradually stabilized, and the training of the Transformer model is completed.
[0082] For further reference, Figure 3 As shown, in the embodiment of the present invention, the visual conversion model obtained by training using pre-built labeled samples includes steps S101-S105:
[0083] S101, splitting the labeled samples into images to obtain a set of blocks of the labeled samples;
[0084] S102, performing one-dimensional mapping on each tile in the tile set using a preset mapping matrix to obtain a tile vector set;
[0085] S103, using a pre-built fully connected layer to connect and arrange the tile vector set to obtain a two-dimensional vector;
[0086] S104, performing a column feature extraction operation and a row feature extraction operation on the two-dimensional vector respectively, to obtain a position relationship feature vector and a channel relationship feature vector respectively;
[0087] S105 , using the Gaussian error linear loss function, the position relationship feature vector and the channel relationship feature vector to train and obtain a visual conversion model.
[0088] The visual conversion model described in the embodiment of the present invention is not an attention-based convolutional neural network in a traditional medical image recognition model, but a model of a multi-perception layer mechanism.
[0089] Specifically, in an embodiment of the present invention, according to a preset patch (tile area, such as 16*16) size, the medical image is divided into S tiles without overlap, and then each tile is one-dimensionally mapped according to a preset mapping matrix to obtain a one-dimensional vector of length C. In an embodiment of the present invention, the S one-dimensional vectors are connected according to a fully connected layer to obtain a two-dimensional vector of S*C.
[0090] Assuming that the input medical image size is 240*240*3, and the Patch selected by the model is 16*16, then a medical image can be divided into (240*240) / (16*16)=225 patches; combined with the number of channels of the medical image being 3, each patch contains 16*16*3=768 values, and these 768 values are flattened as the input of the MLP, where the number of neurons in the output layer of the MLP is 128. In this way, each patch can obtain a feature vector of length 128, and a two-dimensional vector of 225*128 can be obtained by combining.
[0091] Then, the token-mixing MLPs and channel-mixing MLPs in the mixing layer are used to perform column and row feature extraction on the two-dimensional vector respectively, wherein the token-mixing MLPs is used to extract spatial domain information from the two-dimensional vector, and the channel-mixing MLPs is used to extract channel domain information from the two-dimensional vector.
[0092] Then, the embodiment of the present invention trains the visual conversion model through Gaussian Error Linear Units (GELU) and the position relationship feature vector and the channel relationship feature vector. The GELU is a commonly used loss function in training the Transformer model, which will not be described in detail here.
[0093] S2. Extracting encoder weight parameters in the visual conversion model, and migrating the encoder weight parameters to a pre-built local encoder to obtain a local update encoder.
[0094] For more details, please refer to Figure 4 As shown, in the embodiment of the present invention, extracting the encoder weight parameters in the visual conversion model and migrating the encoder weight parameters to the pre-built local encoder to obtain the local update encoder includes steps S21-S22:
[0095] S21, using the torch.save function package to obtain the storage address of the encoder weight parameters of the visual conversion model;
[0096] S22. According to the storage address, using the torch.load function package, migrate the encoder weight parameters in the storage address to the pre-built local encoder to obtain a local update encoder.
[0097] In an embodiment of the present invention, pytorch is used to migrate the encoder weight parameters in the model, wherein pytorch is an open source Python machine learning library, which includes multiple model processing function packages, such as torch.load and torch.save.
[0098] In the embodiment of the present invention, the encoder weight parameters are saved according to torch.save(path), where path is the file save path. Then, according to model=torch.load(path), the encoder weight parameters are migrated to the pre-built local encoder to obtain a local update encoder. In order to maintain the flexible operation capability of the model, the present invention uses torch.save(model.module.state_dict()) to only save the values of the encoder weight parameters.
[0099] S3. Call a preset downstream medical task decoder, connect the local update encoder and the downstream medical task decoder, and obtain a pre-trained medical image recognition model.
[0100] With the development of medical technology, there are now various medical tasks such as classification and segmentation, such as UNETR segmentation.
[0101] The embodiment of the present invention can obtain a pre-built UNETRDecoder to connect with the local update encoder, wherein the local update encoder migrates the medical knowledge in the mask autoencoder, so that a pre-trained medical image recognition model can be obtained.
[0102] In addition, in the embodiment of the present invention, in the step S3, the local encoder can be directly the Encoder in the UNETR segmentation model, so that the solution does not need to take out the data of the UNETRDecoder and directly obtains the pre-trained medical image recognition model.
[0103] S4. Using the task sample set corresponding to the downstream medical task decoder, the pre-trained medical image recognition model is trained to obtain a trained medical image recognition model.
[0104] In an embodiment of the present invention, it is only necessary to use the task sample set corresponding to the downstream medical task decoder to fine-tune the above-mentioned pre-trained medical image recognition model to obtain a trained medical image recognition model, wherein only a small number of labeled sample sets and task sample sets are required, which greatly reduces the cost of manual labeling and increases the speed of model building.
[0105] For more details, please refer to Figure 5 As shown, in an embodiment of the present invention, the task sample set corresponding to the downstream medical task decoder is used to train the pre-trained medical image recognition model to obtain a trained medical image recognition model, including steps S41-S45:
[0106] S41, using a pre-trained medical image recognition model to perform model prediction on a task sample set corresponding to the downstream medical task decoder to obtain a prediction result set;
[0107] S42, calculating the loss value of the predicted result set and the actual result of each task sample in the task sample set according to a preset loss function;
[0108] S43, minimizing the loss value to obtain the network parameters when the loss value is the smallest, and using the network parameters to perform network back propagation to update the pre-trained medical image recognition model;
[0109] S44, determining whether the loss value is less than a preset threshold;
[0110] When the loss value is greater than or equal to the preset threshold, return to the step of using the pre-trained medical image recognition model to perform model prediction on the task sample set corresponding to the downstream medical task decoder to obtain a prediction result set, and iteratively update the pre-trained medical image recognition model;
[0111] When the loss value is less than the preset threshold, S45, a trained medical image recognition model is obtained.
[0112] The embodiment of the present invention fine-tunes the pre-trained medical image recognition model through a task sample set to obtain a prediction result set; then, through a preset loss function, according to the prediction result set, the loss value of the model is calculated, and the network parameters when the loss value is the smallest are measured, and then, through the BP neural network structure pre-constructed in the pre-trained medical image recognition model, the network parameters are back-propagated to update the pre-trained medical image recognition model. When the loss value reaches a preset threshold, such as 0.15, it indicates that the prediction result of the pre-trained medical image recognition model is basically the true result, and the training process is completed to obtain the medical image recognition model.
[0113] The embodiment of the present invention obtains a visual conversion model through semi-supervised learning through a mask autoencoder pre-training method, wherein the mask autoencoder adopts a multi-layer perception mechanism, can divide the image into blocks, and splice each block through a fully connected layer, and finally extracts the features of the columns and rows of the splicing results based on the image position relationship and channel domain relationship to obtain the medical image features. The present invention changes the previous method of extracting image features through the attention mechanism, simplifies the complexity of the medical image recognition model, and increases the real-time performance; in addition, the sample order of magnitude required for semi-supervised learning is lower, which is conducive to the training process of the medical image recognition model. Therefore, the training method of a medical image recognition model provided by the embodiment of the present invention can construct a medical image recognition model through a method with a simpler structure and fewer training samples.
[0114] like Figure 6 , which is a functional module diagram of a training device for a medical image recognition model provided by an embodiment of the present invention.
[0115] The training device 100 of the medical image recognition model of the present invention can be installed in an electronic device. According to the functions to be implemented, the training device 100 of the medical image recognition model can include a semi-supervised learning module 101, a knowledge transfer module 102, a model splicing module 103 and a fine-tuning training module 104. The module of the present invention can also be called a unit, which refers to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, which are stored in the memory of the electronic device.
[0116] In this embodiment, the functions of each module / unit are as follows:
[0117] The semi-supervised learning module 101 is used to perform semi-supervised training on the pre-constructed visual conversion model using the pre-constructed labeled sample set and the unlabeled sample set according to the mask autoencoder pre-training method to obtain a trained visual conversion model;
[0118] The knowledge transfer module 102 is used to extract encoder weight parameters in the visual conversion model and transfer the encoder weight parameters to a pre-built local encoder to obtain a local update encoder;
[0119] The model splicing module 103 is used to call a preset downstream medical task decoder, connect the local update encoder and the downstream medical task decoder, and obtain a pre-trained medical image recognition model;
[0120] The fine-tuning training module 104 is used to train the pre-trained medical image recognition model using the task sample set corresponding to the downstream medical task decoder to obtain a trained medical image recognition model.
[0121] In detail, each module described in the training device 100 of the medical image recognition model in the embodiment of the present application is used in the same manner as described above. Figures 1 to 5 The training method of the medical image recognition model described in the present invention is the same as the technical means and can produce the same technical effect, so I will not go into details here.
[0122] like Figure 7 , is a schematic diagram of the structure of an electronic device 1 for implementing a training method for a medical image recognition model provided by an embodiment of the present invention.
[0123] The electronic device 1 may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a training program for a medical image recognition model.
[0124] In some embodiments, the processor 10 may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips. The processor 10 is the control core (ControlUnit) of the electronic device 1, and uses various interfaces and lines to connect various components of the entire electronic device, and executes or executes programs or modules stored in the memory 11 (for example, executing a training program for a medical image recognition model, etc.), and calls data stored in the memory 11 to execute various functions of the electronic device and process data.
[0125] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of an electronic device, such as a mobile hard disk of the electronic device. In other embodiments, the memory 11 may also be an external storage device of an electronic device, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device. Further, the memory 11 may also include both an internal storage unit of the electronic device and an external storage device. The memory 11 can not only be used to store application software and various types of data installed in the electronic device, such as the code of the training program of the medical image recognition model, but also can be used to temporarily store data that has been output or is to be output.
[0126] The communication bus 12 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to realize connection and communication between the memory 11 and at least one processor 10, etc.
[0127] The communication interface 13 is used for communication between the above-mentioned electronic device 1 and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device and other electronic devices. The user interface may be a display (Display), an input unit (such as a keyboard (Keyboard)), and optionally, the user interface may also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode, organic light-emitting diode) touch device, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device and to display a visual user interface.
[0128] Figure 7 Only an electronic device with components is shown, and those skilled in the art will understand that Figure 7The structure shown does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0129] For example, although not shown, the electronic device 1 may also include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so that the power management device can realize functions such as charging management, discharging management, and power consumption management. The power source may also include any components such as one or more DC or AC power sources, recharging devices, power failure detection circuits, power converters or inverters, power status indicators, etc. The electronic device 1 may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be repeated here.
[0130] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.
[0131] The training program of the medical image recognition model stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can achieve:
[0132] According to the mask autoencoder pre-training method, a pre-constructed visual conversion model is semi-supervised trained using a pre-constructed labeled sample set and an unlabeled sample set to obtain a trained visual conversion model;
[0133] Extracting encoder weight parameters in the visual conversion model, and migrating the encoder weight parameters to a pre-built local encoder to obtain a local update encoder;
[0134] Calling a preset downstream medical task decoder, connecting the local update encoder and the downstream medical task decoder to obtain a pre-trained medical image recognition model;
[0135] The pre-trained medical image recognition model is trained using the task sample set corresponding to the downstream medical task decoder to obtain a trained medical image recognition model.
[0136] Specifically, the specific implementation method of the processor 10 for the above instructions can refer to the description of the relevant steps in the corresponding embodiment of the accompanying drawings, which will not be repeated here.
[0137] Furthermore, if the module / unit integrated in the electronic device 1 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).
[0138] The present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor of an electronic device, the computer program can implement:
[0139] According to the mask autoencoder pre-training method, a pre-constructed visual conversion model is semi-supervised trained using a pre-constructed labeled sample set and an unlabeled sample set to obtain a trained visual conversion model;
[0140] Extracting encoder weight parameters in the visual conversion model, and migrating the encoder weight parameters to a pre-built local encoder to obtain a local update encoder;
[0141] Calling a preset downstream medical task decoder, connecting the local update encoder and the downstream medical task decoder to obtain a pre-trained medical image recognition model;
[0142] The pre-trained medical image recognition model is trained using the task sample set corresponding to the downstream medical task decoder to obtain a trained medical image recognition model.
[0143] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0144] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0145] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional modules.
[0146] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0147] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present invention is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any attached figure mark in the claims should not be regarded as limiting the claims involved.
[0148] The blockchain referred to in the present invention is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the blockchain underlying platform, platform product service layer, and application service layer.
[0149] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0150] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the system claim can also be implemented by one unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.
[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. A training method for a medical image recognition model, characterized in that: The method comprises: According to the mask autoencoder pre-training method, a pre-constructed visual conversion model is semi-supervised trained using a pre-constructed labeled sample set and an unlabeled sample set to obtain a trained visual conversion model; Extracting encoder weight parameters in the visual conversion model, and migrating the encoder weight parameters to a pre-built local encoder to obtain a local update encoder; Calling a preset downstream medical task decoder, connecting the local update encoder and the downstream medical task decoder to obtain a pre-trained medical image recognition model; Using the task sample set corresponding to the downstream medical task decoder, the pre-trained medical image recognition model is trained to obtain a trained medical image recognition model; The visual conversion model is a model of a multi-perception layer mechanism. The pre-constructed visual conversion model is semi-supervisedly trained using a pre-constructed labeled sample set and an unlabeled sample set according to the mask autoencoder pre-training method to obtain a trained visual conversion model, including: using pre-constructed labeled samples to train the visual conversion model; masking each sample data in the pre-constructed unlabeled sample set according to a preset masking ratio to obtain a masked data set; using the visual conversion model to identify the masked data set to obtain a recognition result and accuracy probability distribution corresponding to each masked data. number; obtain the recognition result corresponding to the masked data with an accuracy probability score greater than or equal to a preset qualified threshold value as the unlabeled sample, and use the recognition result as the pseudo-label of the unlabeled sample to obtain the pseudo-labeled sample; calculate the average value of each of the accuracy probability scores, and determine the convergence of the average value; when the average value does not converge, import the pseudo-labeled sample into the labeled sample, and return to the above step of using the pre-constructed labeled samples to train the visual conversion model, and update the visual conversion model; when the average value converges, obtain the trained visual conversion model; The method of using pre-constructed labeled samples to train a visual conversion model includes: splitting the labeled samples into images to obtain a set of blocks of the labeled samples; using a preset mapping matrix to perform one-dimensional mapping on each block in the block set to obtain a set of block vectors; using a pre-constructed fully connected layer to connect and arrange the set of block vectors to obtain a two-dimensional vector; performing column feature extraction operations and row feature extraction operations on the two-dimensional vectors to obtain a position relationship feature vector and a channel relationship feature vector respectively; and using a Gaussian error linear loss function, the position relationship feature vector and the channel relationship feature vector to train a visual conversion model.
2. The training method of the medical image recognition model according to claim 1, characterized in that: The step of training the pre-trained medical image recognition model using the task sample set corresponding to the downstream medical task decoder to obtain a trained medical image recognition model includes: Using a pre-trained medical image recognition model to perform model prediction on a task sample set corresponding to the downstream medical task decoder to obtain a prediction result set; According to a preset loss function, the loss value of the predicted result set and the actual result of each task sample in the task sample set is calculated; Minimize the loss value to obtain the network parameters when the loss value is the smallest, and use the network parameters to perform network back propagation to update the pre-trained medical image recognition model; Determine whether the loss value is less than a preset threshold; When the loss value is greater than or equal to the preset threshold, return to the step of using the pre-trained medical image recognition model to perform model prediction on the task sample set corresponding to the downstream medical task decoder to obtain a prediction result set, and iteratively update the pre-trained medical image recognition model; When the loss value is less than the preset threshold, a trained medical image recognition model is obtained.
3. The training method of the medical image recognition model according to claim 1, characterized in that: The extracting encoder weight parameters in the visual conversion model and migrating the encoder weight parameters to a pre-built local encoder to obtain a local update encoder includes: Use the torch.save function package to obtain the storage address of the encoder weight parameters of the visual conversion model; According to the storage address, the encoder weight parameters in the storage address are migrated to the pre-built local encoder using the torch.load function package to obtain a local update encoder.
4. A training device for a medical image recognition model, used to implement the training method for a medical image recognition model as claimed in any one of claims 1 to 3, characterized in that: The device comprises: A semi-supervised learning module is used to perform semi-supervised training on a pre-constructed visual conversion model using a pre-constructed labeled sample set and an unlabeled sample set according to a mask autoencoder pre-training method to obtain a trained visual conversion model; A knowledge transfer module, used for extracting encoder weight parameters in the visual conversion model, and migrating the encoder weight parameters to a pre-built local encoder to obtain a local update encoder; A model splicing module, used to call a preset downstream medical task decoder, connect the local update encoder and the downstream medical task decoder, and obtain a pre-trained medical image recognition model; The fine-tuning training module is used to train the pre-trained medical image recognition model using the task sample set corresponding to the downstream medical task decoder to obtain a trained medical image recognition model.
5. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the training method of the medical image recognition model as described in any one of claims 1 to 3.
6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the training method of the medical image recognition model as described in any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Clothing image appearance attribute modification method based on deep learning
CN112861884A
Text recognition model training method and device and text recognition method and device
CN114399769A