OCTA image classification structure training method based on self-supervised learning
Through a self-supervised learning method, OCTA images are trained in classification structures, and vascular features are extracted using a three-dimensional random mask feature encoder, which solves the problem of ignoring three-dimensional structure and lacking labeling information in the prior art, and achieves a more accurate classification of retinal diseases.
Patent Information
- Application Number
- CN202210887658.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-26
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-07-26
AI Technical Summary
The existing OCTA image analysis methods ignore the stratified information in the three-dimensional structure of retinal blood vessels, and due to the lack of labeling information for disease classification, a large number of potential disease characteristics cannot be utilized, limiting the accuracy of disease analysis.
The OCTA image classification structure training method based on self-supervised learning is adopted. By self-supervised learning of the label-free B-scan OCTA image sequence, the three-dimensional image sequence is reconstructed, and the blood vessel feature information is extracted using the three-dimensional random mask feature encoder to generate a two-dimensional fused OCTA feature image. Then, the tagged en-face OCTA image is used to fine-tune the model weight parameters to achieve accurate classification of OCTA images.
By fully utilizing the label-free three-dimensional B-scan OCTA image sequence information, more accurate two-dimensional fusion OCTA feature images are extracted, which improves the classification accuracy of retinal diseases, especially when only en-face OCTA images are used, which significantly improves the accuracy of renal disease classification.
Smart Images

Figure CN115410032B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of image processing technology, and specifically relates to an OCTA image classification structure training method based on self-supervised learning. Background Art
[0002] Currently, Optical Coherence Tomography Angiography (OCTA) is a relatively new non-invasive imaging method in recent years, which can effectively show the subtle changes in the capillary network in the human retinal plexus. Industry insiders have found that a variety of medical clinical retinal diseases can be analyzed from OCTA images, so OCTA has been used to evaluate a series of retinal vascular diseases, including diabetic retinopathy, diabetic nephropathy, etc. However, the common OCTA image analysis method can only obtain a shallow en-face OCTA image and a deep en-face OCTA image, ignoring a lot of useful information in other layers of the three-dimensional structure of retinal blood vessels. In addition, there are a large number of OCTA image data that contain potential characteristics of the disease, but they cannot be used because there is no annotation information for disease classification. This has led to the limitations of conventional methods in disease analysis.
[0003] Therefore, a structural training method for OCTA image classification based on self-supervised learning is needed. Summary of the invention
[0004] 1. Technical issues to be resolved
[0005] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present application provides an OCTA image classification structure training method based on self-supervised learning.
[0006] (II) Technical solution
[0007] In order to achieve the above objectives, this application adopts the following technical solutions:
[0008] In a first aspect, an embodiment of the present invention further provides an OCTA image classification structure training method based on self-supervised learning, the method comprising:
[0009] A10, performing self-supervised learning on the pre-established model based on a pre-given B-scan OCTA image sequence with no label information, until the reconstruction error between the reconstructed B-scan OCTA image sequence and the given B-scan OCTA image sequence, and the reconstruction error between the reconstructed OCTA feature image and the fused OCTA feature image meet preset conditions;
[0010] Among them, for each user's B-scan OCTA image sequence, the pre-established model processes the input pre-processed B-scan OCTA image sequence to generate a two-dimensional fused OCTA feature image, and the two-dimensional fused OCTA feature image is used to reconstruct the three-dimensional image sequence to obtain a reconstructed B-scan OCTA image sequence; the two-dimensional fused OCTA feature image is scaled, sampled, and features are extracted before reconstruction to obtain a reconstructed OCTA feature image;
[0011] A20. Fine-tune the two-dimensional random mask feature combination processing unit in the model after self-supervised learning for the given en-face OCTA image with labeled information to obtain a two-dimensional random mask feature combination processing unit for classifying the OCTA image of any user, and use the two-dimensional random mask feature combination processing unit as the trained OCTA image classification structure.
[0012] Optionally, for each user's B-scan OCTA image sequence, A10 includes:
[0013] A11, preprocessing each image in a predetermined B-scan OCTA image sequence without label information to obtain a preprocessed B-scan OCTA image sequence;
[0014] A12, extracting vascular feature information from the preprocessed B-scan OCTA image sequence based on the three-dimensional random mask feature encoder in the established model, and generating a two-dimensional fused OCTA feature image;
[0015] Wherein, the three-dimensional random mask feature encoder comprises: a first random mask unit, a three-dimensional autoencoder;
[0016] The first random mask unit is used to divide the entire preprocessed B-scan OCTA image sequence into 64 stereo image blocks according to the size of 75 pixels in length, 75 pixels in width, and 160 pixels in height, and sample with a probability of 0.5 according to a sampling strategy that obeys uniform distribution to obtain a non-masked part and a masked part, and input the non-masked part into a three-dimensional autoencoder;
[0017] The three-dimensional autoencoder is used to extract vascular features from the non-masked part of the input and output a two-dimensional fused OCTA feature image.
[0018] Optionally, the three-dimensional autoencoder includes: a plurality of self-attention encoding blocks, each of which is an encoding block based on a self-attention mechanism.
[0019] The calculation formula of the self-attention mechanism is:
[0020]
[0021] In the above formula, Q, K, and V represent the non-masked part and the randomly initialized matrix W respectively. Q , W K , W V The result of multiplication, K T is the result of K transpose, d k represents the length of vector K, and Softmax represents the flexible maximum function.
[0022] Optionally, A10 also includes:
[0023] A13, the two-dimensional fused OCTA feature image is input into the three-dimensional decoder of the model for reconstruction, and the three-dimensional decoder outputs a reconstructed B-scan OCTA image sequence;
[0024] A14, obtaining a reconstruction error between the reconstructed B-scan OCTA image sequence and the original B-scan OCTA image sequence;
[0025]
[0026] In formula (1), L mse represents the reconstruction error, L represents the length of the original B-scan OCTA image sequence, H represents the height of the original B-scan OCTA image sequence, W represents the width of the original B-scan OCTA image sequence, and y lhw,true is the true value of each pixel in the original B-scanOCTA image sequence, y lhw,pre The pixel value of each point in the reconstructed B-scan OCTA image sequence output by the three-dimensional decoder;
[0027] The original B-scan OCTA image sequence is a preprocessed B-scan OCTA image sequence to which the reconstructed B-scan OCTA image sequence belongs.
[0028] Optionally, A10 also includes:
[0029] The two-dimensional random mask feature combination processing unit includes: a two-dimensional random mask feature encoding module, a fully connected layer, and a softmax layer connected in sequence;
[0030] A15, scaling the two-dimensional fused OCTA feature image, and inputting the scaled image into the two-dimensional random mask feature encoding module in the established model to extract the fused OCTA feature block;
[0031] The two-dimensional random mask feature encoding module includes: a second random mask unit and a two-dimensional autoencoder;
[0032] The second random mask unit is used to divide the scaled image into Q' plane image blocks according to the specified size, and sample with a probability of 0.3 to 0.5 according to a sampling strategy that obeys uniform distribution to obtain a new non-masked part and a new masked part, and input the new non-masked part into the two-dimensional autoencoder;
[0033] The two-dimensional autoencoder extracts features from the new non-masked part of the input and outputs a fused OCTA feature block;
[0034] A16. Input the fused OCTA feature block and the new mask part into the two-dimensional decoder, and the output of the two-dimensional decoder is used as the reconstructed OCTA feature image.
[0035] Optionally, obtaining a reconstruction error between the reconstructed OCTA feature image and the fused OCTA feature image;
[0036]
[0037] In the above formula, L tmse represents the size of the reconstruction error, H t Indicates the height of the fused OCTA feature image, W t Indicates the width of the fused OCTA feature image, y hw,true represents the pixel value in the fused OCTA feature image, y hw,pre Represents the pixel value in the reconstructed OCTA feature image;
[0038] When the reconstruction errors in formula (1) and formula (2) meet the preset conditions, the self-supervised learning ends.
[0039] Optionally, the two-dimensional random mask feature combination processing unit includes: a two-dimensional random mask feature encoding module, a fully connected layer, and a softmax layer connected in sequence;
[0040] A20 includes: directly inputting the given en-face OCTA image with label information into the two-dimensional random mask feature encoding module;
[0041] The two-dimensional random mask feature encoding module extracts the features of the en-face OCTA image, and inputs the extracted features into the fully connected layer and the softmax layer connected in sequence to obtain the classification results;
[0042] Among them, the cross entropy function is used as the loss function to calculate the classification error. When the classification error meets the preset end condition, the fine-tuning training ends;
[0043]
[0044] In formula (3), L ceIndicates the error size of the classification result, c indicates the number of categories, y i,true Represents the true category of the image obtained based on the label information, y i,pre Represents the classification result output by the softmax layer.
[0045] Optionally, A15 performs scaling processing on the two-dimensional fused OCTA feature image, including:
[0046] The two-dimensional fused OCTA feature image was scaled to 224 pixels in length and 224 pixels in width by random cropping.
[0047] In a second aspect, an embodiment of the present invention further provides an OCTA image classification method based on self-supervised learning, the method comprising:
[0048] For the OCTA image to be classified, the OCTA image to be classified is input into the trained two-dimensional random mask feature combination processing unit to obtain a classification result;
[0049] The two-dimensional random mask feature combination processing unit is obtained by training through any of the OCTA image classification structure training methods based on self-supervised learning described in the first aspect above.
[0050] In a third aspect, the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the OCTA image classification structure training method based on self-supervised learning as described in any one of the first aspect above; or, executes the steps of the OCTA image classification method based on self-supervised learning as described in any one of the second aspect above.
[0051] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the OCTA image classification structure training method based on self-supervised learning as described in any one of the first aspect above are implemented; or, the steps of the OCTA image classification method based on self-supervised learning as described in any one of the second aspect above are implemented.
[0052] (III) Beneficial effects
[0053] The technical solution provided by this application may have the following beneficial effects:
[0054] 1) Based on existing studies showing that human retinal lesions may be correlated with kidney lesions, the method of the embodiment of the present invention provides a basis for disease analysis based on human retinal en-face OCTA images, making the classification results more accurate and the classification accuracy higher.
[0055] 2) The method of the embodiment of the present invention takes into account the relationship between a three-dimensional image sequence consisting of multiple B-scan OCTA images obtained by taking multiple times of the same eye and a single en-face OCTA image, and uses self-supervised learning technology to process the B-scan OCTA image sequence. Thus, a three-dimensional random mask feature encoder is designed, which can fully utilize the unlabeled three-dimensional B-scan OCTA image sequence information and obtain a more accurate two-dimensional fused OCTA feature image.
[0056] 3) In the embodiment of the present invention, based on the fusion of OCTA feature images, self-supervised learning technology is used to extract the lesion features of the retinal area, and the weight parameters of the model are fine-tuned and classified using the labeled en-face OCTA images based on the extracted features. The above method can assist in clinical diagnosis. The accuracy of kidney disease classification can be greatly improved when only en-face OCTA images are used. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] FIG1( a ) is a schematic diagram of an example of a label-free B-scan OCTA image of the present invention;
[0058] FIG1( b ) is a schematic diagram of an example of an image sequence consisting of multiple label-free B-scan OCTA images;
[0059] FIG1( c ) is a schematic diagram of an example of a labeled en-face OCTA image of the present invention;
[0060] Figure 2 A schematic diagram of a three-dimensional random mask feature encoding module described in the present invention;
[0061] FIG3( a) is a schematic diagram of a self-attention encoding block described in the present invention;
[0062] FIG3( b ) is a schematic diagram of the calculation process of the self-attention mechanism described in the present invention;
[0063] FIG3( c ) is a schematic diagram of the multi-layer perceptron block structure in the self-attention encoding block described in the present invention;
[0064] Figure 4 This is a schematic diagram of the structure of the three-dimensional decoder described in the present invention;
[0065] FIG5( a ) is a schematic diagram of a two-dimensional random mask feature encoding module described in the present invention;
[0066] FIG5( b ) is a schematic diagram of the structure of a two-dimensional decoder described in the present invention;
[0067] Figure 6A schematic diagram of the OCTA image classification structure training process based on self-supervised learning described in the present invention;
[0068] Figure 7 It is a flowchart of the OCTA image classification structure training method based on self-supervised learning of the present invention. DETAILED DESCRIPTION
[0069] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below in conjunction with the accompanying drawings through specific implementation methods. It is understood that the specific embodiments described below are only used to explain the relevant inventions, rather than to limit the invention. It should also be noted that the embodiments and features in the embodiments of this application can be combined with each other in the absence of conflict; for ease of description, only the parts related to the invention are shown in the accompanying drawings.
[0070] The core solution of the embodiment of the present invention is: inputting the B-scan (i.e., structural cross-section) OCTA image sequence that does not contain the patient's kidney disease label information into the model, using a self-supervised method to perform pre-training and extract results containing en-face (i.e., planar image) OCTA image features; then inputting the en-face OCTA images with kidney disease labels into the model to fine-tune the weight parameters of the model and perform classification, thereby obtaining a trained model.
[0071] Embodiment 1
[0072] Figure 6 and Figure 7 They are respectively a flow chart of a method for training an OCTA image classification structure based on self-supervised learning in an embodiment of the present application. The method of this embodiment can be executed by any computing device, and the computing device can be implemented in the form of software and / or hardware, such as Figure 6 or Figure 7 As shown, the method comprises the following steps:
[0073] A10, performing self-supervised learning on the pre-established model based on a pre-given B-scan OCTA image sequence with no label information, until the reconstruction error between the reconstructed B-scan OCTA image sequence and the given B-scan OCTA image sequence, and the reconstruction error between the reconstructed OCTA feature image and the fused OCTA feature image meet preset conditions;
[0074] Among them, for each user's B-scan OCTA image sequence, the pre-established model processes the input pre-processed B-scan OCTA image sequence to generate a two-dimensional fused OCTA feature image, and the two-dimensional fused OCTA feature image is used to reconstruct the three-dimensional image sequence to obtain a reconstructed B-scan OCTA image sequence; the two-dimensional fused OCTA feature image is scaled, sampled, and features are extracted before reconstruction to obtain a reconstructed OCTA feature image;
[0075] A20. Fine-tune the two-dimensional random mask feature combination processing unit in the model after self-supervised learning for the given en-face OCTA image with labeled information to obtain a two-dimensional random mask feature combination processing unit for classifying the OCTA image of any user, and use the two-dimensional random mask feature combination processing unit as the trained OCTA image classification structure.
[0076] The two-dimensional random mask feature combination processing unit of this embodiment may include: a two-dimensional random mask feature encoding module, a fully connected layer, and a softmax layer connected in sequence.
[0077] The method of this embodiment takes into account the relationship between a three-dimensional image sequence consisting of multiple B-scan OCTA images obtained by taking multiple times of the same eye and a single en-face OCTA image, and uses self-supervised learning technology to process the B-scan OCTA image sequence. Thus, a three-dimensional random mask feature encoder is designed, which can fully utilize the unlabeled three-dimensional B-scan OCTA image sequence information and obtain a more accurate two-dimensional fused OCTA feature image.
[0078] For each user's B-scan OCTA image sequence, the above step A10 may include the following steps:
[0079] A11, preprocessing each image in a predetermined B-scan OCTA image sequence without label information to obtain a preprocessed B-scan OCTA image sequence;
[0080] A12, extracting vascular feature information from the preprocessed B-scan OCTA image sequence based on the three-dimensional random mask feature encoder in the established model, and generating a two-dimensional fused OCTA feature image;
[0081] A13, the two-dimensional fused OCTA feature image is input into the three-dimensional decoder of the model for reconstruction, and the three-dimensional decoder outputs a reconstructed B-scan OCTA image sequence;
[0082] A14, obtaining a reconstruction error between the reconstructed B-scan OCTA image sequence and the original B-scan OCTA image sequence;
[0083] A15, scaling the two-dimensional fused OCTA feature image, and inputting the scaled image into the two-dimensional random mask feature encoding module in the established model to extract the fused OCTA feature block;
[0084] A16. Input the fused OCTA feature block and the new mask part (two-dimensional image block) into the two-dimensional decoder, and the output of the two-dimensional decoder is used as the reconstructed OCTA feature image.
[0085] For each user's B-scan OCTA image sequence, the above step A20 may include the following steps:
[0086] A20 includes: directly inputting the given en-face OCTA image with label information into the two-dimensional random mask feature encoding module;
[0087] The two-dimensional random mask feature encoding module extracts the features of the en-face OCTA image, and inputs the extracted features into the fully connected layer and the softmax layer connected in sequence to obtain the classification results;
[0088] In this embodiment, based on the fusion of OCTA feature images, self-supervised learning technology is used to extract the lesion features of the retinal area, and the weight parameters of the model are fine-tuned and classified using labeled en-face OCTA images based on the extracted features. The above method can assist in clinical diagnosis. It can greatly improve the accuracy of kidney disease classification when only en-face OCTA images are used, making the classification results more accurate and the classification accuracy higher.
[0089] Embodiment 2
[0090] To better understand the method of the above embodiment 1, the following Figure 1(a) to Figure 6 The above method is described in detail with examples.
[0091] For the sake of convenience, the following steps are explained using the image or image sequence of one patient as an example, and the final training process is composed of image / image sequence data of multiple patients.
[0092] Step 01: Multiple B-scan OCTA images of the same eye of a patient are combined into an image sequence and input into the model as a three-dimensional image sample.
[0093] Each image in the B-scan OCTA image sequence does not contain information about the patient's kidney disease, that is, unlabeled data.
[0094] The three-dimensional B-scan OCTA image sequence is processed through a three-dimensional random mask feature encoding module to extract the three-dimensional information and obtain a two-dimensional image, which is called a fused OCTA feature image.
[0095] In order to improve the representation effect of the two-dimensional fused OCTA feature image on the three-dimensional B-scan OCTA image sequence, the fused OCTA feature image is input into the three-dimensional decoder branch of the model for reconstruction and restoration. The result obtained by reconstruction and restoration is consistent with the shape and size of the B-scan OCTA image sequence. This result is called the reconstructed B-scan OCTA image sequence. The difference between the reconstructed B-scan OCTA image sequence and the original B-scan OCTA image sequence is measured using the reconstruction error.
[0096] Step 02: Input the fused OCTA feature image obtained in step 01 into the two-dimensional random mask feature encoding module of the model for training. The result is called the fused OCTA feature block. Then input the fused OCTA feature block into the two-dimensional decoder for reconstruction and restoration. The output result is consistent with the shape and size of the fused OCTA feature image. The result is called the reconstructed OCTA feature image. The difference between the reconstructed OCTA feature image and the fused OCTA feature image is measured using the reconstruction error.
[0097] The training process of the above two steps is all self-supervised learning. After the training is completed, only the two-dimensional random mask feature encoding module needs to be retained as the classification model, which can be used for classification analysis of en-face OCTA images. The en-face OCTA image with labels is input into the two-dimensional random mask feature encoding module for weight parameter fine-tuning training. The feature map output by the two-dimensional random mask feature encoding module is then input into a classifier composed of multiple fully connected layers, and training is performed in a supervised manner to finally obtain the classification results of kidney disease. The classifier here may include: a fully connected layer and a softmax layer.
[0098] (I) B-scan OCTA image preprocessing
[0099] The input of the self-supervised learning stage of this method is a B-scan OCTA image sequence of the human retina that does not contain the user's kidney disease label information. A single B-scan OCTA image in the B-scan OCTA image sequence is obtained by shooting the sagittal plane of the user's human retina. Multiple B-scan OCTA images can be obtained by shooting multiple times along the Y-axis direction of the spatial coordinate. Multiple B-scan OCTA images can constitute a B-scan OCTA image sequence.
[0100] Each B-scan OCTA image contains various layered information of the retina from the superficial to the deep layers. Figure 1(a) shows a single B-scan OCTA image, and Figure 1(b) shows a sequence of B-scan OCTA images of any user, which are combined along the Y-axis direction of the spatial coordinate in the order of shooting time.
[0101] The preprocessing operations for each B-scan OCTA image in the original B-scan OCTA image sequence of each user include but are not limited to: histogram equalization, image binarization, and image filtering based on connected region screening.
[0102] The user's original B-scan OCTA image sequence is numbered starting from 001 in the order of shooting time, and the image format is converted. All images are converted into 8-bit grayscale images, and the pixel grayscale value is limited to between 0 and 255. The number of pixels at each grayscale level in each image is counted, and then the pixel grayscale in each image is histogram equalized.
[0103]
[0104] Among them, s represents the gray value after conversion, k represents the gray level, r represents the gray value before conversion, n represents the current gray level, and p r Represents the distribution function of r.
[0105] In addition, in order to further reduce the impact of isolated noise on B-scan OCTA images and increase the proportion of effective features in the image without reducing the image quality, a filtering algorithm based on connected regions is used to filter the isolated noise points. The filtering algorithm uses a sliding window to operate on the B-scan OCTA image pixel by pixel to filter out the isolated connected pixel areas within the sliding window range. The sliding window is defined as:
[0106]
[0107] Among them, i and j represent the coordinates of the center pixel, and D represents the connected area within the current sliding window.
[0108] (II) Extracting features of B-scan OCTA image sequences
[0109] The preprocessed B-scan OCTA image sequence is used as the input data of the self-supervised learning stage. A three-dimensional random mask feature encoder is used to extract the vascular feature information of the image. The three-dimensional random mask feature encoder includes: a first random mask unit and a three-dimensional self-encoder.
[0110] The first random mask unit is used to divide the preprocessed B-scan OCTA image sequence into 64 stereo image blocks with a length of 75 pixels, a width of 75 pixels, and a height of 160 pixels, and sample them with a probability of 0.5 according to a sampling strategy that obeys uniform distribution. After random masking, the non-masked part (i.e., the stereo image block retained after sampling) is obtained and input into the three-dimensional autoencoder, and the remaining masked part is discarded.
[0111] The three-dimensional autoencoder includes: multiple self-attention encoding blocks, that is, multiple self-attention encoding blocks are stacked. The three-dimensional autoencoder extracts vascular features from the non-masked part of the input and finally outputs a two-dimensional fused OCTA feature image. Figure 2 It shows the flow chart of the B-scan OCTA image sequence outputting the fused OCTA feature image through the three-dimensional random mask feature encoder. Figure 3(a) shows the schematic diagram of the structure of the self-attention encoding block. Figure 3(b) shows the schematic diagram of the calculation process of the self-attention mechanism. Figure 3(c) shows the schematic diagram of the multi-layer perceptron block structure in the self-attention encoding block.
[0112] The calculation formula of the self-attention mechanism is as follows:
[0113]
[0114] In the above formula, Q, K, and V represent the non-masked part and the randomly initialized matrix W respectively. Q , W K , W V The result of multiplication, K T is the result of K transpose, d k Represents the length of vector K, and Softmax represents the flexible maximum function. The above formula describes the calculation process of the self-attention mechanism, which will be used in the self-attention encoding block.
[0115] After the features are extracted by the three-dimensional random mask feature encoder, the vascular structure information in the B-scan OCTA image sequence will be included in the two-dimensional fused OCTA feature image.
[0116] In order to make the two-dimensional fused OCTA feature image more accurately represent the vascular structure information in the three-dimensional B-scan OCTA image sequence, the fused OCTA feature image is then input into the three-dimensional decoder branch of the model for reconstruction and restoration. The output result of the three-dimensional decoder is the reconstructed B-scan OCTA image sequence.
[0117] The 3D decoder is implemented through upsampling layers and 3D convolutional layers. During the upsampling process, the length and width of the fused OCTA feature image are kept unchanged. The gradual increase of upsampling channels can lead to a gradual increase in the height of the image until the final height is the same as the original B-scan OCTA image sequence. Figure 4 It indicates that the fused OCTA feature images are outputted by a 3D decoder to reconstruct the B-scan OCTA image sequence.
[0118] The reconstruction error between the reconstructed B-scan OCTA image sequence and the original B-scan OCTA image sequence (ie, the preprocessed B-scan OCTA image sequence) is calculated using MSE (mean square error).
[0119]
[0120] In the above formula, L mse represents the size of the reconstruction error, L represents the length of the B-scan OCTA image sequence, H represents the height of the B-scan OCTA image sequence, W represents the width of the B-scan OCTA image sequence, and y lhw,true is the true value of each pixel in the B-scan OCTA image sequence, y lhw,pre It is the pixel value of each point in the reconstructed B-scan OCTA image sequence output by the 3D decoder. Figure 4 It is a schematic diagram showing the reconstruction of B-scan OCTA image sequence by fusion of OCTA feature images through the output of a three-dimensional decoder.
[0121] (III) Extracting features of fused OCTA feature images
[0122] The fused OCTA feature image obtained in the previous step is scaled to a size of 224 pixels in length and 224 pixels in width by random cropping. The scaled image is then input into the two-dimensional random mask feature encoding module in the two-dimensional random mask feature combination processing unit in the model to extract potential disease features.
[0123] Among them, the two-dimensional random mask feature encoding module includes: a second random mask unit and a two-dimensional autoencoder.
[0124] The second random mask unit is used to divide the fused OCTA feature image into 196 planar image blocks of size 16 pixels long and 16 pixels wide, and sample them with a probability of 0.5 using a sampling strategy that obeys a uniform distribution. After random masking, the new non-masked part (the new non-masked part is a planar image block) is retained and input into the two-dimensional autoencoder, and the remaining masked part is reserved for subsequent decoding.
[0125] The above-mentioned two-dimensional autoencoder is composed of stacking multiple self-attention encoding blocks, extracting features from the new non-masked part of the input, i.e., the plane image block, and finally outputting a fused OCTA feature block with the same shape, size, and number as the non-masked part. Figure 5(a) shows a schematic diagram of the fused OCTA feature block output by the two-dimensional autoencoder of the fused OCTA feature image. The self-attention encoding block used by the two-dimensional autoencoder is consistent with the form of the self-attention encoding block mentioned above, and the specific weight parameters may be different.
[0126] Then, the fused OCTA feature block and the mask part output by the second random mask unit are sequentially input into the two-dimensional decoder, and the output of the two-dimensional decoder is called the reconstructed OCTA feature image. FIG5( b ) is a schematic diagram showing the reconstructed OCTA feature image by fusion of the OCTA feature block and the mask part through the output of the two-dimensional decoder.
[0127] The above-mentioned sequential input can be understood as follows: the mask part includes multiple small image blocks, the first small image block in the upper left corner is numbered 1, and the subsequent small image blocks are numbered 2, 3, 4, and so on until the last small image block in the lower right corner. The small image blocks are input into the two-dimensional decoder in sequence according to this number, that is, the sequential input is realized.
[0128] The reconstruction error between the reconstructed OCTA feature image and the fused OCTA feature image was calculated using the mean square error (MSE).
[0129]
[0130] In the above formula, L tmse represents the size of the reconstruction error, H t Indicates the height of the fused OCTA feature image, W t Indicates the width of the fused OCTA feature image, y hw,true represents the pixel value in the fused OCTA feature image, y hw,pre Represents the pixel value in the reconstructed OCTA feature image.
[0131] (IV) Extracting features and classifying labeled en-face OCTA images
[0132] Repeat the above (ii) and (iii) until the two reconstruction errors in the two steps no longer decrease and the self-supervised learning process is terminated. All image data used in the above steps are unlabeled image data without patient kidney disease information, with the purpose of accurately learning the intrinsic characteristics of vascular planar structure information, so it is a self-supervised learning stage. After the self-supervised learning stage, the two-dimensional random mask feature encoding module of the two-dimensional random mask feature combination processing unit in (ii) above can already fully extract vascular features, and only the two-dimensional random mask feature combination processing unit needs to be retained.
[0133] Next, the en-face OCTA images with the patient's kidney disease classification label information are used as input data for the fine-tuning stage. The en-face OCTA images with label information are directly input into the two-dimensional random mask feature combination processing unit. The two-dimensional random mask feature encoding module in the two-dimensional random mask feature combination processing unit extracts features and then continues to input them into the fully connected layer, and finally connects to the softmax layer to obtain the classification result of kidney disease.
[0134] The above training process is actually a fine-tuning process of the weight parameters. The cross entropy function is used as the loss function to calculate the classification error. The fine-tuning process ends when the classification error no longer decreases.
[0135]
[0136] In the above formula, L ce represents the error of the classification result, c represents the two categories of diabetic nephropathy and non-diabetic nephropathy, y i,true Represents the true category of the image obtained based on the label information, y i,pre Represents the classification result output by the softmax layer.
[0137] Embodiment 3
[0138] An embodiment of the present invention provides an OCTA image classification method based on self-supervised learning, the method comprising:
[0139] For the OCTA image to be classified, the OCTA image to be classified is input into the trained two-dimensional random mask feature combination processing unit to obtain a classification result;
[0140] The two-dimensional random mask feature combination processing unit is obtained by training the OCTA image classification structure training method based on self-supervised learning described in any of the above embodiments.
[0141] That is to say, after the fine-tuning process of part (iv) in the above embodiment 2 is completed, an en-face OCTA image of a patient to be diagnosed is input. The image is sequentially passed through the two-dimensional random mask feature encoding module and the fully connected layer to extract the features of the relevant diseases, and then the softmax layer outputs the final disease classification result and the corresponding disease probability to provide analysis basis for doctors.
[0142] Embodiment 4
[0143] This embodiment provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the OCTA image classification structure training method based on self-supervised learning as described in any one of the above embodiments are implemented, or the steps of the OCTA image classification method based on self-supervised learning are executed.
[0144] The method disclosed in the above embodiment of the present invention can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a ready-made programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The methods, steps, and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present invention can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software units in a decoding processor. The software unit may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0145] It should be noted that in the claims, any figure marks between brackets should not be understood as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "one" or "an" preceding a component does not exclude the presence of multiple such components. In addition, it should be noted that in the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" and the like refers to the specific features, structures, materials or characteristics described in conjunction with the embodiment or example included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0146] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments after knowing the basic creative concept. Therefore, the claims should be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.
[0147] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention should also include these modifications and variations.
Claims
1. A method for training OCTA image classification structure based on self-supervised learning, characterized in that: The method includes: A10, performing self-supervised learning on the pre-established model based on a pre-given B-scan OCTA image sequence with no label information, until the reconstruction error between the reconstructed B-scan OCTA image sequence and the given B-scan OCTA image sequence, and the reconstruction error between the reconstructed OCTA feature image and the fused OCTA feature image meet preset conditions; Among them, for each user's B-scan OCTA image sequence, the pre-established model processes the input pre-processed B-scan OCTA image sequence to generate a two-dimensional fused OCTA feature image, and the two-dimensional fused OCTA feature image is used to reconstruct the three-dimensional image sequence to obtain a reconstructed B-scan OCTA image sequence; the two-dimensional fused OCTA feature image is scaled, sampled, and features are extracted before reconstruction to obtain a reconstructed OCTA feature image; A10 includes: A11, preprocessing each image in a predetermined B-scan OCTA image sequence without label information to obtain a preprocessed B-scan OCTA image sequence; A12, extracting vascular feature information from the preprocessed B-scan OCTA image sequence based on the three-dimensional random mask feature encoder in the established model, and generating a two-dimensional fused OCTA feature image; A13, the two-dimensional fused OCTA feature image is input into the three-dimensional decoder of the model for reconstruction, and the three-dimensional decoder outputs a reconstructed B-scan OCTA image sequence; A14, obtaining a reconstruction error between the reconstructed B-scan OCTA image sequence and the original B-scan OCTA image sequence; Formula (1): In formula (1), L mse represents the reconstruction error, L represents the length of the original B-scan OCTA image sequence, H represents the height of the original B-scan OCTA image sequence, W represents the width of the original B-scan OCTA image sequence, and y lhw,true is the true value of each pixel in the original B-scan OCTA image sequence, y lhw,pre The pixel value of each point in the reconstructed B-scan OCTA image sequence output by the three-dimensional decoder; The original B-scan OCTA image sequence is a pre-processed B-scan OCTA image sequence to which the reconstructed B-scan OCTA image sequence belongs; A20. Fine-tune the two-dimensional random mask feature combination processing unit in the model after self-supervised learning for the given en-face OCTA image with labeled information to obtain a two-dimensional random mask feature combination processing unit for classifying the OCTA image of any user, and use the two-dimensional random mask feature combination processing unit as the trained OCTA image classification structure.
2. The training method according to claim 1, characterized in that: B-scan OCTA image sequence for each user, The three-dimensional random mask feature encoder comprises: a first random mask unit, a three-dimensional autoencoder; The first random mask unit is used to divide the entire preprocessed B-scan OCTA image sequence into 64 stereo image blocks according to the size of 75 pixels in length, 75 pixels in width, and 160 pixels in height, and sample with a probability of 0.5 according to a sampling strategy that obeys uniform distribution to obtain a non-masked part and a masked part, and input the non-masked part into a three-dimensional autoencoder; The three-dimensional autoencoder is used to extract vascular features from the non-masked part of the input and output a two-dimensional fused OCTA feature image.
3. The training method according to claim 2, characterized in that: The three-dimensional autoencoder includes: a plurality of self-attention encoding blocks, each of which is an encoding block based on the self-attention mechanism. The calculation formula of the self-attention mechanism is: ; In the above formula, Q, K, and V represent the non-masked part and the randomly initialized matrix W respectively. Q , W K , W V The result of multiplication, K T is the result of K transposition, d k represents the length of vector K, and Softmax represents the flexible maximum function.
4. The training method according to claim 1, characterized in that: The A10 also includes: The two-dimensional random mask feature combination processing unit includes: a two-dimensional random mask feature encoding module, a fully connected layer, and a softmax layer connected in sequence; A15, scaling the two-dimensional fused OCTA feature image, and inputting the scaled image into the two-dimensional random mask feature encoding module in the established model to extract the fused OCTA feature block; The two-dimensional random mask feature encoding module includes: a second random mask unit and a two-dimensional autoencoder; The second random mask unit is used to divide the scaled image into Q' plane image blocks according to the specified size, and sample with a probability of 0.3-0.5 according to a sampling strategy that obeys uniform distribution to obtain a new non-masked part and a new masked part, and input the new non-masked part into the two-dimensional autoencoder; The two-dimensional autoencoder extracts features from the new non-masked part of the input and outputs a fused OCTA feature block; A16. Input the fused OCTA feature block and the new mask part into the two-dimensional decoder, and the output of the two-dimensional decoder is used as the reconstructed OCTA feature image.
5. The training method according to claim 4, characterized in that: Obtaining the reconstruction error between the reconstructed OCTA feature image and the fused OCTA feature image; Formula (2) In the above formula, L tmse represents the size of the reconstruction error, H t Indicates the height of the fused OCTA feature image, W t Indicates the width of the fused OCTA feature image, y hw,true represents the pixel value in the fused OCTA feature image, y hw,pre Represents the pixel value in the reconstructed OCTA feature image; When the reconstruction errors in formula (1) and formula (2) meet the preset conditions, the self-supervised learning ends.
6. The training method according to any one of claims 1 to 5, characterized in that: The two-dimensional random mask feature combination processing unit includes: a two-dimensional random mask feature encoding module, a fully connected layer, and a softmax layer connected in sequence; A20 includes: directly inputting the given en-face OCTA image with label information into the two-dimensional random mask feature encoding module; The two-dimensional random mask feature encoding module extracts the features of the en-face OCTA image, and inputs the extracted features into the fully connected layer and the softmax layer connected in sequence to obtain the classification results; Among them, the cross entropy function is used as the loss function to calculate the classification error. When the classification error meets the preset conditions, the fine-tuning training ends; Formula (3) In formula (3), L ce Indicates the error size of the classification result, C indicates the number of categories, y i,true Represents the true category of the image obtained based on the label information, y i,pre Represents the classification result output by the softmax layer.
7. The training method according to claim 4, characterized in that: A15 scales the two-dimensional fused OCTA feature image, including: The two-dimensional fused OCTA feature image was scaled to 224 pixels in length and 224 pixels in width by random cropping.
8. An OCTA image classification method based on self-supervised learning, characterized in that: The method includes: For the OCTA image to be classified, the OCTA image to be classified is input into the trained two-dimensional random mask feature combination processing unit to obtain a classification result; The two-dimensional random mask feature combination processing unit is obtained by training through the OCTA image classification structure training method based on self-supervised learning described in any one of claims 1 to 6 above.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the OCTA image classification structure training method based on self-supervised learning as described in any one of claims 1 to 7 above are implemented, and the OCTA image classification method based on self-supervised learning as described in claim 8 above are executed.
Citation Information
Patent Citations
Earth observation image semantic segmentation method based on self-supervised learning
CN112308860A
Active saliency target detection method based on semi-supervised learning
CN112598053A