A classification method, device, electronic device and medium for a target image
By constructing and training an image classification model including feature extraction, fusion and classification components, and mapping the three-dimensional corneal images into two-dimensional depth maps for classification, the problem of low classification accuracy of medical images is solved and the accurate classification of corneal images is achieved.
Patent Information
- Application Number
- CN202210609664.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-05-31
AI Technical Summary
In the prior art, the classification accuracy of image classification models is not high when classifying medical images.
By obtaining a three-dimensional sample corneal image, mapping it into a two-dimensional sample depth map of several channels, and forming a sample data set. The initial image classification model is constructed, including a first component for feature extraction, a second component for feature fusion, and a third component for feature classification. The initial image classification model is trained using the sample data set to obtain the target image classification model for corneal image classification. Map the 3D target corneal image into a 2D target depth map and enter the target image classification model to output the target classification results.
The accurate classification of three-dimensional target corneal images is achieved, the accuracy of the image classification model in medical image classification is improved, and the problem of low classification accuracy in the prior art is solved.
Smart Images

Figure CN114842270B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and in particular, to a method, device, electronic device and medium for classifying target images. Background Art
[0002] With the continuous enhancement of computer computing power and the continuous development of artificial intelligence algorithms, image classification technology based on deep learning has begun to be applied in various fields. An image classification model trained based on a large-scale dataset can classify images. Among them, the classification accuracy of the image classification model often depends on the quality and difficulty of the dataset. How to improve the accuracy of the image classification model is a problem that needs to be studied. Currently, the image classification model can classify natural images relatively accurately, such as classifying natural images into animal images, landscape images or building images, etc. However, due to the different imaging methods of medical images and natural images, and the fact that medical images contain more classification features, the classification accuracy of the image classification model when classifying medical images is not high. Summary of the Invention
[0003] The present invention provides a method, device, electronic device and medium for classifying target images to solve the problem of low classification accuracy of the image classification model in the prior art.
[0004] The method for classifying target images provided by the present invention includes:
[0005] Obtain a three-dimensional sample corneal image, map it into two-dimensional sample depth maps of a plurality of channels, and form a sample dataset;
[0006] Construct an initial image classification model, and use the sample dataset to train the initial image classification model to obtain a target image classification model for corneal image classification. The initial image classification model includes a first component for feature extraction, a second component for feature fusion, and a third component for feature classification;
[0007] Obtain a three-dimensional target corneal image, map it into two-dimensional target depth maps of a plurality of channels, and input the two-dimensional target depth maps into the target image classification model to output a target classification result.
[0008] Optionally, the step of using the sample dataset to train the initial image classification model to obtain a target image classification model for corneal image classification includes:
[0009] Perform channel splicing on the two-dimensional sample depth maps, and use the first component to extract features from the two-dimensional sample depth maps after channel splicing to obtain sample features;
[0010] Based on the second component, use the self-attention mechanism to perform cross-channel feature fusion on the sample features to obtain fused features;
[0011] Use the third component to classify the fused features to obtain a classification result;
[0012] Use the cross-entropy loss function to obtain the classification error between the classification result and the preset result, and use the backpropagation of the classification error to update the initial image classification model to obtain the target image classification model.
[0013] Optionally, the step of based on the second component, using the self-attention mechanism to perform cross-channel feature fusion on the sample features to obtain fused features includes:
[0014] Obtain the position information of the sample features to obtain position features;
[0015] Perform cross-channel feature fusion on the sample features according to the self-attention mechanism and the position features to obtain fused features.
[0016] Optionally, the sample features are composed of multi-dimensional matrices, and the step of using the self-attention mechanism to perform cross-channel feature fusion on the sample features to obtain fused features includes:
[0017] Split the sample features into several matrices according to a preset splitting rule to obtain sub-sample features;
[0018] Respectively obtain the correlation features between two sub-sample features;
[0019] Determine the target sub-sample features according to the attention mechanism and the correlation features;
[0020] Perform cross-channel feature fusion on the target sub-sample features to obtain fused features.
[0021] Optionally, the mathematical expression of the target sub-sample feature z i is:
[0022]
[0023] where z i is the i-th target sub-sample feature, i is the label of the target sub-sample feature, x is the sample feature, x = (x1, x2,..., x n ), x1 is the first sub-sample feature, x2 is the second sub-sample feature, x n is the n-th sub-sample feature, n is the total number of sub-sample features, j is the label of the sub-sample feature, x j is the j-th sub-sample feature, x i is the i-th sub-sample feature, α ij is the sub-sample feature xi The associated feature with the sub-sample feature x j , V is the input matrix of the self-attention mechanism, and W v is the weight matrix corresponding to V;
[0024] The sub-sample feature x i and the sub-sample feature x j The mathematical expression of the associated feature is:
[0025]
[0026] where, e ij is the associated data between the sub-sample feature x i and the sub-sample feature x k , and e ik is the associated data between the sub-sample feature x i and the sub-sample feature x k ;
[0027] The mathematical expression of e ij is:
[0028] e ij =(x i W Q )(x j W K ) T ;
[0029] where, Q and K are both input matrices of the self-attention mechanism, W Q is the weight matrix corresponding to Q, W K is the weight matrix corresponding to K, and T is the transpose of the matrix.
[0030] Optionally, the cross-channel feature fusion of the sample features by using the self-attention mechanism to obtain the fusion features includes:
[0031] Updating the target sub-sample feature, the associated feature, and the associated data according to the position feature;
[0032] The mathematical expression of the updated target sub-sample feature is:
[0033]
[0034] where, z i ’ is the updated i-th target sub-sample feature, a′ ij is the associated feature between the updated sub-sample feature x i and the sub-sample feature, is the position feature weight of the sub-sample feature x i corresponding to V and the sub-sample feature x j , e’ij The updated sub-sample feature x i and the sub-sample feature x k associated data;
[0035] The updated sub-sample feature x i and the sub-sample feature x j The mathematical expression of the associated feature is:
[0036]
[0037] where, e’ ij is the associated data between the updated sub-sample feature x i and the sub-sample feature x j e’ ik is the associated data between the updated sub-sample feature x i and the sub-sample feature x k associated data;
[0038] e’ ij The mathematical expression is:
[0039]
[0040] where, is the position feature weight of the sub-sample feature x corresponding to K i and the sub-sample feature x j associated data.
[0041] Optionally, the first component includes a convolutional neural network model, the second component includes a transformer model, and the third component includes a softmax classification model.
[0042] The present invention also provides a classification device for a target image, including:
[0043] A data acquisition module, configured to acquire a three-dimensional sample corneal image, map it into a two-dimensional sample depth map of a plurality of channels, and form a sample data set;
[0044] A model training module, configured to construct an initial image classification model, train the initial image classification model using the sample data set, and obtain a target image classification model for corneal image classification. The initial image classification model includes a first component for feature extraction, a second component for feature fusion, and a third component for feature classification;
[0045] An image classification module, configured to acquire a three-dimensional target corneal image, map it into a two-dimensional target depth map of a plurality of channels, and input the two-dimensional target depth map into the target image classification model to output a target classification result. The data acquisition module, the model training module, and the image classification module are connected.
[0046] The present invention also provides an electronic device, including: a processor and a memory;
[0047] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the classification method of the target image.
[0048] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the classification method of the target image as described above is implemented.
[0049] Advantages of the present invention: In the classification method of the target image in the present invention, first, a three-dimensional sample corneal image is obtained, mapped into two-dimensional sample depth maps of a plurality of channels, and a sample data set is formed; then an initial image classification model including a first component, a second component, and a third component is constructed, and the initial image classification model is trained using the sample data set to obtain a target image classification model for corneal image classification; the obtained three-dimensional target corneal image is mapped into two-dimensional target depth maps of a plurality of channels, and the two-dimensional target depth maps are input into the target image classification model to output a target classification result, thereby realizing accurate classification of the three-dimensional target corneal image and solving the problem of low classification accuracy of the image classification model in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0051] Figure 1 is a schematic flowchart of the classification method of the target image in the embodiment of the present invention;
[0052] Figure 2 is a schematic flowchart of the method for obtaining the target image classification model in the embodiment of the present invention;
[0053] Figure 3 is a schematic block diagram of the classification device of the target image in the embodiment of the present invention;
[0054] Figure 4 is a schematic structural diagram of the electronic device in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The following specific examples illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0056] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0057] Keratoconus is a congenital eye disease characterized by corneal ectasia, forward protrusion and thinning of the central cornea, presenting a conical shape. Its incidence is about one in two thousand, and it mostly occurs in teenagers. The early features of keratoconus are not clear and it is difficult to diagnose. When potential keratoconus patients undergo myopia surgery, it will induce the early onset of keratoconus and even exacerbate the condition. Therefore, before myopia surgery, it is necessary to perform corneal topographic modeling on the subject through a Pentacam three-dimensional anterior segment analyzer, and then make a judgment based on the data indicators statistically analyzed in the device. However, the large sample database referred to and compared by the Pentacam diagnostic system when calculating data indicators mostly comes from European ethnic groups, and it is not targeted for Asians with a smaller corneal radius. Therefore, in clinical use in our country, there are often many false positive cases diagnosed, far higher than the incidence of keratoconus. With the continuous enhancement of computer computing power and the continuous development of artificial intelligence algorithms, image classification technology based on deep learning has begun to be applied in various fields. Moreover, the image classification technology itself has also made great progress. If the accuracy of early diagnosis of keratoconus can be improved through the cutting-edge technology of artificial intelligence, the error rate and false positive rate can be reduced, providing a real and effective guiding role for early screening and treatment, then this will be a very meaningful work. However, currently, the classification accuracy of three-dimensional corneal images is relatively low. To solve the above problems, this application provides a classification method, device, electronic device, and medium for target images.
[0058] To illustrate the technical solutions described in the present invention, the following specific examples are used for illustration.
[0059] Figure 1 It is a schematic flowchart of the classification method for target images provided by the present invention in an embodiment.
[0060] As Figure 1As shown, the classification method of the above target image includes steps S110 - S130:
[0061] S110, obtain a three - dimensional sample corneal image, map it into two - dimensional sample depth maps of several channels, and form a sample data set.
[0062] First of all, it should be noted that the three - dimensional sample corneal image can be collected by a photographing device, and then the collected three - dimensional sample corneal image is mapped into two - dimensional sample depth maps of seven types (seven channels), namely, the two - dimensional sample depth map representing the height of the anterior corneal surface, the two - dimensional sample depth map representing the height of the posterior corneal surface, the two - dimensional sample depth map representing the curvature of the anterior corneal surface, the two - dimensional sample depth map representing the curvature of the posterior corneal surface, the two - dimensional sample depth map representing the refractive power of the whole cornea, the two - dimensional sample depth map representing the corneal thickness data, and the two - dimensional sample depth map representing the anterior chamber depth. Specifically, it can be taken by the Scheimpflug camera built in Pentacam in a uniformly rotating state, and then multiple corneal - related data can be directly exported in the system. After further screening out the files related to keratoconus, the circular corneal data in the files are filled and transformed into a rectangle, and then these seven - type (seven - channel) two - dimensional sample depth maps can be obtained.
[0063] It should be noted that in the selection of samples, since the symptoms of keratoconus often occur in one eye first and then affect both eyes as a whole, when collecting three - dimensional sample corneal images of keratoconus patients, usually only one eye is selected as a sample. To prevent the problem of class imbalance, when collecting samples, try to ensure the balance of positive and negative samples, that is, the number ratio of positive cases, negative cases, and false - positive cases is 1:1:1.
[0064] It should be understood that in the step of forming the sample data set, the sample data set is formed according to the two - dimensional sample depth maps of several channels. Before forming the sample data set, it is also necessary to label the three - dimensional sample corneal image, labeling normal corneas and keratoconus corneas.
[0065] S120, construct an initial image classification model, train the initial image classification model using the sample data set, and obtain a target image classification model for corneal image classification.
[0066] It should be noted that the initial image classification model includes a first component for feature extraction, a second component for feature fusion, and a third component for feature classification.
[0067] It can be understood that when training an initial image classification model using a sample data set to obtain a target image classification model, the sample data set can be divided into a training data set, a validation data set, and a test data set according to a certain ratio. For example, the sample data set is divided into a training data set, a validation data set, and a test data set according to a ratio of 6:2:2. The training set and the validation set are divided in a 4-fold cross-validation manner to obtain the average error of different models and determine appropriate hyperparameters. According to a certain ratio, all data is divided into three parts: a training set, a validation set, and a test set. The training set and the validation set are divided in an N-fold cross-validation manner to obtain the average error of different models and determine appropriate hyperparameters. After the hyperparameters are determined, the training set and the validation set are combined to train the final model, and finally the generalization ability of the model is tested using the test set.
[0068] It should be noted that in the validation process, cases where doctors judge as negative based on prior knowledge but pentacam judges as positive are collected and input into the trained target image classification model respectively to detect whether the individual has keratoconus. By training the initial image classification model with this type of sample image, the discrimination ability of the target classification model for such false positive cases is improved.
[0069] It should be understood that for the implementation method of training an initial image classification model using a sample data set to obtain a target image classification model for corneal image classification, please refer to Figure 2 , Figure 2 which is a schematic flowchart of the method for obtaining a target image classification model in an embodiment of the present invention.
[0070] As Figure 2 shown, the method for obtaining a target image classification model may include the following steps S210 - S240:
[0071] S210, perform channel splicing on the two-dimensional sample depth map, and use a first component to extract features from the two-dimensional sample depth map after channel splicing to obtain sample features.
[0072] It should be noted that the two-dimensional sample depth map is subjected to channel splicing to obtain a sample matrix; the first component is used to extract features from the sample matrix to obtain sample features composed of multi-dimensional matrices. Specifically, the first component can be a convolutional neural network model. When using the convolutional neural network model to extract features from the two-dimensional sample depth map after channel splicing, multiple convolutions and skip connections can be used to extract features therefrom. The skip connection ensures the backpropagation of gradients, solves the problem of gradient disappearance when the network is deeper, and speeds up the training process. The convolution operation is responsible for obtaining local region features respectively, such as local height, curvature, and thickness. Specifically, when using the convolutional neural network model to extract features from the sample matrix, three convolutions and skip connections can be used to extract features therefrom, and the features are reduced in dimension and then increased in dimension in the form of a bottleneck layer, which improves the specificity of the features better while reducing the computational amount; the convolution kernel sizes of the three convolutions can be 1x1, 3x3, and 1x1 respectively. S220, based on the second component, the self-attention mechanism is used to perform cross-channel feature fusion on the sample features to obtain fused features.
[0073] It should be noted that the implementation method of using the self-attention mechanism to perform cross-channel feature fusion on the sample features to obtain fused features may include splitting the sample features into several matrices according to a preset splitting rule to obtain sub-sample features; respectively obtaining the correlation features between two sub-sample features; determining the target sub-sample features according to the attention mechanism and the correlation features; performing cross-channel feature fusion on the target sub-sample features to obtain fused features.
[0074] It should be noted that the implementation method of using the self-attention mechanism to perform cross-channel feature fusion on the sample features to obtain fused features may also include obtaining the position information of the sample features to obtain position features; performing cross-channel feature fusion on the sample features according to the self-attention mechanism and the position features to obtain fused features. When obtaining the position information of the sample features, not only the absolute position information of each sub-sample feature in the sample features can be obtained, but also the relative position information of each sub-sample feature in the sample features can be obtained.
[0075] The target sub-sample feature z i is mathematically expressed as:
[0076]
[0077] where z i is the i-th target sub-sample feature, i is the label of the target sub-sample feature, x is the sample feature, x = (x1, x2,..., x n ), x1 is the first sub-sample feature, x2 is the second sub-sample feature, x nis the nth subsample feature, where n is the total number of subsample features, j is the label of the subsample feature, and x j is the jth subsample feature, x i is the ith subsample feature, α ij is the subsample feature x i and the associated feature of the subsample feature x j . V is the input matrix of the self-attention mechanism, and W v is the weight matrix corresponding to V;
[0078] The mathematical expression of the associated feature between the subsample feature x i and the subsample feature x j is:
[0079]
[0080] The associated feature can be calculated by the softmax function. The main role of the softmax function here is to convert the weights from any real number to positive numbers and normalize them. e ij can measure the mutual correlation between the subsample feature x i and the subsample feature x j .
[0081] Among them, e ij is the associated data between the subsample feature x i and the subsample feature x k . e ik is the associated data between the subsample feature x i and the subsample feature x k ;
[0082] The mathematical expression of e ij is:
[0083] e ij =(x i W Q )(x j W K ) T ;
[0084] Among them, Q and K are both input matrices of the self-attention mechanism, W Q is the weight matrix corresponding to Q, W K is the weight matrix corresponding to K, and T is the transpose of the matrix.
[0085] Using the self-attention mechanism to perform cross-channel fusion on the sample features, the obtained fusion features include: updating the target subsample features, associated features, and associated data according to the position features;
[0086] The mathematical expression of the updated target subsample feature is:
[0087]
[0088] Among them, z i ’ is the updated i-th target sub-sample feature, a′ ij is the updated sub-sample feature x i and the associated feature of the sub-sample feature, is the sub-sample feature x corresponding to V i and the sub-sample feature x j 's position feature weight, e’ ij is the updated sub-sample feature x i and the sub-sample feature x k 's associated data;
[0089] The updated sub-sample feature x i and the sub-sample feature x j 's mathematical expression of the associated feature is:
[0090]
[0091] Among them, e’ ij is the updated sub-sample feature x i and the sub-sample feature x j 's associated data, e’ ik is the updated sub-sample feature x i and the sub-sample feature x k 's associated data;
[0092] e’ ij 's mathematical expression is:
[0093]
[0094] Among them, is the sub-sample feature x corresponding to K i and the sub-sample feature x j 's position feature weight.
[0095] It should be understood that the second component includes a Transformer model. The sample features extracted by the first component are input into the Transformer model for feature encoding and decoding, as well as feature screening and fusion based on the self-attention mechanism of multi-head self-attention (MHSA). After passing through multiple Transformer modules, mean pooling is performed on all feature maps, which are then connected to a fully connected network. Different from recurrent neural networks, as a non-recursive Transformer structure cannot implicitly consider the order of elements in a sequence, position information may be lost in many tasks, and it is necessary to explicitly provide encoded position information. Therefore, the present invention explicitly performs absolute position encoding and relative position encoding for the model, enabling the self-attention operation to not only focus on content information but also be able to focus on the absolute or relative distances between features at different positions, thereby effectively associating cross-object information with position perception. The deep learning model with the Transformer structure is applied to the classification of keratoconus to reduce the classification error rate and solve the problem of too high false positive rate in the actual application of the Pentacam system.
[0096] It can be understood that the multi-head self-attention operation is a self-attention mechanism that can perform matrix calculations in a parallel manner, which can better capture the internal correlation of data and learn long-distance dependence relationships. When a feature map is input into the multi-head self-attention network, the self-attention mechanism will output a correlation matrix, representing the correlation between any two channels including itself; then, the correlation matrix (V, Q, K) is applied to the input feature map (the sample features in this application), and cross-channel feature fusion between any channel and all channels can be achieved. By using the self-attention mechanism, long-distance dependence relationships are better fused to make full use of multi-dimensional features and improve the classification accuracy. Using the multi-head self-attention network helps the model achieve parallel calculations of multiple sets of data and speeds up the calculation efficiency compared with traditional networks. By using the Contraction, Expansion, and self-attention mechanisms inside the Transformer model, long-distance cross-channel feature fusion is achieved. Absolute position encoding and relative position encoding information are added during feature fusion, thereby improving the classification accuracy of the target classification model.
[0097] S230, classify the fused features using a third component to obtain a classification result.
[0098] It should be noted that the third component includes a softmax classification model. The fused features obtained after cross-channel feature fusion processing by the second component are input into the softmax classification model. The softmax layer will output two probability values, and the class corresponding to the maximum output probability is the sample class predicted by the model. Specifically, the sample features can be passed through multiple Transformer modules, then average pooling is performed on all feature maps, and it is connected to a fully connected network to obtain the fused features and finally input into the softmax classification model to output the target classification result.
[0099] The mathematical expression of the cross-entropy loss function L is:
[0100]
[0101] where y i represents the label of sample i, the positive class is 1, and the negative class is 0; p i represents the probability that sample i is predicted as the positive class.
[0102] S240. The cross-entropy loss function is used to obtain the classification error between the classification result and the preset result, and the initial image classification model is updated by backpropagation of the classification error to obtain the target image classification model.
[0103] It should be understood that the preset result is the labeled class of the above three-dimensional sample corneal image. The model parameters are iteratively updated through the cross-entropy loss function and the stochastic gradient descent algorithm. Finally, the model parameters are fixed and reach a convergent state.
[0104] S130. Obtain a three-dimensional target corneal image, map it into two-dimensional target depth maps of several channels, and input the two-dimensional target depth maps into the target image classification model to output the target classification result.
[0105] It should be understood that the three-dimensional target corneal image is the three-dimensional corneal image to be classified, and the three-dimensional target corneal image can be obtained by using the acquisition method of the above three-dimensional sample corneal image. After obtaining the three-dimensional target corneal image, it is also necessary to map it into two-dimensional target depth maps of several channels. The acquisition method of the two-dimensional target depth maps can refer to the above two-dimensional sample depth maps and will not be elaborated here. Inputting the two-dimensional target depth maps into the target image classification model means inputting the two-dimensional target depth maps into the first component, the second component, and the third component in sequence to output the target classification result.
[0106] This embodiment provides a method for classifying target images. First, a three-dimensional sample corneal image is obtained, mapped into two-dimensional sample depth maps of several channels, and a sample data set is formed. Then, an initial image classification model including a first component, a second component, and a third component is constructed, and the initial image classification model is trained using the sample data set to obtain a target image classification model for classifying corneal images. The obtained three-dimensional target corneal image is mapped into two-dimensional target depth maps of several channels, and the two-dimensional target depth maps are input into the target image classification model to output a target classification result, thereby achieving accurate classification of the three-dimensional target corneal image and solving the problem of low classification accuracy of the image classification model in the prior art. It provides more accurate results for the preoperative preparation of myopia surgery and provides a true and effective guiding role for the early screening of keratoconus.
[0107] Based on the same inventive concept as the above method for classifying target images, correspondingly, this embodiment also provides a device for classifying target images.
[0108] Figure 3 It is a schematic diagram of the modules of the device for classifying target images provided by the present invention.
[0109] As Figure 3 shown, the above device for classifying target images includes: a data acquisition module 31, a model training module 32, and an image classification module 33.
[0110] Among them, the data acquisition module is used to obtain a three-dimensional sample corneal image, map it into two-dimensional sample depth maps of several channels, and form a sample data set;
[0111] The model training module is used to construct an initial image classification model, train the initial image classification model using the sample data set, and obtain a target image classification model for classifying corneal images. The initial image classification model includes a first component for feature extraction, a second component for feature fusion, and a third component for feature classification;
[0112] The image classification module is used to obtain a three-dimensional target corneal image, map it into two-dimensional target depth maps of several channels, and input the two-dimensional target depth maps into the target image classification model to output a target classification result.
[0113] In the exemplary classification device for the target image, first, a three-dimensional sample corneal image is obtained, mapped into two-dimensional sample depth maps of a plurality of channels, and a sample data set is formed; then, an initial image classification model including a first component, a second component, and a third component is constructed, and the initial image classification model is trained using the sample data set to obtain a target image classification model for corneal image classification; the obtained three-dimensional target corneal image is mapped into two-dimensional target depth maps of a plurality of channels, and the two-dimensional target depth maps are input into the target image classification model to output a target classification result, thereby achieving accurate classification of the three-dimensional target corneal image and solving the problem of low classification accuracy of the image classification model in the prior art.
[0114] In some exemplary embodiments, the model training module includes:
[0115] A feature extraction unit, configured to perform channel splicing on the two-dimensional sample depth map, and use the first component to extract features from the two-dimensional sample depth map after channel splicing to obtain sample features;
[0116] A feature fusion unit, configured to perform cross-channel feature fusion on the sample features based on the second component using a self-attention mechanism to obtain fused features;
[0117] A feature classification unit, configured to classify the fused features using the third component to obtain a classification result;
[0118] A model update unit, configured to obtain the classification error between the classification result and a preset result using a cross-entropy loss function, and update the initial image classification model by backpropagating the classification error to obtain a target image classification model.
[0119] In some exemplary embodiments, the feature fusion unit includes:
[0120] A position feature subunit, configured to obtain the position information of the sample features to obtain position features;
[0121] A feature fusion subunit, configured to perform cross-channel feature fusion on the sample features according to the self-attention mechanism and the position features to obtain fused features.
[0122] In some exemplary embodiments, the feature fusion unit further includes:
[0123] A feature splitting subunit, configured to split the sample features into a plurality of matrices according to a preset splitting rule to obtain sub-sample features;
[0124] An associated feature subunit, configured to respectively obtain the associated features between two sub-sample features;
[0125] A sample feature subunit, configured to determine a target sub-sample feature according to an attention mechanism and associated features;
[0126] A feature fusion subunit, configured to perform cross-channel feature fusion on the target sub-sample feature to obtain a fusion feature.
[0127] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the methods in any one of the embodiments of this embodiment are implemented.
[0128] In one embodiment, please refer to Figure 4 , this embodiment also provides an electronic device 400, including a memory 401, a processor 402, and a computer program stored on the memory and executable on the processor. When the processor 402 executes the computer program, the steps of the method described in any one of the above embodiments are implemented.
[0129] For the computer-readable storage medium in this embodiment, those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to the computer program. The foregoing computer program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above method embodiments are executed; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program codes.
[0130] The electronic device provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication with each other. The memory is used to store a computer program, the communication interface is used for communication, and the processor and the transceiver are used to run the computer program to enable the electronic device to execute each step of the above method.
[0131] In this embodiment, the memory may include a random access memory (Random Access Memory, abbreviated as RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0132] The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0133] In the above embodiments, the reference in the specification to "this embodiment", "an embodiment", "another embodiment", "in some exemplary embodiments" or "other embodiments" means that the specific features, structures or characteristics described in connection with the embodiments are included in at least some embodiments, but not necessarily all embodiments. The repeated occurrences of "this embodiment", "an embodiment", "another embodiment" do not necessarily all refer to the same embodiment.
[0134] In the above embodiments, although the present invention has been described in conjunction with specific embodiments of the present invention, many substitutions, modifications and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other storage structures (e.g., dynamic RAM (DRAM)) may be used in the embodiments discussed. The embodiments of the present invention are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims.
[0135] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.
[0136] The present invention can be used in numerous general-purpose or special-purpose computing system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.
[0137] The present invention may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including storage devices.
[0138] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Any person familiar with this technology may modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those of ordinary skill in the art without departing from the spirit and technical ideas disclosed by the present invention shall still be covered by the claims of the present invention.
Claims
1. A classification method for a target image, characterized in that, Comprising: Obtain a three-dimensional sample corneal image, map it into two-dimensional sample depth maps of a number of channels, and form a sample data set; Construct an initial image classification model, train the initial image classification model using the sample data set, and obtain a target image classification model for corneal image classification. The initial image classification model includes a first component for feature extraction, a second component for feature fusion, and a third component for feature classification; Obtain a three-dimensional target corneal image, map it into two-dimensional target depth maps of a number of channels, and input the two-dimensional target depth maps into the target image classification model to output a target classification result; The training of the initial image classification model using the sample data set to obtain a target image classification model for corneal image classification includes: Perform channel splicing on the two-dimensional sample depth maps, and use the first component to extract features from the spliced two-dimensional sample depth maps of channels to obtain sample features; the first component is a convolutional neural network model. When using the convolutional neural network model to extract features from the spliced two-dimensional sample depth maps of channels, perform feature extraction on it by using multiple convolutions and skip connections; Based on the second component, use the self-attention mechanism to perform cross-channel feature fusion on the sample features to obtain fused features; Use the third component to classify the fused features to obtain a classification result; Use the cross-entropy loss function to obtain the classification error between the classification result and the preset result, and use the backpropagation of the classification error to update the initial image classification model to obtain the target image classification model; The using the self-attention mechanism to perform cross-channel feature fusion on the sample features based on the second component to obtain fused features includes: Obtain the position information of the sample features to obtain position features; Perform cross-channel feature fusion on the sample features according to the self-attention mechanism and the position features to obtain fused features; The sample features are composed of multi-dimensional matrices. The using the self-attention mechanism to perform cross-channel feature fusion on the sample features to obtain fused features includes: Split the sample features into a number of matrices according to a preset splitting rule to obtain sub-sample features; Respectively obtain the correlation features between two sub-sample features; Determine the target sub-sample features according to the attention mechanism and the correlation features; Perform cross-channel feature fusion on the target sub-sample features to obtain fused features; The target sub-sample feature z i is mathematically expressed as: Among them, z i is the i-th target sub-sample feature, i is the label of the target sub-sample feature, x is the sample feature, x = (x1, x2,..., x n ), x1 is the first sub-sample feature, x2 is the second sub-sample feature, x n is the n-th sub-sample feature, n is the total number of sub-sample features, j is the label of the sub-sample feature, x j is the j-th sub-sample feature, x i is the i-th sub-sample feature, a ij is the correlation feature between the sub-sample feature x i and the sub-sample feature x j , V is the input matrix of the self-attention mechanism, W V is the weight matrix corresponding to V; Sub-sample feature x i The associated feature with the sub-sample feature x j has a mathematical expression as follows: where e ij is the associated data of the sub-sample feature x i and the sub-sample feature x k ; e ik is the associated data of the sub-sample feature x i and the sub-sample feature x k ; e ij The mathematical expression of e ij = (x i W Q )(x j W K ) T ; Among them, both Q and K are input matrices of the self-attention mechanism, and W Q is the weight matrix corresponding to Q, and W K is the weight matrix corresponding to K, and T is the transpose of the matrix.
2. The classification method of the target image according to claim 1, wherein The using the self-attention mechanism to perform cross-channel feature fusion on the sample features to obtain fused features includes: Update the target sub-sample features, correlation features, and correlation data according to the position features; The mathematical expression of the updated target sub-sample features is: Among them, z i ′ is the updated feature of the i-th target subsample, a′ ij is the updated subsample feature x i and the associated feature of the subsample feature, is the subsample feature x corresponding to V i and the position feature weight of the subsample feature x j ; Updated subsample feature x i The associated feature with the subsample feature x j has the following mathematical expression: Among them, e′ ij is the updated sub-sample feature x i and the associated data of the sub-sample feature x j ; e′ ik is the updated sub-sample feature x i and the associated data of the sub-sample feature x k ; e′ ij The mathematical expression of Among them, is the sub-sample feature x corresponding to K i and the sub-sample feature x j is the position feature weight of 3. The classification method of the target image according to claim 1, characterized in that, The first component includes a convolutional neural network model, the second component includes a transformer model, and the third component includes a softmax classification model.
4. An apparatus for classifying a target image of the classification method of the target image according to any one of claims 1-3, characterized in that, Comprising: A data acquisition module for obtaining a three-dimensional sample corneal image, mapping it into two-dimensional sample depth maps of a number of channels, and forming a sample data set; A model training module, configured to build an initial image classification model, train the initial image classification model using the sample data set, and obtain a target image classification model for corneal image classification. The initial image classification model includes a first component for feature extraction, a second component for feature fusion, and a third component for feature classification. An image classification module, configured to obtain a three-dimensional target corneal image, map it into a two-dimensional target depth map of a plurality of channels, and input the two-dimensional target depth map into the target image classification model to output a target classification result. The data acquisition module, the model training module, and the image classification module are connected to each other.
5. An electronic device, characterized in that, It includes a processor, a memory, and a communication bus. The communication bus is configured to connect the processor and the memory. The processor is configured to execute the computer program stored in the memory to implement the classification method of the target image according to any one of claims 1-3.
6. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program is configured to cause the computer to execute the classification method of the target image according to any one of claims 1-3.
Citation Information
Patent Citations
Semantic segmentation method and system for automatic driving, electronic equipment and medium
CN114022858A