Image classification method and device, electronic equipment, readable storage medium and chip

Through the comparative learning method, the problem of poor robustness in the classification process of three-dimensional reconnaissance images is solved, and higher discrimination and robustness are achieved.

CN119963883APending Publication Date: 2025-05-09CHINA SHIPBUILDING ZHIHAI INNOVATION RES INST CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411957739.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The robustness of the three-dimensional reconnaissance image classification process is poor and is greatly affected by changes in spatial postures and other changes.

Method used

The contrast learning method is used to classify and learn three-dimensional reconnaissance images, and learn local features and global features through feature learning and contrast loss, determine optimization parameters, and improve the robustness of image classification.

Benefits of technology

Through simultaneous learning of local features and global features, the discriminant and robustness of image classification can be improved, and the compatibility of features and classifiers can be enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963883A_ABST
    Figure CN119963883A_ABST
Patent Text Reader

Abstract

The invention provides an image classification method and device, electronic equipment, a readable storage medium and a chip, and the method comprises the steps: obtaining a first three-dimensional image and at least one feature extraction network; performing data enhancement processing on the first three-dimensional image to determine a second three-dimensional image; respectively determining local areas corresponding to the first three-dimensional image and the second three-dimensional image; performing feature extraction on local areas of the first three-dimensional image and the second three-dimensional image, and determining local feature parameters corresponding to a feature extraction network; determining a left loss parameter and a right loss parameter according to the at least one local feature parameter; determining a global loss parameter according to the at least one local feature parameter; determining classification optimization parameters in a mode of combining the left loss parameter, the right loss parameter and the global loss parameter with feature learning; and determining a classification result according to the classification optimization parameter. Through the scheme of the invention, the robustness and compatibility of three-dimensional reconnaissance image classification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional image processing, and in particular to an image classification method, device, electronic device, readable storage medium and chip. Background Art

[0002] At present, in the field of three-dimensional reconnaissance image processing, classification tasks are one of the core research tasks. Compared with two-dimensional image data, three-dimensional image data has a more complex structure and more dimensional information. Three-dimensional reconnaissance images contain depth information, which requires not only processing and analyzing the information within the layer, but also processing the depth between layers, which is an important feature information. Changes in the spatial posture of three-dimensional reconnaissance images will cause large changes in image data, thereby affecting the robustness of the image classification process. Summary of the invention

[0003] In view of this, the present invention aims to solve the problem of poor robustness in the classification process of three-dimensional reconnaissance images.

[0004] Specifically, the present invention is achieved through the following technical solutions:

[0005] A first aspect of the present invention provides an image classification method.

[0006] A second aspect of the present invention provides an image classification device.

[0007] A third aspect of the present invention provides an electronic device.

[0008] A fourth aspect of the present invention provides a readable storage medium.

[0009] A fifth aspect of the present invention provides a chip.

[0010] The image classification method provided by the present invention includes: acquiring a first three-dimensional image and at least one feature extraction network; performing data enhancement processing on the first three-dimensional image to determine a second three-dimensional image; respectively determining local areas corresponding to the first three-dimensional image and the second three-dimensional image; performing feature extraction on the local areas of the first three-dimensional image and the second three-dimensional image to determine local feature parameters corresponding to the feature extraction network; determining a left loss parameter and a right loss parameter according to at least one local feature parameter; determining a global loss parameter according to at least one local feature parameter; determining a classification optimization parameter by combining the left loss parameter, the right loss parameter and the global loss parameter with feature learning; and determining a classification result according to the classification optimization parameter.

[0011] In some technical solutions, optionally, respectively determining the local areas corresponding to the first three-dimensional image and the second three-dimensional image includes: determining a first plane coordinate system with the center point of the first three-dimensional image as the coordinate origin; determining the local area corresponding to the first three-dimensional image according to the first plane coordinate system, the local area of ​​the first three-dimensional image including a first left area and a first right area; determining a second plane coordinate system with the center point of the second three-dimensional image as the coordinate origin; determining the local area corresponding to the second three-dimensional image according to the second plane coordinate system, the local area of ​​the second three-dimensional image including a second left area and a second right area.

[0012] In some technical solutions, optionally, feature extraction is performed on local areas of the first three-dimensional image and the second three-dimensional image, and determining local feature parameters corresponding to the feature extraction network includes: acquiring at least one feature extraction network; determining weight parameters corresponding to the feature extraction network, and the weight parameters are shared between at least one feature extraction network; determining at least one first feature extraction network corresponding to the first three-dimensional image and at least one second feature extraction network corresponding to the second three-dimensional image; performing feature extraction on the first left area and the first right area according to the first feature extraction network, respectively, and determining local feature parameters corresponding to the first feature extraction network, and the local feature parameters include a first left feature and a first right feature; performing feature extraction on the second left area and the second right area according to the second feature extraction network, respectively, and determining local feature parameters corresponding to the second feature extraction network, and the local feature parameters include a second left feature and a second right feature.

[0013] In some technical solutions, optionally, a left loss parameter and a right loss parameter are determined according to at least one local feature parameter, including: determining the left loss parameter according to a first left feature and a second left feature; determining the right loss parameter according to a first right feature and a second right feature.

[0014] In some technical schemes, optionally, determining a global loss parameter based on at least one local feature parameter includes: obtaining a feature module; splicing a first left feature and a first right feature according to the feature module to determine a first global feature; splicing a second left feature and a second right feature according to the feature module to determine a second global feature; and determining a global loss parameter through the first global feature and the second global feature based on contrastive learning loss.

[0015] In some technical solutions, optionally, determining the classification result according to the classification optimization parameters includes: obtaining a classifier; inputting the classification optimization parameters into the classifier to optimize the classifier; and determining the classification result corresponding to the first three-dimensional image through the classifier.

[0016] The second aspect of the present invention provides an image classification device, including: an acquisition module, used to acquire a first three-dimensional image and at least one feature extraction network; an image enhancement module, used to perform data enhancement processing on the first three-dimensional image to determine a second three-dimensional image; a region division module, used to respectively determine local regions corresponding to the first three-dimensional image and the second three-dimensional image; a feature extraction module, used to perform feature extraction on local regions of the first three-dimensional image and the second three-dimensional image, and determine local feature parameters corresponding to the feature extraction network; a parameter determination module, used to determine a left loss parameter and a right loss parameter based on at least one local feature parameter; determine a global loss parameter based on at least one local feature parameter; a classification optimization module, used to determine classification optimization parameters by combining the left loss parameter, the right loss parameter and the global loss parameter with feature learning; and a result determination module, used to determine a classification result based on the classification optimization parameters.

[0017] An embodiment of the third aspect of the present invention provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction implements the steps in the first aspect when executed by the processor.

[0018] An embodiment of the fourth aspect of the present invention provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps in the first aspect are implemented.

[0019] An embodiment of the fifth aspect of the present invention provides a chip, the chip includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the steps in the first aspect.

[0020] The technical solution provided by the present invention brings at least the following beneficial effects:

[0021] The present invention proposes an image classification method, device, electronic device, readable storage medium and chip, which classify and learn three-dimensional reconnaissance images based on a contrastive learning (CL) method, learn local features of partitions in the three-dimensional reconnaissance image and global features of the entire reconnaissance area through feature learning and contrast loss, determine optimization parameters, and a classifier classifies images according to the optimization parameters to enhance the compatibility between features extracted from the three-dimensional reconnaissance image and the classifier, thereby improving the robustness of three-dimensional reconnaissance image classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0024] Figure 1 A schematic diagram of a flow chart of an image classification method provided by an embodiment of the present invention;

[0025] Figure 2 A schematic diagram of a portion of the flow chart of the image classification method provided by an embodiment of the present invention;

[0026] Figure 3 A schematic diagram of a portion of the flow chart of the image classification method provided by an embodiment of the present invention;

[0027] Figure 4 A schematic diagram of a portion of the flow chart of the image classification method provided by an embodiment of the present invention;

[0028] Figure 5 A schematic diagram of a portion of the flow chart of the image classification method provided by an embodiment of the present invention;

[0029] Figure 6 A schematic diagram of a portion of the flow chart of the image classification method provided by an embodiment of the present invention;

[0030] Figure 7 A schematic block diagram of the structure of an image classification device provided by an embodiment of the present invention;

[0031] Figure 8 A schematic block diagram of the structure of an electronic device provided by an embodiment of the present invention;

[0032] Fig. 9 A schematic diagram of a flow chart of an image classification method provided by an embodiment of the present invention;

[0033] Fig.10 A schematic block diagram of the structure of a feature module provided in an embodiment of the present invention.

[0034] in, Figure 7 and Figure 8 The corresponding relationship between the component names and numbers in is as follows:

[0035] 900: image classification device; 902: acquisition module; 904: image enhancement module; 906: region division module; 908: feature extraction module; 910: parameter determination module; 912: classification optimization module; 914: result determination module; 1000: electronic device; 1109: memory; 1110: processor. DETAILED DESCRIPTION

[0036] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0037] See also Figure 1 The first aspect of the present invention provides an image classification method, comprising the following steps:

[0038] Step S100: acquiring a first three-dimensional image and at least one feature extraction network;

[0039] Step S102: performing data enhancement processing on the first three-dimensional image to determine a second three-dimensional image;

[0040] Step S104: determining local areas corresponding to the first three-dimensional image and the second three-dimensional image respectively;

[0041] Step S106: extracting features from local areas of the first three-dimensional image and the second three-dimensional image, and determining local feature parameters corresponding to the feature extraction network;

[0042] Step S108: determining a left loss parameter and a right loss parameter according to at least one local feature parameter;

[0043] Step S110: determining a global loss parameter according to at least one local feature parameter;

[0044] Step S112: Determine the classification optimization parameters by combining the left loss parameter, the right loss parameter and the global loss parameter with feature learning;

[0045] Step S114: Determine the classification result according to the classification optimization parameters.

[0046] According to the image classification method provided by the present invention, classification learning is performed on three-dimensional reconnaissance images based on a contrastive learning (CL) method. The basic idea is to learn a hidden feature space and maximize the consistency between different enhanced views of the same image by comparing the consistency between different images therein. The self-supervised contrastive learning method used in the present invention adopts a two-stage learning method. The first stage adopts contrastive learning loss to learn features, and the second stage adopts cross entropy loss and other learning classifiers. Specifically, a three-dimensional reconnaissance image and at least one feature extraction network are obtained. In this embodiment, four feature extraction networks for local feature extraction of the three-dimensional reconnaissance image are determined, which are sub-network RG(·), sub-network RP(·), sub-network LG(·) and sub-network LP(·). The feature extraction networks are all small three-dimensional dense convolutional networks (3DDenseNet). After data augmentation processing is performed on the acquired three-dimensional reconnaissance image (i.e., the first three-dimensional image), the generated image is determined to be the second three-dimensional image. The second three-dimensional image is determined by data augmentation to effectively improve the generalization ability of the model and reduce overfitting. The first three-dimensional image and the acquired second three-dimensional image are used as two input quantities input1 (i.e., the first three-dimensional image) and input2 (i.e., the second three-dimensional image) for comparative learning. In order to obtain comprehensive features related to the classification indicators in the entire reconnaissance area, input1 and input2 are divided into regions respectively, that is, each three-dimensional reconnaissance image is evenly divided into two left and right regions, with the purpose of extracting more detailed shallow local features in the divided regions (i.e., local regions). The determined feature extraction network is used to extract features from the local regions corresponding to the divided input1 and input2, that is, the sub-network RG(·) is used to extract features from the right local region of input1, and the sub-network LG(·) is used to extract features from the left local region of input1; the sub-network RP(·) is used to extract features from the right local region of input2, and the sub-network RP(·) is used to extract features from the left local region of input2, and the local features of the lung regions corresponding to the four regions of interest (ROI) are determined, and the local feature parameters corresponding to the four local features of the lung regions are determined by projecting the local features of the lung regions to determine the feature vectors.The local features of the partitioned lung area are mapped to the entire space, and the deep global features of the reconnaissance area are determined by connection. The deep global features are projected to determine the feature vector, and the global feature parameters corresponding to the deep global features are determined. Then, supervised contrast learning loss is used on the regularized feature vector to determine the left loss parameters corresponding to the left area of ​​input1 and input2, the right loss parameters corresponding to the right area of ​​input1 and input2, and the global loss parameters corresponding to the global features. Through the above feature learning method, the shallow local features from the local areas of input1 and input2 and the deep global features of the entire reconnaissance area are effectively learned. Since the local features come from the shallower network, the overlapping area between the receptive fields corresponding to each pixel in the image is smaller and has a higher resolution, which ensures that the information it contains is more detailed and specific; the global features come from the deeper network. With the gradual increase in the number of downsampling times and the number of convolution layers in the forward propagation of the model, the receptive field (Receptive Field) continues to increase, and the overlapping area between the receptive fields corresponding to each pixel in the feature map is also larger, making the captured information more global. The global feature vector from input1 is used for classifier learning, and the classifier is optimized based on the cross entropy loss. The acquired three-dimensional reconnaissance image is classified by the optimized classifier to determine the classification result corresponding to the three-dimensional reconnaissance image, for example, object classification, scene classification, geometric shape classification, and dynamic and static state classification of objects in the image.

[0047] It can be understood that the image classification method is determined by an end-to-end reconnaissance area contrast learning hybrid network, and the local features and global features are learned simultaneously to improve the discriminability and robustness of image classification. More detailed shallow local features can be extracted from local areas. Under the constraints of local features, the global features are richer, and the information contained in the image classification process is more detailed and specific.

[0048] In some embodiments, optionally, Figure 2 As shown, respectively determining the local areas corresponding to the first three-dimensional image and the second three-dimensional image includes:

[0049] Step S1042: determining a first plane coordinate system with the center point of the first three-dimensional image as the coordinate origin;

[0050] Step S1044: determining a local area corresponding to the first three-dimensional image according to the first plane coordinate system, where the local area of ​​the first three-dimensional image includes a first left area and a first right area;

[0051] Step S1046: determining a second plane coordinate system with the center point of the second three-dimensional image as the coordinate origin;

[0052] Step S1048: determining a local area corresponding to the second three-dimensional image according to the second plane coordinate system, wherein the local area of ​​the second three-dimensional image includes a second left area and a second right area.

[0053] In this embodiment, in order to obtain comprehensive features related to the classification index in the entire reconnaissance area and to ensure that the fine-grained information is not ignored during the analysis process, the two inputs of each three-dimensional reconnaissance image are evenly divided into two left and right regions, and then more detailed shallow local features are extracted from the regions. The fine-grained information refers to the specific details provided at a higher precision and more detailed level during the image data processing process, such as the specific surface texture, edge details or substructure features of the reconnaissance target. By determining the fine-grained information, the resolution and discrimination in the three-dimensional reconnaissance image classification process can be improved. Specifically, the center point of the first three-dimensional image is determined as the origin of coordinates, the plane coordinate system constructed is the first plane coordinate system, the plane coordinate system is parallel to the horizontal plane, the y-axis of the plane coordinate system is determined as the symmetry reference axis, all points in the first three-dimensional image with x-axis values ​​less than zero are divided into the left area (i.e., the first left area), all points in the first three-dimensional image with x-axis values ​​greater than zero are divided into the right area (i.e., the first right area), and the areas of the first left area and the first right area are equal; the center point of the second three-dimensional image is determined as the origin of coordinates, the plane coordinate system constructed is the second plane coordinate system, the plane coordinate system is parallel to the horizontal plane, the y-axis of the plane coordinate system is determined as the symmetry reference axis, all points in the second three-dimensional image with x-axis values ​​less than zero are divided into the left area (i.e., the second left area), all points in the second three-dimensional image with x-axis values ​​greater than zero are divided into the right area (i.e., the second right area), and the areas of the second left area and the second right area are equal.

[0054] In some embodiments, optionally, Figure 3 As shown, feature extraction is performed on local areas of the first three-dimensional image and the second three-dimensional image to determine local feature parameters corresponding to the feature extraction network, including:

[0055] Step S1062: Obtain at least one feature extraction network;

[0056] Step S1064: determining a weight parameter corresponding to the feature extraction network, where the weight parameter is shared between at least one feature extraction network;

[0057] Step S1066: determining at least one first feature extraction network corresponding to the first three-dimensional image and at least one second feature extraction network corresponding to the second three-dimensional image;

[0058] Step S1068: performing feature extraction on the first left region and the first right region respectively according to the first feature extraction network, and determining local feature parameters corresponding to the first feature extraction network, where the local feature parameters include the first left feature and the first right feature;

[0059] Step S1070: performing feature extraction on the second left region and the second right region respectively according to the second feature extraction network, and determining local feature parameters corresponding to the second feature extraction network, where the local feature parameters include the second left feature and the second right feature.

[0060] In this embodiment, the network used for local feature extraction is composed of four feature extraction networks with the same structure, namely, sub-network RG(·), sub-network RP(·), sub-network LG(·) and sub-network LP(·). Each feature extraction network is a small 3DDenseNet network, and each DenseNet is composed of three dense block structures. In each Dense Block, the input of each convolution layer is the union of the outputs of all previous convolution layers, and the features learned by this layer will also be directly passed to all subsequent layers as input. In order to ensure the consistency of local features at the same position, sub-network RG(·), sub-network RP(·), sub-network LG(·) and sub-network LP(·) share weights with each other, which is achieved by the operator setting the weight parameters and sharing the weight parameters among sub-network RG(·), sub-network RP(·), sub-network LG(·) and sub-network LP(·). Among them, sub-network RG(·) and sub-network LG(·) are the first feature extraction network, and sub-network RG(·) is used to extract features from the right local area of ​​input1, and sub-network LG(·) is used to extract features from the left local area of ​​input1; sub-network RP(·) and sub-network LP(·) are the second feature extraction network, and sub-network RP(·) is used to extract features from the right local area of ​​input2, and sub-network RP(·) is used to extract features from the left local area of ​​input2 to determine the local features of the lung area corresponding to four regions of interest (ROI), namely, the first left feature, the first right feature, the second left feature and the second right feature.

[0061] In some embodiments, optionally, Figure 4 As shown, determining the left loss parameter and the right loss parameter according to at least one local feature parameter includes:

[0062] Step S1082: determining a left loss parameter according to the first left feature and the second left feature;

[0063] Step S1084: Determine the right loss parameter according to the first right feature and the second right feature.

[0064] In this embodiment, for the local feature learning of the left and right lung regions, the local features (first left feature, first right feature, second left feature, and second right feature) are mapped into feature vectors through the projection module, corresponding to the first left feature vector, the first right feature vector, the second left feature vector, and the second right feature vector, respectively. The above feature vectors are regularized so that the inner product can be used for distance measurement. The supervised contrast learning loss is used for the regularized feature vector, and the local contrast learning loss formula for local feature learning is as follows:

[0065]

[0066] Where L SC R (z i ) is the supervised contrastive learning loss for local feature learning in the right region, z i is the eigenvector. Ri is the anchor sample feature vector, Yes Ri The set of all positive sample feature vectors, Yes Ri The number of all positive samples. Indicates z Rj The eigenvector is the set of eigenvectors belonging to the positive samples, that is, z Ri The feature vector of the positive sample; z Rk is the number of batches except z Ri All sample feature vectors except itself. i, j, k are subscripts. L SC L (z i ) is the supervised contrastive learning loss for local feature learning in the left region, z i is the eigenvector. Li is the anchor sample feature vector, Yes Li The set of all positive sample feature vectors, Yes Li The number of all positive samples. Indicates z Lj The eigenvector is the set of eigenvectors belonging to the positive samples, that is, z Li The feature vector of the positive sample; z Lk is the number of batches except z Li All sample feature vectors except itself. i, j, k are subscripts.

[0067] In some embodiments, optionally, Figure 5 As shown, determining the global loss parameter according to at least one local feature parameter includes:

[0068] Step S1102: Acquire feature modules;

[0069] Step S1104: splicing the first left feature and the first right feature according to the feature module to determine a first global feature;

[0070] Step S1106: splicing the second left feature and the second right feature according to the feature module to determine a second global feature;

[0071] Step S1108: Based on the contrastive learning loss, determine the global loss parameter by using the first global feature and the second global feature.

[0072] In this embodiment, the feature acquisition module (Con-Block module) maps the local features of the partition to the entire space and further obtains the deep global features of the reconnaissance area. The structure of the Con-Block module is as follows: Fig.10 As shown, it includes a BN layer (BN Layer), a linear rectification layer (ReLU Layer), a convolution layer (Conv Layer), a dense block (Dense Block), and an average pooling layer (Average Pooling Layer). Among them, the BN layer (BN Layer), the linear rectification layer (ReLU Layer) and the convolution layer (Conv Layer) constitute a BN linear convolution (BN-ReLU-Convolution) layer; the BN layer (BN Layer), the linear rectification layer (ReLU Layer) and the average pooling layer (Average Pooling Layer) constitute a BN linear average pooling (BN-ReLU-Average pooling) layer. The first left feature and the first right feature are spliced ​​through the Con-Block module to determine the global feature of the reconnaissance area corresponding to input1 (i.e., the first global feature), and the second left feature and the second right feature are spliced ​​through the Con-Block module to determine the global feature of the reconnaissance area corresponding to input2 (i.e., the second global feature). The projection module is used to learn global features of the entire area. The projection module maps the first global feature and the second global feature into the first global feature vector and the second global feature vector, respectively regularizes the first global feature vector and the second global feature vector, and then uses the supervised contrastive learning loss on the regularized feature vector. The contrastive learning loss formula is as follows:

[0073]

[0074] Where L SC G (z i ) is the supervised contrastive learning loss for global feature learning, z iis the eigenvector. Gi is the anchor sample feature vector, Yes Gi The set of all positive sample feature vectors, Yes Gi The number of all positive samples. Indicates z Gj The eigenvector is the set of eigenvectors belonging to the positive samples, that is, z Gi The feature vector of the positive sample; z Gk is the number of batches except z Gi All sample feature vectors except itself. i, j, k are subscripts. In some embodiments, optionally, Figure 6 As shown, the classification results are determined according to the classification optimization parameters, including:

[0075] Step S1122: obtaining a classifier;

[0076] Step S1124: inputting the classification optimization parameters into the classifier to optimize the classifier;

[0077] Step S1126: Determine a classification result corresponding to the first three-dimensional image by a classifier.

[0078] In this embodiment, a classifier for classifying a three-dimensional reconnaissance image is determined. The goal of the classifier learning phase is to learn a classifier with low bias performance based on the effective features obtained from the feature learning phase. The model uses the global feature vector from input1 for classifier learning. In the proposed deep learning network model, a fully connected layer and a sigmoid activation function are applied to the global feature vector to obtain the classification. The classifier is optimized by a cross entropy loss, and the cross entropy loss function is shown as follows:

[0079]

[0080] Where n is the number of samples, y is the true label, is the predicted output.

[0081] The overall loss function is:

[0082] Loos=0.5L SC R +0.5L SC L +L SC G +L CE ;

[0083] Among them, L SC R and L SC L denote the supervised contrast loss for learning local features of right and left partitions, respectively, and L SC Grepresents the supervised contrast loss for global feature learning of the entire reconnaissance area, L CE represents the cross entropy loss used in the classifier learning phase.

[0084] After the classifier is optimized according to the loss function, the classifier can classify at least one acquired first three-dimensional image and determine a classification result.

[0085] In a specific embodiment, the flowchart of the image classification method is as follows: Fig. 9 As shown, Input1 and Input2 are obtained, where Input2 is obtained by data enhancement of Input1. Four ROIs are generated for each input as the input of the feature extraction network. The sub-network RG (·) (Subnet.RG) and the sub-network LP (·) (Subnet.LG) are used to extract features from Input1 respectively, and the first right feature r in the local image features is determined. RG and the first left feature r LG ; Extract features of Input2 through sub-network RP(·)(Subnet.RP) and sub-network LP(·)(Subnet.LP) respectively, and determine the second right feature r in the local image feature RP and the second left feature r LP ; Based on the local projection method (Local projection heads), local image features are mapped and the projection module H r (·) The first right feature r RG and the second right feature r RP Mapped to the first right eigenvector z RG and the second right eigenvector z RP , through the projection module H l (·) The first left feature r LG and the second left feature r LP Mapped to the first left eigenvector z LG and the second left eigenvector z LP ; Perform L2 regularization on the eigenvector, and determine the local left region contrast loss (Lcoal SC loss of left regions) parameter L according to the regularized first left eigenvector and the regularized second left eigenvector SC_L , according to the regularized first right eigenvector and the regularized second right eigenvector, the local right region contrast loss (Lcoal SC loss of right regions) parameter L is determined SC_RConcatenate local image features and determine global image features through the Con-Block module GG and the global image feature r PG , based on the global projection method (Global projection heads), the global image features are mapped through the projection module H G (·) The global image feature r GG and r GP Mapped to the feature vector z GG and z GP Next, we calculate the feature vector z GG and z GP Perform L2 regularization, and then use supervised contrast learning loss on the regularized feature vector to determine the global contrast loss (Global SC loss of the whole regions) parameter L SC_G . And through the logic module f c (·) Perform logits processing to determine the cross entropy loss function L corresponding to the cross entropy loss classifier learning (CE loss for classfier learning) CE .

[0086] In a specific embodiment, the image classification method in the present application includes the following steps:

[0087] a) In order to obtain comprehensive features related to the classification index in the entire reconnaissance area and to avoid the relevant fine-grained information from being submerged in the analysis process, the present invention evenly divides the two inputs of each 3D reconnaissance image into left and right regions, and then extracts more detailed shallow local features in the partitions. Where input2 is obtained from input1 through data enhancement. Fig. 9 As shown, four ROIs are generated for each input as input to the feature extraction network.

[0088] b) The network used for local feature extraction in the model consists of four small 3D DenseNet networks with the same structure, namely sub-network RG(·), sub-network RP(·), sub-network LG(·) and sub-network LP(·). Each DenseNet consists of three dense block structures. In each Dense Block, the input of each convolution layer is the union of the outputs of all previous convolution layers, and the features learned in this layer will be directly passed to all subsequent layers as input. Each Dense Block is connected to the subsequent dense blocks via a Transition Block, which consists of a batch normalization layer (Batch Normalization Layer, BN), a ReLU activation layer, a convolution layer (Convolution Layer) and a pooling layer (Pooling Layer).

[0089] c) In order to ensure the consistency of local features at the same position, sub-network RG(·) and sub-network RP(·), sub-network LG(·) and sub-network LP(·) share weights with each other. The four ROIs (RG, RP, LG and LP) generated above are respectively input into the corresponding sub-networks (sub-network RG(·), sub-network RP(·), sub-network LG(·) and sub-network LP(·)) for partition local feature extraction. Through the above operations, we obtain the lung area local features r corresponding to the four ROIs. RG , r LG , r RP and r LP .

[0090] r RG =RG(RG), r LG =LG(LG);

[0091] r RP =RP(RP), r LP =LP(LP);

[0092] d) In order to map the local features of the partition to the entire space and further obtain the deep global features of the reconnaissance area, the present invention designs a Con-Block module. The structural schematic diagram of the feature module (i.e., the Con-Block module) is as follows: Fig.10 As shown in the figure, it consists of a BN-ReLU-Convolution layer (including BN Layer, ReLU Layer and Conv Layer), a dense block DenseBlock and a BN-ReLU-Average pooling layer (including BN Layer, RELU Layer and Average poolingLayer).RG and r LG Then it is sent to the Con-Block module to obtain the global feature r of the reconnaissance area GG , splicing local features r RP and r LP Then it is sent to the Con-Block module to obtain another global feature r of the reconnaissance area after data enhancement PG The local features and global features obtained above will be used in the local and global feature learning stages based on contrastive learning and the subsequent classification stage.

[0093] e) The feature learning stage learns the local features of the partition and the global features of the entire reconnaissance area based on the SC loss. The purpose of the feature learning stage is to learn a feature space with intra-class compactness and inter-class separability so that the learned features have better effective discriminability.

[0094] f) The proposed network model uses the projection module. After the input passes through the feature extraction network to obtain image features, the projection module is used to map these features into feature vectors that are more suitable for contrastive learning loss. In addition, studies have shown that this projection module can effectively improve feature quality. For local feature learning of the left and right lung regions, the projection module H r (·) The local image feature r RG and r RP Mapped to the feature vector z RG and z RP , through the projection module H l (·) The local image feature r LG and r LP Mapped to the feature vector z LG and z LP . This application will use a multilayer perceptron (MLP) to implement the projection module. Next, we L2 regularize the feature vector so that the inner product can be used for distance measurement. Then, supervised contrastive learning loss is used on the regularized feature vector. The local contrastive learning loss design for local feature learning is as follows:

[0095]

[0096] Where L SC R (z i ) is the supervised contrastive learning loss for local feature learning in the right region, z i is the eigenvector. Ri is the anchor sample feature vector, Yes Ri The set of all positive sample feature vectors, Yes RiThe number of all positive samples. Indicates z Rj The eigenvector is the set of eigenvectors belonging to the positive samples, that is, z Ri The feature vector of the positive sample; z Rk is the number of batches except z Ri All sample feature vectors except itself. i, j, k are subscripts.

[0097] L SC L (z i ) is the supervised contrastive learning loss for local feature learning in the left region, z i is the eigenvector. Li is the anchor sample feature vector, Yes Li The set of all positive sample feature vectors, Yes Li The number of all positive samples. Indicates z Lj The eigenvector is the set of eigenvectors belonging to the positive samples, that is, z Li The feature vector of the positive sample; z Lk is the number of batches except z Li All sample feature vectors except itself. i, j, k are subscripts.

[0098] g) For global feature learning of the entire lung area, the projection module H G (·) The global image feature r GG and r GP Mapped to the feature vector z GG and z GP Next, we calculate the feature vector z GG and z GP Perform L2 regularization, and then use supervised contrastive learning loss on the regularized feature vector. The contrastive learning loss formula for global feature learning is as follows:

[0099]

[0100] Where L SC G (z i ) is the supervised contrastive learning loss for global feature learning, z i is the eigenvector. Gi is the anchor sample feature vector, Yes Gi The set of all positive sample feature vectors, Yes Gi The number of all positive samples. Indicates z Gj The eigenvector is the set of eigenvectors belonging to the positive samples, that is, z GiThe feature vector of the positive sample; z Gk is the number of batches except z Gi All sample feature vectors except itself. i, j, k are subscripts.

[0101] h) Through the above-mentioned feature learning stage, the shallow local features from the partition and the deep global features of the entire reconnaissance area are effectively learned. Among them, the local features come from the shallower network. In the feature map stage, the overlapping area between the receptive fields corresponding to each pixel in the map is smaller and has a higher resolution, which ensures that the information it contains is more detailed and specific; the global features come from the deeper network. As the number of downsampling times and the number of convolution layers gradually increase during the forward propagation of the model, the receptive field continues to increase, and the overlapping area between the receptive fields corresponding to each pixel in the feature map is also larger, the information granularity and resolution are reduced, but the captured information is more global.

[0102] i) Based on the proposed network, the global high-dimensional features from the entire reconnaissance area will be effectively learned, and under the constraint of contrastive learning, the effective features of the entire reconnaissance area will be implicitly restricted to the focused part; under the constraint of local features, the global features are richer and more specific to the classification task.

[0103] j) The goal of the classifier learning phase is to learn a classifier with low bias performance based on the effective features obtained from the feature learning phase. The model uses the global feature vector from input1 for classifier learning. In the proposed deep learning network model, a fully connected layer and a Sigmoid activation function are applied to the global feature vector to obtain the classification. Finally, we optimize the classifier based on the cross entropy loss. The cross entropy loss function is as follows:

[0104]

[0105] Where n is the number of samples, y is the true label, is the predicted output.

[0106] k) The overall loss function of the entire framework is as follows:

[0107] Loss = 0.5L SC R +0.5L SC L +L SC G +L CE ;

[0108] Among them, L SC R and L SC L denote the supervised contrast loss for learning local features of right and left partitions, respectively, and L SC Grepresents the supervised contrast loss for global feature learning of the entire reconnaissance area, L CE represents the cross entropy loss used in the classifier learning phase.

[0109] like Figure 7 As shown, the second aspect of the present invention provides an image classification device 900, which includes: an acquisition module 902, used to acquire a first three-dimensional image and at least one feature extraction network; an image enhancement module 904, used to perform data enhancement processing on the first three-dimensional image to determine a second three-dimensional image; a region division module 906, used to determine local regions corresponding to the first three-dimensional image and the second three-dimensional image respectively; a feature extraction module 908, used to perform feature extraction on local regions of the first three-dimensional image and the second three-dimensional image, and determine local feature parameters corresponding to the feature extraction network; a parameter determination module 910, used to determine a left loss parameter and a right loss parameter according to at least one local feature parameter; determine a global loss parameter according to at least one local feature parameter; a classification optimization module 912, used to determine a classification optimization parameter by combining a left loss parameter, a right loss parameter and a global loss parameter with feature learning; a result determination module 914, used to determine a classification result according to the classification optimization parameter.

[0110] The image classification device 900 provided by the present invention implements the image classification method of the first aspect, wherein the acquisition module 902 acquires a three-dimensional reconnaissance image and at least one feature extraction network, and the acquisition module 902 may acquire at least one three-dimensional reconnaissance image by means of a camera or laser scanning; the image enhancement module 904 performs data augmentation processing on the acquired three-dimensional reconnaissance image, and determines that the generated image is a second three-dimensional image, and determines the second three-dimensional image by means of data augmentation to effectively improve the generalization ability of the model and reduce overfitting; the region division module 906 is used to divide the first three-dimensional reconnaissance image to determine a first right region and a first left region, wherein the first right region and the first left region are equal; and divide the second three-dimensional reconnaissance image to determine a second right region and a second left region, wherein the second right region and the second left region are equal, and the region division module 906 evenly divides the two inputs of each three-dimensional reconnaissance image into two left and right regions, thereby improving the fine-grainedness in the image classification process; the feature extraction The extraction module 908 extracts the local features of the lung area of ​​the two input images through sub-network RG(·), sub-network RP(·), sub-network LG(·) and sub-network LP(·), and determines the first left feature, the first right feature, the second left feature and the second right feature; the parameter determination module 910 determines the left loss parameter and the right loss parameter corresponding to the local features of the lung area through the projection module and regularization processing, and determines the global loss parameter corresponding to the entire reconnaissance area; the classification optimization module 912 applies the full connection layer and the Sigmoid activation function to the global feature vector in the proposed deep learning network model to obtain the classification, optimizes the classifier based on the cross entropy loss (Cross Entropy Loss), and determines the classification optimization parameters; finally, the classification optimization parameters are input into the classifier through the result determination module 914, and the classifier is controlled to classify the acquired three-dimensional reconnaissance image to determine the classification result.

[0111] like Figure 8 As shown, the third aspect of the present invention provides an electronic device 1000, including a processor 1110, a memory 1109, and a program or instruction stored in the memory 1109 and executable on the processor 1110. When the program or instruction is executed by the processor 1110, the various processes of the embodiment of the above-mentioned image classification method are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0112] Wherein, the processor 1110 is used to obtain a first three-dimensional image and at least one feature extraction network;

[0113] Optionally, the processor 1110 is further configured to perform data enhancement processing on the first three-dimensional image to determine a second three-dimensional image;

[0114] Optionally, the processor 1110 is further configured to respectively determine local areas corresponding to the first three-dimensional image and the second three-dimensional image;

[0115] Optionally, the processor 1110 is further configured to perform feature extraction on local areas of the first three-dimensional image and the second three-dimensional image, and determine local feature parameters corresponding to the feature extraction network;

[0116] Optionally, the processor 1110 is further configured to determine a left loss parameter and a right loss parameter according to at least one local feature parameter; and determine a global loss parameter according to at least one local feature parameter;

[0117] Optionally, the processor 1110 is further configured to determine the classification optimization parameter by combining the left loss parameter, the right loss parameter and the global loss parameter with feature learning;

[0118] Optionally, the processor 1110 is further configured to determine a classification result according to classification optimization parameters.

[0119] In a fourth aspect, the present invention provides a readable storage medium on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the embodiment of the above-mentioned image classification method is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. In addition, the data storage capacity corresponding to the image classification method in this application and the data processing speed of each step in the image classification method are increased by the readable storage medium.

[0120] The methods may be implemented in a variety of different ways depending on the specific features and / or example applications. For example, the methods may be implemented by a combination of hardware, firmware, and / or software. For example, in a hardware implementation, the processor may be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, electronic devices, other device units for performing the above functions, and / or combinations thereof.

[0121] A computer readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer readable storage medium may be an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above devices, but is not limited thereto. A non-exhaustive list of more specific examples of computer readable storage media includes: portable computer floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory card, floppy disk, encoding mechanical device (such as a punch card or a groove with a raised structure with instructions recorded) and any suitable combination of the above devices. The computer readable storage medium used herein should not be understood as a transmission signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium, or an electrical signal transmitted through a wire, etc.

[0122] The processor is the processor in the electronic device in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0123] In a fifth aspect, the present invention provides a chip, the chip includes a processor and a communication interface, the communication interface is coupled to the processor, the processor is used to run a program or instruction, implement each process of the above-mentioned image classification method embodiment, and can achieve the same technical effect, to avoid repetition, no further description is given here. In addition, the chip is used to improve the data processing speed of each step in the image classification method in this application.

[0124] Although this specification includes many specific implementation details, these should not be interpreted as limiting the scope of any invention or the scope of protection claimed, but are mainly used to describe the features of the specific embodiments of specific inventions. Certain features described in multiple embodiments in this specification may also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although the features may work as above in certain combinations and even initially claim protection, one or more features from the claimed combination may be removed from the combination in some cases, and the claimed combination may point to a sub-combination or a variation of a sub-combination.

[0125] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that these operations be performed in the particular order shown or performed sequentially, or requiring that all illustrated operations be performed to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.

[0126] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In some implementations, multitasking and parallel processing may be advantageous.

[0127] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0128] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.

Claims

1. An image classification method, characterized in that: include: acquiring a first three-dimensional image and at least one feature extraction network; performing data enhancement processing on the first three-dimensional image to determine a second three-dimensional image; respectively determining local areas corresponding to the first three-dimensional image and the second three-dimensional image; Performing feature extraction on the local areas of the first three-dimensional image and the second three-dimensional image to determine local feature parameters corresponding to the feature extraction network; Determine a left loss parameter and a right loss parameter according to at least one of the local feature parameters; Determining a global loss parameter based on at least one of the local feature parameters; Determine the classification optimization parameter by combining the left loss parameter, the right loss parameter and the global loss parameter with feature learning; A classification result is determined according to the classification optimization parameters.

2. The image classification method according to claim 1, characterized in that: The determining of the local areas corresponding to the first three-dimensional image and the second three-dimensional image respectively includes: Determine a first plane coordinate system with the center point of the first three-dimensional image as the coordinate origin; determining a local area corresponding to the first three-dimensional image according to the first plane coordinate system, wherein the local area of ​​the first three-dimensional image includes a first left area and a first right area; Determine a second plane coordinate system with the center point of the second three-dimensional image as a coordinate origin; A local area corresponding to the second three-dimensional image is determined according to the second plane coordinate system, and the local area of ​​the second three-dimensional image includes a second left area and a second right area.

3. The image classification method according to claim 2, characterized in that: The step of extracting features from the local areas of the first three-dimensional image and the second three-dimensional image and determining local feature parameters corresponding to the feature extraction network includes: Obtain at least one feature extraction network; Determining weight parameters corresponding to the feature extraction networks, wherein the weight parameters are shared between at least one of the feature extraction networks; determining at least one first feature extraction network corresponding to the first three-dimensional image and at least one second feature extraction network corresponding to the second three-dimensional image; Performing feature extraction on the first left region and the first right region respectively according to the first feature extraction network, and determining local feature parameters corresponding to the first feature extraction network, wherein the local feature parameters include a first left feature and a first right feature; According to the second feature extraction network, feature extraction is performed on the second left region and the second right region respectively, and local feature parameters corresponding to the second feature extraction network are determined, where the local feature parameters include a second left feature and a second right feature.

4. The image classification method according to claim 3, characterized in that: The determining of the left loss parameter and the right loss parameter according to at least one of the local feature parameters comprises: Determine a left loss parameter according to the first left feature and the second left feature; A right loss parameter is determined according to the first right feature and the second right feature.

5. The image classification method according to claim 3, characterized in that: Determining the global loss parameter according to at least one of the local feature parameters comprises: Get feature module; splicing the first left feature and the first right feature according to the feature module to determine a first global feature; splicing the second left feature and the second right feature according to the feature module to determine a second global feature; Based on contrastive learning loss, a global loss parameter is determined by the first global feature and the second global feature.

6. The image classification method according to claim 1, characterized in that: Determining the classification result according to the classification optimization parameter includes: Get the classifier; Inputting the classification optimization parameters into the classifier to optimize the classifier; A classification result corresponding to the first three-dimensional image is determined by the classifier.

7. An image classification device, characterized in that: include: An acquisition module, configured to acquire a first three-dimensional image and at least one feature extraction network; An image enhancement module, configured to perform data enhancement processing on the first three-dimensional image to determine a second three-dimensional image; A region division module, used to determine local regions corresponding to the first three-dimensional image and the second three-dimensional image respectively; a feature extraction module, configured to extract features from the local regions of the first three-dimensional image and the second three-dimensional image, and determine local feature parameters corresponding to the feature extraction network; A parameter determination module, configured to determine a left loss parameter and a right loss parameter according to at least one of the local feature parameters; and determine a global loss parameter according to at least one of the local feature parameters; A classification optimization module, used for determining classification optimization parameters by combining the left loss parameter, the right loss parameter and the global loss parameter with feature learning; The result determination module is used to determine the classification result according to the classification optimization parameters.

8. An electronic device, characterized in that: The method comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the image classification method as claimed in any one of claims 1 to 6.

9. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the image classification method according to any one of claims 1 to 6 are implemented.

10. A chip, characterized in that: The chip includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or an instruction to implement the steps of the image classification method according to any one of claims 1 to 6.