Object recognition and object recognition model training method, device, equipment and medium
By employing a fusion network of 2D and 3D features in 3D images, the problem of low object recognition accuracy caused by the neglect of connections between 3D voxel images is solved, and higher object recognition accuracy is achieved.
Patent Information
- Application Number
- CN202310803923.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-30
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2043-06-30
AI Technical Summary
When recognizing objects in 3D images, the spatial relationships between 3D voxel maps are ignored, resulting in low object recognition accuracy. In particular, key information of larger objects is incomplete, while smaller objects are easily affected by noise.
The method employs a two-dimensional global feature extraction network, a feature processing network, a three-dimensional encoding/decoding network, and a multi-dimensional feature fusion network. By cropping a three-dimensional image into multiple voxel images and using Gaussian weight codes to weight overlapping regions, the method combines two-dimensional and three-dimensional features to improve the accuracy of object recognition.
By integrating two-dimensional and three-dimensional features, the spatial context information of objects is effectively recovered, improving the accuracy of object recognition in three-dimensional images.
Smart Images

Figure CN116778470B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image data processing, and particularly relates to an object recognition method and an object recognition model training method, device and medium. BACKGROUND
[0002] The data contained in a three-dimensional image is much more than that in a two-dimensional image. When an object in the three-dimensional image is recognized, the entire three-dimensional image data is too large, and thus the entire three-dimensional image data needs to be split into multiple three-dimensional voxel graphs (also referred to as voxel blocks) for object recognition. However, object recognition is performed on each three-dimensional voxel graph respectively, and the connection between the three-dimensional voxel graphs is ignored, and the spatial context information of the three-dimensional voxel graphs is lost. For a large-size object, the key information in the object is easily split into multiple three-dimensional voxel graphs, and the loss of the spatial context information causes the key information of the object in each three-dimensional voxel graph to be incomplete, the correlation between the key information of the object is lost, and the accuracy of object recognition is low. For a small-size object, the pixels of the object are less, and the features that can be extracted are also relatively less, and thus the object is easily interfered by noise, and especially when the small-size object is split into multiple three-dimensional voxel graphs, the accuracy of object recognition is low. SUMMARY
[0003] Embodiments of the present application aim to provide an object recognition method and an object recognition model training method, device and medium to improve the accuracy of object recognition in a three-dimensional image. The specific technical solutions are as follows.
[0004] In a first aspect, an object recognition method is provided, and the method comprises the following steps.
[0005] Obtaining a three-dimensional image to be recognized, and obtaining a pre-trained object recognition model, wherein the object recognition model comprises a two-dimensional global feature extraction network, a two-dimensional feature processing network, a three-dimensional encoding / decoding network and a multi-dimensional feature fusion network.
[0006] Cutting the three-dimensional image to be recognized into multiple three-dimensional voxel graphs to be recognized.
[0007] Inputting the three-dimensional image to be recognized and the multiple three-dimensional voxel graphs to be recognized into the object recognition model to obtain an object recognition result.
[0008] Optionally, the two-dimensional feature processing network comprises a two-dimensional encoding / decoding network.
[0009] Alternatively, the two-dimensional feature processing network comprises a two-dimensional encoding network and does not comprise a two-dimensional decoding network.
[0010] Optionally, the inputting the to-be-recognized three-dimensional image and the plurality of to-be-recognized three-dimensional voxel maps into the object recognition model to obtain an object recognition result comprises:
[0011] extracting a target two-dimensional global feature of the to-be-recognized three-dimensional image by using the two-dimensional global feature extraction network;
[0012] performing feature coding on the target two-dimensional global feature by using the two-dimensional feature processing network to obtain a plurality of target two-dimensional image coding features;
[0013] performing feature coding on the to-be-recognized three-dimensional voxel map by using the three-dimensional coding network to obtain a plurality of second target three-dimensional image coding features;
[0014] fusing the target two-dimensional image coding features and the second target three-dimensional image coding features by using the multi-dimensional feature fusion network to obtain a plurality of target fusion features;
[0015] decoding each target fusion feature and at least one second target three-dimensional image coding feature by using the three-dimensional decoding network to obtain a second object prediction result of the to-be-recognized three-dimensional voxel map, wherein the object recognition result comprises the second object prediction result of each to-be-recognized three-dimensional voxel map.
[0016] Optionally, the cropping the to-be-recognized three-dimensional image into a plurality of to-be-recognized three-dimensional voxel maps comprises:
[0017] cropping the to-be-recognized three-dimensional image into a plurality of to-be-recognized three-dimensional voxel maps with overlapping by using a sliding window method, wherein there is a pixel area overlap between adjacent to-be-recognized three-dimensional voxel maps;
[0018] The method further comprises:
[0019] splicing the second object prediction result of each to-be-recognized three-dimensional voxel map and weighting the pixel area with overlapping by using a Gaussian weight code to obtain an object recognition result.
[0020] Optionally, the two-dimensional global feature extraction network comprises a three-dimensional convolution block layer and a pooling layer, and the extracting a target two-dimensional global feature of the to-be-recognized three-dimensional image by using the two-dimensional global feature extraction network comprises:
[0021] inputting the to-be-recognized three-dimensional image into the two-dimensional global feature extraction network, coding and mapping the to-be-recognized three-dimensional image to a high dimension by using the three-dimensional convolution block layer to obtain a high-dimensional first target three-dimensional image coding feature;
[0022] performing a pooling operation on the high-dimensional first target three-dimensional image coding feature by using the pooling layer to obtain a target two-dimensional global feature of the to-be-recognized three-dimensional image.
[0023] Optionally, the pooling operation on the high-dimensional first target three-dimensional image coding features by using the pooling layer obtains the target two-dimensional global feature of the three-dimensional image to be recognized, and the method further comprises:
[0024] The pooling operation on the high-dimensional first target three-dimensional image coding features by using the pooling layer obtains the target two-dimensional global feature of the three-dimensional image to be recognized, and the method further comprises:
[0025] The pooling operation on the high-dimensional first target three-dimensional image coding features by using the pooling layer obtains the target two-dimensional global feature of the three-dimensional image to be recognized, and the method further comprises: I global The target two-dimensional global feature of the three-dimensional image to be recognized is G input The three-dimensional image to be recognized is V input The high-dimensional first target three-dimensional image coding feature is F The pooling operation on the high-dimensional first target three-dimensional image coding features by using the pooling layer obtains the target two-dimensional global feature of the three-dimensional image to be recognized, and the method further comprises:
[0026] Optionally, the two-dimensional global feature extraction network further comprises a two-dimensional convolution block layer, and after the pooling operation on the high-dimensional first target three-dimensional image coding features by using the pooling layer obtains the target two-dimensional global feature of the three-dimensional image to be recognized, the method further comprises:
[0027] The two-dimensional global feature of the three-dimensional image to be recognized is reduced in dimension by using the two-dimensional convolution block layer.
[0028] Optionally, the two-dimensional coding network comprises k two-dimensional coding layers, and the three-dimensional coding network comprises K three-dimensional coding layers, and the fusion of the target two-dimensional image coding feature and the second target three-dimensional image coding feature by using the multi-dimensional feature fusion network obtains a plurality of target fusion features, and the method further comprises:
[0029] The target two-dimensional image coding feature output by the lth two-dimensional coding layer is fused with the second target three-dimensional image coding feature output by the lth three-dimensional coding layer by using the multi-dimensional feature fusion network to obtain the lth target fusion feature, wherein l∈[1,K-1].
[0030] Optionally, the three-dimensional decoding network comprises K three-dimensional decoding layers and a three-dimensional convolution layer, and the decoding of each target fusion feature and at least one second target three-dimensional image coding feature by using the three-dimensional decoding network obtains the second object prediction result of the three-dimensional voxel image to be recognized, and the method further comprises:
[0031] The target image feature output by the Kth three-dimensional coding layer is decoded by using the first three-dimensional decoding layer.
[0032] decoding the target image feature output by the (t-1)th three-dimensional decoding layer and the (K+1-t)th target fusion feature by using the tth three-dimensional decoding layer, wherein t [2, K];
[0033] performing convolution operation on the target image feature output by the Kth three-dimensional decoding layer by using a three-dimensional convolution layer to obtain the recognition result corresponding to the to-be-recognized three-dimensional voxel graph.
[0034] Optionally, the fusion of the target two-dimensional image coding feature output by the lth two-dimensional encoding layer and the second target three-dimensional image coding feature output by the lth three-dimensional encoding layer to obtain the lth target fusion feature by using the multi-dimensional feature fusion network comprises:
[0035] extracting the target two-dimensional image coding feature output by the lth two-dimensional encoding layer by using the multi-dimensional feature fusion network to obtain the lth target two-dimensional image conversion feature, wherein the lth target two-dimensional image conversion feature has the same number of channels as the lth second target three-dimensional image coding feature;
[0036] determining the feature region corresponding to the lth target two-dimensional image conversion feature of the to-be-recognized three-dimensional voxel graph according to the position information of the to-be-recognized three-dimensional voxel graph in the to-be-recognized three-dimensional image, to obtain the lth target two-dimensional image region feature;
[0037] adjusting the size of the lth target two-dimensional image region feature to the size of the lth second target three-dimensional image coding feature to obtain the lth target two-dimensional image region feature after size adjustment;
[0038] fusing the lth target two-dimensional image region feature after size adjustment and the lth second target three-dimensional image coding feature to obtain the lth target fusion feature.
[0039] Optionally, the determination of the feature region corresponding to the to-be-recognized three-dimensional voxel graph in the lth target two-dimensional image conversion feature according to the position information of the to-be-recognized three-dimensional voxel graph in the to-be-recognized three-dimensional image to obtain the lth target two-dimensional image region feature comprises:
[0040] obtaining the position information (i, j, k) of the to-be-recognized three-dimensional voxel graph in the to-be-recognized three-dimensional image;
[0041] determining the feature region (i l , j l ) corresponding to the to-be-recognized three-dimensional voxel graph in the lth target two-dimensional image conversion feature according to the following formula:
[0042]
[0043]
[0044] wherein, the size of the convolution kernel in the two-dimensional encoding layer is (K V ,K W ,K H ), and the step length of the pooling calculation in the two-dimensional encoding layer is S.
[0045] In a second aspect, the embodiments of the present application provide a method for training an object recognition model, the method comprising:
[0046] obtaining a sample three-dimensional image, an object label of the sample three-dimensional image, and an object recognition model to be trained, wherein the object recognition model comprises a two-dimensional global feature extraction network, a two-dimensional feature processing network, a three-dimensional encoding / decoding network, and a multi-dimensional feature fusion network, and the two-dimensional feature processing network comprises a two-dimensional encoding / decoding network;
[0047] cropping the sample three-dimensional image into a plurality of sample three-dimensional voxel images;
[0048] inputting the sample three-dimensional image and the sample three-dimensional voxel images into the object recognition model to obtain a first object prediction result of the sample three-dimensional image and a second prediction result of the sample three-dimensional voxel images, wherein the first object prediction result of the sample three-dimensional image is an output of the two-dimensional decoding network, and the second prediction result of the sample three-dimensional voxel images is an output of the three-dimensional decoding network;
[0049] determining a two-dimensional network loss according to the first object prediction result and the object label, and determining a three-dimensional network loss according to the second object prediction result and the object label;
[0050] adjusting parameters of the two-dimensional global feature extraction network, the two-dimensional encoding / decoding network, and the three-dimensional encoding / decoding network according to the two-dimensional network loss and the three-dimensional network loss;
[0051] selecting other sample three-dimensional image data to continue training until a preset ending condition is met, and obtaining a trained object recognition model.
[0052] Optionally, the cropping of the sample three-dimensional image into a plurality of sample three-dimensional voxel images comprises:
[0053] in a case where the sample three-dimensional image is a positive sample, cropping a region containing an object in the sample three-dimensional image into m sample three-dimensional voxel images, and cropping a remaining region into n-m sample three-dimensional voxel images, wherein the positive sample is used to indicate that the sample three-dimensional image contains an object, and m≤n.
[0054] In a case that the sample three-dimensional image is a negative sample, the sample three-dimensional image is cropped into q sample three-dimensional voxel images, the negative sample is used to indicate that the sample three-dimensional image does not contain an object, and q < n.
[0055] Optionally, the inputting the sample three-dimensional image and the sample three-dimensional voxel image into the object recognition model to obtain a first object prediction result of the sample three-dimensional image and a second prediction result of the sample three-dimensional voxel image comprises:
[0056] extracting a sample two-dimensional global feature of the sample three-dimensional image by using a two-dimensional global feature extraction network;
[0057] performing feature coding on the sample two-dimensional global feature by using a two-dimensional coding network to obtain a plurality of sample two-dimensional image coding features;
[0058] performing decoding on each of the sample two-dimensional image coding features by using a two-dimensional decoding network to obtain the first object prediction result of the sample three-dimensional image;
[0059] performing feature coding on the sample three-dimensional voxel image by using a three-dimensional coding network to obtain a plurality of second sample three-dimensional image coding features;
[0060] fusing the sample two-dimensional image coding features and the second sample three-dimensional image coding features by using a multi-dimensional feature fusion network to obtain a plurality of sample fusion features;
[0061] performing decoding on each of the sample fusion features and at least one of the second sample three-dimensional image coding features by using a three-dimensional decoding network to obtain the second object prediction result of the sample three-dimensional voxel image.
[0062] Optionally, the two-dimensional coding network comprises K two-dimensional coding layers, the three-dimensional coding network comprises K three-dimensional coding layers, and the fusing the sample two-dimensional image coding features and the second sample three-dimensional image coding features by using the multi-dimensional feature fusion network to obtain the plurality of sample fusion features comprises:
[0063] fusing a sample two-dimensional image coding feature output by an lth two-dimensional coding layer and a second sample three-dimensional image coding feature output by an lth three-dimensional coding layer by using the multi-dimensional feature fusion network to obtain an lth sample fusion feature, where l ∈ [1, K-1].
[0064] Optionally, the three-dimensional decoding network comprises K three-dimensional decoding layers and a three-dimensional convolution layer, and the performing decoding on each of the sample fusion features and at least one of the second sample three-dimensional image coding features by using the three-dimensional decoding network to obtain the second object prediction result of the sample three-dimensional voxel image comprises:
[0065] decoding a sample image feature output by the Kth three-dimensional encoding layer by using a first three-dimensional decoding layer;
[0066] decoding the sample image feature output by the (t-1)th three-dimensional decoding layer and the (K+1-t)th sample fusion feature by using a tth three-dimensional decoding layer, where t∈[2,K];
[0067] performing convolution operation on the sample image feature output by the Kth three-dimensional decoding layer by using a three-dimensional convolution layer to obtain a second object prediction result of the sample three-dimensional voxel image.
[0068] In a third aspect, an object recognition device is provided, and the device comprises:
[0069] a first obtaining module, configured to obtain a three-dimensional image to be recognized and a pre-trained object recognition model, wherein the object recognition model comprises a two-dimensional global feature extraction network, a two-dimensional feature processing network, a three-dimensional encoding / decoding network and a multi-dimensional feature fusion network;
[0070] a first cropping module, configured to crop the three-dimensional image to be recognized into a plurality of three-dimensional voxel images to be recognized;
[0071] a recognition module, configured to input the three-dimensional image to be recognized and the plurality of three-dimensional voxel images to be recognized into the object recognition model to obtain an object recognition result.
[0072] Optionally, the two-dimensional feature processing network comprises a two-dimensional encoding / decoding network.
[0073] Or, the two-dimensional feature processing network comprises a two-dimensional encoding network and does not comprise a two-dimensional decoding network.
[0074] Optionally, the recognition module is specifically configured to:
[0075] extract target two-dimensional global features of the three-dimensional image to be recognized by using the two-dimensional global feature extraction network;
[0076] perform feature encoding on the target two-dimensional global features by using the two-dimensional feature processing network to obtain a plurality of target two-dimensional image encoding features;
[0077] perform feature encoding on the three-dimensional voxel images to be recognized by using the three-dimensional encoding network to obtain a plurality of second target three-dimensional image encoding features;
[0078] fuse the target two-dimensional image encoding features and the second target three-dimensional image encoding features by using the multi-dimensional feature fusion network to obtain a plurality of target fusion features;
[0079] decoding each of the target fusion feature and at least one second target three-dimensional image coding feature by using the three-dimensional decoding network to obtain a second object prediction result of the three-dimensional voxel image to be recognized, wherein the object recognition result comprises the second object prediction result of each three-dimensional voxel image to be recognized.
[0080] Optionally, the first cropping module is configured to crop the three-dimensional image to be recognized into a plurality of three-dimensional voxel images to be recognized by using a sliding window method, wherein there is a pixel region overlap between adjacent three-dimensional voxel images to be recognized.
[0081] The apparatus further comprises a module configured to perform the following operation: splicing the second object prediction result of each three-dimensional voxel image and weighting the pixel region overlap by using a Gaussian weight code to obtain an object recognition result.
[0082] Optionally, the recognition module is specifically configured to:
[0083] inputting the three-dimensional image to be recognized into the two-dimensional global feature extraction network, encoding and mapping the three-dimensional image to be recognized to a high dimension by using the three-dimensional convolution block layer to obtain a high-dimensional first target three-dimensional image coding feature;
[0084] performing a pooling operation on the high-dimensional first target three-dimensional image coding feature by using the pooling layer to obtain a target two-dimensional global feature of the three-dimensional image to be recognized.
[0085] Optionally, the recognition module is specifically configured to:
[0086] performing processing on the high-dimensional first target three-dimensional image coding feature by using the pooling layer and an aggregation function to obtain a target two-dimensional global feature of the three-dimensional image to be recognized.
[0087] wherein the aggregation function is I global is the target two-dimensional global feature of the three-dimensional image to be recognized, input is the three-dimensional image to be recognized, ConvBlock(G input ) is the high-dimensional first target three-dimensional image coding feature, indicates an average pooling in the axial direction.
[0088] Optionally, the two-dimensional global feature extraction network further comprises a two-dimensional convolution block layer, and the apparatus further comprises a module configured to perform the following operation:
[0089] performing dimension reduction on the target two-dimensional global feature of the three-dimensional image to be recognized by using the two-dimensional convolution block layer.
[0090] Optionally, the two-dimensional coding network comprises k two-dimensional coding layers, the three-dimensional coding network comprises K three-dimensional coding layers, and the identification module is specifically configured to:
[0091] The multi-dimensional feature fusion network is used to fuse the target two-dimensional image coding feature output by the lth two-dimensional coding layer and the second target three-dimensional image coding feature output by the lth three-dimensional coding layer to obtain an lth target fusion feature, where l [1, K-1].
[0092] Optionally, the three-dimensional decoding network comprises K three-dimensional decoding layers and a three-dimensional convolution layer, and the identification module is specifically configured to:
[0093] The first three-dimensional decoding layer is used to decode the target image feature output by the Kth three-dimensional coding layer;
[0094] The tth three-dimensional decoding layer is used to decode the target image feature output by the t-1th three-dimensional decoding layer and the K+1-tth target fusion feature, where t [2, K];
[0095] The three-dimensional convolution layer is used to perform convolution operation on the target image feature output by the Kth three-dimensional decoding layer to obtain the identification result corresponding to the to-be-identified three-dimensional voxel image.
[0096] Optionally, the identification module is specifically configured to:
[0097] The multi-dimensional feature fusion network is used to extract features from the target two-dimensional image coding feature output by the lth two-dimensional coding layer to obtain an lth target two-dimensional image conversion feature, where the channel number of the lth target two-dimensional image conversion feature is the same as that of the lth second target three-dimensional image coding feature;
[0098] According to the position information of the currently input to-be-identified three-dimensional voxel image in the to-be-identified three-dimensional image, the feature region corresponding to the lth target two-dimensional image conversion feature of the currently input to-be-identified three-dimensional voxel image is determined to obtain an lth target two-dimensional image region feature;
[0099] The size of the lth target two-dimensional image region feature is adjusted to the size of the lth second target three-dimensional image coding feature to obtain an lth target two-dimensional image region feature after size adjustment;
[0100] The lth target two-dimensional image region feature after size adjustment is fused with the lth second target three-dimensional image coding feature to obtain an lth target fusion feature.
[0101] Optionally, the identification module is specifically configured to:
[0102] The position information (i, j, k) of the currently input to-be-identified three-dimensional voxel image in the to-be-identified three-dimensional image is obtained.
[0103] The feature region (i, j) corresponding to the current input three-dimensional voxel map in the lth target two-dimensional image conversion feature is determined according to the following formula: l , j l ):
[0104]
[0105]
[0106] wherein the size of the convolution kernel in the two-dimensional encoding layer is (K V , K W , K H ), and the step length of the pooling calculation in the two-dimensional encoding layer is S.
[0107] In a fourth aspect, an object recognition model training apparatus is provided, and the apparatus comprises:
[0108] A second acquisition module is configured to acquire a sample three-dimensional image, an object label of the sample three-dimensional image, and an object recognition model to be trained, wherein the object recognition model comprises a two-dimensional global feature extraction network, a two-dimensional feature processing network, a three-dimensional encoding / decoding network, and a multi-dimensional feature fusion network, and the two-dimensional feature processing network comprises a two-dimensional encoding / decoding network.
[0109] A second cropping module is configured to crop the sample three-dimensional image into a plurality of sample three-dimensional voxel maps.
[0110] A first determination module is configured to input the sample three-dimensional image and the sample three-dimensional voxel maps into the object recognition model to obtain a first object prediction result of the sample three-dimensional image and a second prediction result of the sample three-dimensional voxel maps, wherein the first object prediction result of the sample three-dimensional image is the output of the two-dimensional decoding network, and the second prediction result of the sample three-dimensional voxel maps is the output of the three-dimensional decoding network.
[0111] A second determination module is configured to determine a two-dimensional network loss according to the first object prediction result and the object label, and determine a three-dimensional network loss according to the second object prediction result and the object label.
[0112] An adjustment module is configured to adjust the parameters of the two-dimensional global feature extraction network, the two-dimensional encoding / decoding network, and the three-dimensional encoding / decoding network according to the two-dimensional network loss and the three-dimensional network loss.
[0113] A training module is configured to select other sample three-dimensional image data for continuous training until a preset ending condition is met, and obtain a trained object recognition model.
[0114] Optionally, the second cropping module is specifically used for:
[0115] In a case where the sample three-dimensional image is a positive sample, cropping a region containing an object in the sample three-dimensional image into m sample three-dimensional voxel images, and cropping a remaining region into n-m sample three-dimensional voxel images, the positive sample is used to indicate that the sample three-dimensional image contains an object, and m≤n;
[0116] In a case where the sample three-dimensional image is a negative sample, cropping the sample three-dimensional image into q sample three-dimensional voxel images, the negative sample is used to indicate that the sample three-dimensional image does not contain an object, and q<n.
[0117] Optionally, the first determining module is specifically used for:
[0118] extracting sample two-dimensional global features of the sample three-dimensional image by using a two-dimensional global feature extraction network;
[0119] performing feature coding on the sample two-dimensional global features by using a two-dimensional coding network to obtain a plurality of sample two-dimensional image coding features;
[0120] performing decoding on each of the sample two-dimensional image coding features by using a two-dimensional decoding network to obtain a first object prediction result of the sample three-dimensional image;
[0121] performing feature coding on the sample three-dimensional voxel images by using a three-dimensional coding network to obtain a plurality of second sample three-dimensional image coding features;
[0122] fusing the sample two-dimensional image coding features and the second sample three-dimensional image coding features by using a multi-dimensional feature fusion network to obtain a plurality of sample fusion features;
[0123] performing decoding on each of the sample fusion features and at least one of the second sample three-dimensional image coding features by using a three-dimensional decoding network to obtain a second object prediction result of the sample three-dimensional voxel images.
[0124] Optionally, the two-dimensional coding network includes K two-dimensional coding layers, the three-dimensional coding network includes K three-dimensional coding layers, and the first determining module is specifically used for:
[0125] fusing sample two-dimensional image coding features output by an lth two-dimensional coding layer and second sample three-dimensional image coding features output by an lth three-dimensional coding layer by using a multi-dimensional feature fusion network to obtain an lth sample fusion feature, where l∈[1,K-1].
[0126] Optionally, the three-dimensional decoding network includes K three-dimensional decoding layers and a three-dimensional convolution layer, and the first determining module is specifically used for:
[0127] decoding the sample image features output by the Kth three-dimensional encoding layer using a first three-dimensional decoding layer;
[0128] decoding the sample image features output by the (t-1)th three-dimensional decoding layer and the (K+1-t)th sample fusion features using the tth three-dimensional decoding layer, where t∈[2,K];
[0129] performing convolution operation on the sample image features output by the Kth three-dimensional decoding layer using a three-dimensional convolution layer to obtain the second object prediction result of the sample three-dimensional voxel map.
[0130] In a fifth aspect, an electronic device is provided, including a processor and a memory;
[0131] The memory is configured to store a computer program.
[0132] The processor is configured to execute the program stored in the memory to implement the object recognition method or the object recognition model training method.
[0133] In a sixth aspect, a computer readable storage medium is also provided, which stores a computer program. The computer program is executed by a processor to implement the object recognition method or the object recognition model training method.
[0134] The embodiments of the present application have the following beneficial effects:
[0135] The object recognition method and the object recognition model training method, the device and the medium provided by the embodiments of the present application obtain a sample three-dimensional image, object annotation of the sample three-dimensional image, a plurality of sample three-dimensional voxel maps and an object recognition model to be trained. The object recognition model includes a two-dimensional global feature extraction network, a two-dimensional feature processing network, a three-dimensional encoding / decoding network and a multi-dimensional feature fusion network. The sample three-dimensional image and the sample three-dimensional voxel map are input into the object recognition model to obtain a first object prediction result of the sample three-dimensional image and a second prediction result of the sample three-dimensional voxel map. The parameters of the object recognition model are adjusted according to the first object prediction result, the second prediction result and the object annotation. In this way, the trained object recognition model is obtained. Then, the object in the three-dimensional image to be recognized is recognized based on the trained object recognition model, and the accuracy of object recognition is improved.
[0136] Of course, implementing any product or method of the present application does not necessarily require all the advantages described above. BRIEF DESCRIPTION OF DRAWINGS
[0137] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the description of the embodiments or the prior art will be briefly introduced. Obviously, the accompanying drawings in the following description only represent some embodiments of the present application, and other embodiments can be obtained by those of ordinary skill in the art based on these drawings.
[0138] Figure 1 A schematic diagram of an object recognition model training method according to an embodiment of the present application;
[0139] Figure 2 A schematic diagram of a two-dimensional global feature extraction network according to an embodiment of the present application;
[0140] Figure 3 A schematic diagram of an object recognition model according to an embodiment of the present application;
[0141] Figure 4 A schematic diagram of a multi-dimensional feature fusion network according to an embodiment of the present application;
[0142] Figure 5 A schematic diagram of an object recognition method according to an embodiment of the present application;
[0143] Figure 6 A schematic diagram of an object recognition device according to an embodiment of the present application;
[0144] Figure 7 A schematic diagram of an object recognition model training device according to an embodiment of the present application;
[0145] Figure 8 A schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0146] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments only represent some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application are within the scope of protection of the present application.
[0147] First, the professional terms in the present application are explained.
[0148] Convolutional Neural Networks (CNN) is a kind of feedforward neural network containing convolution calculation and having a deep structure, and is one of the representative algorithms of deep learning. Convolutional Neural Networks has the ability of representation learning, and can perform translation-invariant classification on input information according to its hierarchical structure. Unlike conventional neural networks, the neurons in each layer of the convolutional neural network will only be connected to a small block (receptive field) in the previous layer, rather than taking a full connection mode.
[0149] Gaussian weight code, because the accuracy of the network will decrease at the edge of the block, directly assembling all the block prediction results will have obvious edge effect, and there is overlap when cutting the block, therefore, before assembling the block, first use Gaussian weight code to give different weights to the block center and edge results, after regularization, the edge effect can be eliminated.
[0150] In the related art, the maximum density projection of a three-dimensional image in the sagittal direction is taken as two-dimensional global information, and three-dimensional voxel image information is fused to realize object recognition based on deep learning. However, the global information obtained by the maximum density projection method will lose a lot of feature information, and thus the accuracy of object recognition is not high.
[0151] Therefore, the embodiment of the present application provides an object recognition model training method, referring to Figure 1 The method comprises the following steps:
[0152] In S101, a sample three-dimensional image, an object label of the sample three-dimensional image, and an object recognition model to be trained are obtained, wherein the object recognition model comprises a two-dimensional global feature extraction network, a two-dimensional feature processing network, a three-dimensional encoding / decoding network, and a multi-dimensional feature fusion network, and the two-dimensional feature processing network comprises a two-dimensional encoding / decoding network.
[0153] In a possible embodiment, one sample three-dimensional image is selected from a plurality of sample three-dimensional images to train the object recognition model. The sample three-dimensional image corresponds to an object label, and the object label represents a pixel point set of an object in the sample three-dimensional image. For example, when the sample three-dimensional image is a brain MRI image, the object label represents a pixel point set of a tumor in the brain MRI image. The object label of the sample three-dimensional image can be obtained by artificial labeling or from a related open source database. In a possible embodiment,
[0154] The encoding / decoding network in the object recognition model can be a U-net network, specifically, a ResUNet, a DenseUNet, or the like. The U-net network is a segmentation network specially proposed for medical images, and mainly consists of a max-pooling layer, a deconvolution layer, and a ReLU (Linear rectification function) nonlinear activation function.
[0155] In a possible embodiment, before using the sample three-dimensional image, a preprocessing operation can be performed on the sample three-dimensional image. For example, taking a brain MRI image as an example, at least one preprocessing operation such as skull removal, registration, or window width and window level adjustment can be performed on the brain MRI image to obtain a preprocessed brain MRI image. In this way, irrelevant information in the image is eliminated, real useful information is recovered, the detectability of relevant information is enhanced, and the data is maximally simplified, thereby improving the reliability of object recognition.
[0156] In a possible embodiment, when the number of sample three-dimensional images is small, the number of sample three-dimensional images can be increased through data enhancement. Optionally, rotation, scaling transformation, flipping transformation, translation transformation, or the like can be used to perform data enhancement on the sample three-dimensional image to obtain a data-enhanced sample three-dimensional image. In this way, the overfitting phenomenon of the network can be reduced, and a network with stronger generalization ability can be obtained through transformation of the training data, thereby better adapting to application scenarios.
[0157] S102, the sample three-dimensional image is cropped into a plurality of sample three-dimensional voxel images.
[0158] Because the entire three-dimensional image data is too large, using the three-dimensional image directly for object recognition model training will occupy a large amount of video memory. Therefore, when training the object recognition model, the sample three-dimensional image can be cropped into a plurality of sample three-dimensional voxel images, and the object recognition model is trained in units of sample three-dimensional voxel images. In one example, the sample three-dimensional image is cropped into a plurality of sample three-dimensional voxel images, which can be randomly cropped into a plurality of sample three-dimensional voxel images according to random sizes; in other examples, the sample three-dimensional image can also be randomly cropped into a plurality of sample three-dimensional voxel images according to preset size, which can be customized according to actual conditions, for example, the size of a three-dimensional convolution kernel.
[0159] In a possible embodiment, the sample three-dimensional image is cropped into a plurality of sample three-dimensional voxel images with overlaps by using a sliding window method, and there is a pixel area overlap between adjacent sample three-dimensional voxel images. Cropping the sample three-dimensional image with overlaps can avoid missing information during the cropping process.
[0160] For the sample three-dimensional image being a positive sample, the sample three-dimensional image is optionally cropped into multiple sample three-dimensional voxel images, including: in the case that the sample three-dimensional image is a positive sample, cropping a region containing an object in the sample three-dimensional image into m sample three-dimensional voxel images, and cropping the remaining region into n-m sample three-dimensional voxel images, the positive sample is used to indicate that the sample three-dimensional image contains an object, and m≤n.
[0161] For the sample three-dimensional image being a negative sample, the sample three-dimensional image is optionally cropped into multiple sample three-dimensional voxel images, including: in the case that the sample three-dimensional image is a negative sample, cropping the sample three-dimensional image into q sample three-dimensional voxel images, the negative sample is used to indicate that the sample three-dimensional image does not contain an object, and q is less than n.
[0162] Specifically, during training, the sample three-dimensional image is randomly cropped into n sample three-dimensional voxel images, which are saved for training, so as to increase the number of training samples while ensuring the diversity of the sample three-dimensional voxel images. For the sample three-dimensional image containing an object, the region where the object is located needs to be cropped into m sample three-dimensional voxel images during random cropping, and the remaining region can be cropped into n-m sample three-dimensional voxel images. For the sample three-dimensional image not containing an object, it is directly randomly cropped into q sample three-dimensional voxel images, where the ratio of q to n is determined according to the positive and negative sample ratio, and q is generally less than n. For example, taking a brain MRI image as the sample three-dimensional image, during training, the brain MRI image is randomly cropped into n brain MRI blocks, for the brain MRI image containing a tumor, the region where the tumor is located is cropped into m brain MRI blocks, and the remaining region is cropped into n-m brain MRI blocks; for the brain MRI image not containing a tumor, the brain MRI image is cropped into q brain MRI blocks. The above-mentioned way of obtaining sample three-dimensional voxel images can ensure that the proportion of positive and negative samples in the obtained samples is a preset proportion, avoiding the situation that the training of the object recognition model is not in place due to too few positive samples and too many negative samples, and thus the accuracy of object recognition is low.
[0163] S103, inputting the sample three-dimensional image and the sample three-dimensional voxel image into the object recognition model to obtain a first object prediction result of the sample three-dimensional image and a second prediction result of the sample three-dimensional voxel image, the first object prediction result of the sample three-dimensional image being an output of the two-dimensional decoding network, and the second prediction result of the sample three-dimensional voxel image being an output of the three-dimensional decoding network.
[0164] For example, taking a brain MRI image as the sample three-dimensional image, the predicted mask label based on the brain MRI image and the real mask label Y gt determining the two-dimensional network loss L global ; the predicted mask label based on the brain MRI block and the real mask label Ypatch determining a three-dimensional network loss L patch .
[0165] S104, determining a two-dimensional network loss according to the first object prediction result and the object label, and determining a three-dimensional network loss according to the second object prediction result and the object label.
[0166] For example, the network loss can use Cross Entropy Loss, Focal Loss, DiceLoss, but not limited to.
[0167] S105, adjusting the parameters of the two-dimensional global feature extraction network, the two-dimensional encoding / decoding network and the three-dimensional encoding / decoding network according to the two-dimensional network loss and the three-dimensional network loss.
[0168] In a possible embodiment, adjusting the parameters of the two-dimensional global feature extraction network, the two-dimensional encoding / decoding network and the three-dimensional encoding / decoding network according to the two-dimensional network loss and the three-dimensional network loss comprises:
[0169] According to the two-dimensional network loss and the three-dimensional network loss, the total loss function is determined by the following formula:
[0170] L total = L global + μL patch
[0171] Wherein, L total is the total loss function, μ is a weight coefficient for balancing the loss, L global is the two-dimensional network loss, and L patch is the three-dimensional network loss.
[0172] Adjusting the parameters of the two-dimensional global feature extraction network, the two-dimensional encoding / decoding network and the three-dimensional encoding / decoding network based on the total loss function.
[0173] In another possible embodiment, adjusting the parameters of the two-dimensional global feature extraction network, the two-dimensional encoding / decoding network and the three-dimensional encoding / decoding network according to the two-dimensional network loss and the three-dimensional network loss comprises:
[0174] Adjusting the parameters of the two-dimensional encoding / decoding network according to the two-dimensional network loss; and adjusting the parameters of the three-dimensional encoding / decoding network according to the three-dimensional network loss.
[0175] S106, selecting other sample three-dimensional image data to continue training until a preset ending condition is met, and obtaining a trained object recognition model.
[0176] In a possible embodiment, the preset ending condition can be set according to actual conditions, for example, the loss (total loss) of the object recognition model converges, or reaches a preset training number, etc.
[0177] In one example, a preset proportion of the obtained sample three-dimensional images can be used for training of the object recognition model. The training process: for any sample three-dimensional image used for training data, the sample three-dimensional image is cropped to obtain a plurality of sample three-dimensional voxel images in the manner provided by the embodiments of the present application, the sample three-dimensional image and the plurality of sample three-dimensional voxel images are used to train the object recognition model, and the parameters of the determined object recognition model are stored.
[0178] The remaining sample data other than that used as training data is used as a test set, and the test process: for any sample three-dimensional image used for test data, the sample three-dimensional image is cropped with overlapping using the sliding window method block reasoning strategy with a preset step size. The trained object recognition model is used for reasoning, only the output of the three-dimensional encoding / decoding network is used as the recognition result, all the sample three-dimensional voxel images are spliced together, the edge effect is eliminated using the Gaussian weight code, and the object recognition result is obtained.
[0179] In the embodiments of the present application, sample three-dimensional images and object labels thereof, a plurality of sample three-dimensional voxel images, and an object recognition model to be trained are obtained. The object recognition model includes a two-dimensional global feature extraction network, a two-dimensional feature processing network, a three-dimensional encoding / decoding network, and a multi-dimensional feature fusion network. The sample three-dimensional image and the sample three-dimensional voxel image are input into the object recognition model to obtain a first object prediction result of the sample three-dimensional image and a second prediction result of the sample three-dimensional voxel image. The parameters of the object recognition model are adjusted according to the first object prediction result, the second prediction result, and the object label. In this way, a trained object recognition model is obtained. Then, the object recognition model is used to recognize objects in a three-dimensional image to be recognized. The use of the object recognition model for object recognition establishes a global connection of three-dimensional voxel images in the global by using the globality of two-dimensional global features, which can reduce the problem of global context information loss of three-dimensional voxel images, and further improve the accuracy of object recognition.
[0180] In one possible embodiment, the sample three-dimensional image and the sample three-dimensional voxel image are input into the object recognition model to obtain a first object prediction result of the sample three-dimensional image and a second prediction result of the sample three-dimensional voxel image, including:
[0181] S1031, extracting sample two-dimensional global features of the sample three-dimensional image using the two-dimensional global feature extraction network.
[0182] Since object recognition models trained solely on block-based methods (dividing a sample 3D image into sample 3D voxel maps) cannot capture the spatial context information of the sample 3D voxel maps, it is necessary to obtain the sample 2D global features of the sample 3D image and fuse them with the corresponding sample 3D voxel maps to supplement the spatial context information. The object recognition model includes a 2D global feature extraction network, which can be used to extract the sample 2D global features of the sample 3D image.
[0183] In one example, see Figure 2 The two-dimensional global feature extraction network includes three-dimensional convolutional block layers and pooling layers. This embodiment of the invention illustrates the process of extracting two-dimensional global features from a sample three-dimensional image using the two-dimensional global feature extraction network in the following two ways:
[0184] In the first approach, a sample 3D image is input into a 2D global feature extraction network. The 3D image is encoded and mapped to a high dimension using a 3D convolutional block layer to obtain the high-dimensional first sample 3D image encoding features. The high-dimensional first sample 3D image encoding features are then pooled using a pooling layer to obtain the sample 2D global features of the sample 3D image.
[0185] The process of convolutional encoding in a 3D convolutional block layer is the process of increasing the channel dimension of the image (mapping it to a higher dimension). The dimension of the 3D image encoding features of the high-dimensional sample can be customized according to the actual situation. For example, it can be 2 times, 4 times or 8 times the dimension of the sample 3D image, so as to obtain the 3D image encoding features of the first high-dimensional sample.
[0186] In one example, a sample 3D image can be represented as H×W×V×C, where H represents height, W represents width, V represents the number of slices, and C represents the number of channels. During pooling, the averaging / maximizing operation is performed on the slice number dimension. For example, H×W×V×C' will be pooled into H×W×1×C', where 1 can be directly omitted as H×W×C', thus obtaining the sample's 2D global features. The slices mentioned above can be sagittal, coronal, or axial. For example, axial can be chosen because objects to be identified in the axial direction are usually displayed more clearly. The specific slice direction can be customized according to the actual situation. The number of channels can generally be 1 for input, i.e., intensity value or grayscale value. During dimensionality increase, it can be in multiples of 2, increasing the dimension by 2 times each time. In one example, to ensure that the length and width of the 3D image remain unchanged during the dimensionality increase stage, a 1×1×1 convolution can be used directly. Of course, other sizes of convolution can also be selected according to the actual situation, but regardless of the transformation method used, same padding (i.e., output size = input size) must be guaranteed.
[0187] Understandably, a two-dimensional global feature extraction network may include one or more three-dimensional convolutional blocks, which consist of convolution, group normalization, and activation functions. Pooling layers are used to perform pooling operations on the input features; these operations can be average pooling or max pooling.
[0188] Specifically, pooling layers are used to perform pooling operations on the encoded features of the high-dimensional first sample 3D image to obtain the sample's 2D global features, including:
[0189] Pooling layers are used to process the encoded features of the high-dimensional first sample 3D image using aggregation functions, thereby obtaining the sample 2D global features of the sample 3D image;
[0190] The aggregation function can be: I global For the sample three-dimensional image, the sample two-dimensional global features, input For the sample 3D image, ConvBlock(G) input ) represents the encoded features of the high-dimensional first sample 3D image. This indicates that average pooling is performed along the axial direction. This example only uses the average pooling algorithm along the axial direction; max pooling, as well as pooling in the sagittal or coronal planes, can also be used, all of which are within the scope of this application.
[0191] Optionally, the aggregation function can also be used for global For the sample three-dimensional image, the sample two-dimensional global features, inpur For the sample 3D image, convBlock(G input ) represents the encoded features of the high-dimensional first sample 3D image. This indicates that max pooling is performed along the axial direction.
[0192] For example, taking a sample 3D image as a brain MRI map, the brain MRI map is input into a 2D global feature extraction network. The 3D convolutional block in the 2D global feature extraction network encodes the brain MRI map and maps it to a high dimension. In the high dimension, an aggregation function is used to obtain the 2D global features of the brain MRI map. The sample 2D global features of the sample 3D image are obtained and fused with the sample 3D voxel map corresponding to the sample 3D image to supplement spatial context information.
[0193] In the second approach, a sample 3D image is input into a 2D global feature extraction network. The 3D image is then encoded using a 3D convolutional block layer to obtain the encoded features of the first sample 3D image. A pooling layer is then used to perform a pooling operation on the encoded features of the first sample 3D image to obtain the sample 2D global features of the sample 3D image.
[0194] In the second method, the implementation of pooling operation on the encoded features of the first sample three-dimensional image using a pooling layer to obtain the sample two-dimensional global features of the sample three-dimensional image is the same as that in the first method, and will not be described in detail here.
[0195] In this embodiment of the application, compared with directly using aggregation functions, mapping the encoded features to a higher dimension can reduce the loss of feature information.
[0196] Optionally, the two-dimensional global feature extraction network also includes two-dimensional convolutional block layers. These layers can compress the dimensionality of the two-dimensional global features of the sample three-dimensional image to save GPU memory and reduce computation. For example, taking a brain MRI image as an example, the two-dimensional convolutional block layers can compress the dimensionality of the two-dimensional global features of the brain MRI image to reduce the amount of data in the two-dimensional global features.
[0197] S1032 uses a two-dimensional coding network to encode the two-dimensional global features of the samples, thereby obtaining two-dimensional image coding features of multiple samples.
[0198] A two-dimensional coding network consists of multiple coding layers used to encode the two-dimensional global features of a sample. Each coding layer can output the encoded features of a sample's two-dimensional image. The number of coding layers in a two-dimensional coding network can be customized according to actual needs. See the example below. Figure 3 The object recognition model includes a two-dimensional coding network. Taking the three-dimensional image of the sample as a brain MRI image as an example, the two-dimensional coding network is used to encode the two-dimensional global features of the brain MRI image to obtain multiple two-dimensional feature maps.
[0199] Optionally, the two-dimensional coding network includes k two-dimensional coding layers. For example, k can be 4, that is, the two-dimensional coding network includes a first two-dimensional coding layer, a second two-dimensional coding layer, a third two-dimensional coding layer and a fourth two-dimensional coding layer, and uses k two-dimensional coding layers to encode the two-dimensional global features of the sample.
[0200] For example, the first two-dimensional coding layer is used to encode the two-dimensional global features of the sample, and the a-th coding layer is used to encode the sample image features output by the (a-1)-th two-dimensional coding layers, resulting in multiple sample two-dimensional image coding features, a∈[2,K]. Multi-layer coding is used to encode the sample two-dimensional global features so that the first object prediction result of the sample three-dimensional image can be obtained based on the multiple sample two-dimensional image coding features. Furthermore, the object recognition model can be trained based on the first object prediction result of the sample three-dimensional image, thereby improving the recognition accuracy of the object recognition model.
[0201] S1033 uses a two-dimensional decoding network to decode the encoded features of each sample's two-dimensional image to obtain the first object prediction result of the sample's three-dimensional image.
[0202] The two-dimensional decoding network includes multiple decoding layers, which are used to decode the encoded features of the sample two-dimensional image. The number of decoding layers in the two-dimensional decoding network can be customized according to actual needs and can be the same as the number of encoding layers in the two-dimensional encoding network. In one example, the two-dimensional decoding network may end with a convolutional layer for classification, thereby obtaining a first object prediction result for the sample three-dimensional image. Furthermore, in other embodiments, a classifier can also be connected after the two-dimensional decoding network to classify the features output by the two-dimensional decoding network to obtain a first object prediction result for the sample three-dimensional image.
[0203] In one example, see Figure 3 The object recognition model includes a two-dimensional decoding network and a two-dimensional convolutional layer. The two-dimensional decoding network decodes the encoded features of each sample's two-dimensional image to obtain the first object prediction result for the sample's three-dimensional image. This can be achieved by: using the two-dimensional convolutional layer to perform a convolution operation on the sample image features output by the two-dimensional decoding network to obtain an initial prediction result; then, a tiling operation is performed on the initial prediction result, i.e., reshaping it in three-dimensional space, to obtain the first object prediction result for the sample's three-dimensional image. In one example, because the two-dimensional encoding / decoding network uses two-dimensional convolutional kernels, the initial prediction result output by the two-dimensional convolutional layer can be represented as (Batch, C, W, H), where Batch represents the number of images used in one iteration, C represents the number of output channels, W represents the width, and H represents the height. The initial prediction result is then reshaped into (Batch, N_class, V, W, H), i.e., tiling, to obtain the first object prediction result for the sample's three-dimensional image, where N_class represents the number of channels in each slice, and V represents the number of slices. If the sample's three-dimensional image is a brain MRI image, the first object prediction result is the predicted mask label of the brain MRI image.
[0204] Optionally, the two-dimensional decoding network includes k two-dimensional decoding layers. For example, k can be 4, meaning the two-dimensional decoding network includes a first two-dimensional decoding layer, a second two-dimensional decoding layer, a third two-dimensional decoding layer, and a fourth two-dimensional decoding layer. The two-dimensional decoding network decodes the encoded features of each sample two-dimensional image to obtain the first object prediction result of the sample three-dimensional image. This can be achieved by using the t-th two-dimensional decoding layer to decode the sample image features output by the (k+1-t)-th two-dimensional encoding layer and the sample image features output by the (t-1)-th two-dimensional decoding layer, and then using a two-dimensional convolutional layer to perform convolution and tiling operations on the sample image features output by the k-th two-dimensional decoding layer, outputting the first object prediction result of the sample three-dimensional image. Obtaining the first object prediction result of the sample three-dimensional image allows for subsequent training of the object recognition model based on this first object prediction result, improving the recognition accuracy of the object recognition model.
[0205] S1034, a three-dimensional coding network is used to encode the features of the sample three-dimensional voxel map to obtain multiple second sample three-dimensional image coding features.
[0206] The 3D encoding network comprises multiple encoding layers for feature encoding of sample 3D voxel images. Each encoding layer can output an encoded feature of a second sample 3D image. The number of encoding layers in the 3D encoding network can be customized according to actual needs. See the example below. Figure 3 The object recognition model includes a three-dimensional coding network. Taking the brain MRI image as an example, the three-dimensional coding network is used to encode the features of the brain MRI image blocks to obtain multiple three-dimensional feature maps.
[0207] Optionally, the three-dimensional encoding network includes k three-dimensional encoding layers. For example, k can be 4, that is, the three-dimensional encoding network includes a first three-dimensional encoding layer, a second three-dimensional encoding layer, a third three-dimensional encoding layer, and a fourth three-dimensional encoding layer. The k three-dimensional encoding layers are used to encode the features of the sample three-dimensional voxel map.
[0208] In one example, the first two-dimensional coding layer is used to encode the features of the sample three-dimensional voxel image, and the a-th coding layer is used to encode the sample image features output by the (a-1)-th three-dimensional coding layers, resulting in multiple second sample three-dimensional image coding features, a∈[2,K]. Multi-layer coding is used to encode the features of the sample three-dimensional voxel image so that the second object prediction result of the sample three-dimensional voxel image can be obtained based on the multiple second sample three-dimensional image coding features. Furthermore, the object recognition model can be trained based on the second object prediction result of the sample three-dimensional voxel image, thereby improving the recognition accuracy of the object recognition model.
[0209] S1035 utilizes a multi-dimensional feature fusion network to fuse the two-dimensional image coding features of the sample with the three-dimensional image coding features of the second sample, resulting in multiple sample fusion features.
[0210] Here, common fusion methods found in relevant feature fusion techniques can be employed to fuse the encoded features of the sample 2D image with the encoded features of the second sample 3D image. For example, concatenation, product, and summation can be used for feature fusion. Specifically, the `concat` function can be used to concatenate (fuse) the encoded features of the sample 2D image with the encoded features of the second sample 3D image, or a weighted fusion method can be used to fuse the encoded features of the sample 2D image with the encoded features of the second sample 3D image, thus obtaining the fused features. For example... Figure 4 As shown, a multi-dimensional feature fusion network is used to fuse the two-dimensional image coding features of the sample obtained through the above steps with the three-dimensional image coding features of the second sample to obtain multiple sample fusion features.
[0211] In one example, the two-dimensional image coding features output by coding layers of the same number of layers can be fused with the three-dimensional image coding features of the second sample. In one possible implementation, a multi-dimensional feature fusion network is used to fuse the two-dimensional image coding features of the sample with the three-dimensional image coding features of the second sample to obtain multiple sample fusion features, including: using a multi-dimensional feature fusion network to fuse the two-dimensional image coding features of the sample output by the l-th two-dimensional coding layer with the three-dimensional image coding features of the second sample output by the l-th three-dimensional coding layer to obtain the l-th sample fusion feature, where l∈[1,K-1].
[0212] In both the 2D and 3D coding layers, the 2D image coding features output from the same number of coding layers are fused with the 3D image coding features of the second sample to obtain the fused feature. The features output from the last coding layer can be directly used as the input to the first decoding layer without further fusion.
[0213] The 2D image encoding features of the sample are global features of the 3D image of the sample; while the 3D image encoding features of the second sample are only features of one of the 3D voxel images among multiple sample 3D voxel images, that is, the 3D image encoding features of the second sample are partial features of the 3D image of the sample; therefore, when performing feature fusion, the 3D image encoding features of the second sample can be aligned with the 2D image encoding features of the sample. Optionally, a multi-dimensional feature fusion network is used to fuse the 2D image encoding features of the sample output from the l-th 2D encoding layer with the 3D image encoding features of the second sample output from the l-th 3D encoding layer to obtain the l-th sample fusion feature, including:
[0214] Step 1: Use a multi-dimensional feature fusion network to extract features from the two-dimensional image encoding features of the output of the l-th two-dimensional encoding layer to obtain the transformation features of the l-th sample two-dimensional image. The transformation features of the l-th sample two-dimensional image have the same number of channels as the encoding features of the l-th second sample three-dimensional image.
[0215] The two-dimensional convolutional blocks in the multidimensional feature fusion network are used to further extract the coded features of the sample two-dimensional image to obtain the sample two-dimensional image transformation features; so that the number of channels of the l-th sample two-dimensional image transformation feature is the same as that of the l-th second sample three-dimensional image coding feature.
[0216] Step 2: Based on the position information of the current input sample 3D voxel map in the sample 3D image, determine the feature region corresponding to the current input sample 3D voxel map in the transformation feature of the l-th sample 2D image, and obtain the region feature of the l-th sample 2D image.
[0217] The sample 3D voxel images are input into the 3D encoding / decoding network. Based on the position information of the currently input sample 3D voxel image in the sample 3D image, the feature region corresponding to the currently input sample 3D voxel image is found in the 2D image transformation features, and the sample 2D image region features are obtained.
[0218] Step 3: Copy and adjust the size of the two-dimensional image region feature of the l-th sample to the size of the three-dimensional image encoding feature of the l-th sample, and obtain the two-dimensional image region feature of the l-th sample after size adjustment;
[0219] Step four involves fusing the 2D image region features of the l-th sample with the 3D image encoding features of the l-th second sample to obtain the fused feature of the l-th sample. This allows the 3D decoding network to utilize both global and local information simultaneously during decoding.
[0220] S1036, the three-dimensional decoding network is used to decode the fusion features of each sample and the three-dimensional image encoding features of at least one second sample to obtain the second object prediction result of the three-dimensional voxel map of the sample.
[0221] See Figure 3 The object recognition model includes a 3D decoding network and a 3D decoding layer. The 3D decoding network decodes the fused features of each sample and the encoded features of at least one second sample 3D image to obtain the second object prediction result of the sample 3D voxel map. This includes using a 3D convolutional layer to perform a convolution operation on the sample image features output by the 3D decoding network to obtain the second object prediction result of the sample 3D voxel map. Taking a brain MRI image as an example, the second object prediction result is the predicted mask label of the sample 3D voxel map.
[0222] In one example, the final step of the 3D decoding network may include a 3D convolutional layer for classification, thereby obtaining a second object prediction result for the sample 3D voxel image. Optionally, the 3D decoding network includes k 3D decoding layers, for example, k can be 4, that is, the 3D decoding network includes a first 3D decoding layer, a second 3D decoding layer, a third 3D decoding layer, and a fourth 3D decoding layer. The 3D decoding network is used to decode the fusion features of each sample and at least one second sample 3D image encoding feature to obtain the second object prediction result for the sample 3D voxel image, including: decoding the sample image features output by the Kth 3D encoding layer using the first 3D decoding layer; decoding the sample image features output by the (t-1)th 3D decoding layer and the fusion features of the (K+1-t)th sample using the tth 3D decoding layer, where t∈[2,K]; and performing a convolution operation on the sample image features output by the Kth 3D decoding layer using a 3D convolutional layer to obtain the second object prediction result for the sample 3D voxel image, which can be the prediction mask label of the sample 3D voxel image.
[0223] In other examples, a classifier can be added after the 3D decoding network to classify the data output by the network, thereby obtaining a second object prediction result from the sample 3D voxel image. This second object prediction result is then used to train the object recognition model, improving its accuracy.
[0224] In this embodiment, a two-dimensional global feature extraction network is used to extract sample two-dimensional global features of a three-dimensional image. These sample two-dimensional global features are then input into a two-dimensional encoding / decoding network. The three-dimensional voxel map corresponding to the three-dimensional image is input into a three-dimensional encoding / decoding network. A multi-dimensional feature fusion network is used to fuse the features output by the two-dimensional / three-dimensional encoding network. The fused features are then decoded and classified to train the object recognition model. The trained object recognition model is then used to perform object recognition on the three-dimensional image to be recognized, thereby improving the accuracy of object recognition.
[0225] In one possible implementation, based on the position information of the currently input sample 3D voxel map in the sample 3D image, the feature region corresponding to the currently input sample 3D voxel map in the transformed feature of the l-th sample 2D image is determined, and the feature of the l-th sample 2D image region is obtained, including:
[0226] Obtain the position information (i, j, k) of the current input sample 3D voxel image in the sample 3D image; determine the feature region (i, j, k) corresponding to the current input sample 3D voxel image in the transformed feature of the l-th sample 2D image according to the following formula. l j l ):
[0227]
[0228]
[0229] The size of the convolution kernel in the two-dimensional coding layer is (K V K W K H In the two-dimensional coding layer, the step size of pooling calculation is S.
[0230] For example, the multidimensional feature fusion network includes two-dimensional convolutional block layers. For the l-th encoding layer, the ith two-dimensional encoding layer outputs a two-dimensional feature map F. g l The I-th 3D coding layer outputs a 3D feature map M. p l The dimensions are B×C. g ×W×H and B×C p ×V p ×M p ×H p B represents the batch size, C g and C p Representing two-dimensional feature maps F respectively g l and 3D feature map M p l The number of channels, W and H represent F g l Width and height, V p M p H p Representing M p l The depth, width, and height of the feature map F are calculated using two-dimensional convolutional block layers in a multidimensional feature fusion network. g l Perform dimensionality reduction on convolutional blocks to make C g =C g The resulting size is B×C p Two-dimensional feature map F of ×W×H g l Assuming the 3D indices of the current sample's 3D voxel image in the sample's 3D image are i, j, k, after encoding by the i-th encoding layer, the 3D indices of the current sample's 3D voxel image in the sample's 3D image are:
[0231]
[0232]
[0233]
[0234] Among them, (K) VK W K H ) represents the kernel size, and S represents the stride of the pooling layer.
[0235] Using i l j l From the two-dimensional feature map F g l Mapping and cropping the 3D feature map M p l The corresponding dimension is B×C p ×W p ×H p Two-dimensional region features F p l The two-dimensional region feature F p l Using formula (1), cut, copy, and stitch them together to form a three-dimensional feature map M. p l Feature maps of the same size, i.e., size B×C p ×V p ×W p ×H p Three-dimensional and two-dimensional feature maps F p l The three-dimensional two-dimensional feature map F will be obtained using formula (2). p l and 3D feature map M p l We perform a weighted summation to obtain the i-th fused feature, with a size of B×C. p ×V p ×W p ×H p The fusion feature map.
[0236]
[0237] Among them, F g,3D l For V p Two-dimensional region features F p l The feature map formed by stacking them together, F p l For the three-dimensional feature map M p l Two-dimensional region features at the corresponding location.
[0238] F l =ωF g,3D l +(1-ω)M p l (2)
[0239] Among them, Fl Let ω be the i-th fusion feature, and F be the weight coefficient. g,3D l For V p Two-dimensional region features F p l The feature map formed by stacking them together, M p l Output a 3D feature map for the I-th 3D coding layer.
[0240] In other embodiments, the I-th fusion feature can also be obtained by splicing:
[0241]
[0242] Among them, F l For the i-th fusion feature, F g,3D l For V p Two-dimensional region features F p l The feature map formed by stacking them together, M p l Output a 3D feature map for the I-th 3D coding layer. Indicates splicing.
[0243] Two-dimensional feature map F g l The receptive field increases, allowing each pixel to capture information from more pixels, thus increasing the F... p l Although the features are cropped from the entire image, they still contain rich global information. Furthermore, the fused features are fed into the I-th 3D decoding layer via skip connections, allowing both global and local information to be utilized during decoding.
[0244] This application also provides an object recognition method, see [link to relevant documentation]. Figure 5 The method includes:
[0245] S501, acquire the 3D image to be identified and the pre-trained object recognition model, wherein the object recognition model includes a 2D global feature extraction network, a 2D feature processing network, a 3D encoding / decoding network and a multi-dimensional feature fusion network;
[0246] The 3D image to be identified is a 3D image, which can be a brain MRI image or other 3D images. The object recognition model is trained using the scheme provided in the previous embodiment, and will not be described in detail here.
[0247] S502, crop the 3D image to be identified into multiple 3D voxel images to be identified;
[0248] For example, the brain MRI image to be identified can be cropped into multiple brain MRI patches, for example, using... Figure 3 The object recognition model in the model performs 3D voxel image cropping on the 3D image to be recognized, resulting in multiple 3D voxel images to be recognized. The specific cropping process can be found in the process of cropping the sample 3D image into multiple sample 3D voxel images when training the object recognition model, and will not be elaborated here.
[0249] S503: Input the 3D image to be identified and multiple 3D voxel images to be identified into the object recognition model to obtain the object recognition result.
[0250] For example, by inputting a brain MRI image to be identified into the object recognition model, tumors in the brain MRI image can be identified.
[0251] The trained object recognition model is used to identify objects in the 3D image to be recognized. In this way, the global nature of the 2D global features is used to establish a global connection between the 3D voxel map, which can reduce the problem of loss of global contextual information in the 3D voxel map and thus improve the accuracy of object recognition.
[0252] In one possible implementation, the two-dimensional feature processing network includes a two-dimensional encoding / decoding network; or, the two-dimensional feature processing network includes a two-dimensional encoding network but does not include a two-dimensional decoding network.
[0253] In one possible embodiment, the second object prediction results of each three-dimensional voxel image to be identified are spliced together, and Gaussian weight codes are used to weight overlapping pixel regions to obtain the object recognition result.
[0254] In the process of using the object recognition model, the output of the 3D encoding / decoding network branch can be used as the final object recognition result. Therefore, the 2D decoding network does not play a role in actual use. In order to reduce the amount of data in the object recognition model and reduce the consumption of computing resources, the 2D decoding network can be removed.
[0255] This application also provides an object recognition device, see [link to relevant documentation]. Figure 6 The device includes:
[0256] The first acquisition module 601 is used to acquire the three-dimensional image to be identified and to acquire a pre-trained object recognition model, wherein the object recognition model includes a two-dimensional global feature extraction network, a two-dimensional feature processing network, a three-dimensional encoding / decoding network and a multi-dimensional feature fusion network.
[0257] The first cropping module 602 is used to crop the three-dimensional image to be identified into multiple three-dimensional voxel images to be identified.
[0258] The recognition module 603 is used to input the three-dimensional image to be recognized and multiple three-dimensional voxel images to be recognized into the object recognition model to obtain the object recognition result.
[0259] Optionally, the two-dimensional feature processing network includes a two-dimensional encoding / decoding network;
[0260] Alternatively, a two-dimensional feature processing network may include a two-dimensional encoding network but exclude a two-dimensional decoding network.
[0261] Optional, the recognition module is specifically used for:
[0262] A two-dimensional global feature extraction network is used to extract the target two-dimensional global features of the three-dimensional image to be identified.
[0263] A two-dimensional feature processing network is used to encode the two-dimensional global features of the target, resulting in multiple target two-dimensional image encoded features.
[0264] A three-dimensional coding network is used to encode the features of the three-dimensional voxel image to be identified, resulting in multiple encoded features of the three-dimensional image of the second target.
[0265] A multi-dimensional feature fusion network is used to fuse the two-dimensional image coding features of the target with the three-dimensional image coding features of the second target to obtain multiple target fusion features;
[0266] The fusion features of each target and the 3D image coding features of at least one second target are decoded using a 3D decoding network to obtain the second object prediction result of the 3D voxel image to be identified. The object recognition result includes the second object prediction result of each 3D voxel image to be identified.
[0267] Optionally, the first cropping module is used to crop the three-dimensional image to be identified into multiple overlapping three-dimensional voxel images to be identified using the sliding window method, wherein there is pixel region overlap between adjacent three-dimensional voxel images to be identified.
[0268] The device also includes a module that performs the following operations: stitching together the second object prediction results of each three-dimensional voxel map and using Gaussian weight codes to weight overlapping pixel regions to obtain object recognition results.
[0269] Optional, the recognition module is specifically used for:
[0270] The three-dimensional image to be identified is input into the two-dimensional global feature extraction network. The three-dimensional convolutional block layer encodes the three-dimensional image to be identified and maps it to a high dimension to obtain the high-dimensional first target three-dimensional image encoding features.
[0271] Pooling layers are used to perform pooling operations on the encoded features of the high-dimensional first target 3D image to obtain the target 2D global features of the 3D image to be identified.
[0272] Optional, the recognition module is specifically used for:
[0273] Pooling layers are used to process the encoded features of the high-dimensional first target 3D image using aggregation functions to obtain the target 2D global features of the 3D image to be identified;
[0274] Wherein, the aggregation function is I global For the target two-dimensional global features of the three-dimensional image to be identified, input For the 3D image to be identified, ConvBlock(G) input ) represents the encoded features of the high-dimensional first target 3D image. This indicates that average pooling is performed along the axial direction.
[0275] Optionally, the two-dimensional global feature extraction network further includes two-dimensional convolutional block layers, and the apparatus further includes modules for performing the following operations:
[0276] Two-dimensional convolutional block layers are used to reduce the dimensionality of the target two-dimensional global features in the three-dimensional image to be identified.
[0277] Optionally, the two-dimensional coding network includes k two-dimensional coding layers, and the three-dimensional coding network includes K three-dimensional coding layers. The recognition module is specifically used for:
[0278] The target two-dimensional image coding features output by the l-th two-dimensional coding layer and the target three-dimensional image coding features output by the l-th three-dimensional coding layer are fused using a multi-dimensional feature fusion network to obtain the l-th target fusion feature, where l∈[1,K-1].
[0279] Optionally, the 3D decoding network includes K 3D decoding layers and 3D convolutional layers, and a recognition module, specifically used for:
[0280] The target image features output by the Kth 3D encoding layer are decoded using the first 3D decoding layer;
[0281] The target image features output by the (t-1)th 3D decoding layer and the fused features of the (K+1-t)th target are decoded using the t-th 3D decoding layer, where t∈[2,K];
[0282] The target image features output by the Kth 3D decoding layer are convolved using a 3D convolutional layer to obtain the recognition result corresponding to the 3D voxel image to be recognized.
[0283] Optional, the recognition module is specifically used for:
[0284] The target two-dimensional image encoding features output by the l-th two-dimensional encoding layer are extracted using a multi-dimensional feature fusion network to obtain the l-th target two-dimensional image transformation feature. The l-th target two-dimensional image transformation feature has the same number of channels as the l-th target three-dimensional image encoding feature.
[0285] Based on the position information of the currently input three-dimensional voxel image to be identified in the three-dimensional image to be identified, the feature region corresponding to the currently input three-dimensional voxel image to be identified in the first target two-dimensional image transformation feature is determined, and the feature of the first target two-dimensional image region is obtained.
[0286] The size of the two-dimensional image region feature of the l-th target is adjusted to the size of the three-dimensional image encoding feature of the l-th second target, thus obtaining the size-adjusted two-dimensional image region feature of the l-th target;
[0287] The 2D image region features of the l-th target after size adjustment are fused with the 3D image coding features of the l-th second target to obtain the fused feature of the l-th target.
[0288] Optional, the recognition module is specifically used for:
[0289] Obtain the position information (i, j, k) of the currently input 3D voxel map to be identified in the 3D image to be identified;
[0290] The feature region (i) corresponding to the l-th target 2D image transformation feature of the currently input 3D voxel image to be identified is determined according to the following formula. l j l ):
[0291]
[0292]
[0293] The size of the convolution kernel in the two-dimensional coding layer is (K V ,K W ,K H In the two-dimensional coding layer, the step size of pooling calculation is S.
[0294] This application also provides an object recognition device, see [link to relevant documentation]. Figure 7 The device includes:
[0295] The second acquisition module 701 is used to acquire a sample 3D image, object annotations of the sample 3D image, and an object recognition model to be trained. The object recognition model includes a 2D global feature extraction network, a 2D feature processing network, a 3D encoding / decoding network, and a multi-dimensional feature fusion network. The 2D feature processing network includes a 2D encoding / decoding network.
[0296] The second cropping module 702 is used to crop the sample 3D image into multiple sample 3D voxel images;
[0297] The first determining module 703 is used to input the sample three-dimensional image and the sample three-dimensional voxel image into the object recognition model to obtain the first object prediction result of the sample three-dimensional image and the second prediction result of the sample three-dimensional voxel image. The first object prediction result of the sample three-dimensional image is the output of the two-dimensional decoding network, and the second prediction result of the sample three-dimensional voxel image is the output of the three-dimensional decoding network.
[0298] The second determination module 704 is used to determine the two-dimensional network loss based on the first object prediction result and object annotation; and to determine the three-dimensional network loss based on the second object prediction result and object annotation.
[0299] The adjustment module 705 is used to adjust the parameters of the two-dimensional global feature extraction network, the two-dimensional encoding / decoding network, and the three-dimensional encoding / decoding network according to the two-dimensional network loss and the three-dimensional network loss.
[0300] Training module 706 is used to select other sample 3D image data to continue training until the preset termination condition is met, so as to obtain the trained object recognition model.
[0301] Optional, a second cropping module, specifically used for:
[0302] When the sample 3D image is a positive sample, the region containing the object in the sample 3D image is cropped into m sample 3D voxel images, and the remaining region is cropped into nm sample 3D voxel images. Positive samples are used to indicate that the sample 3D image contains an object, where m≤n.
[0303] When the sample 3D image is a negative sample, the sample 3D image is cropped into q sample 3D voxel images. The negative samples are used to indicate that the sample 3D image does not contain objects. <n。
[0304] Optional, the first determining module is specifically used for:
[0305] Two-dimensional global feature extraction network is used to extract two-dimensional global features of sample three-dimensional images;
[0306] A two-dimensional coding network is used to encode the two-dimensional global features of the samples to obtain the two-dimensional image coding features of multiple samples.
[0307] The two-dimensional decoding network is used to decode the encoded features of each sample's two-dimensional image to obtain the first object prediction result of the sample's three-dimensional image;
[0308] A three-dimensional coding network is used to encode the features of the three-dimensional voxel map of the sample to obtain multiple coded features of the three-dimensional image of the second sample.
[0309] A multidimensional feature fusion network is used to fuse the two-dimensional image coding features of the sample with the three-dimensional image coding features of the second sample to obtain multiple sample fusion features;
[0310] The fusion features of each sample and the 3D image encoding features of at least one second sample are decoded using a 3D decoding network to obtain the second object prediction result of the sample 3D voxel map.
[0311] Optionally, the two-dimensional coding network includes K two-dimensional coding layers, and the three-dimensional coding network includes K three-dimensional coding layers. The first determining module is specifically used for:
[0312] The multidimensional feature fusion network is used to fuse the two-dimensional image coding features of the sample output by the l-th two-dimensional coding layer with the three-dimensional image coding features of the second sample output by the l-th three-dimensional coding layer to obtain the l-th sample fusion feature, where l∈[1,K-1].
[0313] Optionally, the 3D decoding network includes K 3D decoding layers and 3D convolutional layers. The first determining module is specifically used for:
[0314] The first 3D decoding layer is used to decode the sample image features output by the Kth 3D encoding layer;
[0315] The sample image features output by the (t-1)th 3D decoding layer and the fusion features of the (K+1-t)th sample are decoded using the t-th 3D decoding layer, where t∈[2,K];
[0316] The second object prediction result of the sample 3D voxel map is obtained by performing a convolution operation on the sample image features output by the Kth 3D decoding layer using a 3D convolutional layer.
[0317] This application also provides an electronic device, such as... Figure 8 As shown, it includes:
[0318] Memory 802 is used to store computer programs;
[0319] When processor 801 executes a program stored in memory 802, it performs the following steps:
[0320] The process involves acquiring a 3D image to be identified and a pre-trained object recognition model, which includes a 2D global feature extraction network, a 2D feature processing network, a 3D encoding / decoding network, and a multi-dimensional feature fusion network.
[0321] The 3D image to be identified is cropped into multiple 3D voxel images to be identified;
[0322] The object recognition model is input into the 3D image to be recognized and multiple 3D voxel images to be recognized to obtain the object recognition result.
[0323] Furthermore, the aforementioned electronic device may also include a communication bus and / or a communication interface, with the processor 801, communication interface, and memory 802 communicating with each other via the communication bus.
[0324] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0325] The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0326] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0327] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0328] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the above-described object recognition and object recognition model training methods.
[0329] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the object recognition and object recognition model training methods in the above embodiments.
[0330] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0331] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0332] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, electronic devices, and storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0333] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.
Claims
1. A method of object recognition, characterized by, The method comprises: acquiring a three-dimensional image to be recognized, and acquiring a pre-trained object recognition model, wherein the object recognition model comprises a two-dimensional global feature extraction network, a two-dimensional feature processing network, a three-dimensional encoding / decoding network, and a multi-dimensional feature fusion network; cropping the three-dimensional image to be recognized into a plurality of three-dimensional voxel images to be recognized; inputting the three-dimensional image to be recognized and the plurality of three-dimensional voxel images to be recognized into the object recognition model to obtain an object recognition result, comprising: extracting target two-dimensional global features of the three-dimensional image to be recognized by using the two-dimensional global feature extraction network; performing feature coding on the target two-dimensional global features by using the two-dimensional feature processing network to obtain a plurality of target two-dimensional image coding features; performing feature coding on the three-dimensional voxel images to be recognized by using the three-dimensional encoding network to obtain a plurality of second target three-dimensional image coding features; fusing the target two-dimensional image coding features and the second target three-dimensional image coding features by using the multi-dimensional feature fusion network to obtain a plurality of target fusion features; decoding each target fusion feature and at least one second target three-dimensional image coding feature by using the three-dimensional decoding network to obtain a second object prediction result of the three-dimensional voxel image to be recognized, wherein the object recognition result comprises the second object prediction result of each three-dimensional voxel image to be recognized.
2. The method of claim 1, wherein, The two-dimensional feature processing network comprises a two-dimensional encoding / decoding network. Or, the two-dimensional feature processing network comprises a two-dimensional encoding network and does not comprise a two-dimensional decoding network.
3. The method of claim 1, wherein, The cropping of the three-dimensional image to be recognized into a plurality of three-dimensional voxel images to be recognized comprises: cropping the three-dimensional image to be recognized into a plurality of three-dimensional voxel images to be recognized with overlapping by using a sliding window method, wherein there is a pixel area overlap between adjacent three-dimensional voxel images to be recognized. The method further comprises: splicing the second object prediction results of each three-dimensional voxel image to be recognized, and weighting the pixel areas with overlapping by using a Gaussian weight code to obtain an object recognition result.
4. The method of claim 1, wherein, The two-dimensional global feature extraction network comprises a three-dimensional convolution block layer and a pooling layer, and the extraction of target two-dimensional global features of the three-dimensional image to be recognized by using the two-dimensional global feature extraction network comprises: inputting the three-dimensional image to be recognized into the two-dimensional global feature extraction network, encoding and mapping the three-dimensional image to be recognized to a high dimension by using the three-dimensional convolution block layer to obtain a high-dimensional first target three-dimensional image coding feature; performing a pooling operation on the high-dimensional first target three-dimensional image coding feature by using the pooling layer to obtain target two-dimensional global features of the three-dimensional image to be recognized.
5. The method of claim 4, wherein, The pooling operation on the high-dimensional first target three-dimensional image coding feature by using the pooling layer to obtain target two-dimensional global features of the three-dimensional image to be recognized comprises: processing the high-dimensional first target three-dimensional image coding feature by using the pooling layer using an aggregation function to obtain target two-dimensional global features of the three-dimensional image to be recognized. wherein the aggregation function is , is a target two-dimensional global feature of the three-dimensional image to be recognized, is the three-dimensional image to be recognized, is the high-dimensional first target three-dimensional image coding feature, indicates that average pooling is performed in the axial direction.
6. The method of claim 4, wherein, The two-dimensional global feature extraction network further comprises a two-dimensional convolution block layer, and after the high-dimensional first target three-dimensional image coding feature is pooled by the pooling layer to obtain the target two-dimensional global feature of the three-dimensional image to be recognized, the method further comprises: The two-dimensional convolution block layer is used to reduce the dimension of the target two-dimensional global feature of the three-dimensional image to be recognized.
7. The method of claim 2, wherein, The two-dimensional coding network comprises k two-dimensional coding layers, and the three-dimensional coding network comprises K three-dimensional coding layers. The target two-dimensional image coding feature and the second target three-dimensional image coding feature are fused by the multi-dimensional feature fusion network to obtain a plurality of target fusion features, comprising: The target two-dimensional image coding feature output by the lth two-dimensional coding layer and the second target three-dimensional image coding feature output by the lth three-dimensional coding layer are fused by the multi-dimensional feature fusion network to obtain the lth target fusion feature, wherein l∈[1,K-1].
8. The method of claim 7, wherein, The three-dimensional decoding network comprises K three-dimensional decoding layers and a three-dimensional convolution layer. The target fusion features and at least one second target three-dimensional image coding feature are decoded by the three-dimensional decoding network to obtain the second object prediction result of the three-dimensional voxel image to be recognized, comprising: The target image feature output by the Kth three-dimensional coding layer is decoded by the first three-dimensional decoding layer; The target image feature output by the t-1th three-dimensional decoding layer and the K+1-tth target fusion feature are decoded by the tth three-dimensional decoding layer, wherein t∈[2,K]; The target image feature output by the Kth three-dimensional decoding layer is convolved by the three-dimensional convolution layer to obtain the recognition result corresponding to the three-dimensional voxel image to be recognized.
9. The method of claim 7, wherein, The target two-dimensional image coding feature output by the lth two-dimensional coding layer and the second target three-dimensional image coding feature output by the lth three-dimensional coding layer are fused by the multi-dimensional feature fusion network to obtain the lth target fusion feature, comprising: The target two-dimensional image coding feature output by the lth two-dimensional coding layer is extracted by the multi-dimensional feature fusion network to obtain the lth target two-dimensional image conversion feature, wherein the channel number of the lth target two-dimensional image conversion feature is the same as that of the lth second target three-dimensional image coding feature; According to the position information of the current input three-dimensional voxel image to be recognized in the three-dimensional image to be recognized, the feature region corresponding to the current input three-dimensional voxel image to be recognized in the lth target two-dimensional image conversion feature is determined to obtain the lth target two-dimensional image region feature; The size of the lth target two-dimensional image region feature is adjusted to the size of the lth second target three-dimensional image coding feature to obtain the size-adjusted lth target two-dimensional image region feature; The size-adjusted lth target two-dimensional image region feature is fused with the lth second target three-dimensional image coding feature to obtain the lth target fusion feature.
10. The method of claim 9, wherein, According to the position information of the current input three-dimensional voxel image to be recognized in the three-dimensional image to be recognized, the feature region corresponding to the current input three-dimensional voxel image to be recognized in the lth target two-dimensional image conversion feature is determined to obtain the lth target two-dimensional image region feature, comprising: Obtaining position information (i, j, k) of a current input three-dimensional voxel image in a three-dimensional image to be recognized; The feature region corresponding to the current input three-dimensional voxel image to be recognized in the lth target two-dimensional image conversion feature is determined according to the following formula: ): ; ; In the two-dimensional coding layer, the size of the convolution kernel is , and the step length of the pooling calculation in the two-dimensional coding layer is S. 11.A method for training an object recognition model, the method comprising: The method comprises: Obtaining a sample three-dimensional image, object annotation of the sample three-dimensional image and an object recognition model to be trained, wherein the object recognition model comprises a two-dimensional global feature extraction network, a two-dimensional feature processing network, a three-dimensional encoding / decoding network and a multi-dimensional feature fusion network, the two-dimensional feature processing network comprises a two-dimensional encoding / decoding network; Cutting the sample three-dimensional image into a plurality of sample three-dimensional voxel images; Extracting sample two-dimensional global features of the sample three-dimensional image by using the two-dimensional global feature extraction network; Encoding the sample two-dimensional global features by using a two-dimensional encoding network to obtain a plurality of sample two-dimensional image encoding features; Decoding each sample two-dimensional image encoding feature by using a two-dimensional decoding network to obtain a first object prediction result of the sample three-dimensional image; Encoding the sample three-dimensional voxel image by using a three-dimensional encoding network to obtain a plurality of second sample three-dimensional image encoding features; Fusing the sample two-dimensional image encoding features and the second sample three-dimensional image encoding features by using a multi-dimensional feature fusion network to obtain a plurality of sample fusion features; Decoding each sample fusion feature and at least one second sample three-dimensional image encoding feature by using a three-dimensional decoding network to obtain a second object prediction result of the sample three-dimensional voxel image; Determining a two-dimensional network loss according to the first object prediction result and the object annotation, and determining a three-dimensional network loss according to the second object prediction result and the object annotation; Adjusting parameters of the two-dimensional global feature extraction network, the two-dimensional encoding / decoding network and the three-dimensional encoding / decoding network according to the two-dimensional network loss and the three-dimensional network loss; Selecting other sample three-dimensional image data to continue training until a preset ending condition is met to obtain a trained object recognition model.
12. The method of claim 11, wherein, The cutting of the sample three-dimensional image into a plurality of sample three-dimensional voxel images comprises: In the case that the sample three-dimensional image is a positive sample, a region containing an object in the sample three-dimensional image is cropped into m sample three-dimensional voxel images, and the remaining region is cropped into n-m sample three-dimensional voxel images, the positive sample is used to indicate that the sample three-dimensional image contains an object, ; In the case that the sample three-dimensional image is a negative sample, the sample three-dimensional image is cropped into q sample three-dimensional voxel images, the negative sample is used to indicate that the sample three-dimensional image does not contain an object, q n.
13. The method of claim 11, wherein, The two-dimensional encoding network comprises K two-dimensional encoding layers, the three-dimensional encoding network comprises K three-dimensional encoding layers, and the fusing of the sample two-dimensional image encoding features and the second sample three-dimensional image encoding features by using the multi-dimensional feature fusion network to obtain a plurality of sample fusion features comprises: Fusing sample two-dimensional image encoding features output by an lth two-dimensional encoding layer and second sample three-dimensional image encoding features output by an lth three-dimensional encoding layer by using a multi-dimensional feature fusion network to obtain an lth sample fusion feature, wherein l [1, K-1].
14. The method of claim 13, wherein, The three-dimensional decoding network comprises K three-dimensional decoding layers and a three-dimensional convolution layer, and the decoding of each sample fusion feature and at least one second sample three-dimensional image encoding feature by using the three-dimensional decoding network to obtain a second object prediction result of the sample three-dimensional voxel image comprises: Decoding sample image features output by the Kth three-dimensional encoding layer by using the first three-dimensional decoding layer; decoding the sample image features output by the (t-1)th three-dimensional decoding layer and the (K+1-t)th sample fusion feature by using the tth three-dimensional decoding layer, where t∈[2,K]; performing convolution operation on the sample image features output by the Kth three-dimensional decoding layer by using a three-dimensional convolution layer to obtain the second object prediction result of the sample three-dimensional voxel map.
15. An object recognition apparatus characterized by comprising: The device comprises: a first acquisition module configured to acquire a three-dimensional image to be recognized and a pre-trained object recognition model, wherein the object recognition model comprises a two-dimensional global feature extraction network, a two-dimensional feature processing network, a three-dimensional encoding / decoding network, and a multi-dimensional feature fusion network; a first cropping module configured to crop the three-dimensional image to be recognized into a plurality of three-dimensional voxel maps to be recognized; an identification module configured to input the three-dimensional image to be recognized and the plurality of three-dimensional voxel maps to be recognized into the object recognition model to obtain an object recognition result, including: extracting target two-dimensional global features of the three-dimensional image to be recognized by using the two-dimensional global feature extraction network; performing feature encoding on the target two-dimensional global features by using the two-dimensional feature processing network to obtain a plurality of target two-dimensional image encoding features; performing feature encoding on the three-dimensional voxel maps to be recognized by using the three-dimensional encoding network to obtain a plurality of second target three-dimensional image encoding features; fusing the target two-dimensional image encoding features and the second target three-dimensional image encoding features by using the multi-dimensional feature fusion network to obtain a plurality of target fusion features; and decoding each target fusion feature and at least one second target three-dimensional image encoding feature by using the three-dimensional decoding network to obtain a second object prediction result of the three-dimensional voxel map to be recognized, wherein the object recognition result comprises the second object prediction result of each three-dimensional voxel map to be recognized.
16. An object recognition model training apparatus, comprising: The device comprises: a second acquisition module configured to acquire a sample three-dimensional image, object labels of the sample three-dimensional image, and an object recognition model to be trained, wherein the object recognition model comprises a two-dimensional global feature extraction network, a two-dimensional feature processing network, a three-dimensional encoding / decoding network, and a multi-dimensional feature fusion network, and the two-dimensional feature processing network comprises a two-dimensional encoding / decoding network; a second cropping module configured to crop the sample three-dimensional image into a plurality of sample three-dimensional voxel maps; and a training module configured to train the object recognition model by using the sample three-dimensional image, the object labels of the sample three-dimensional image, and the plurality of sample three-dimensional voxel maps. The first determining module is configured to extract sample two-dimensional global features of the sample three-dimensional image by using a two-dimensional global feature extraction network; encode the sample two-dimensional global features by using a two-dimensional encoding network to obtain a plurality of sample two-dimensional image encoding features; decode each of the sample two-dimensional image encoding features by using a two-dimensional decoding network to obtain a first object prediction result of the sample three-dimensional image; encode the sample three-dimensional voxel image by using a three-dimensional encoding network to obtain a plurality of second sample three-dimensional image encoding features; fuse the sample two-dimensional image encoding features and the second sample three-dimensional image encoding features by using a multi-dimensional feature fusion network to obtain a plurality of sample fusion features; and decode each of the sample fusion features and at least one of the second sample three-dimensional image encoding features by using a three-dimensional decoding network to obtain a second object prediction result of the sample three-dimensional voxel image. The second determining module is configured to determine a two-dimensional network loss according to the first object prediction result and the object label, and determine a three-dimensional network loss according to the second object prediction result and the object label. The adjusting module is configured to adjust parameters of the two-dimensional global feature extraction network, the two-dimensional encoding / decoding network, and the three-dimensional encoding / decoding network according to the two-dimensional network loss and the three-dimensional network loss. The training module is configured to select other sample three-dimensional image data for continuous training until a preset ending condition is met, and obtain a trained object recognition model.
17. An electronic device, comprising: The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the method in any one of claims 1-10 or 11-14. The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the method in any one of claims 1-10 or 11-14. The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the method in any one of claims 1-10 or 11-14.
18. A computer-readable storage medium, characterized in that,
Citation Information
Patent Citations
Three-dimensional medical image segmentation method based on contextual information fusion
CN114882219A