Image recognition method, recognition model training method, device, and electronic device

By training a second recognition model using rotated and sorted cubes from 3D images, the method addresses inefficiencies in existing 3D image recognition model training, resulting in improved training speed and accuracy.

CN111046855BActive Publication Date: 2025-07-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010043334.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-21
Filing Date
2020-01-15
Publication Date
2025-07-15
Estimated Expiration
2040-04-23

AI Technical Summary

Technical Problem

Existing methods for training 3D image recognition models require a large number of 3D image samples, leading to inefficient model training processes.

Method used

The method involves training a second recognition model using rotated and sorted cubes extracted from 3D sample images, which are then used to enhance the efficiency of the first recognition model by sharing convolutional blocks, thereby improving the training speed.

Benefits of technology

This approach significantly enhances the training efficiency of the first recognition model by leveraging the training of the second model, allowing for faster and more effective identification of 3D images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111046855B_ABST
    Figure CN111046855B_ABST
Patent Text Reader

Abstract

The present invention discloses a picture recognition method, a recognition model training method, a device and an electronic device. Among them, the method includes: obtaining a target 3D picture to be recognized; inputting the target 3D picture to be recognized into a first recognition model, where the first recognition model is used to recognize the target 3D picture to be recognized to obtain the picture type of the target 3D picture to be recognized, and the convolutional blocks of the first recognition model are the same as those of the second recognition model, and the second recognition model is a model obtained by training an original recognition model using target training samples, and the target training samples include cubes obtained by rotating and sorting N target cubes obtained from 3D sample pictures, and N is a natural number greater than 1; obtaining the first type of the target 3D picture to be recognized output by the first recognition model. The present invention solves the technical problem of low model training efficiency in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and in particular, to a method for image recognition, a method for training an identification model, a device, and an electronic device. Background Art

[0002] In the related art, when identifying the type of a 3D image, it is usually necessary to use a large number of 3D picture samples to train a 3D model, and then the trained 3D model can be used to identify the type of the 3D image.

[0003] However, if the above method is used, it takes a lot of time to train the model, resulting in the problem of low training efficiency of the model.

[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of the present invention provide a method for image recognition, a method for training an identification model, a device, and an electronic device, so as to at least solve the technical problem of low model training efficiency in the related art.

[0006] According to one aspect of the embodiments of the present invention, a method for image recognition is provided, including: obtaining a target 3D picture to be recognized; inputting the target 3D picture to be recognized into a first recognition model, where the first recognition model is used to recognize the target 3D picture to be recognized to obtain the picture type of the target 3D picture to be recognized, the convolutional blocks of the first recognition model are the same as those of a second recognition model, and the second recognition model is a model obtained by training an original recognition model with a target training sample, and the target training sample includes cubes obtained by rotating and sorting N target cubes obtained from 3D sample pictures, and N is a natural number greater than 1; obtaining a first type of the target 3D picture output by the first recognition model.

[0007] According to another aspect of the embodiments of the present invention, a method for training an identification model is further provided, including: obtaining 3D sample pictures, and segmenting N target cubes from the 3D sample pictures; performing a predetermined operation on the N target cubes to obtain a target training sample, where the predetermined operation includes rotating and sorting the N target cubes; using the target training sample to train the original recognition model to obtain a second recognition model, where the original recognition model is used to output an identification result of the target training sample, and when the probability that the identification result satisfies a first objective function is greater than a first threshold, the original recognition model is determined as the second recognition model.

[0008] According to another aspect of the embodiments of the present invention, there is also provided an image recognition device, including: a first acquisition unit, configured to acquire a target 3D image to be recognized; a first input unit, configured to input the target 3D image to be recognized into a first recognition model, wherein the first recognition model is used to recognize the target 3D image to be recognized to obtain the image type of the target 3D image to be recognized, the convolutional blocks of the first recognition model are the same as those of the second recognition model, and the second recognition model is a model obtained by training an original recognition model using target training samples, and the target training samples include cubes obtained by rotating and sorting N target cubes acquired from 3D sample images, and N is a natural number greater than 1; a second acquisition unit, configured to acquire the first type of the target 3D image output by the first recognition model.

[0009] As an optional example, the device further includes: a third acquisition unit, configured to acquire the 3D sample image before acquiring the target 3D image to be recognized; a first determination unit, configured to determine an original cube from the 3D sample image; a splitting unit, configured to split the original cube into the N target cubes.

[0010] As an optional example, the device further includes: a second determination unit, configured to determine a first target cube from the N target cubes before acquiring the target 3D image to be recognized; a rotation unit, configured to rotate the first target cube by a first angle; a sorting unit, configured to sort the first target cube after being rotated by the first angle and other target cubes among the N target cubes to obtain the target training samples.

[0011] As an optional example, the device further includes: a second input unit, configured to input the target training samples into the original recognition model after sorting the first target cube after being rotated by the first angle and other target cubes among the N target cubes to obtain the target training samples, so as to train the original recognition model to obtain the second recognition model.

[0012] As an optional example, N is the cube of a positive integer greater than 1, and the splitting unit includes: a splitting module, configured to keep a space of M voxels between two adjacent target cubes, and split the N target cubes from the original cube, where M is a positive integer greater than 0 and less than J - 1, and J is the side length of the target cube.

[0013] As an alternative example, the above device further includes: a fourth acquisition unit, configured to acquire the recognition result output by the original recognition model after recognizing the target training sample before acquiring the target 3D picture to be recognized, where the recognition result includes various sorting orders of the target cube in the target training sample and the probability of the rotation angle of each target cube; a third determination unit, configured to determine the original recognition model as the second recognition model when the probability that the recognition result satisfies the first objective function is greater than the first threshold.

[0014] As an alternative example, the above device further includes: a fourth determination unit, configured to determine the convolutional block of the second recognition model as the convolutional block of the first recognition model before acquiring the target 3D picture to be recognized; a training unit, configured to train the first recognition model using the first training sample until the accuracy of the first recognition model is greater than the second threshold, where the first training sample includes the first 3D picture and the type of the first 3D picture.

[0015] According to another aspect of the embodiments of the present invention, there is also provided a recognition model training device, including: a segmentation unit, configured to acquire a 3D sample picture and segment N target cubes from the 3D sample picture; a processing unit, configured to perform a predetermined operation on the N target cubes to obtain a target training sample, where the predetermined operation includes rotating and sorting the N target cubes; a training unit, configured to train an original recognition model using the target training sample to obtain a second recognition model, where the original recognition model is used to output a recognition result of the target training sample, and when the probability that the recognition result satisfies the first objective function is greater than the first threshold, determine the original recognition model as the second recognition model.

[0016] According to another aspect of the embodiments of the present invention, there is also provided a model training method, including:

[0017] Convert the original three-dimensional picture data into three-dimensional cube training samples, where the three-dimensional cube training samples include multiple micro-cubes; sequentially perform a first operation and a second operation on the multiple micro-cubes to obtain target training samples, where the first operation is used to change the order of the multiple micro-cubes, and the second operation is used to change the direction of the first object micro-cubes among the multiple micro-cubes; use the target training samples for training to obtain a pre-trained network model, where the pre-trained network model is used to extract features in the original three-dimensional picture data and is also used to identify the data structure in the original three-dimensional picture data; migrate a target fully connected layer matching the target picture recognition task to the pre-trained network model to obtain a first recognition model; input the target three-dimensional picture data to be recognized into the first recognition model to obtain a recognition result, where the recognition result includes the abnormal area in the target three-dimensional picture data.

[0018] According to another aspect of the embodiments of the present invention, there is also provided a model training device, including: a conversion unit, configured to convert the original three-dimensional picture data into three-dimensional cube training samples, where the three-dimensional cube training samples include multiple micro-cubes; a first execution unit, configured to sequentially perform a first operation and a second operation on the multiple micro-cubes to obtain target training samples, where the first operation is used to change the order of the multiple micro-cubes, and the second operation is used to change the direction of the first object micro-cubes among the multiple micro-cubes; a training unit, configured to use the target training samples for training to obtain a pre-trained network model, where the pre-trained network model is used to extract features in the original three-dimensional picture data and is also used to identify the data structure in the original three-dimensional picture data; a migration unit, configured to migrate a target fully connected layer matching the target picture recognition task to the pre-trained network model to obtain a first recognition model; an input unit, configured to input the target three-dimensional picture data to be recognized into the first recognition model to obtain a recognition result, where the recognition result includes the abnormal area in the target three-dimensional picture data.

[0019] As an optional example, the first execution unit includes: an arrangement module, configured to perform permutation and combination on the multiple micro-cubes to obtain K types of micro-cube combinations; a first determination module, configured to determine a target micro-cube combination from the K types of micro-cube combinations; a second determination module, configured to determine the first object micro-cubes from the target micro-cube combination; an execution module, configured to perform a rotation operation on the first object micro-cubes to obtain the target training samples.

[0020] As an alternative example, the above-mentioned device further includes: a determination unit, configured to determine, after sequentially performing a first operation and a second operation on the above-mentioned multiple micro-cubes, a second target micro-cube from the above-mentioned multiple micro-cubes; a second execution unit, configured to perform a third operation on the above-mentioned second target micro-cube to update the above-mentioned target training sample, where the above-mentioned third operation is used to occlude a partial area of the above-mentioned second target micro-cube.

[0021] As an alternative example, the above-mentioned second execution unit includes: an operation module, configured to multiply the above-mentioned second target micro-cube by a target matrix, where the above-mentioned target matrix is a three-dimensional matrix having the same size as the above-mentioned second target micro-cube.

[0022] As an alternative example, the above-mentioned conversion unit includes: a conversion module, configured to convert the above-mentioned original three-dimensional picture data into an original three-dimensional cube; a splitting module, configured to split the above-mentioned original three-dimensional cube into multiple original micro-cubes; an extraction module, configured to extract the above-mentioned multiple micro-cubes from the above-mentioned multiple original micro-cubes.

[0023] As an alternative example, the above-mentioned splitting module includes: a sub-splitting module, configured to split the above-mentioned original three-dimensional cube to obtain the above-mentioned multiple original micro-cubes, where a space of M voxels is maintained between two adjacent ones of the above-mentioned original micro-cubes, and the M is a positive integer greater than 0 and less than J-1, and the J is the side length of the above-mentioned original micro-cube.

[0024] As an alternative example, the above-mentioned device further includes: a construction unit, configured to construct an original network model of the above-mentioned pre-trained network model before training by using the above-mentioned target training sample to obtain the pre-trained network model, where the objective function in the fully-connected layer of the above-mentioned original network model includes: a first loss function corresponding to the above-mentioned first operation, a second loss function corresponding to the above-mentioned second operation, and a third loss function corresponding to the above-mentioned third operation; the above-mentioned training unit includes: an input module, configured to input the above-mentioned target training sample into the above-mentioned original network model for training to obtain the above-mentioned pre-trained network model, where the output result of the objective function in the fully-connected layer of the above-mentioned pre-trained network model has reached a convergence condition.

[0025] As an alternative example, the above-mentioned transfer unit includes: an acquisition module, configured to acquire the above-mentioned target picture recognition task to be currently processed; a third determination module, configured to determine a target fully-connected layer matching the above-mentioned target picture recognition task; a replacement module, configured to replace the fully-connected layer of the above-mentioned pre-trained network model with the above-mentioned target fully-connected layer to obtain the above-mentioned first recognition model, where the above-mentioned first recognition model is used to perform the above-mentioned target picture recognition task.

[0026] According to another aspect of the embodiments of the present invention, there is also provided a storage medium storing a computer program, wherein the computer program is configured to execute the above-mentioned picture recognition method when running.

[0027] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the above-mentioned processor executes the above-mentioned picture recognition method through the computer program.

[0028] In the embodiments of the present invention, the method includes: obtaining a target 3D picture to be recognized; inputting the target 3D picture to be recognized into a first recognition model, wherein the first recognition model is used to recognize the target 3D picture to be recognized to obtain the picture type of the target 3D picture to be recognized, the convolutional blocks of the first recognition model are the same as those of the second recognition model, the second recognition model is a model obtained by training an original recognition model with target training samples, the target training samples include cubes obtained by rotating and sorting N target cubes extracted from 3D sample pictures, and N is a natural number greater than 1; obtaining the first type of the target 3D picture output by the first recognition model. Since in the above method, the cubes extracted from 3D pictures are used to train the second recognition model in advance, the training efficiency of the second recognition model is improved. Further, the convolutional blocks of the second recognition model are used as the convolutional blocks of the first recognition model, and the first recognition model is used to recognize 3D pictures, achieving the effect of greatly improving the training efficiency of the first recognition model and solving the technical problem of low model training efficiency in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the illustrative embodiments and descriptions of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0030] Figure 1 is a schematic diagram of an application environment of an optional picture recognition method according to an embodiment of the present invention;

[0031] Figure 2 is a schematic flowchart of an optional picture recognition method according to an embodiment of the present invention;

[0032] Figure 3 is a schematic diagram of an optional picture recognition method according to an embodiment of the present invention;

[0033] Figure 4 is a schematic diagram of another optional picture recognition method according to an embodiment of the present invention;

[0034] Figure 5 It is a schematic diagram of another optional picture recognition method according to an embodiment of the present invention;

[0035] Figure 6 It is a schematic diagram of another optional picture recognition method according to an embodiment of the present invention;

[0036] Figure 7 It is a schematic diagram of another optional picture recognition method according to an embodiment of the present invention;

[0037] Figure 8 It is a schematic diagram of another optional picture recognition method according to an embodiment of the present invention;

[0038] Figure 9 It is a schematic diagram of another optional picture recognition method according to an embodiment of the present invention;

[0039] Figure 10 It is a schematic flowchart of an optional recognition model training method according to an embodiment of the present invention;

[0040] Figure 11 It is a schematic structural diagram of an optional picture recognition device according to an embodiment of the present invention;

[0041] Figure 12 It is a schematic structural diagram of an optional recognition model training device according to an embodiment of the present invention;

[0042] Figure 13 It is a schematic flowchart of an optional model training method according to an embodiment of the present invention;

[0043] Figure 14 It is a schematic diagram of an optional model training method according to an embodiment of the present invention;

[0044] Figure 15 It is a schematic structural diagram of an optional model training device according to an embodiment of the present invention;

[0045] Figure 16 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present invention;

[0046] Figure 17 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present invention;

[0047] Figure 18 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present invention. Detailed implementation manners

[0048] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0049] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0050] Magnetic Resonance Imaging (MRI for short): A type of medical imaging.

[0051] Computed Tomography (CT for short): A type of medical imaging that can be used for the examination of various diseases.

[0052] Convolution neural network (CNN for short)

[0053] Multimodal Brain Tumor Segmentation (BRATS for short)

[0054] Feature map: The feature map obtained after the image and the filter are convolved. The feature map can be convolved with the filter to generate a new feature map.

[0055] Siamese network: It contains several convolutional neural networks with the same structure, and the weight parameters can be shared among the networks

[0056] Hamming distance: The Hamming distance, which measures the number of different characters at the corresponding positions of two strings

[0057] ImageNet: ImageNet is a large visual database for software research on visual object recognition, containing more than 14 million images and their corresponding annotation information.

[0058] ResNet: Residual Neural Network, a convolutional neural network based on residual learning. VGG: A deep convolutional neural network developed by the computer vision team at the University of Oxford and researchers at Google DeepMind.

[0059] one-hot: One-hot encoding mainly uses an N-bit status register to encode N states. Each state has an independent register bit, and only one bit is valid at any time.

[0060] Fully convolutional network (FCN for short): A convolutional network most commonly used in image segmentation technology, consisting entirely of convolutional layers and pooling layers.

[0061] According to one aspect of the embodiments of the present invention, a method for image recognition is provided. Optionally, as an alternative implementation, the above-mentioned method for image recognition can be but is not limited to being applied to an environment such as Figure 1 shown.

[0062] Figure 1 In the figure, human-computer interaction can be carried out between user 102 and user device 104. The user device 104 includes a memory 106 for storing interaction data and a processor 108 for processing interaction data. The user device 104 can perform data interaction with the server 112 through the network 110. The server 112 includes a database 114 for storing interaction data and a processing engine 116 for processing interaction data. The user device 104 includes the above-mentioned first recognition model. The user device 104 can obtain the target 3D image 104-2 to be recognized, recognize the target 3D image 104-2, and output the first type 104-4 of the target 3D image 104-2.

[0063] Optionally, the above-mentioned method for image recognition can be but is not limited to being applied to terminals that can calculate data, such as mobile phones, tablet computers, laptops, PCs, etc. The above-mentioned network can include but is not limited to wireless networks or wired networks. Among them, the wireless network includes: Bluetooth, WIFI, and other networks that implement wireless communication. The above-mentioned wired network can include but is not limited to: wide area network, metropolitan area network, local area network. The above-mentioned server can include but is not limited to any hardware device that can perform calculations.

[0064] Optionally, as an alternative implementation, as Figure 2As shown in the figure, the above-mentioned image recognition method includes:

[0065] S202, obtaining a target 3D image to be recognized;

[0066] S204, inputting the target 3D image to be recognized into a first recognition model, where the first recognition model is used to recognize the target 3D image to be recognized to obtain the image type of the target 3D image to be recognized. The convolutional blocks of the first recognition model are the same as those of the second recognition model. The second recognition model is a model obtained by training an original recognition model with target training samples. The target training samples include cubes obtained by rotating and sorting N target cubes obtained from 3D sample images, and N is a natural number greater than 1;

[0067] S206, obtaining the first type of the target 3D image output by the first recognition model.

[0068] Optionally, the above-mentioned image recognition method can be but is not limited to being applied in the field of image recognition. For example, applying the above method to the process of recognizing the type of 3D images. Such as in the process of recognizing the type of diseases in 3D disease images. For example, when recognizing the type of cerebral hemorrhage, after obtaining a 3D disease image (the 3D disease image can be an MRI image or a CT image), input the 3D disease image into the first recognition model, and use the first model to recognize the 3D disease image and output the first type of the 3D disease image. For example, the first type can be healthy, or aneurysm, arteriovenous malformation, moyamoya disease, hypertension, etc.

[0069] In the above method, since the second recognition model is pre-trained with cubes extracted from 3D images, the training efficiency of the second recognition model is improved. Further, using the convolutional blocks of the second recognition model as the convolutional blocks of the first recognition model and using the first recognition model to recognize 3D images achieves the effect of greatly improving the training efficiency of the first recognition model.

[0070] Optionally, in the above method, before obtaining the target 3D image, the second recognition model needs to be trained first. During training, first, 3D sample images need to be obtained. The 3D sample images are images without label annotation. After obtaining the 3D sample images, the original cubes need to be extracted from the 3D sample images, and the original cubes are split into N target cubes.

[0071] Optionally, when extracting the original cube, the geometric center of the 3D sample image can be determined first. After determining the geometric center, using this geometric center as the geometric center of the above-mentioned original cube, and determining the original cube. The side length of the above-mentioned original cube is less than the length of the smallest side of the 3D sample image.

[0072] For example, as Figure 3 shown, for a 3D sample picture 302, first determine the geometric center 304 of the 3D sample picture 302, and then determine the original cube 306 with the geometric center 304 as its geometric center.

[0073] Optionally, after determining the geometric center of the 3D sample picture, a radius r can also be determined, and then a sphere is made with the geometric center of the 3D sample picture as the center and the radius r as the radius. Then, any point is selected from the sphere as the geometric center of the above-mentioned original cube to determine the above-mentioned original cube. It should be noted that the determined original cube is located within the 3D sample picture and will not exceed the range of the 3D sample picture.

[0074] Optionally, after determining the original cube, the original cube needs to be split to obtain N target cubes. When splitting, any method can be used, such as randomly digging out N target cubes from the original cube, or splitting a part of the original cube to obtain N target cubes. Or, the original cube is evenly split into N target cubes, where N is the cube of a positive integer. Taking N as 8 as an example, as Figure 4 shown, an original cube 404 is split in the directions indicated by the arrows of 402-1, 402-2, and 402-3 to obtain 8 target cubes ( Figure 4 The splitting method in is only an example). Or, when splitting, there is a gap of M voxels between every two adjacent cubes. For example, taking M as 2 as an example, as Figure 5 shown, the original cube 502 is split into 8 target cubes 504. The side length of the original cube 502 is 10 voxels, then the side length of the target cube 504 is 4 voxels.

[0075] Optionally, after obtaining N target cubes, the first target cube among the N target cubes can also be rotated by a first angle, such as rotating 90 degrees, 180 degrees, etc. There can be one or more first target cubes, and the rotation angle of each first target cube can be the same or different. The rotated first target cube and the remaining unrotated target cubes are sorted, and the sorting can be random sorting. After sorting, the target training sample is obtained.

[0076] After obtaining the target training samples, use the target training samples to train the original recognition model, and the original recognition model outputs the probability of which rotation and arrangement order the target cube in the target training samples has. The above probability may satisfy the first objective function or may not satisfy the first objective function. The first objective function can be a loss function. If the above probability satisfies the first objective function, it means that the recognition result of the original recognition model is correct. If the above probability does not satisfy the first objective function, it means that the recognition result of the original recognition model is incorrect. When the probability that the recognition result satisfies the first objective function is greater than the first threshold, determine the original recognition model as the second recognition model. It shows that the accuracy of the second recognition model is greater than the first threshold. For example, the accuracy reaches more than 99.95%.

[0077] Using the above training method greatly improves the efficiency of training the second recognition model.

[0078] Optionally, after training the second recognition model, the convolutional blocks in the second recognition model can be obtained, and the convolutional blocks can be used as the convolutional blocks of the first recognition model, and the first training samples are used to train the first recognition model. The first training samples are 3D pictures including picture types. After the recognition accuracy of the first recognition model is greater than the second threshold, the first recognition model can be put into use. For example, recognizing the disease type of 3D pictures. As Figure 6 As shown, on the display interface 602 of the terminal, there is a selection button 602-1, the user can select the target 3D picture 604 to be recognized, the terminal recognizes the target 3D picture 604 to be recognized, and outputs the first type 606 of the target 3D picture to be recognized.

[0079] The following is illustrated with a specific example.

[0080] For example, when recognizing brain diseases, obtain the publicly available BRATS-2018 brain glioma segmentation dataset and the intracerebral hemorrhage classification dataset collected from cooperative hospitals, and the above data are used as experimental data.

[0081] The BRATS-2018 dataset includes MRI images of 285 patients. Each patient's MRI image includes 4 different modalities, namely T1, T1Gd, T2, and FLAIR. The data of different modalities have been co-registered, and the size of each image is 240x240x155.

[0082] The intracerebral hemorrhage dataset includes 1486 brain CT scan images of intracerebral hemorrhage. The types of intracerebral hemorrhage are aneurysm, arteriovenous malformation, moyamoya disease, and hypertension. The size of each CT image is 230x270x30.

[0083] Use the above pictures for training the second recognition model. As Figure 7As shown, for a picture, the original cube is extracted from the picture and the original cube is split into target cubes. For the specific method of selecting the original cube, please refer to the above example and will not be repeated here. After selecting the original cube, in order to encourage the network to learn high-level semantic feature information rather than low-level statistical feature information of pixel distribution through the proxy task of Rubik's Cube restoration, we reserve a random interval of less than 10 voxels between two adjacent target cubes when cutting the original cube to obtain the target cubes, and then perform [-1,1] normalization on the voxels in each target cube to obtain the target training samples.

[0084] After obtaining the target training samples, the second recognition model needs to be trained. As Figure 7 shown, the Siamese network includes X sub-networks that share weights with each other, where X represents the number of target cubes. In the experiment, an eight-in-one Siamese network with 8 target cube inputs was used. Each sub-network has the same network structure and shares weights with each other. The backbone structure of each sub-network can use various existing types of 3D CNNs. In the experiment, a 3D VGG network was used. The output feature maps of the last fully connected layer of all sub-networks are stacked and then input into different branches, which are used for the spatial rearrangement task of the target cube and the rotation judgment task of the target cube. The above feature map is the content output by any one network in the convolutional model.

[0085] 1. Rearrangement of the target cube

[0086] For the Rubik's Cube restoration task proposed in this scheme, the first step is to rearrange the target cubes. Taking the second-order Rubik's Cube as an example, as Figure 7 shown, it has a total of 2 x 2 x 2 = 8 target cubes. We first need to generate all permutation and combination sequences P=(P1, P2,..., P8!) of the 8 target cubes. These permutation sequences control the complexity of the Rubik's Cube restoration task. If two permutation sequences are too similar to each other, the learning process of the network will become very simple and it is difficult to learn complex feature information. To ensure the effectiveness of learning, the Hamming distance is used as a measurement index, and K sequences with greater differences from each other are selected in turn. For each training input data of Rubik's Cube restoration, a random one is selected from the K sequences, such as (2, 5, 8, 4, 1, 7, 3, 6), and then the 8 cut target cubes are rearranged in the order of this sequence, and then the rearranged target cubes are input into the network in turn. Finally, the goal that the network needs to learn is to judge which of these K sequences the input sequence belongs to. Therefore, for the rearrangement of the target cube, the loss function is as follows:

[0087]

[0088] In the above formula, lj represents the true label one-hot label of the sequence, and pj represents the predicted probability for each sequence output by the network.

[0089] 2. Rotation of the target cube

[0090] In the 3D Rubik's Cube restoration task, a new operation is added, that is, the rotation of the target cube. Through this operation, the network can learn the rotation-invariant features of 3D image blocks.

[0091] The target cube is usually of a cube structure. If a target cube rotates freely in space, there will be 3 (rotation axes, x, y, z axes) x 2 (rotation directions, clockwise, counterclockwise) x 4 (rotation angles, 0°, 90°, 180°, 270°) = 24 different possibilities. To reduce the complexity of the task, the rotation options of the target cube are restricted, and it is stipulated that the target cube can only rotate 180° in the horizontal or vertical direction. As Figure 2 shown, cubelets 3 and 4 are rotated 180° horizontally, and cubelets 5 and 7 are rotated 180° vertically. After the rotated cubelets are input into the network, the network needs to judge what form of rotation each target cube has undergone. Therefore, for the cubelet rotation task, its loss function is as follows:

[0092]

[0093] In the formula, M represents the number of target cubes, g i hor represents the one-hot label of the vertical rotation of the target cube, and g i ver represents the one-hot label of the horizontal rotation of the target cube, r i hor , r i ver respectively represent the predicted output probabilities of the network in the vertical and horizontal directions.

[0094] According to the previous definition, the objective function of the model is the linear weighted sum of the permutation loss function and the rotation loss function. The overall loss function of the model is as follows:

[0095] loss = a * loss p + b * loss R (3)

[0096] Where a and b are the weights of two loss functions respectively, controlling the degree of mutual influence between the two sub-tasks. Setting both weight values to 0.5 in the experiment can achieve better pre-training results.

[0097] After the above training, a second recognition model can be obtained. The accuracy of the second recognition model is greater than the first threshold.

[0098] At this time, the convolutional blocks of the second recognition model can be extracted and fine-tuned for use in other target tasks.

[0099] For example, extract the convolutional blocks of the second recognition model for the first recognition model to identify the types of 3D pictures. For the classification task, only the fully connected layer behind the CNN network needs to be retrained, and the convolutional layers before the fully connected layer can be fine-tuned using a smaller learning rate.

[0100] Or use the convolutional blocks of the above second recognition model for the segmentation task. For the segmentation task, the pre-trained network can use the fully convolutional neural network (FCN) commonly used in image segmentation tasks, such as the 3D U-Net structure, as Figure 8 shown. However, since the pre-training in the form of Rubik's Cube restoration in the early stage can only be applied to the downsampling stage of U-Net, the network parameters in the upsampling stage of U-Net still need to be randomly initialized during training. To avoid the impact of a large number of parameter initializations on the pre-training effect in the early stage, the Dense Upsampling Convolution (DUC) module is used to replace the original transposed convolution to upsample the feature map and restore it to the original input size of the image. The structure of the DUC module is as Figure 9 shown. Where C represents the number of channels, d represents the expansion factor. H is the length of the feature map, and W is the width of the feature map.

[0101] Through this embodiment, since the second recognition model is pre-trained using the cubes extracted from 3D pictures in advance, the training efficiency of the second recognition model is improved. Further, using the convolutional blocks of the second recognition model as the convolutional blocks of the first recognition model to identify 3D pictures with the first recognition model achieves the effect of greatly improving the training efficiency of the first recognition model.

[0102] As an alternative implementation, before obtaining the target 3D picture to be recognized, it further includes:

[0103] S1. Obtain the 3D sample pictures;

[0104] S2. Determine the original cubes from the 3D sample pictures;

[0105] S3. Split the original cube into the N target cubes.

[0106] Optionally, in this solution, the 3D sample image and the target 3D image can be the same image. That is, after training the second recognition model using the 3D sample image and using the second convolutional block as the convolutional block of the first recognition model, the 3D sample image can be input into the first recognition model, and the first recognition model can identify the type of the 3D sample image. When the 3D sample image is input into the second recognition model, the type of the 3D sample image does not need to be input.

[0107] Through this embodiment and the above method, before using the first recognition model, N target cubes are obtained to train the second recognition model, which improves the training efficiency of training the second recognition model and further improves the training efficiency of the first recognition model.

[0108] As an optional implementation, N is the cube of a positive integer greater than 1, and the splitting of the original cube into the N target cubes includes:

[0109] S1. Keep a space of M voxels between two adjacent target cubes, and split the N target cubes from the original cube. M is a positive integer greater than 0 and less than J - 1, and J is the side length of the target cube.

[0110] Optionally, when determining the N target cubes, keeping a space of M voxels between two adjacent target cubes can enable the second recognition model to learn high-level semantic feature information rather than low-level statistical feature information of pixel distribution, improving the training efficiency of the second recognition model and further improving the training efficiency of the first recognition model.

[0111] As an optional implementation, before obtaining the target 3D image to be recognized, it further includes:

[0112] S1. Determine a first target cube from the N target cubes;

[0113] S2. Rotate the first target cube by a first angle;

[0114] S3. Sort the first target cube after rotating the first angle and the other target cubes among the N target cubes to obtain the target training sample.

[0115] Optionally, the above sorting can be randomly sorting the N target cubes. The above rotation can rotate multiple first target cubes among the N target cubes. The rotation can rotate at any angle.

[0116] Through this embodiment and the above method, before using the first recognition model, after obtaining N target cubes, the first target cube among the N target cubes is rotated, which improves the training efficiency of training the second recognition model and further improves the training efficiency of the first recognition model.

[0117] As an alternative implementation, after sorting the first target cube after rotating the first angle among the N target cubes to obtain the target training sample, it further includes:

[0118] S1. Input the target training sample into the original recognition model to train the original recognition model to obtain the second recognition model.

[0119] Through this embodiment and the above method, the training efficiency of training the second recognition model is improved, and the training efficiency of the first recognition model is further improved.

[0120] As an alternative implementation, before obtaining the target 3D picture to be recognized, it further includes:

[0121] S1. Obtain the recognition result output by the original recognition model after recognizing the target training sample, where the recognition result includes various sorting orders of the target cubes in the target training sample and the probability of the rotation angle of each target cube.

[0122] S2. When the probability that the recognition result satisfies the first objective function is greater than the first threshold, determine the original recognition model as the second recognition model.

[0123] Optionally, the training of the second recognition model cannot continue indefinitely. When the recognition accuracy of the second recognition model is greater than a certain value, it is considered that the second recognition model meets the requirements, and thus the training is stopped.

[0124] Through this embodiment, by setting a selection condition to stop the training of the second recognition model, the training efficiency of training the second recognition model is improved.

[0125] As an alternative implementation, before obtaining the target 3D picture to be recognized, it further includes:

[0126] S1. Determine the convolutional blocks of the second recognition model as the convolutional blocks of the first recognition model.

[0127] S2. Use the first training sample to train the first recognition model until the accuracy of the first recognition model is greater than the second threshold, where the first training sample includes the first 3D picture and the type of the first 3D picture.

[0128] Optionally, when training the first recognition model, first sample pictures with labels can be input. Then, the first recognition model is trained until the recognition accuracy of the first recognition model is greater than a second threshold, and then the first recognition model can be put into use.

[0129] In this embodiment, by training the first recognition model before using it, the training efficiency of training the first recognition model is improved.

[0130] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0131] According to another aspect of the embodiments of the present invention, a method for training a recognition model is further provided. As Figure 10 shown, the method includes:

[0132] S1002, obtaining 3D sample pictures, and segmenting N target cubes from the 3D sample pictures;

[0133] S1004, performing a predetermined operation on the N target cubes to obtain target training samples, where the predetermined operation includes rotating and sorting the N target cubes;

[0134] S1006, using the target training samples to train an original recognition model to obtain a second recognition model, where the original recognition model is used to output a recognition result of the target training samples, and when the probability that the recognition result satisfies a first objective function is greater than a first threshold, the original recognition model is determined as the second recognition model.

[0135] Optionally, the above method can be but is not limited to being applied in the process of model training. When training the original recognition model, N target cubes are extracted from a 3D sample picture, and the N cubes obtained after rotating and sorting the N target cubes are used as target training samples and input into the original recognition model. The specific extraction, rotation, and sorting methods can refer to the methods in the above embodiments, which will not be elaborated in this embodiment. When training the original recognition model, the original recognition model outputs the probability of which rotation and arrangement order the target cubes in the target training samples have. The above probability may or may not satisfy the first objective function. The first objective function can be a loss function. If the above probability satisfies the first objective function, it indicates that the recognition result of the original recognition model is correct. If the above probability does not satisfy the first objective function, it indicates that the recognition result of the original recognition model is incorrect. When the probability that the recognition result satisfies the first objective function is greater than the first threshold, the current original recognition model is determined as a trained and mature model.

[0136] Through the above method, the training efficiency of the original recognition model can be greatly improved.

[0137] Optionally, after training the mature original recognition model, the convolutional blocks of the original recognition model can be extracted, and after adding new fully connected layers, a new recognition model can be formed, and the new recognition model can be used to recognize other people. The new recognition model can have a high recognition accuracy after being trained with a small number of samples. For example, applying the new recognition model to the process of recognizing the type of 3D pictures, or applying the new recognition model to tasks such as the segmentation of 3D pictures, which will not be elaborated here.

[0138] According to another aspect of the embodiments of the present invention, there is also provided an image recognition device for implementing the above image recognition method. As Figure 11 shown, the device includes:

[0139] (1) A first acquisition unit 1102, configured to acquire a target 3D picture to be recognized;

[0140] (2) A first input unit 1104, configured to input the target 3D picture to be recognized into a first recognition model, where the first recognition model is used to recognize the target 3D picture to be recognized to obtain the picture type of the target 3D picture to be recognized, the convolutional blocks of the first recognition model are the same as those of the second recognition model, and the second recognition model is a model obtained by training the original recognition model with target training samples, and the target training samples include the cubes obtained after rotating and sorting N target cubes acquired from a 3D sample picture, and N is a natural number greater than 1;

[0141] (3) The second acquisition unit 1106 is configured to acquire the first type of the target 3D picture to be recognized output by the first recognition model.

[0142] Optionally, the above picture recognition device can be applied to, but is not limited to, the field of picture recognition. For example, the above method is applied to the process of recognizing the type of 3D pictures. Such as in the process of recognizing the type of diseases in 3D disease pictures. For illustration, when recognizing the type of cerebral hemorrhage, after obtaining a 3D disease picture, the 3D disease picture is input into the first recognition model, and the first model is used to recognize the 3D disease picture and output the first type of the 3D disease picture. For example, the first type can be healthy, or aneurysm, arteriovenous malformation, moyamoya disease, hypertension, etc.

[0143] In the above method, since the second recognition model is trained in advance using the cubes extracted from 3D pictures, the training efficiency of the second recognition model is improved. Further, the convolutional blocks of the second recognition model are used as the convolutional blocks of the first recognition model, and the first recognition model is used to recognize 3D pictures, achieving the effect of greatly improving the training efficiency of the first recognition model.

[0144] Optionally, in the above method, before obtaining the target 3D picture, the second recognition model needs to be trained first. During training, first, 3D sample pictures need to be obtained. The 3D sample pictures are pictures without label annotations. After obtaining the 3D sample pictures, the original cubes need to be extracted from the 3D sample pictures, and the original cubes are split into N target cubes.

[0145] Optionally, when extracting the original cubes, the geometric center of the 3D sample picture can be determined first. After determining the geometric center, using this geometric center as the geometric center of the above original cube, and determining the original cube. The side length of the above original cube is less than the length of the smallest side of the 3D sample picture.

[0146] For example, as Figure 3 shown, for a 3D sample picture 302, first determine the geometric center 304 of the 3D sample picture 302, and then determine the original cube 306 with the geometric center 304 as the geometric center.

[0147] Optionally, after determining the geometric center of the 3D sample picture, a radius r can also be determined, and then a sphere is made with the geometric center of the 3D sample picture as the center and the radius r as the radius, and then any point in the sphere is selected as the geometric center of the above original cube to determine the above original cube. It should be noted that the determined original cube is located within the 3D sample picture and does not exceed the range of the 3D sample picture.

[0148] Optionally, after determining the original cube, the original cube needs to be split to obtain N target cubes. When splitting, any method can be used, such as randomly digging out N target cubes from the original cube, or splitting a part of the original cube to obtain N target cubes. Or, the original cube is evenly split into N target cubes, where N is the cube of a positive integer. Taking N as 8 as an example, as Figure 4 shown, split an original cube 404 along the directions indicated by the arrows of 402-1, 402-2, and 402-3 to obtain 8 target cubes ( Figure 4 The splitting method in is only an example). Or, when splitting, there is a gap of M voxels between every two adjacent cubes. For example, taking M as 2 as an example, as Figure 5 shown, split the original cube 502 into 8 target cubes 504. The side length of the original cube 502 is 10 voxels, then the side length of the target cube 504 is 4 voxels.

[0149] Optionally, after obtaining N target cubes, the first target cube among the N target cubes can also be rotated by a first angle, such as rotating 90 degrees, rotating 180 degrees, etc. There can be one or more first target cubes, and the rotation angle of each first target cube can be the same or different. Sort the rotated first target cubes and the remaining unrotated target cubes. The sorting can be random sorting, and after sorting, the target training samples are obtained.

[0150] After obtaining the target training samples, use the target training samples to train the original recognition model, and the original recognition model outputs the probability of which rotation and arrangement order the target cubes in the target training samples have. The above probability may satisfy the first objective function or may not satisfy the first objective function. The first objective function can be a loss function. If the above probability satisfies the first objective function, it means that the recognition result of the original recognition model is correct. If the above probability does not satisfy the first objective function, it means that the recognition result of the original recognition model is incorrect. When the probability that the recognition result satisfies the first objective function is greater than the first threshold, determine the original recognition model as the second recognition model. It shows that the accuracy of the second recognition model is greater than the first threshold. For example, the accuracy reaches more than 99.95%.

[0151] Using the above training method greatly improves the efficiency of training the second recognition model.

[0152] Optionally, after the second recognition model is trained, the convolutional block in the second recognition model can be obtained, and the convolutional block can be used as the convolutional block of the first recognition model, and the first recognition model can be trained using the first training sample. The first training sample is a 3D picture including the picture type. After the recognition accuracy of the first recognition model is greater than the second threshold, the first recognition model can be put into use. Such as recognizing the disease type of a 3D picture. As Figure 6 shown, a selection button 602-1 is displayed on the display interface 602 of the terminal. The user can select the target 3D picture 604 to be recognized. The terminal recognizes the target 3D picture 604 to be recognized and outputs the first type 606 of the target 3D picture to be recognized.

[0153] Through this embodiment, since the cube extracted from the 3D picture is used to train the second recognition model in advance, the training efficiency of the second recognition model is improved. Further, the convolutional block of the second recognition model is used as the convolutional block of the first recognition model, and the first recognition model is used to recognize the 3D picture, achieving the effect of greatly improving the training efficiency of the first recognition model.

[0154] As an optional implementation, the device further includes:

[0155] (1) A third acquisition unit, configured to acquire the 3D sample picture before acquiring the target 3D picture to be recognized;

[0156] (2) A first determination unit, configured to determine the original cube from the 3D sample picture;

[0157] (3) A splitting unit, configured to split the original cube into the N target cubes.

[0158] Optionally, in this solution, the 3D sample picture and the target 3D picture can be the same picture. That is, after training the second recognition model using the 3D sample picture and using the second convolutional block as the convolutional block of the first recognition model, the 3D sample picture can be input into the first recognition model, and the first recognition model can recognize the type of the 3D sample picture. When the 3D sample picture is input into the second recognition model, the type of the 3D sample picture does not need to be input.

[0159] Through this embodiment, by the above method, N target cubes are obtained to train the second recognition model before using the first recognition model, improving the training efficiency of training the second recognition model and further improving the training efficiency of the first recognition model.

[0160] As an optional implementation, the N is the cube of a positive integer greater than 1, and the splitting unit includes:

[0161] (1) The splitting module is used to keep a space of M voxels between two adjacent target cubes, and split the N target cubes from the original cube, where M is a positive integer greater than 0 and less than J - 1, and J is the side length of the target cube.

[0162] Optionally, when determining the N target cubes, having a space of M voxels between two adjacent target cubes can enable the second recognition model to learn high-level semantic feature information rather than low-level statistical feature information of pixel distribution, improving the training efficiency of the second recognition model and further improving the training efficiency of the first recognition model.

[0163] As an alternative implementation, the device further includes:

[0164] (1) The second determination unit is used to determine a first target cube from the N target cubes before obtaining the target 3D picture to be recognized.

[0165] (2) The rotation unit is used to rotate the first target cube by a first angle.

[0166] (3) The sorting unit is used to sort the first target cube after rotating the first angle and the other target cubes among the N target cubes to obtain the target training sample.

[0167] Optionally, the above sorting can randomly sort the N target cubes. The above rotation can rotate multiple first target cubes among the N target cubes. The rotation can be at any angle.

[0168] Through this embodiment, by the above method, before using the first recognition model, after obtaining the N target cubes, rotating the first target cube among the N target cubes improves the training efficiency of training the second recognition model and further improves the training efficiency of the first recognition model.

[0169] As an alternative implementation, the device further includes:

[0170] (1) The second input unit is used to input the target training sample into the original recognition model after sorting the first target cube after rotating the first angle and the other target cubes among the N target cubes to obtain the target training sample, so as to train the original recognition model to obtain the second recognition model.

[0171] Through this embodiment, by the above method, the training efficiency of training the second recognition model is improved, and the training efficiency of the first recognition model is further improved.

[0172] As an alternative embodiment, the device further comprises:

[0173] (1) A fourth acquisition unit, configured to, before acquiring the target 3D picture to be recognized, acquire the recognition result output after the original recognition model recognizes the target training sample, wherein the recognition result includes various sorting orders of the target cube in the target training sample and the probability of the rotation angle of each target cube;

[0174] (2) A third determination unit, configured to determine the original recognition model as the second recognition model when the probability that the recognition result satisfies the first objective function is greater than the first threshold.

[0175] Optionally, the training of the second recognition model cannot continue indefinitely. When the recognition accuracy of the second recognition model is greater than a certain value, it is considered that the second recognition model meets the requirements, and thus the training is stopped.

[0176] Through this embodiment, by setting a selection condition to stop the training of the second recognition model, the training efficiency of the second recognition model is improved.

[0177] As an alternative embodiment, the device further comprises:

[0178] (1) A fourth determination unit, configured to, before acquiring the target 3D picture to be recognized, determine the convolutional blocks of the second recognition model as the convolutional blocks of the first recognition model;

[0179] (2) A training unit, configured to train the first recognition model using the first training sample until the accuracy of the first recognition model is greater than the second threshold, wherein the first training sample includes the first 3D picture and the type of the first 3D picture.

[0180] Optionally, when training the first recognition model, a first sample picture with a label can be input, and then the first recognition model is trained until the recognition accuracy of the first recognition model is greater than the second threshold, and then the first recognition model can be put into use.

[0181] Through this embodiment, by training the first recognition model before using the first recognition model, the training efficiency of the first recognition model is improved.

[0182] According to another aspect of the embodiments of the present invention, there is also provided an identification model training device for implementing the above identification model training method. As Figure 12 shown, the device comprises:

[0183] (1) A segmentation unit 1202 for obtaining a 3D sample picture and segmenting N target cubes from the 3D sample picture;

[0184] (2) A processing unit 1204 for performing a predetermined operation on the N target cubes to obtain a target training sample, where the predetermined operation includes rotating and sorting the N target cubes;

[0185] (3) A training unit 1206 for training an original recognition model using the target training sample to obtain a second recognition model, where the original recognition model is used to output a recognition result of the target training sample, and when the probability that the recognition result satisfies a first objective function is greater than a first threshold, the original recognition model is determined as the second recognition model.

[0186] Optionally, the above device can be but is not limited to being applied in the process of model training. When training the original recognition model, N target cubes are extracted from a 3D sample picture, and the N cubes obtained after rotating and sorting the N target cubes are used as the target training sample and input into the original recognition model. The specific extraction, rotation, and sorting methods can refer to the methods in the above embodiments, which will not be elaborated in this embodiment. When training the original recognition model, the original recognition model outputs the probability of which rotation and arrangement order the target cubes in the target training sample have. The above probability may or may not satisfy the first objective function. The first objective function can be a loss function. If the above probability satisfies the first objective function, it means that the recognition result of the original recognition model is correct. If the above probability does not satisfy the first objective function, it means that the recognition result of the original recognition model is incorrect. When the probability that the recognition result satisfies the first objective function is greater than the first threshold, the current original recognition model is determined as a trained and mature model.

[0187] Through the above method, the training efficiency of the original recognition model can be greatly improved.

[0188] Optionally, after training a mature original recognition model, the convolutional blocks of the original recognition model can be extracted, and after adding new fully connected layers, a new recognition model is formed, and the new recognition model can be used to recognize other people. The new recognition model can have a high recognition accuracy after being trained with a small number of samples. For example, applying the new recognition model to the process of recognizing the type of 3D pictures, or applying the new recognition model to tasks such as segmentation of 3D pictures, which will not be elaborated here.

[0189] According to another aspect of the embodiments of the present invention, a model training method is further provided. Optionally, as Figure 13 shown, the above model training method includes:

[0190] S1302. Convert the original three-dimensional image data into three-dimensional cube training samples, where the three-dimensional cube training samples include a plurality of micro-cubes;

[0191] S1304. Perform a first operation and a second operation on the plurality of micro-cubes in sequence to obtain target training samples, where the first operation is used to change the order of the plurality of micro-cubes, and the second operation is used to change the direction of a first object micro-cube among the plurality of micro-cubes;

[0192] S1306. Use the target training samples for training to obtain a pre-trained network model, where the pre-trained network model is used to extract features from the original three-dimensional image data and is also used to identify the data structure in the original three-dimensional image data;

[0193] S1308. Transfer a target fully connected layer matching the target image recognition task to the pre-trained network model to obtain a first recognition model;

[0194] S1310. Input the target three-dimensional image data to be recognized into the first recognition model to obtain a recognition result, where the recognition result includes the abnormal area in the target three-dimensional image data. Optionally, the above-mentioned original three-dimensional data can be but is not limited to the 3D sample image in this solution, the above three-dimensional cube training samples can be but is not limited to the original cube in this solution, the above plurality of micro-cubes can be but is not limited to the N target cubes in this solution, the above first operation can be but is not limited to a sorting operation, the above second operation can be but is not limited to a rotation operation, the above first object micro-cube can be but is not limited to the first target cube in this solution, the above pre-trained network model can be but is not limited to the second recognition model in this solution, and replace the target fully connected layer of the second recognition model to obtain the first recognition model in this solution. The above target three-dimensional image data can be but is not limited to the target 3D image in this solution.

[0195] In this solution, after obtaining the original three-dimensional image data, convert the data into three-dimensional cube training samples, then change the order of the plurality of micro-cubes and the direction of the first object micro-cube in the three-dimensional cube training samples, and use the adjusted data as target training samples to obtain a training network model. At this time, the training efficiency of the training network model is improved. Replace the fully connected layer of the training network model to obtain a first recognition model, and the first recognition model can be used to identify the abnormal area in the three-dimensional image data.

[0196] As an optional implementation, the performing of the first operation and the second operation on the plurality of micro-cubes includes:

[0197] Perform permutations and combinations on the multiple micro-cubes to obtain K combinations of micro-cubes;

[0198] Determine a target micro-cube combination from the K combinations of micro-cubes;

[0199] Determine the first object micro-cube from the target micro-cube combination;

[0200] Perform a rotation operation on the first object micro-cube to obtain the target training sample.

[0201] Optionally, the above-mentioned performing permutations and combinations on multiple micro-cubes means performing permutations and combinations on N target cubes. Performing a rotation operation on the first object micro-cube means performing a rotation operation on the first target cube.

[0202] As an optional implementation, after successively performing a first operation and a second operation on the multiple micro-cubes, it further includes:

[0203] Determine a second object micro-cube from the multiple micro-cubes;

[0204] Perform a third operation on the second object micro-cube to update the target training sample, where the third operation is used to occlude a partial area of the second object micro-cube.

[0205] Optionally, after performing the second operation on the first object micro-cube, a third operation can also be performed on the second object micro-cube among the multiple micro-cubes. The third operation can be an occlusion operation. For example, for the second object micro-cube, generate a 3D matrix (Ran) with the same magic cube size as the second object micro-cube, and then multiply the second object micro-cube by the 3D matrix to obtain a new small cube. The Ran matrix is filled with values 0 or 1. This step can be regarded as randomly covering some areas in a cell. The new small cube is the occluded second object micro-cube.

[0206] As an optional implementation, the performing the third operation on the second object micro-cube includes:

[0207] Multiply the second object micro-cube by a target matrix, where the target matrix is a three-dimensional matrix with the same size as the second object micro-cube.

[0208] Optionally, the target matrix is the 3D matrix with the same magic cube size as the above-mentioned second object micro-cube.

[0209] As an optional implementation, the converting the original three-dimensional picture data into a three-dimensional cube training sample includes:

[0210] Convert the original three-dimensional picture data into an original three-dimensional cube;

[0211] Split the original three-dimensional cube into a plurality of original micro-cubes;

[0212] Extract the plurality of micro-cubes from the plurality of original micro-cubes.

[0213] As an alternative embodiment, the splitting the original three-dimensional cube into a plurality of original micro-cubes includes:

[0214] Split the original three-dimensional cube to obtain the plurality of original micro-cubes, wherein there is a gap of M voxels between two adjacent original micro-cubes, and M is a positive integer greater than 0 and less than J-1, and J is the side length of the original micro-cube.

[0215] As an alternative embodiment, before training with the target training sample to obtain a pre-trained network model, it further includes: constructing an original network model of the pre-trained network model, wherein the objective function in the fully connected layer of the original network model includes: a first loss function corresponding to the first operation, a second loss function corresponding to the second operation, and a third loss function corresponding to the third operation;

[0216] The training with the target training sample to obtain a pre-trained network model includes: inputting the target training sample into the original network model for training to obtain the pre-trained network model, wherein the output result of the objective function in the fully connected layer of the pre-trained network model has reached the convergence condition.

[0217] The above original network model can be the original recognition model in this solution.

[0218] As an alternative embodiment, the migrating a target fully connected layer matching the target picture recognition task for the pre-trained network model to obtain a first recognition model includes:

[0219] Obtain the target picture recognition task to be processed currently;

[0220] Determine a target fully connected layer matching the target picture recognition task;

[0221] Replace the fully connected layer of the pre-trained network model with the target fully connected layer to obtain the first recognition model, wherein the first recognition model is used to execute the target picture recognition task.

[0222] The following is illustrated with a specific example.

[0223] First, preprocess the data. During data preprocessing, in order to capture 3D voxel information and understand the internal characteristics of 3D medical imaging data. In this case, the texture information near the cube boundary may interfere with network training. Therefore, to avoid interference, this information can be skipped or excluded. Therefore, during the process of cutting the magic cube, after leaving a gap between two adjacent magic cubes (target cubes), perform [-1,1] normalization on the voxels within each magic cube.

[0224] The network structure in this solution can be as Figure 14 shown. The Siamese network includes M sub-networks that share weights with each other, where M represents the number of magic cubes. In the 2×2×2 magic cube division setting, an eight-in-one Siamese network with 8 magic cube inputs is used, and in the 3×3×3 magic cube division setting, an eight-in-one Siamese network with 27 magic cube inputs is used. Each sub-network has the same network structure and shares weights with each other. Figure 14 The Siamese network in

[0225] includes 8 sub-networks. The backbone structure of each sub-network can use a 3D CNN, or 3D Resnet or 3D VGG network. Stack the outputs of the last fully connected layer of all sub-networks and then input them into different branches, which are used for the tasks of magic cube rearrangement, rotation judgment, and occlusion judgment respectively.

[0226] Rotation of the magic cube: Since there are 3 (axes) × 2 (directions) × 4 (angles) = 24 free rotation methods during the rearrangement operation. To reduce the complexity of the task, only two types of rotations are allowed, namely 180° horizontal and vertical cube rotations. During this process, randomly select the small cube to be rotated and the rotation direction. For example, as Figure 14 shown, small cubes 6 and 7 are horizontally rotated, and small cube 4 is vertically rotated. In order to be able to orient the rotation of the small cubes, the network needs to discover and identify whether each small cube rotates and how it rotates. This task can be regarded as a multi-label classification task. A 1×M (M is the number of small cubes) vector is used to represent one rotation category, and the corresponding position of the rotated cube is 1, otherwise it is 0. Therefore, the prediction task can be described by two 1×M vectors (r), which represent the horizontal and vertical rotation probabilities of each cube respectively. The rotation loss function can be:

[0227]

[0228] In the formula, M represents the number of target magic cubes, g represents the label information. r represents the rotation vector, r h represents horizontal rotation, r v represents vertical rotation, gih 、ri h 、gi v 、ri v are 1×M dimensional vectors with values of 0 or 1, and i is a positive integer.

[0229] Occlusion of the magic cube: After the cube is rotated, in the magic cube restoration task, a data augmentation method for self-supervised learning is introduced. Specifically, a magic cube is selected and a 3D matrix (Ran) with the same magic cube size is generated, and then they are multiplied to obtain a new small cube (the second cube). The Ran matrix is filled with values of 0 or 1. This step can be regarded as randomly covering some areas in a cell.

[0230] Since only a part of a small cube is covered in this step, there is an obvious difference between the selected small cube and other small cubes. Therefore, it is also expected that the network can identify which small cube is partially occluded. This problem can be regarded as a classification task with M labels. The prediction can be described as a 1×M vector. Then the occlusion loss can be defined as:

[0231]

[0232] where lc represents the one-hot label of the cube occlusion, C is the covering occlusion vector, and i is a positive integer. The occlusion operation can force the network to capture more detailed information and learn more about the local area.

[0233] According to the previous definition, the objective function of the model is the linear weighted sum of the permutation loss function and the rotation loss function. The overall loss function of the model is as follows:

[0234] loss = a * loss1 + b * loss2 + c * loss3(6)

[0235] where a, b, and c are the weights of the three loss functions respectively, which control the degree of mutual influence between the three subtasks. In the experiment, the three weight values are set to 1:1:1, which can make the pre-training achieve better results.

[0236] After the above training, the second recognition model can be obtained. The accuracy of the second recognition model is greater than the first threshold.

[0237] At this time, the convolutional blocks of the second recognition model can be extracted and fine-tuned for use in other target tasks.

[0238] The network pre-trained on the magic cube restoration task can capture the hidden features of 3D medical images, identify the basic structure of 3D medical image data, and obtain a powerful feature representation. After some iterations on the proxy task, the network can be transferred to the target task.

[0239] For the 3D medical image classification task, transfer the pre-trained network except for the final fully-connected layer and add another new fully-connected layer. Then, the network can be fine-tuned in the target classification task. For the segmentation task of 3D medical image data, the weights of the pre-trained network can only be transferred to the encoder part (downsampling stage) of the fully convolutional neural network (FCN), such as 3D U-Net. The decoder part (upsampling stage) of the fully convolutional neural network still needs to be randomly initialized. After pre-training the encoder part, the network can capture useful information, which can provide the gradient direction for the decoder part. Experiments show that this pre-training task can improve the segmentation performance compared with training from scratch.

[0240] Since the previous Rubik's Cube restoration-style pre-training can only be applied to the downsampling stage of U-Net, the network parameters of the upsampling stage of U-Net still need to be randomly initialized during training. To avoid the impact of a large number of parameter initializations on the previous pre-training effect, the Dense Upsampling Convolution (DUC) module is used to replace the original transposed convolution to upsample the feature map and restore it to the original input size of the image. The structure of the DUC module is as Figure 9 shown. Among them, C represents the number of channels, d represents the expansion factor. H is the length of the feature map, and W is the width of the feature map.

[0241] Through this embodiment, the effect of improving the model training efficiency is achieved.

[0242] According to another aspect of the embodiments of the present invention, a model training device is further provided. Optionally, as Figure 15 shown, the above model training device includes:

[0243] (1) A conversion unit 1502, configured to convert the original three-dimensional picture data into three-dimensional cube training samples, where the three-dimensional cube training samples include a plurality of micro-cubes;

[0244] (2) A first execution unit 1504, configured to sequentially perform a first operation and a second operation on the plurality of micro-cubes to obtain target training samples, where the first operation is used to change the order of the plurality of micro-cubes, and the second operation is used to change the direction of the first object micro-cube among the plurality of micro-cubes;

[0245] (3) A training unit 1506, configured to perform training using the target training samples to obtain a pre-trained network model, where the pre-trained network model is used to extract features from the original three-dimensional picture data and is also used to identify the data structure in the original three-dimensional picture data;

[0246] (4) A migration unit 1508, configured to migrate a target fully-connected layer matching the target picture recognition task for the pre-trained network model, so as to obtain a first recognition model;

[0247] (5) An input unit 1510, configured to input target three-dimensional picture data to be recognized into the first recognition model, so as to obtain a recognition result, where the recognition result includes an abnormal area in the target three-dimensional picture data.

[0248] Optionally, the above-mentioned original three-dimensional data may but is not limited to the 3D sample pictures in this solution, the above-mentioned three-dimensional solid training samples may but is not limited to the original cube in this solution, the above-mentioned multiple micro-cubes may but is not limited to the N target cubes in this solution, the above-mentioned first operation may but is not limited to a sorting operation, the above-mentioned second operation may but is not limited to a rotation operation, the above-mentioned first object micro-cube may but is not limited to the first target cube in this solution, the above-mentioned pre-trained network model may but is not limited to the second recognition model in this solution, and by replacing the target fully-connected layer for the second recognition model, the first recognition model in this solution is obtained. The above-mentioned target three-dimensional picture data may but is not limited to the target 3D pictures in this solution.

[0249] In this solution, after obtaining the original three-dimensional picture data, the data is converted into a three-micro-cube training sample, and then the order of multiple micro-cubes in the changed three-dimensional cube training sample and the direction of the first object micro-cube are adjusted, and the adjusted data is used as the target training sample to obtain a training network model. At this time, the training efficiency of the training network model is improved. Replace the fully-connected layer of the training network model to obtain a first recognition model, and the first recognition model can be used to recognize abnormal areas in three-dimensional picture data.

[0250] According to another aspect of the embodiments of the present invention, an electronic device for implementing the above-mentioned picture recognition method is further provided. As Figure 16 shown, the electronic device includes a memory 1602 and a processor 1604. A computer program is stored in the memory 1602, and the processor 1604 is configured to execute the steps in any one of the above method embodiments through the computer program.

[0251] Optionally, in this embodiment, the above-mentioned electronic device may be at least one network device among multiple network devices in a computer network.

[0252] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the following steps through the computer program:

[0253] S1, obtain a target 3D picture to be recognized;

[0254] S2. Input the target 3D picture to be recognized into the first recognition model, where the first recognition model is used to recognize the target 3D picture to be recognized to obtain the picture type of the target 3D picture to be recognized. The convolutional blocks of the first recognition model are the same as those of the second recognition model. The second recognition model is a model obtained by training the original recognition model with target training samples, and the target training samples include cubes obtained by rotating and sorting N target cubes obtained from 3D sample pictures, where N is a natural number greater than 1;

[0255] S3. Obtain the first type of the target 3D picture output by the first recognition model.

[0256] Optionally, those of ordinary skill in the art can understand that Figure 16 The structure shown is only schematic. The electronic device can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 16 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, etc.) than those shown in Figure 16 or have a different configuration from that shown in Figure 16 Shown.

[0257] Among them, the memory 1602 can be used to store software programs and modules, such as the program instructions / modules corresponding to the picture recognition method and device in the embodiments of the present invention. The processor 1604 executes various functional applications and data processing by running the software programs and modules stored in the memory 1602, that is, implements the above-mentioned picture recognition method. The memory 1602 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, a flash memory, or other non-volatile solid-state memories. In some instances, the memory 1602 may further include a memory remotely set relative to the processor 1604, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. Among them, the memory 1602 can specifically but not limitedly be used to store information such as the target 3D picture to be recognized. As an example, as Figure 16 Shown, the above-mentioned memory 1602 may include but not limited to the first acquisition unit 1102, the first input unit 1104, and the second acquisition unit 1106 in the above-mentioned picture recognition device. In addition, it may further include but not limited to other module units in the above-mentioned picture recognition device, which will not be elaborated in this example.

[0258] Optionally, the above-mentioned transmission device 1606 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired network and a wireless network. In one example, the transmission device 1606 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through network cables, so as to communicate with the Internet or a local area network. In one example, the transmission device 1606 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0259] In addition, the above-mentioned electronic device further includes: a display 1608, which is used to display the first type of 3D picture to be recognized; and a connection bus 1610, which is used to connect each module component in the above-mentioned electronic device.

[0260] According to another aspect of the embodiments of the present invention, an electronic device for implementing the above-mentioned recognition model training method is further provided. As Figure 17 shown, the electronic device includes a memory 1702 and a processor 1704. A computer program is stored in the memory 1702, and the processor 1704 is configured to execute the steps in any one of the above method embodiments through the computer program.

[0261] Optionally, in this embodiment, the above-mentioned electronic device may be located in at least one network device among multiple network devices of a computer network.

[0262] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the following steps through a computer program:

[0263] S1. Obtain a 3D sample picture, and segment N target cubes from the 3D sample picture;

[0264] S2. Perform a predetermined operation on the N target cubes to obtain a target training sample, where the predetermined operation includes rotating and sorting the N target cubes;

[0265] S3. Use the target training sample to train the original recognition model to obtain a second recognition model, where the original recognition model is used to output the recognition result of the target training sample. When the probability that the recognition result satisfies the first objective function is greater than the first threshold, the original recognition model is determined as the second recognition model.

[0266] Optionally, those of ordinary skill in the art can understand that Figure 17The structure shown is only illustrative. The electronic device can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a personal digital assistant, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 17 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, etc.) than those shown Figure 17 in the figure, or have a different configuration from that Figure 17 shown in the figure.

[0267] Among them, the memory 1702 can be used to store software programs and modules, such as the program instructions / modules corresponding to the recognition model training method and device in the embodiments of the present invention. The processor 1704 executes various functional applications and data processing by running the software programs and modules stored in the memory 1702, that is, implements the above-mentioned recognition model training method. The memory 1702 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 1702 may further include a memory remotely disposed relative to the processor 1704, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. Among them, the memory 1702 can specifically but not limitedly be used to store information such as 3D sample pictures. As an example, as Figure 17 shown, the above-mentioned memory 1702 may include, but is not limited to, the segmentation unit 1202, the processing unit 1204, and the training unit 1206 in the above-mentioned recognition model training device. In addition, it may further include, but is not limited to, other module units in the above-mentioned recognition model training device, which will not be elaborated in this example.

[0268] Optionally, the above-mentioned transmission device 1706 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired network and a wireless network. In one instance, the transmission device 1706 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and a router through a network cable, so as to communicate with the Internet or a local area network. In one instance, the transmission device 1706 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0269] In addition, the above-mentioned electronic device further includes: a display 1708, which is used to display the training accuracy of the original recognition model, etc.; and a connection bus 1710, which is used to connect each module component in the above-mentioned electronic device.

[0270] According to another aspect of the embodiments of the present invention, an electronic device for implementing the above-mentioned recognition model training method is further provided. As Figure 18 shown, the electronic device includes a memory 1802 and a processor 1804. A computer program is stored in the memory 1802, and the processor 1804 is configured to execute the steps in any of the above method embodiments through the computer program.

[0271] Optionally, in this embodiment, the above-mentioned electronic device may be at least one network device among multiple network devices in a computer network.

[0272] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the following steps through a computer program:

[0273] S1, obtain 3D sample pictures, and segment N target cubes from the 3D sample pictures;

[0274] S2, perform a predetermined operation on the N target cubes to obtain a target training sample, where the predetermined operation includes rotating and sorting the N target cubes;

[0275] S3, use the target training sample to train the original recognition model to obtain a second recognition model, where the original recognition model is used to output the recognition result of the target training sample. When the probability that the recognition result satisfies the first objective function is greater than the first threshold, the original recognition model is determined as the second recognition model.

[0276] Optionally, those of ordinary skill in the art can understand that Figure 18 the structure shown is only schematic. The electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and a mobile Internet device (MID), a PAD and other terminal devices. Figure 18 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, etc.) than those shown in Figure 18 , or have a different configuration from that shown in Figure 18 .

[0277] Among them, the memory 1802 can be used to store software programs and modules, such as the program instructions / modules corresponding to the recognition model training method and device in the embodiments of the present invention. The processor 1804 executes various functional applications and data processing by running the software programs and modules stored in the memory 1802, that is, implements the above-mentioned recognition model training method. The memory 1802 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 1802 may further include a memory remotely disposed relative to the processor 1804, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof. Among them, the memory 1802 can specifically but not limitedly be used to store information such as 3D sample pictures. As an example, as Figure 18 shown, the above memory 1802 may include but is not limited to the conversion unit 1502, the first execution unit 1504, the training unit 1506, the migration unit 1508, and the input unit 1510 in the above recognition model training device. In addition, it may also include but is not limited to other module units in the above recognition model training device, which will not be elaborated in this example.

[0278] Optionally, the above transmission device 1806 is used to receive or send data via a network. Specific examples of the above network may include a wired network and a wireless network. In one instance, the transmission device 1806 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable so as to communicate with the Internet or a local area network. In one instance, the transmission device 1806 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0279] In addition, the above electronic device further includes: a display 1808, which is used to display the training accuracy of the original recognition model, etc.; and a connection bus 1810, which is used to connect each module component in the above electronic device.

[0280] According to another aspect of the embodiments of the present invention, a storage medium is further provided, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0281] Optionally, in this embodiment, the above storage medium may be configured to store a computer program for executing the following steps:

[0282] S1, obtain a target 3D picture to be recognized;

[0283] S2. Input the target 3D image to be recognized into the first recognition model. The first recognition model is used to recognize the target 3D image to be recognized to obtain the image type of the target 3D image to be recognized. The convolutional blocks of the first recognition model are the same as those of the second recognition model. The second recognition model is a model obtained by training the original recognition model using target training samples. The target training samples include cubes obtained by rotating and sorting N target cubes obtained from 3D sample images, and N is a natural number greater than 1;

[0284] S3. Obtain the first type of the target 3D image to be recognized output by the first recognition model.

[0285] Alternatively, optionally, in this embodiment, the above storage medium may be set to store a computer program for executing the following steps:

[0286] S1. Obtain 3D sample images and segment N target cubes from the 3D sample images;

[0287] S2. Perform a predetermined operation on the N target cubes to obtain target training samples. The predetermined operation includes rotating and sorting the N target cubes;

[0288] S3. Use the target training samples to train the original recognition model to obtain a second recognition model. The original recognition model is used to output the recognition result of the target training samples. When the probability that the recognition result satisfies the first objective function is greater than the first threshold, the original recognition model is determined as the second recognition model.

[0289] Alternatively, optionally, in this embodiment, the above storage medium may be set to store a computer program for executing the following steps:

[0290] S1. Convert the original three-dimensional image data into three-dimensional cube training samples, where the three-dimensional cube training samples include a plurality of micro-cubes;

[0291] S2. Perform a first operation and a second operation on the plurality of micro-cubes in sequence to obtain target training samples. The first operation is used to change the order of the plurality of micro-cubes, and the second operation is used to change the direction of the first object micro-cube among the plurality of micro-cubes;

[0292] S3. Use the target training samples for training to obtain a pre-trained network model. The pre-trained network model is used to extract the features in the original three-dimensional image data and is also used to recognize the data structure in the original three-dimensional image data;

[0293] S4. Transfer a target fully-connected layer matching the target picture recognition task to the pre-trained network model to obtain a first recognition model;

[0294] S5. Input the target three-dimensional picture data to be recognized into the first recognition model to obtain a recognition result, where the recognition result includes the abnormal area in the target three-dimensional picture data.

[0295] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the above-mentioned various methods can be completed by instructing the relevant hardware of the terminal device through a program, and this program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0296] The serial numbers of the above-mentioned embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.

[0297] If the integrated unit in the above-mentioned embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above-mentioned computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing one or more computer devices (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0298] In the above-mentioned embodiments of the present invention, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0299] In the several embodiments provided by the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the units or modules can be in an electrical or other form.

[0300] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0301] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.

[0302] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A method for image recognition, characterized in that, Including: Determine a first target cube from N target cubes, where N is a natural number greater than 1; rotate the first target cube by a first angle; sort the first target cube after rotating the first angle and other target cubes among the N target cubes to obtain a target training sample; Obtain a target 3D picture to be recognized; Input the target 3D picture to be recognized into a first recognition model, where the first recognition model is used to recognize the target 3D picture to be recognized to obtain the picture type of the target 3D picture to be recognized, the convolutional blocks of the first recognition model are the same as those of a second recognition model, and the second recognition model is a model obtained by training an original recognition model using the target training sample; Obtain a first type of the target 3D picture output by the first recognition model.

2. The method according to claim 1, characterized in that Before obtaining the target 3D picture to be recognized, it further includes: Obtain 3D sample pictures; Determine an original cube from the 3D sample pictures; Split the original cube into the N target cubes.

3. The method according to claim 2, wherein N is the cube of a positive integer greater than 1, and the splitting of the original cube into the N target cubes includes: keeping a space of M voxels between two adjacent target cubes, and splitting the N target cubes from the original cube, where M is a positive integer greater than 0 and less than J - 1, and J is the side length of the target cube.

4. The method according to claim 1, characterized in that, After sorting the first target cube after rotating the first angle and other target cubes among the N target cubes to obtain the target training sample, it further includes: Input the target training sample into the original recognition model to train the original recognition model to obtain the second recognition model.

5. The method according to claim 1, wherein Before obtaining the target 3D picture to be recognized, it further includes: Obtain a recognition result output by the original recognition model after recognizing the target training sample, where the recognition result includes various sorting orders of the target cubes in the target training sample and the probabilities of the rotation angles of each target cube; When the probability that the recognition result satisfies a first objective function is greater than a first threshold, determine the original recognition model as the second recognition model.

6. The method according to claim 1, wherein Before obtaining the target 3D picture to be recognized, it further includes: Determine the convolutional blocks of the second recognition model as the convolutional blocks of the first recognition model; train the first recognition model using a first training sample until the accuracy of the first recognition model is greater than a second threshold, where the first training sample includes a first 3D picture and the type of the first 3D picture.

7. A method for training an identification model, characterized in that, Including: Obtain 3D sample pictures, and segment N target cubes from the 3D sample pictures; Determine a first target cube from N target cubes, where N is a natural number greater than 1; rotate the first target cube by a first angle; sort the first target cube after rotating the first angle and other target cubes among the N target cubes to obtain a target training sample; Use the target training sample to train an original recognition model to obtain a second recognition model, where the original recognition model is used to output a recognition result of the target training sample, and when the probability that the recognition result satisfies a first objective function is greater than a first threshold, determine the original recognition model as the second recognition model.

8. An image recognition device, characterized in that, Comprising: A first acquisition unit, configured to determine a first target cube from N target cubes, where N is a natural number greater than 1; rotate the first target cube by a first angle; sort the first target cube after rotating the first angle and other target cubes among the N target cubes to obtain a target training sample; acquire a target 3D image to be recognized; A first input unit, configured to input the target 3D image to be recognized into a first recognition model, where the first recognition model is used to recognize the target 3D image to be recognized to obtain the image type of the target 3D image to be recognized, and the convolutional blocks of the first recognition model are the same as those of the second recognition model, and the second recognition model is a model obtained by training the original recognition model with the target training sample; A second acquisition unit, configured to acquire a first type of the target 3D image output by the first recognition model.

9. An identification model training device, characterized in that, Comprising: A segmentation unit, configured to acquire a 3D sample image and segment N target cubes from the 3D sample image; A processing unit, configured to determine a first target cube from N target cubes, where N is a natural number greater than 1; rotate the first target cube by a first angle; sort the first target cube after rotating the first angle and other target cubes among the N target cubes to obtain a target training sample; A training unit, configured to use the target training sample to train an original recognition model to obtain a second recognition model, where the original recognition model is used to output a recognition result of the target training sample, and when the probability that the recognition result satisfies a first objective function is greater than a first threshold, determine the original recognition model as the second recognition model.

10. A model training method, characterized in that, Comprising: Convert the original three-dimensional image data into a three-dimensional cube training sample, where the three-dimensional cube training sample includes a plurality of micro-cubes; Perform a first operation and a second operation on the plurality of micro-cubes in sequence to obtain a target training sample, where the first operation is used to change the order of the plurality of micro-cubes, and the second operation is used to change the direction of a first object micro-cube among the plurality of micro-cubes; Train using the target training samples to obtain a pre-trained network model, where the pre-trained network model is used to extract features from the original three-dimensional picture data and is also used to identify the data structure in the original three-dimensional picture data; Migrate a target fully connected layer that matches the target picture recognition task for the pre-trained network model to obtain a first recognition model; Input the target three-dimensional picture data to be recognized into the first recognition model to obtain a recognition result, where the recognition result includes the abnormal area in the target three-dimensional picture data.

11. The method according to claim 10, wherein The performing the first operation and the second operation on the multiple micro-cubes includes: Perform permutations and combinations on the multiple micro-cubes to obtain K types of micro-cube combinations; Determine a target micro-cube combination from the K types of micro-cube combinations; Determine the first object micro-cube from the target micro-cube combination; Perform a rotation operation on the first object micro-cube to obtain the target training sample.

12. The method according to claim 10, wherein After successively performing the first operation and the second operation on the multiple micro-cubes, it further includes: Determine a second object micro-cube from the multiple micro-cubes; Perform a third operation on the second object micro-cube to update the target training sample, where the third operation is used to occlude a part of the area of the second object micro-cube.

13. The method according to claim 12, characterized in that, The performing the third operation on the second object micro-cube includes: Multiply the second object micro-cube by a target matrix, where the target matrix is a three-dimensional matrix of the same size as the second object micro-cube.

14. The method according to claim 10, wherein The converting the original three-dimensional picture data into a three-dimensional cube training sample includes: Convert the original three-dimensional picture data into an original three-dimensional cube; Split the original three-dimensional cube into multiple original micro-cubes; Extract the multiple micro-cubes from the multiple original micro-cubes.

15. The method according to claim 14, wherein The splitting the original three-dimensional cube into multiple original micro-cubes includes: Split the original three-dimensional cube to obtain the multiple original micro-cubes, where there is a gap of M voxels between two adjacent original micro-cubes, M is a positive integer greater than 0 and less than J - 1, and J is the side length of the original micro-cube.

16. The method according to claim 12, wherein Before training using the target training samples to obtain a pre-trained network model, it further includes: constructing an original network model of the pre-trained network model, where the objective function in the fully connected layer of the original network model includes: a first loss function corresponding to the first operation, a second loss function corresponding to the second operation, and a third loss function corresponding to the third operation; The training using the target training samples to obtain a pre-trained network model includes: Input the target training samples into the original network model for training to obtain the pre-trained network model, where the output result of the objective function in the fully connected layer of the pre-trained network model has reached the convergence condition.

17. The method according to claim 16, characterized in that, The above-mentioned is to transfer the target fully-connected layer of the pre-trained network model to match the target image recognition task, and the steps to obtain the first recognition model include: Obtain the current target image recognition task to be processed; Determine the target fully-connected layer that matches the target image recognition task; Replace the fully-connected layer of the pre-trained network model with the target fully-connected layer to obtain the first recognition model, where the first recognition model is used to perform the target image recognition task.

18. A model training device, characterized in that, It includes: A conversion unit for converting the original three-dimensional picture data into three-dimensional cube training samples, where the three-dimensional cube training samples include a plurality of micro-cubes; A first execution unit for sequentially performing a first operation and a second operation on the plurality of micro-cubes to obtain a target training sample, where the first operation is used to change the order of the plurality of micro-cubes, and the second operation is used to change the direction of the first object micro-cube among the plurality of micro-cubes; A training unit for training using the target training sample to obtain a pre-trained network model, where the pre-trained network model is used to extract features from the original three-dimensional picture data and is also used to identify the data structure in the original three-dimensional picture data; A migration unit for migrating the target fully-connected layer that matches the target image recognition task for the pre-trained network model to obtain the first recognition model; An input unit for inputting the target three-dimensional picture data to be recognized into the first recognition model to obtain a recognition result, where the recognition result includes the abnormal area in the target three-dimensional picture data.

19. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method described in any one of claims 1 to 6 or 7 or 10 to 17 through the computer program.

Citation Information

Patent Citations

  • Proximal femur segmentation method and device, computer device and storage medium

    CN108764241A