Three-dimensional image classification method, device, computer equipment and medium

By dividing the three-dimensional image into image blocks and converting it into tensor data, and using encoder and fully connected neural network for classification, the problem that Vision Transformer in the prior art is not able to process three-dimensional images, and the accurate classification of three-dimensional images is achieved.

CN115048974BActive Publication Date: 2025-05-13SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210334686.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-05-13
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

The Vision Transformer in the prior art can only classify two-dimensional images and cannot effectively process and classify three-dimensional nuclear magnetic resonance imaging images.

Method used

By dividing the three-dimensional images to be classified into image blocks and converting these image blocks into tensor data, the encoder unit and fully connected neural network unit are used for classification. Specific steps include image slicing, tensor transformation, encoding and calculation of fully connected neural networks to realize the classification of three-dimensional images.

Benefits of technology

It realizes the accurate classification of three-dimensional images, solves the problem that the existing Vision Transformer cannot handle three-dimensional images, and expands the application scope of image classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115048974B_ABST
    Figure CN115048974B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional image classification method, device, computer equipment and medium. The method comprises the following steps: obtaining an image to be classified, wherein the image to be classified is a three-dimensional image to be classified; segmenting the image to be classified by an image segmentation pair to generate image blocks to be converted; performing tensor conversion on the image blocks to be converted to form a converted image tensor; and calculating the category of the image to be classified according to the converted image tensor by the encoder unit and the fully connected neural network unit. The present application realizes image classification of the three-dimensional image to be classified by a convolutional neural network by segmenting the three-dimensional image to be classified into image blocks to be converted, converting the image blocks to be converted into tensor data, and classifying the image blocks to be converted according to the tensor data. The method solves the problem that the existing Vision Transformer in the prior art can only classify two-dimensional images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of brain age recognition, and in particular to a three-dimensional image classification method, device, computer equipment and medium. Background Art

[0002] Magnetic resonance imaging (MRI) is also called magnetic resonance imaging. This technology uses the principle of nuclear magnetic resonance. Based on the different attenuation of the energy released in different structural environments inside the material, the electromagnetic waves emitted can be detected by applying an external gradient magnetic field to know the position and type of the atomic nuclei that make up the object. Based on this, the internal structure image of the object can be drawn. The application of this technology in clinical diagnosis and scientific research has become a reality, greatly promoting the development of disciplines such as medicine and cognitive neurology.

[0003] At present, most convolutional neural networks for image classification can only recognize two-dimensional images and cannot directly process three-dimensional magnetic resonance imaging. At the same time, most networks for image classification cannot do without convolution operations. In 2017, the article "Attention is All You Need" by the Google Machine Translation Team swept the NLP field with its eye-catching title and amazing results, and neural networks became one of the most commonly used models in the NLP field. Recently, the computer vision field has also tried to use the Transformer model and produced some gratifying results, namely Vision Transformer. However, the Vision Transformer in the existing technology can only be used for the classification of two-dimensional images.

[0004] Therefore, how to achieve accurate image classification through the original Vision Transformer has become an urgent problem to be solved. Summary of the invention

[0005] Based on this, it is necessary to address the above problems and propose a three-dimensional image classification method, device, computer equipment and medium that can realize three-dimensional image classification through Transformer and neural network.

[0006] A three-dimensional image classification method includes a neural network, wherein the neural network includes an encoder unit and a fully connected neural network unit, wherein the encoder unit is connected to the fully connected neural network unit, and the method includes:

[0007] Acquire an image to be classified, wherein the image to be classified is a three-dimensional image to be classified;

[0008] Segmenting the image to be classified by image segmentation to generate image blocks to be converted;

[0009] Performing tensor conversion on the image block to be converted to form a converted image tensor;

[0010] The category of the image to be classified is inferred based on the transformed image tensor through the encoder unit and the fully connected neural network unit.

[0011] Furthermore, the step of performing tensor conversion on the image block to be converted to form a converted image tensor specifically includes:

[0012] Performing an expansion operation on the image block to be converted by a one-dimensional operation of the image block to obtain a one-dimensional expanded tensor of the image block to be converted;

[0013] Compressing the one-dimensional expanded tensor by tensor linear projection to form a compressed one-dimensional tensor;

[0014] A tensor to be added is obtained, and the tensor to be added is added to the compressed one-dimensional tensor, thereby generating the converted image tensor.

[0015] Further, the tensor to be added is obtained including a one-dimensional category tensor and a learning position tensor;

[0016] Then the step of obtaining the tensor to be added and adding the tensor to be added to the compressed one-dimensional tensor to generate the converted image tensor specifically includes:

[0017] Adding the one-dimensional category tensor to the compressed one-dimensional tensor to obtain the compressed one-dimensional tensor having the one-dimensional category tensor, and recording it as the category added tensor;

[0018] The learned position tensor is added to the category-added tensor to obtain the category-added tensor having the learned position tensor, and the tensor is recorded as the converted image tensor.

[0019] Furthermore, the step of inferring the category of the image to be classified according to the transformed image tensor through the encoder unit and the fully connected neural network unit specifically includes:

[0020] Encoding and translating the transformed image tensor by the encoder unit to generate an encoded image tensor recognizable by the fully connected neural network unit;

[0021] The encoded image tensor is input into the fully connected neural network unit for inference, thereby obtaining classified image category information.

[0022] Furthermore, the step of encoding and translating the transformed image tensor by the encoder unit to generate an encoded image tensor recognizable by the fully connected neural network unit specifically includes:

[0023] Multiply the transformed image tensors by three transformation matrices respectively, and generate a query vector q, a key vector k, and a value vector v for each transformed image tensor;

[0024] Multiply the query vector q by the key vector k to obtain a score value that matches the number of the transformed image tensor;

[0025] Scaling the score value to obtain a scaled score value;

[0026] Using a softmax function to calculate a normalized score value of the scale adjustment score value, thereby calculating a normalized score value;

[0027] Multiplying the normalized score value by the value vector v to obtain a product vector that matches the quantity of the transformed image tensor;

[0028] Summing all the product vectors to generate a matrix vector that matches the quantity of the transformed image tensor;

[0029] Arranging the matrix components of the plurality of matrices generated in the same way to obtain a combined matrix, and then multiplying the combined matrix by a transformation matrix, thereby generating a coding matrix;

[0030] The category tensor of the first column in the encoding matrix is ​​extracted and recorded as the encoded image tensor.

[0031] Furthermore, the step of scaling the score value to obtain the scale-adjusted score value specifically includes:

[0032] Dividing the score by the square root of the length of the query vector q to obtain a scale adjustment score;

[0033] Furthermore, a weight matrix and an activation function are provided in the fully connected neural network unit;

[0034] Then the step of inputting the encoded image tensor into the fully connected neural network unit for calculation, thereby obtaining the classified image category information, specifically includes:

[0035] Inputting the coded image tensor into the weight matrix for inference to obtain a matrix inference tensor;

[0036] Inputting the matrix inference tensor into the activation function for inference, thereby obtaining a classification tensor;

[0037] Obtaining a preset classification condition, and calculating a probability score of the classification tensor meeting the preset classification condition;

[0038] Determining whether the probability score is higher than a preset threshold;

[0039] If so, the category information that matches the preset classification condition is recorded as the classified image category information corresponding to the encoded image tensor.

[0040] A three-dimensional image classification device includes a neural network, wherein the neural network includes an encoder unit and a fully connected neural network unit, wherein the encoder unit is connected to the fully connected neural network unit, and the device includes:

[0041] A classification image acquisition unit, used to acquire an image to be classified, wherein the image to be classified is a three-dimensional image to be classified;

[0042] An image segmentation unit, used for segmenting the image to be classified by image segmentation pairs to generate image blocks to be converted;

[0043] A tensor conversion unit, used for performing tensor conversion on the image block to be converted to form a converted image tensor;

[0044] A classification unit is used to infer the category of the image to be classified according to the transformed image tensor through the encoder unit and the fully connected neural network unit.

[0045] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps:

[0046] Acquire an image to be classified, wherein the image to be classified is a three-dimensional image to be classified;

[0047] Segmenting the image to be classified by image segmentation to generate image blocks to be converted;

[0048] Performing tensor conversion on the image block to be converted to form a converted image tensor;

[0049] The category of the image to be classified is inferred based on the transformed image tensor through the encoder unit and the fully connected neural network unit.

[0050] A computer readable medium stores a computer program, which, when executed by a processor, causes the processor to perform the following steps:

[0051] Acquire an image to be classified, wherein the image to be classified is a three-dimensional image to be classified;

[0052] Segmenting the image to be classified by image segmentation to generate image blocks to be converted;

[0053] Performing tensor conversion on the image block to be converted to form a converted image tensor;

[0054] The category of the image to be classified is inferred based on the transformed image tensor through the encoder unit and the fully connected neural network unit.

[0055] The above-mentioned three-dimensional image classification method, device, computer equipment and medium, by dividing the three-dimensional image to be classified into image blocks to be converted, converting the image blocks to be converted into tensor data, and classifying the image blocks to be converted according to the tensor data, realizes the image classification of the three-dimensional image to be classified by a convolutional neural network composed of an encoder unit and a fully connected neural network unit, thereby solving the problem in the prior art that the existing VisionTransformer can only classify two-dimensional images but cannot classify three-dimensional images. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0057] in:

[0058] Figure 1 A method flow chart of a three-dimensional image classification method in one embodiment;

[0059] Figure 2 is a structural block diagram of a three-dimensional image classification device in one embodiment;

[0060] Figure 3 1 is a structural block diagram of a computer device in one embodiment. DETAILED DESCRIPTION

[0061] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0062] like Figure 1 As shown, a three-dimensional image classification method includes a convolutional neural network, the convolutional neural network includes an encoder unit and a fully connected neural network unit, the encoder unit is connected to the fully connected neural network unit, and the method includes:

[0063] S1. Acquire an image to be classified, where the image to be classified is a three-dimensional image to be classified;

[0064] As described in the above step S1, the background system obtains the image to be classified. It can be understood that the image to be classified is a three-dimensional brain magnetic resonance image of the patient obtained by an external magnetic resonance device, wherein the background system is generally a background server. In addition, the background system can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The present invention does not limit this.

[0065] S2, segmenting the image to be classified by image segmentation to generate image blocks to be converted;

[0066] As described in the above step S2, the background system divides the image to be classified of the three-dimensional shape according to a preset ratio, thereby dividing it into a plurality of image blocks, and records the image blocks as image blocks to be converted. It can be understood that the preset ratio can be set by the user, and the specific division ratio of the preset ratio is not limited in this application;

[0067] S3, performing tensor conversion on the image block to be converted to form a converted image tensor;

[0068] As described in the above step S3, the background system performs tensor conversion on the image block to be converted, thereby converting the image block to be converted into a tensor image that can be compressed and classified by the encoder unit and the fully connected neural network unit, and records the converted tensor image as the converted image tensor.

[0069] S4. Infer the category of the image to be classified according to the transformed image tensor through the encoder unit and the fully connected neural network unit.

[0070] As described in the above step S4, the background system identifies the classification category corresponding to the converted image tensor according to the converted image tensor, thereby realizing the function of classifying the converted image tensor.

[0071] This embodiment uses the above method to implement image classification of the three-dimensional image to be classified by a convolutional neural network composed of an encoder unit and a fully connected neural network unit, by dividing the three-dimensional image to be classified into image blocks to be converted, converting the image blocks to be converted into tensor data, and classifying the image blocks to be converted according to the tensor data, thereby solving the problem in the prior art that the existing Vision Transformer can only classify two-dimensional images but cannot classify three-dimensional images.

[0072] In one embodiment, the step S3 specifically includes:

[0073] S31. Expand the image block to be converted by one-dimensionalizing the image block to obtain a one-dimensional expanded tensor of the image block to be converted, compress the one-dimensional expanded tensor by tensor linear projection to form a compressed one-dimensional tensor, obtain the tensor to be added, and add the tensor to be added to the compressed one-dimensional tensor, thereby generating the converted image tensor.

[0074] As described in the above embodiment, the background system expands the image block to be converted through an image one-dimensionalization operation, thereby converting the image block to be converted into one dimension, and obtaining a one-dimensional expanded tensor of the image block to be converted with various three-dimensional properties, which is recorded as a one-dimensional expanded tensor. Then, the background system linearly projects the one-dimensional expanded tensor to compress it, and obtains a plurality of compressed one-dimensional tensors with reduced scale, which are recorded as the compressed one-dimensional tensors. At the same time, the background system obtains the pre-stored tensor to be added and adds it to the compressed one-dimensional tensor, thereby generating the converted image tensor.

[0075] This embodiment uses the above method to compress the image block to be converted and then add the added tensor to generate the converted image tensor that can be recognized and converted by the encoder unit and the fully connected neural network unit, thereby preventing the encoder unit and the fully connected neural network unit from making errors when compressing and adding the image block to be converted.

[0076] In one embodiment, the tensor to be added includes a one-dimensional category tensor and a learning position tensor;

[0077] Then the step of obtaining the tensor to be added and adding the tensor to be added to the compressed one-dimensional tensor to generate the converted image tensor specifically includes:

[0078] S32. Add the one-dimensional category tensor to the compressed one-dimensional tensor to obtain the compressed one-dimensional tensor with the one-dimensional category tensor, and record it as the category-added tensor; add the learned position tensor to the category-added tensor to obtain the category-added tensor with the learned position tensor, and record it as the converted image tensor.

[0079] As described in the above embodiment, the background system obtains the one-dimensional category tensor in the tensor to be added, and adds the one-dimensional category tensor to the compressed one-dimensional tensor, thereby obtaining the compressed one-dimensional tensor with the one-dimensional category tensor, and records it as the category-added tensor. After that, the background system obtains the learning position tensor in the tensor to be added, and adds the learning position tensor to the compressed one-dimensional tensor, thereby obtaining the category-added tensor with the learning position tensor, and records it as the converted image tensor.

[0080] This embodiment uses the above method to generate the converted image tensor by adding the one-dimensional category tensor and the category-added tensor to the compressed one-dimensional tensor, thereby allowing the encoder unit to add categories to the converted image tensor, so that the encoder unit can perform subsequent inference on the converted image tensor based on the one-dimensional category tensor and the category-added tensor.

[0081] In one embodiment, the step of inferring the category of the image to be classified according to the transformed image tensor by the encoder unit and the fully connected neural network unit specifically includes:

[0082] S33. Encode and translate the transformed image tensor through the encoder unit to generate an encoded image tensor that can be recognized by the fully connected neural network unit, and input the encoded image tensor into the fully connected neural network unit for inference to obtain classified image category information.

[0083] As described in the above embodiment, the background system inputs the converted image tensor into the encoder unit for inference, and then encodes and translates the converted image tensor through the encoder unit, thereby generating an encoded image tensor that can be recognized by the fully connected neural network unit. Thereafter, the background system inputs the encoded image tensor into the fully connected neural network unit for inference, thereby obtaining classified image category information.

[0084] Afterwards, the background system inputs the converted image tensor into the encoder unit for calculation, thereby obtaining multiple one-dimensional tensors of unchanged size and recognizable by the fully connected neural network unit, and records them as the encoded image tensors. Afterwards, the background system inputs the encoded image tensor into the fully connected neural network unit for calculation, thereby obtaining classified image category information, thereby achieving the effect of classifying the encoded image tensor.

[0085] In one embodiment, the step of encoding and translating the transformed image tensor by the encoder unit to generate an encoded image tensor recognizable by the fully connected neural network unit specifically includes:

[0086] S34. Multiply the transformed image tensor by three transformation matrices respectively, generate a query vector q, a key vector k, and a value vector v for each transformed image tensor, multiply the query vector q by the key vector k to obtain a score value that matches the number of the transformed image tensors, scale the score value to obtain a scale-adjusted score value, use the softmax function to infer the normalized score value of the scale-adjusted score value, thereby inferring the normalized score value, multiply the normalized score value by the value vector v to obtain a product vector that matches the number of the transformed image tensors, sum all the product vectors to generate a matrix component vector that matches the number of the transformed image tensors, arrange the matrix component vectors generated by the same method to obtain a combined matrix, and then multiply it by a transformation matrix to generate a coding matrix, extract the category tensor of the first column in the coding matrix, and record it as the coded image tensor.

[0087] As described in the above embodiment, the background system records the converted image tensor as xi, and then multiplies xi by three preset weight matrices. The weight matrices in this embodiment are respectively expressed as:

[0088]

[0089] At this time, the backend system can obtain the query vector q: Key vector k: and the value vector v: At this time, dk is set to 3 by default in the encoder unit;

[0090] Then, the background system calculates M scoring values ​​α for the query vector q. In this embodiment, the scoring value calculation method is specifically performed as follows:

[0091] α i,j =qi·kj / √dk, j=1,…,M;

[0092] Then the backend system rescales the M score values, and at the same time the backend system calculates the softmax of the M score values ​​α to obtain M normalized score values Said The specific manifestations are:

[0093]

[0094] Afterwards, the background system will Multiply by M value vectors v respectively i Get the matrix component vector z i , the matrix forms the vector zi The specific manifestations are:

[0095]

[0096] Then the backend system combines the M matrices into a vector z i Construct a matrix The matrix The scales are as follows:

[0097] Query Matrix Key Matrix Value Matrix

[0098] It can be seen that at this time, the background system generates M matrices After that, the background system will convert all the matrices The combined matrix is ​​obtained by connecting the first and the last, and then multiplied by a transformation matrix to generate the encoding matrix. Then, the category tensor of the first column in the encoding matrix is ​​extracted and recorded as the encoded image tensor.

[0099] In one embodiment, the step of adjusting the length of the query vector q according to the score value specifically includes:

[0100] S35 . Divide the score value by the square root of the length of the query vector q to obtain a scale adjustment score value.

[0101] As described in the above embodiment, since the score value α outputs the similarity between the query vector q and the key vector k, the two vectors have the same length. The longer their lengths are, the greater the variance after dot product. Therefore, it is necessary to divide by the square root of the length of the query vector q, so that the gradient of the output scale-adjusted score value in machine learning is more stable.

[0102] In one embodiment, a weight matrix and an activation function are provided in the fully connected neural network unit;

[0103] Then the step of inputting the encoded image tensor into the fully connected neural network unit for calculation, thereby obtaining the classified image category information, specifically includes:

[0104] S36. Input the encoded image tensor into the weight matrix for calculation to obtain a matrix-calculated tensor, input the matrix-calculated tensor into the activation function for calculation to obtain a classification tensor, obtain a preset classification condition, calculate the probability score of the classification tensor reaching the preset classification condition, and determine whether the probability score is higher than a preset threshold. If so, record the category information that matches the preset classification condition as the classified image category information corresponding to the encoded image tensor.

[0105] As described in the above embodiment, the background system inputs the encoded image tensor into the weight matrix in the fully connected neural network unit for inference. It can be understood that there are n neurons in the fully connected neural network unit, and m×n weight matrices are arranged between the neurons in each layer, so the weight matrix is ​​specifically expressed as Then the matrix inference tensor output by this layer of neural network through the authority matrix is ​​collectively expressed as p=pA,

[0106] Afterwards, the background system inputs the matrix inference tensor into the activation function for inference, thereby obtaining a classification tensor. It can be understood that the activation function is a Relu activation function, then the background system infers the matrix inference tensor according to the Relu activation function, thereby outputting the classification tensor. Afterwards, the background system obtains a classification score that matches the preset classification condition, and determines the probability value of the classification tensor reaching the classification score in the subsequent inference process, and records it as the probability score of the preset classification condition. Afterwards, the background system obtains a preset threshold that matches the preset classification condition, and determines whether the probability score is higher than the preset threshold. If so, the category information that matches the preset classification condition is recorded as the classified image category information corresponding to the encoded image tensor. At this time, the background system can classify the image to be classified according to the classified image category information, thereby completing the process of classifying the image to be classified.

[0107] like Figure 2 As shown, a three-dimensional image classification device includes a neural network, wherein the neural network includes an encoder unit and a fully connected neural network unit, wherein the encoder unit is connected to the fully connected neural network unit, and the device includes:

[0108] The classification image acquisition unit 1 is used to acquire an image to be classified, wherein the image to be classified is a three-dimensional image to be classified;

[0109] An image segmentation unit 2 is used to segment the image to be classified by image segmentation pairs to generate image blocks to be converted;

[0110] A tensor conversion unit 3, used for performing tensor conversion on the image block to be converted to form a converted image tensor;

[0111] The classification unit 4 is used to infer the category of the image to be classified according to the transformed image tensor through the encoder unit and the fully connected neural network unit.

[0112] The above-mentioned units are used to execute the above-mentioned three-dimensional image classification device, and will not be introduced one by one here.

[0113] Figure 3 FIG. 1 shows an internal structure diagram of a computer device in an embodiment. The computer device may be a server, including but not limited to a high-performance computer and a high-performance computer cluster. Figure 3 As shown, the computer device includes a processor, a memory and a network interface connected via a system bus. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor may implement the employee status determination method. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor may execute the three-dimensional image classification method.

[0114] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0115] In one embodiment, the customer behavior identification method provided by the present application can be implemented in the form of a computer program. The computer program can be Figure 3 The computer device shown in the figure is run. The memory of the computer device can store various program templates constituting the automatic mail classification and aggregation device, such as the classification image acquisition unit 1, the image segmentation unit 2, the tensor conversion unit 3 and the classification unit 4.

[0116] like Figure 3 As shown, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:

[0117] An image to be classified is obtained, where the image to be classified is a three-dimensional image to be classified, the image to be classified is segmented by image segmentation to generate image blocks to be converted, the image blocks to be converted are tensor-converted to form a converted image tensor, and the category of the image to be classified is inferred according to the converted image tensor by the encoder unit and the fully connected neural network unit.

[0118] From the above embodiments, it can be seen that the greatest beneficial effect of the present invention is that by dividing the three-dimensional image to be classified into image blocks to be converted, converting the image blocks to be converted into tensor data, and classifying the image blocks to be converted according to the tensor data, a neural network composed of an encoder unit and a fully connected neural network unit can be used to classify the three-dimensional image to be classified, thereby solving the problem in the prior art that the existing VisionTransformer can only classify two-dimensional images but cannot classify three-dimensional images.

[0119] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0120] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0121] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A three-dimensional image classification method, characterized in that: A neural network includes an encoder unit and a fully connected neural network unit, and the encoder unit is connected to the fully connected neural network unit. The method includes: Acquire an image to be classified, wherein the image to be classified is a three-dimensional image to be classified; Segmenting the image to be classified by image segmentation to generate image blocks to be converted; Performing tensor conversion on the image block to be converted to form a converted image tensor; Inferring the category of the image to be classified according to the transformed image tensor through the encoder unit and the fully connected neural network unit; The step of inferring the category of the image to be classified according to the transformed image tensor by the encoder unit and the fully connected neural network unit specifically includes: Encoding and translating the transformed image tensor by the encoder unit to generate an encoded image tensor recognizable by the fully connected neural network unit; Inputting the encoded image tensor into the fully connected neural network unit for inference, thereby obtaining classified image category information; The step of encoding and translating the transformed image tensor by the encoder unit to generate an encoded image tensor recognizable by the fully connected neural network unit specifically includes: Multiply the transformed image tensors by three transformation matrices respectively, and generate a query vector q, a key vector k, and a value vector v for each transformed image tensor; Multiply the query vector q by the key vector k to obtain a score value that matches the number of the transformed image tensor; Scaling the score value to obtain a scaled score value; Using a softmax function to calculate a normalized score value of the scale adjustment score value, thereby calculating a normalized score value; Multiplying the normalized score value by the value vector v to obtain a product vector that matches the quantity of the transformed image tensor; Summing all the product vectors to generate a matrix vector that matches the quantity of the transformed image tensor; Arranging the matrix components of the plurality of matrices generated in the same way to obtain a combined matrix, and then multiplying the combined matrix by a transformation matrix, thereby generating a coding matrix; The category tensor of the first column in the encoding matrix is ​​extracted and recorded as the encoded image tensor.

2. The three-dimensional image classification method according to claim 1, characterized in that: The step of performing tensor conversion on the image block to be converted to form a converted image tensor specifically includes: Performing an expansion operation on the image block to be converted by a one-dimensional operation of the image block to obtain a one-dimensional expanded tensor of the image block to be converted; Compressing the one-dimensional expanded tensor by tensor linear projection to form a compressed one-dimensional tensor; A tensor to be added is obtained, and the tensor to be added is added to the compressed one-dimensional tensor, thereby generating the converted image tensor.

3. The three-dimensional image classification method according to claim 2, characterized in that: The tensor to be added includes a one-dimensional category tensor and a learning position tensor; Then the step of obtaining the tensor to be added and adding the tensor to be added to the compressed one-dimensional tensor to generate the converted image tensor specifically includes: Adding the one-dimensional category tensor to the compressed one-dimensional tensor to obtain the compressed one-dimensional tensor having the one-dimensional category tensor, and recording it as the category added tensor; The learned position tensor is added to the category-added tensor to obtain the category-added tensor having the learned position tensor, and the tensor is recorded as the converted image tensor.

4. The three-dimensional image classification method according to claim 1, characterized in that: The step of scaling the score value to obtain the scale-adjusted score value specifically includes: The score is divided by the square root of the length of the query vector q to obtain a scaled score.

5. The three-dimensional image classification method according to claim 1, characterized in that: The fully connected neural network unit is provided with a weight matrix and an activation function; Then the step of inputting the encoded image tensor into the fully connected neural network unit for calculation, thereby obtaining the classified image category information, specifically includes: Inputting the coded image tensor into the weight matrix for inference to obtain a matrix inference tensor; Inputting the matrix inference tensor into the activation function for inference, thereby obtaining a classification tensor; Obtaining a preset classification condition, and calculating a probability score of the classification tensor meeting the preset classification condition; Determining whether the probability score is higher than a preset threshold; If so, the category information that matches the preset classification condition is recorded as the classified image category information corresponding to the encoded image tensor.

6. A computer device, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps: Acquire an image to be classified, wherein the image to be classified is a three-dimensional image to be classified; Segmenting the image to be classified by image segmentation to generate image blocks to be converted; Performing tensor conversion on the image block to be converted to form a converted image tensor; Inferring the category of the image to be classified according to the transformed image tensor through an encoder unit and a fully connected neural network unit; The step of inferring the category of the image to be classified according to the transformed image tensor by the encoder unit and the fully connected neural network unit specifically includes: Encoding and translating the transformed image tensor by the encoder unit to generate an encoded image tensor recognizable by the fully connected neural network unit; Inputting the encoded image tensor into the fully connected neural network unit for inference, thereby obtaining classified image category information; The step of encoding and translating the transformed image tensor by the encoder unit to generate an encoded image tensor recognizable by the fully connected neural network unit specifically includes: Multiply the transformed image tensors by three transformation matrices respectively, and generate a query vector q, a key vector k, and a value vector v for each transformed image tensor; Multiply the query vector q by the key vector k to obtain a score value that matches the number of the transformed image tensor; Scaling the score value to obtain a scaled score value; Using a softmax function to calculate a normalized score value of the scale adjustment score value, thereby calculating a normalized score value; Multiplying the normalized score value by the value vector v to obtain a product vector that matches the quantity of the transformed image tensor; Summing all the product vectors to generate a matrix vector that matches the quantity of the transformed image tensor; Arranging the matrix components of the plurality of matrices generated in the same way to obtain a combined matrix, and then multiplying the combined matrix by a transformation matrix, thereby generating a coding matrix; The category tensor of the first column in the encoding matrix is ​​extracted and recorded as the encoded image tensor.

7. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the three-dimensional image classification method according to any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the three-dimensional image classification method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Image coding method and device and electronic equipment

    CN113507608A

  • Robotic medical system

    US11045271B1