3D medical image recognition method, device, equipment, and computer program

By decomposing 3D medical image features into 2D views for feature extraction and fusion, the method addresses the computational intensity and efficiency issues in current 3D medical image recognition, achieving improved recognition efficiency.

JP7744519B2Active Publication Date: 2025-09-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024531536
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-02-28
Filing Date
2022-12-16
Publication Date
2025-09-25
Estimated Expiration
2042-12-16

AI Technical Summary

Technical Problem

Current methods for recognizing 3D medical images are computationally intensive and require large amounts of data for pre-training, leading to low recognition efficiency and complexity.

Method used

The method involves decomposing 3D medical image features into 2D image features in different views through view rearrangement, followed by semantic feature extraction and fusion, reducing computational complexity and improving recognition efficiency.

Benefits of technology

This approach simplifies feature extraction by performing it on 2D image features in different views, thereby reducing computational complexity and enhancing the recognition efficiency of 3D medical images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007744519000059
    Figure 0007744519000059
  • Figure 0007744519000060
    Figure 0007744519000060
  • Figure 0007744519000061
    Figure 0007744519000061
Patent Text Reader

Abstract

A method, apparatus, device, and computer program for 3D medical image recognition, related to the field of artificial intelligence. The method includes: in the i-th feature extraction process, performing view rearrangement processing on the (i - 1)-th 3D medical image feature to obtain 2D image features, where the (i - 1)-th 3D medical image feature is the feature obtained by performing the (i - 1)-th feature extraction on the 3D medical image, and different 2D image features are the features in different views of the (i - 1)-th 3D medical image feature; performing semantic feature extraction processing on each 2D image feature to obtain image semantic features in different views; performing feature fusion processing on the image semantic features in different views to obtain the i-th 3D medical image feature; and performing image recognition processing based on the i-th 3D medical image feature obtained by the i-th feature extraction to obtain the image recognition result of the 3D medical image, where i is a positive integer that increases sequentially, 1 < i ≤ I, and I is a positive integer.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to a Chinese patent application bearing application number .202210191770.3, filed with the China Patent Office on February 28, 2022, the entire contents of which are incorporated herein by reference.

[0002] This application relates to the field of artificial intelligence, and in particular to a method, device, and machine for recognizing three-dimensional (3D) medical images. instrumentality and computer programs Mu Regarding. [Background technology]

[0003] In the medical field, computer vision technology is used to recognize 3D medical images and help predict disease conditions.

[0004] Currently, in the recognition process of 3D medical images, image analysis can be performed on 3D medical images using image density prediction methods, where the density prediction method is a method for predicting each pixel in the image. In the related art, when density prediction is performed on 3D medical images, image recognition is performed based on the entire 3D medical image to obtain an image recognition result.

[0005] However, methods that directly perform image recognition based on 3D medical images are computationally intensive, have low recognition efficiency, and require a large amount of data for pre-training, making them complex. Summary of the Invention

[0006] The embodiments of the present application provide a method, apparatus, device, computer-readable storage medium and computer program product for recognizing 3D medical images, which can improve the recognition efficiency of 3D medical images and reduce the calculation complexity.

[0007] An embodiment of the present application provides a method for recognizing 3D medical images executed by a computing device, the method comprising: In the i-th feature extraction process, a view rearrangement process is performed on the 3D medical image features of the (i - 1)-th time to obtain 2D image features, wherein the 3D medical image features of the (i - 1)-th time are features obtained by performing the (i - 1)-th feature extraction on the 3D medical image, and different 2D image features are features in different views of the 3D medical image features of the (i - 1)-th time; Performing semantic feature extraction processing on each of the 2D image features to obtain image semantic features in different views; Performing feature fusion processing on the image semantic features in different views to obtain the 3D medical image features of the i-th time; Performing image recognition processing based on the 3D medical image features of the i-th time obtained by the i-th feature extraction to obtain the image recognition result of the 3D medical image, where i is a positive integer that increases sequentially, 1 < i ≤ I, and I is a positive integer.

[0008] An embodiment of the present application provides a recognition device for 3D medical images, and the device includes: A view rearrangement module configured to perform a view rearrangement process on the 3D medical image features of the (i - 1)-th time in the i-th feature extraction process to obtain 2D image features, wherein the 3D medical image features of the (i - 1)-th time are features obtained by performing the (i - 1)-th feature extraction on the 3D medical image, and different 2D image features are features in different views of the 3D medical image features of the (i - 1)-th time; A feature extraction module configured to perform semantic feature extraction processing on each of the 2D image features to obtain image semantic features in different views; A feature fusion module configured to perform feature fusion processing on the image semantic features in different views to obtain the 3D medical image features of the i-th time; An image recognition module configured to perform image recognition processing based on the 3D medical image features of the i-th time obtained by the i-th feature extraction to obtain the image recognition result of the 3D medical image, where i is a positive integer that increases sequentially, 1 < i ≤ I, and I is a positive integer.

[0009] An embodiment of the present application provides a computer device, the computer device comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, code set or instruction set is loaded and executed by the processor to realize the 3D medical image recognition method described in the above aspect.

[0010] An embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, code set or instruction set is loaded and executed by the processor to realize the 3D medical image recognition method described in the above aspect.

[0011] An embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program including computer instructions stored in a computer-readable storage medium, a processor of a computer device reading the computer instructions from the computer-readable storage medium, and the processor executing the computer instructions to realize the method for recognizing 3D medical images according to the above aspect.

[0012] The technical solutions provided by the embodiments of the present application have at least the following beneficial effects:

[0013] In the embodiment of the present application, in each feature extraction step, the 3D medical image features are first divided into 2D image features in different views by view rearrangement, and then feature extraction is performed on the 2D image features to obtain image semantic features in different views, and the image semantic features in different views are then fused to obtain feature-extracted 3D medical image features. In this process, feature extraction is performed on 2D image features in different views. Therefore, compared to the related art method of directly extracting 3D image features for image recognition, the embodiment of the present application uses a simplified local computing unit to perform feature extraction in different views, thereby reducing computational complexity and improving the recognition efficiency of 3D medical images. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a schematic diagram illustrating the principle of a method for recognizing 3D medical images according to an embodiment of the present application. [Figure 2] FIG. 1 is a schematic diagram of an implementation environment according to an embodiment of the present application. [Figure 3] 1 is a flowchart of a method for recognizing 3D medical images according to an embodiment of the present application; [Figure 4] 1 is a flowchart of a method for recognizing 3D medical images according to an embodiment of the present application; [Figure 5] 1 is a schematic diagram illustrating the overall structure of an image recognition structure according to an embodiment of the present application. [Figure 6] FIG. 1 is a schematic diagram illustrating the structure of a spatial feature extraction process according to an embodiment of the present application. [Figure 7] FIG. 1 is a schematic diagram illustrating the structure of a semantic feature extraction process according to an embodiment of the present application. [Figure 8] FIG. 1 is a schematic diagram illustrating the structure of a feature fusion process according to an embodiment of the present application; [Figure 9] FIG. 1 is a schematic diagram illustrating the structure of a TR-MLP network according to an embodiment of the present application. [Figure 10] FIG. 1 is a schematic diagram illustrating the structure of a skip connection fusion network according to an embodiment of the present application. [Figure 11]1 is a block diagram showing the structure of a 3D medical image recognition device according to an embodiment of the present application; [Figure 12] FIG. 1 is a schematic diagram illustrating the structure of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION

[0015] In order to more clearly describe the technical solutions of the embodiments of the present application, the drawings used in the description of the embodiments are briefly introduced above. Obviously, the above drawings are only some embodiments of the present application, and those skilled in the art can also obtain other related drawings based on these drawings without any creative efforts.

[0016] In order to more clearly describe the objectives, technical solutions and advantages of the present application, the following describes the embodiments of the present application in more detail with reference to the drawings.

[0017] Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology in computer science that seeks to understand the nature of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI is the study of design principles and implementation methods for various intelligent machines that can perceive, reason, and make decisions.

[0018] Artificial intelligence technology is a comprehensive field that encompasses a wide range of fields, including both hardware and software technologies. Basic AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, medical image processing technology, operation / interaction systems, mechatronics, etc. Artificial intelligence software technologies mainly include computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0019] Computer vision (CV) is the science that studies how machines "see," using cameras and computers instead of human eyes to identify, measure, and otherwise manipulate targets. It also performs graphics processing to make computer processing more suitable for human observation or for transmitting images to machines for detection. As a scientific field, CV studies related theories and techniques to build artificial intelligence systems that can extract information from images or multidimensional data. CV technologies typically include image processing, image recognition, image segmentation, image semantic understanding, image retrieval, video processing, video semantic understanding, video content / action recognition, three-dimensional (3D) object reconstruction, 3D technology, virtual reality, augmented reality, and simultaneous localization and mapping (SLAM). It also includes common biological feature recognition techniques such as facial recognition and fingerprint recognition.

[0020] The 3D medical image recognition method according to the embodiments of the present application, i.e., the application of computer vision technology in the field of image recognition, performs feature extraction on 2D image features corresponding to 3D medical image features in different views, thereby reducing computational complexity and improving the recognition efficiency of 3D medical images.

[0021] For example, as shown in FIG. 1, in the i-th feature extraction process, first, view rearrangement is performed on the i-1th 3D medical image feature 101 obtained by the i-1th feature extraction to obtain a first 2D image feature 102 in the first view, a second 2D image feature 103 in the second view, and a third 2D image feature 104 in the third view, respectively; semantic feature extraction is performed on the first 2D image feature 102, the second 2D image feature 103, and the third 2D image feature 104 in different views, respectively, to obtain a first image semantic feature 105, a second image semantic feature 106, and a third image semantic feature 107; and then these three are fused to obtain the i-th 3D image semantic feature 108.

[0022] To decompose 3D medical image features into 2D image features in different views, feature extraction is performed on the 2D image features, which is advantageous to reduce the computational complexity and improve the recognition efficiency of 3D medical images.

[0023] The method according to the embodiment of the present application can be applied to the image recognition process of any 3D medical image, for example, to recognize the category to which each part of the 3D medical image belongs, and to assist in the analysis of lesions and organs.

[0024] The computing device for 3D medical image recognition according to the embodiments of the present application may be various types of terminal devices or servers, where the server may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services, and the terminal may be, but is not limited to, a smartphone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc.

[0025] Taking a server as an example, it may be a server cluster deployed in the cloud, and artificial intelligence cloud servers (AIaaS: AI as a Service) can be opened to users. The AIaaS platform divides several common AI services and provides them on the cloud as independent services or packaged services. This service model is similar to an AI theme mall, and all users can access one or more artificial intelligence services provided using the AIaaS platform through an application programming interface.

[0026] For example, one of the artificial intelligence cloud services may be a 3D medical image recognition service, i.e., a server in the cloud encapsulates a program for 3D medical image recognition provided by an embodiment of the present application. A user invokes the 3D medical image recognition service in the cloud service through a terminal (where a client such as a lesion analysis client is running). The server deployed in the cloud then invokes the encapsulated 3D medical image recognition program, decomposes 3D medical image features into 2D image features in different views, performs feature extraction on the 2D image features, and recognizes the 3D medical image to obtain image recognition results. The image recognition results can then be used to assist doctors and researchers in diagnosing diseases, conducting follow-up examinations, and researching treatment methods. For example, an auxiliary diagnosis can be performed based on an edema index included in the image recognition results, determining whether the target object has inflammation, trauma, allergies, or excessive water intake.

[0027] It should be noted that the 3D medical image recognition method according to the embodiments of the present application is not intended to directly obtain a disease diagnosis result or a health condition, and a disease diagnosis result or a health condition cannot be directly obtained based on the image recognition result. That is, the image recognition result is not directly used for disease diagnosis, but is used only as intermediate data for predicting a patient's disease and assisting doctors and researchers in diagnosing diseases, conducting follow-up examinations, and researching treatment methods.

[0028] 2 is a schematic diagram of an implementation environment according to an embodiment of the present application, including a terminal 210 and a server 220. Here, data communication between the terminal 210 and the server 220 is performed via a communication network, which in some embodiments may be a wired network or a wireless network, and may be at least one of a local area network (LAN), a metropolitan area network (MAN), and a wide area network (WAN).

[0029] The terminal 210 is an electronic device that executes a 3D medical image recognition program, and the electronic device may be a smartphone, a tablet computer, a personal computer, etc., and the embodiments of the present application are not limited thereto. When a 3D medical image needs to be recognized, the 3D medical image can be input into the program of the terminal 210, and the terminal 210 uploads the 3D medical image to the server 220. The server 220 executes the 3D medical image recognition method according to the embodiments of the present application to perform image recognition, and feeds back the image recognition result to the terminal 210.

[0030] The server 220 may be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), big data and artificial intelligence platforms.

[0031] In some embodiments, server 220 is configured to provide image recognition services to applications installed on terminal 210. In some embodiments, server 220 is provided with an image recognition network for classifying 3D medical images transmitted by terminal 210.

[0032] Of course, in some embodiments, the image recognition network may be located on the terminal 210, and the terminal 210 locally implements the 3D medical image recognition method (i.e., the image recognition network) according to the embodiments of the present application without the intervention of the server 220. Correspondingly, the image recognition network may complete training on the terminal 210, and the embodiments of the present application are not limited thereto. For convenience of explanation, the following embodiments will be described taking the case where the 3D medical image recognition method is executed by a computer device as an example.

[0033] Referring to FIG. 3, FIG. 3 is a flowchart of a 3D medical image recognition method according to an embodiment of the present application, which includes the following steps:

[0034] In step 301, in the i-th feature extraction process, a view rearrangement process is performed on the i-1th 3D medical image feature to obtain a 2D image feature, the i-1th 3D medical image feature is a feature obtained by performing the i-1th feature extraction on the 3D medical image, and the different 2D image features are features in different views of the i-1th 3D medical image feature.

[0035] Here, the 3D medical image features are features obtained by extraction from a 3D medical image to be recognized, which may be a 3D medical image such as a computed tomography (CT) image, a magnetic resonance imaging (MRI) image, or a positron emission computed tomography (PET) image.

[0036] Here, the first 3D medical image features are features obtained by performing feature extraction on the initial 3D medical image features, and the initial 3D medical image features are features obtained by performing an initial embedding process on the 3D medical image. Here, the initial embedding process is intended to obtain low-dimensional initial 3D medical image features by mapping high-dimensional data, such as the 3D medical image, into a low-dimensional space.

[0037] In the embodiment of the present application, 3D medical images are recognized through multiple feature extraction processes, where the same feature extraction network is used for each feature extraction process, and the input of the feature extraction network in each feature extraction process is determined based on the output result of the previous feature extraction network, i.e., in the i-th feature extraction process, feature extraction is performed based on the (i-1)th 3D medical image features.

[0038] Because 3D medical image features are 3D data, directly extracting features from the entire 3D medical image requires a large amount of calculation and results in a complex process. Therefore, in this embodiment, in each feature extraction process, the 3D medical image features are first divided. That is, in the i-th feature extraction process, view rearrangement is performed on the features obtained by the (i-1)th feature extraction. Here, view rearrangement divides the 3D medical image features into 2D image features in different views, and then performs feature extraction based on the 2D image features in the different views, thereby reducing computational complexity.

[0039] In some embodiments, the view rearrangement process is realized by performing view rearrangement processing on multiple dimensions of the (i-1)th 3D medical image feature to obtain 2D image features in multiple views, i.e., performing alignment and combination processing on multiple dimensions of the (i-1)th 3D medical image feature to obtain multiple views and extracting 2D image features in each view respectively.

[0040] In one embodiment, view rearrangement is performed on the (H,W,D) dimension of the 3D medical image feature for the (i-1)th time, to obtain 2D image features in three views, (H,W), (H,D), and (W,D), where each view corresponds to one 2D direction in the 3D medical image feature. Different 2D image features are image features corresponding to different 2D image slices, where the 2D image slices are 2D images in 2D space after view rearrangement is performed on the 3D medical image.

[0041] In addition, in the i-th feature extraction process, when feature extraction is performed based on the i-1th 3D medical image features, upsampling or downsampling may be performed on the i-1th 3D medical image features. In this case, view rearrangement is performed on the i-1th 3D medical image features after upsampling or downsampling to obtain 2D image features.

[0042] In step 302, a semantic feature extraction process is performed on each 2D image feature to obtain image semantic features in different views.

[0043] For example, after obtaining each 2D image feature, semantic feature extraction is performed on the 2D image feature to learn image information in the corresponding 2D image slice, where the process of performing semantic feature extraction on the 2D image feature includes learning 2D image slice spatial information and image semantic learning based on the corresponding view.

[0044] After performing semantic feature extraction on each 2D image feature, we can obtain image semantic features corresponding to different views, i.e., we obtain image semantic features corresponding to three views: (H,W), (H,D), and (W,D).

[0045] In step 303, a feature fusion process is performed on the image semantic features in different views to obtain the i-th 3D medical image feature.

[0046] In one embodiment, after obtaining the image semantic features in different views, the feature extraction process is completed by fusing the image semantic features in different views, obtaining the 3D medical image features of the i-th time, and then, based on the 3D medical image features of the i-th time, the feature extraction process of the (i + 1)-th 3D medical image features can be performed.

[0047] In the embodiments of the present application, by performing feature fusion on the image semantic features in different views, the intensive aggregation of rich semantics of full-view learning is realized, and the feature learning process of 3D medical image features is completed.

[0048] In step 304, image recognition processing is performed based on the i-th 3D medical image features obtained by the i-th feature extraction, and an image recognition result of the 3D medical image is obtained. i is a positive integer that increases sequentially, 1 < i ≤ I, and I is a positive integer.

[0049] After passing through multiple feature extraction processes and the feature extraction process is completed, after the i-th feature extraction process is completed, image recognition is performed based on the i-th 3D medical image features.

[0050] As described above, in the embodiments of the present application, in each feature extraction stage, first, view rearrangement is performed on the 3D medical image features to divide them into 2D image features in different views, feature extraction is performed on each of the 2D image features, and image semantic features in different views are obtained. Thereby, the image semantic features in different views are fused to obtain the 3D image semantic features after feature extraction. In this process, since feature extraction is performed on the 2D image features in different views, compared with the method of directly extracting 3D image features in the related art, the embodiments of the present application can reduce the computational complexity and improve the recognition efficiency of 3D medical images by performing feature extraction in different views using a simplified local calculation unit.

[0051] In some embodiments, in the process of performing feature extraction on 2D image features in different views, each 2D image feature is divided to learn features corresponding to a local window, and then learn context features of slices corresponding to each 2D image feature, thereby obtaining image semantic features in different views. Exemplary embodiments are described below.

[0052] Referring to FIG. 4, FIG. 4 is a flowchart of a 3D medical image recognition method according to an embodiment of the present application, which includes the following steps:

[0053] In step 401, in the i-th feature extraction process, a view rearrangement process is performed on the (i-1)-th 3D medical image feature to obtain a 2D image feature.

[0054] After acquiring a 3D medical image, a patch embedding process is first performed on the 3D medical image, for example, using a convolutional stem structure to perform a rudimentary embedding process and obtain initial 3D medical image features. Then, multiple feature extraction processes can be performed using the initial 3D medical image features as a starting point. Here, the convolutional stem is the initial convolution layer of a convolutional neural network. The patch embedding process is used to obtain low-dimensional initial 3D medical image features by mapping high-dimensional data, such as the 3D medical image, into a low-dimensional space.

[0055] In an embodiment of the present application, the feature extraction process includes a feature encoding process and a feature decoding process. Here, the feature encoding process includes a downsampling process for 3D medical image features, i.e., a process of reducing the dimensionality of the 3D medical image features, and the feature decoding process includes an upsampling process for 3D medical image features, i.e., a process of increasing the dimensionality of the 3D medical image features. Here, the downsampling process uses 3D convolution with a kernel size of 3 and a stride of 2 to downsample by 2 each time. The upsampling process uses 3D transposed convolution with a kernel size of 2 and a stride of 2 to upsample by 2 each time. After multiple feature encoding and feature decoding processes, the resulting 3D medical image features are used to recognize medical images. Here, each feature extraction process is implemented using the same Transformer-Multilayer Perceptron (TR-MLP) structure.

[0056] For example, as shown in FIG. 5, the size C iA 3D medical image of 2C×H / 4×W / 4×D / 8 is input, and an initial embedding process (Patch Embedding) 501 is first performed, where the image block (Patch) size is 2×2, and 3D medical image features of C×H / 4×W / 4×D / 4 are obtained. The 3D medical image features of C×H / 4×W / 4×D / 4 are input to a first TR-MLP block to perform a first feature extraction. After the first feature extraction is completed, the obtained first 3D medical image features are downsampled to obtain 3D medical image features of 2C×H / 8×W / 8×D / 8. The 3D medical image features of 2C×H / 8×W / 8×D / 8 are input to a second TR-MLP block to perform a second feature extraction to obtain second 3D medical image features. Then, the second 3D medical image features are input into the third TR-MLP block to perform the third feature extraction, and after the third extraction is completed, the 3D medical image features obtained in the third extraction are downsampled to 8C×H / 32×W / 32×D / 32, and then an upsampling process is performed. Here, the feature extraction process performed in the TR-MLP block 502 and the feature extraction process in the immediately preceding TR-MLP block are feature encoding processes, and then a feature decoding process is performed.

[0057] Note that each feature encoding or decoding process is realized by a view rearrangement process, a semantic feature process, and a feature fusion process.

[0058] 3 can be realized by steps 402 and 403 in FIG.

[0059] In step 402, a spatial feature extraction process is performed on the 2D image features to obtain 2D image spatial features.

[0060] After obtaining the 2D image features corresponding to each view, spatial feature extraction is first performed on the 2D image features. Here, the spatial feature extraction process is a process of learning the features of each corresponding 2D image slice. Here, in the process of spatial feature extraction based on the three views, the network parameters are shared, i.e., the network parameters are the same. This process may include steps 402a to 402c (not shown).

[0061] In step 402a, a window division process is performed on the 2D image features to obtain local 2D image features corresponding to N windows, where the N windows do not overlap each other and N is a positive integer greater than 1.

[0062] In this process, a window-based multi-head self-attention (W-MSA) network structure is mainly used to model the long-range and local spatial semantic information in 2D image slices. Here, when processing 2D image features using the W-MSA network structure, first, a window division process is performed on the 2D image feature Z, and local 2D image features Z corresponding to N non-overlapping windows are generated. i The division process is shown in Equation 1.

[0063]

number

[0064] Here, M is the window size set in W-MSA, and HW is the size of the 2D image feature, that is, the 2D image size obtained by cutting in the (H, W) view.

[0065] Then, we perform attention calculation based on the window and obtain the output result, i.e., local 2D image space features.

[0066] Attention processing is realized by an attention mechanism. In cognitive science, an attention mechanism is used to selectively pay attention to some of all information and ignore other information. The attention mechanism enables a neural network to pay attention to some inputs, i.e., select specific inputs. When computing power is limited, the attention mechanism is a resource allocation method used as a primary means to solve the information overload problem, allocating computing resources to more important tasks. However, the embodiments of the present application are not limited to the form of the attention mechanism. For example, the attention mechanism may be multi-head attention, key-value pair attention, structured attention, etc.

[0067] In step 402b, a feature extraction process is performed on the N local 2D image features to obtain a 2D image window feature.

[0068] The local 2D image features Z corresponding to each of the N non-overlapping windows are i After obtaining N local 2D image features, feature extraction is performed on each local 2D image feature to obtain N 2D image window features. Here, the feature extraction processing method includes the following steps:

[0069] Step 1: Perform self-attention processing on N local 2D image features to obtain self-attention features of the N local 2D image features.

[0070] First, we perform self-attention processing on each local 2D image feature. Here, the self-attention processing process is multi-head self-attention processing, where each local 2D image feature corresponds to multiple self-attention heads.

[0071] For example, a self-attention process is performed based on the query term Q, key term K, and value term V corresponding to the local 2D image features to obtain self-attention features of the N local 2D image features.

[0072] Here, the query item (Q,Query), key item (K,Key), and value item (V,Value) corresponding to the k-th self-attention head are respectively:

number

[0073]

number

[0074] Here, RPE is relative position encoding information, i.e., window position encoding, and represents window-sensitive spatial position information.

[0075] In this case, the self-attention feature corresponding to the kth self-attention head includes features corresponding to N windows, as shown in Equation 3.

[0076]

number

[0077] Step 2: Perform feature fusion processing on the N local 2D image characteristic self-attention features to obtain the first image window internal features.

[0078] After obtaining the self-attention features corresponding to each self-attention head in each window, the self-attention features corresponding to all self-attention heads are merged and linearly mapped by the parameter matrix to realize the feature fusion process and obtain the corresponding internal features of the first image window. This process is shown in Equation 4.

[0079]

number

[0080] where W H is the parameter matrix and Concat represents the merging operation.

[0081] In some embodiments, before performing self-attention processing based on the W-MSA structure, we first calculate the l-th local 2D image feature from view v.

number

number

[0082] For example, as shown in FIG. 6, first,

number

number

number

[0083]

number

[0084] Step 3: Perform convolution processing on the first image window internal features to obtain the first image window interaction features.

[0085] Here, the W-MSA structure is used to learn the features of each divided local 2D image feature. To further enhance the learning of 2D image features, a depthwise separable convolutional block (DWConv2D) structure with a kernel size of 5 is used to perform convolution processing, thereby increasing local learning between spatially adjacent windows. For example, the internal features of the first image window are input to the DWConv2D network and convolution processing is performed to obtain the first image window interaction features.

[0086] In some embodiments, DWConv2D may also include a residual structure, i.e., the first image window intra-feature after convolution and the first image window intra-feature are fused to generate the first image window interaction feature, as shown in Equation 6.

number

[0087]

number

[0088] For example, as shown in FIG. 6, the first image window internal features

number

number

number

[0089] Step 4: Perform feature extraction processing on the first image window interaction feature by multi-layer perceptron (MLP) to obtain 2D image window feature.

[0090] To further enhance the learning of the 2D image slices in the corresponding views, the first image window interaction features after the convolution process are normalized using BN, and the channel features, i.e., the features of the 2D image slices in the corresponding views, are learned using a multilayer perceptron (MLP), and the 2D image window features are then obtained.

number

[0091]

number

[0092] Here, MLP stands for multi-layer perceptron structure.

[0093] In step 402c, a window rearrangement process is performed on the N windows, and a 2D image window feature extraction process is performed on each of the N windows after the window rearrangement to obtain 2D image spatial features, which are used to change the spatial positions of the N windows.

[0094] After the window self-attention learning is performed using the W-MSA structure, it is necessary to further learn the cross-window image feature information. Therefore, in one possible embodiment, the window rearrangement is performed on the N windows, and the 2D image window features after the window rearrangement are re-learned.

[0095] For example, by using a shuffle operation to rearrange the windows, spatial information can be perturbed and the interaction between cross-window information can be enhanced. After the window rearrangement, 2D image window features corresponding to the N windows are learned to obtain the final 2D image spatial features. Here, the method may include the following steps:

[0096] Step 1: Perform self-attention processing on the 2D image window features corresponding to each of the N windows after window rearrangement to obtain self-attention features corresponding to each of the N windows.

[0097] First, perform self-attention processing on the 2D image window features corresponding to each of the N windows after window rearrangement to obtain self-attention features, where the method can refer to the above steps and will not be repeated here.

[0098] Step 2: Perform feature fusion processing on the N self-attention features to obtain the second image window interior features.

[0099] Here, the process of obtaining the features within the second image window by feature fusion can refer to the process of obtaining the features within the first image window by fusion, and will not be described again here.

[0100] Step 3: perform a position inversion process on the internal features of the second image window, and perform a convolution process on the internal features of the second image window after the position inversion to obtain second image window interaction features.

[0101] For example, by perturbing the window positions again, we can use the W-MSA structure to perform window self-attention learning again to strengthen cross-window information learning, and then perform position inversion on the internal features of the second image window, i.e., restore the position information corresponding to each window to its original position to obtain the second image window interaction features.

[0102] For example, as shown in Figure 6, first, BN normalization processing is performed on the 2D image window features, then window rearrangement (Transpose) is performed, and feature learning (including self-attention processing and feature fusion processing) is performed on the 2D image window features corresponding to each of the N windows after window rearrangement based on the W-MSA structure to obtain a second image window interaction feature, and then the positions of the N windows are reversed again to restore the position information corresponding to each window. This processing is shown in Equation 8.

[0103]

number

[0104] where:

number

number

[0105] Then, after the position inversion, the convolution process is performed again using DWConv2D to obtain the second image window interaction feature, which can be referred to the process of obtaining the first image window interaction feature by convolution process in the above step, and will not be described again here.

[0106] Exemplarily, as shown in FIG.

number

number

[0107]

number

[0108] Step 4: Perform feature extraction processing on the second image window interaction feature by MLP to obtain 2D image space feature.

[0109] For example, after the convolution process, we again use MLP to perform channel learning to obtain the final 2D image space features.

[0110] For example, as shown in FIG. 6, first, the second image window interaction feature

number

number

number

[0111]

number

[0112] Performing spatial feature extraction on 2D image features to obtain 2D image spatial features is the Full-View Slice Spatial Shuffle Block (FVSSSB) process, and the entire process is shown in Figure 6. This allows the 2D image features to be fully learned and accurate 2D image spatial features to be extracted, facilitating subsequent accurate image recognition.

[0113] In step 403, based on the main view and auxiliary view, a semantic feature extraction process is performed on the 2D image space features to obtain image semantic features, where the main view is a view corresponding to the 2D image features, and the auxiliary view is a 3D view different from the main view.

[0114] Since the 2D image spatial features only represent features corresponding to the 2D view (i.e., the main view), spatial feature extraction is performed on each 2D image feature to obtain the 2D image spatial features, and then the remaining semantic information of the remaining third view (i.e., the auxiliary view) is captured to perform supplementary learning of the information. Here, the process of performing semantic feature extraction on the 2D image spatial features to obtain image semantic features is a slice-aware volume context mixing (SAVCM) process, where the network parameters of the SAVCM network are shared in each view, i.e., the network parameters are the same. This process may include the following steps:

[0115] In step 403a, a feature fusion process is performed on the 2D image spatial feature and the position-encoding feature to obtain a first image semantic feature, and the position-encoding feature is used to indicate the position information corresponding to the 2D image feature.

[0116] In one possible embodiment, first, each 2D image space feature

number

number

[0117] For example, as shown in FIG. 7, the 2D image spatial feature and the position coding feature are fused to generate a first image semantic feature.

number

[0118]

number

[0119] Here, APE s teeth,

number

[0120] In step 403b, in the main view, perform semantic feature extraction on the first image semantic features by MLP to obtain main image semantic features.

[0121] In one possible embodiment, semantic feature extraction is performed on the main view and the auxiliary view, respectively, where the main view refers to a view corresponding to the 2D image features, and the auxiliary view is a 3D view different from the main view. For example,

number

[0122] For example, we use residual axial multi-layer perceptron (axial-MLP) to extract semantic features from the first image in the main view, and then extract the main image semantic features.

number

number

[0123] In step 403c, in the auxiliary view, perform semantic feature extraction on the first image semantic features by MLP to obtain auxiliary image semantic features.

[0124] The semantic feature extraction is performed based on the main view, and at the same time, the semantic feature extraction is performed based on the auxiliary view using MLP for the first image semantic feature, and the auxiliary image semantic feature is extracted.

number

[0125] In step 403d, a feature fusion process is performed on the main image semantic features and the auxiliary image semantic features to obtain image semantic features.

[0126] For example, after obtaining the main image semantic features and the auxiliary image semantic features, feature fusion is performed on both of them to obtain the image semantic features. In one possible embodiment, as shown in FIG.

number

number

[0127]

number

[0128] Here, Axial-MLP represents the axial multi-layer perceptron operation, Concat represents the merging operation, and MLPcp represents the feature fusion operation.

[0129] 3 can be realized by steps 404 and 405 in FIG.

[0130] In step 404, a fusion process is performed on the image semantic features and the view features to obtain view image semantic features.

[0131] In the feature fusion process, we first extract the image semantic features of each view.

number

number

[0132]

number

[0133] In step 405, a feature fusion process is performed on each view image semantic feature to obtain the i-th 3D medical image feature.

[0134] Next, the full view features of the three channels

number

number

number

[0135] Here, Concat represents a merging operation, LN represents a normalization operation, and MLPva represents a mapping operation.

[0136] As shown in Fig. 8, we first fuse each image semantic feature with the APE coding, and then stitch together the three views to obtain the final 3D medical image features.

[0137] 3 can be realized by steps 406 and 407 in Fig. 4. Here, the feature extraction process includes a feature encoding process or a feature decoding process, where the feature encoding process includes a downsampling process for the 3D medical image features, i.e., a process of reducing the dimension of the 3D medical image features, and the feature decoding process includes an upsampling process for the 3D medical image features, i.e., a process of increasing the dimension of the 3D medical image features.

[0138] In step 406, if the upsampling result reaches the original size, the 3D medical image feature obtained by extraction is determined as the I-th 3D medical image feature obtained by the I-th feature extraction.

[0139] In one possible embodiment, when the upsampling result reaches the original size of the 3D medical image, it is determined to be the Ith feature extraction process.

number

number

number

[0140] In step 407, an image recognition process is performed based on the I-th 3D medical image feature to obtain an image recognition result.

[0141] Finally, image recognition is performed based on the I-th 3D medical image features, and then image registration, classification, etc. can be performed on the 3D medical image.

[0142] In one possible embodiment, the TR-MLP network structure first calculates the 3D medical image features Z i For the (H,W,D) dimensions of the image, view rearrangement is performed to rearrange the image into three 2D image slices of (H,W), (H,D), and (W,D), each corresponding to one 2D slice direction in the 3D. After rearrangement, the full-view 2D image slices are fully learned using FVSSB to obtain 2D image features. Next, slice-sensitive context mixing (SAVCM) is used to capture the remaining image semantic information along the third view. Finally, a view-sensitive aggregator is used to aggregate the rich semantic information learned from the full-view learning, ultimately resulting in 3D medical image features in the Transformer-MLP block.

number

[0143] In the embodiment of the present application, full-view 2D spatial information is first learned, and then the remaining image semantics on the third view are learned. After that, the full-view semantics are fused to realize the context-sensitive ability of 3D medical image features and greatly enhance the inductive bias ability, thereby improving the accuracy of 3D medical image recognition. The computationally intensive 3D convolutional neural network (3D CNN) and pure visual transformer are replaced with a simplified local visual transformer-MLP computation unit, thereby reducing computational complexity and improving recognition efficiency.

[0144] Here, the feature extraction process includes a feature encoding process or a feature decoding process, and the extraction process includes a self-attention processing process, where the self-attention processing process calculates self-attention based on Q, K, and V. In one possible embodiment, to fuse multi-scale visual features, the features of the feature encoding process (realized by the encoding device) and the features of the feature decoding process (realized by the decoding device) are fused to obtain Q, K, and V values ​​in the feature decoding process.

[0145] In some embodiments, the K value in the t-th feature decoding process is obtained by fusion based on the K value in the t-1-th feature decoding and the K value in the corresponding feature encoding process, the V value in the t-th feature decoding process is obtained by fusion based on the V value in the t-1-th feature decoding and the V value in the corresponding feature encoding process, and the Q value in the t-th decoding process is the Q value in the t-1-th feature decoding.

[0146] In one possible embodiment, the resolution of the input features of the t-th feature decoding and the output features of the corresponding encoding process are the same. That is, skip connection fusion is performed on image features with the same resolution. For example, as shown in FIG. 5, the resolution corresponding to the second feature decoding process is 4C×H / 16×W / 16×D / 16, and the feature encoding process corresponding to the skip connection fusion is the final encoding process, which also has the resolution of 4C×H / 16×W / 16×D / 16. When performing skip connection fusion, skip connection fusion is performed on the features input in the second feature decoding (i.e., the features obtained by upsampling the output features of the first feature decoding) and the output features of the final feature encoding process.

[0147] The output feature of the feature encoding process corresponding to the t-th feature decoding is E v Let the input feature of the t-th feature decoding process be D v Here, v refers to a certain view. That is, skip connection fusion is performed on different views. First, E v , D v Convolution is performed on the encoder feature E using standard convolution (PWConv2D) with a kernel size of 1. Here, in the feature decoding process, the Q value is obtained only from the previous feature decoding process, and in the skip connection fusion of the encoder and decoder, fusion is performed only on the K value and the V value. Therefore, as shown in Figure 10, the encoder feature E is obtained using PWConv2D. v Divide the original number of channels into two,

number

[0148]

number

[0149] As shown in Figure 10, PWConv2D is used to obtain the decoder feature D v Divide the original number of channels into three,

number

[0150]

number

[0151] after that,

number

[0152]

number

[0153] where:

number

[0154]

number

[0155] Here, CrossMerge represents a skip connection merging operation.

[0156] In the embodiments of the present application, a skip connection fusion network is introduced, and skip connection fusion is performed on the features corresponding to the encoding device and the decoding device, so as to fuse multi-scale information and enrich the semantic learning of image features.

[0157] FIG. 11 is a block diagram showing the structure of a 3D medical image recognition device according to an embodiment of the present application. As shown in FIG. 11, the device includes a view rearrangement module 1101, a feature extraction module 1102, a feature fusion module 1103, and an image recognition module 1104. The view rearrangement module 1101 is configured to perform view rearrangement processing on the 3D medical image features of the (i - 1)th time in the ith feature extraction process to obtain 2D image features. The 3D medical image features of the (i - 1)th time are features obtained by performing the (i - 1)th feature extraction on the 3D medical image. Different 2D image features are features in different views of the 3D medical image features of the (i - 1)th time. The feature extraction module 1102 is configured to perform semantic feature extraction processing on each of the 2D image features to obtain image semantic features in different views. The feature fusion module 1103 is configured to perform feature fusion processing on the image semantic features in different views to obtain 3D medical image features of the ith time. The image recognition module 1104 is configured to perform image recognition processing based on the 3D medical image features of the Ith time obtained by the Ith feature extraction to obtain the image recognition result of the 3D medical image. Here, i is a positive integer that increases sequentially, 1 < i ≤ I, and I is a positive integer.

[0158] In some embodiments, the feature extraction module 1102 a first extraction unit configured to perform spatial feature extraction processing on the 2D image features to obtain 2D image spatial features, and and a second extraction unit that performs semantic feature extraction processing on the 2D image spatial features based on a main view and an auxiliary view to obtain the image semantic features, wherein the main view is a view corresponding to the 2D image features and the auxiliary view is a 3D view different from the main view.

[0159] In some embodiments, the first extraction unit is further configured to perform the following steps: performing a window division process on the 2D image features to obtain local 2D image features corresponding to each of N windows, where the N windows do not overlap each other and N is a positive integer greater than 1; performing a feature extraction process on the N local 2D image features to obtain 2D image window features; and performing a window rearrangement process on the N windows to obtain 2D image spatial features corresponding to each of the N windows after the window rearrangement, where the window rearrangement is used to change the spatial positions of the N windows.

[0160] In some embodiments, the first extraction unit further comprises: The system is configured to perform the steps of: performing a self-attention process on the N local 2D image features to obtain self-attention features corresponding to each of the N local 2D image features; performing a feature fusion process on the N self-attention features to obtain first image window interior features; performing a convolution process on the first image window interior features to obtain first image window interaction features; and performing a feature extraction process on the first image window interaction features by a multi-layer perceptron (MLP) to obtain the 2D image window features.

[0161] In some embodiments, the first extraction unit further comprises: The system is configured to perform the following steps: performing self-attention processing on the 2D image window features corresponding to each of the N windows after window rearrangement to obtain self-attention features corresponding to each of the N windows; performing feature fusion processing on the N self-attention features to obtain second image window internal features; performing position inversion processing on the second image window internal features and performing convolution processing on the second image window internal features after position inversion to obtain second image window interaction features; and performing feature extraction processing on the second image window interaction features using a multi-layer perceptron (MLP) to obtain the 2D image space features.

[0162] In some embodiments, the first extraction unit further comprises: The method is configured to perform a self-attention process based on a query term Q, a key term K, and a value term V corresponding to the local 2D image features to obtain self-attention features of the N local 2D image features.

[0163] In some embodiments, the feature extraction process includes a feature encoding process or a feature decoding process, wherein the K value in the t-th feature decoding process is obtained by fusion based on the K value in the t-1-th feature decoding and the K value in the corresponding feature encoding process, the V value in the t-th feature decoding process is obtained by fusion based on the V value in the t-1-th feature decoding and the V value in the corresponding feature encoding process, and the Q value in the t-th decoding process is the Q value in the t-1-th feature decoding.

[0164] In some embodiments, the second extraction unit further comprises: The system is configured to perform the following steps: performing a feature fusion process on the 2D image spatial features and position-coding features to obtain first image semantic features, wherein the position-coding features are used to indicate position information corresponding to the 2D image features; performing a semantic feature extraction process on the first image semantic features by an MLP in the main view to obtain main image semantic features; performing a semantic feature extraction process on the first image semantic features by an MLP in the auxiliary view to obtain auxiliary image semantic features; and performing a feature fusion process on the main image semantic features and the auxiliary image semantic features to obtain the image semantic features.

[0165] In some embodiments, the feature fusion module 1103 comprises: a first fusion unit configured to perform a fusion process on the image semantic features and view features to obtain view image semantic features; a second fusion unit configured to perform a feature fusion process on each of the view image semantic features to obtain the i-th 3D medical image feature.

[0166] In some embodiments, the feature extraction module 1102 further comprises: The feature extraction networks are configured to perform semantic feature extraction processing on the 2D image features in each view using feature extraction networks corresponding to the same network parameters, thereby obtaining the image semantic features in different views.

[0167] In some embodiments, the feature extraction process comprises a feature encoding process or a feature decoding process, wherein the feature encoding process comprises a downsampling process for 3D medical image features, and wherein the feature decoding process comprises an upsampling process for 3D medical image features.

[0168] The image recognition module 1104 a determining unit configured to determine the 3D medical image feature obtained by extraction as the I-th 3D medical image feature obtained by the I-th feature extraction when the upsampling result reaches the original size; The method further includes a recognition unit configured to perform image recognition processing based on the I-th 3D medical image features to obtain the image recognition result.

[0169] In some embodiments, the 3D medical image is a CT image, an MRI image, or a PET image.

[0170] As described above, in the embodiment of the present application, in each feature extraction step, 3D medical image features are first divided into 2D image features in different views by performing view rearrangement, and then feature extraction is performed on the 2D image features to obtain image semantic features in different views, and the image semantic features in different views are then fused to obtain feature-extracted 3D medical image features. In this process, feature extraction is performed on 2D image features in different views. Therefore, compared to the related art's method of directly extracting 3D image features, the embodiment of the present application uses a simplified local computing unit to perform feature extraction in different views, thereby reducing computational complexity and improving the recognition efficiency of 3D medical images.

[0171] It should be noted that the device provided in the above embodiments is only described as an example of dividing the above functional modules, and in actual applications, the above functions can be completed by allocating them to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the above-described functions. In addition, the device provided in the above embodiments belongs to the same concept as the method embodiments of the terminal device, and the implementation process thereof can be referred to the method embodiments, and will not be repeated here.

[0172] 12, which is a schematic diagram showing the structure of a computer device according to an embodiment of the present application. Specifically, the computer device 1200 includes a central processing unit (CPU) 1201, a system memory 1204 including a random access memory 1202 and a read-only memory 1203, and a system bus 1205 connecting the system memory 1204 and the CPU 1201. The computer device 1200 further includes a basic input / output system (I / O system) 1206 that supports information transmission between devices within the computer device, and a mass storage device 1207 that stores an operating system 1213, applications 1214, and other program modules 1215.

[0173] The basic input / output system 1206 includes a display 1208 for displaying information and input devices 1209, such as a mouse or keyboard, for inputting information by a user. Both the display 1208 and the input devices 1209 are connected to the central processing unit 1201 via an input / output controller 1210, which is connected to the system bus 1205. The basic input / output system 1206 may further include an input / output controller 1210 for receiving and processing input from a number of other devices, such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 1210 may provide output to a display, a printer, or other type of output device.

[0174] The mass storage device 1207 is connected to the central processing unit 1201 via a mass storage controller (not shown) connected to the system bus 1205. The mass storage device 1207 and its associated computer device-readable media provide non-volatile storage for the computer device 1200. That is, the mass storage device 1207 may include a computer-readable medium (not shown), such as a hard disk or a read-only compact disk drive.

[0175] Without loss of generality, the computer-readable media may include computer storage media and communication media. Computer storage media includes any method or technological implementation of volatile and nonvolatile, removable and non-removable media for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes random access memory (RAM), read-only memory (ROM), flash memory or other solid-state storage technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, tape cartridge, tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that the computer device is not limited to the above. The system memory 1204 and mass storage device 1207 may be collectively referred to as memory.

[0176] The memory stores one or more programs, which are configured to be executed by one or more central processing units 1201, and which include instructions for implementing the above-mentioned methods, and the central processing unit 1201, by executing the one or more programs, implements the methods provided in each of the above-mentioned method embodiments.

[0177] According to various embodiments of the present application, the computing device 1200 may also be implemented via a network, such as the Internet, to connect to a remote computing device connected to the network, i.e., the computing device 1200 may be connected to a network 1212 via a network interface unit 1211 connected to the system bus 1205, or may use the network interface unit 1211 to connect to another type of network or remote computer system (not shown).

[0178] The memory further includes one or more programs, the one or more programs being stored in the memory, and the one or more programs including steps executed by a computer device in the methods provided in the embodiments of the present application.

[0179] An embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, code set or instruction set is loaded and executed by the processor to realize the 3D medical image recognition method described in the above aspect.

[0180] An embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program including computer instructions stored in a computer-readable storage medium, a processor of a computer device reading the computer instructions from the computer-readable storage medium, and the processor executing the computer instructions to realize the method for recognizing 3D medical images according to the above aspect.

[0181] Those skilled in the art will understand that all or some of the steps in the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, which may be a computer-readable storage medium included in the memory in the above embodiments, or a computer-readable storage medium that is not built into the terminal and exists separately. An embodiment of the present application provides a computer-readable storage medium, which stores at least one instruction, at least one program, code set, or instruction set, and which is loaded and executed by the processor to realize the 3D medical image recognition method according to the above aspect.

[0182] In some embodiments, the computer-readable storage medium may include a ROM, a RAM, a solid-state hard disk (SSD), an optical disk, etc. Here, the RAM may include a resistance random access memory (ReRAM) and a dynamic random access memory (DRAM). The numbers of the above embodiments of the present application are for illustrative purposes only and do not represent the superiority or inferiority of the embodiments.

[0183] Those skilled in the art can understand that all or part of the steps in the above embodiments may be performed by hardware, or may be completed by instructing relevant hardware through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may be a read-only memory, a magnetic disk, an optical disk, etc.

[0184] The above are only some examples of the present application, and are not intended to limit the present application, and any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. 1. A method for recognizing three-dimensional (3D) medical images implemented by a computing device, comprising: In the i-th feature extraction process, a step of performing a view rearrangement process on the (i-1)th 3D medical image feature to obtain a two-dimensional (2D) image feature, wherein the (i-1)th 3D medical image feature is a feature obtained by performing the (i-1)th feature extraction on the 3D medical image, and the different 2D image feature is a feature in a different view of the (i-1)th 3D medical image feature; performing a semantic feature extraction process on each of the 2D image features to obtain image semantic features in different views; performing a feature fusion process on the image semantic features in different views to obtain an i-th 3D medical image feature; a step of performing image recognition processing based on the Ith 3D medical image feature obtained by the Ith feature extraction, and obtaining an image recognition result for the 3D medical image, where i is a sequentially increasing positive integer, 1 < i ≦ I, and I is a positive integer.

2. performing semantic feature extraction on each of the 2D image features to obtain image semantic features in different views, performing a spatial feature extraction process on the 2D image features to obtain 2D image spatial features; performing semantic feature extraction processing on the 2D image spatial features based on a main view and an auxiliary view to obtain the image semantic features, wherein the main view is a view corresponding to the 2D image features, and the auxiliary view is a 3D view different from the main view; The method for recognizing 3D medical images according to claim 1 .

3. The step of performing spatial feature extraction processing on the 2D image features to obtain 2D image spatial features includes: performing a window division process on the 2D image features to obtain local 2D image features corresponding to each of N windows, where the N windows do not overlap each other and N is a positive integer greater than 1; performing a feature extraction process on the N local 2D image features to obtain 2D image window features; performing a window rearrangement process on the N windows, and performing a feature extraction process on the 2D image window features corresponding to each of the N windows after the window rearrangement to obtain 2D image spatial features, wherein the window rearrangement is used to change the spatial positions of the N windows; The method for recognizing 3D medical images according to claim 2.

4. The step of performing a feature extraction process on the 2D image window features corresponding to each of the N windows after the window rearrangement to obtain 2D image space features includes: performing a self-attention process on the 2D image window features corresponding to each of the N windows after the window rearrangement to obtain self-attention features corresponding to each of the N windows; performing a feature fusion process on the N self-attention features to obtain second image window interior features; performing a position inversion process on the second image window internal features, and performing a convolution process on the second image window internal features after the position inversion to obtain second image window interaction features; performing feature extraction processing on the second image window interaction features by a multi-layer perceptron (MLP) to obtain the 2D image space features; The method for recognizing 3D medical images according to claim 3 .

5. The step of performing a feature extraction process on the N local 2D image features to obtain a 2D image window feature includes: performing a self-attention process on the N local 2D image features to obtain self-attention features corresponding to each of the N local 2D image features; performing a feature fusion process on the N self-attention features to obtain a first image window interior feature; performing a convolution process on the first image window internal features to obtain first image window interaction features; performing feature extraction processing on the first image window interaction feature by a multi-layer perceptron (MLP) to obtain the 2D image window feature; The method for recognizing 3D medical images according to claim 3 .

6. The step of performing a self-attention process on the N local 2D image features to obtain a self-attention feature corresponding to each of the N local 2D image features includes: performing self-attention processing based on a query term Q, a key term K, and a value term V corresponding to the local 2D image features to obtain self-attention features of the N local 2D image features; The method for recognizing 3D medical images according to claim 5.

7. the feature extraction process includes a feature encoding process or a feature decoding process, wherein the K value of the key item K in the t-th feature decoding process is obtained by fusion based on the K value in the t-1-th feature decoding and the K value in the corresponding feature encoding process; the V value of the value item V in the t-th feature decoding process is obtained by fusion based on the V value in the t-1-th feature decoding and the V value in the corresponding feature encoding process; and the Q value of the query item Q in the t-th decoding process is the Q value in the t-1-th feature decoding. The method for recognizing 3D medical images according to claim 6.

8. performing a semantic feature extraction process on the 2D image space features based on the main view and the auxiliary view to obtain the image semantic features, performing a feature fusion process on the 2D image spatial features and the position-coding features to obtain first image semantic features, wherein the position-coding features are used to indicate position information corresponding to the 2D image features; performing semantic feature extraction processing on the first image semantic features by MLP in the main view to obtain main image semantic features; performing a semantic feature extraction process on the first image semantic features by the MLP in the auxiliary view to obtain auxiliary image semantic features; performing a feature fusion process on the main image semantic features and the auxiliary image semantic features to obtain the image semantic features; The method for recognizing 3D medical images according to claim 2.

9. performing a feature fusion process on the image semantic features in the different views to obtain an i-th 3D medical image feature, performing a fusion process on the image semantic features and view features to obtain view image semantic features; performing a feature fusion process on each of the view image semantic features to obtain the i-th 3D medical image feature; The method for recognizing 3D medical images according to claim 1 .

10. performing a semantic feature extraction process on each of the 2D image features to obtain image semantic features in different views, performing semantic feature extraction processing on the 2D image features in each view using feature extraction networks corresponding to the same network parameters, to obtain the image semantic features in different views; The method for recognizing 3D medical images according to claim 1 .

11. the feature extraction process comprises a feature encoding process or a feature decoding process, the feature encoding process comprises a downsampling process for 3D medical image features, and the feature decoding process comprises an upsampling process for 3D medical image features; The method for recognizing a 3D medical image includes: performing image recognition processing based on the I-th 3D medical image feature obtained by the I-th feature extraction; and before obtaining an image recognition result for the 3D medical image, When the upsampling result reaches the original size, the 3D medical image feature obtained by the extraction is determined as the I-th 3D medical image feature obtained by the I-th feature extraction. The method for recognizing 3D medical images according to claim 1 .

12. The 3D medical image is a computed tomography (CT) image, a magnetic resonance imaging (MRI) image, or a positron emission tomography (PET) image. The method for recognizing 3D medical images according to claim 1 .

13. 1. A three-dimensional (3D) medical image recognition device, comprising: a view rearrangement module configured to perform a view rearrangement process on an (i-1)th 3D medical image feature in an i-th feature extraction process to obtain a 2D image feature, wherein the (i-1)th 3D medical image feature is a feature obtained by performing an (i-1)th feature extraction on the 3D medical image, and different 2D image features are features in different views of the (i-1)th 3D medical image feature; a feature extraction module configured to perform a semantic feature extraction process on each of the 2D image features to obtain image semantic features in different views; a feature fusion module configured to perform a feature fusion process on the image semantic features in different views to obtain an i-th 3D medical image feature; an image recognition module configured to perform image recognition processing based on the Ith 3D medical image feature obtained by the Ith feature extraction, and obtain an image recognition result for the 3D medical image, where i is a sequentially increasing positive integer, 1 < i ≦ I, and I is a positive integer.

14. A computer device comprising:

13. A computing device comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and wherein the at least one instruction, the at least one program, code set or instruction set is loaded and executed by the processor to implement the method for recognition of 3D medical images according to any one of claims 1 to 12.

15. A computer program comprising:

13. A computer program product comprising computer instructions, the computer instructions being stored in a computer-readable storage medium, the computer instructions being read by a processor of a computing device from the computer-readable storage medium, and the processor executing the computer instructions to realize the method for recognizing 3D medical images according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Three-dimensional medical image unsupervised feature extraction system

    CN113889235A

  • Image-processing method and apparatus for object detection or identification

    EP3815617A1