Task processing method and system based on three-dimensional medical image, terminal and medium
By extracting the global spatial structure and two-dimensional slice feature vectors from three-dimensional medical images, calculating weights, and fusing them to generate image fusion feature vectors, the problem of existing models being unable to effectively utilize three-dimensional image information is solved, achieving more accurate task processing results.
Patent Information
- Application Number
- CN202511038878.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-11
AI Technical Summary
Existing medical big data models cannot effectively utilize the spatial structure and hierarchical information in 3D medical images, resulting in limited analytical performance and difficulty in achieving accurate task results.
By acquiring task instructions and 3D medical images, global spatial structure feature vectors and 2D slice feature vectors are extracted. The weights of the 2D slices are calculated and weighted and aggregated to generate image fusion feature vectors, which are then input into a large language model for task processing.
It enriches the feature vector representation of 3D medical images, improves the accuracy of task results, and especially enhances the analytical performance in complex multimodal tasks.
Smart Images

Figure CN120931996A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a task processing method, system, terminal, and medium based on three-dimensional medical images. Background Technology
[0002] With the widespread application of 3D medical images (such as CT and MRI) in clinical diagnosis and treatment, the demand for their precise analysis is constantly increasing. Existing medical AI models are mostly task-specific and difficult to generalize to complex multimodal tasks. In recent years, the rise of Multimodal Large Language Models (MLLMs) has provided new ideas for medical image analysis. However, most existing medical MLLMs only process 2D images and cannot effectively utilize the spatial structure and hierarchical information contained in 3D images, resulting in limited analytical performance. The root of this limitation lies in the fact that if a 3D encoder is used alone to extract global spatial feature vectors, slice details are easily overlooked; if only a 2D encoder is used for layer-by-layer processing, it is difficult to model cross-slice spatial relationships. This deficiency in representation ability severely restricts the accuracy of task results.
[0003] Therefore, existing technologies have shortcomings and need to be improved and developed. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a task processing method, system, terminal and medium based on three-dimensional medical images, in order to address the above-mentioned deficiencies of the prior art and solve the problem that the prior art seriously restricts the accuracy of task results.
[0005] The technical solution adopted by this invention to solve the technical problem is as follows:
[0006] In a first aspect, embodiments of the present invention provide a task processing method based on three-dimensional medical images, the method comprising:
[0007] Acquire task instructions and three-dimensional medical images to be processed, wherein the three-dimensional medical images are composed of a continuous sequence of two-dimensional slices;
[0008] The task instructions are processed into task semantic feature vectors, and a global spatial structure feature vector and a two-dimensional slice feature vector sequence are obtained based on the three-dimensional medical image.
[0009] The weights of each two-dimensional slice are calculated based on the task semantic feature vector. The two-dimensional slice feature vector sequence is weighted and aggregated to generate a global two-dimensional feature vector. The global two-dimensional feature vector is fused with the global spatial structure feature vector to obtain the image fusion feature vector.
[0010] The image fusion feature vector and the task semantic feature vector are input into the large language model to obtain the task result.
[0011] In one implementation, a sequence of global spatial structure feature vectors and two-dimensional slice feature vectors is obtained based on the three-dimensional medical image, including:
[0012] The three-dimensional medical image is input into the global spatial structure feature vector extraction branch and the slice feature vector extraction branch;
[0013] A global spatial feature vector is generated by the first feature vector extraction module in the global spatial structure feature vector extraction branch, and the global spatial feature vector is output through a three-dimensional connector.
[0014] Each two-dimensional slice is processed by the second feature vector extraction module in the slice feature vector extraction branch to generate a two-dimensional slice feature vector sequence, and the two-dimensional slice feature vector sequence is output through a two-dimensional connector.
[0015] In one implementation, the weights of each two-dimensional slice are calculated based on the task semantic feature vector, including:
[0016] The task semantic feature vector, the global spatial structure feature vector, and the two-dimensional slice feature vector sequence are input into the inter-slice attention scoring module.
[0017] The attention scoring module is used to process and output the weights of each two-dimensional slice.
[0018] In one implementation, the attention scoring module is used to process and output the weights of each two-dimensional slice, including:
[0019] Within the inter-slice attention scoring module, the local feature vector corresponding to each two-dimensional slice is extracted from the global spatial structure feature vector, and the local feature vector and the two-dimensional slice feature vector of each two-dimensional slice are concatenated to obtain the slice fusion feature vector corresponding to each two-dimensional slice.
[0020] The task semantic feature vector is pooled to obtain the task pooling vector;
[0021] The original attention score for each 2D slice is obtained by performing a dot product between the fused feature vector and the task pooling vector.
[0022] All the original attention scores are processed to obtain a slice weight distribution vector, where each weight in the slice weight distribution vector represents the importance of the corresponding two-dimensional slice to the current task.
[0023] In one implementation, all the raw attention scores are processed to obtain a slice weight distribution vector, including:
[0024] The softmax function is used to normalize all the original attention scores to obtain the slice weight distribution vector; or,
[0025] The Sigmoid activation function is used to convert all the original attention scores into slice weight distribution vectors.
[0026] In one implementation, fusing the global two-dimensional feature vector with the global spatial structure feature vector to obtain an image fusion feature vector includes:
[0027] The global two-dimensional feature vector and the global spatial structure feature vector are input into a preset fusion module;
[0028] The fusion module is used to concatenate the global two-dimensional feature vector and the global spatial structure feature vector along the feature vector dimension to obtain the image fusion feature vector.
[0029] In one embodiment, the first feature vector extraction module is a three-dimensional image encoder or a neural network structure with the ability to extract global spatial structure feature vectors, and the second feature vector extraction module is a two-dimensional image encoder or a neural network structure with the ability to extract two-dimensional slice feature vector sequences.
[0030] Secondly, embodiments of the present invention also provide a task processing system based on three-dimensional medical images, the system comprising:
[0031] The data acquisition module is used to acquire task instructions and three-dimensional medical images to be processed, wherein the three-dimensional medical images are composed of a continuous two-dimensional slice sequence;
[0032] The feature vector extraction module is used to process the task instructions into task semantic feature vectors and obtain global spatial structure feature vectors and two-dimensional slice feature vector sequences based on the three-dimensional medical images.
[0033] The feature vector fusion module is used to calculate the weight of each two-dimensional slice based on the task semantic feature vector, aggregate the two-dimensional slice feature vector sequence to generate a global two-dimensional feature vector, and fuse the global two-dimensional feature vector with the global spatial structure feature vector to obtain the image fusion feature vector.
[0034] The task result output module is used to input the image fusion feature vector and the task semantic feature vector into the large language model to obtain the task result.
[0035] Thirdly, embodiments of the present invention also provide a terminal, the terminal comprising: a memory, a processor, and a task processing program based on three-dimensional medical images stored in the memory and executable on the processor, wherein the task processing program based on three-dimensional medical images, when executed by the processor, implements the steps of the task processing method based on three-dimensional medical images as described above.
[0036] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a task processing program based on three-dimensional medical images, the task processing program being executable to implement the steps of the task processing method based on three-dimensional medical images as described above.
[0037] The beneficial effects of this invention are as follows: This invention acquires task instructions and a three-dimensional medical image to be processed, the three-dimensional medical image being composed of a continuous sequence of two-dimensional slices; processes the task instructions into task semantic feature vectors, and obtains a global spatial structure feature vector and a sequence of two-dimensional slice feature vectors based on the three-dimensional medical image; calculates the weights of each two-dimensional slice based on the task semantic feature vector, and aggregates the two-dimensional slice sequence to generate a global two-dimensional feature vector; fuses the global two-dimensional feature vector with the global spatial structure feature vector to obtain an image fusion feature vector; inputs the image fusion feature vector and the task semantic feature vector into a large language model to obtain the task result. This invention effectively enriches the feature vector representation of three-dimensional medical images by fusing the global spatial structure and global two-dimensional feature vectors, thereby improving the accuracy of the task result. Attached Figure Description
[0038] Figure 1 This is a flowchart of a preferred embodiment of the task processing method based on three-dimensional medical images in this invention.
[0039] Figure 2 This is a schematic diagram of the data processing of the inter-slice attention scoring module in this invention.
[0040] Figure 3 This is a schematic diagram of a data processing flow according to the present invention.
[0041] Figure 4 This is a schematic diagram of a preferred embodiment of the task processing system based on three-dimensional medical images in this invention.
[0042] Figure 5 This is a block diagram of the terminal principle of the present invention. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0044] With the widespread application of 3D medical images (such as CT and MRI) in clinical diagnosis and treatment, the demand for their precise analysis is constantly increasing. Existing medical AI models are mostly task-specific and difficult to generalize to complex multimodal tasks. In recent years, the rise of Multimodal Large Language Models (MLLMs) has provided new ideas for medical image analysis. However, most existing medical MLLMs only process 2D images and cannot effectively utilize the spatial structure and hierarchical information contained in 3D images, resulting in limited analytical performance. The root of this limitation lies in the fact that if a 3D encoder is used alone to extract global spatial feature vectors, slice details are easily overlooked; if only a 2D encoder is used for layer-by-layer processing, it is difficult to model cross-slice spatial relationships. This deficiency in representation ability severely restricts the accuracy of 3D medical imaging tasks.
[0045] To address the aforementioned deficiencies in existing technologies, this invention provides a task processing method, system, terminal, and storage medium based on three-dimensional medical images. The method includes: acquiring task instructions and a three-dimensional medical image to be processed, wherein the three-dimensional medical image is composed of a sequence of continuous two-dimensional slices; processing the task instructions into task semantic feature vectors, and obtaining a global spatial structure feature vector and a sequence of two-dimensional slice feature vectors based on the three-dimensional medical image; calculating the weights of each two-dimensional slice based on the task semantic feature vectors, weighted aggregating the two-dimensional slice sequence to generate a global two-dimensional feature vector, fusing the global two-dimensional feature vector with the global spatial structure feature vector to obtain an image fusion feature vector; and inputting the image fusion feature vector and the task semantic feature vector into a large language model to obtain the task result. This invention effectively enriches the feature vector representation of three-dimensional medical images by fusing the global spatial structure and global two-dimensional feature vectors, thereby improving the accuracy of three-dimensional medical image tasks.
[0046] Please see Figure 1 The task processing method based on three-dimensional medical images described in this embodiment of the invention includes the following steps:
[0047] Step S100: Obtain task instructions and three-dimensional medical images to be processed, wherein the three-dimensional medical images are composed of a continuous two-dimensional slice sequence.
[0048] Specifically, task instruction x T The commands are in natural language format, such as "Please generate a lung CT diagnostic report" or "Are there any lung nodules?" Users can customize the task commands as needed.
[0049] Please see Figure 1 The task processing method based on three-dimensional medical images described in this embodiment of the invention further includes the following steps:
[0050] Step S200: Process the task instruction into a task semantic feature vector, and obtain a global spatial structure feature vector and a two-dimensional slice feature vector sequence based on the three-dimensional medical image.
[0051] Specifically, using a pre-trained text encoder f T Task instruction x T Processed into task semantic feature vector z T This process can be represented as: Where R is the set of real numbers, d T This represents the dimension of the task semantic feature vector. The task semantic feature vector subsequently guides the aggregation process of image features.
[0052] In one implementation, a sequence of global spatial structure feature vectors and two-dimensional slice feature vectors is obtained based on the three-dimensional medical image, including:
[0053] The three-dimensional medical image is input into the global spatial structure feature vector extraction branch and the slice feature vector extraction branch;
[0054] A global spatial feature vector is generated by the first feature vector extraction module in the global spatial structure feature vector extraction branch, and the global spatial feature vector is output through a three-dimensional connector.
[0055] Each two-dimensional slice is processed by the second feature vector extraction module in the slice feature vector extraction branch to generate a two-dimensional slice feature vector sequence, and the two-dimensional slice feature vector sequence is output through a two-dimensional connector.
[0056] Specifically, this invention comprises two branches, which extract the global spatial structure feature vector of the three-dimensional medical image and the two-dimensional slice vector of each two-dimensional slice in the three-dimensional medical image, respectively. The first feature vector extraction module is a three-dimensional image encoder or a neural network structure capable of extracting global spatial structure feature vectors, and the second feature vector extraction module is a two-dimensional image encoder or a neural network structure capable of extracting two-dimensional slice feature vector sequences.
[0057] When the first feature vector extraction module is a 3D image encoder The process of extracting the global spatial structure feature vector can be represented as follows: Where L is the number of spatial units and d3 is the length of the global spatial structure feature vector.
[0058] When the first feature vector extraction module is a neural network structure with the ability to extract global spatial structure feature vectors, it can be any one of Swin-UNETR, nnUNet-3D, or ViT3D.
[0059] When the second feature vector extraction module is a two-dimensional image encoder The process of processing each two-dimensional slice to generate the original slice feature vector set can be represented as:
[0060]
[0061] Where d2 is the length of the two-dimensional feature vector, and the sequence of two-dimensional slice feature vectors is represented as follows: N is the total number of two-dimensional slices.
[0062] When the second feature vector extraction module is a neural network structure capable of extracting two-dimensional slice feature vector sequences, it can be any one of ResNet-50, Swin-Tiny, DenseNet, and RadImageNet.
[0063] Please see Figure 1 The task processing method based on three-dimensional medical images described in this embodiment of the invention further includes the following steps:
[0064] Step S300: Calculate the weight of each two-dimensional slice based on the task semantic feature vector, aggregate the two-dimensional slice sequence in a weighted manner to generate a global two-dimensional feature vector, and fuse the global two-dimensional feature vector with the global spatial structure feature vector to obtain an image fusion feature vector.
[0065] Specifically, three-dimensional medical images consist of a continuous sequence of two-dimensional slices, typically numerous, but only a few slices contain crucial pathological information. Therefore, this invention first calculates the weight of each two-dimensional slice based on the task semantic feature vector. The weight represents the importance of the corresponding two-dimensional slice to the current task. A high weight means that the two-dimensional slice is highly likely to contain crucial pathological information, while a low weight means that the two-dimensional slice is highly unlikely to contain pathological information or has a low probability of containing pathological information. After obtaining the weight of each two-dimensional slice, a weighted aggregation is performed to obtain a global two-dimensional feature vector. This global two-dimensional feature vector no longer represents a single level but highlights the task-related information in the two-dimensional feature sequence vector. The global two-dimensional feature vector is fused with the global spatial structure feature vector to obtain a comprehensive representation that integrates task-related visual information and spatial structure information, i.e., an image fusion feature vector. Feature fusion is a crucial step in this invention. If one only knows that a certain two-dimensional slice has abnormal pixels (two-dimensional features) but does not know its specific location and extent within the organ (spatial structure), or only knows the spatial structure information but not the specific information of the two-dimensional slice, it is insufficient to make a comprehensive and accurate judgment. The fusion of feature vectors takes into account both the subtle textures and dynamic changes within the lesion and the macroscopic morphology, location, size, and proximity to surrounding key structures in three-dimensional space, which can effectively improve the accuracy of task results.
[0066] In one implementation, the weights of each two-dimensional slice are calculated based on the task semantic feature vector, including:
[0067] The task semantic feature vector, the global spatial structure feature vector, and the two-dimensional slice feature vector sequence are input into the inter-slice attention scoring module.
[0068] The attention scoring module is used to process and output the weights of each two-dimensional slice.
[0069] Specifically, the task semantic feature vector, global spatial structure feature vector, and two-dimensional slice feature vector sequence are first input into the inter-slice attention scoring module (TG-IS module). The inter-slice attention scoring module then extracts the local feature vector corresponding to each two-dimensional slice from the global spatial structure feature vector. Finally, each local feature vector is concatenated with the corresponding two-dimensional slice feature vector to obtain the slice fusion feature vector for each two-dimensional slice. This process can be represented as: in, To extract the global spatial structure feature vector z 3D The local feature vector extracted is the one corresponding to the j-th two-dimensional slice, where j = 1, ..., N, and N is the total number of two-dimensional slices. Let be the two-dimensional slice feature vector corresponding to the j-th two-dimensional slice, and d be the dimension of the slice fusion feature vector, where d = d3 + d2 or d is normalized to a unified dimension through linear mapping.
[0070] Next, the task semantic feature vector is pooled to obtain the task pooling vector. This process can be represented as:
[0071] Then, the feature vector and task pooling vector of each 2D slice are fused and multiplied by a dot product to obtain the original attention score for each 2D slice. This process is represented as...
[0072] Finally, all the raw attention scores are processed to obtain a slice weight distribution vector. Each weight in the slice weight distribution vector represents the importance of the corresponding two-dimensional slice to the current task. The slice weight distribution vector can be obtained by normalizing all the raw attention scores using the softmax function, or by converting all the raw attention scores into a slice weight distribution vector using the sigmoid activation function.
[0073] The process of obtaining the slice weight distribution vector using the softmax function can be represented as:
[0074]
[0075] in, s is the original attention score of the j-th slice. j is the weight of the j-th slice, and k is the index variable used to traverse all two-dimensional slices.
[0076] Slice weight distribution vector s = [s1, ..., s2] N The value represents the importance of each 2D slice to the current task. The Softmax function is used to generate a probability distribution, which essentially forces relative competition among all slices.
[0077] The method using the sigmoid activation function can be expressed as: This approach no longer forces the normalization of slice attention, but allows multiple slices to obtain relatively independent and non-exclusive attention weights, making it suitable for medical image tasks with multiple lesions and a wide distribution of information.
[0078] In one implementation, such as Figure 2As shown, the inter-slice attention scoring module extracts the local feature vector corresponding to each 2D slice from the global spatial structure feature vector. Each local feature vector is then concatenated with the corresponding 2D slice feature vector to obtain the slice fusion feature vector for each 2D slice. Next, the inter-slice attention scoring module performs pooling on the task semantic feature vector to obtain the task pooling vector. The fusion feature vector of each 2D slice and the task pooling vector are then multiplied by a dot product to obtain the original attention score for each 2D slice. Finally, the softmax function is used to normalize all the original attention scores to obtain the slice weight distribution vector.
[0079] After obtaining the slice weight distribution vector, the two-dimensional slice sequence is weighted and aggregated to generate a global two-dimensional feature vector. Specifically, based on the slice attention weight vector s j Aggregating the features of each two-dimensional slice yields a global two-dimensional feature vector. This process can be represented as:
[0080] In one implementation, the global two-dimensional feature vector is fused with the global spatial structure feature vector to obtain an image fusion feature vector, including:
[0081] The global two-dimensional feature vector and the global spatial structure feature vector are input into a preset fusion module;
[0082] The fusion module is used to concatenate the global two-dimensional feature vector and the global spatial structure feature vector along the feature vector dimension to obtain the image fusion feature vector.
[0083] Specifically, in the preset fusion module, the process of concatenating the global two-dimensional feature vector and the global spatial structure feature vector along the feature vector dimension can be represented as:
[0084]
[0085] In this way, an image fusion feature vector can be generated that combines lesion details (such as two-dimensional texture and density) and three-dimensional spatial information (such as location, shape and extent).
[0086] Please see Figure 1 The task processing method based on three-dimensional medical images described in this embodiment of the invention further includes the following steps:
[0087] Step S400: Input the image fusion feature vector and the task semantic feature vector into the large language model to obtain the task result.
[0088] Specifically, the image fusion feature vector and the task semantic feature vector are input into a large language model. After processing by the large language model, the lesion structure and details in the image fusion feature vector are analyzed, and the task type is identified based on the semantic features, outputting the task result x. R This process can be represented as: x R =f LLM ([Z I Z T ]). Among them, f LLM This is a large language model inference module (such as LLaMA, GPT structure). The result of this task is in the form of natural language text, which can be an automatic diagnosis report of medical images, a question-and-answer response to a doctor's question, a structured multi-label description, etc., without any restrictions.
[0089] Furthermore, to enable large language models to better understand the positional order between two-dimensional slices and improve their ability to perceive slice order, the two-dimensional slice features can be improved. Or slice fusion feature vector Introducing slice index position embedding vector PE j .
[0090] When in two-dimensional slice feature vector Introducing slice index position embedding vector PE jAt this point, the location-enhanced 2D slice feature vector can be obtained. Subsequent processing will then be based on this location-enhanced 2D slice feature vector. Specifically, the task semantic feature vector, the global spatial structure feature vector, and the sequence of location-enhanced 2D slice feature vectors (containing all location-enhanced 2D slice feature vectors) are input into the inter-slice attention scoring module. The attention scoring module processes these vectors and outputs the weights of each 2D slice. This process includes: extracting the local feature vector corresponding to each 2D slice from the global spatial structure feature vector using the inter-slice attention scoring module, and concatenating the local feature vector of each 2D slice with the location-enhanced 2D slice feature vector to obtain the second slice fusion feature vector corresponding to each 2D slice; pooling the task semantic feature vector using the inter-slice attention scoring module to obtain the task pooling vector; performing a dot product between each second slice fusion feature vector and the task pooling vector to obtain the second attention raw score for each 2D slice; and processing all the second attention raw scores to obtain the second slice weight distribution vector, where each weight represents the importance of the corresponding 2D slice to the current task. All raw scores for the second attention are processed to obtain the second slice weight distribution vector, including: normalizing all raw scores using the softmax function; or converting all raw scores for the second attention into a second slice weight distribution vector using the sigmoid activation function. Based on the second slice weight distribution, all position-enhanced two-dimensional slice features are weighted and aggregated to obtain the second global two-dimensional feature vector. The second global two-dimensional feature vector and the global spatial feature vector are fused to obtain the second image fusion feature vector. The second image fusion feature vector and task semantic features are input into the large language model to obtain the second task result.
[0091] When fusing feature vectors in slices Introducing slice index position embedding vector PE j At this time, the location-enhanced slice fusion feature vector can be obtained, and this process can be represented as follows:
[0092] Among them, the slice index position embedding vector PE jThis can be achieved using methods such as sine and cosine positional encoding, learnable positional embedding, and relative positional encoding. Subsequent processing will then be based on the position-enhanced slice fusion feature vector. Specifically, the dot product of each position-enhanced slice fusion feature vector and the task pooling vector is performed to obtain the original third attention score for each 2D slice. All the original third attention scores are processed to obtain the third slice weight distribution vector, where each weight represents the importance of the corresponding 2D slice to the current task. This processing includes: normalizing all the original third attention scores using the softmax function; or converting all the original third attention scores into a third slice weight distribution vector using the sigmoid activation function. The weights in the third slice weight distribution vector are used to weight and aggregate the 2D slice feature vector sequence to generate a third global 2D feature vector. This third global 2D feature vector is then fused with the global spatial structure feature vector to obtain the third image fusion feature vector. Finally, the third image fusion feature vector and the task semantic feature vector are input into the large language model to obtain the third task result.
[0093] The above approach, by introducing slice index position embedding vectors into two-dimensional slice features or slice fusion feature vectors, helps large language models understand the impact of slice order on tasks, and is particularly effective in applications where sequence modeling (such as temporal and anatomical gradations) is important.
[0094] To further model long-range dependencies and contextual semantics between slices, a Transformer encoder can be set within the attention scoring module. After generating the location-enhanced slice fusion feature vector, it can be input into the Transformer encoder, which outputs the second enhanced slice feature vector. This process can be represented as: This Transformer encoder employs a multi-head self-attention mechanism to achieve non-local information aggregation, making it particularly suitable for analyzing structurally complex organs (such as the brain and lungs). Subsequent processing is based on the feature vector of the second enhanced slice. Specifically, the fused feature vector of each second enhanced slice is multiplied by the task pooling vector to obtain the raw fourth attention score for each 2D slice. All raw fourth attention scores are processed to obtain the fourth slice weight distribution vector, where each weight represents the importance of the corresponding 2D slice to the current task. The processing of all raw fourth attention scores to obtain the fourth slice weight distribution vector includes: normalizing all raw fourth attention scores using the softmax function; or converting all raw fourth attention scores into a fourth slice weight distribution vector using the sigmoid activation function. The four slice feature vectors are weighted and aggregated using the weights in the fourth slice weight distribution vector to generate the fourth global two-dimensional feature vector. The fourth global two-dimensional feature vector is then fused with the global spatial structure feature vector to obtain the fourth image fusion feature vector. The fourth image fusion feature vector and the task semantic feature vector are then input into the large language model to obtain the fourth task result.
[0095] The above approach can adapt to different clinical image scenarios and task requirements, and also provides a convenient foundation for subsequent model iteration and deployment.
[0096] A data processing flow of the present invention is as follows: Figure 3As shown, the task instructions are input into a text encoder to obtain a task semantic feature vector. The 3D medical image is input into a global spatial structure feature vector extraction branch and a slice feature vector extraction branch. The 3D image encoder in the global spatial structure feature vector extraction branch generates a global spatial feature vector, which is then output via a 3D connector. Each 2D slice is processed by a 2D image encoder in the slice feature vector extraction branch to generate a 2D slice feature vector sequence, which is then output via a 2D connector. The task semantic feature vector, global spatial feature vector, and 2D slice feature vector sequence are then input into a text-guided inter-slice attention scoring module for processing to obtain the weight of each 2D slice. The weighted aggregation of the 2D slice feature vector sequence generates a global 2D feature vector. This global 2D feature vector is fused with the global spatial structure feature vector to obtain an image fusion feature vector. The image fusion feature vector and the task semantic feature vector are then input into a large language model to obtain the task result. This invention combines global spatial structure feature vectors and two-dimensional slice feature vector sequences to extract complementary information from different perspectives. Combined with a text-guided slice scoring mechanism, it achieves fine perception and task response of three-dimensional medical images and is widely applicable to report generation and question-answering tasks.
[0097] In one embodiment, such as Figure 4 As shown, based on the above-described task processing method based on three-dimensional medical images, the present invention also provides a task processing system (Med-2E3) based on three-dimensional medical images, the system comprising:
[0098] The data acquisition module 100 is used to acquire task instructions and three-dimensional medical images to be processed, wherein the three-dimensional medical images are composed of a continuous two-dimensional slice sequence;
[0099] The feature vector extraction module 200 is used to process the task instructions into task semantic feature vectors and obtain global spatial structure feature vectors and two-dimensional slice feature vector sequences based on the three-dimensional medical images.
[0100] The feature vector fusion module 300 is used to calculate the weight of each two-dimensional slice based on the task semantic feature vector, weighted aggregate the two-dimensional slice feature vector sequence to generate a global two-dimensional feature vector, and fuse the global two-dimensional feature vector with the global spatial structure feature vector to obtain an image fusion feature vector.
[0101] The task result output module 400 is used to input the image fusion feature vector and the task semantic feature vector into the large language model to obtain the task result.
[0102] In one embodiment, the system further includes:
[0103] The image input unit is used to input the three-dimensional medical image into the global spatial structure feature vector extraction branch and the slice feature vector extraction branch;
[0104] The first extraction unit is used to generate a global spatial feature vector through the first feature vector extraction module in the global spatial structure feature vector extraction branch, and output the global spatial feature vector through the three-dimensional connector.
[0105] The second extraction unit is used to process each two-dimensional slice through the second feature vector extraction module in the slice feature vector extraction branch, generate a two-dimensional slice feature vector sequence, and output the two-dimensional slice feature vector sequence through a two-dimensional connector.
[0106] In one embodiment, the system further includes:
[0107] The first feature vector input unit is used to input the task semantic feature vector, the global spatial structure feature vector and the two-dimensional slice feature vector sequence into the inter-slice attention scoring module;
[0108] The weight generation unit is used to process the attention scoring module and output the weights of each two-dimensional slice.
[0109] In one embodiment, the system further includes:
[0110] The slice fusion feature vector generation unit is used to extract the local feature vector corresponding to each two-dimensional slice from the global spatial structure feature vector within the slice attention scoring module, and to concatenate the local feature vector of each two-dimensional slice with the two-dimensional slice feature vector to obtain the slice fusion feature vector corresponding to each two-dimensional slice.
[0111] The task pooling vector generation unit is used to perform pooling processing on the task semantic feature vector to obtain the task pooling vector;
[0112] The score generation unit is used to perform a dot product between the fused feature vector and the task pooling vector of each two-dimensional slice to obtain the original attention score of each two-dimensional slice.
[0113] The score processing unit processes all the original attention scores to obtain a slice weight distribution vector, where each weight in the slice weight distribution vector represents the importance of the corresponding two-dimensional slice to the current task.
[0114] In one embodiment, the system further includes:
[0115] A normalization unit is used to normalize all the original attention scores using the softmax function to obtain a slice weight distribution vector; or,
[0116] The score transformation unit is used to convert all the original attention scores into slice weight distribution vectors using the Sigmoid activation function.
[0117] In one embodiment, the system further includes:
[0118] The second feature vector input module is used to input the global two-dimensional feature vector and the global spatial structure feature vector into a preset fusion module;
[0119] The feature vector fusion module is used to concatenate the global two-dimensional feature vector and the global spatial structure feature vector along the feature vector dimension to obtain an image fusion feature vector.
[0120] In one embodiment, such as Figure 5 As shown, based on the task processing method of three-dimensional medical images, the present invention also provides a terminal, including a processor 10 and a memory 20. Figure 5 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0121] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code for installing the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a task processing program 30 based on three-dimensional medical images, which can be executed by the processor 10 to implement the task processing method based on three-dimensional medical images in this application.
[0122] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the task processing method based on three-dimensional medical images.
[0123] This invention also provides a computer-readable storage medium storing a task processing program based on three-dimensional medical images. When the task processing program based on three-dimensional medical images is executed by a processor, it implements the steps of any of the task processing methods based on three-dimensional medical images provided in this invention.
[0124] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0125] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the above system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0126] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0127] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0128] In the embodiments provided by this invention, it should be understood that the disclosed system / terminal device and method can be implemented in other ways. For example, the system / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units described above is only a logical functional division, and in actual implementation, it can be divided in other ways. For example, multiple units or components can be combined or integrated into another system, or some feature vectors can be ignored or not executed.
[0129] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical feature vectors. Such modifications or substitutions do not mean that the essence of the corresponding technical solutions deviates from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A task processing method based on three-dimensional medical images, wherein the feature vector is, The method includes: Acquire task instructions and three-dimensional medical images to be processed, wherein the three-dimensional medical images are composed of a continuous sequence of two-dimensional slices; The task instructions are processed into task semantic feature vectors, and a global spatial structure feature vector and a two-dimensional slice feature vector sequence are obtained based on the three-dimensional medical image. The weights of each two-dimensional slice are calculated based on the task semantic feature vector. The two-dimensional slice feature vector sequence is weighted and aggregated to generate a global two-dimensional feature vector. The global two-dimensional feature vector is fused with the global spatial structure feature vector to obtain the image fusion feature vector. The image fusion feature vector and the task semantic feature vector are input into the large language model to obtain the task result.
2. The task processing method based on three-dimensional medical images according to claim 1, wherein the feature vector is, Based on the three-dimensional medical images, a global spatial structure feature vector and a two-dimensional slice feature vector sequence are obtained, including: The three-dimensional medical image is input into the global spatial structure feature vector extraction branch and the slice feature vector extraction branch; A global spatial feature vector is generated by the first feature vector extraction module in the global spatial structure feature vector extraction branch, and the global spatial feature vector is output through a three-dimensional connector. Each two-dimensional slice is processed by the second feature vector extraction module in the slice feature vector extraction branch to generate a two-dimensional slice feature vector sequence, and the two-dimensional slice feature vector sequence is output through a two-dimensional connector.
3. The task processing method based on three-dimensional medical images according to claim 1, wherein the feature vector is, The weights of each two-dimensional slice are calculated based on the task's semantic feature vector, including: The task semantic feature vector, the global spatial structure feature vector, and the two-dimensional slice feature vector sequence are input into the inter-slice attention scoring module. The attention scoring module is used to process and output the weights of each two-dimensional slice.
4. The task processing method based on three-dimensional medical images according to claim 3, wherein the feature vector is, The attention scoring module is used to process and output the weights of each two-dimensional slice, including: Within the inter-slice attention scoring module, the local feature vector corresponding to each two-dimensional slice is extracted from the global spatial structure feature vector, and the local feature vector and the two-dimensional slice feature vector of each two-dimensional slice are concatenated to obtain the slice fusion feature vector corresponding to each two-dimensional slice. The task semantic feature vector is pooled to obtain the task pooling vector; The original attention score for each 2D slice is obtained by performing a dot product between the fused feature vector and the task pooling vector. All the original attention scores are processed to obtain a slice weight distribution vector, where each weight in the slice weight distribution vector represents the importance of the corresponding two-dimensional slice to the current task.
5. The task processing method based on three-dimensional medical images according to claim 4, wherein the feature vector is, All the original attention scores are processed to obtain the slice weight distribution vector, including: The softmax function is used to normalize all the original attention scores to obtain the slice weight distribution vector; or, The Sigmoid activation function is used to convert all the original attention scores into slice weight distribution vectors.
6. The task processing method based on three-dimensional medical images according to claim 1, wherein the feature vector is, The global two-dimensional feature vector is fused with the global spatial structure feature vector to obtain an image fusion feature vector, including: The global two-dimensional feature vector and the global spatial structure feature vector are input into a preset fusion module; The fusion module is used to concatenate the global two-dimensional feature vector and the global spatial structure feature vector along the feature vector dimension to obtain the image fusion feature vector.
7. The task processing method based on three-dimensional medical images according to claim 2, wherein the feature vector is, The first feature vector extraction module is a three-dimensional image encoder or a neural network structure with the ability to extract global spatial structure feature vectors, and the second feature vector extraction module is a two-dimensional image encoder or a neural network structure with the ability to extract two-dimensional slice feature vector sequences.
8. A task processing system based on three-dimensional medical images, wherein the feature vector is, include: The data acquisition module is used to acquire task instructions and three-dimensional medical images to be processed, wherein the three-dimensional medical images are composed of a continuous two-dimensional slice sequence; The feature vector extraction module is used to process the task instructions into task semantic feature vectors, and obtain global spatial structure feature vectors and two-dimensional slice feature vector sequences based on the three-dimensional medical images. The feature vector fusion module is used to calculate the weight of each two-dimensional slice based on the task semantic feature vector, aggregate the two-dimensional slice feature vector sequence to generate a global two-dimensional feature vector, and fuse the global two-dimensional feature vector with the global spatial structure feature vector to obtain the image fusion feature vector. The task result output module is used to input the image fusion feature vector and the task semantic feature vector into the large language model to obtain the task result.
9. A terminal whose feature vector is, The terminal includes: a memory, a processor, and a task processing program based on three-dimensional medical images stored in the memory and executable on the processor. When the task processing program based on three-dimensional medical images is executed by the processor, it implements the steps of the task processing method based on three-dimensional medical images as described in any one of claims 1-7.
10. A computer-readable storage medium, wherein its characteristic vector is, The computer-readable storage medium stores a task processing program based on three-dimensional medical images. When the task processing program based on three-dimensional medical images is executed by a processor, it implements the steps of the task processing method based on three-dimensional medical images as described in any one of claims 1-7.