Three-dimensional image data processing method, data processing method, device and storage medium
Through the dynamic memory update mechanism and three-dimensional convolution integrating multi-view and multi-stage image features, the problem of incomplete feature extraction in three-dimensional image data processing in the existing technology is solved, and more accurate and comprehensive feature expression is achieved, which improves the effect of medical image analysis.
Patent Information
- Application Number
- CN202510230436.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-02-28
AI Technical Summary
In the three-dimensional image data processing, especially in the field of medical imaging, it is difficult to effectively integrate multi-view and multi-stage spatial information, resulting in insufficient comprehensive and accurate feature extraction, and lack of fine adjustment in the joint training of multi-modal data, affecting the learning effect of the model.
The dynamic memory update mechanism is adopted, by extracting local features of multiple perspectives and fusion using memory tensors, combining three-dimensional convolution and gating mechanisms, the memory information is dynamically adjusted, and the image features of different perspectives and stages are integrated.
It improves the comprehensiveness and accuracy of feature expression of three-dimensional image data, enhances the learning ability and stability of the model in multi-dimensional data, and can more accurately capture and express key information in physiological processes.
Smart Images

Figure CN119723004B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a three-dimensional image data processing method, a data processing method, a device and a storage medium. Background Art
[0002] In today's digital age, numerous fields rely on the efficient processing and analysis of various data types, and 3D imaging data processing technology plays a crucial role. From medicine to industrial manufacturing, virtual reality, geological exploration, and many other industries, 3D imaging data is increasingly being used. The depth and breadth of information it contains significantly impacts decision-making accuracy and work efficiency.
[0003] For example, in the medical field, 3D imaging data contains a wealth of information and is crucial for early disease diagnosis, condition assessment, and treatment plan development. Accurate and comprehensive analysis of this 3D medical imaging data can help doctors more accurately identify lesions and determine the stage of disease progression, thereby improving treatment effectiveness and patient survival rates. Another example is that in industrial manufacturing, 3D imaging data is used for product inspection, quality control, and reverse engineering. By analyzing 3D images of products, surface defects and internal structural problems can be detected to ensure that product quality meets standards. In reverse engineering, 3D imaging data can be used to quickly obtain a 3D model of the product, improving design and manufacturing efficiency.
[0004] Therefore, how to efficiently process and analyze three-dimensional image data and accurately extract valuable image features has become a key issue in many fields. Summary of the Invention
[0005] The present invention provides a three-dimensional image data processing method, a data processing method, a device and a storage medium, so as to at least solve the problem of poor accuracy and comprehensiveness in extracting image features in the related art.
[0006] The present invention provides a three-dimensional image data processing method, comprising the following steps: obtaining three-dimensional images from multiple perspectives in three-dimensional image data; extracting local features of the three-dimensional images from multiple perspectives, and using the local features from multiple perspectives to update a memory tensor, wherein the memory tensor is used to store and fuse the local features from multiple perspectives; and aggregating the updated memory tensors to obtain image features of the three-dimensional image.
[0007] The present invention also provides a data processing method, comprising the following steps: obtaining the data to be processed and task requirements input by a user, wherein the data to be processed includes at least one of three-dimensional image data and text data; inputting the data to be processed and the task requirements into a data processing model, and the data processing model outputs an analysis result corresponding to the task requirement, wherein the data processing model includes an image encoding module, a text encoding module and a feature alignment module, the image encoding module processes the three-dimensional image data based on the above-mentioned three-dimensional image data processing method to obtain image features of the three-dimensional image, the text encoding module processes the text data based on the text encoder to obtain text features, and the feature alignment module is used to perform feature alignment on the image features and the text features, and determine the analysis result corresponding to the task requirement based on the result of the feature alignment.
[0008] The present invention also provides a three-dimensional image data processing device, including: a first acquisition module, used to obtain three-dimensional images of multiple perspectives in three-dimensional image data; an extraction module, used to extract local features of the three-dimensional images of multiple perspectives, and use the local features of multiple perspectives to update a memory tensor, wherein the memory tensor is used to store and fuse the local features of multiple perspectives and enhance the image feature expression of the three-dimensional image; an aggregation module, used to aggregate the updated memory tensor to obtain the image features of the three-dimensional image.
[0009] The present invention also provides a data processing device, including: a second acquisition module, used to obtain the data to be processed and task requirements input by the user, wherein the data to be processed includes at least one of three-dimensional image data and text data; an input module, used to input the data to be processed and the task requirements into a data processing model, and the data processing model outputs the analysis results corresponding to the task requirements, wherein the data processing model includes an image encoding module, a text encoding module and a feature alignment module, the image encoding module processes the three-dimensional image data based on the above-mentioned three-dimensional image data processing device to obtain image features of the three-dimensional image, the text encoding module processes the text data based on the text encoder to obtain text features, and the feature alignment module is used to perform feature alignment on the image features and the text features, and determine the analysis results corresponding to the task requirements based on the results of the feature alignment.
[0010] The present invention also provides an electronic device, comprising: a memory for storing a computer program; a processor for implementing the steps of the above-mentioned three-dimensional image data processing method, or the steps of the above-mentioned data processing method when executing the computer program.
[0011] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the three-dimensional image data processing method or the steps of the above-mentioned data processing method are implemented.
[0012] The present invention also provides a computer program product, including a computer program, which implements the steps of the three-dimensional image data processing method or the steps of the above-mentioned data processing method when the computer program is executed by a processor.
[0013] Through the present invention, by extracting local features of three-dimensional images from multiple perspectives, and using the local features of multiple perspectives to update the memory tensor, the updated memory tensor is aggregated to obtain the image features of the three-dimensional image. By dynamically updating the memory tensor using local features from different perspectives, the spatial information of three-dimensional images from different perspectives can be effectively captured and integrated, thereby fully enhancing the image feature expression capability of the three-dimensional image, and improving the comprehensiveness and accuracy of the image feature expression. Therefore, technical problems such as the poor accuracy and comprehensiveness of image feature extraction in related technologies can be solved, and the technical effect of enhancing the image feature expression capability of the three-dimensional image and improving the comprehensiveness and accuracy of the image feature expression can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0015] Figure 1 A flowchart of a three-dimensional image data processing method provided by an embodiment of the present invention;
[0016] Figure 2 A flowchart of a data processing method provided by an embodiment of the present invention;
[0017] Figure 3 A flowchart of a data processing method provided by a specific embodiment of the present invention;
[0018] Figure 4 A schematic diagram of a three-dimensional image data processing device provided by an embodiment of the present invention;
[0019] Figure 5 A schematic diagram of a data processing device provided by an embodiment of the present invention;
[0020] Figure 6 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0022] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.
[0023] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0024] Before describing the technical solution of the present invention, some technologies related to the solution of the present invention are first introduced.
[0025] In traditional computer vision, multi-view image analysis primarily studies the spatial and geometric relationships between images captured from different angles. These methods are often applied to tasks such as video understanding, 3D rendering, and image segmentation, integrating geometric features such as depth, shape, and texture to improve scene understanding and object classification. However, these methods mostly focus on analyzing spatial variations and lack attention to temporal dynamics or physiological dimensions, which are particularly important in medical imaging.
[0026] It should be noted that the solution of the present invention can process three-dimensional images, three-dimensional videos or text data in multiple fields, without any specific limitation. The following technical solutions of the present invention are described using the processing of medical data (such as three-dimensional medical imaging data and medical text data) in the medical field as an example.
[0027] In the field of medical imaging, multi-view analysis not only focuses on spatial perspectives but also needs to incorporate temporal dynamics and physiological and pathological variations of biological processes, such as the distribution of contrast agents within the body. These variations present unique challenges, as they require simultaneous distinction between normal and pathological structures. This is significantly different from typical computer vision tasks, as medical imaging requires analyzing cross-stage and cross-modal correlations to extract clinically meaningful insights. The complexity of this problem requires customized approaches to capture these characteristics.
[0028] Medical imaging diagnostic reports provide rich clinical information in text form by integrating spatial, temporal, and pathological correlations in multi-stage images. By mapping diagnostic reports to imaging stages, domain-specific knowledge can be deeply mined. Vision-language models have obvious advantages in combining image and text data, and can support tasks such as classification, retrieval, and generating explanations. In multi-stage radiological image analysis, temporal factors such as the dynamic changes of contrast agents are crucial for clinical decision-making, and this information is usually contained in diagnostic reports. By combining images with text descriptions, such models can learn the correlations between stages and extract key information about multi-stage dynamics, thereby improving the interpretation of complex imaging patterns and generating clinically meaningful conclusions.
[0029] Medical image-text alignment combines visual data (such as X-rays, MRI (Magnetic Resonance Imaging), and CT (Computed Tomography) images) with domain-specific textual knowledge (such as clinical notes and diagnostic reports) to improve the understanding and diagnosis of medical conditions. Pre-trained vision-language models (such as BiomedCLIP) have shown significant potential on large-scale biomedical image and text datasets, supporting tasks such as image retrieval, classification, and explanation generation. Furthermore, knowledge distillation techniques (such as the Self-Evolving Visual Transformer for Chest X-ray Diagnosis) incorporate medical expertise into visual models, enabling them to simultaneously interpret visual patterns and clinical context, achieving excellent performance in highly specialized domains. Models such as LLaVA-Med (Large Language and Vision Assistant for Medicine) further extend this concept, processing both clinical text and medical images simultaneously and explaining predictions based on medical knowledge, thereby improving model interpretability. While addressing data limitations and noisy annotation issues, recent advances in multimodal representation learning have also enhanced cross-modal retrieval and diagnostic reasoning capabilities.
[0030] For example, medical image analysis methods fail to fully exploit the rich temporal and spatial information inherent in contrast-enhanced images. Traditional multi-view methods primarily focus on geometric correlations but often overlook subtle differences in temporal and physiological changes, limiting their performance in clinical diagnosis and treatment.
[0031] The solution of the present invention can solve at least one of the following problems in the medical field:
[0032] 1. Related technologies face challenges in extracting features from high-dimensional data, particularly limitations in deeply mining multidimensional spatial information and integrating cross-view image data. In medical imaging in particular, effectively fusing information from different anatomical planes (such as axial, sagittal, and coronal) and extracting comprehensive features based on this information remains a pressing challenge. Related technologies often fail to fully capture the deep relationships between high-dimensional features, impacting the comprehensiveness of feature representation and the parsing power of the resulting model.
[0033] 2. Address several current technical challenges in multimodal image analysis, particularly accuracy and consistency issues when processing heterogeneous data (i.e., text and image data) and high-dimensional information. Related technologies typically rely on a single loss function for cross-modal alignment in the joint training of multimodal data such as images and text. However, this approach often fails to fully balance the contributions of each modality, especially when processing multi-viewpoint and cross-stage information in medical images. The lack of fine-tuning of the features of different modalities results in suboptimal model learning.
[0034] An embodiment of the present invention provides a three-dimensional image data processing method. The method is described in detail in conjunction with the execution flow of the three-dimensional image data processing method. The three-dimensional image data processing method of this embodiment is described using three-dimensional medical image processing in the medical field as an example.
[0035] like Figure 1 As shown, the three-dimensional image data processing method includes the following steps:
[0036] In step S101 , three-dimensional images of multiple viewing angles are acquired from three-dimensional image data.
[0037] The multiple perspectives may be, for example, different anatomical planes in the medical field, such as axial, coronal, and sagittal planes.
[0038] It is understandable that the embodiment of the present invention can obtain three-dimensional images of multiple viewing angles in the three-dimensional image data so as to perform feature processing thereon subsequently.
[0039] In step S102, local features of the three-dimensional images from multiple perspectives are extracted, and the memory tensor is updated using the local features from multiple perspectives, wherein the memory tensor is used to store and fuse the local features from multiple perspectives.
[0040] Among them, three-dimensional images from multiple perspectives include three-dimensional images from multiple stages. For example, if the current perspective is the axial perspective, the axial perspective can also include multiple stages, such as the plain scan stage, the arterial phase stage, the perfusion phase, etc.; the memory tensor is a shared, dynamically updateable data structure used to store and fuse information from all perspectives. It can also be regarded as a long-term memory representation of the input data by the model, which can transmit and integrate information at different time points (i.e., different stages) or between perspectives, better capture and utilize the complementary information between different perspectives, and avoid the deviation or information loss that may be caused by a single perspective.
[0041] It can be understood that the embodiments of the present invention can extract local features of three-dimensional images from multiple perspectives, and use the local features of multiple perspectives to update the memory tensor to fuse the information between multiple perspectives, avoid the deviation or information loss caused by using only a single perspective, and generate a more comprehensive and accurate image feature representation.
[0042] In addition, the embodiment of the present invention can extract local features of three-dimensional images from multiple perspectives through the Transformer network. Since local features have higher robustness than global features, they can more accurately locate and describe image details, and the data processing volume is also smaller. Therefore, the embodiment of the present invention only uses local features when updating the memory tensor, thereby reducing the data processing volume, improving data processing efficiency, and generating a more comprehensive and accurate image feature representation.
[0043] In an embodiment of the present invention, the memory tensor is updated using local features of multiple perspectives, including: obtaining local features of the three-dimensional image at the current stage under the current perspective; using the local features of the current stage under the current perspective to update the memory tensor stored in the previous stage under the current perspective, until the memory tensors at multiple stages under the current perspective are updated; after updating the memory tensor using the local features of multiple stages of the current perspective, using the local features of multiple stages of the next perspective to update the memory tensor stored in the current perspective, until the memory tensors of multiple perspectives are updated, to obtain the final updated memory tensor.
[0044] It can be understood that the embodiments of the present invention can obtain the local features of the three-dimensional image at the current stage under the current perspective, and use the local features of the current perspective at the current stage to update the memory tensor stored in the previous stage of the current perspective, until the memory tensors of multiple stages under the current perspective are updated, and then use the local features of multiple stages of the next perspective to update the memory tensor stored in the current perspective, until the local features of multiple perspectives are updated to the memory tensor, and the final updated memory tensor is obtained to fuse the spatial information and temporal information of the three-dimensional images of multiple perspectives and multiple stages, thereby improving the utilization efficiency of spatial and temporal information in the three-dimensional image processing process, the accuracy of feature extraction and the comprehensiveness of image feature expression, wherein multiple stages represent data at multiple time points.
[0045] Furthermore, it's important to note that in the medical field, three-dimensional medical imaging data at multiple stages contains rich information on changes in physiological characteristics. For example, by imaging the heart at different stages, physiological activities such as the heart's beating cycle and the contraction and relaxation of the myocardium are reflected in the images. By capturing local features from images at different stages and updating the memory tensor, it's possible to record changes in the heart's physiological state at different moments, such as cyclical changes in myocardial thickness and changes in cardiac chamber volume, thereby expressing the heart's physiological characteristics. In the diagnosis of cardiovascular disease, this capture of changes in the heart's physiological characteristics over time helps doctors determine whether heart function is normal and whether there are problems such as myocardial pathology. Furthermore, three-dimensional medical images from multiple perspectives showcase the physiological structure of human organs and tissues from different angles.
[0046] In some medical imaging procedures, contrast agents and other techniques are used to observe physiological processes. For example, in angiography, the flow and distribution of contrast agents within blood vessels reflect the physiological characteristics of blood circulation. When processing image data, the dynamic changes of the contrast agent at different stages of the image (such as the time it takes for the contrast agent to reach different vascular locations and its concentration distribution) are captured by local features. Through the updating and aggregation of memory tensors, the physiological characteristics of human blood circulation are expressed. This capture and expression of dynamic information about physiological processes is of great significance for the diagnosis of vascular diseases (such as stenosis and blockage).
[0047] By fusing local features of multi-view and multi-stage 3D medical images into a memory tensor, the resulting image features can comprehensively reflect the physiological characteristics of the human body. The memory tensor integrates information from different viewpoints and stages, allowing the generated image features to include not only the spatial structure of organs and tissues, but also information about how physiological processes change over time. This comprehensive representation capability enables a more comprehensive and accurate representation of the human body's physiological state, subsequently providing doctors with richer diagnostic evidence and helping to improve their understanding of physiological characteristics and the accuracy of disease diagnosis.
[0048] Specifically, for example, the current perspective is the first perspective V1, and the first perspective includes three stages, the first stage Sv1, the second stage Sv2, and the third stage Sv3. The initial memory tensor M0 is set to empty, and the local features of Sv1 of V1 are updated to the memory tensor. Then, the memory tensor M1 stored in Sv1 is updated according to the local features of Sv2 of V1 to obtain the memory tensor M2 stored in Sv2. The memory tensor M2 stored in Sv2 is updated according to the local features of Sv3 to obtain the memory tensor M3 stored in Sv3. And so on, the local features of each stage of the second perspective V2 are updated to the memory tensor M3 finally stored in V1. And so on, the updated memory tensors of all stages of all perspectives are obtained.
[0049] In an embodiment of the present invention, the local features of the current stage of the current perspective are used to update the memory tensor stored in the previous stage of the current perspective, including: using a gating mechanism and the local features of the current stage to update the memory tensor stored in the previous stage, and using the updated memory tensor of the previous stage as the memory tensor stored in the current stage, wherein the gating mechanism includes an update gate and a reset gate, the update gate is used to control the fusion ratio between the local features of the current stage and the memory tensor stored in the previous stage, and the reset gate is used to control the reset level of the memory tensor stored in the previous stage in the current stage.
[0050] It is understandable that the embodiment of the present invention can utilize the gating mechanism and the local features (F t ) Update the memory tensor stored in the previous stage (also called the previous memory M prev ), and the memory tensor updated in the previous stage is used as the memory tensor stored in the current stage, where the gating mechanism includes the update gate (z t ) and reset gate (r t ), the update gate is used to control the fusion ratio of the local features of the current stage and the memory tensor stored in the previous stage, and the reset gate is used to control the reset level of the memory tensor stored in the previous stage in the current stage, that is, it is used to control the role of the previous memory in the update of the current stage. The reset gate and the update gate jointly determine the degree of fusion of the local features of the current stage and the previous memory.
[0051] In an embodiment of the present invention, before using the gating mechanism and the local features of the current stage to update the memory tensor stored in the previous stage, it includes: calculating the value of the update gate and the value of the reset gate based on the local features of the current stage and the memory tensor stored in the previous stage; updating the memory tensor stored in the previous stage based on the value of the update gate, the value of the reset gate and the local features of the current stage, wherein the value of the update gate is the fusion ratio, and the value of the reset gate is the reset level.
[0052] It can be understood that the embodiment of the present invention can calculate the values of the update gate and the reset gate based on the local features of the current stage and the memory tensor stored in the previous stage, and update the memory tensor stored in the previous stage based on the value of the update gate, the value of the reset gate and the local features of the current stage, wherein the value of the update gate is the fusion ratio and the value of the reset gate is the reset level to determine the degree of fusion of the local features of the current stage and the memory tensor stored in the previous stage.
[0053] In an embodiment of the present invention, the value of the update gate and the value of the reset gate are calculated based on the local features of the current stage and the memory tensor stored in the previous stage, including: performing a three-dimensional convolution operation on the local features of the current stage to obtain a first feature; performing a three-dimensional convolution operation on the memory tensor stored in the previous stage to obtain a second feature; combining the first feature and the second feature to obtain a first combined feature, inputting the first combined feature into a first target activation function, and the first target activation function outputting the value of the update gate and the value of the reset gate.
[0054] The first target activation function may be a Sigmoid activation function, which is not specifically limited.
[0055] It can be understood that the values of the update gate and the reset gate of the embodiment of the present invention can be calculated through a 3D convolution operation. The three-dimensional convolution operation can be performed on the local features of the current stage and the memory tensor stored in the previous stage respectively to obtain the first feature and the second feature respectively. The first feature and the second feature are combined to obtain the first combined feature, and the first combined feature and the second combined feature are input into the first target activation function. The first target activation function outputs the value of the update gate and the value of the reset gate.
[0056] It should be noted that the values of the update gate and the reset gate in the embodiment of the present invention are between 0 and 1.
[0057] Among them, the update gate (z t ) and reset gate (r t ) is calculated using the following formula:
[0058] z t =σ(Conv3d z (F t )+Conv3d z (M prev ))⊙valid_view;
[0059] r t =σ(Conv3d r (F t )+Conv3d r (M prev ))⊙valid_view;
[0060] In the above formula, σ(·) represents the element-wise Sigmoid activation function, ⊙ represents the element-wise multiplication (this operation will be broadcasted on the batch dimension), and F t is the local feature of the current stage (also called the input feature of the current stage), M prev The memory tensor stored in the previous stage (also known as previous memory), Conv3d r and Conv3d z For three-dimensional (3D) convolution operations, valid_view is the validity mask of each batch of data. B is the batch size, D, H, and W are the spatial dimensions of the local features: depth, height, and width, respectively. C is the number of input channels. The update gate controls the ratio of the fusion between the input features of the current stage and the features in the previous memory. The update gate value is calculated through a 3D convolution operation, which takes into account the input features of the current stage and the features in the previous memory. It also combines the validity mask in each batch example to ensure that only data with valid viewpoints are processed.
[0061] The three-dimensional convolution (Conv3d) of the embodiment of the present invention is a convolution operation for processing 3D data, which is generally applied to images, videos and other data with three-dimensional structures. In convolutional neural networks (CNNs), 2D convolution is more widely used to process planar images, while 3D convolution can be extended to three dimensions of depth, width and height, thereby better processing space, time and other three-dimensional information. Therefore, the medical data processing of the present invention adopts 3D convolution operation.
[0062] The specific execution steps of the 3D convolution operation include:
[0063] Assume that there are the following input data (i.e., the local features of the current stage of the present invention or the memory tensor stored in the previous stage) and convolution kernel, the input tensor Among them, B is the batch size (batchsize), C in is the number of input channels (e.g. the number of grayscale channels of medical images), D, H, and W are the depth, height, and width of the input data, respectively;
[0064] Convolution kernel Among them, K d ,K h ,K w Represents the size of the convolution kernel in depth, height, and width, respectively. in is the number of input channels, C out is the number of output channels.
[0065] In 3D convolution, the convolution kernel K slides on the input tensor X (performing a sliding window operation) to generate the output tensor Y (i.e., the first feature and the second feature output by the embodiment of the present invention). Among them, D′, H′, W′ are the output depth, height, and width after the convolution operation. The output size is determined by setting the stride and padding.
[0066] For each position (b,c out ,d′,h′,w′), the convolution operation can be expressed as:
[0067]
[0068] The above formula describes how the convolution filter and the input tensor interact during the convolution process to generate the output tensor.
[0069] This is an element in the output tensor Y, representing the output channel c in batch b. out The corresponding depth d', height h', width w' position value. In other words, this is the result of the convolution operation;
[0070] This is a part of the input tensor X, that is, the position of the current convolution window on the input data. The convolution window (i.e., the convolution kernel) will slide on the input tensor, so d′+k d -1, h′+k h -1, and w′+k w -1These expressions represent the corresponding positions of the convolution kernel current position and the input data;
[0071] This is an element in the convolution kernel K, which indicates the specific position of the convolution kernel in the depth, width, and height directions, and distinguishes the input channel c in and output channel c out .
[0072] The 3D convolution operation described above specifically processes the 3D features of 3D medical images (including depth, width, and height). Specifically, each feature in 3D data often involves information from multiple dimensions. This information aggregation and updating requires convolution to capture the relationships between different spatial locations.
[0073] In an embodiment of the present invention, a gating mechanism and local features of the current stage are used to update a memory tensor stored in the previous stage, including: resetting the memory tensor stored in the previous stage using a reset gate; performing a three-dimensional convolution operation on the reset memory tensor stored in the previous stage to obtain a third feature; combining the first feature and the third feature to obtain a second combined feature, inputting the second combined feature into a second target activation function, and the second target activation function outputting a candidate memory tensor; and using an update gate to fuse the candidate memory tensor and the memory tensor stored in the previous stage.
[0074] The second target activation function is a nonlinear activation function, such as a hyperbolic tangent function.
[0075] It can be understood that the embodiment of the present invention can use the reset gate to reset the memory tensor stored in the previous stage, and perform a three-dimensional convolution operation on the reset memory tensor stored in the previous stage to obtain a third feature, combine the first feature and the third feature to obtain a second combined feature, input the second combined feature into the second target activation function, and the second target activation function outputs a candidate memory tensor. The candidate memory tensor and the memory tensor stored in the previous stage are fused using the update gate to obtain the updated memory tensor stored in the previous stage, that is, the memory tensor stored in the current stage.
[0076] Among them, the reset gate determines how to weight the current input and the previous memory, thereby adjusting the degree of memory retention and reset. When the value of the reset gate is close to 1, the previous memory will have a stronger influence on the current candidate memory; when it is close to 0, the previous memory has a weaker influence on the current update.
[0077] Specifically, the candidate memory tensor (also referred to as candidate memory) of the present invention is calculated by combining the local features of the current stage (also known as input features) with the memory tensor stored in the previous stage after reset (also known as previous memory). When generating candidate memories, it is first necessary to reset the previous memory so that it only retains the part of the information required for the current input. The reset operation is completed by multiplying with the reset gate, which adjusts how much previous memory information is retained.
[0078] Then, the candidate memory M candidate Generated through a 3D convolution operation, it not only considers the current input features, but also generates new memory representations by interacting with previous memories. Candidate memories are usually activated nonlinearly (such as the hyperbolic tangent function tanh) to enhance their expressive power in order to capture complex spatial and temporal features.
[0079] Candidate memory M candidate According to the current feature F t and the previous memory r after reset t ⊙M prev The calculation formula is:
[0080] M candidate =tanh(Conv3d candidate (F t +r t ⊙M prev ))⊙valid_view;
[0081] Among them, M candidate For candidate memory, Conv3d candidate is a 3D convolution operation, F t is the current input feature, r t To reset the gate value, M prev is the previous memory, r t ⊙M prev is the previous memory after reset, valid_view is the validity mask of each batch of data, and ⊙ represents element-by-element multiplication (the operation will be broadcasted on the batch dimension).
[0082] After generating the candidate memory, the previous memory is fused with the candidate memory, and the update gate controls the influence ratio of the input features of the current stage and the previous memory. Specifically, when z t When it approaches 1, the new candidate memory will have a greater impact on the final memory. t As it approaches 0, the previous memory is retained more. Through such a gating mechanism, the model can flexibly and dynamically adjust the memory between different time steps, thereby effectively capturing key information in multi-view images or multi-stage data processing.
[0083] The final memory update is calculated using the following formula: t =(1-z t )⊙M prev +z t ⊙M candidate .This formula uses the update gate z t To control the contribution ratio of the current feature and candidate memory to the final memory. t When it approaches 1, more information comes from the candidate memory M candidate ; When z t When it approaches 0, the memory update tends to keep the previous memory M prev .
[0084] In addition, it should be noted that after the memory update, layer normalization (LayerNorm) is performed. This operation is applied to the hidden dimension C, that is, the channel dimension, and the formula is expressed as:
[0085] M t =LayerNorm(M t );
[0086] Layer normalization helps stabilize the learning process, ensuring the numerical stability of memory and the efficiency of subsequent model training.
[0087] Finally, the updated memory tensor M t The shape of the input will be the same as that of the input, i.e. [B×C×D×H×W], so as to be used for subsequent feature aggregation and alignment tasks.
[0088] In the implementation of the present invention, after extracting the local features of the three-dimensional images of multiple viewing angles, the method further includes: using a viewing angle mask to filter out invalid viewing angles; and resetting the local features corresponding to the invalid viewing angles to zero.
[0089] It can be understood that the embodiments of the present invention can use the perspective mask to filter out invalid perspectives and reset the local features corresponding to the invalid perspectives to zero. Compared with the simple superposition of perspective or spatial information processing in traditional methods, the present invention uses dynamic masks to eliminate invalid perspectives, avoiding the interference of irrelevant data on the memory update process, thereby further improving the accuracy of image feature representation.
[0090] In summary, the embodiments of the present invention update the memory tensor through a dynamic memory update mechanism, which not only takes into account spatial relationships but also dynamically adjusts the memory information of each perspective through three-dimensional local features. This combines the current perspective with the previously memorized three-dimensional features, thereby enhancing the learning ability and stability of subsequent models in multi-dimensional data. This mechanism can dynamically update memory between different time steps and different perspectives, ensuring the effective fusion of cross-perspective information. In addition, a three-dimensional gating mechanism is introduced during the memory update process. By dynamically adjusting the three-dimensional update gate and reset gate, it accurately filters important information and suppresses noise and invalid features during the integration of multi-perspective and multi-stage image features.
[0091] The dynamic memory update mechanism is applicable to three-dimensional spatial data and can dynamically adjust the depth, width, and height dimensions to enhance the fusion effect of image features. The specific update process is as follows:
[0092] 1. At each stage of each view, the current input features (i.e., the local features of the current stage) are interacted with the previous memory (the memory tensor stored in the previous stage) to generate new candidate memories. This process is achieved through convolution operations. The convolution kernel extracts the spatial and temporal features between the current input and the historical memory. Under the action of the convolution kernel, these features can span the three dimensions of depth, width, and height, effectively capturing the associated information at each spatial location.
[0093] 2. 3D convolution processes the input features, extracting spatial and temporal correlations. The convolution kernel is designed to slide in three-dimensional space, capturing the spatial relationships of the input data in depth, width, and height. Through this convolution operation, the model can identify local features in the image and aggregate these features for further memory updates. This operation not only models the spatial structure of the image but also takes into account the relationships between different time steps, providing rich information for the dynamic memory network.
[0094] 3. When processing multi-view data, a validity mask is introduced to ensure that only valid viewpoint information contributes to memory updates. This is particularly important because, during multi-view image processing, some views may not be updated due to missing data or quality issues. The validity mask dynamically filters out valid viewpoint data, ensuring that only meaningful input information is used for memory updates. This is crucial for ensuring the effective fusion of 3D medical imaging data and preventing invalid data from interfering with the memory update process.
[0095] 4. The entire process, by combining 3D convolution, dynamic memory updating, and validity masking, fully considers information from all dimensions of three-dimensional space, ensuring the effective integration of multi-view and multi-stage medical imaging data in memory updating. This approach not only accurately captures spatial and temporal features, but also ensures the screening and enhancement of important information in complex data environments.
[0096] In step S103 , the updated memory tensors are aggregated to obtain image features of the three-dimensional image.
[0097] It can be understood that the embodiments of the present invention can aggregate the updated memory tensors to obtain the image features of the three-dimensional image. Through memory updating and three-dimensional convolution operations, meaningful features in the three-dimensional image can be screened out from different perspectives and dimensions, thereby enhancing the image feature expression capability of the three-dimensional image.
[0098] In the implementation of the present invention, the updated memory tensors are aggregated to obtain image features of the three-dimensional image, including: performing global average pooling on the updated memory tensors; aggregating the memory tensors after global average pooling to obtain global features of the three-dimensional image; and projecting and normalizing the global features to obtain image features of the three-dimensional image.
[0099] It can be understood that the embodiment of the present invention can perform global average pooling on the updated memory tensor, aggregate the memory tensors after global average pooling to obtain the global features of the three-dimensional image, project and normalize the global features, and obtain the image features of the three-dimensional image to ensure the consistency and stability of the image features between different perspectives, and eliminate unnecessary differences between different perspectives.
[0100] It should be noted that the three-dimensional image data processing method of the embodiment of the present invention can also directly process multi-stage three-dimensional images of the same perspective in the absence of multi-perspective data. The processing flow is the same, and the local features of each stage are used to update the memory tensor until all stages are updated to obtain the corresponding image feature representation.
[0101] The following describes a three-dimensional image data processing method according to an embodiment of the present invention through a specific embodiment, which specifically includes a dynamic memory update mechanism (also known as storage update) and a 3D convolution operation, taking the application of this method in an image encoder as an example.
[0102] The image encoder's storage update method flexibly extracts local features through the encoding and memory update mechanism of multi-view image data. The core of this method is to utilize image information from different viewpoints and gradually update the memory through a gating mechanism to ultimately obtain an image feature representation that includes different viewpoints. First, the input data is flattened and processed through a visual encoder for each viewpoint to extract local feature tokens. Then, a view mask is used to filter invalid viewpoints, reorganize the local features into a 3D grid, and update the memory tensor. Finally, an aggregated image feature representation is generated for each batch of samples through global 3D average pooling and projection normalization. This method effectively captures key information in multi-view data and provides powerful feature representation capabilities for image analysis tasks.
[0103] Definition: Batch size: B; Number of views (maximum number of views): V; Number of input channels: C; Spatial dimensions of each 3D view: (D, H, W); Spatial dimensions of the 3D patch grid: (D′, H′, W′) (indicates the number of patches in each dimension); Total number of tokens returned for each view: T; Hidden layer dimension (token embedding dimension): H s .
[0104] enter: -Multi-view 3D data; M∈{0,1} B×V - View mask (1 for valid views, 0 for filled / invalid views).
[0105] Output: Output of the visual encoder: Each view is encoded as a sequence of tokens of length T, and each token has an embedding dimension of H s .
[0106] 1. View flattening and processing through the visual encoder.
[0107] 1.1 Flatten the input data.
[0108] To process each view independently, we first flatten the input data X so that all views can be processed along the batch dimension. The resized tensor X flat The shape is:
[0109]
[0110] Among them, X flat [i] corresponds to the data of the i-th view in the batch. By flattening, the image data of all views are combined into a form that is easier to process.
[0111] 1.2 Processing by visual encoder
[0112] Next, the flattened data is fed into the visual encoder. The function of the visual encoder is to encode the data of each view into a series of tokens, where each token corresponds to a local or global feature of the image. After encoding, the output tensor Z is of the shape:
[0113]
[0114] Here, T is the number of tokens per view, H s is the embedding dimension of each token. Through the visual encoder, the spatial information of each view is converted into a set of embedded tokens of fixed dimension.
[0115] 1.3 Reshape into perspective format
[0116] After encoding, the tensor Z is reshaped into B×V to restore the perspective information:
[0117]
[0118] At this time, the tensor Z view Organized in a B×V shape, each batch contains token sequences from multiple views.
[0119] 2. Extract the patch token.
[0120] The encoding result of each view contains a special [CLS] token for representing global information and several patch tokens for representing local features of the image.
[0121] Remove one [CLS] token, and the remaining T-1 tokens represent the local features of the view, which are called patch tokens. view Extract these patch tokens and get:
[0122]
[0123] These patch tokens capture the characteristics of each local region in the image so as to understand the image at multiple scales.
[0124] 3. Applying a View Mask
[0125] To avoid processing image information of invalid views, a view mask M is introduced, which indicates which views are valid (value 1) and which are invalid (value 0). For each invalid view, its corresponding patch token is set to zero.
[0126] The masking operation for the [patch] token is as follows:
[0127] patch′ b,v =patch b,v ×M[b,v], when M[b,v]=0, the [patch] token of the invalid view is set to zero, thereby preventing invalid data from interfering with subsequent feature extraction.
[0128] 4. Reorganize the patch tokens into a 3D grid.
[0129] To capture the local spatial features of the image, the patch tokens are reorganized into a 3D grid. First, verify that the number of patch tokens matches the expected D′×H′×W′:
[0130] T-1 = D′×H′×W′ Then, reshape the patch token into a shape of B×V×D′×H′×W′×H s 5D tensor and rearrange the dimensions so that the data can adapt to subsequent 3D operations:
[0131]
[0132] The shape of the rearranged patch token is H s ×D′×H′×W′, which can effectively capture the local spatial features of the image.
[0133] 5. Update your memory.
[0134] In the storage update method, the memory tensor M is gradually updated by the patch token of each view (mem) , the memory tensor stores information from different viewpoints. During processing, we first extract the patch token from the current viewpoint and combine it with the view mask M b,v , ensuring that only data from valid perspectives is used for memory updates.
[0135] The memory update operation is as follows:
[0136]
[0137] Among them, UpdateMemory is an operation that updates the memory through a gating mechanism, which gradually enhances the memory content based on the current patch token and the previous memory tensor.
[0138] 6. Global average pooling and final projection.
[0139] After all perspectives are processed, the final memory tensor M (mem) Perform global 3D average pooling to obtain the aggregate representation agg of each batch of samples b :
[0140]
[0141] After global average pooling, agg b Perform projection and normalization to obtain the final eigenvector
[0142]
[0143] The image encoder's memory update method encodes 3D image data from different viewpoints into a unified image feature representation through a series of sophisticated steps. This process leverages the spatial information of multi-view images and combines it with a dynamic memory update mechanism to effectively extract global and local features. Specifically, the memory update method steps are as follows:
[0144] First, the input multi-view 3D image data is flattened. This flattening process processes the image data from each view independently, better capturing the characteristics of each view. The flattened data is then passed to the visual encoder, which converts the data from each view into a series of tokens, each of which represents a local or global feature of the image in a high-dimensional embedding space.
[0145] Next, we extract global information tokens ([CLS] tokens) and corresponding local information tokens (patch tokens) from the token sequence output by the visual encoder. This allows for separate processing of global and local features, providing a foundation for feature aggregation in subsequent steps. Furthermore, we introduce a view mask to filter out invalid viewpoints, ensuring that only valid viewpoint information is included in the subsequent memory update process.
[0146] Then, the extracted patch tokens are reorganized into a 3D grid. The main purpose of this step is to enable the network to effectively capture the spatial relationship between local regions in the image by mapping local features into a spatial grid, thereby enhancing the modeling ability of spatial features.
[0147] After this, the memory tensor is updated using the patch tokens from each view. The memory update process gradually enhances the image's feature representation by combining the local features of the current view with the previously remembered information. This memory update operation allows the network to share information across multiple views, thereby obtaining more comprehensive image features.
[0148] Finally, after processing all viewpoints, the final memory tensor undergoes global 3D average pooling. This step aggregates features from all viewpoints to obtain a global feature representation for each batch of samples. These aggregated features are then projected and normalized to obtain the final image representation. This process ensures consistency and stability of image features across viewpoints while eliminating unnecessary differences between different viewpoints.
[0149] Through the above steps, the image encoder can flexibly update and aggregate data from different viewpoints, capturing key information and ultimately generating a precise image feature representation. This approach not only processes multi-view data but also dynamically adjusts the information in memory to adapt to different data and task requirements. Therefore, this method has broad application prospects in multi-view image processing, medical image analysis, and other tasks that require cross-view information integration.
[0150] The three-dimensional image data processing method proposed in an embodiment of the present invention can identify three-dimensional images of multiple perspectives in three-dimensional image data and extract local features of three-dimensional images of multiple perspectives. It can capture changes in details in the three-dimensional image and use the local features of multiple perspectives to update the memory tensor, aggregate the updated memory tensor, and obtain the image features of the three-dimensional image. By dynamically updating the memory tensor using local features of different perspectives, it can transmit information between different perspectives, effectively capture and integrate the spatial information of three-dimensional image data from different perspectives, thereby fully enhancing the image feature expression capability of the three-dimensional image and improving the comprehensiveness and accuracy of the image feature expression.
[0151] The 3D image data processing method of the above embodiment focuses on processing 3D image data, while the data processing method of the following embodiment focuses on comprehensive processing of 3D image data and text data. The embodiments may refer to each other for any incomplete details.
[0152] It should be noted that the data processing method described in the following embodiments takes the processing of three-dimensional medical imaging data and medical text data in the medical field as an example.
[0153] like Figure 2 As shown, the data processing method includes the following steps:
[0154] In step S201 , the data to be processed and the task requirements input by the user are obtained, wherein the data to be processed includes at least one of three-dimensional image data and text data.
[0155] Among them, the task requirements include at least one of a retrieval task, a generation task and a matching task. The retrieval task includes retrieving the corresponding medical text based on the three-dimensional medical image, or retrieving the corresponding three-dimensional medical image based on the medical text. The generation task includes generating a medical text describing the three-dimensional medical image data based on the three-dimensional medical image data. The matching task includes inputting multiple three-dimensional medical images and multiple medical texts and matching the correspondence between the two.
[0156] In step S202, the data to be processed and the task requirements are input into the data processing model, and the medical data processing model outputs the analysis results corresponding to the task requirements, wherein the data processing model includes an image encoding module, a text encoding module and a feature alignment module. The image encoding module processes the three-dimensional image data based on the above-mentioned three-dimensional image data processing method to obtain the image features of the three-dimensional image. The text encoding module processes the text data based on the text encoder to obtain text features. The feature alignment module is used to align the image features and the text features, and determine the analysis results corresponding to the task requirements based on the results of the feature alignment.
[0157] It can be understood that the embodiment of the present invention can input the data to be processed and the task requirements into the data processing model, and the data processing model outputs the analysis results corresponding to the task requirements. The data processing model includes three modules, namely, an image encoding module, a text encoding module, and a feature alignment module, wherein:
[0158] The image encoding module processes the 3D image data based on the above-mentioned 3D image data processing method to obtain image features of the 3D image. The specific application entity of the image encoding module may be the image encoder in the above-mentioned embodiment, that is, the image encoder is used to process the 3D image.
[0159] The text encoding module processes text data based on a text encoder to obtain text features. The text encoder can be BERT (Bidirectional Encoder Representations from Transformers).
[0160] The feature alignment module is used to align image features and text features, and determine the analysis results corresponding to the task requirements based on the results of the feature alignment.
[0161] Through the above-mentioned data processing method, effective cross-modal data fusion can be achieved according to the user's task requirements and input data through the collaborative work of the image encoding module, text encoding module and feature alignment module in the medical data processing model. Moreover, since the image encoding module fuses three-dimensional image features from different perspectives, the integration and update mechanism of image features is optimized, further improving the alignment efficiency and accuracy between cross-modal data, that is, the accuracy and consistency of the fusion between multimodal data, thereby improving the performance of the data processing model under various tasks.
[0162] For example, if the user's task requirement is a retrieval task and the input data is three-dimensional medical imaging data, the input three-dimensional medical imaging data can be processed by the image encoding module of the medical data processing model to obtain corresponding image features. The text encoding module is used to process the medical text data stored in the preset medical database (which stores three-dimensional medical imaging data and medical text data) to obtain multiple text features. The feature alignment module is used to align the image features output by the image encoding module and the multiple text features output by the text encoding module to obtain a text feature with the highest feature alignment and use it as the analysis result of this task requirement, that is, the most matching medical text data is retrieved based on the input three-dimensional medical imaging data.
[0163] In an embodiment of the present invention, before inputting the data to be processed and the task requirements into the data processing model, it also includes: obtaining a training data set for the data processing model, wherein the training data set includes medical text data, three-dimensional image data, and a matching relationship between the text data and the three-dimensional image data; and training the data processing model using the training data set and the target loss function.
[0164] It is understandable that the present invention can utilize a training data set and a target loss function to train a data processing model to improve the model's processing capabilities under various tasks.
[0165] In an embodiment of the present invention, a data processing model is trained using a training data set and a target loss function, including: dividing the training data set into multiple batches of training data; iteratively training the data processing model using the multiple batches of training data; during the iterative training process, calculating the loss value of the data processing model using the target loss function, and updating the model parameters of the data processing model according to the loss value.
[0166] Among them, the objective loss function is:
[0167] L total =λ1L MLM +λ2L retrieval +λ3L LM ;
[0168] Among them, λ1, λ2, and λ3 are hyperparameters used to control the relative importance of each loss function. They can also be understood as the weights corresponding to each loss function. MLM is the text encoding loss function, L retrieval is the retrieval loss function, L LM Generate loss function for text, L total is the target loss function.
[0169] It can be understood that the implementation of the present invention can divide the training data set into multiple batches of training data, and use multiple batches of training data to iteratively train the data processing model. During the iterative training process, the target loss function is used to calculate the loss value of the data processing model, and the model parameters of the data processing model are updated according to the loss value, thereby improving the processing capability and effect of subsequent models.
[0170] In the implementation of the present invention, the model parameters of the data processing model are updated according to the loss value, including: calculating the gradient of the model parameters in the data processing model according to the loss value; and updating the model parameters according to the gradient of the model parameters and the learning rate of the model parameters.
[0171] It is understandable that the embodiment of the present invention can calculate the gradient of the model parameters in the data processing model based on the loss value, and update the model parameters based on the gradient of the model parameters and the learning rate of the model parameters.
[0172] It should be noted that the size of the training batch plays a role in balancing memory usage, computing speed, and model performance during the training process. The specific model parameter gradients and model parameter learning rates are calculated as follows.
[0173] Training batch: The batch size directly affects the accuracy and memory consumption of gradient calculation. During training, the gradient calculation formula for each batch is:
[0174]
[0175] Where N is the batch size, L(θ,x i ,y i ) is the loss function for each sample, θ is the parameter of the model, x i and y i Increasing the batch size can reduce gradient noise and make the model more stable, but it will also increase video memory consumption.
[0176] Model parameter gradients: When video memory limitations are strict, a gradient accumulation strategy is used to simulate large-batch training. The gradients of multiple small batches are accumulated before performing a parameter update. Assuming the cumulative number of steps for each update is K, the parameter gradient for each update is:
[0177]
[0178] Gradient accumulation can be used to simulate the effects of large batch training with smaller batch sizes, but at the cost of increased computation time.
[0179] Learning rate: The choice of learning rate has a significant impact on the convergence speed and stability of training. Excessively high learning rates can lead to unstable training, while excessively low learning rates can cause the model to converge too slowly. Using an appropriate learning rate scheduling strategy can effectively improve training efficiency.
[0180] Learning rate decay: Using the cosine decay strategy to adjust the learning rate can gradually reduce the learning rate during training to avoid fluctuations caused by excessive learning rate in the later stages. The cosine decay formula is as follows:
[0181]
[0182] Where η0 is the initial learning rate, t is the current training step, and T is the total number of training steps. Through cosine decay, the learning rate is initially high and gradually decreases over time, effectively avoiding training instability caused by excessively high learning rates.
[0183] Learning rate warmup: To prevent instability caused by excessive gradients at the beginning of training, learning rate warmup is used to stabilize the initial training process by gradually increasing the learning rate. Setting warmup_ratio = 0.03 means that during the initial training phase, the learning rate is gradually increased from 0 to the initial learning rate. The warmup process can be represented as follows:
[0184]
[0185] Among them, T warmup is the number of warmup steps, η0 is the initial learning rate, and t is the current step.
[0186] In addition, it should be noted that the embodiment of the present invention also balances training time, memory usage and model performance through optimizer configuration during the training process.
[0187] The optimizer is responsible for adjusting the model's parameters during training to reduce the loss function. In joint training, the AdamW (Adaptive Moment Estimation with Weight Decay Optimizer) optimizer is widely used because it effectively handles sparse gradients and combines weight decay to prevent overfitting. The update rule for the AdamW optimizer is as follows:
[0188]
[0189] Among them, θ tis the current model parameter, θ t-1 is the model parameter at the previous moment t-1, η is the learning rate, m t is the first moment estimate of the gradient, v t is the second-order moment estimate of the gradient, ∈ is a minimum value used to avoid zero division errors, and λ is the weight decay coefficient. Weight decay reduces the risk of overfitting by applying L2 regularization to the model parameters. The specific form is:
[0190]
[0191] in, is the L2 norm of the parameter. Proper optimizer configuration and weight decay coefficient selection can accelerate model convergence and improve its generalization ability on test data. When training data is scarce or noisy, strong weight decay can effectively prevent overfitting.
[0192] Optimizer selection and learning rate: The AdamW optimizer combines adaptive learning rate and L2 regularization. By automatically adjusting the learning rate and introducing weight decay during training, it can accelerate convergence and improve the final performance of the model.
[0193] In an embodiment of the present invention, before using the target loss function to calculate the loss value of the data processing model, it also includes: obtaining the true value of the text data in the training data set and the output value of the data processing model, and constructing a text encoding loss function based on the difference between the output value and the true value; obtaining the true value of the matching relationship between the text data and the three-dimensional image data in the training data set and the output matching relationship of the data processing model, and constructing a retrieval loss function based on the difference between the output matching relationship and the true value of the matching relationship; obtaining the text true value of the text data corresponding to the three-dimensional image data in the training data set and the output text of the data processing model, and constructing a text generation loss function based on the difference between the output text and the text true value; constructing a target loss function based on the text encoding loss function, the retrieval loss function, the text generation loss function and their corresponding hyperparameters.
[0194] Since related technologies usually rely on a single loss function for cross-modal alignment in the joint training process of multimodal data such as images and text, this method often cannot fully balance the contributions of each modality, especially when processing multi-view and cross-stage information in medical images. The lack of fine-tuning of different modal features results in the model's learning effect not being optimal. Therefore, the embodiment of the present invention can jointly train medical text data and three-dimensional medical imaging data, using different loss functions to promote the alignment of image features and text features, thereby improving the model's processing capabilities for multimodal data and improving the accuracy and consistency of multimodal data fusion.
[0195] Text encoding loss, retrieval loss, and text generation loss are key loss functions in multimodal learning. The text encoding loss helps train the model to understand and encode textual information, the retrieval loss facilitates mutual retrieval between images and text, and the text generation loss optimizes the ability to generate text from images. By jointly optimizing these loss functions, effective joint training of medical text and multimodal medical imaging data is possible, improving model performance in visual understanding and language generation tasks. The following are the three main loss functions: text encoding loss, retrieval loss, and text generation loss, as well as their mathematical expressions.
[0196] 1. Text Encoding Loss
[0197] The text encoding loss aims to optimize the learning of text encoders (such as BERT) so that they can accurately understand text information and effectively align it with image features. In joint training, the Masked Language Modeling (MLM) loss function is used to optimize the text encoder, especially to predict masked words. The specific loss function is:
[0198]
[0199] Among them, w m Indicates the covered word, w \m is the unmasked word, v is the image feature, P θ is the probability model for predicting masked words, D is the training data set, It is to find the expectation of the samples (w,v) sampled from the training data set D.
[0200] 2. Retrieval Loss
[0201] Retrieval loss is used to train the model to achieve mutual retrieval between images and texts in multimodal data. In this case, the goal is to optimize the model so that it can retrieve the correct text based on the image content, and vice versa. To this end, contrastive loss can be used to optimize the image-text similarity calculation. For a pair of image and text (v i ,w j ), the retrieval loss is calculated as:
[0202]
[0203] Among them, sim(v i ,w j ) represents the image v i and text w jThe similarity between them (usually using cosine similarity), τ is the temperature parameter, D is the training data set, k is the negative sample, is the sample obtained from the training data set D (v i ,w j )Seek expectations.
[0204] 3. Text Generation Loss.
[0205] The text generation loss is mainly used to optimize the model to generate accurate text descriptions given an image. In this case, the goal is to minimize the difference between the generated text and the real text. The autoregressive language model loss can be used for training. The specific loss function is:
[0206]
[0207] Among them, w t represents the tth word of the generated text, w <t is the word generated previously, v is the image feature, P θ It is a probability model that generates the current text word based on the image and the previous text, D is the training data set, It is to find the expectation of the samples (v, w) sampled from the training data set D.
[0208] Joint training objective: During the joint training process, the above three loss functions can be weighted and combined into the final target loss function:
[0209] L total =λ1L MLM +λ2L retrieval +λ3L LM ;
[0210] Among them, λ1, λ2, λ3 are hyperparameters that control the relative importance of each loss function. By optimizing this total loss function, the model can simultaneously optimize the encoding of images and text, making it perform well on both retrieval and generation tasks.
[0211] According to the data processing method proposed in the embodiment of the present invention, the user's input data and task requirements can be obtained. Through the collaborative work of the image encoding module, text encoding module and feature alignment module in the data processing model, effective fusion of cross-modal data is achieved. Moreover, since the image encoding module fuses three-dimensional image features from different perspectives, the integration and update mechanism of image features is optimized, and the alignment efficiency and accuracy between cross-modal data are further improved, thereby improving the performance and effect of the data processing model under various tasks.
[0212] In summary, the data processing method of the embodiment of the present invention mainly includes two aspects: one is the processing of multi-view 3D image data, and the other is to optimize the alignment of image and text features through joint training to improve the performance and effect of the data processing model in various tasks. Specifically, it includes:
[0213] 1. Extract image features through 3D convolution operations, specifically by meticulously modeling spatial information in the three dimensions of depth, width, and height. Compared to traditional methods that primarily rely on two-dimensional convolution or single-viewpoint information, the method of this invention enhances the model's ability to capture spatial information by fusing three-dimensional information from multiple perspectives, enabling the effective integration of fine-grained features obtained from different spatial locations. This approach can more accurately represent spatial relationships in images, especially for three-dimensional medical imaging data, where it can deeply explore three-dimensional structural information.
[0214] 2. When processing 3D image data, a dynamic memory update mechanism is used to integrate the feature information of each viewpoint, fully capturing detailed features from different viewpoints, different spatial dimensions, and different stages. This further enhances the model's expressiveness in multiple viewpoints and dimensions, ensuring the effective fusion of cross-viewpoint information. Specifically:
[0215] A 3D gating mechanism has been introduced. By dynamically adjusting the 3D update gate and reset gate, it accurately filters important information and suppresses noise and invalid features during the integration of multi-view and multi-stage medical image features. This mechanism is particularly suitable for 3D spatial data, dynamically adjusting the depth, width, and height dimensions to enhance the fusion of image features.
[0216] 3. Apply visual masks to ensure that only valid perspective information is involved in memory updates during 3D image processing and model training. Compared to the simple superposition of perspective or spatial information processing in traditional methods, the present invention uses dynamic masks to eliminate invalid perspectives, avoiding the interference of irrelevant data on the memory update process and achieving efficient fusion of multi-view image features. This mask not only operates in the spatial dimension, but also takes into account the impact of temporal dynamic information (i.e., multi-stage information) on memory updates, and can subsequently improve the training efficiency and effect of the model.
[0217] 4. We propose an efficient multimodal feature aggregation framework to effectively integrate 3D image features and text information. By optimizing the alignment of image and text features in a shared space through a joint loss function, we specifically enhance the fusion of information from different perspectives and stages in 3D space.
[0218] 5. By optimizing the joint loss function, combining text encoding loss, retrieval loss, and text generation loss, and especially balancing the optimization goals of different tasks in three-dimensional spatial data, the model's performance in multi-view medical imaging and text tasks is enhanced.
[0219] 6. Through the three-dimensional dynamic memory mechanism and feature fusion strategy, the present invention solves the problem of information fusion between different perspectives and stages, ensuring that image and text data can be accurately aligned in three-dimensional space, improving the efficiency and accuracy of multimodal training, and showing stronger performance especially when processing high-dimensional medical imaging data.
[0220] The following describes a method for jointly training multimodal medical images and medical texts according to an embodiment of the present invention through a specific embodiment, that is, a method for training a data processing model in the medical field. Figure 3 As shown, for example, multiple three-dimensional medical images are input, including the intracranial axial plain scan stage, axial arterial stage, and axial perfusion stage, and medical text is input. The three-dimensional features of the three-dimensional medical image features of different stages and the text features of the medical text are aligned and fused, and the medical data processing model is trained and optimized through the joint loss function.
[0221] The data processing method of this embodiment for the medical field enables researchers to more comprehensively and accurately integrate complex medical imaging data and clinical text information, thereby revealing deep-seated pathophysiological mechanisms and promoting innovative developments in medical research and clinical practice. Researchers can also more deeply analyze key features in complex imaging patterns, for example, combining multi-stage imaging to analyze the dynamic changes of contrast agents, locate the progression stage of lesions, and evaluate the effectiveness of treatments. This capability provides new technical support for the diagnosis and treatment of complex diseases such as cancer, cardiovascular disease, or neurodegenerative diseases.
[0222] Furthermore, through multimodal fusion, this invention can significantly improve the efficiency and accuracy of diagnostic report generation, reduce the time cost of manual image interpretation and report writing, and enable doctors to focus more on personalized treatment for their patients. Furthermore, through an efficient multi-view image encoder and text alignment model, researchers can explore the temporal dynamics of disease progression and construct disease-related imaging biomarkers and text description networks. For example, by analyzing the association between multi-stage images and a patient's medical history, early diagnostic clues for specific diseases can be revealed, providing new research directions for precision medicine.
[0223] In healthcare systems, this invention holds particular clinical value. Doctors can leverage the image-text alignment information generated by medical data processing models to quickly identify key features of a patient's disease progression and develop comprehensive treatment plans based on multimodal information. This improves treatment accuracy and reduces misdiagnoses and delays. In settings with limited medical resources, it can also help optimize diagnostic processes, making healthcare more efficient and accessible.
[0224] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0225] An embodiment of the present invention further provides a three-dimensional image data processing device.
[0226] like Figure 4 As shown, the three-dimensional image data processing device 10 includes: a first acquisition module 101 , an extraction module 102 and an aggregation module 103 .
[0227] Among them, the first acquisition module 101 is used to obtain three-dimensional images of multiple perspectives in the three-dimensional image data; the extraction module 102 is used to extract local features of the three-dimensional images of multiple perspectives, and use the local features of multiple perspectives to update the memory tensor, wherein the memory tensor is used to store and fuse the local features of multiple perspectives; the aggregation module 103 is used to aggregate the updated memory tensor to obtain the image features of the three-dimensional image.
[0228] In an embodiment of the present invention, the three-dimensional images of multiple perspectives include three-dimensional images of multiple stages, and the extraction module 102 is further used to: obtain local features of the three-dimensional image at the current stage under the current perspective; use the local features at the current stage under the current perspective to update the memory tensor stored in the previous stage under the current perspective, until the memory tensors at the multiple stages under the current perspective are updated; after using the local features of the multiple stages of the current perspective to update the memory tensor, use the local features of the multiple stages of the next perspective to update the memory tensor stored in the current perspective, until the memory tensors of the multiple perspectives are updated, and the final updated memory tensor is obtained.
[0229] In an embodiment of the present invention, the extraction module 102 is further used to: use the gating mechanism and the local features of the current stage to update the memory tensor stored in the previous stage, and use the updated memory tensor of the previous stage as the memory tensor stored in the current stage, wherein the gating mechanism includes an update gate and a reset gate, the update gate is used to control the fusion ratio between the local features of the current stage and the memory tensor stored in the previous stage, and the reset gate is used to control the reset level of the memory tensor stored in the previous stage in the current stage.
[0230] In an embodiment of the present invention, the three-dimensional image data processing device 10 of the embodiment of the present invention further includes: a calculation module, which is used to calculate the value of the update gate and the value of the reset gate according to the local features of the current stage and the memory tensor stored in the previous stage before using the gating mechanism and the local features of the current stage to update the memory tensor stored in the previous stage; based on the value of the update gate, the value of the reset gate and the local features of the current stage, the memory tensor stored in the previous stage is updated, wherein the value of the update gate is the fusion ratio, and the value of the reset gate is the reset level.
[0231] In an embodiment of the present invention, the computing module is further used to: perform a three-dimensional convolution operation on the memory tensor stored in the previous stage to obtain a second feature; combine the first feature and the second feature to obtain a first combined feature, input the first combined feature into a first target activation function, and the first target activation function outputs the value of the update gate and the value of the reset gate.
[0232] In an embodiment of the present invention, the extraction module 102 is further used to: reset the memory tensor stored in the previous stage using a reset gate; perform a three-dimensional convolution operation on the reset memory tensor stored in the previous stage to obtain a third feature; combine the first feature and the third feature to obtain a second combined feature, input the second combined feature into the second target activation function, and the second target activation function outputs a candidate memory tensor; and use an update gate to fuse the candidate memory tensor and the memory tensor stored in the previous stage.
[0233] In the embodiment of the present invention, the 3D image data processing device 10 of the embodiment of the present invention further includes a screening module.
[0234] The screening module is used to extract local features of three-dimensional images from multiple perspectives, and then use the perspective mask to screen out invalid perspectives; and reset the local features corresponding to the invalid perspectives to zero.
[0235] In an embodiment of the present invention, the aggregation module 103 is further used to: perform global average pooling on the updated memory tensor; aggregate the memory tensors after global average pooling to obtain global features of the three-dimensional image; project and normalize the global features to obtain image features of the three-dimensional image.
[0236] It should be noted that, for the description of the features in the embodiment corresponding to the three-dimensional image data processing device, reference can be made to the relevant description of the embodiment corresponding to the three-dimensional image data processing method, and no further details will be given here.
[0237] An embodiment of the present invention further provides a data processing device.
[0238] like Figure 5 As shown, the data processing device 20 includes: a second acquisition module 201 and an input module 202.
[0239] Among them, the second acquisition module 201 is used to obtain the data to be processed and task requirements input by the user, wherein the data to be processed includes at least one of three-dimensional image data and text data; the input module 202 is used to input the data to be processed and the task requirements into the data processing model, and the data processing model outputs the analysis results corresponding to the task requirements, wherein the data processing model includes an image encoding module, a text encoding module and a feature alignment module, the image encoding module processes the three-dimensional image data based on the above-mentioned three-dimensional image data processing device 10 to obtain the image features of the three-dimensional image, the text encoding module processes the text data based on the text encoder to obtain text features, and the feature alignment module is used to perform feature alignment on the image features and the text features, and determine the analysis results corresponding to the task requirements based on the results of the feature alignment.
[0240] In the embodiment of the present invention, the data processing device 20 of the embodiment of the present invention further includes: a training module.
[0241] Among them, the training module is used to obtain a training data set for the data processing model before inputting the data to be processed and the task requirements into the data processing model, wherein the training data set includes text data, three-dimensional image data, and the matching relationship between the text data and the three-dimensional image data; and the data processing model is trained using the training data set and the target loss function.
[0242] In an embodiment of the present invention, the training module is further used to: divide the training data set into multiple batches of training data; iteratively train the data processing model using multiple batches of training data; during the iterative training process, use the target loss function to calculate the loss value of the data processing model, and update the model parameters of the data processing model according to the loss value.
[0243] In an embodiment of the present invention, the training module is further used to: calculate the gradient of the model parameters in the data processing model according to the loss value; and update the model parameters according to the gradient of the model parameters and the learning rate of the model parameters.
[0244] In this embodiment of the present invention, the objective loss function is:
[0245] L total =λ1L MLM +λ2L retrieval +λ3L LM ;
[0246] Among them, λ1, λ2, and λ3 are hyperparameters used to control the relative importance of each loss function. MLM is the text encoding loss function, L retrieval is the retrieval loss function, L LM Generate loss function for text, L total is the target loss function.
[0247] It should be noted that, for the description of the features in the embodiments corresponding to the data processing device, reference can be made to the relevant description of the embodiments corresponding to the data processing method, and no further details will be given here.
[0248] An embodiment of the present invention further provides an electronic device, such as Figure 6 As shown, it includes a memory 601 and a processor 602, the memory 601 stores a computer program, and the processor 602 is configured to run the computer program to execute the steps in any of the above-mentioned three-dimensional image data processing method embodiments, or the steps in any of the above-mentioned data processing method embodiments.
[0249] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps of any one of the above-mentioned three-dimensional image data processing method embodiments, or the steps of any one of the above-mentioned data processing method embodiments when run.
[0250] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0251] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned three-dimensional image data processing method embodiments, or the steps in any of the above-mentioned data processing method embodiments.
[0252] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0253] The above is a detailed introduction to a medical data processing method provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only intended to help understand the method and core concept of the present invention. It should be noted that, for those skilled in the art, without departing from the principles of the present invention, several improvements and modifications may be made to the present invention, and such improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A three-dimensional image data processing method, characterized in that: The following steps are involved: Acquire three-dimensional images of multiple perspectives from the three-dimensional image data, wherein the three-dimensional images of the multiple perspectives include three-dimensional images of multiple stages; Extract local features of the three-dimensional images of the multiple perspectives, and use the local features of the multiple perspectives to update a memory tensor, wherein the memory tensor is used to store and fuse the local features of the multiple perspectives; the use of the local features of the multiple perspectives to update the memory tensor includes: obtaining the local features of the three-dimensional image of the current stage of the current perspective; using the local features of the current stage of the current perspective to update the memory tensor stored in the previous stage of the current perspective, until the memory tensors of the multiple stages of the current perspective are updated; after updating the memory tensor using the local features of the multiple stages of the current perspective, using the local features of the multiple stages of the next perspective to update the memory tensor stored in the current perspective, until the memory tensors of the multiple perspectives are updated, The memory tensor after the final update; the using the local features of the current stage of the current view to update the memory tensor stored in the previous stage of the current view, including: using the gating mechanism and the local features of the current stage of the current view to update the memory tensor stored in the previous stage of the current view, and using the updated memory tensor of the previous stage of the current view as the memory tensor stored in the current stage of the current view, wherein the gating mechanism includes an update gate and a reset gate, the update gate is used to control the fusion ratio between the local features of the current stage of the current view and the memory tensor stored in the previous stage of the current view, and the reset gate is used to control the reset level of the memory tensor stored in the previous stage of the current view at the current stage of the current view; The updated memory tensors are aggregated to obtain image features of the three-dimensional image.
2. The three-dimensional image data processing method according to claim 1, characterized in that: Before updating the memory tensor stored in the previous stage of the current view using the gating mechanism and the local features of the current stage of the current view, the method includes: Calculating a value of the update gate and a value of the reset gate according to the local features of the current stage of the current perspective and the memory tensor stored in the previous stage of the current perspective; According to the value of the update gate, the value of the reset gate and the local features of the current stage of the current perspective, the memory tensor stored in the previous stage of the current perspective is updated, wherein the value of the update gate is the fusion ratio, and the value of the reset gate is the reset level.
3. The three-dimensional image data processing method according to claim 2, characterized in that: The calculating the value of the update gate and the value of the reset gate according to the local features of the current stage of the current perspective and the memory tensor stored in the previous stage of the current perspective includes: Performing a three-dimensional convolution operation on the local features of the current stage of the current perspective to obtain a first feature; Performing a three-dimensional convolution operation on the memory tensor stored in the previous stage of the current perspective to obtain a second feature; The first feature and the second feature are combined to obtain a first combined feature, and the first combined feature is input into a first target activation function, which outputs a value of the update gate and a value of the reset gate.
4. The three-dimensional image data processing method according to claim 3, wherein: The updating of the memory tensor stored in the previous stage of the current perspective by using the gating mechanism and the local features of the current stage of the current perspective includes: Resetting the memory tensor stored in the previous stage of the current perspective using the reset gate; Perform a three-dimensional convolution operation on the reset memory tensor stored in the previous stage to obtain the third feature; Combining the first feature and the third feature to obtain a second combined feature, inputting the second combined feature into a second target activation function, and the second target activation function outputting a candidate memory tensor; The candidate memory tensor and the memory tensor stored in the previous stage of the current perspective are fused using the update gate.
5. The three-dimensional image data processing method according to claim 1, wherein: After extracting the local features of the 3D images from multiple perspectives, the following steps are also included: Use the view mask to filter out invalid view angles; The local features corresponding to the invalid viewing angle are reset to zero.
6. The three-dimensional image data processing method according to claim 1, characterized in that: The aggregating the updated memory tensor to obtain the image features of the three-dimensional image includes: Performing global average pooling on the updated memory tensor; aggregating the memory tensors after global average pooling to obtain the global features of the three-dimensional image; The global features are projected and normalized to obtain image features of the three-dimensional image.
7. A data processing method, characterized in that: The following steps are involved: Acquiring data to be processed and task requirements input by a user, wherein the data to be processed includes at least one of three-dimensional image data and text data; The data to be processed and the task requirements are input into a data processing model, and the data processing model outputs an analysis result corresponding to the task requirements, wherein the data processing model includes an image encoding module, a text encoding module and a feature alignment module, the image encoding module processes the three-dimensional image data based on the three-dimensional image data processing method according to any one of claims 1 to 6 to obtain image features of the three-dimensional image, the text encoding module processes the text data based on a text encoder to obtain text features, and the feature alignment module is used to perform feature alignment on the image features and the text features, and determine the analysis result corresponding to the task requirements based on the result of the feature alignment.
8. The data processing method according to claim 7, characterized in that: Before inputting the data to be processed and the task requirements into the three-dimensional data processing model, the method further includes: Acquire a training data set for the three-dimensional data processing model, wherein the training data set includes text data, three-dimensional image data, and a matching relationship between the text data and the three-dimensional image data; The data processing model is trained using the training data set and the target loss function.
9. The data processing method according to claim 8, characterized in that: The step of training the data processing model using the training data set and the target loss function includes: Dividing the training data set into multiple batches of training data; Iteratively training the data processing model using the multiple batches of training data; During the iterative training process, the target loss function is used to calculate the loss value of the data processing model, and the model parameters of the data processing model are updated according to the loss value.
10. The data processing method according to claim 9, characterized in that: The updating of the model parameters of the data processing model according to the loss value includes: Calculating the gradient of the model parameters in the data processing model according to the loss value; The model parameters are updated according to the gradient of the model parameters and the learning rate of the model parameters.
11. The data processing method according to any one of claims 8 to 10, characterized in that: The objective loss function is: L total =λ1L MLM +λ2L retrieval +λ3L LM ; Among them, λ1, λ2, λ3 are hyper parameters, L MLM is the text encoding loss function, L retrieval is the retrieval loss function, L LM Generate loss function for text, L total is the target loss function.
12. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the three-dimensional image data processing method according to any one of claims 1 to 6, or the steps of the data processing method according to any one of claims 7 to 11, when executing the computer program.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the three-dimensional image data processing method according to any one of claims 1 to 6, or the steps of the data processing method according to any one of claims 7 to 11 are implemented.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the three-dimensional image data processing method according to any one of claims 1 to 6 or the steps of the data processing method according to any one of claims 7 to 11 are implemented.
Citation Information
Patent Citations
Radiology clinical medical image report generation method and system
CN119028509A