Three-dimensional medical image generation method based on two-stage diffusion

CN122798992APending Publication Date: 2026-09-22QINGDAO RES INST OF BEIHANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610825769.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

显存消耗过高且生成结果存在块间解剖结构不连续、不一致的伪影

Benefits of technology

[0011]本公开的上述各个实施例具有如下有益效果:通过本公开的一些实施例的基于两阶段扩散的三维医学影像生成方法,通过层级生成架构与全局条件控制手段,实现在有限计算资源下高效生成高分辨率、结构一致的三维医学影像的效果。具体来说,造成显存消耗过高且生成结果存在块间解剖结构不连续、不一致的伪影的原因在于:显存消耗过高且生成结果存在块间解剖结构不连续、不一致的伪影。现有直接生成高分辨率三维潜在表示的方法,导致计算机的图形处理器面临巨大的显存压力,无法在资源有限的硬件环境下执行。现有简单的分块生成策略中,由于计算机在生成每个子块时缺乏有效的全局信息传递机制,导致最终拼接生成的完整图像存在块间解剖结构不连续、不一致的伪影。基于此,本公开的一些实施例的基于两阶段扩散的三维医学影像生成方法,首先,响应于检测到目标显示设备的图形处理器内存容量低于预设单次峰值内存阈值,生成第一噪声张量。由此,可以使计算机在资源受限的硬件环境下,智能地触发低显存消耗的高分辨率图像生成流程。其次,基于预设扩散时间步、图像类别标签和第一阶段扩散变换器,对上述第一噪声张量进行第一阶段条件去噪处理,以生成缩略图潜在表示。由此,可以生成一个包含完整全局解剖结构信息的低分辨率“草图”,为后续的高分辨率分块生成提供统一的结构规划与约束条件。然后,确定目标潜在表示空间,以及将上述目标潜在表示空间划分为预设数量的子块空间,得到子块空间集合。由此,可以将高维生成任务分解为多个可独立处理的低维子任务。再然后,基于上述缩略图潜在表示和第二阶段扩散变换器,对上述子块空间集合中的各个子块空间执行第二阶段条件去噪处理,得到子块潜在表示集合。由此,可以通过跨注意力机制实时参考上述全局缩略图的指导,并通过三维坐标位置编码明确该子块的绝对空间位置,从而在生成过程中确保每个子块的内容与全局结构相符且在空间上精准对齐,从根本上避免了块间结构错乱与不连续伪影的产生。接着,将上述子块潜在表示集合中的各个子块潜在表示进行拼接,得到目标潜在表示。由此,可以将所有在全局条件约束下生成的、位置明确的子块,无损地组装成一个完整、连贯的高分辨率潜在表示。最后,对上述目标潜在表示进行解码,以生成三维医学影像,以及将上述三维医学影像发送至上述目标显示设备进行显示。由此,可以输出最终的高质量三维医学图像,完成从条件到像素的完整生成过程,并将结果可视化以供使用。该实施方式实现了降低峰值显存占用、同时确保生成影像解剖结构一致且无块间伪影的技术效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122798992A_ABST
    Figure CN122798992A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a three-dimensional medical image generation method based on two-stage diffusion. A specific embodiment of the method comprises: in response to detecting that the memory capacity of a graphics processor is lower than a preset single-peak memory threshold, generating a first noise tensor; generating a thumbnail latent representation based on a first-stage diffusion transformer; dividing a target latent representation space into a preset number of sub-block spaces to obtain a set of sub-block spaces; performing second-stage conditional denoising processing on the set of sub-block spaces based on the thumbnail latent representation and a second-stage diffusion transformer to obtain a set of sub-block latent representations; splicing the set of sub-block latent representations to obtain a target latent representation; decoding the target latent representation to generate a three-dimensional medical image; and sending the three-dimensional medical image to a target display device for display. This embodiment achieves the technical effect of reducing peak memory usage while ensuring that the generated image has consistent anatomical structure and no inter-block artifacts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of computer technology, and more specifically to a method for generating three-dimensional medical images based on two-stage diffusion. Background Technology

[0002] High-resolution 3D medical images (such as computed tomography (CT) scans and magnetic resonance imaging (MRI) images are of great value in clinical diagnosis, surgical planning, and medical research. However, acquiring large-scale, high-quality 3D medical image data faces many challenges, including high data acquisition costs, a huge workload for expert annotation, and strict requirements for patient privacy protection. Currently, 3D medical image generation typically employs the following methods: Direct generation of high-resolution 3D latent representations: Existing methods usually directly perform diffusion generation on high-resolution 3D latent representations (e.g., 64×64×64). Simple spatial block generation strategy: To alleviate memory pressure, another intuitive approach is to divide the high-resolution volumetric data space into multiple sub-blocks, perform diffusion generation on each sub-block independently, and finally stitch them together to obtain the complete result.

[0003] However, when using the above method to generate 3D medical images, the following technical problems often arise: Excessive memory consumption and artifacts resulting from discontinuous and inconsistent inter-block anatomical structures in the generated image are problematic. Existing methods for directly generating high-resolution 3D latent representations place enormous demands on the computer's graphics processor, making them unsuitable for resource-constrained hardware environments. Furthermore, existing simple block-based generation strategies lack an effective global information transfer mechanism when generating each sub-block, leading to artifacts of discontinuous and inconsistent inter-block anatomical structures in the final stitched image.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure propose a method, apparatus, electronic device, and computer-readable medium for generating three-dimensional medical images based on two-stage diffusion to address one or more of the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide a method for generating three-dimensional medical images based on two-stage diffusion. The method includes: generating a first noise tensor in response to detecting that the memory capacity of the graphics processor of a target display device is lower than a preset single-peak memory threshold; performing a first-stage conditional denoising process on the first noise tensor based on a preset diffusion time step, image category labels, and a first-stage diffusion transformer to generate a thumbnail latent representation; determining a target latent representation space and dividing the target latent representation space into a preset number of sub-block spaces to obtain a set of sub-block spaces; performing a second-stage conditional denoising process on each sub-block space in the set of sub-block spaces based on the thumbnail latent representation and the second-stage diffusion transformer to obtain a set of sub-block latent representations; stitching together the sub-block latent representations in the set of sub-block latent representations to obtain a target latent representation; decoding the target latent representation to generate a three-dimensional medical image; and sending the three-dimensional medical image to the target display device for display.

[0008] Secondly, some embodiments of this disclosure provide a three-dimensional medical image generation apparatus based on two-stage diffusion. The apparatus includes: a detection unit configured to generate a first noise tensor in response to detecting that the memory capacity of the graphics processor of a target display device is lower than a preset single-peak memory threshold; a first denoising processing unit configured to perform a first-stage conditional denoising processing on the first noise tensor based on a preset diffusion time step, an image category label, and a first-stage diffusion transformer to generate a thumbnail latent representation; a determination unit configured to determine a target latent representation space and divide the target latent representation space into a preset number of sub-block spaces to obtain a set of sub-block spaces; a second denoising processing unit configured to perform a second-stage conditional denoising processing on each sub-block space in the set of sub-block spaces based on the thumbnail latent representation and the second-stage diffusion transformer to obtain a set of sub-block latent representations; a stitching unit configured to stitch together the sub-block latent representations in the set of sub-block latent representations to obtain a target latent representation; and a decoding unit configured to decode the target latent representation to generate a three-dimensional medical image and send the three-dimensional medical image to the target display device for display.

[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0011] The above embodiments of this disclosure have the following beneficial effects: The two-stage diffusion-based 3D medical image generation method of some embodiments of this disclosure, through a hierarchical generation architecture and global condition control, achieves the effect of efficiently generating high-resolution, structurally consistent 3D medical images under limited computing resources. Specifically, the reason for excessive memory consumption and artifacts of discontinuous and inconsistent inter-block anatomical structures in the generated results is that: Existing methods for directly generating high-resolution 3D latent representations cause the computer's graphics processor to face enormous memory pressure, making it impossible to execute in a resource-constrained hardware environment. In existing simple block generation strategies, due to the lack of an effective global information transmission mechanism when generating each sub-block, the final stitched complete image has artifacts of discontinuous and inconsistent inter-block anatomical structures. Based on this, the two-stage diffusion-based 3D medical image generation method of some embodiments of this disclosure first generates a first noise tensor in response to detecting that the graphics processor memory capacity of the target display device is lower than a preset single-peak memory threshold. Thus, the computer can intelligently trigger a high-resolution image generation process with low memory consumption in a resource-constrained hardware environment. Secondly, based on the preset diffusion time step, image category labels, and the first-stage diffusion transformer, the first noise tensor undergoes first-stage conditional denoising to generate a thumbnail latent representation. This generates a low-resolution "sketch" containing complete global anatomical information, providing a unified structural plan and constraints for subsequent high-resolution block generation. Then, the target latent representation space is determined, and this target latent representation space is divided into a preset number of sub-block spaces, resulting in a set of sub-block spaces. This decomposes the high-dimensional generation task into multiple independently processable low-dimensional sub-tasks. Next, based on the thumbnail latent representation and the second-stage diffusion transformer, second-stage conditional denoising is performed on each sub-block space in the sub-block space set, resulting in a set of sub-block latent representations. This allows for real-time reference to the global thumbnail via a cross-attention mechanism, and the absolute spatial position of each sub-block is determined through 3D coordinate position encoding. This ensures that the content of each sub-block conforms to the global structure and is precisely aligned spatially during generation, fundamentally avoiding structural errors and discontinuities between blocks. Next, the latent representations of each sub-block in the aforementioned set of sub-block latent representations are concatenated to obtain the target latent representation. This allows all the location-defined sub-blocks generated under global constraints to be losslessly assembled into a complete, coherent, high-resolution latent representation. Finally, the target latent representation is decoded to generate a 3D medical image, which is then sent to the target display device for display. This process outputs a final high-quality 3D medical image, completing the entire generation process from conditions to pixels, and visualizing the results for use.This implementation achieves the technical effect of reducing peak video memory usage while ensuring consistent anatomical structure of the generated image and eliminating inter-block artifacts. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0013] Figure 1 This is a flowchart of some embodiments of the two-stage diffusion-based three-dimensional medical image generation method disclosed herein; Figure 2 This is a schematic diagram of the structure of some embodiments of the three-dimensional medical image generation apparatus based on two-stage diffusion according to the present disclosure. Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0018] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0019] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0020] Figure 1 A flow 100 of some embodiments of a two-stage diffusion-based three-dimensional medical image generation method according to this disclosure is shown. This two-stage diffusion-based three-dimensional medical image generation method includes the following steps: Step 101: In response to detecting that the graphics processor memory capacity of the target display device is lower than a preset single peak memory threshold, a first noise tensor is generated.

[0021] In some embodiments, the execution entity (e.g., a server) of the two-stage diffusion-based 3D medical image generation method can generate a first noise tensor in response to detecting that the graphics processor memory capacity of the target display device is lower than a preset single-peak memory threshold. The target display device can be a general computer monitor or a medical imaging-specific display. For example, the general computer monitor can be an LCD, LED, or OLED screen connected to a desktop or laptop computer. The medical imaging-specific display can be a high-resolution grayscale or color display conforming to medical imaging display standards such as DICOM. The preset single-peak memory threshold can be the maximum amount of memory estimated by the computer's graphics processor to be required when performing a complete high-resolution 3D image generation task. For example, the preset single-peak memory threshold can be 32GB. The first noise tensor can be a five-dimensional tensor with dimensions [B, C, D, H, W]. B can be a preset batch size, representing the number of data samples processed simultaneously. For example, B can be 16, 32, or 64. C can be a preset number of channels. For example, C can be 8. D, H, and W can be the spatial dimensions of the first noise tensor in terms of depth, height, and width. For example, D, H, W above could be 16, 16, 16.

[0022] In some optional implementations of certain embodiments, the execution entity may generate a first noise tensor in response to detecting that the graphics processor memory capacity of the target display device is lower than a preset single-peak memory threshold: The first step involves determining the space size of a first noise tensor based on a preset space size, in response to the detection that the memory capacity of the graphics processor of the target display device is lower than a preset single-peak memory threshold. In practice, the execution entity can call the Application Programming Interface (API) provided by the operating system or hardware driver to obtain the memory capacity of the graphics processor of the target display device. Then, in response to the detection that the memory capacity is lower than the preset single-peak memory threshold, the preset space size is read and assigned to the space size of the first noise tensor. The preset space size can be a predefined set of parameters stored in the computer, specifying the size of the first noise tensor in the three spatial dimensions of depth, height, and width. For example, the preset space size could be [16, 16, 16].

[0023] The second step involves generating a tensor conforming to a standard Gaussian distribution, based on the pseudo-random number generation function and the aforementioned first noise tensor space size, as the first noise tensor. In practice, the executing entity can call a pseudo-random number generation function from a computer programming library, passing the aforementioned first noise tensor space size as a shape parameter to the pseudo-random number generation function to generate a tensor conforming to a standard Gaussian distribution, the aforementioned preset batch size, and the aforementioned preset number of channels, as the first noise tensor. The aforementioned computer programming library can be the `random` library in NumPy or PyTorch. The aforementioned pseudo-random number generation function can be the `torch.randn` function.

[0024] Optionally, the first-stage diffusion converter and the second-stage diffusion converter can be obtained through the following steps: The first step is to acquire the original 3D medical images. In practice, the aforementioned entity can retrieve the original 3D medical images from a Picture Archiving and Communication System (PACS). These original 3D medical images can be high-resolution 3D volumetric data directly acquired by medical imaging equipment such as computed tomography (CT) or magnetic resonance imaging (MRI). For example, the original 3D medical images could be images with a resolution of 512×512×512.

[0025] The second step involves encoding the original 3D medical image using a preset encoder to obtain a latent representation of the original image. In practice, the executing entity can input the original 3D medical image into the preset encoder to generate the latent representation of the original image. The preset encoder can be a VQ-VAE encoder. The latent representation of the original image can be the latent representation obtained after compressing the original 3D medical image using the VQ-VAE encoder. As an example, the shape of the latent representation of the original image can be [B, C, 64, 64, 64].

[0026] The third step involves downsampling the original image latent representation to obtain a thumbnail of the original image. In practice, the execution entity can perform trilinear interpolation on the original image latent representation to obtain the thumbnail. The thumbnail can be a low-resolution latent representation obtained by downsampling the original image latent representation. As an example, the shape of the thumbnail can be [B, C, 16, 16, 16].

[0027] The fourth step involves dividing the original image latent representation into a predetermined number of original sub-block latent representations to obtain a set of original sub-block latent representations. The predetermined number can be 8. For example, the executing entity can divide the original image latent representation of shape [B, C, 64, 64, 64] into 8 original sub-block latent representations of shape [B, C, 32, 32, 32] to obtain a set of original sub-block latent representations.

[0028] Fifth, based on the denoising diffusion probability model, the original image thumbnail, and the original sub-block latent representation set, the original first-stage diffusion transformer and the original second-stage diffusion transformer are trained to obtain the first-stage diffusion transformer and the second-stage diffusion transformer. In practice, firstly, the execution entity can input the original image thumbnail into the denoising diffusion probabilistic models to achieve forward propagation, obtaining a noisy original image thumbnail and the added real noise. Then, the execution entity can input the noisy original image thumbnail, a preset diffusion time step, and image category labels into the original first-stage diffusion transformer to obtain the predicted noise. Then, the execution entity can determine the first loss (e.g., mean squared error) between the predicted noise and the real noise. Subsequently, through backpropagation, the gradient of the first loss with respect to the parameters of the original first-stage diffusion transformer is determined, and the original first-stage diffusion transformer is trained using the parameter gradient and an optimizer (such as AdamW) to obtain the first-stage diffusion transformer. Secondly, the aforementioned execution entity can input the latent representations of each original sub-block in the aforementioned set of latent representations of the original sub-blocks into the aforementioned denoising diffusion probability model to achieve denoising processing of each latent representation of the original sub-blocks, obtaining each denoised original sub-block latent representation and each sub-block's true noise. Next, the aforementioned denoised original sub-block latent representations are input into the aforementioned original second-stage diffusion transformer to obtain each sub-block's predicted noise. Then, the aforementioned execution entity can determine the second loss (e.g., mean squared error) between each sub-block's predicted noise and each sub-block's true noise. Subsequently, through a backpropagation algorithm, the gradient of the second loss with respect to the parameters of the aforementioned original first-stage diffusion transformer is determined, and the aforementioned parameter gradient and an optimizer (such as AdamW) are used to train the aforementioned original second-stage diffusion transformer to obtain the second-stage diffusion transformer. The aforementioned preset diffusion time step can be an index value of the specific number of steps in the forward or backward propagation process in the diffusion model. For example, the aforementioned preset diffusion time step can be an integer in the range of 0 to 1000. The aforementioned image category label can be discrete label information used to control the semantic category to which the generated image belongs. For example, the semantic categories mentioned above could be brain magnetic resonance imaging (MRI) or lung computed tomography (CT). The original first-stage diffusion transformer described above has the same network structure as the original first-stage diffusion transformer described above; the original first-stage diffusion transformer is a model obtained by training the original first-stage diffusion transformer. Similarly, the original second-stage diffusion transformer described above has the same network structure as the original second-stage diffusion transformer described above; the original second-stage diffusion transformer is a model obtained by training the original second-stage diffusion transformer.

[0029] Both the first-stage diffusion transformer and the second-stage diffusion transformer mentioned above include at least one diffusion transformer sub-block. As an example, both the first-stage and second-stage diffusion transformers can be composed of at least one (e.g., 12) diffusion transformer sub-blocks connected in series. Each diffusion transformer sub-block includes an adaptive layer normalization module, a multi-head self-attention layer, and a feedforward neural network. The adaptive layer normalization module can include at least one linear layer. This module can take an input feature representation (e.g., a word sequence) and a conditional vector generated by fusing the diffusion time step and category label as input, and output a modulated feature representation sequence. The multi-head self-attention layer can be structured as multiple parallel Scaled Dot-Product Attention heads followed by a linear projection layer. This multi-head self-attention layer can take the feature representation sequence as input and output the self-attention result. The feedforward neural network can be a multilayer perceptron with a structure of a fully connected network-linear activation function (such as GELU or Swish)-fully connected network. The aforementioned feedforward neural network can take the self-attention result as input and the nonlinearly transformed feature representation as output. The original second-stage diffusion transformer also includes a cross-attention module. This cross-attention module can include a scaled dot-product attention head. It can take a query vector, key vector, and value vector as input and a cross-attention context vector as output. This cross-attention module is used to interact with the sub-block features generated by the diffusion transformer sub-blocks and the global information contained in the original image thumbnail. As an example, the diffusion transformer sub-block can be a diffusion transformer base module (DiTBlock) with an Adaptive Layer Normalization Zero (adaLN-Zero) structure.

[0030] Step 102: Based on the preset diffusion time step, image category label and first-stage diffusion transformer, perform first-stage conditional denoising on the first noise tensor to generate a thumbnail latent representation.

[0031] In some embodiments, the execution entity may perform a first-stage conditional denoising process on the first noise tensor based on a preset diffusion time step, image category labels, and a first-stage diffusion transformer to generate a thumbnail latent representation. The thumbnail latent representation may be a feature tensor containing complete global anatomical structure information, generated after the first-stage denoising process. For example, the thumbnail latent representation may be a tensor of size [B, C, 16, 16, 16]. The complete global anatomical structure information may be the overall anatomical layout and topological relationships of the target three-dimensional medical image compressed and represented in the thumbnail latent representation output after the first-stage diffusion transformer denoises the random noise. This complete global anatomical structure information can be used to represent the spatial relative positions, shapes, proportions, and connections between different anatomical components (such as organs and tissue regions).

[0032] In some optional implementations of certain embodiments, the execution entity may perform a first-stage conditional denoising process on the first noise tensor based on a preset diffusion time step, image category label, and a first-stage diffusion transformer to generate a thumbnail latent representation: The first step involves encoding the preset diffusion time step into a time step vector and the image category labels into label vectors. In practice, the execution entity can input the preset diffusion time step into the first embedding layer to output the time step vector. Next, the image category labels are input into the second embedding layer to output the label vector. The first embedding layer can be a pre-trained parameter matrix of dimension [T, D]. T represents the total number of time steps in the diffusion process. For example, T can be 1000, and the preset diffusion time step can be 200. In response to the execution entity inputting the preset diffusion time step 200 into the first embedding layer, the 200th row vector is retrieved from the parameter matrix using an indexing operation as the time step vector. D represents the total dimension of the time step vector. The second embedding layer can be a pre-trained parameter matrix of dimension [K, D]. K represents the total number of image category labels. As an example, the first embedding layer can be a parameter matrix of shape [1000, 768]. For a preset diffusion time step t=250 input to the first embedding layer, the execution entity can retrieve the vector in the 250th row of the first embedding layer as the time step vector using an index. Secondly, the second embedding layer can be a parameter matrix of shape [10, 768]. For an input image category label cls=3 (e.g., corresponding to "brain MRI"), the execution entity can retrieve the vector in the 3rd row of the second embedding layer as the category label vector using an index. Finally, the execution entity can add the 768-dimensional time step vector and the 768-dimensional category label vector to obtain a 768-dimensional second condition vector.

[0033] The second step is to add the time step vector and the label vector together to obtain the condition vector. In practice, the execution entity can perform element-wise addition on the time step vector and the label vector to obtain the condition vector.

[0034] The third step involves performing three-dimensional voxel partitioning on the first noise tensor to obtain a first noise voxel set. In practice, the execution entity can perform a three-dimensional convolution operation on the first noise tensor to divide it into multiple non-overlapping, fixed-size small cubes in space, thus obtaining the first noise voxel set. As an example, the execution entity can divide the 16^3 first noise tensor into 2^3 small cubes, resulting in 8×8×8=512 small cubes.

[0035] The fourth step involves mapping each first noise voxel in the aforementioned first noise voxel set to a first noise word, resulting in a first noise word sequence. In practice, the execution entity can perform a linear transformation operation on each first noise voxel in the aforementioned first noise voxel set to output the first noise word sequence. As an example, the execution entity can arrange all the values ​​contained within the first noise voxel into a one-dimensional vector (e.g., a one-dimensional vector of length 2^3). Then, a pre-trained weight matrix is ​​multiplied by the aforementioned one-dimensional vector to obtain the first noise words.

[0036] The fifth step is to add sine and cosine positional codes to the above first noise word sequence to obtain a position-coded word sequence. In practice, the execution entity can use the sine and cosine positional encoding method to determine the positional code of each first noise word in the above first noise word sequence, and then add each first noise word to its corresponding positional code element by element to generate a position-coded word sequence.

[0037] Step 6: Based on the aforementioned conditional vector, perform conditional feature transformation operations on each positional encoded lexical in the aforementioned positional encoded lexical sequence to obtain a conditional feature transformed lexical sequence. In practice, the aforementioned execution entity can input the aforementioned conditional vector and the aforementioned positional encoded lexical sequence into the aforementioned diffusion transformer sub-block to generate the conditional feature transformed lexical sequence. The aforementioned conditional feature transformation operations include: adaptive layer normalization, multi-head self-attention processing, and nonlinear transformation.

[0038] Step 7: Map and recombine the conditional feature transformation lexical sequence to obtain the thumbnail latent representation. In practice, the execution entity can input the conditional feature transformation lexical sequence into a linear projection layer to generate the thumbnail latent representation. This linear projection layer can be a fully connected layer. As an example, the linear projection layer maps the feature dimensions of each conditional feature transformation lexical in the sequence to a 2^3*C dimension, obtaining the projected feature sequence. Then, the projected feature sequence is rearranged according to its positional order in the original 3D space to generate a tensor of shape [B, C, 16, 16, 16], which serves as the thumbnail latent representation.

[0039] Step 103: Determine the target latent representation space and divide the target latent representation space into a preset number of sub-block spaces to obtain a set of sub-block spaces.

[0040] In some embodiments, the execution entity can determine a target latent representation space and divide the target latent representation space into a predetermined number of sub-block spaces to obtain a set of sub-block spaces. In practice, the execution entity can read the three-dimensional spatial size parameters of the original image latent representation as the target latent representation space. For example, the shape of the target latent representation space can be [64, 64, 64]. Then, the target latent representation space is divided into 2×2×2 sub-block spaces to obtain a set of sub-block spaces. The shape of the sub-block spaces can be [32, 32, 32].

[0041] Step 104: Based on the thumbnail latent representation and the second-stage diffusion transformer, perform the second-stage conditional denoising process on each sub-block space in the sub-block space set to obtain the sub-block latent representation set.

[0042] In some embodiments, the execution entity may perform second-stage conditional denoising processing on each sub-block space in the sub-block space set based on the thumbnail latent representation and the second-stage diffusion transformer to obtain a sub-block latent representation set.

[0043] In addressing the technical challenges of the aforementioned background technologies, and considering the application scenario—high-fidelity medical image synthesis in resource-constrained environments (e.g., research institutions or hospitals typically use mid-range or consumer-grade graphics processors for localized research. These devices have limited video memory (e.g., 24GB or less), but research requires generating extremely high-resolution (e.g., 512³ voxels or higher) 3D images)—the following technical problems often arise: During block generation, the lack of an effective global information transfer and position-aware mechanism leads to discontinuous and inconsistent anatomical structures and artifacts in the generated high-resolution 3D medical images. To meet the following requirements for this application scenario: adaptability to resource-constrained devices and artifact-free images, we have decided to adopt the following solution: In some optional implementations of certain embodiments, the execution entity may perform second-stage conditional denoising processing on each sub-block space in the sub-block space set based on the thumbnail latent representation and the second-stage diffusion transformer to obtain the sub-block latent representation set: The first step is to generate second noise data conforming to a standard Gaussian distribution for each sub-block space in the aforementioned sub-block space set, thus obtaining the second noise data set. In practice, the aforementioned execution entity can call the aforementioned pseudo-random number generation function to generate a random number tensor that conforms to the shape of each sub-block space and in which each value independently follows a standard Gaussian distribution (mean of 0 and variance of 1), as the second noise data, thus obtaining the second noise data set.

[0044] The second step is to determine the spatial coordinate codes of each sub-block space in the aforementioned sub-block space set, thus obtaining a set of spatial coordinate codes. In practice, the aforementioned execution entity can determine the spatial coordinate codes of each sub-block space using the aforementioned Sinusoidal Positional Encoding method, thereby obtaining a set of spatial coordinate codes.

[0045] The third step involves encoding and fusing the preset diffusion time step and the image category labels to obtain a condition vector. In practice, the executing entity can input the preset diffusion time step into the first embedding layer to obtain a time step vector. Then, the image category labels are input into the second embedding layer to obtain a label vector. Finally, the time step vector and the label vector are added element-wise to obtain the condition vector.

[0046] The fourth step involves converting the second noise data set into a second noise word sequence set using the second-stage diffusion transformer described above. In practice, the executing entity can input the second noise data set into the second-stage diffusion transformer to generate the second noise word sequence set. As an example, the second-stage diffusion transformer performs the following operations: First, it divides each second noise data point in the second noise data set into fixed-size sub-block noise data sets, obtaining each sub-block noise data set. Then, it linearly projects each sub-block noise data point in each sub-block noise data set into each second noise word sequence, obtaining the second noise word sequence set.

[0047] Fifth, based on the aforementioned spatial coordinate encoding set and the aforementioned second noise word sequence set, a query vector sequence set is generated. In practice, the executing entity can add each second noise word sequence in the aforementioned second noise word sequence set to the corresponding spatial coordinate encoding in the aforementioned spatial coordinate encoding set element by element to obtain each fused feature sequence. Then, each of the aforementioned fused feature sequences is determined as the query vector sequence set.

[0048] Step 6: The thumbnail latent representation is converted into a thumbnail word sequence using a thumbnail encoder, and this sequence is then determined as a key vector sequence and a value vector sequence. In practice, the execution entity can input the thumbnail latent representation into the thumbnail encoder to generate the thumbnail word sequence. Then, the thumbnail word sequence is simultaneously determined as a key vector sequence and a value vector sequence. The thumbnail encoder can be an encoder that converts input data into fixed-length representation vectors. For example, it can be an encoder that takes a tensor of shape [B, C, 16, 16, 16] as input and outputs a sequence containing 512 feature vectors. The thumbnail encoder can include a 3D voxel block embedding layer and a positional encoding module. The 3D voxel block embedding layer can be a 3D convolutional layer with a kernel size and stride equal to 2. For example, both the kernel size and stride can be 2. The positional encoding module can be used to add spatial location information to the feature sequence.

[0049] Step 7: Based on the aforementioned conditional vector, key vector sequence, value vector sequence, and second-stage diffusion transformer, a forward propagation operation is performed on the aforementioned second noisy word sequence set to generate a sub-block latent representation set. In practice, the executing entity can input the aforementioned conditional vector, key vector sequence, and value vector sequence into the aforementioned second-stage diffusion transformer to obtain the sub-block latent representation set. The sub-block latent representation can be a three-dimensional tensor with shape [B, C, 32, 32, 32]. The forward propagation operation includes inputting the aforementioned second noisy word sequence set into at least one diffusion transformer sub-block, and, according to a preset layer frequency, using the output features generated by the at least one diffusion transformer sub-block during processing as a query vector, performing cross-attention interaction with the aforementioned key vector sequence and value vector sequence. The preset layer frequency can be a preset index list. For example, the preset layer frequency can be [3, 6, 9, 12], indicating that in the processing sequence, the output features of the 3rd, 6th, 9th, and 12th layer diffusion transformer sub-blocks will be used to perform cross-attention interaction.

[0050] The first to eighth steps and related content described above, as an inventive point of this disclosure, combined with step "106" below, solve the technical problem that "during block generation, the lack of an effective global information transmission and position awareness mechanism leads to discontinuities and inconsistent artifacts in the anatomical structures between blocks in the generated high-resolution 3D medical image." The factors leading to discontinuities and inconsistent artifacts in the anatomical structures between blocks in the generated high-resolution 3D medical image are often as follows: During block generation, the lack of an effective global information transmission and position awareness mechanism leads to discontinuities and inconsistent artifacts in the anatomical structures between blocks in the generated high-resolution 3D medical image. If these factors are resolved, it is possible to achieve the effect of displaying high-resolution (e.g., 512×512×512) 3D medical images with consistent anatomical structures and no inter-block artifacts under limited computing resources (e.g., approximately 24GB of video memory). To achieve this effect, firstly, second noise data conforming to a standard Gaussian distribution is generated for each sub-block space in the aforementioned sub-block space set, resulting in a second noise data set. This provides a random starting point for the diffusion model denoising process for each high-resolution sub-block to be generated, enabling the computer to initialize multiple independent sub-block generation tasks in parallel or serially. Second, the spatial coordinate encoding of each sub-block in the aforementioned sub-block spatial set is determined, resulting in a spatial coordinate encoding set. This allows the generation of a unique embedding vector identifying the absolute position of each sub-block in the entire three-dimensional space. Third, the preset diffusion time step and the aforementioned image category labels are encoded and fused to obtain a conditional vector. This allows the diffusion transformer to control the generation progress based on the time step information during denoising and guide the generated content to conform to the specified anatomical structure category based on the category information, thereby achieving accurate conditional image generation. Fourth, the second-stage diffusion transformer converts the aforementioned second noise data set into a second noise word sequence set. This maps the original noisy pixel data into a series of feature vectors rich in semantics that the model can process. Fifth, a query vector sequence set is generated based on the aforementioned spatial coordinate encoding set and the aforementioned second noise word sequence set. This allows the spatial location information of each sub-block to be fused with its content feature information. The fused query vector sequence can simultaneously query the global conditions based on "where is the current sub-block" and "what are the initial features of the current sub-block" in subsequent cross-attention calculations, making the calculation of attention weights more accurate and targeted. Sixth, the thumbnail latent representation is converted into a thumbnail term sequence using a thumbnail encoder, and this thumbnail term sequence is then identified as a key vector sequence and a value vector sequence. Thus, the low-resolution thumbnail generated in the first stage, containing a global anatomical blueprint, can be encoded into a set of queryable feature sequences.The aforementioned sequences, serving as keys and values, constitute a "global knowledge base" within the cross-attention mechanism. This provides a stable and consistent source of structural conditions for the generation of each sub-block, ensuring that the generation of all sub-blocks is constrained by the same set of global planning. Seventh, based on the aforementioned condition vectors, key vector sequences, value vector sequences, and the second-stage diffusion transformer, a forward propagation operation is performed on the aforementioned second set of noise lexical sequences to generate a set of latent representations for sub-blocks. This allows for the gradual removal of components from the noise that do not conform to the target, iteratively generating high-quality sub-block content that meets both local detail requirements and global structural and positional constraints. Finally, combined with step "Step 106" below, decoding and image display are performed, achieving the effect of displaying high-resolution (e.g., 512×512×512) three-dimensional medical images with consistent anatomical structures and no inter-block artifacts, even with limited computing resources (e.g., approximately 24GB of video memory).

[0051] In addressing the technical challenges of the aforementioned scenarios, the application scenario—microscopic or sub-millimeter precision neurosurgical planning—often presents the following technical issues: Existing generation methods, particularly the feature transformation and modulation mechanisms within their diffusion transformer models, cannot ensure that the global conditional information obtained from the thumbnail is deeply, stably, and pixel-level precisely integrated into the generation process of each sub-block. This results in blurred internal structures within sub-blocks and subtle discontinuities between them (e.g., insufficient clarity of delicate nerves and microvessels, or texture discontinuities at stitching boundaries that affect the interpretation of microscopic structures). Considering the following requirements for this application scenario: adaptability to pixel-level precision, adaptability to resource-constrained devices, and consistent image content structure, we have decided to adopt the following solution: In some optional implementations of certain embodiments, the execution entity may perform a forward propagation operation on the second noise lexical sequence set based on the condition vector, the key vector sequence, the value vector sequence, and the second-stage diffusion transformer to generate a sub-block latent representation set: The first step involves using the aforementioned second set of noisy word sequences as input data, and inputting it into a second-stage diffusion transformer that includes at least one diffusion transformer sub-block to perform feature transformation operations: The first feature transformation operation involves conditionally modulating the input data using the aforementioned adaptive layer normalization module and conditional vector to obtain a set of modulated feature representations. In practice, the executing entity can input the conditional vector and the input data into the adaptive layer normalization module of the second-stage diffusion transformer to obtain the set of modulated feature representations.

[0052] The second feature transformation operation involves performing multi-head self-attention processing on the aforementioned modulated feature representation set to obtain a self-attention result set. In practice, the executing entity can input the aforementioned modulated feature representation set into the multi-head self-attention layer of the second-stage diffusion transformer to perform multi-head self-attention processing and obtain a self-attention result set.

[0053] The third feature transformation operation, in response to the diffusion transformer sub-block's level conforming to a preset level frequency, generates a preliminary fused feature set based on the aforementioned self-attention result set, key vector sequence, and value vector sequence. In practice, in response to the current diffusion transformer sub-block's level conforming to the preset level frequency, the executing entity can perform the following cross-attention interactions: First, input the aforementioned self-attention results, key vector sequence, and value vector sequence into a ScaledDot-Product Attention unit to generate a cross-attention result set. Second, add the aforementioned cross-attention result set element-wise to the aforementioned self-attention result set to obtain the preliminary fused features. The second-stage diffusion transformer consists of multiple (e.g., L) diffusion transformer sub-blocks connected in series. The level of the aforementioned diffusion transformer sub-block can be its specific position within the second-stage diffusion transformer, such as layer 1, layer 2, ..., layer L. The preset level frequency can be a preset index list. For example, the preset layer frequencies mentioned above can be [3, 6, 9, 12], indicating that in the processing sequence, the output features of the 3rd, 6th, 9th, and 12th layer diffusion transformer sub-blocks will be used to perform cross-attention interaction.

[0054] As an example, the Scaled Dot-Product Attention unit described above can perform the following operations: First, the dot product of each self-attention result in the self-attention result set and the key vector sequence is divided by a preset scaling factor, and then normalized using a softmax function to obtain an attention weight matrix set. Then, a weighted sum is performed on each attention weight matrix in the attention weight matrix set and the value vector sequence to obtain a cross-attention result set. The preset scaling factor can be the square root of the feature dimension of the query vector. For example, the feature dimension could be 1024, and the preset scaling factor could be 32.

[0055] Optionally, in response to the fact that the level of the diffusion transformer sub-block does not conform to the preset level frequency, the above self-attention result set is determined as the preliminary fusion feature set.

[0056] The fourth feature transformation operation involves performing a nonlinear transformation on the aforementioned preliminary fused feature set to obtain a nonlinear transformed feature set. In practice, the aforementioned execution entity can input the aforementioned self-attention result set into the feedforward neural network of the aforementioned second-stage diffusion transformer to perform a nonlinear transformation operation and obtain a nonlinear transformed feature set.

[0057] The fifth feature transformation operation involves performing a residual concatenation between the aforementioned nonlinear transformation feature set and the aforementioned modulated feature representation set to obtain a residual concatenation result set. This residual concatenation result set is then subjected to layer normalization to obtain the output feature set. In practice, the executing entity can add the aforementioned nonlinear transformation feature set and the aforementioned modulated feature representation set element-wise to obtain the residual concatenation result set.

[0058] The lowest level responds to the diffusion converter sub-blocks, outputting a set of residual connection results. The second-stage diffusion converter consists of multiple (e.g., L) diffusion converter sub-blocks connected in series. The level of these diffusion converter sub-blocks can be their specific location within the second-stage diffusion converter, such as layer 1, layer 2, ..., layer L. Layer L can be the lowest level.

[0059] Optionally, in response to the fact that the level of the diffusion transformer sub-block is not the lowest level, the above residual connection result set is used as input data to perform the above execution feature transformation operation.

[0060] The second step is to perform layer normalization on the residual connection result set to obtain the output feature set. In practice, the execution entity can perform layer normalization (LN) on the residual connection result set to obtain the output feature set.

[0061] The third step involves linearly mapping the output feature set to the sub-block space to obtain a linearly projected feature set. In practice, the execution entity can input the output feature set into a second linear projection layer to obtain the linearly projected feature set. This second linear projection layer can be a fully connected layer without an activation function. For example, the second linear projection layer can be used to map each output feature in the output feature set to a 2^3*C dimension.

[0062] The fourth step involves reorganizing the aforementioned linear projection feature set into a three-dimensional structure to obtain a set of sub-block latent representations. In practice, the execution entity can perform the following operations for each linear projection feature in the aforementioned linear projection feature set: First, arrange the linear projection features into a three-dimensional grid comprising multiple feature vectors of dimension 2^3*C; second, reorganize each feature vector of dimension 2^3*C in the aforementioned three-dimensional grid into a three-dimensional sub-block of [C, 2, 2, 2], and then stitch the various three-dimensional sub-blocks together to obtain the sub-block latent representation. As an example, the shape of the aforementioned sub-block latent representation can be [B, C, 32, 32, 32]. Finally, the sub-block latent representations corresponding to each linear projection feature are determined as the set of sub-block latent representations.

[0063] The first to fourth steps and related content described above, as an inventive point of this disclosure, combined with step "106" below, solve the technical problem that "in existing generation methods, the feature transformation and modulation mechanism inside the diffusion transformer model cannot ensure that the global condition information obtained from the thumbnail is deeply, stably, and pixel-level accurately integrated into the generation process of each sub-block, resulting in blurred internal structures of sub-blocks and subtle discontinuities between sub-blocks (e.g., insufficient clarity of delicate nerves and microvessels, or texture discontinuities at splicing boundaries that affect the interpretation of microstructures)." The factors leading to blurred internal structures of sub-blocks and subtle discontinuities between sub-blocks are often as follows: In existing generation methods, the feature transformation and modulation mechanism inside the diffusion transformer model cannot ensure that the global condition information obtained from the thumbnail is deeply, stably, and pixel-level accurately integrated into the generation process of each sub-block, resulting in blurred internal structures of sub-blocks and subtle discontinuities between sub-blocks (e.g., insufficient clarity of delicate nerves and microvessels, or texture discontinuities at splicing boundaries that affect the interpretation of microstructures). If the above factors are addressed, it is possible to display 3D medical images with sub-millimeter precision structures under limited computational resources. To achieve this, firstly, the aforementioned second set of noisy word sequences is used as input data, and a second-stage diffusion transformer, including at least one diffusion transformer sub-block, is input to perform feature transformation operations. This allows for deep fusion of features modulated across attention mechanisms and rich in global structural information with features representing the initial state and detail basis of the current sub-block. This ensures that subsequent operations are based on "global structure guidance" and "local detail basis." Secondly, the aforementioned residual connection result set is subjected to layer normalization to obtain the output feature set. This alleviates the gradient vanishing problem in deep networks and ensures the numerical stability of the deep network output. Thirdly, the aforementioned output feature set is linearly mapped to the aforementioned sub-block space to obtain a linear projection feature set. This allows high-dimensional abstract features, after a series of complex transformations, to be mapped back to a space matching the data dimension of the target sub-block's latent representation through a linear projection layer. Fourthly, the aforementioned linear projection feature set undergoes 3D structural reconstruction to obtain the sub-block latent representation set. Therefore, based on the order of each feature vector in the sequence (corresponding to its three-dimensional spatial position), the two-dimensional linear projection feature sequence can be rearranged and combined into a three-dimensional voxel tensor. Finally, combined with step "106" below, decoding and image display can be performed, achieving the effect of displaying a three-dimensional medical image with sub-millimeter precision structure under limited computing resources.

[0064] Step 105: Concatenate the potential representations of each sub-block in the sub-block potential representation set to obtain the target potential representation.

[0065] In some embodiments, the execution entity can concatenate the various sub-block latent representations in the aforementioned sub-block latent representation set to obtain the target latent representation. In practice, the execution entity can concatenate the various sub-block latent representations in the aforementioned spatial coordinate encoding set to obtain the target latent representation. For example, the aforementioned sub-block latent representation set contains eight sub-block latent representations with the shape [C, 32, 32, 32]. In the aforementioned spatial coordinate encoding set, each sub-block latent representation is associated with a spatial coordinate code. For example, sub-block latent representation A has coordinates (0,0,0), sub-block latent representation B has coordinates (0,0,1), ..., sub-block latent representation H has coordinates (1,1,1). In three-dimensional space, the various sub-block latent representations are concatenated to obtain the target latent representation. The shape of the aforementioned target latent representation can be [C, 64, 64, 64].

[0066] Step 106: Decode the target latent representation to generate a three-dimensional medical image, and send the three-dimensional medical image to the target display device for display.

[0067] In some embodiments, the execution entity may decode the potential representation of the target to generate a three-dimensional medical image and send the three-dimensional medical image to the target display device for display.

[0068] In some optional implementations of certain embodiments, the aforementioned execution entity may decode the target latent representation to generate a three-dimensional medical image and send the three-dimensional medical image to a target display device for display by performing the following steps: The first step is to input the aforementioned latent representation of the target into a preset decoder to obtain the decoded image. As an example, the preset decoder can be a VQ-VAE decoder.

[0069] The second step involves format conversion of the decoded image to obtain a three-dimensional medical image. In practice, the executing entity can convert the decoded image into the data format required by the target display device to obtain a three-dimensional medical image. As an example, the executing entity can quantize the floating-point pixel values ​​in the decoded image into 8-bit unsigned integers.

[0070] The third step is to send the aforementioned three-dimensional medical images to the target display device for display. In practice, the executing entity can send the aforementioned three-dimensional medical images to the target display device via an internal computer bus (such as PCIe) and external interfaces (such as HDMI, DisplayPort) to control the rendering pipeline of the display hardware to render the aforementioned three-dimensional medical images into a visual image on a two-dimensional screen.

[0071] The above embodiments of this disclosure have the following beneficial effects: The two-stage diffusion-based 3D medical image generation method of some embodiments of this disclosure, through a hierarchical generation architecture and global condition control, achieves the effect of efficiently generating high-resolution, structurally consistent 3D medical images under limited computing resources. Specifically, the reason for excessive memory consumption and artifacts of discontinuous and inconsistent inter-block anatomical structures in the generated results is that: Existing methods for directly generating high-resolution 3D latent representations cause the computer's graphics processor to face enormous memory pressure, making it impossible to execute in a resource-constrained hardware environment. In existing simple block generation strategies, due to the lack of an effective global information transmission mechanism when generating each sub-block, the final stitched complete image has artifacts of discontinuous and inconsistent inter-block anatomical structures. Based on this, the two-stage diffusion-based 3D medical image generation method of some embodiments of this disclosure first generates a first noise tensor in response to detecting that the graphics processor memory capacity of the target display device is lower than a preset single-peak memory threshold. Thus, the computer can intelligently trigger a high-resolution image generation process with low memory consumption in a resource-constrained hardware environment. Secondly, based on the preset diffusion time step, image category labels, and the first-stage diffusion transformer, the first noise tensor undergoes first-stage conditional denoising to generate a thumbnail latent representation. This generates a low-resolution "sketch" containing complete global anatomical information, providing a unified structural plan and constraints for subsequent high-resolution block generation. Then, the target latent representation space is determined, and this target latent representation space is divided into a preset number of sub-block spaces, resulting in a set of sub-block spaces. This decomposes the high-dimensional generation task into multiple independently processable low-dimensional sub-tasks. Next, based on the thumbnail latent representation and the second-stage diffusion transformer, second-stage conditional denoising is performed on each sub-block space in the sub-block space set, resulting in a set of sub-block latent representations. This allows for real-time reference to the global thumbnail via a cross-attention mechanism, and the absolute spatial position of each sub-block is determined through 3D coordinate position encoding. This ensures that the content of each sub-block conforms to the global structure and is precisely aligned spatially during generation, fundamentally avoiding structural errors and discontinuities between blocks. Next, the latent representations of each sub-block in the aforementioned set of sub-block latent representations are concatenated to obtain the target latent representation. This allows all the location-defined sub-blocks generated under global constraints to be losslessly assembled into a complete, coherent, high-resolution latent representation. Finally, the target latent representation is decoded to generate a 3D medical image, which is then sent to the target display device for display. This process outputs a final high-quality 3D medical image, completing the entire generation process from conditions to pixels, and visualizing the results for use.This implementation achieves the technical effect of reducing peak video memory usage while ensuring consistent anatomical structure of the generated image and eliminating inter-block artifacts.

[0072] Continue to refer to Figure 2 As a response to the above Figure 1 The present disclosure provides some embodiments of a two-stage diffusion-based three-dimensional medical image generation device, which are similar to the implementation of the method shown. Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0073] like Figure 2 As shown, the two-stage diffusion-based three-dimensional medical image generation 200 in some embodiments includes: a detection unit 201, a first denoising processing unit 202, a determination unit 203, a second denoising processing unit 204, a stitching unit 205, and a decoding unit 206. The detection unit 201 is configured to generate a first noise tensor in response to detecting that the graphics processor memory capacity of the target display device is lower than a preset single peak memory threshold; the first denoising processing unit 202 is configured to perform a first-stage conditional denoising process on the first noise tensor based on a preset diffusion time step, image category label, and a first-stage diffusion transformer to generate a thumbnail latent representation; the determination unit 203 is configured to determine a target latent representation space and divide the target latent representation space into a preset number of sub-block spaces to obtain a set of sub-block spaces; the second denoising processing unit 204 is configured to perform a second-stage conditional denoising process on each sub-block space in the set of sub-block spaces based on the thumbnail latent representation and the second-stage diffusion transformer to obtain a set of sub-block latent representations; the stitching unit 205 is configured to stitch together each sub-block latent representation in the set of sub-block latent representations to obtain a target latent representation; and the decoding unit 206 is configured to decode the target latent representation to generate a three-dimensional medical image and send the three-dimensional medical image to the target display device for display.

[0074] It is understandable that the units described in the device 200 are related to the reference. Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units contained therein, and will not be repeated here.

[0075] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0076] like Figure 3As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0077] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.

[0078] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.

[0079] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0080] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0081] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently without being assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: generate a first noise tensor in response to detecting that the graphics processor memory capacity of the target display device is lower than a preset single-peak memory threshold; perform a first-stage conditional denoising process on the first noise tensor based on a preset diffusion time step, image category labels, and a first-stage diffusion transformer to generate a thumbnail latent representation; determine a target latent representation space and divide the target latent representation space into a preset number of sub-block spaces to obtain a set of sub-block spaces; perform a second-stage conditional denoising process on each sub-block space in the set of sub-block spaces based on the thumbnail latent representation and a second-stage diffusion transformer to obtain a set of sub-block latent representations; concatenate the sub-block latent representations in the set of sub-block latent representations to obtain a target latent representation; decode the target latent representation to generate a three-dimensional medical image, and send the three-dimensional medical image to the target display device for display.

[0082] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0083] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0084] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a detection unit, a first denoising processing unit, a determination unit, a second denoising processing unit, a splicing unit, and a decoding unit. The names of these units do not necessarily limit the specific unit; for example, the detection unit may also be described as "a unit that generates a first noise tensor in response to detecting that the graphics processor memory capacity of the target display device is lower than a preset single-peak memory threshold."

[0085] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0086] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A method for generating three-dimensional medical images based on two-stage diffusion, comprising: In response to detecting that the graphics processor memory capacity of the target display device is lower than a preset single peak memory threshold, a first noise tensor is generated; Based on a preset diffusion time step, image category labels, and a first-stage diffusion transformer, the first noise tensor is subjected to a first-stage conditional denoising process to generate a thumbnail latent representation. Determine the target latent representation space and divide the target latent representation space into a predetermined number of sub-block spaces to obtain a set of sub-block spaces; Based on the thumbnail latent representation and the second-stage diffusion transformer, the second-stage conditional denoising process is performed on each sub-block space in the sub-block space set to obtain the sub-block latent representation set. The target latent representation is obtained by concatenating the latent representations of each sub-block in the sub-block latent representation set. The target latent representation is decoded to generate a three-dimensional medical image, and the three-dimensional medical image is sent to the target display device for display.

2. The method according to claim 1, wherein, Decoding the target latent representation to generate a three-dimensional medical image, and sending the three-dimensional medical image to the target display device for display, includes: The latent representation of the target is input into a preset decoder to obtain a decoded image; The decoded image is converted to a different format to obtain a three-dimensional medical image; The three-dimensional medical image is sent to the target display device for display.

3. The method according to claim 1, wherein, The step of generating a first noise tensor in response to detecting that the graphics processor memory capacity of the target display device is lower than a preset single peak memory threshold includes: In response to detecting that the graphics processor memory capacity of the target display device is lower than a preset single peak memory threshold, a first noise tensor space size is determined based on a preset space size; Based on the pseudo-random number generation function and the size of the first noise tensor space, a tensor conforming to a standard Gaussian distribution is generated as the first noise tensor.

4. The method according to claim 1, wherein, The first-stage diffusion converter and the second-stage diffusion converter are obtained through the following steps: Acquire raw 3D medical images; The original three-dimensional medical image is encoded by a preset encoder to obtain the latent representation of the original image; The latent representation of the original image is downsampled to obtain a thumbnail of the original image; The original image latent representation is divided into a preset number of original sub-block latent representations to obtain a set of original sub-block latent representations; Based on the denoising diffusion probability model, the original image thumbnail, and the original sub-block latent representation set, the original first-stage diffusion transformer and the original second-stage diffusion transformer are trained to obtain the first-stage diffusion transformer and the second-stage diffusion transformer.

5. The method according to claim 4, wherein, Both the original first-stage diffusion transformer and the original second-stage diffusion transformer include at least one diffusion transformer sub-block. The diffusion transformer sub-block includes an adaptive layer normalization module, a multi-head self-attention layer, and a feedforward neural network. The original second-stage diffusion transformer also includes a cross-attention module, wherein the cross-attention module is used to interact with the sub-block features generated by the diffusion transformer sub-block and the global information contained in the original image thumbnail.

6. The method according to claim 1, wherein, The first-stage conditional denoising process, based on a preset diffusion time step, image category labels, and a first-stage diffusion transformer, is performed on the first noise tensor to generate a thumbnail latent representation, including: The preset diffusion time step is encoded into a time step vector, and the image category label is encoded into a label vector; The condition vector is obtained by adding the time step vector and the label vector. The first noise tensor is divided into three-dimensional voxels to obtain the first noise voxel set; Map each first noise voxel in the first noise voxel set to a first noise word to obtain a first noise word sequence; Add sine and cosine position codes to the first noisy word sequence to obtain a position-coded word sequence; Based on the conditional vector, conditional feature transformation is performed on each positional encoded lexical in the positional encoded lexical sequence to obtain a conditional feature transformed lexical sequence. The conditional feature transformation operation includes: adaptive layer normalization, multi-head self-attention processing, and nonlinear transformation. The conditional feature transformation lexical sequence is mapped and recombined to obtain a thumbnail potential representation.

7. A three-dimensional medical image generation device based on two-stage diffusion, comprising: The detection unit is configured to generate a first noise tensor in response to detecting that the graphics processor memory capacity of the target display device is lower than a preset single peak memory threshold; The first denoising processing unit is configured to perform a first-stage conditional denoising process on the first noise tensor based on a preset diffusion time step, image category label and a first-stage diffusion transformer to generate a thumbnail latent representation. The determining unit is configured to determine the target latent representation space and divide the target latent representation space into a preset number of sub-block spaces to obtain a set of sub-block spaces; The second denoising processing unit is configured to perform second-stage conditional denoising processing on each sub-block space in the sub-block space set based on the thumbnail latent representation and the second-stage diffusion transformer to obtain a sub-block latent representation set. The splicing unit is configured to splice the potential representations of each sub-block in the set of potential representations of sub-blocks to obtain the target potential representation; The decoding unit is configured to decode the target latent representation to generate a three-dimensional medical image, and to send the three-dimensional medical image to the target display device for display.

8. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 6.

9. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 6.