A Laparoscopic Image Generation Method Based on the Stable Diffusion Model
By sampling different time steps as cache points in the diffusion model and using Kalman filtering technology for noise prediction, the problem of slow inference speed of diffusion model in laparoscopic image generation is solved, achieving more efficient image generation.
Patent Information
- Application Number
- CN202510437623.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The diffusion model is slow in inference in laparoscopic image generation, resulting in high computational and time costs.
By sampling different time steps as cache points, the Kalman filtering hyperparameters are calculated, and the Kalman filtering technology is used to predict noise during the denoising process to reduce redundant calculations.
The inference speed of diffusion models in laparoscopic image generation is significantly improved, and the generation rate is increased by 2 times while maintaining the high quality of the image.
Smart Images

Figure CN119963441B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image generation technology, and particularly to a laparoscopic image generation method based on a Stable Diffusion model. Background Art
[0002] The Stable Diffusion (SD) model, as a generative deep learning model, has shown great potential and broad application prospects in the field of image generation. By learning a large amount of image data, it can generate high-quality and diverse images according to given text prompts or other conditions. In the aspect of medical image generation, the SD model also shows good application prospects. For example, it can be used to generate images in laparoscopic surgery, providing richer data support for computer-aided intervention (CAI) systems and helping to improve the safety and accuracy of surgeries. However, although the SD model can generate high-quality images, it faces the problems of high computational cost and slow inference speed, which are mainly attributed to the following aspects:
[0003] 1. Multi-step denoising process. Diffusion models usually need to go through a denoising process of multiple time steps to gradually generate images, which leads to a large amount of computational overhead.
[0004] 2. Complex model structure. In order to generate high-quality images, diffusion models usually adopt relatively complex network structures, such as UNet, etc. Although these complex network structures can improve the quality of image generation, they also increase the computational complexity and time cost.
[0005] To solve the problem of slow inference speed of diffusion models, researchers have proposed a variety of acceleration methods, mainly including the following two categories:
[0006] Reducing the number of sampling steps. By optimizing the sampling strategy or using specific algorithms to reduce the required number of sampling steps, thereby accelerating the generation speed. For example, the DDIM (https: / / arxiv.org / abs / 2010.02502) method converts random noise into the initial image and only needs to perform one model evaluation, greatly reducing the number of sampling steps. In addition, there are also some methods based on consistency models, such as Consistency Models and Latent Consistency Models, which can generate high-quality images with fewer sampling steps by performing consistency modeling in the latent space.
[0007] Optimize model structure and computational efficiency. Optimize the network structure of the diffusion model to reduce redundant computation and improve the model's reasoning efficiency. For example, the DeepCache method (https: / / arxiv.org / abs / 2312.00858) exploits the inherent temporal redundancy of adjacent steps in the reverse denoising process of the diffusion model to cache and retrieve features across adjacent denoising stages, thereby reducing redundant computation. This method speeds up Stable Diffusion v1.5 by 2.3 times without any additional training, and the CLIP score drops by only 0.05.
[0008] In summary, although the diffusion model has important application value in the field of image generation, its slow reasoning speed needs to be solved urgently. By combining the image generation capability of the stable diffusion model with the existing acceleration methods, it is expected that the generation speed will be significantly improved while maintaining the quality of image generation, providing more effective technical support for fields with high data set requirements such as computer-assisted intervention. Summary of the invention
[0009] In view of this, the present application provides a laparoscopic image generation method based on a stable diffusion model.
[0010] The present application discloses a laparoscopic image generation method based on a stable diffusion model, which comprises:
[0011] Step 1: Sample n different time steps as cache points; n is a positive integer;
[0012] Step 2: Based on the cache points, calculate the hyperparameters of the Kalman filter at all time steps by generating N laparoscopic images; N is a positive integer;
[0013] Step 3: Convert the designated prompt words of the laparoscopic image into semantic vectors, randomly generate a Gaussian noise matrix, use the semantic vector and Gaussian noise matrix as the initial input of the noise prediction network, and generate the laparoscopic image based on the hyperparameters of the Kalman filter at the corresponding time step during the denoising process.
[0014] Furthermore, the step 1 comprises:
[0015] In order to obtain n different time steps, arrive Take n linear distribution values in the range and then pass them through the power function After transformation, the distribution value after transformation is Within the range, where x represents the n linearly distributed values obtained, T represents the total number of time steps in the denoising process, center represents the central position of the sampling point, pow represents the non-linear degree of the sampling point, total_numbers represents the number of points to be sampled, and then convert the obtained distribution values into integers and remove duplicates. Add the center offset to the de-duplicated values to obtain the final time steps as cache points. By gradually adjusting the pow value, ensure that the number of generated cache points is equal to n.
[0016] Further, step 2 includes:
[0017] Step 21: Use a text encoder to convert the specified prompt words of the laparoscopic image into a semantic vector, and randomly initialize a three-dimensional Gaussian noise matrix , during initialization, t = T, which serves as the initial input to the noise prediction network;
[0018] Step 22: Use the noise prediction network without cache to output the noise at the current time step according to the input and the semantic vector, and subtract it from the input to obtain the input for the next time step ; is the result of denoising the Gaussian noise matrix at time step t + 1. At t = T, is the randomly initialized Gaussian noise matrix , is the input for the next time step; where the value of t ranges from T to 1, and T is a positive integer;
[0019] Step 23: Use the noise prediction network with cache to output the biased noise at the current time step t according to the input and the semantic vector, and use the noise output by the noise prediction network without cache, the noise predicted for the current time step at the previous time step t + 1, to calculate the Kalman filter hyperparameters and ;
[0020] Step 24: Repeat steps 22 and 23 a total of T times, where the subscript t of the input starts from T and decreases by 1 each time in the repeated process, that is, the values of t are T, T - 1, T - 2,..., 2, 1 in sequence;
[0021] Step 25: Repeat steps 21 and 24 a total of N times, where N is the number of laparoscopic images to be generated in the pre-computation process.
[0022] Further, step 22 includes:
[0023] Step 221: The noise prediction network of the Stable Diffusion model takes as input and fully executes the main branch, i.e., all downsampling layers , intermediate layers , and upsampling layers , and automatically caches the output of the second layer of the upsampling layer as . During the calculation, the semantic vector is fused into the network calculation through the cross-attention mechanism, and the output of the last layer of the upsampling layer is used as the original predicted noise of the Stable Diffusion model noise prediction network at time step ;
[0024] Step 222: Use the default sampler PNDM of the Stable Diffusion model to subtract the original predicted noise from the input to obtain the input for the next time step;
[0025] Step 223: Based on the input for the next time step, obtain the early prediction value of the noise for the next time step.
[0026] Furthermore, the said Step 222 includes:
[0027] The expression of the input for the next time step, that is, the PNDM sampler formula is:
[0028]
[0029] In the formula, and respectively represent the sampling results of the Stable Diffusion model at time steps and , that is, the inputs of the Stable Diffusion model at time steps t - 1 and t, and represent the noise levels at time steps t and , represents the original predicted noise of the Stable Diffusion model noise prediction network at time step ;
[0030] The said Step 223 includes:
[0031] Obtain the early prediction value of the noise for the next time step through the following formula:
[0032]
[0033]
[0034] Perform an equality operation on the above two formulas regarding to derive:
[0035]
[0036]
[0037] .
[0038] Furthermore, the said step 23 includes:
[0039] Step 231: The noise prediction network of the Stable Diffusion model takes as the input, only calculates the first layer of the downsampling layer , and jumps to connect the result output by the first layer with the result cached most recently, and they are jointly used as the input of the last layer of the upsampling layer. When performing calculations, the semantic vector is fused into the network calculations through the cross-attention mechanism, and finally the biased noise of the Stable Diffusion model at the current time step is calculated;
[0040] Step 232: Use the predicted value in advance of the noise of the current time step at the previous time step, the original predicted noise obtained in step 221, and the biased noise obtained in step 231 to calculate the Kalman filter hyperparameters and at the current time step t.
[0041] Furthermore, in the said step 232, the calculation formulas for the Kalman filter hyperparameters and at the current time step t are respectively:
[0042]
[0043]
[0044] In the formula, N is the number of laparoscopic images to be generated for pre-calculation, The initial value is 0.
[0045] Furthermore, the said step 3 includes:
[0046] Step 31: Use the text encoder to convert the laparoscopic image with the specified prompt "CholecT45" into a semantic vector, and randomly initialize a three-dimensional Gaussian noise matrix , during initialization, , which serves as the initial input to the noise prediction network;
[0047] Step 32: Use the noise prediction network to obtain the noise of the input ;
[0048] Step 33: Use Kalman filtering to estimate , and obtain the optimal estimated value , and use the sampler to obtain the input for the next time step;
[0049] Step 34: Repeat Step 32 and Step 33 a total of T times, where the subscript of the input in Step 32 starts from T and is subtracted by 1 each time during the repetition process, i.e., T, T - 1, T - 2,..., 2, 1, until finally obtaining the predicted image latent space representation ;
[0050] Step 35: Use the decoder to generate the final laparoscopic image from .
[0051] Furthermore, the said Step 32 includes:
[0052] Noise prediction network calculation, which is divided into two cases: The first case is to completely calculate and cache the time steps of partial outputs in the stable diffusion model noise prediction network. At the time step serving as the cache point, the stable diffusion model noise prediction network completely executes the main branch, i.e., all downsampling layers , intermediate layers , and upsampling layers , and automatically caches the output of the second layer of the upsampling layer as , and takes the output of the last layer of the upsampling layer as the original predicted noise of the stable diffusion model noise prediction network at the current time step ; The second case is the time steps that do not require complete execution of network calculations, i.e., non-cache points. For non-cache points, when the non-cache point is the time step , the stable diffusion model noise prediction network only calculates the first layer of the downsampling layer , and jumps and connects the result of the first layer output with the result cached most recently, which together serve as the input to the last layer of the upsampling layer , and calculate to obtain the stable diffusion model at the current time step The original prediction noise of ; the input in both cases is , will use the cross-attention mechanism to fuse semantic vectors into network calculations.
[0053] Further, the step 33 comprises:
[0054] Use Previous Time Step The noise of the current time step of the forecast As the current time step The predicted value of
[0055] Based on the previous time step The state variance , current time step Noise variance of the predicted values , get the current time step The state prediction variance :
[0056]
[0057] In the formula, Indicates the current time step The state prediction variance at time , Represents the previous time step The state variance at Indicates the current time step The noise variance of the predicted value when ;
[0058] Calculate the Kalman gain at time step t , which is used to balance the prediction error and observation error:
[0059]
[0060] In the formula, Indicates the current time step The state prediction variance at time , Indicates the noise variance of the observation value at the current time step;
[0061] The noise obtained in step 32 As the observation value, combined with the Kalman gain correction prediction state at time step t, the optimal estimate of the noise after filtering is obtained, that is, the filtering output result :
[0062]
[0063] In the formula, represents the Kalman gain at time step t, Represents the previous time step The noise of the current time step of the forecast, Represents the current time step of the observed value, i.e., the output result of the noise prediction network; update the state value variance , to reflect the new uncertainty:
[0064]
[0065] In the formula, represents the state prediction variance at the current time step . Take the optimal estimated value of the filtered noise as the gradient actually used by the sampler at time step . After sampling by the PNDM sampler, we get , and send it to the noise prediction network of the stable diffusion model at the next time step.
[0066] Due to the above technical solutions, the present application has the following advantages: The present invention significantly improves the inference speed of the diffusion model in laparoscopic image generation while maintaining the high quality of the generated images. Specifically:
[0067] 1. Improved generation rate: Through the feature map caching and Kalman filtering techniques, the generation rate is increased by 2 times compared with the original stable diffusion model, significantly reducing the time required to generate images.
[0068] 2. Improved image quality: Use the FID and KID metrics to evaluate the quality of the generated images. In the vid01&vid49 styles, the FID using only the feature map caching technique is 74.05, and the KID is 0.06683±0.00174. After introducing the pre-computed Kalman filtering technique, the FID is 71.42, and the KID is 0.06344±0.00174. The results show that the generated results of this generation method are closer to the real data CholecT45 distribution.
[0069] 3. Robustness and stability: The Kalman filtering technique improves the accuracy of noise prediction, enhances the robustness and stability of the model, and further improves the quality and consistency of the generated images. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0071] Figure 1 is a schematic flowchart of a laparoscopic image generation method based on a stable diffusion model according to an embodiment of the present application;
[0072] Figure 2 Schematic diagram of the process of diffusion model image denoising based on Kalman filter according to an embodiment of the present application;
[0073] Figure 3 Schematic diagram of the local process of pre - calculating Kalman hyperparameters according to an embodiment of the present application;
[0074] Figure 4 Schematic diagram of the overall process of pre - calculating Kalman hyperparameters according to an embodiment of the present application;
[0075] Figure 5 Schematic diagram of the overall process of diffusion model image denoising based on Kalman filter according to an embodiment of the present application. Detailed implementation manners
[0076] The present application will be further described in conjunction with the accompanying drawings and embodiments. The described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art shall fall within the scope of protection of the embodiments of the present application.
[0077] The present invention aims to solve the problem of slow inference speed of the diffusion model in laparoscopic image generation in the prior art. Although the diffusion model can generate high - quality images, its multi - step denoising process and complex model structure lead to high computational costs and slow inference speed, which limits its use in scenarios with limited computing resources. Therefore, the goal of the present invention is to optimize the inference process of the diffusion model, improve the efficiency of image generation, while maintaining the accuracy of image generation, and provide better data support for computer - aided intervention (CAI) systems.
[0078] See Figure 1 and Figure 2 , an embodiment of a method for generating laparoscopic images based on a stable diffusion model provided by the present application includes:
[0079] Step 1: Sample n different time steps as cache points; n is a positive integer;
[0080] Step 2: Based on the cache points, calculate the hyperparameters of Kalman filter for all time steps through the process of generating N laparoscopic images; N is a positive integer;
[0081] Step 3: Convert the laparoscopic image specified prompt word into a semantic vector, randomly generate a Gaussian noise matrix, use the semantic vector and the Gaussian noise matrix as the initial input of the noise prediction network, and generate laparoscopic images based on the hyperparameters of Kalman filter for the corresponding time steps during the denoising process.
[0082] Optionally, the Step 1 includes:
[0083] To obtain n different time steps, n linearly distributed values are taken from the range between and , and then transformed through the power function . The transformed distributed values are within the range of . Among them, x represents the n linearly distributed values taken, T represents the total number of time steps in the denoising process, center represents the center position of the sampling point, pow represents the non - linear degree of the sampling point, total_numbers represents the number of points to be sampled. Then, the obtained distributed values are converted to integers and duplicates are removed. The values after removing duplicates are added with the center offset to obtain the final time steps as cache points. By gradually adjusting the pow value, it is ensured that the number of generated cache points is equal to n.
[0084] Optionally, referring to Figure 4 , step 2 includes:
[0085] Step 21: Use a text encoder to convert the specified prompt word of the laparoscopic image into a semantic vector, and randomly initialize a three - dimensional Gaussian noise matrix . When initializing, t = T, which serves as the initial input of the noise prediction network;
[0086] Step 22: Use the noise prediction network without cache to output the noise at the current time step according to the input and the semantic vector, and subtract it from the input to obtain the input for the next time step; is the result of denoising the Gaussian noise matrix at the time step t + 1. When t = T, is the randomly initialized Gaussian noise matrix , is the input for the next time step. Among them, the value of t ranges from T to 1, and T is a positive integer;
[0087] Step 23: Use the noise prediction network with cache to output the biased noise at the current time step t according to the input and the semantic vector. Utilize the noise output by the noise prediction network without cache and the noise predicted for the current time step at the previous time step t + 1 to calculate the Kalman filter hyperparameters and for the current time step t;
[0088] Step 24: Repeat step 22 and step 23 a total of T times. Among them, the subscript t of the input starts from T and decreases by 1 each time in the repeating process, that is, the values of t are successively T, T - 1, T - 2,..., 2, 1;
[0089] Step 25: Repeat Step 21 and Step 24 for a total of N times, where N is the number of laparoscopic images to be generated in the pre-computation process.
[0090] Optionally, participate in Figure 3 , and the said Step 22 includes:
[0091] Step 221: The noise prediction network of the StableDiffusion model takes as the input, fully executes the main branch, that is, all downsampling layers , intermediate layers , and upsampling layers , and automatically caches the output of the second layer of the upsampling layer as . During the calculation, the semantic vector is fused into the network calculation through the cross-attention mechanism, and the output of the last layer of the upsampling layer is used as the original predicted noise of the StableDiffusion model noise prediction network at time step ;
[0092] Step 222: Use the default sampler PNDM of the StableDiffusion model to subtract the original predicted noise from the input to obtain the input for the next time step;
[0093] Step 223: According to the input for the next time step, obtain the early prediction value of the noise for the next time step.
[0094] Optionally, the said Step 222 includes:
[0095] The expression of the input for the next time step, that is, the PNDM sampler formula is:
[0096]
[0097] In the formula, and respectively represent the sampling results of the StableDiffusion model at time steps and , that is, the inputs of the StableDiffusion model at time steps t-1 and t, and represent the noise levels at time steps t and , represents the original predicted noise of the StableDiffusion model noise prediction network at time step ;
[0098] Step 223 includes:
[0099] Obtain the advanced prediction value of the noise at the next time step through the following formula :
[0100]
[0101]
[0102] Perform an equality operation on the above two formulas regarding and deduce:
[0103]
[0104]
[0105] .
[0106] Optionally, referring to Figure 3 , Step 23 includes:
[0107] Step 231: The noise prediction network of the stable diffusion model takes as the input, only calculates the first layer of the downsampling layer , and skips the downsampling layer as well as the upsampling layer , and jumps the result output by the first layer to connect with the result of the most recent cache , and they are jointly used as the input of the last layer of the upsampling layer . When performing calculations, fuse the semantic vector into the network calculation through the cross-attention mechanism, and finally calculate the biased noise of the stable diffusion model at the current time step ;
[0108] Step 232: Use the advanced prediction value of the noise at the previous time step for the current time step , the original predicted noise obtained in Step 221 , the biased noise obtained in Step 231 to calculate the Kalman filter hyperparameters and at the current time step t.
[0109] Optionally, in Step 232, the calculation formulas for the Kalman filter hyperparameters and at the current time step t are respectively:
[0110]
[0111]
[0112] Wherein, N is the number of laparoscopic images to be generated in the pre - calculation, and the initial value is 0.
[0113] Optionally, referring to Figure 5 , step 3 includes:
[0114] Step 31: Use a text encoder to convert the laparoscopic image specified prompt "CholecT45" into a semantic vector, and randomly initialize a three - dimensional Gaussian noise matrix , during initialization, , as the initial input of the noise prediction network;
[0115] Step 32: Use the noise prediction network to obtain the noise of the input ;
[0116] Step 33: Use Kalman filtering to estimate , obtain the optimal estimated value , and use a sampler to obtain the input for the next time step;
[0117] Step 34: Repeat step 32 and step 33 for a total of T times, where the subscript of the input in step 32 starts from T and is subtracted by 1 each time during the repetition process, that is, T, T - 1, T - 2,..., 2, 1, until finally obtaining the predicted image latent space representation ;
[0118] Step 35: Use a decoder to generate the final laparoscopic image from .
[0119] Optionally, step 32 includes:
[0120] The noise prediction network calculation is divided into two cases: The first case is to fully calculate and cache the time steps of part of the output in the stable diffusion model noise prediction network. At the time step used as the cache point, the stable diffusion model noise prediction network fully executes the main branch, that is, all downsampling layers , intermediate layers , and upsampling layers , and automatically caches the output of the second layer of the upsampling layer as , and uses the output of the last layer of the upsampling layer as the output of the stable diffusion model noise prediction network at the current time step The second case is the time step that does not require the full execution of the network calculation, that is, the non-cached point. For the non-cached point, when the non-cached point is the time step When the stable diffusion model noise prediction network only calculates the first layer of the downsampling layer , and compare the output of the first layer with the most recently cached result Skip connection, together as the last layer of the upsampling layer The stable diffusion model is calculated based on the input of the current time step. The original prediction noise of ; the input in both cases is , will use the cross-attention mechanism to fuse semantic vectors into network calculations.
[0121] Optionally, the step 33 includes:
[0122] Use Previous Time Step The noise of the current time step of the forecast As the current time step The predicted value of
[0123] Based on the previous time step The state variance , current time step Noise variance of the predicted values , get the current time step The state prediction variance :
[0124]
[0125] In the formula, Indicates the current time step The state prediction variance at time , Represents the previous time step The state variance at Indicates the current time step The noise variance of the predicted value when ;
[0126] Calculate the Kalman gain at time step t , which is used to balance the prediction error and observation error:
[0127]
[0128] In the formula, Indicates the current time step The state prediction variance at time , Indicates the noise variance of the observation value at the current time step;
[0129] The noise obtained in step 32 As the observed value, the Kalman gain at time step t is combined to correct the predicted state, and the optimal estimated value of the noise after filtering is obtained, that is, the filtered output result. :
[0130]
[0131] In the formula, represents the Kalman gain at time step t, represents the noise at the current time step predicted at the previous time step , represents the observed value at the current time step , that is, the output result of the noise prediction network; update the variance of the state value , to reflect the new uncertainty:
[0132]
[0133] In the formula, represents the state prediction variance at the current time step . The optimal estimated value of the filtered noise is used as the gradient actually used by the sampler at time step . After sampling by the PNDM sampler, is obtained and sent to the noise prediction network of the stable diffusion model at the next time step.
[0134] The Kalman filter of the present application combines the noise prediction and caching techniques of the stable diffusion model. By pre-computing the Kalman filter hyperparameters, the noise predicted by the noise prediction network of the stable diffusion model at each time step is corrected during the denoising process, realizing more accurate noise prediction.
[0135] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present application can still be modified or equivalently replaced, and any modification or equivalent replacement that does not depart from the spirit and scope of the present application should be covered within the protection scope of the claims of the present application.
Claims
1. A laparoscopic image generation method based on a stable diffusion model, characterized in that: include: Step 1: Sample n different time steps as cache points; n is a positive integer; Step 2: Based on the cache points, calculate the hyperparameters of the Kalman filter at all time steps by generating N laparoscopic images; N is a positive integer; Step 3: Convert the designated prompt words of the laparoscopic image into semantic vectors, randomly generate a Gaussian noise matrix, use the semantic vectors and Gaussian noise matrix as the initial input of the noise prediction network, and generate the laparoscopic image based on the hyperparameters of the Kalman filter at the corresponding time step during the denoising process; The step 2 comprises: Step 21: Use the text encoder to convert the specified prompt words of the laparoscopic image into semantic vectors and randomly initialize a three-dimensional Gaussian noise matrix ,When initialized, t=T, serves as the initial input of the noise prediction network; Step 22: Use the uncached noise prediction network to predict the input and semantic vector, outputs the noise of the current time step, and converts it from the input Subtract it from the input to get the next time step ; is the Gaussian noise matrix for the t+1 time step The denoising result, at t=T, is a randomly initialized Gaussian noise matrix , is the input of the next time step; where t ranges from T to 1, and T is a positive integer; Step 23: Use the cached noise prediction network to predict the input And the semantic vector, output the biased noise of the current time step t , using the noise of the unbuffered noise prediction network output , the current time step noise predicted by the previous time step t+1 , calculate the Kalman filter hyperparameters for the current time step t and ; Step 24: Repeat steps 22 and 23 for a total of T times, where input The value of the subscript t starts from T, and decreases by 1 each time the process is repeated, that is, the value of t is T, T-1, T-2, ..., 2, 1 in sequence; Step 25: repeat steps 21 and 24 for a total of N times, where N is the number of laparoscopic images that need to be generated in the pre-calculation process; The step 22 comprises: Step 221: Stabilize the diffusion model noise prediction network As input, the main branch is fully executed, i.e. all downsampling layers , Middle Layer , and upsampling layers , and automatically cache the second layer of the upsampling layer The output is , the semantic vector is integrated into the network calculation through the cross attention mechanism during calculation, and the last layer of the upsampling layer The output of the stable diffusion model noise prediction network is used at time step The original prediction noise ; Step 222: Use the default sampler PNDM of the stable diffusion model to convert the original prediction noise From input Subtract it from the input to get the next time step ; Step 223: Based on the input of the next time step , get the advance prediction value of the noise at the next time step ; The step 222 comprises: Input for the next time step The expression of , that is, the PNDM sampler formula is: In the formula, and Respectively represent the time step and The sampling result of the stable diffusion model at time step t-1 and t, that is, the input of the stable diffusion model at time step t-1 and t, and Indicates that at time step t and The noise level at represents the stable diffusion model noise prediction network at time step The original prediction noise of The step 223 includes: The advance prediction value of the noise at the next time step is obtained by the following formula: : The above Equalizing the two formulas, we can deduce: The step 23 comprises: Step 231: Stabilize the diffusion model noise prediction network As input, only the first layer of the downsampling layer is calculated , and compare the output of the first layer with the most recently cached result Skip connection, together as the last layer of the upsampling layer The semantic vector is integrated into the network calculation through the cross attention mechanism during the calculation, and the stable diffusion model is finally calculated at the current time step. Biased noise ; Step 232: Use the previous time step to the current time step The advance prediction value of the noise , the original prediction noise obtained in step 221 , the biased noise obtained in step 231 , calculate the Kalman filter hyperparameters for the current time step t and ; In step 232, the Kalman filter hyperparameters of the current time step t are and The calculation formulas are: Where N is the number of laparoscopic images that need to be generated for pre-calculation. The initial value is 0; The step 3 comprises: Step 31: Use the text encoder to convert the laparoscopic image specified prompt word "CholecT45" into a semantic vector and randomly initialize a three-dimensional Gaussian noise matrix , when initialized, , as the initial input of the noise prediction network; Step 32: Use the noise prediction network to get input Noise ; Step 33: Use Kalman filter to Make an estimate and get the best estimate , and use the sampler to get the input for the next time step ; Step 34: Repeat steps 32 and 33 for a total of T times, where the input of step 32 is The subscript of starts from T, and 1 is subtracted each time the process is repeated, that is, T, T-1, T-2, ..., 2, 1, until the predicted image latent space representation is finally obtained ; Step 35: Using the decoder, Generate final laparoscopic images; The step 32 comprises: The noise prediction network calculation is divided into two cases: the first case is to fully calculate and cache the time step of some outputs in the stable diffusion model noise prediction network, and the time step as the cache point In the stable diffusion model noise prediction network, the main branch is fully executed, that is, all downsampling layers , Middle Layer , and upsampling layers , and automatically cache the second layer of the upsampling layer The output is , and the last layer of the upsampling layer The output of the stable diffusion model noise prediction network is used as the current time step The second case is the time step that does not require the full execution of the network calculation, that is, the non-cached point. For the non-cached point, when the non-cached point is the time step When the stable diffusion model noise prediction network only calculates the first layer of the downsampling layer , and compare the output of the first layer with the most recently cached result Skip connection, together as the last layer of the upsampling layer The stable diffusion model is calculated based on the input of the current time step. The original prediction noise of ; the input in both cases is , both use the cross-attention mechanism to fuse semantic vectors into network calculations; The step 33 comprises: Use Previous Time Step The noise of the current time step of the forecast As the current time step The predicted value of Based on the previous time step The state variance , current time step Noise variance of the predicted values , get the current time step The state prediction variance : In the formula, Indicates the current time step The state prediction variance at time , Represents the previous time step The state variance at Indicates the current time step The noise variance of the predicted value when ; Calculate the Kalman gain at time step t , which is used to balance the prediction error and observation error: In the formula, Indicates the current time step The state prediction variance at time , Indicates the noise variance of the observation value at the current time step; The noise obtained in step 32 As the observation value, combined with the Kalman gain correction prediction state at time step t, the optimal estimate of the noise after filtering is obtained, that is, the filtering output result : In the formula, represents the Kalman gain at time step t, Represents the previous time step The noise of the current time step of the forecast, Indicates the current time step The observed value, that is, the output result of the noise prediction network; update the state value variance , to reflect the new uncertainty: In the formula, Indicates the current time step The state prediction variance at time , the optimal estimate of the filtered noise As a sampler at time step The actual gradient used is obtained after sampling by the PNDM sampler , and sent to the stable diffusion model noise prediction network for the next time step.
2. The method according to claim 1, characterized in that The step 1 comprises: In order to obtain n different time steps, arrive Take n linear distribution values in the range and then pass them through the power function After transformation, the distribution value after transformation is In the range, x represents the n linear distribution values obtained, T represents the total time step of the denoising process, center represents the center position of the sampling point, pow represents the nonlinear degree of the sampling point, and total_numbers represents the number of points to be sampled. The obtained distribution value is then converted to an integer and deduplicated. The deduplicated value is added to the center offset to obtain the final time step as the cache point. By gradually adjusting the pow value, ensure that the number of cache points generated is equal to n.
Citation Information
Patent Citations
Nerve intelligent auxiliary recognition system for laparoscopic colorectal cancer operation
CN115187596A
Assisting medical procedures with luminescence images processed in limited informative regions identified in corresponding auxiliary images
US20220409057A1