Image generation method and device, electronic equipment, storage medium and program product

Through the methods of feature extraction and noise gradient control, the problems of noise distortion and low efficiency in image generation are solved, and high-quality and efficient image generation is achieved.

CN120672604AActive Publication Date: 2025-09-19UBTECH ROBOTICS CORP LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510687414.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-19
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

In the image generation process of existing technologies, local features of the image are easily distorted or details are lost under the influence of global noise, resulting in unstable quality of the generated image and low iterative processing efficiency.

Method used

The noise gradient is obtained through feature extraction, the noise amplitude negatively correlated with the noise gradient is injected, denoising is performed, and the position information is fused to generate a high-quality image.

Benefits of technology

The quality and efficiency of image generation are improved, the original detail information is retained, the interference of the denoising process is reduced, and the accuracy and speed of image generation are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672604A_ABST
    Figure CN120672604A_ABST
Patent Text Reader

Abstract

The invention provides an image generation method and device, electronic equipment, a storage medium and a program product. The method comprises the steps of obtaining at least one first instance in a first image; performing feature extraction on the first instance to obtain a first instance feature; determining a noise gradient of the first instance based on the first instance feature; obtaining a noise amplitude which is negatively correlated with the noise gradient, and injecting noise which is positively correlated with the noise amplitude into the first instance to obtain a first noise image; performing denoising processing on the first noise image to obtain a first generated image; and performing fusion processing on the first generated image and the position information in the first instance to obtain a second image including the first instance. According to the invention, the quality and efficiency of image generation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to an image generation method, device, electronic device, storage medium, and program product. Background Art

[0002] With the continuous development of computer technology, users' demand for image generation has increased. Image generation technology can recover or create high-fidelity images from noisy, low-resolution or defective data based on text or images through algorithms or models, and is applied in many fields such as medical treatment, artistic creation, and industrial inspection.

[0003] In the image generation process of related technologies, images are uniformly processed, and random Gaussian noise is gradually applied to the image and then denoised to obtain images of different styles. During the generation process, local features of the image may be distorted or details may be lost under the influence of global noise, resulting in unstable quality of the generated image. The iterative processing of the entire image affects the efficiency of image generation. Summary of the Invention

[0004] The embodiments of the present application provide an image generation method, device, electronic device, storage medium, and program product, which can improve the quality and efficiency of image generation.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] The present invention provides an image generation method, which includes:

[0007] acquiring at least one first instance in a first image;

[0008] Performing feature extraction on the first instance to obtain first instance features;

[0009] determining a noise gradient of the first instance based on the first instance feature;

[0010] Obtaining a noise amplitude negatively correlated with the noise gradient, and injecting noise positively correlated with the noise amplitude into the first instance to obtain a first noise image;

[0011] performing denoising processing on the first noisy image to obtain a first generated image;

[0012] The first generated image is fused with the position information in the first instance to obtain a second image including the first instance.

[0013] An embodiment of the present application provides an image generating device, comprising:

[0014] an image segmentation module, configured to obtain at least one first instance in the first image;

[0015] An image generation module is configured to perform feature extraction on the first instance to obtain first instance features; determine a noise gradient of the first instance based on the first instance features; obtain a noise amplitude negatively correlated with the noise gradient, and inject noise positively correlated with the noise amplitude into the first instance to obtain a first noise image; perform denoising on the first noise image to obtain a first generated image; and fuse the first generated image with position information in the first instance to obtain a second image including the first instance.

[0016] An embodiment of the present application provides an electronic device, comprising:

[0017] a memory for storing computer executable instructions or computer programs;

[0018] The processor is used to implement the image generation method provided in the embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.

[0019] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the image generation method provided in the embodiment of the present application when executed by a processor.

[0020] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the image generation method provided in the embodiment of the present application is implemented.

[0021] The embodiments of the present application have the following beneficial effects:

[0022] At least one first instance in a first image is acquired and feature extraction is performed on the first instance to obtain first instance features. This instance acquisition ensures that subsequent processing focuses on each first instance independently. Based on the first instance features, the noise gradient of the first instance is determined and quantified. Different noise injection strategies are designed for different instances. The noise injection strategies are implemented by acquiring a noise amplitude that is negatively correlated with the noise gradient and injecting noise that is positively correlated with the noise amplitude into the first instance. The noise injected into the first instance is adjusted based on the noise amplitude of each first instance, improving the efficiency of noise injection and preserving the original detail information of the first instance in the first noisy image. Denoising is performed on the first noisy image to obtain a first generated image, reducing interference during the denoising process and improving the accuracy and speed of the first generated image. The first generated image is then fused with position information contained in the first instance to restore blurred first instance boundaries in the first generated image. This ensures that the position of the resulting second image is accurate and avoids misalignment, thereby improving the quality of the generated second image. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 Schematic diagram of an application mode of the image generation method provided in an embodiment of the present application;

[0024] Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application;

[0025] Figure 3A This is a schematic diagram of the first process of the image generation method provided in an embodiment of the present application;

[0026] Figure 3B 2 is a schematic diagram of a second flow chart of the image generation method provided in an embodiment of the present application;

[0027] Figure 3C 3 is a schematic diagram of a third flow chart of the image generation method provided in an embodiment of the present application;

[0028] Figure 3D 4 is a schematic diagram of a fourth flow chart of the image generation method provided in an embodiment of the present application;

[0029] Figure 4 Schematic diagram of the structure of the instance segmentation model provided in the embodiment of the present application;

[0030] Figure 5 This is a schematic diagram of the instance segmentation result provided by the embodiment of the present application;

[0031] Figure 6 It is a schematic diagram of the noise state change of the image;

[0032] Figure 7 It is a structural diagram of the denoising model provided in the embodiment of the present application.

[0033] It should be pointed out that the above-mentioned "first" and "second" are only used to distinguish different solutions, and do not represent the degree of distinction between the advantages and disadvantages of the solutions or the priority in the implementation process. DETAILED DESCRIPTION

[0034] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0035] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0036] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0037] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0038] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0039] In the embodiments of this application, the collection and processing of relevant data (for example, image data) should be strictly in accordance with the requirements of relevant laws and regulations when applied in practice, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.

[0040] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0041] 1) Image Generation: This refers to the technology of using algorithms or models to generate high-fidelity, high-resolution visual content (e.g., photos, videos, artwork) from noisy, low-resolution, defective data or abstract descriptions (e.g., text, images). By learning the distribution of real data, it gradually restores or creates new visual information. In the embodiments of this application, image generation is to generate higher quality and higher precision images based on the original image.

[0042] 2) Instance Segmentation: This involves using computer vision algorithms to separate multiple independent objects (instances) from their background in complex images and accurately label the boundaries and location of each instance. Convolutional neural networks (e.g., Mask R-CNN) extract multi-level image features, generate candidate region proposals, and perform bounding box regression and mask prediction to achieve refined pixel-level instance separation.

[0043] 3) Diffusion Model: A generative deep learning technique based on probability theory and stochastic processes. Diffusion Model uses an iterative process of incremental Gaussian noise injection and inverse denoising to map low-dimensional noise (e.g., random noise, low-resolution images) to high-dimensional real-world data distributions (e.g., natural images), generating visual content with semantic consistency and detail fidelity.

[0044] 4) Noising: This is the forward pass of the diffusion model. It involves gradually injecting Gaussian noise at the initial stage of image generation, transforming a clear image (or low-resolution image) into a pure noisy image. Its core goal is to control the noise intensity at each time step, gradually distorting the image and providing a reversible path for subsequent denoising.

[0045] 5) Denoising: Denoising is the reverse process of the diffusion model, gradually recovering structure from pure noise. A neural network is used to predict the noise distribution at each step and reversely restore image details.

[0046] With the continuous development of computer technology, users' demand for image generation has increased. Image generation technology can recover or create high-fidelity images from noisy, low-resolution or defective data based on text or images through algorithms or models. Related technologies process images uniformly. During the image generation process, all areas (whether complex instances or simple backgrounds) are subjected to the same noise processing. Local features of the image may be distorted or details may be lost under the influence of global noise, resulting in unstable quality of the generated image. The image generation process requires multiple steps, each of which requires processing the entire image and gradually denoising and restoring it in each iteration. The iterative processing of the entire image affects the efficiency of image generation. On high-resolution or large-scale datasets, image generation takes a long time and consumes a lot of computing resources.

[0047] Embodiments of the present application provide an image generation method, an image generation device, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the quality and efficiency of image generation.

[0048] The following describes exemplary applications of the electronic devices provided in the embodiments of the present application. The devices provided in the embodiments of the present application can be implemented as various types of terminals, such as laptops, tablet computers, desktop computers, set-top boxes, smartphones, smart speakers, smart watches, smart TVs, and in-vehicle terminals. They can also be implemented as servers. The following describes exemplary applications when the devices are implemented as terminals or servers.

[0049] See also Figure 1 , Figure 1 This is a schematic diagram of an application mode of the image generation method provided in an embodiment of the present application, for example, to support an image generation application. Figure 1 The server 200, network 300, terminal device 400 and database 500 are involved. The terminal device 400 is connected to the server 200 via the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0050] In some embodiments, the embodiments of the present application can be implemented collaboratively by a server and a terminal device. For example, the user may be a person skilled in the art, server 200 is a server for generating a model image, terminal device 400 is a terminal operated by the user, and the original image to be generated is stored in database 500. Terminal device 400 sends an image generation request to server 200. Server 200 receives the image generation request and, using the image generation method provided in the embodiments of the present application, uses the original image to be generated as the first image, optimizes the accuracy of the first image, and generates a second image as the optimized new image. The optimized new image is then sent to terminal device 400.

[0051] The image generation method provided in the embodiments of the present application can be applied to various scenarios requiring image optimization, such as medical imaging, industrial inspection, security, etc., as illustrated below with examples.

[0052] 1) In the field of medical imaging, for example, a terminal device receives an image generation request from medical personnel, and the server optimizes the original diagnostic image or case picture through an image generation method to generate a clearer and more precise image to assist medical personnel in making a medical diagnosis.

[0053] 2) In the field of industrial inspection, for example, when a terminal device receives an image generation request from a technician, the server optimizes the original mechanical device image through an image generation method to generate a clearer and more precise part image to assist the technician in quality inspection of device production.

[0054] 3) In the security field, for example, when a terminal device receives an image generation request from a worker, the server optimizes the original images taken in low light or bad weather conditions through image generation methods to generate clearer, high-precision facial or license plate images, assisting workers in traffic safety control.

[0055] See also Figure 2 , Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application, Figure 2 The server 200 shown includes: at least one processor 410, a memory 450 and at least one network interface 420. The various components in the server 200 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not described in detail. Figure 2 Various buses are labeled as bus system 440 .

[0056] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0057] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0058] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0059] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0060] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0061] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420 . Exemplary network interfaces 420 include Bluetooth, Wireless LAN (WiFi), and Universal Serial Bus (USB).

[0062] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2 An image generation device 455 stored in memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: an image segmentation module 4551, an image generation module 4552, and a model training module 4553. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.

[0063] In some embodiments, the terminal or server can implement the image generation method provided by the embodiment of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be commands, machine instructions or software instructions at the microprogram level. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (APPlication, APP); it can also be a small program that can be embedded in any APP, that is, a program that can be run only by downloading it to a browser environment. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.

[0064] The image generation method provided in the embodiment of the present application will be described in conjunction with the exemplary application and implementation of the electronic device provided in the embodiment of the present application.

[0065] The following describes the image generation method provided by the embodiments of the present application. As previously mentioned, the electronic device implementing the image generation method of the embodiments of the present application can be a terminal, a server, or a combination of the two. Therefore, the execution entity of each step will not be repeated below.

[0066] See also Figure 3A , Figure 3A This is a first flow chart of the image generation method provided in the embodiment of the present application, which will be combined with Figure 3A The steps shown are explained, Figure 3A The executive body is Figure 1Server 200 in.

[0067] In step 301 , at least one first instance in a first image is acquired.

[0068] In some embodiments, see Figure 3B , Figure 3B 2 is a schematic diagram of a second flow chart of the image generation method provided in an embodiment of the present application; Figure 3A Step 301 can be performed by Figure 3B Steps 3011 to 3015 are implemented as described below.

[0069] In step 3011, a global feature map of the first image is obtained.

[0070] In some embodiments, a pre-trained convolutional neural network utilizes network layers of different depths to extract features from the first image. A shallow network extracts local features of the first image, and a deep network extracts global semantics of the first image. Multi-level features of the first image are obtained through layer-by-layer convolution of the convolutional neural network. The pre-trained convolutional neural network can be a backbone network. The multi-level features of the first image are aligned at different scales of resolution to obtain aligned multi-scale features. Subsequently, the aligned multi-scale features are pooled to obtain a global feature map of the first image that includes semantic information and spatial information of the first image. The semantic information includes the category and shape of the object, and the spatial information includes the position and size of the object. This information provides a basis for subsequent target detection and instance segmentation.

[0071] In step 3012, at least one candidate region is identified from the global feature map.

[0072] In some embodiments, each pixel in the global feature map is regarded as a feature point, and the row coordinates and column coordinates (i.e., row number and column number) of the feature point represent the position in the matrix of the global feature map. In the global feature map, the first image is divided into local areas with a fixed step size, and each local area is associated with the feature point in the global feature map to determine the mapping relationship, the width and height of the first image are divided by the downsampling rate respectively, and the result of dividing by the downsampling rate is rounded down to obtain the width and height of the global feature map. Based on the mapping relationship, each feature point of the global feature map is converted to the coordinate system of the first image according to the downsampling rate, and the geometric center of the local area corresponding to the feature point converted to the first image is used as the reference point of the rectangular anchor frame. Combined with the preset reference scale and aspect ratio, multi-size candidate frames are generated, and the actual width and height of the candidate frame are calculated according to the aspect ratio through square root transformation to ensure that the area of ​​the rectangular anchor frame is consistent under different ratios. A rectangular anchor frame covering multiple scales and multiple forms is generated at the position of each feature point to form at least one candidate area that is strictly aligned with the spatial distribution of the global feature map.

[0073] For example, the first image size is 1280 pixels (width) × 720 pixels (height), the downsampling rate (step size) S = 32, the preset reference scale is 32 pixels, the aspect ratio is 1:1 (square) 1:2 (vertical rectangle) and 2:1 (horizontal rectangle), and three rectangular anchor frames corresponding to the aspect ratios are generated for each feature point. After dividing the width and height of the first image by the downsampling rate, the result of dividing by the downsampling rate is rounded down to obtain the width and height (i.e., size) of the global feature map. According to the size of the global feature map 40×22 It is determined that the global feature map contains 40×22=880 feature points. Taking the feature point (2, 3) as an example (located in the 3rd row and 4th column of the feature map), the feature point (2, 3) is mapped to an area with the upper left corner (32×2, 32×3)=(64, 96) and the lower right corner (32×3, 32×4)=(96, 128) in the coordinate system of the first image as diagonal vertices. The area covers 32×32 pixels.

[0074] Each feature point of the global feature map is converted to the coordinate system of the first image according to the downsampling rate. The horizontal coordinate is 32×(2+0.5)=80 and the vertical coordinate is 32×(3+0.5)=112. That is, the center coordinate of the rectangular anchor box is (80, 112) in the coordinate system of the first image. When the aspect ratio of the corresponding rectangular anchor box is 1:1, the width is Height is The horizontal direction of the rectangular anchor frame is [80-16, 80+16] (i.e. 64 to 96), and the vertical direction is [112-16, 112+16] (i.e. 96 to 128), which completely overlaps with the local area in the first image. When the aspect ratio of the corresponding rectangular anchor frame is 1:2, the width is Height is The horizontal direction of the rectangular anchor box is 68.69~91.31, and the vertical direction is 89.37~134.63. If rectangular anchor boxes corresponding to three aspect ratios are generated for each feature point, a total of 40×22×3=2640 rectangular anchor boxes are generated, and each rectangular anchor box corresponds to a candidate region.

[0075] In step 3013, the local feature map corresponding to the candidate region in the global feature map is determined.

[0076] In some embodiments, the different levels contained in the global feature map constitute a multi-scale feature pyramid. According to the width and height of the candidate region in the first image, the corresponding level of the candidate region in the multi-scale feature pyramid is calculated by a preset formula. By calculating the ratio of the area of ​​the candidate region to the preset reference area, the mapping level is dynamically adjusted in combination with a logarithmic function. The smaller candidate region is mapped to the high-resolution level, and the larger one is mapped to the low-resolution level. For example, if the area of ​​the candidate region is w×h, the mapping level k can be expressed as Here, k0 is the base level (e.g., layer 3), S0 is the base size (e.g., 224 pixels), and the calculated result is rounded down to an integer. The target level is determined based on the rounded-down result. Smaller candidate regions (e.g., 32×32) are mapped to higher-resolution levels (e.g., layer 2, downsampled to 16), while larger regions (e.g., 256×256) are mapped to lower-resolution levels (e.g., layer 5, downsampled to 128).

[0077] Through the candidate region alignment operation, the candidate region is aligned with the global feature map of the corresponding level. After alignment, the candidate region is projected onto the global feature map of the corresponding level. A bounding box is defined on the projected candidate region to define the range of the candidate region. The bounding box is evenly divided into a fixed number of grid cells, and the geometric center of each grid cell is determined. Based on the bilinear interpolation method, the weighted pixel values ​​of the four adjacent feature points on the global feature map of the geometric center of each grid cell are calculated. Bilinear interpolation is a commonly used interpolation method used to find the pixel values ​​corresponding to the coordinates of the candidate region bounding box in the feature map. After aligning the feature grid through bilinear interpolation, a local feature map of fixed size is obtained to ensure that candidate regions of different sizes can be processed uniformly. The local feature map is a standardized representation of the candidate region in the deep feature space.

[0078] For example, the area of ​​the candidate region is 64×128=8192, the calculated mapping level k=2, the corresponding downsampling rate is 16, the size of the level 2 projection is 4×6 (64 / 16=4, 128 / 16=6), and the floating-point coordinate area in the corresponding level 2 feature map is 4.0×8.0. The projection area 4.0×8.0 is evenly divided into target sizes (such as 7×7 grid cells), with the width of each grid cell being 4.0 / 7≈0.571 and the height of each grid cell being 8.0 / 7≈1.143. Taking the grid cell in the 3rd row and 4th column as an example, the horizontal coordinate of the geometric center of the grid cell is determined to be 0.571×(3-0.5)=1.428, and the vertical coordinate is determined to be 1.143×(4-0.5)=4.000. The adjacent feature points are (1, 4), (1, 5), (2, 4), and (2, 5). The weighted pixel values ​​of the four adjacent feature points in the global feature map of the geometric center of each grid cell are calculated based on the bilinear interpolation method. The interpolated weighted pixel values ​​are aggregated to obtain the interpolation result F=0.572·F (1,4) +0.428·F(F (2,4) , the continuous feature representation of the interpolated value F on the local feature map integrates the spatial context information of the four surrounding points, and the interpolation results of all 7×7 grid cells are arranged into a fixed matrix to form a standardized local feature map.

[0079] In step 3014, category prediction is performed based on the local feature map to obtain the confidence of the candidate area, and regression is performed based on the local feature map to obtain the position offset. The position of the candidate area is corrected based on the position offset to obtain the corrected candidate area.

[0080] In some embodiments, a classification branch in a pre-trained classification and regression model is used to predict the category of an instance of a candidate region based on a local feature map. The fully connected layer and activation function of the classification branch output a category probability distribution for each candidate region, and the highest value in the category probability distribution is used as the confidence score of the candidate region. A bounding box regression is performed on the candidate region based on the local feature map using the regression branch. The fully connected layer and linear activation function of the regression branch predict the bounding box offset to obtain an offset relative to the original candidate region. The coordinates of the original candidate region are corrected based on the offset obtained by regression to obtain a corrected candidate region.

[0081] For example: assuming that the category probability distribution of a local feature map is [cat: 0.7, dog: 0.2, background: 0.1], the highest value 0.7 in the category probability distribution is used as the confidence of the candidate region, corresponding to the category "cat"; set the original candidate region to (x=50, y=60, w=20, h=30), x and y are the center point coordinates, w represents the width of the candidate region, h represents the height of the candidate region, the offset of the regression prediction is Δx=0.5, Δy=-0.2, Δw=0.1, Δh=0.3, and adjust the position of the candidate region according to the offset. The new x is 60 (50+20*0.5), the new y is 54 (60+30*(-0.2)), and the new w is approximately 22.1 (20*e 0.1 ), the new h is about 40.5(30*e 0.3 ), the corrected candidate region is (x=60,y=54,w≈22.1,h≈40.5).

[0082] In step 3015, a first instance is determined based on the corrected candidate region and the confidence level.

[0083] In some embodiments, based on the revised candidate regions and the confidence of the revised candidate regions, the top N revised candidate regions with the highest category probability distribution are retained during preliminary filtering, where N is a preset threshold for the number of candidate regions to exclude low-probability candidate regions, and non-maximum suppression (NMS) is performed on the remaining candidate regions after preliminary filtering, and overlapping revised candidate regions are eliminated based on a preset intersection-over-union threshold. Non-maximum suppression is an algorithm in computer vision that is used to remove overlapping candidate regions in target detection tasks and retain the best detection results. Multiple revised candidate regions may overlap on the same target, and non-maximum suppression selects the best candidate region from the overlapping revised candidate regions as the final revised candidate region. The intersection-over-union (IoU) between any two corrected candidate regions is calculated. If the IoU exceeds a preset IoU threshold (e.g., 0.5), the two corrected candidate regions are determined to overlap on the same target. Only the corrected candidate region with higher confidence is retained to achieve redundant detection of the target. The above process is iterated until all corrected candidate regions are determined. The first instance is determined based on the retained corrected candidate regions and their corresponding confidence levels.

[0084] In some embodiments, see Figure 4 , Figure 4It is a structural diagram of the instance segmentation model provided in an embodiment of the present application; instance segmentation of the first image is implemented by the instance segmentation model, the first image 401 is input into the instance segmentation model, the candidate region in the first image is identified by the region detection module 402, the category of the candidate region is predicted and the coordinates of the candidate region are adjusted by the convolution layer 403 including multiple convolutions, and the category probability distribution and the corrected candidate region are obtained, the highest value in the category probability distribution is used as the classification result 4031, that is, the confidence of the corrected candidate region, the first instance 404 is determined from the first image according to the corrected candidate region and the confidence, and the instance segmentation process performed on the first image 401 is completed, each first instance 404 is clearly marked in the first image 401, and the first image 401 will be divided into multiple non-overlapping first instances 404.

[0085] Through the embodiments of the present application, a global feature map that integrates semantic and spatial information is generated to provide rich contextual support for instance positioning, and candidate regions covering targets of different scales are generated. The feature grid is aligned through bilinear interpolation, and the candidate regions are mapped to local feature maps of fixed size to ensure standardized expression of targets of different sizes in the feature space. Based on the local feature map, category prediction and bounding box regression are performed simultaneously to determine the confidence of the target candidate region and perform position correction. The candidate box with high confidence and precise positioning is retained as the first instance, effectively balancing the detection recall rate and accuracy.

[0086] In some embodiments, see Figure 3C , Figure 3C 3 is a schematic diagram of a third flow chart of the image generation method provided in an embodiment of the present application; Figure 3B Step 3015 can be performed by Figure 3C Steps 30151 to 30156 are implemented as described below.

[0087] In step 30151, the corrected candidate regions whose confidence levels are lower than a preset confidence threshold are removed, and the remaining corrected candidate regions after the removal are divided into multiple groups, wherein the confidence levels of the corrected candidate regions in a group are within the confidence interval corresponding to the group.

[0088] In some embodiments, candidate areas in the revised candidate areas with confidence levels lower than the confidence threshold are filtered out according to a preset confidence threshold, thereby eliminating interference from false detection or fuzzy predictions, and the remaining revised candidate areas after removal are divided into multiple groups according to preset confidence intervals, and the confidence levels of the revised candidate areas in a group are the same or similar.

[0089] For example, there are 10 corrected candidate regions with confidences of 0.95, 0.92, 0.88, 0.85, 0.80, 0.75, 0.70, 0.65, 0.60, and 0.55, respectively. The preset confidence threshold is 0.7. The confidences of all corrected candidate regions are traversed, and the corrected candidate regions with confidences lower than 0.7 are removed. After removal, the remaining corrected candidate regions and their confidences are 0.95, 0.92, 0.88, 0.85, 0.80, 0.75, and 0.70. The preset confidence intervals are interval 1: [0.90, 1.00], interval 2: [0.80, 0.89], interval 3: [0.70, 0.79], and interval 4: [0.65, 0.69]. The remaining corrected candidate regions are divided into corresponding interval groups according to their confidence levels. The group corresponding to interval 1 includes corrected candidate regions corresponding to confidence levels of 0.95 and 0.92, and the group corresponding to interval 2 includes corrected candidate regions corresponding to confidence levels of 0.88 and 0.85.

[0090] In step 30152, the following processing is performed for each group to determine the intersection-over-union (IoU) of any two revised candidate regions in the group. When the IoU exceeds the preset IoU threshold, the revised candidate region with higher confidence is retained. The following processing is iteratively performed: determining the IoU between the retained revised candidate region and any remaining revised candidate region in the group. When the IoU exceeds the preset IoU threshold, the revised candidate region with higher confidence is retained. When the iteration ends, the revised candidate region finally retained is used as the target candidate region in the group.

[0091] In some embodiments, the revised candidate regions are sorted from high to low by confidence, and the current processing index is initialized to the first position in the group. The candidate region corresponding to the current index is then selected as the base candidate region. All subsequent revised candidate regions are traversed, and the intersection over union (IoU) between the current base candidate region and the uncompared revised candidate regions is calculated. The IoU is a metric that measures the degree of overlap between two regions. When the IoU exceeds a preset IoU threshold, the two revised candidate regions are determined to be highly overlapping and pointing to the same target, and the revised candidate region with the higher confidence is retained. After completing one round of traversal, the current index is shifted back one position, and the above iterative process is repeated to determine the IoU between the retained revised candidate region and any remaining revised candidate region in the group. When the IoU exceeds the preset IoU threshold, the revised candidate region with the higher confidence is retained. This process continues until all revised candidate regions have been processed and the iteration ends. By successively retaining the candidate regions with the higher confidence, it is ensured that only non-overlapping revised candidate regions with decreasing confidence are retained in the group, and the target candidate region in the candidate region group is ultimately determined.

[0092] For example, setting the intersection-and-union threshold to 0.5, the four corrected candidate regions of the same target category (such as "cat") are arranged in descending order of confidence: Box1 (confidence 0.9, coordinates [10, 20, 50, 60]), Box2 (confidence 0.8, coordinates [12, 22, 52, 62]), Box3 (confidence 0.7, coordinates [15, 25, 55, 65]), Box4 (confidence 0.6, coordinates [100, 150, 200, 200]), the current processing index points to Box1, the first round of traversal benchmark candidate region is Box1, the intersection-and-union ratio of Box1 and Box2 is 0.85, which is greater than the intersection-and-union threshold, and the confidence of Box1 is higher than that of Box2, so Box2 is removed and Box1 is retained. The intersection-and-union ratio of Box1 and Box3 is 0.72, which is greater than the intersection-and-union threshold, and the confidence of Box1 is higher than that of Box3, so Box3 is removed and Box1 is retained. The intersection-over-union ratio of Box1 and Box4 is 0.01, which is less than the intersection-over-union ratio threshold. Box4 is retained and Box1 and Box4 are used as target candidate regions in the candidate region grouping.

[0093] In step 30153, the maximum value of the confidence of the target candidate region is used as the confidence of the first instance.

[0094] In some embodiments, all confidence levels in the target candidate region are traversed, and the maximum confidence level is used as the confidence level of the first instance. For example, if the confidence levels of Box 1 and Box 4 are traversed, and the confidence level of Box 1 is the largest, the confidence level of Box 1, 0.9, is used as the confidence level of the first instance.

[0095] In step 30154, a first mask of the target candidate region is obtained, the first mask is upsampled, and the upsampled first mask is binarized to obtain a second mask.

[0096] In some embodiments, an instance segmentation model is used to obtain a first mask corresponding to the target candidate region. The first mask is a low-resolution probability map with a lower resolution than the first image. The shape and size of the first mask correspond to the target candidate region and are used to represent the shape and location of the instances included in the candidate region. The pixel values ​​of the first mask are typically between 0 and 1, indicating the probability of each pixel belonging to the instance. The first mask is then upsampled. Upsampling is a technique for magnifying a low-resolution image to a high-resolution image. Each pixel in the first mask is mapped to a corresponding position in a high-resolution grid. Interpolation weights are calculated based on the coordinate distances of the four surrounding pixels. The weighted summation yields the probability value for each point in the high-resolution grid, ensuring that the size of the first mask is consistent with the spatial resolution of the original input image. The upsampled first mask is then binarized. Binarization renders the entire image visually distinct in black and white. A fixed pixel value threshold is selected, and pixels in the upsampled mask with values ​​greater than or equal to the pixel value threshold are set to 1, while pixels with values ​​less than the pixel value threshold are set to 0. This results in a second mask with pixel values ​​of only 0 and 1, indicating whether each pixel belongs to the target candidate region.

[0097] For example, the first mask of the target candidate area is 28×28, and the probability value range is [0, 1]. After bilinear interpolation, a 224×224 high-resolution probability map is generated. The probability of the edge area is gradually smoothed. The pixel value threshold is set to 0.5. The pixel values ​​in the first mask with a probability value greater than or equal to the set pixel threshold are set to 1, and the pixel values ​​in the first mask with a probability value less than the set pixel threshold are set to 0. The second mask is 224×224, which is a matrix containing 0 and 1.

[0098] In step 30155, the second mask is mapped to the image coordinate system of the first image according to the target candidate region to obtain a third mask.

[0099] In some embodiments, the bounding box coordinates of the target candidate region are extracted, the spatial coverage of the bounding box coordinates in the first image is determined, each pixel of the second mask is linearly mapped to the global coordinate system of the first image, and the second mask is scaled and adjusted according to the size of the target candidate region to ensure that the second mask and the target candidate region are spatially matched. All pixels of the second mask are traversed, and when the pixel value of a pixel is 1, the coordinate position corresponding to the pixel position in the first image is marked. After traversing all pixels, a third mask that is spatially aligned with the first image is obtained. The third mask is used to represent the precise position of the target candidate region in the first image.

[0100] In step 30156, a first instance is obtained based on the target candidate region, the confidence of the first instance, and the third mask.

[0101] In some embodiments, coordinate information of a bounding box is extracted from the target candidate area as the spatial positioning of the first instance, and the third mask is used as the pixel segmentation attribute of the first instance. The spatial positioning of the first instance, the pixel segmentation attribute of the first instance, and the confidence of the first instance are bound through a data structure to form a complete first instance. The bounding box and the third mask are associated through coordinates, and the confidence is independently stored as a scalar value.

[0102] In some embodiments, see Figure 5 , Figure 5 This is a schematic diagram of the instance segmentation result provided by the embodiment of the present application; Figure 5 shows the result of instance segmentation of the first image, including multiple first instances obtained by segmentation. The first instances do not overlap, and the categories of the first instances can be people, animals, objects, or background. Each first instance includes a bounding box corresponding to the corrected candidate region and a confidence level. The spatial positioning of the bounding box of the corrected candidate region corresponding to the first instance 5011 reflects the coverage of the target candidate region. The identified instance region included in the bounding box reflects the pixel segmentation properties of the first instance, that is, the third mask of the first instance. If the information displayed for the first instance 5011 is (person, 1.0), it means that the category of the instance in the bounding box of the first instance 5011 is person, and the confidence level is 1.0. If the information displayed for the first instance 5012 is (bottle, 0.99), it means that the category of the instance in the bounding box of the first instance 5012 is bottle, and the confidence level is 0.99. If the information displayed for the first instance 5021 is (umbrella, 0.98), it means that the class of the instance in the bounding box of the first instance 5021 is an umbrella, and the confidence level is 0.98. If the information displayed for the first instance 5022 is (backpack, 1.0), it means that the class of the instance in the bounding box of the first instance 5022 is a backpack, and the confidence level is 1.0. If the information displayed for the first instance 5031 is (motorcycle, 1.0), it means that the class of the instance in the bounding box of the first instance 5031 is a motorcycle, and the confidence level is 1.0. If the information displayed for the first instance 5032 is (motorcycle, 1.0), it means that the class of the instance in the bounding box of the first instance 5032 is a motorcycle, and the confidence level is 1.0. If the information displayed by the first instance 5041 is (elephant, 1.0), it means that the category of the instance in the bounding box of the first instance 5041 is elephant, and the confidence is 1.0; if the information displayed by the first instance 5042 is (person, 1.0), it means that the category of the instance in the bounding box of the first instance 5042 is person, and the confidence is 1.0.

[0103] Through the embodiments of the present application, low-quality candidate areas are removed through confidence threshold filtering and grouping mechanism, false detection interference is reduced, redundant boxes are suppressed iteratively based on intersection-over-union, ensuring that only the optimal candidate area is retained for the same target, and positioning ambiguity caused by overlapping detection is resolved. Through the maximum confidence distribution strategy, the model's highest probability prediction of target existence is solidified as instance attributes, enhancing the credibility of the results. The first mask is upsampled and converted into a high-precision binary mask to achieve pixel-level segmentation detail restoration. Through coordinate mapping and data structure binding, the target candidate area, confidence and mask are integrated into a spatially aligned, attribute-complete first instance, providing a basis for subsequent image generation.

[0104] Continue to see Figure 3A In step 302, feature extraction is performed on the first instance to obtain first instance features.

[0105] In some embodiments, feature extraction is performed on the first instance, and the first instance feature includes feature information such as texture, edge, shape, etc. of the first instance. Feature extraction can be achieved through a convolutional neural network or other feature extraction models, which is not limited in this embodiment of the present application.

[0106] In step 303 , a noise gradient of the first instance is determined based on the first instance feature.

[0107] In some embodiments, step 303 can be implemented by the following method: determining the transverse gradient and the longitudinal gradient of the first instance feature; determining the partial derivative of the transverse gradient and the partial derivative of the longitudinal gradient, and determining the sum of the squares of the partial derivative of the transverse gradient and the partial derivative of the longitudinal gradient; determining the square root of the sum of the squares as the noise gradient of the first instance.

[0108] In some embodiments, the gradient of the first instance feature map is calculated by an automatic derivation mechanism, the corresponding derivation function is called, and the noise gradient of the first instance feature map is calculated according to the chain rule. The chain rule is a method for backpropagating gradients in neural networks, which is achieved by accumulating local gradients layer by layer, and the gradient is backpropagated from the loss function to the input. The calculated noise gradient is a tensor with the same dimension as the first instance feature map, each element of which represents the rate of change of the first instance feature map with respect to the input at the corresponding position. A convolution operation is performed on the first instance feature to extract the gradient information in the horizontal and vertical directions. The horizontal gradient and the vertical gradient of the first instance feature map are calculated based on the gradient information, and the partial derivative of the horizontal gradient and the partial derivative of the vertical gradient are determined. The square sum of the two partial derivatives is calculated, and the square root of the square sum is used as the noise gradient of the first instance feature. This can be achieved by formula (1), which is described in detail below.

[0109]

[0110] in, is the noise gradient of the first instance feature, is the first instance feature, is the partial derivative of the transverse gradient of the first instance feature in the horizontal direction, is the partial derivative of the longitudinal gradient of the first instance feature in the vertical direction, It is the square root of the sum of the squares of the partial derivatives of the transverse and longitudinal gradients.

[0111] Through the embodiments of this application, the horizontal and vertical gradients of the feature map are calculated to quantify the horizontal and vertical changes of the feature, reflecting the spatial rate of change of the feature, helping to identify important information such as the edge, texture, and shape of the feature. By determining the degree of change of the first instance feature quantification by the noise gradient, it provides an important basis for subsequent noise injection and denoising.

[0112] In step 304, a noise amplitude negatively correlated with the noise gradient is obtained, and noise positively correlated with the noise amplitude is injected into the first instance to obtain a first noise image.

[0113] In some embodiments, the noise amplitude in step 304 can be obtained by the following method: determining the average value of the noise gradient and the variance of the noise gradient; determining the sum of the average value and the variance, multiplying the inverse of the sum by a preset adjustment parameter to obtain a noise amplitude that is negatively correlated with the noise gradient.

[0114] In some embodiments, the average value of the noise gradient and the variance of the noise gradient are calculated, and the average value and the variance of the noise gradient are added to obtain the sum of the average value and the variance to determine the complexity corresponding to the noise gradient. The inverse of the sum is multiplied by a preset adjustment parameter to obtain a noise amplitude that is negatively correlated with the noise gradient. This can be achieved by formula (2), which is described in detail below.

[0115]

[0116] Here, γ is a preset adjustment parameter used to adjust the overall scale of the noise amplitude. The preset adjustment parameter is selected based on balancing the differences in noise amplitude between different instances. is the first instance feature, β i is the noise amplitude, which is a parameter that controls the intensity of noise injection. It is the sum of the mean and variance of the noise gradient. A lower noise amplitude means that during the noise injection process, relatively less noise is added to the first instance, and the image is less disturbed during the denoising process, which is more conducive to preserving the original detail information of the instance.

[0117] For example, if the noise amplitude for facial features has large mean and variance values, the face is considered complex. When processing complex instances, a lower noise amplitude is used so that the facial texture and facial details remain clearly discernible after denoising. A higher noise amplitude indicates more added noise, requiring a greater degree of image restoration and reconstruction during denoising. When both the mean and variance of the noise amplitude are small, the face is considered simple. For instances with simple textures (e.g., backgrounds or monochrome areas), a higher noise amplitude can speed up generation and increase the diversity of the generated images.

[0118] Through the embodiments of the present application, the mean and variance of the noise gradient are calculated to reflect the overall level of the noise gradient and the degree of variation in the noise gradient, which together describe the complexity of the first instance feature. The noise amplitude, which is negatively correlated with the noise gradient, is obtained, and the intensity of the noise injection is adaptively adjusted based on the complexity of the first instance feature. This helps preserve the detailed information of the complex instance during the noise injection process, speeds up the generation of simple instances, and increases the diversity of the generated images, thereby improving the quality and efficiency of image generation overall.

[0119] In some embodiments, the first noise image in step 304 can be obtained by the following method: iterate t and perform the following processing: according to the noise ratio of the tth time step determined by the scheduling algorithm, inject noise that is positively correlated with the noise amplitude into the noise image of the t-1th time step to obtain the noise image of the tth time step, t is a positive integer variable that increases starting from 1, the value of t satisfies 2≤t≤T, T is a preset time step threshold, and the noise image of the 1st time step is the first instance; when the value of t is iterated to T, the noise image of the Tth time step is used as the first noise image.

[0120] In some embodiments, a time step is the smallest unit after dividing continuous time into discrete intervals, and each time step corresponds to a state update. According to the noise ratio of the t-th time step determined by the scheduling algorithm, the scheduling algorithm is a strategy for controlling the noise injection process, and the noise ratio is dynamically adjusted according to the time step t. The scheduling algorithm can be linear scheduling, cosine scheduling and quadratic scheduling, which are not limited by the embodiment of the present application. Noise that is positively correlated with the noise amplitude is injected into the noise image of the t-1th time step to obtain the noise image of the t-1th time step. The noise injection is performed on each pixel of the image by adding the noise value to the original pixel value. Iterative execution time step t, t is a positive integer variable that increases from 1, the value of t satisfies 2≤t≤T, T is a preset time step threshold, the noise image of the 1st time step is the first instance, when the value of t is iterated to T, the noise image of the Tth time step is used as the first noise image, and the noise image of the tth time step can be implemented by formula (3), which is described in detail below.

[0121]

[0122] in, is the first instance x i The noise image at time step t, is the noise proportion in the tth time step that is positively correlated with the noise amplitude, is the first instance x i The noise image at the t-1th time step, that is, the noise image of the previous state at the tth time step, ∈ i Represents independently sampled Gaussian noise. When the time step t reaches the preset time step threshold T, the noise image at the tth time step is used as the first noise image.

[0123] Through the embodiments of the present application, the noise ratio is dynamically adjusted according to the scheduling algorithm at each time step t, and noise positively correlated with the noise amplitude is injected into the noise image of the previous time step. The noise injection process is precisely controlled, and the noise intensity is adaptively adjusted according to the feature complexity of the instance, making the entire noise injection process more efficient and accurate.

[0124] In some embodiments, the first noise image, the first generated image, and the second image are implemented by calling the first model, see Figure 3D , Figure 3D is a fourth flow chart of the image generation method provided in the embodiment of the present application; Figure 3A Before step 304, the Figure 3D Steps 3041 to 3044 in the training of the first model are described in detail below.

[0125] In step 3041, at least one second instance of the image sample is obtained.

[0126] In some embodiments, instance segmentation is performed on the image sample to obtain at least one second instance. The principle of instance segmentation is the same as that in step 301 above and will not be repeated here.

[0127] In step 3042 , a second generated image including a second instance is acquired through the first model.

[0128] In some embodiments, feature extraction is performed on the second instance using the first model to obtain second instance features. A noise gradient of the second instance is determined based on the second instance features, and a noise amplitude negatively correlated with the noise gradient is obtained. Noise is injected into the second instance based on the noise amplitude to obtain a noise image corresponding to the second instance. The noise image corresponding to the second instance is denoised to obtain a second generated image corresponding to the second instance. The principles for determining the noise amplitude and obtaining the noise image are the same as those of step 304 above. The denoising process for the second instance is the same as that of step 305 above and will not be further described here.

[0129] In step 3043, the second generated image is fused with the position information in the second instance to obtain a third image.

[0130] In some embodiments, the second instance includes position information, and the second generated image is fused with the position information in the second instance to obtain a third image. The principle of obtaining the third image is the same as the principle of obtaining the second image in the above step 306, and will not be repeated here.

[0131] In step 3044, a first loss is determined based on the image sample and the third image, and a second loss is determined based on the second instance and the second generated image, and the first model is optimized using the first loss and the second loss.

[0132] In some embodiments, a first loss is determined based on the image sample and the third image, and the first loss is a global loss between image features. A second loss is determined based on the second instance and the second generated image, and the second loss is a pixel-level loss for the instance. The first model is optimized by backpropagation through the first loss and the second loss, and the network parameters in the first model are updated. The above optimization process is iterated, and the optimization of the first model is stopped when the iteration meets the preset number of iterations to obtain the optimized first model. The optimized first model is used to generate high-quality and high-precision images based on the input image. The global loss and pixel-level loss can achieve diversity optimization of the global features and local features of the first model.

[0133] Through the embodiments of the present application, the global image features are optimized by the first loss, the accuracy of the local instances is optimized by the second loss, the first model is trained hierarchically in both global and local aspects, the overall features and local features are balanced, the authenticity and accuracy of the images generated by the first model are improved, and the generalization ability and noise resistance of the first model are improved.

[0134] In some embodiments, the first loss in step 3044 can be achieved by the following method: obtaining image sample features and third image features; determining a first difference between the image sample features and the third image features; and taking the mean of the sum of squares of the first differences as the first loss.

[0135] In some embodiments, feature extraction is performed on the image sample to obtain image sample features, and feature extraction is performed on the third image to obtain third image features, a first difference between the image sample features and the third image features is determined, and the mean of the sum of squares of the first differences is used as the first loss. This can be achieved through formula (4), which is described in detail below.

[0136]

[0137] in, It’s the first loss. is the third image, x is the real image sample, φ is the extracted image feature, φ(x) is the image sample feature, is the third image feature, It is to calculate the mean square error between the third image feature and the sample image feature.

[0138] In some embodiments, the second loss in step 3044 can be implemented by: determining a second difference between pixels at the same position in the second instance and the second generated image; calculating the sum of the squares of the second differences for all pixels, and taking the mean of the sum of the squares of the second differences as the second loss.

[0139] In some embodiments, determining a second difference between pixels at the same position in the second instance and the second generated image, where the second difference is a pixel difference, the second instance and the second generated image contain multiple instances, calculating the sum of the squares of the second differences of all pixels, and taking the mean of the sum of the squares of the second differences as the second loss can be achieved through formula (5), which is described in detail below.

[0140]

[0141] in, The second loss, is the i-th instance in the second generated image, x i is the real image of the i-th second instance, It is the sum of the mean squared errors between all instances of the second generated image and the true image of the second instance.

[0142] Through the embodiments of the present application, combined with global feature alignment and instance-level pixel consistency optimization, the balance between the overall structural fidelity and the restoration of local details of the generated image is improved. The first loss is based on the mean square error constraint of the feature space to constrain the global semantic consistency, ensuring that the generated image is highly matched with the real sample in high-level feature distribution, avoiding overall structural distortion or semantic shift. The second loss aligns the generated instance (generated image) with the local details of the real instance through pixel-level constraints, strengthening the refined reconstruction of key areas.

[0143] In step 305, denoising is performed on the first noisy image to obtain a first generated image.

[0144] In some embodiments, the first noise image is denoised by a trained denoising network. The denoising network can be a symmetrical U-shaped convolutional neural network (U-Net), which is not limited in the embodiments of the present application. The first noise image is input into the denoising network. The structure of the denoising network includes an encoder-decoder architecture. The encoder is responsible for extracting the features of the first noise image, and the decoder gradually reconstructs a clear first generated image based on the features of the first noise image. The encoder reduces the resolution of the first noise image and extracts the features of the first noise image through convolutional layers and downsampling operations. At each layer of the decoder, a connection is made with the feature map of the corresponding layer of the encoder, and the detail information extracted in the encoder is reintroduced into the decoder to help reconstruct the image more accurately. At the output layer of the decoder, a denoised image is output. The denoising process is a step-by-step refinement process that requires multiple iterations. In each iteration, a denoised image is output based on the current noise image. The denoising intensity corresponds to the noise amplitude of the added noise. When the number of iterations reaches a preset threshold or the loss function converges, the denoising process ends and the first generated image is obtained.

[0145] In some embodiments, see Figure 6 , Figure 6 is a schematic diagram of the noise state change of the image; the first noise image 601 is a pure noise image obtained after the noise is iteratively injected for T time steps, X T The first noise image 601 can be obtained by the above step 304. The first noise image 601 is subjected to the reverse process based on the Markov chain to generate the first generated image 603. t-1 represents the noise image 602 of the corresponding intermediate state at the state of the t-1th time step, and X0 represents the first generated image 603 finally generated. θ (x t-1 |x t ) is the probability distribution of the reverse process, which represents the probability of generating the image state at time step t-1 from the image state at time step t, q(x t |x t-1 ) is the probability distribution of the forward process, which represents the probability of injecting noise from the image state at time step t-1 to generate the image state at time step t, corresponding to the noise injection process in step 304. The reverse process starts from the initial first noisy image 601 and gradually removes the noise to finally generate the first generated image 603, which is composed of a predefined noise distribution p θ The noise distribution of the denoised noise corresponds to the noise distribution of the injected noise, and is consistent with the noise injected in the above step 304.

[0146] In some embodiments, see Figure 7 , Figure 7This is a schematic diagram of the structure of the denoising model provided in an embodiment of the present application. A first noisy image 701 is input into the denoising network. The encoder 702 undergoes a series of convolutional layers and downsampling operations, gradually reducing the resolution of the first noisy image 701 and extracting features of the first noisy image 701, such as texture, edges, and shape. At the innermost layer of the encoder 702, the resulting feature map contains the most abstract representation of the first noisy image 701. Each layer of the decoder 703 is connected to the feature map of the corresponding layer in the encoder 702, reintroducing the detailed information extracted in the encoder 702 into the decoder 703 to help reconstruct the image more accurately. At the output layer of the decoder 703, a denoised image is output. The denoising process is a step-by-step refinement process that requires multiple iterations. In each iteration, a denoised image is output based on the current noisy image. The denoising intensity corresponds to the noise amplitude of the added noise. When the number of iterations reaches a preset threshold or the loss function converges, the denoising process ends, and the resulting first generated image 704 is the final result of denoising the first noisy image 701.

[0147] In step 306 , the first generated image is fused with the position information in the first instance to obtain a second image including the first instance.

[0148] In some embodiments, the position information of the first instance is obtained, and the position information can determine an area of ​​the same size as the input image, the pixel value of the area where the first instance is located is 1, and the pixel value of the background area is 0, that is, the mask corresponding to the first instance is obtained. The first generated image is adjusted to the same size as the mask of the first instance by an interpolation method to ensure that the first generated image and the mask are spatially aligned. The resized first generated image is placed pixel by pixel at the position indicated by the mask in the second image, and each pixel value in the mask is traversed, and the pixel value in the first generated image is copied to the same position in the second image. The transition between the first generated image and the surrounding background or the adjacent first instance is achieved by a smoothing algorithm. The smoothing algorithm is a type of mathematical method used to eliminate local mutations in images and achieve natural transitions. It can make the fusion boundary visually continuous and natural while retaining the details of the generated area. By using image fusion technology, while maintaining the details of the first generated image, the boundary area is made consistent with the surrounding environment in terms of lighting, color, etc. Image fusion technology is the process of merging multiple images (or image areas) from the same scene or different sources into a high-quality, more information-rich image. Through direct overlay, weighted fusion, Poisson fusion or a combination of the above, the first generated image is fused according to the mask to obtain a second image that is optimized and restored in the first instance area.

[0149] The image production method provided by the embodiments of the present application has the following beneficial effects:

[0150] The algorithm extracts a global feature map that combines semantic and spatial information. It then combines a region proposal network with multi-scale anchors to generate candidate regions covering different objects. Local feature maps are aligned via bilinear interpolation to achieve standardized representation. The candidate regions are then accurately positioned based on category prediction and bounding box regression. Non-maximum suppression is then used to select high-confidence first instances, covering objects of different scales and reducing redundant candidate regions. This improves candidate recall and localization efficiency, ultimately enhancing candidate localization accuracy. The algorithm calculates noise amplitude based on feature gradients and dynamically adjusts the noise injection intensity. Low noise is used to preserve detail in complex instances, while high noise is used to accelerate generation in simple regions, reducing unnecessary computation and balancing detail accuracy with processing efficiency. The denoising network gradually reconstructs the image, ensuring that the denoising process restores clear structure while preserving edges and texture. A first loss based on global features ensures high-level semantic consistency between the generated image and the sample. A second loss based on instance-level pixels enforces local detail alignment, simultaneously optimizing overall structure and local accuracy, improving the visual realism of the generated image. The generated image is fused with the original image using a position mask. A smoothing algorithm is used to smooth the transition between the instance and the surrounding area, resulting in a high-quality optimized image that ensures consistent style and structure across different instances.

[0151] The following continues to describe the exemplary structure of the image generating device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the image generation device 455 of the memory 450 may include: an image segmentation module 4551, used to obtain at least one first instance in the first image; an image generation module 4552, used to perform feature extraction on the first instance to obtain first instance features; determine the noise gradient of the first instance based on the first instance features; obtain the noise amplitude negatively correlated with the noise gradient, inject noise positively correlated with the noise amplitude into the first instance to obtain a first noise image; perform denoising on the first noise image to obtain a first generated image; and perform fusion processing on the first generated image and the position information in the first instance to obtain a second image including the first instance.

[0152] In some embodiments, the image segmentation module 4551 is also used to obtain a global feature map of the first image; identify at least one candidate region from the global feature map; determine the local feature map corresponding to the candidate region in the global feature map; perform category prediction based on the local feature map to obtain the confidence of the candidate region, and perform regression based on the local feature map to obtain a position offset, correct the position of the candidate region based on the position offset to obtain a corrected candidate region; and determine the first instance based on the corrected candidate region and the confidence.

[0153] In some embodiments, the image segmentation module 4551 is also used to remove the corrected candidate areas that are below a preset confidence threshold, and divide the remaining corrected candidate areas after the removal into multiple groups, and the confidence of the corrected candidate areas in a group is in the confidence interval corresponding to the group; for each group, determine the target candidate area in the group; take the maximum value of the confidence in the target candidate area as the confidence of the first instance; obtain a first mask of the target candidate area, upsample the first mask, and binarize the upsampled first mask to obtain a second mask; map the second mask to the image coordinate system of the first image according to the target candidate area to obtain a third mask; obtain the first instance according to the target candidate area, the confidence of the first instance and the third mask.

[0154] In some embodiments, the image segmentation module 4551 is also used to perform the following processing for each group: determine the intersection-over-union ratio of any two revised candidate regions in the group; when the intersection-over-union ratio exceeds a preset intersection-over-union ratio threshold, retain the revised candidate region with higher confidence; iteratively perform the following processing: determine the intersection-over-union ratio of the retained revised candidate region with any remaining revised candidate region in the group; when the intersection-over-union ratio exceeds a preset intersection-over-union ratio threshold, retain the revised candidate region with higher confidence; and when the iteration ends, use the final retained revised candidate region as the target candidate region in the group.

[0155] In some embodiments, the image generation module 4552 is further used to determine the average value of the noise gradient and the variance of the noise gradient; determine the sum of the average value and the variance, multiply the inverse of the sum by a preset adjustment parameter to obtain a noise amplitude that is negatively correlated with the noise gradient.

[0156] In some embodiments, the image generation module 4552 is also used to iterate t to perform the following processing: according to the noise ratio of the tth time step determined by the scheduling algorithm, the noise positively correlated with the noise amplitude is injected into the noise image of the t-1th time step to obtain the noise image of the tth time step, t is a positive integer variable increasing from 1, the value of t satisfies 2≤t≤T, T is a preset time step threshold, and the noise image of the 1st time step is the first instance; when the value of t is iterated to T, the noise image of the Tth time step is used as the first noise image.

[0157] In some embodiments, the image generation module 4552 is also used to determine the transverse gradient and the longitudinal gradient of the first instance feature; determine the partial derivative of the transverse gradient and the partial derivative of the longitudinal gradient, determine the sum of the squares of the partial derivative of the transverse gradient and the partial derivative of the longitudinal gradient; and determine the square root of the sum of the squares as the noise gradient of the first instance.

[0158] In some embodiments, the first noise image, the first generated image and the second image are realized by calling the first model. The model training module 4553 is also used to obtain at least one second instance in the image sample; obtain a second generated image including the second instance through the first model; fuse the second generated image with the position information in the second instance to obtain a third image; determine the first loss based on the image sample and the third image, and determine the second loss based on the second instance and the second generated image, and optimize the first model through the first loss and the second loss.

[0159] In some embodiments, the model training module 4553 is also used to obtain image sample features and a third image feature; determine a first difference between the image sample features and the third image feature; use the mean of the sum of squares of the first differences as a first loss; determine a second difference between pixels at the same position in the second instance and the second generated image; calculate the sum of squares of the second differences of all pixels, and use the mean of the sum of squares of the second differences as a second loss.

[0160] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the image generation method described in the present invention.

[0161] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the image generation method provided by the embodiment of the present application, for example, Figure 3A The image generation method is shown.

[0162] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0163] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0164] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0165] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0166] In summary, through the embodiment of the present application, the first instance in the first image is processed independently, the noise gradient of the first instance is determined based on the characteristics of the first instance, and different noise injection strategies are designed for different instances. The noise amplitude that is negatively correlated with the noise gradient is obtained, and the amount of noise injection is adjusted according to the noise amplitude of each first instance, thereby improving the noise addition efficiency and preserving the original detail information of the first instance. The denoising process corresponds to the noise amplitude of the denoising process, reducing the interference of global noise in the denoising process, and improving the accuracy and generation speed of the first generated image. The position information of the first generated image and the first instance is aligned, the blurred first instance boundary in the first generated image is restored, the second image is avoided from being misaligned, the image generation quality is improved, and the consistency of the style and structure of different instances in the second image is ensured.

[0167] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. An image generation method, characterized in that: The method comprises: acquiring at least one first instance in a first image; Performing feature extraction on the first instance to obtain first instance features; determining a noise gradient of the first instance based on the first instance feature; Obtaining a noise amplitude negatively correlated with the noise gradient, and injecting noise positively correlated with the noise amplitude into the first instance to obtain a first noise image; performing denoising processing on the first noisy image to obtain a first generated image; The first generated image is fused with the position information in the first instance to obtain a second image including the first instance.

2. The method according to claim 1, characterized in that The obtaining of a noise amplitude negatively correlated with the noise gradient includes: determining a mean value of the noise gradient and a variance of the noise gradient; The sum of the mean value and the variance is determined, and the inverse of the sum is multiplied by a preset adjustment parameter to obtain a noise amplitude that is negatively correlated with the noise gradient.

3. The method according to claim 1 or 2, characterized in that The injecting noise positively correlated with the noise amplitude into the first instance to obtain a first noise image includes: Iteration t performs the following processing: Injecting noise positively correlated with the noise amplitude at the t-th time step, according to the noise ratio at the t-th time step determined by the scheduling algorithm, into the noise image at the t-1-th time step to obtain a noise image at the t-th time step, where t is a positive integer variable that increases from 1, and the value of t satisfies 2≤t≤T, where T is a preset time step threshold, and the noise image at the 1st time step is the first instance; When the value of t is iterated to T, the noise image at the T-th time step is used as the first noise image.

4. The method according to claim 1, wherein The acquiring at least one first instance of the first image comprises: Obtaining a global feature map of the first image; Identifying at least one candidate region from the global feature map; Determine a local feature map corresponding to the candidate region in the global feature map; Performing category prediction based on the local feature map to obtain a confidence score of the candidate region, performing regression based on the local feature map to obtain a position offset, and correcting the position of the candidate region based on the position offset to obtain a corrected candidate region; A first instance is determined according to the corrected candidate region and the confidence level.

5. The method according to claim 4, characterized in that The determining of the first instance according to the corrected candidate region and the confidence level includes: removing the corrected candidate regions whose confidence levels are lower than a preset confidence threshold, and dividing the corrected candidate regions remaining after the removal into a plurality of groups, wherein the confidence levels of the corrected candidate regions in one of the groups are within a confidence interval corresponding to the group; For each of the groups, determining a target candidate region in the group; Taking the maximum value of the confidences of the target candidate region as the confidence of the first instance; Acquire a first mask of the target candidate area, perform upsampling processing on the first mask, and binarize the upsampling first mask to obtain a second mask; mapping the second mask to the image coordinate system of the first image according to the target candidate region to obtain a third mask; The first instance is obtained according to the target candidate region, the confidence of the first instance, and the third mask.

6. The method according to claim 5, wherein determining the target candidate area in the group comprises: The following processing is performed for each group to determine the intersection-over-union (IoU) of any two corrected candidate regions in the group. When the IoU exceeds a preset IoU threshold, the corrected candidate region with a higher confidence level is retained, and the following processing is iteratively performed: Determine the intersection-over-union (IoU) of the retained revised candidate region and any remaining revised candidate region in the group; when the IoU exceeds the preset IoU threshold, retain the revised candidate region with a higher confidence level; and when the iteration ends, use the revised candidate region that is finally retained as the target candidate region in the group.

7. The method according to claim 1, characterized in that The determining the noise gradient of the first instance based on the first instance feature includes: determining a transverse gradient and a longitudinal gradient of the first instance feature; Determine a partial derivative of the transverse gradient and a partial derivative of the longitudinal gradient, and determine a sum of squares of the partial derivative of the transverse gradient and the partial derivative of the longitudinal gradient; The square root of the sum of squares is determined as the noise gradient of the first instance.

8. The method according to claim 1, characterized in that The first noise image, the first generated image, and the second image are realized by calling a first model. Before injecting noise positively correlated with the noise amplitude into the first instance to obtain the first noise image, the method further includes: obtaining at least one second instance of the image sample; acquiring, by the first model, a second generated image including the second instance; fusing the second generated image with the position information in the second instance to obtain a third image; A first loss is determined based on the image sample and the third image, and a second loss is determined based on the second instance and the second generated image, and the first model is optimized by the first loss and the second loss.

9. The method according to claim 8, characterized in that The determining of the first loss according to the image sample and the third image includes: Obtaining image sample features and third image features; determining a first difference between the image sample feature and the third image feature; The mean of the sum of squares of the first differences is taken as the first loss; The determining a second loss according to the second instance and the second generated image includes: determining a second difference between pixels at the same location in the second instance and the second generated image; The sum of squares of the second differences of all pixels is calculated, and the mean of the sum of squares of the second differences is taken as the second loss.

10. An image generating device, characterized in that: The device comprises: an image segmentation module, configured to obtain at least one first instance in the first image; An image generation module is configured to perform feature extraction on the first instance to obtain first instance features; determine a noise gradient of the first instance based on the first instance features; obtain a noise amplitude negatively correlated with the noise gradient, and inject noise positively correlated with the noise amplitude into the first instance to obtain a first noise image; perform denoising on the first noise image to obtain a first generated image; and fuse the first generated image with position information in the first instance to obtain a second image including the first instance.

11. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the image generation method according to any one of claims 1 to 9 when executing the computer-executable instructions or computer program stored in the memory.

12. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the image generation method according to any one of claims 1 to 9 is implemented.

13. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the image generation method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Scene graph generation method based on diffusion model

    CN116958652A

  • Method and device for noise reduction in medical images

    DE102007058498A1

  • Medical image noise reduction based on weighted anisotropic diffusion

    EP2447911A1

  • Denoising method and system for preserving clinically significant structures in reconstructed images using adaptively weighted anisotropic diffusion filter

    US20120106815A1