Image generation method and apparatus, storage medium and electronic device
By determining the replacement area in the target image and performing feature transformation to generate a synthetic map, the problem of insufficient image data in the industrial field is solved, the training effect and image recognition accuracy of the artificial intelligence model are improved, and a realistic target synthetic image is generated.
Patent Information
- Application Number
- PCT/CN2024/140230
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-12-18
- Publication Date
- 2025-07-03
AI Technical Summary
The lack of image data in the industrial field has led to poor training results in product quality detection, affecting the accuracy of image recognition.
By determining the replacement area in the target image, feature extraction and transformation are performed, the target synthetic map is generated, and the image generation model is input for reverse reasoning, and a realistic target synthetic image is generated.
The training effect of artificial intelligence models and the accuracy of image recognition are improved, and the generated images are more realistic, meeting the needs of product quality inspection.
Smart Images

Figure CN2024140230_03072025_PF_FP_ABST
Abstract
Description
Image generation method, device, storage medium and electronic device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on December 26, 2023, with application number 202311804219.2 and application name "Image Generation Method, Device, Storage Medium and Electronic Device", all contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of image processing, and more specifically to an image generation method, device, storage medium and electronic device. Background Art
[0003] As the requirements for product quality in industrial production become increasingly higher, various product quality testing technologies have been widely introduced into related industries.
[0004] In image-based product quality inspection technology, neural network models can be used to examine product appearance and structure to eliminate defective products. However, neural network models that inspect product quality based on appearance require corresponding image data for training. However, image data is often scarce in the industrial sector, making effective image generation technology applicable to industrial production crucial. Summary of the Invention
[0005] The present application has been made in view of the above-mentioned problems.
[0006] In a first aspect, the present application provides an image generation method, the method comprising:
[0007] determining a first replacement region and a second replacement region in the target image;
[0008] Performing feature extraction on the target image to obtain a target feature map of the target image;
[0009] Determining, in the target feature map, a first feature region corresponding to the first replacement region position and a second feature region corresponding to the second replacement region position;
[0010] In the target feature map, transforming the map information of the first feature region according to the map information of the second feature region to obtain a target synthetic map after transforming the map information of the first feature region;
[0011] The target synthetic atlas is input into a trained image generation model to obtain a target synthetic image, and the image generation model is used to perform reverse reasoning based on the target synthetic atlas to obtain the target synthetic image.
[0012] In a possible implementation, in the target feature map, transforming the map information of the first feature region according to the map information of the second feature region includes:
[0013] For each first eigenvalue in the first feature region, determining a second eigenvalue in the second feature region;
[0014] The first eigenvalue in the first feature region is replaced by the second eigenvalue.
[0015] In a possible implementation, determining the second eigenvalue in the second feature area includes: randomly determining the second eigenvalue in the second feature area.
[0016] In a possible implementation, the atlas information of the first feature area is transformed according to the atlas information of the second feature area to obtain a target synthetic atlas after the atlas information of the first feature area is transformed, including: directly replacing the atlas information of the first feature area with the atlas information of the second feature area.
[0017] In one possible implementation, performing feature extraction on a target image to obtain a target feature map of the target image includes: inputting the target image into a feature extraction model to obtain a plurality of target feature maps;
[0018] Determining a first feature area corresponding to a first replacement area position and a second feature area corresponding to a second replacement area position in the target feature map includes: in each target feature map, determining first map information corresponding to the first replacement area position and second map information corresponding to the second replacement area position, wherein the map information of the first feature area includes the first map information in all target feature maps, and the map information of the second feature area includes the second map information in all target feature maps.
[0019] In one possible implementation, the feature extraction model is a label-free knowledge distillation model.
[0020] In one possible implementation, the method further includes:
[0021] Performing feature extraction on the sample image to obtain a sample feature map of the sample image;
[0022] Inputting the sample feature map into the original image generation model to obtain a restored image;
[0023] Based on the difference between the restored image and the sample image, the original image generation model is trained to obtain the trained image generation model.
[0024] In a possible implementation, determining the first replacement region and the second replacement region in the target image includes: performing target detection on the target image to obtain the first replacement region or the second replacement region.
[0025] In one possible implementation, the image generation model is a steady-state diffusion model, and inputting the target synthetic atlas into the trained image generation model to obtain the target synthetic image includes:
[0026] The target synthetic map and the target image are input into a trained steady-state diffusion model, wherein the target synthetic map serves as conditional information of the steady-state diffusion model, and the target image serves as an input image of the steady-state diffusion model.
[0027] A second aspect of the present application further provides an image generating device, the device comprising:
[0028] A first determining module, configured to determine a first replacement area and a second replacement area in a target image;
[0029] A feature extraction module is used to extract features from the target image to obtain a target feature map of the target image;
[0030] a second determining module, configured to determine, in the target feature map, a first feature region corresponding to the first replacement region position and a second feature region corresponding to the second replacement region position;
[0031] a feature transformation module, configured to transform the atlas information of the first feature region in the target feature atlas according to the atlas information of the second feature region, so as to obtain a target synthetic atlas after transforming the atlas information of the first feature region;
[0032] An image generation module is used to input the target synthetic map into a trained image generation model to obtain a target synthetic image, and the image generation model is used to perform reverse reasoning based on the target synthetic map to obtain the target synthetic image.
[0033] The third aspect of the present application further provides a storage medium on which program instructions are stored. The program instructions are used to execute the above-mentioned image generation method when running.
[0034] The fourth aspect of the present application further provides an electronic device, which includes: a processor and a memory, wherein the memory stores computer program instructions, and the computer program instructions are used by the processor to execute the above-mentioned image generation method when running.
[0035] In this technical solution, the feature map corresponding to the target image is modified and used as input to the image generation model to generate a target composite image. This results in a more realistic and effective target composite image. This target composite image helps ensure the training effectiveness of the learning model used for product quality inspection.
[0036] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0038] FIG1 shows a schematic flow chart of an image generation method according to an embodiment of the present application;
[0039] FIG2 shows a schematic flow chart of transforming the atlas information of the first feature region according to the atlas information of the second feature region according to one embodiment of the present application;
[0040] FIG3 shows a schematic flowchart of a training image generation model according to one embodiment of the present application;
[0041] FIG4 shows a schematic diagram of an image generating method according to yet another embodiment of the present application;
[0042] FIG5 shows a schematic block diagram of an image generating device according to an embodiment of the present application;
[0043] FIG6 shows a schematic block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of the present application more apparent, the following is a detailed description of example embodiments of the present application with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the example embodiments described herein. Based on the embodiments of the present application described in this application, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of this application.
[0045] The image generation method of the embodiments of the present application generates different images based on the feature distribution of an image, either by smearing the foreground area of an existing image to generate the background area, or by smearing the background area of an existing image to generate the foreground area. This can provide more realistic images. This image generation method can be applied to application scenarios that require a large number of images, such as training artificial intelligence learning models. Image samples are important input data for training image recognition models, and real image samples can effectively improve the image recognition accuracy of the trained model. For example, in industrial inspection, it is necessary to detect product appearance defects to eliminate defective products. In this case, an artificial intelligence model can perform image recognition on the acquired product appearance images to determine whether the product is defective. Artificial intelligence models require a large amount of image sample data for training when performing image recognition. If the image samples are small and the authenticity is low, the training effect of the artificial intelligence model cannot be guaranteed, which in turn affects the accuracy of image recognition. In industrial inspection, real defect image samples are insufficient to support the training of artificial intelligence models for image recognition.
[0046] To at least partially solve the above problems, an embodiment of the present application provides an image generation method, wherein a different image is generated through image reverse reasoning based on a feature image of a modified image.
[0047] Fig. 1 shows a schematic flow chart of an image generation method according to an embodiment of the present application. As shown in Fig. 1 , the method includes the following steps S110 to S150.
[0048] In step S110 , a first replacement area and a second replacement area are determined in the target image.
[0049] The target image can be any suitable image. For example, the target image can be an RGB image or a grayscale image. The target image can be a static image or any video frame in a dynamic video. The target image can be an image of any suitable size and resolution. The target background image can be the original image directly captured by the image acquisition device, or it can be an image after the original image is preprocessed. The preprocessing operation can include all operations to improve the visual effect of the target image, increase its clarity, or highlight certain features in the image. For example and not limitation, the preprocessing operation can include operations such as digitization, geometric transformation, normalization, filtering, etc. of the original image. The target image can also be a synthesized image without affecting subsequent image processing.
[0050] The first replacement area and the second replacement area are different areas in the target image. The second replacement area is used to replace the first replacement area. The first replacement area and the second replacement area can be determined based on actual needs. If the target composite image is desired to have a reduced foreground area relative to the target image, at least a portion of the foreground area of the target replacement area can be determined as the first replacement area, while at least a portion of the background area of the target replacement area can be determined as the second replacement area. If the target composite image is desired to have a reduced background area relative to the target image, the opposite is true.
[0051] For example, the positions of the first replacement area and the second replacement area may be determined manually in the target image according to one's expectation, thereby obtaining the first replacement area and the second replacement area.
[0052] In step S120 , feature extraction is performed on the target image to obtain a target feature map of the target image.
[0053] The feature map, which can be simply referred to as a feature map, can be one or more. In this step, a convolutional neural network can be used to extract features from the target image to obtain its feature map. The convolutional neural network can include a series of convolution kernels that perform convolution operations on the target image. The feature map can be considered an abstract representation of the original image, where each pixel represents some specific feature.
[0054] Step S130: Determine, in the target feature map, a first feature region corresponding to the first replacement region position and a second feature region corresponding to the second replacement region position.
[0055] Based on the position correspondence, a first feature region is determined in the target feature map according to the position of the first replacement region in the target image, and a second feature region is determined in the target feature map according to the position of the second replacement region in the target image.
[0056] In step S140 , in the target feature atlas, the atlas information of the first feature region is transformed according to the atlas information of the second feature region to obtain a target synthetic atlas after the atlas information of the first feature region is transformed.
[0057] In step S130, after determining the positions corresponding to the first feature region and the second feature region, the atlas information within the first feature region and the atlas information within the second feature region can be obtained. In this step, the atlas information of the first feature region in the target feature map is transformed based on the atlas information of the second feature region in the target feature map, thereby changing the atlas information of the first feature region. It can be understood that the atlas information of the second feature region in the target feature map only serves as the basis for transforming the atlas information within the first feature region and does not change itself.
[0058] The above-mentioned transformation operation can be a replacement operation. In other words, the atlas information of the first feature area is directly replaced by the atlas information of the second feature area. In one possible embodiment, when multiple target feature maps are obtained, the atlas information in the first replacement area can be replaced in each target feature map according to the atlas information in the second replacement area.
[0059] Step S150: input the target synthetic atlas into the trained image generation model to obtain the target synthetic image. The image generation model is used to perform reverse reasoning based on the target synthetic atlas to obtain the target synthetic image.
[0060] The image generation model can be a trained neural network model, which can perform reverse reasoning based on the input feature map to generate an image corresponding to the feature map. It can be understood that the "reverse reasoning" operation is relative to the feature extraction operation. The feature extraction operation refers to extracting features from the input image, that is, the operation performed in step S120. The reverse reasoning operation is to obtain the image corresponding to the image feature based on the input image feature. The image generation model can be an existing or future developed model that can perform reverse reasoning to generate an image based on the input feature map, and this application does not impose any restrictions on this.
[0061] Visually, the target composite image output by the image generation model is an image obtained by transforming the image information in the first replacement region of the target image based on the image information in the second replacement region. This target composite image, generated by the image generation model, avoids the abruptness of the replaced image caused by direct image region replacement, making the resulting image more consistent with expectations and closer to reality.
[0062] In this technical solution, the feature map corresponding to the target image is modified and used as input to the image generation model to generate a target composite image. This results in a more realistic and effective target composite image. This target composite image helps ensure the training effectiveness of the learning model used for product quality inspection.
[0063] Exemplarily, the above-mentioned first replacement area or second replacement area can be obtained by performing target detection on the target image in step S110. Object detection (Object Detection) can be used to determine the position of the target of interest in the image. Depending on whether the target is included in the area, the first replacement area or the second replacement area can be obtained by the target detection model. Exemplarily, the target image can be an image of a chip (Die). Through holes can be provided on the chip, and the area where the through holes are located can be used as the foreground area, while other areas are used as the background area. The area where a through hole is located can be used as the first replacement area, and a part of the background area can be used as the second replacement area. In this example, target detection can be performed on the chip image with the through hole as the target to determine the first replacement area.
[0064] Automatically acquiring the first replacement area or the second replacement area through target detection improves the response speed of the overall system while ensuring the accuracy of the system, thereby improving the efficiency of image generation.
[0065] Exemplarily, step S120, performing feature extraction on the target image to obtain a target feature map of the target image, may include: inputting the target image into a feature extraction model to obtain a plurality of target feature maps. The feature extraction model may be any existing or future developed learning model for feature extraction. Exemplarily, the feature extraction model may be obtained based on sample image training. Alternatively, the feature extraction model may be obtained directly using data modeling without training. The target feature map output by the feature extraction model may be the same size as or different from that of the target image. It is understandable that the target feature map output by the feature extraction model may be scaled to make it the same size as the target image. This facilitates subsequent image processing.
[0066] When multiple target feature maps are obtained, step S130 determines the first feature area corresponding to the first replacement area position and the second feature area corresponding to the second replacement area position in the target feature map, which may include: in each target feature map, determining the first map information corresponding to the first replacement area position and the second map information corresponding to the second replacement area position. The map information of the first feature area includes the first map information in all target feature maps, and the map information of the second feature area includes the second map information in all target feature maps. In this case, the feature values at the same position of each target map can form a feature vector, which can be called the feature vector value of the current point. Thus, multiple target feature maps can be recorded as C×H×W, where C represents the number of target feature maps, and H and W represent the height and width of the target feature map, respectively.
[0067] For a target image, multiple target feature maps are obtained to describe its features from multiple angles. These target feature maps serve as the basis for subsequent image generation, and the final target composite image is better.
[0068] Exemplarily, the feature extraction model is a label-free knowledge distillation model.
[0069] Label-free knowledge distillation is a self-distillation learning model that, without requiring training with labels, can generate feature maps of the same size as an input image. Label-free knowledge distillation can generate multiple feature maps from a single input image. Label-free knowledge distillation removes redundant information from an image, focusing on learning its essential features.
[0070] In the above technical solution, label-free knowledge distillation is used to obtain the target feature map of the target image, which can simplify the model complexity and improve the model's generalization ability. In addition, label-free knowledge distillation does not require image samples for training, thus ensuring the accuracy of feature extraction even when there are only a few target images. Furthermore, the image generation method ensures the realistic effect of the target synthetic image generated even when there are only a few target images.
[0071] Exemplarily, step S140 transforms the atlas information of the first feature area according to the atlas information of the second feature area in the target feature map to obtain a target synthetic map after the atlas information of the first feature area is transformed. This may include: directly replacing the atlas information of the first feature area with the atlas information of the second feature area to obtain a target synthetic map. The second feature area and the first feature area may have the same shape and size. The atlas information of the second feature area may be directly replaced with the atlas information of the first feature area. Specifically, the distance between the second feature area and the first feature area in different directions, such as the horizontal and vertical directions of the feature map, may be calculated based on the respective positions of the first feature area and the second feature area in the target feature map. Then, the second feature area is moved according to the calculated distance to replace the first feature area.
[0072] In the above technical solution, the atlas information of the second feature region is used to directly replace the atlas information of the first feature region. In other words, in this solution, the replacement is performed on a region-by-region basis. This solution can reduce the amount of computational data required to generate the target feature map and speed up image generation.
[0073] Figure 2 shows a schematic flow chart of transforming the atlas information of a first feature region based on the atlas information of a second feature region according to one embodiment of the present application. As shown in Figure 2, step S140, in a target feature atlas, transforms the atlas information of the first feature region based on the atlas information of the second feature region to obtain a target composite atlas after transforming the atlas information of the first feature region, which may include steps S141 and S142.
[0074] In step S141, for each first feature value in the first feature region, a second feature value is determined in the second feature region. In other words, the first feature value is the pixel value in the first feature region of the target feature map, and the second feature value is the pixel value in the second feature region of the target feature map.
[0075] It can be understood that the target feature map is obtained by extracting features from the target image, and the eigenvalues in the target feature map can be considered to represent the pixel values in the target image. The position of the eigenvalue in the target feature map corresponds to the position of the pixel it represents in the target image. In an embodiment where the target image and the target feature map have the same size, a eigenvalue in the target feature map represents a pixel in the target image. The correspondence between the position of the eigenvalue in the target feature map and the position of the pixel it represents in the target image means that the horizontal and vertical coordinates of the two are the same. It can be understood that in an embodiment where the target image and the target feature map have different sizes, for example, the size of the target image is larger than the size of the target feature map, then a eigenvalue in the target feature map may represent multiple pixels in the target image.
[0076] Step S142: Replace the first eigenvalue in the first feature area with the second eigenvalue.
[0077] In this step, each first feature value in the first feature area can be traversed and replaced with the second feature value in the second feature area until all first feature values in the first feature area are replaced.
[0078] In the above technical solution, the replacement is performed based on the characteristic values in the target feature map. As a result, the generated target feature map is more detailed. Furthermore, the target composite image generated by the image generation method is more realistic.
[0079] Exemplarily, when the atlas information in the first feature region is replaced in units of feature values, step S141 may include: randomly determining a second feature value in the second feature region.
[0080] In this example, when traversing the first eigenvalue in the first feature area, a random second eigenvalue is determined in the second feature area instead of the second eigenvalue corresponding to the position of the first eigenvalue in the first feature area, and then the eigenvalue to be replaced in the first feature area is replaced with the randomly selected second eigenvalue.
[0081] In the above technical solution, the first eigenvalue in the first feature region is replaced based on the randomly determined second eigenvalue in the second feature region. Because the second feature region corresponds to a predetermined second replacement region in the target image, randomly extracting eigenvalues from it to obtain the second eigenvalues used for replacement enriches the generated target composite image and reduces the computational effort required to determine the eigenvalue positions during data replacement, thereby improving image generation efficiency.
[0082] FIG3 shows a schematic flow chart of training an image generation model according to an embodiment of the present application. Optionally, the above-mentioned image generation method may further include the step of training the image generation model. As shown in FIG3 , the training image generation model may include the following steps S101 to S103.
[0083] In step S101 , feature extraction is performed on a sample image to obtain a sample feature map of the sample image.
[0084] The sample image may be a real image obtained through detection or a synthesized image. In one embodiment, a trained feature extraction model may be used to extract features from the sample image to obtain a sample feature map.
[0085] Step S101 is similar to the aforementioned step S120 and will not be described again for the sake of brevity.
[0086] Step S102: input the sample feature map into the original image generation model to obtain a restored image.
[0087] The original image generation model can be untrained or incompletely trained. Its image generation performance can be improved through training. The restored image is the image generated by the original image generation model through reverse inference based on the input sample feature map. The restored image is compared with the sample image to determine the difference between the restored image and the sample image.
[0088] Step S103 : training the original image generation model based on the difference between the restored image and the sample image to obtain a trained image generation model.
[0089] The difference between the restored image and the sample image represents the performance of the current image generation model. A larger difference indicates worse performance, and vice versa. Based on the difference between the restored image and the sample image, the image generation model is trained. For example, parameters of the image generation model are modified to gradually improve its image generation performance until a trained image generation model that meets the requirements is obtained.
[0090] The trained image generation model can be used to perform reverse reasoning based on the target synthesis map to obtain a more ideal target synthesis image. Compared with the original image generation model, the images generated by the trained image generation model are more in line with the expected requirements and more efficient.
[0091] Exemplarily, the image generation model is a stable-diffusion model. Step S150 of inputting the target synthetic atlas into the trained image generation model to obtain the target synthetic image includes: inputting the target synthetic atlas and the target image into the trained stable-diffusion model, wherein the target synthetic atlas serves as conditional information for the stable-diffusion model, and the target image serves as an input image for the stable-diffusion model.
[0092] According to an embodiment of the present application, based on the regular distribution of the target image, the distribution information contained in its target feature map is used as a guide to gradually denoise the noisy image and generate a target composite image that matches the feature map. Specifically, during the image generation process in the steady-state diffusion model, noise is gradually added to the image. As noise is added, the steady-state diffusion model gradually learns the characteristics of the target image. Under the guidance of the conditional information provided by the target feature map, the desired target composite image can be generated based on the target feature map.
[0093] The steady-state diffusion model can be a spatial transformer architecture, which may include a U-Net network. The spatial transformer architecture may include a cross-attention module and a basic transformer module. When a target image is received as input for the steady-state diffusion model, the target feature map of the target image is used as conditional information, and the two are modeled using the cross-attention module. It is understood that before the cross-attention module performs modeling, the target feature map can be encoded using an encoder. In the cross-attention module, the target image serves as the query keyword and the target feature map serves as the key value. This allows the cross-attention module to learn the correlation between the corresponding content of the target image and the target feature map. The basic transformer module actually calls the cross-attention module and generates a target composite image based on the correlation learned by the cross-attention module. In summary, when the steady-state diffusion model performs inverse inference on an image, in addition to inputting Gaussian noise and the target image, the target feature map corresponding to the target image is also input as conditional information, which can be used to generate a specified target composite image.
[0094] In this technical solution, a steady-state diffusion model is used as the image generation model, using the input target composite image as conditional information to generate the target composite image. Compared to inputting text information as conditional information into the steady-state diffusion model, the specified image generated based on the input target composite image is more consistent with the desired requirements.
[0095] FIG4 shows a schematic diagram of an image generating method according to yet another embodiment of the present application.
[0096] Exemplarily, before image generation, the steady-state diffusion model can be trained to obtain a trained steady-state diffusion model that can be used for image generation. During training, the features of the sample image can be extracted through the trained unlabeled knowledge distillation model to obtain a sample feature map. For the training of the steady-state diffusion model, the extracted sample feature map can be passed through the encoder in the steady-state diffusion model and then input into the deep network of the U-Net structure to obtain a restored image corresponding to the sample image. Based on the difference between the restored image and the sample image, the steady-state diffusion model is trained to obtain a trained steady-state diffusion model that can be used for image generation. It can be understood that the training process of the above-mentioned steady-state diffusion model can be an iterative process, which can continue until the trained steady-state diffusion model meets the requirements or meets other training completion conditions to obtain a trained steady-state diffusion model that meets the requirements.
[0097] For example, as shown in Figure 4, after acquiring a target image, a trained unlabeled knowledge distillation model is used to obtain multiple target feature maps of the target image. The feature values at the same position in all target feature maps can be combined into a feature vector. The feature set corresponding to the target feature map output by the unlabeled knowledge distillation model can be expressed as C × H × W, where C represents the number of target feature maps, and H and W represent the height and width of the target feature map, respectively.
[0098] For example, as shown in FIG4 , a first replacement area and a second replacement area can be determined in the target image. The first replacement area is the area to be painted. The second replacement area is the area based on which the first replacement area is painted. The first replacement area and the second replacement area have different functions, but can be expressed in the same way. For simplicity, the first replacement area is used as an example for description. The first replacement area can be represented by its position Mask. x,y To express.
[0099] Mask x,y ={(x,y) ROI}
[0100] Among them, (x,y) represents the pixel coordinates in the target image, {(x,y) ROI} refers to the set of pixel coordinates in the first replacement area.
[0101] After obtaining the target feature map through the trained unlabeled knowledge distillation model, the first feature region F(I) corresponding to each target feature map can be determined according to the first replacement region in the target image. ROI .
[0102] For example, according to Mask x,y , the first feature region F(I) can be determined by the following formula: ROI :
[0103] F(I) ROI =R h,w ((h,w)∈Mask x,y )
[0104] Where (h, w) represents the feature value coordinates of the pixels in the first replacement area in the target feature map, R h,w Represents a set of feature value coordinates corresponding to pixels in the first replacement area in the target feature map.
[0105] Similarly, as shown in Figure 4, after obtaining the target feature map through the trained unlabeled knowledge distillation model, the corresponding second feature region in each target feature map can be determined based on the second replacement region in the target image. The data in the second feature region is used to transform the data in the first feature region.
[0106] According to the position of the second replacement area Trget determined in the target image x,y , determine the position F(I) of the second feature region corresponding to the second replacement region in each feature map ROI’ The position of the second replacement area can be set manually or by an algorithm. The second feature area F(I) can be determined as follows: ROI’ .
[0107] F(I) ROI’ =R h,w ((h,w)∈Trget x,y )
[0108] Among them, Trget x,y is the pixel coordinate set in the first replacement area, i.e., the second replacement area. (h, w) represents the feature value coordinates of the pixels in the target feature map corresponding to the second replacement area, R h,w Represents a set of feature value positions corresponding to pixels in the second replacement area in the target feature map.
[0109] Exemplarily, as shown in FIG4 , after determining the positions of the first and second feature regions in each target feature map, the map information of the first feature region is transformed based on the map information of the second feature region to obtain a target composite map after the map information of the first feature region has been transformed. In each target feature map, any second eigenvalue can be randomly determined in the second feature region. The first eigenvalue in the first feature region is replaced based on the second eigenvalue until all first eigenvalues in the first feature region in the target feature map have been replaced, thereby obtaining multiple target feature maps after partial eigenvalue replacement. When the first and second feature regions are of the same size and shape, the data of the first feature region can be directly replaced with the data of the second feature region in each target feature map to obtain multiple replaced target feature maps. Once the multiple replaced target feature maps are obtained, they can be synthesized into a target composite map. Exemplarily, the synthesis of the target composite map can be achieved using algorithms, artificial intelligence, and neural network models. Alternatively, the multiple target feature maps can be directly spliced into a single target composite map.
[0110] For example, as shown in Figure 4, after obtaining the target synthetic atlas, it is combined with Gaussian noise as the input of the trained steady-state diffusion model. The target synthetic image is obtained by reverse reasoning through the U-Net network in the steady-state diffusion model.
[0111] FIG5 shows a schematic block diagram of an image generation apparatus 500 according to an embodiment of the present application. As shown in FIG5 , the image generation apparatus 500 includes a first determination module 510 , a feature extraction module 520 , a second determination module 530 , a feature transformation module 540 , and an image generation module 550 .
[0112] The first determining module 510 is configured to determine a first replacement area and a second replacement area in a target image.
[0113] The feature extraction module 520 is used to extract features from the target image to obtain a target feature map of the target image.
[0114] The second determining module 530 is configured to determine, in the target feature map, a first feature region corresponding to the first replacement region position and a second feature region corresponding to the second replacement region position.
[0115] The feature transformation module 540 is used to transform the atlas information of the first feature region in the target feature atlas according to the atlas information of the second feature region to obtain a target synthetic atlas after the atlas information of the first feature region is transformed.
[0116] The image generation module 550 is used to input the target synthetic map into the trained image generation model to obtain the target synthetic image. The image generation model is used to perform reverse reasoning based on the target synthetic map to obtain the target synthetic image.
[0117] According to another aspect of the present application, an electronic device is also provided. Figure 6 shows a schematic block diagram of an electronic device 600 according to one embodiment of the present application. As shown in Figure 6, electronic device 600 includes a processor 610 and a memory 620, wherein the memory 620 stores computer program instructions, which, when executed by the processor 610, are used to execute the image generation method described above.
[0118] In addition, according to another aspect of the present application, a storage medium is also provided. Program instructions are stored on the storage medium. When the program instructions are executed by a computer or a processor, the computer or processor is caused to perform the corresponding steps of the above-mentioned image generation method of the embodiment of the present application, and is used to implement the corresponding module in the above-mentioned image generation device according to the embodiment of the present application or the corresponding module in the above-mentioned electronic device. The storage medium may, for example, include a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, or any combination of the above-mentioned storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.
[0119] A person skilled in the art can understand the specific implementation and beneficial effects of the above-mentioned image generation device, electronic device and storage medium by reading the above-mentioned detailed description of the image generation method. For the sake of brevity, they will not be described here in detail.
[0120] Example:
[0121] Embodiment 1: A method for generating an image, wherein the method comprises:
[0122] determining a first replacement region and a second replacement region in the target image;
[0123] Performing feature extraction on the target image to obtain a target feature map of the target image;
[0124] Determining, in the target feature map, a first feature region corresponding to the first replacement region position and a second feature region corresponding to the second replacement region position;
[0125] In the target feature map, transforming the map information of the first feature region according to the map information of the second feature region to obtain a target synthetic map after transforming the map information of the first feature region;
[0126] The target synthetic atlas is input into a trained image generation model to obtain a target synthetic image, and the image generation model is used to perform reverse reasoning based on the target synthetic atlas to obtain the target synthetic image.
[0127] Embodiment 2: The method according to embodiment 1, wherein transforming the atlas information of the first feature region according to the atlas information of the second feature region comprises:
[0128] For each first eigenvalue in the first feature region, determining a second eigenvalue in the second feature region;
[0129] The first eigenvalue in the first feature region is replaced by the second eigenvalue.
[0130] Embodiment 3: The method according to embodiment 1 or 2, wherein determining the second characteristic value in the second characteristic region includes:
[0131] A second feature value is randomly determined in the second feature area.
[0132] Embodiment 4: According to the method described in any one of Embodiments 1-3, the step of transforming the atlas information of the first feature area according to the atlas information of the second feature area includes:
[0133] The atlas information of the second feature area is used to directly replace the atlas information of the first feature area.
[0134] Embodiment 5: According to the method described in any one of embodiments 1-4, wherein the step of extracting features from the target image to obtain a target feature map of the target image comprises:
[0135] Inputting the target image into a feature extraction model to obtain a plurality of target feature maps;
[0136] The determining, in the target feature map, a first feature region corresponding to the first replacement region position and a second feature region corresponding to the second replacement region position, includes:
[0137] In each target feature map, the first map information corresponding to the first replacement area position and the second map information corresponding to the second replacement area position are determined, wherein the map information of the first feature area includes the first map information in all target feature maps, and the map information of the second feature area includes the second map information in all target feature maps.
[0138] Example 6: The method described in any one of Examples 1-5, wherein the feature extraction model is a label-free knowledge distillation model.
[0139] Embodiment 7: The method according to any one of embodiments 1 to 6, wherein the method further comprises:
[0140] Performing feature extraction on the sample image to obtain a sample feature map of the sample image;
[0141] Inputting the sample feature map into the original image generation model to obtain a restored image;
[0142] Based on the difference between the restored image and the sample image, the original image generation model is trained to obtain the trained image generation model.
[0143] Embodiment 8: According to the method described in any one of embodiments 1-7, determining the first replacement area and the second replacement area in the target image includes:
[0144] Performing target detection on the target image to obtain the first replacement area or the second replacement area.
[0145] Embodiment 9: The method according to any one of embodiments 1-8, wherein the image generation model is a steady-state diffusion model.
[0146] Inputting the target synthetic atlas into a trained image generation model to obtain a target synthetic image includes:
[0147] The target synthetic map and the target image are input into a trained steady-state diffusion model, wherein the target synthetic map serves as conditional information of the steady-state diffusion model, and the target image serves as an input image of the steady-state diffusion model.
[0148] Embodiment 10: An image generating device, comprising:
[0149] A first determining module, configured to determine a first replacement area and a second replacement area in a target image;
[0150] A feature extraction module is used to extract features from the target image to obtain a target feature map of the target image;
[0151] a second determining module, configured to determine, in the target feature map, a first feature region corresponding to the first replacement region position and a second feature region corresponding to the second replacement region position;
[0152] a feature transformation module, configured to transform the atlas information of the first feature region in the target feature atlas according to the atlas information of the second feature region, so as to obtain a target synthetic atlas after transforming the atlas information of the first feature region;
[0153] An image generation module is used to input the target synthetic map into a trained image generation model to obtain a target synthetic image, and the image generation model is used to perform reverse reasoning based on the target synthetic map to obtain the target synthetic image.
[0154] Embodiment 11: A storage medium having program instructions stored thereon, wherein the program instructions are used to execute the image generation method described in any one of embodiments 1-9 when running.
[0155] Example 12: An electronic device comprising a processor and a memory, characterized in that computer program instructions are stored in the memory, and the computer program instructions are used by the processor to execute the image generation method described in any one of Examples 1-9 when the processor is running.
[0156] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely illustrative and are not intended to limit the scope of the present application. Various changes and modifications may be made therein by those skilled in the art without departing from the scope and spirit of the present application. All such changes and modifications are intended to be included within the scope of the present application as required by the appended claims.
[0157] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0158] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units described is merely a logical function division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another device, or ignoring or not performing some features.
[0159] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.
[0160] Similarly, it should be understood that in order to streamline the present application and aid in understanding one or more of the various inventive aspects, in the description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, this approach of the present application should not be interpreted as reflecting the intention that the application claimed for protection requires more features than those explicitly recited in each claim. More precisely, as reflected in the corresponding claims, the inventive point is that the corresponding technical problem can be solved with fewer features than all the features of a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim itself serving as a separate embodiment of the present application.
[0161] It will be understood by those skilled in the art that, except where mutually exclusive, all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or apparatus disclosed herein may be combined in any combination. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature providing the same, equivalent, or similar purpose.
[0162] Furthermore, those skilled in the art will appreciate that although some embodiments described herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of this application and to form different embodiments. For example, in the claims, any of the claimed embodiments may be used in any combination.
[0163] The various component embodiments of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art will appreciate that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all of the functions of some modules in the image generating device according to the embodiment of the present application. The present application can also be implemented as a device program (e.g., a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0164] It should be noted that the above embodiments illustrate rather than limit the present application, and that a person skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbols placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of appropriately programmed computers. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.
[0165] The above description is merely a specific embodiment or illustration of a specific embodiment of the present application, and the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. The scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. An image generation method, characterized in that, The method includes: Determine a first replacement region and a second replacement region in the target image; Extract features from the target image to obtain a target feature map of the target image; Determine a first feature region corresponding to the position of the first replacement region and a second feature region corresponding to the position of the second replacement region in the target feature map; In the target feature map, transform the map information of the first feature region according to the map information of the second feature region to obtain a target synthesis map after transforming the map information of the first feature region; Input the target synthesis map into a trained image generation model to obtain a target synthesis image, where the image generation model is used to perform reverse inference according to the target synthesis map to obtain the target synthesis image.
2. The method according to claim 1, characterized in that Transforming the map information of the first feature region according to the map information of the second feature region includes: For each first feature value in the first feature region, determine a second feature value in the second feature region; Replace the first feature value in the first feature region with the second feature value.
3. The method according to claim 2, wherein The determining the second feature value in the second feature region includes: Randomly determine a second feature value in the second feature region.
4. The method according to claim 1, characterized in that, The transforming the map information of the first feature region according to the map information of the second feature region includes: Directly replace the map information of the first feature region with the map information of the second feature region.
5. The method according to any one of claims 1 to 4, characterized in that The extracting features from the target image to obtain a target feature map of the target image includes: Input the target image into a feature extraction model to obtain a plurality of target feature maps; The determining a first feature region corresponding to the position of the first replacement region and a second feature region corresponding to the position of the second replacement region in the target feature map includes: In each target feature map, determine first map information corresponding to the position of the first replacement region and second map information corresponding to the position of the second replacement region, where the map information of the first feature region includes the first map information in all target feature maps, and the map information of the second feature region includes the second map information in all target feature maps.
6. The method according to claim 5, wherein The feature extraction model is an unlabeled knowledge distillation model.
7. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Extract features from the sample image to obtain a sample feature map of the sample image; Input the sample feature map into the original image generation model to obtain a restored image; Train the original image generation model based on the difference between the restored image and the sample image to obtain the trained image generation model.
8. The method according to any one of claims 1 to 4, characterized in that The determining a first replacement region and a second replacement region in the target image includes: Perform target detection on the target image to obtain the first replacement region or the second replacement region.
9. The method according to any one of claims 1 to 4, characterized in that The image generation model is a stable diffusion model Inputting the target synthesis map into a trained image generation model to obtain a target synthesis image includes: Inputting the target synthesis map and the target image into a trained steady-state diffusion model, where the target synthesis map serves as conditional information for the steady-state diffusion model, and the target image serves as the input image for the steady-state diffusion model.
10. An image generation device, characterized in that, Including: A first determination module for determining a first replacement region and a second replacement region in the target image; A feature extraction module for extracting features from the target image to obtain a target feature map of the target image; A second determination module for determining a first feature region corresponding to the position of the first replacement region and a second feature region corresponding to the position of the second replacement region in the target feature map; A feature transformation module for transforming the map information of the first feature region according to the map information of the second feature region in the target feature map to obtain a target synthesis map after transforming the map information of the first feature region; An image generation module for inputting the target synthesis map into a trained image generation model to obtain a target synthesis image, where the image generation model is used to perform reverse inference according to the target synthesis map to obtain the target synthesis image.
11. A storage medium, on which program instructions are stored, characterized in that, The program instructions are used to execute the image generation method according to any one of claims 1 to 9 when running.
12. An electronic device, comprising a processor and a memory, characterized in that, Computer program instructions are stored in the memory, and when the computer program instructions are run by the processor, they are used to execute the image generation method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Image processing method and device
CN111354059A
Image recognition model training method and system and computer equipment
CN112836756A
Breast feature extraction and detection model training method and device
CN114820576A
Image recognition method and device, electronic equipment and computer readable storage medium
CN114821614A
Feature map generation method, and target detection model training method and device
CN116612295A