Multi-angle image generation method, apparatus and device, and computer program product
By preprocessing the input image and constructing an image generation data set, multiple angle images are generated using a multi-angle image generation model based on the diffusion process, and post-processing, the problems of insufficient consistency and quality of multi-angle image generation in the prior art are solved, and high consistency and high-quality multi-angle image generation are achieved.
Patent Information
- Application Number
- CN202510101489.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art has limitations when generating multi-angle images, usually focusing only on single angle generation, or lack consistency when generating multiple angles, resulting in incoherent and unnatural phenomena of generated images between different perspectives.
A multi-angle image generation method is adopted to construct an image generation data set by acquiring the input image and preprocessing it, and an initial multiple angle images are generated using a multi-angle image generation model based on the diffusion process, and post-processing is performed to improve image quality.
It achieves a high degree of consistency in the generation of images at different perspectives, and meets the needs of high-precision and high-quality image generation through refined data preprocessing and post-processing optimization operations.
Smart Images

Figure CN119941908A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image generation technology, and in particular to a multi-angle image generation method, device and equipment, and a computer program product. Background Art
[0002] Multi-angle image generation mainly uses algorithms and models to generate images of the same object from different perspectives, which is usually used to enhance the diversity and realism of image generation.
[0003] The multi-angle image generation technologies involved in the existing technologies, such as Unique3D and StableVideo 3D, are mainly used to generate 3D mesh models. L-MAGIC focuses on using language models to control the generation process. However, image generation models, especially those based on diffusion models, have limitations when generating multi-angle images. Existing technologies usually only focus on single-angle generation, or lack consistency when generating multiple angles, resulting in the generated images being incoherent and unnatural between different perspectives. Summary of the invention
[0004] The embodiments of the present application provide a multi-angle image generation method, apparatus and device, and a computer program product to improve the consistency and image quality of multi-angle image generation.
[0005] The present application embodiment adopts the following technical solutions:
[0006] In a first aspect, an embodiment of the present application provides a multi-angle image generation method, the multi-angle image generation method comprising:
[0007] Acquire an input image and preprocess the input image to obtain a preprocessed image;
[0008] constructing an image generation data set based on the preprocessed images;
[0009] Generate initial multiple-angle images using a multi-angle image generation model according to the image generation data set and the multi-angle category information, wherein the multi-angle image generation model is a generation model based on a diffusion process;
[0010] The initial multiple angle images are post-processed to obtain multiple angle images corresponding to the input image.
[0011] Optionally, preprocessing the input image to obtain a preprocessed image includes:
[0012] Performing background removal on the input image to obtain an image after background removal;
[0013] Extracting the main body of the image after the background is removed to obtain a main body image;
[0014] The main image is refined to obtain a refined image, wherein the refined processing includes at least one of cropping, scaling, and adding a margin.
[0015] Optionally, constructing an image generation data set based on the preprocessed image includes:
[0016] Digitally processing the preprocessed image to obtain image tensor data;
[0017] The image tensor data is copied into multiple copies according to multiple angles to be generated, and the multiple copies of image tensor data constitute the image generation data set.
[0018] Optionally, the step of digitally processing the preprocessed image to obtain image tensor data includes:
[0019] Converting the preprocessed image into an image pixel array;
[0020] Normalizing the image pixel values in the image pixel array to obtain a normalized image pixel array;
[0021] Add a white background to the image according to the normalized image pixel array, and combine the images using the Alpha channel parameter to obtain a combined image;
[0022] The combined image is converted into the image tensor data.
[0023] Optionally, generating the initial multiple-angle images using a multi-angle image generation model according to the image generation data set and the multi-angle category information includes:
[0024] Obtaining preset camera angle embedding information of the multi-angle image generation model during the training phase;
[0025] Transforming the preset camera angle embedding information into corresponding category labels to obtain multi-angle category information;
[0026] The multi-angle image generation model is used to generate initial multi-angle images according to the image generation data set and the multi-angle category information.
[0027] Optionally, the post-processing includes at least one of defogging, denoising and super-resolution upscaling.
[0028] Optionally, the multi-angle image generation model is an MVDiffusion model.
[0029] In a second aspect, an embodiment of the present application further provides a multi-angle image generation device, the multi-angle image generation device comprising:
[0030] A preprocessing unit, used to obtain an input image and preprocess the input image to obtain a preprocessed image;
[0031] A construction unit, configured to construct an image generation data set based on the preprocessed image;
[0032] A generating unit, configured to generate initial multiple-angle images using a multi-angle image generating model according to the image generating data set and the multi-angle category information, wherein the multi-angle image generating model is a generating model based on a diffusion process;
[0033] The post-processing unit is used to post-process the initial multiple angle images to obtain multiple angle images corresponding to the input image.
[0034] In a third aspect, an embodiment of the present application further provides a device, including:
[0035] A processor; and a memory arranged to store computer executable instructions, which, when executed, cause the processor to perform any of the multi-angle image generation methods described above.
[0036] In a fourth aspect, an embodiment of the present application further provides a computer program product, comprising a computer program / instructions, which, when executed by a processor, implements any of the aforementioned multi-angle image generation methods.
[0037] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: the multi-angle image generation method of the embodiments of the present application first obtains an input image and preprocesses the input image to obtain a preprocessed image; then an image generation data set is constructed based on the preprocessed image; then, according to the image generation data set and multi-angle category information, an initial multi-angle image is generated using a multi-angle image generation model, and the multi-angle image generation model is a generation model based on a diffusion process; finally, the initial multi-angle images are post-processed to obtain multiple angle images corresponding to the input image. The multi-angle image generation method of the embodiments of the present application covers the whole process design from image data preprocessing to generation data set construction, to multi-angle image generation and image quality optimization, and uses a generation model based on a diffusion process to generate multi-angle images, achieving a high degree of consistency in the generation of images at different viewing angles. Through refined data preprocessing and post-processing optimization operations, the requirements for high-precision and high-quality image generation are met. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0039] Figure 1 A schematic diagram of a process of generating a multi-angle image in an embodiment of the present application;
[0040] Figure 2 This is a schematic diagram of an image preprocessing process in an embodiment of the present application;
[0041] Figure 3 A schematic diagram of a process for constructing an image generation data set in an embodiment of the present application;
[0042] Figure 4 A schematic diagram of a generation process of a multi-angle image generation model in an embodiment of the present application;
[0043] Figure 5 This is a schematic diagram of the structure of a multi-angle image generation device in an embodiment of the present application;
[0044] Figure 6 This is a schematic diagram of the structure of a device in an embodiment of the present application. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0046] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.
[0047] The main technical terms involved in this application include:
[0048] 1) GAN (Generative Adversarial Network): GAN is a neural network composed of a generator and a discriminator. The generator generates false data, while the discriminator tries to distinguish between real data and false data, and improves the quality of generated data through mutual competition. It is often used in tasks such as image generation and video generation.
[0049] 2) VAE (Variational Autoencoder): VAE is a generative model that encodes data into a distribution in a latent space, samples from the distribution, and decodes it into new data, thereby generating new samples similar to the original data. It is widely used in image generation and data dimensionality reduction.
[0050] 3) Diffusion Models: Diffusion models generate new images by gradually adding noise to data and then denoising it. Their advantage is that they can generate high-quality and realistic images. In recent years, they have attracted much attention in the field of image generation.
[0051] 4) MVDiffusion model: A generative model based on the diffusion process, which uses multivariate joint distribution for modeling and can effectively generate images with high consistency and diversity.
[0052] GAN is an early and landmark method in the field of image generation. With its unique adversarial mechanism, it has shown great advantages in image generation, especially in the task of generating images from a single perspective. It consists of a generator and a discriminator, which compete with each other to continuously improve the quality of the generated images. GAN performs well in generating high-resolution, realistic images and can generate images of various styles. However, one challenge of GAN is that the training process may be unstable and prone to problems such as mode collapse, resulting in unsatisfactory generation results. GAN also faces some limitations in multi-angle image generation. Since each generation process is independent, it is difficult to ensure the consistency of perspective when generating images from different angles. For example, if you want to generate images of an object taken from different angles, GAN may generate completely different images from multiple angles, rather than images of the same object observed from different angles.
[0053] VAE is a generative model based on the encoder-decoder architecture. It can generate images with diversity and controllability by learning the potential distribution of data. VAE can maintain a certain continuity when generating images, which has certain advantages for multi-angle image generation. For example, by adjusting the vectors in the latent space, images of different angles can be smoothly generated. However, the quality of images generated by VAE may sometimes be inferior to GAN, especially when generating complex images, there may be problems with blur or unclear details. In addition, when generating multi-angle images, VAE also needs to consider the correspondence between images at different angles to ensure the consistency of perspective.
[0054] The diffusion model is an emerging image generation method that has attracted much attention in recent years due to its high-quality performance in generating images. It introduces a noise diffusion process to gradually transform the image into noise, and then restore the image from the noise. This method can generate images with rich details and realistic effects, but the multi-view images generated based on the traditional diffusion model still have shortcomings in terms of consistency and other aspects.
[0055] In summary, existing image generation technologies cannot simultaneously achieve high quality and multi-view consistency of generated images, especially in the following aspects:
[0056] 1) Unstable quality of multi-view generation: When processing images from different angles, current technologies often generate results with significant quality differences. Some angles may perform well, while other angles may lack details or have obvious flaws. This instability leads to a decline in the overall generation effect, especially in scenes that need to display multiple angles at the same time, which cannot meet high-standard image generation requirements.
[0057] 2) Lack of perspective consistency: Most existing methods cannot guarantee the consistency of geometric structure and visual features when generating images of the same object from different perspectives. For example, an object may look well-proportioned from a certain perspective, but lose its original geometric symmetry or visual feature coherence from another perspective. This inconsistency not only affects the authenticity of the image, but also causes the generated results to appear inconsistent when applied to actual scenes.
[0058] 3) Low training efficiency: In order to achieve the ideal generation effect, a large amount of image data and a long training process are often required, which consumes a lot of computing resources. In practical applications, this inefficient training method not only increases development costs, but also limits the widespread application and popularization of image generation technology. In some scenarios that require fast response and real-time generation, the low efficiency of existing technologies is particularly insufficient.
[0059] Based on this, the present application embodiment provides a multi-angle image generation method, such as Figure 1 As shown, a schematic flow chart of a multi-angle image generation method in an embodiment of the present application is provided, and the multi-angle image generation method at least includes the following steps S110 to S140:
[0060] Step S110, acquiring an input image and preprocessing the input image to obtain a preprocessed image.
[0061] Multi-angle image generation is to generate images of other angles corresponding to the target in the original image. Therefore, when performing multi-angle image generation, it is necessary to obtain the original image first. The requirement for the original image is that the image must contain a clear subject target, for example, it can be an image containing targets such as animals and human faces.
[0062] After obtaining the original image, it is necessary to perform a series of preprocessing on the original image. The preprocessing operation may include, for example, denoising, extracting the subject in the image, adjusting the image size, etc. The purpose of the preprocessing is to improve the image quality and obtain an image with a clear and distinct subject, so that it is more suitable for subsequent image processing and generation tasks. Of course, the specific preprocessing operations can be flexibly set by those skilled in the art according to the actual situation, and are not specifically limited here.
[0063] Step S120: constructing an image generation data set based on the preprocessed image.
[0064] Based on the preprocessed images obtained in the previous steps, an image generation dataset is further constructed. This process may include vectorizing the preprocessed images and copying the image data according to the number of angles to be generated, etc., in order to construct a data structure and form that meets the input requirements of subsequent models.
[0065] Step S130, generating initial multi-angle images using a multi-angle image generation model according to the image generation data set and the multi-angle category information, wherein the multi-angle image generation model is a generation model based on a diffusion process.
[0066] After constructing the image generation dataset, it is also necessary to obtain multi-angle category information. The multi-angle category information is reference information used to guide the multi-angle image generation model to generate images at different angles. The multi-angle category information can be determined based on the preset camera angle information during model training. The angle categories may include, for example, the front, left side, right side, left oblique 45-degree side, right oblique 45-degree side, back, etc.
[0067] The above-mentioned image generation dataset and multi-angle category information are used as the input of the multi-angle image generation model, and the multi-angle category information is used to guide the multi-angle image generation model to generate initial images at multiple different angles.
[0068] The multi-angle image generation model of the embodiment of the present application adopts a generation model based on a diffusion process. The generation model based on a diffusion process here is different from the traditional diffusion model. It can fully capture the intrinsic connection between multi-angle images, so that the generated multi-angle images have a high degree of consistency.
[0069] Step S140, post-processing the initial multiple angle images to obtain multiple angle images corresponding to the input image.
[0070] In order to further improve the quality of the generated multi-angle images, it is necessary to further post-process the multi-angle images initially generated in the above steps to adjust and optimize the details and visual effects of the multi-angle images to meet the usage requirements of the application scenario.
[0071] The multi-angle image generation method of the embodiment of the present application covers the whole process design from image data preprocessing to data set construction, and then to multi-angle image generation and image quality optimization. It generates multi-angle images by using a generation model based on a diffusion process, and achieves a high degree of consistency in the generated images at different viewing angles. Through refined data preprocessing and post-processing optimization operations, it meets the needs of high-precision and high-quality image generation.
[0072] In some embodiments of the present application, the preprocessing of the input image to obtain the preprocessed image includes: removing the background of the input image to obtain an image after the background is removed; extracting the subject of the image after the background is removed to obtain a subject image; performing fine processing on the subject image to obtain a finely processed image, wherein the fine processing includes at least one of cropping, scaling, and adding margins.
[0073] Combination Figure 2 , provides a schematic diagram of an image preprocessing process in an embodiment of the present application. When preprocessing the original image, the original image can be first subjected to background removal processing to obtain an image after background removal. Background removal helps to highlight the main object in the image, reduce the interference of background information, and improve the accuracy and efficiency of subsequent image processing. Then, semantic segmentation and other technologies are used to extract the subject in the image after background removal. The subject extraction further focuses on the key objects or areas in the image to remove unnecessary details and noise.
[0074] The extracted subject image can then be further refined, for example, by cropping, scaling, and adding margins. The purpose of cropping is to remove unnecessary edge portions of the image so that the subject object is more centered or meets specific size requirements. The purpose of scaling the image proportionally is to increase the proportion of the subject in the image by adjusting the size of the image to meet the requirements of subsequent multi-angle image generation. In addition, margins can be added to the processed subject image by adding blank margins around the subject to keep the image size ratio to meet specific output format requirements.
[0075] Through the above preprocessing operations, we can obtain an image with a clean background and a clear and centered subject, which serves as the basis for the subsequent construction of an image generation dataset.
[0076] Through the above preprocessing process, the input image is converted into image data that is more refined, accurate and meets the subsequent input requirements, providing more reliable and high-quality input for subsequent multi-angle image generation tasks.
[0077] In some embodiments of the present application, constructing an image generation dataset based on the preprocessed image includes: digitally processing the preprocessed image to obtain image tensor data; copying the image tensor data into multiple copies according to multiple angles to be generated, and forming the image generation dataset with the multiple copies of image tensor data.
[0078] Combination Figure 3 , a schematic diagram of a process for constructing an image generation data set in an embodiment of the present application is provided. In the stage of constructing the image generation data set, the preprocessed image can be further converted into a digital form that can be processed by the multi-angle image generation model, that is, image tensor data. Image tensor data is usually a multidimensional array structure that contains pixel value information of the image, and these pixel values may represent color, brightness or other image features.
[0079] According to the number of image angles that need to be generated in the multi-angle image generation scenario, the image tensor data obtained above is copied into multiple copies. For example, if the target images of the front, left side, right side, left oblique 45 degree side, right oblique 45 degree side, and back need to be generated, the image tensor data needs to be copied into six copies. These copied multiple copies of the image tensor data are combined to form a data set for subsequent multi-angle image generation tasks.
[0080] In some embodiments of the present application, the digitizing the preprocessed image to obtain image tensor data includes: converting the preprocessed image into an image pixel array; normalizing the image pixel values in the image pixel array to obtain a normalized image pixel array; adding a white background to the image according to the normalized image pixel array, and combining the images using Alpha channel parameters to obtain a combined image; and converting the combined image into the image tensor data.
[0081] When digitizing the preprocessed image, the preprocessed image can be first converted into a NumPy pixel array, in which each element represents a pixel in the image, and its value may represent information such as color and brightness. Then the image pixel values in the NumPy pixel array are normalized. The purpose of normalization is to scale the pixel values to a specific range to eliminate the differences between different images due to factors such as lighting and contrast, which helps to improve the effect of subsequent multi-angle image generation.
[0082] Since the image background is removed in the preprocessing stage, it is also necessary to add a pure white background to the image and combine the images using the Alpha channel parameters. The Alpha channel is usually used to represent the transparency of the image. In this step, a new background layer (pure white) is created for the image, and the Alpha channel value is adjusted to recombine the image with the background layer, thereby achieving the purpose of changing its background while maintaining the image content.
[0083] Finally, the combined image is converted into image tensor data. In this step, the tools provided by traditional deep learning frameworks (such as TensorFlow, PyTorch, etc.) can be used to convert NumPy pixel arrays into image tensor data for subsequent processing.
[0084] The embodiment of the present application can eliminate the illumination and contrast differences between images through normalization processing, thereby improving data quality and consistency, and helping to improve the performance of subsequent models. By adding a white background to the image and recombining the image using the alpha channel parameter, the background of the image subject can be changed to avoid the original background information from interfering with the subsequent multi-angle image generation.
[0085] In some embodiments of the present application, generating initial multiple-angle images using a multi-angle image generation model based on the image generation dataset and the multi-angle category information includes: obtaining preset camera angle embedding information of the multi-angle image generation model during the training phase; transforming the preset camera angle embedding information into a corresponding category label to obtain multi-angle category information; and generating initial multiple-angle images using the multi-angle image generation model based on the image generation dataset and the multi-angle category information.
[0086] Combination Figure 4 , a schematic diagram of the generation process of a multi-angle image generation model in an embodiment of the present application is provided. The generation process of the multi-angle image generation model in the embodiment of the present application is similar to the image generation process of the traditional diffusion model, but it has its own unique features. When the multi-angle image generation model is used to generate multi-angle images, the embodiment of the present application uses a preset camera angle embedding, and the camera angle embedding is transformed according to the preset camera angle when the multi-angle image generation model is trained. The transformed camera angle embedding is input into the multi-angle image generation model as a category label, thereby guiding the image generation process and realizing the generation of multi-angle images.
[0087] In some embodiments of the present application, the post-processing includes at least one of dehazing, denoising and super-resolution upscaling.
[0088] Taking into account that the initial multi-angle images generated by the multi-angle image generation model may have problems such as unclearness, the embodiment of the present application can further adopt a series of post-processing operations such as dehazing, denoising, and super-resolution amplification to adjust and optimize the initially generated multi-angle images, thereby further improving the quality of the multi-angle images.
[0089] Dehazing is designed to reduce or eliminate the blur and low contrast caused by fog in images. Through dehazing, a clearer image can be restored and image details can be more visible. In the field of deep learning, for example, convolutional neural network models such as DehazeNet and AOD-Net can be used to achieve automatic dehazing.
[0090] Denoising aims to remove noise from images, such as blur, noise, and texture, to improve image clarity and quality. This can be achieved through traditional filtering methods (such as mean filtering, median filtering, etc.) or deep learning techniques (such as convolutional neural network denoising models).
[0091] Super-resolution upscaling aims to losslessly upscale low-resolution images to high resolution while maintaining or restoring image detail information. This can be achieved through models such as super-resolution convolutional neural networks (SRCNN) in deep learning technology. These models can restore high-quality image details by learning the mapping relationship between low-resolution images and high-resolution images.
[0092] The embodiments of the present application can significantly improve the overall quality of the generated multi-angle images, improve the visual effects of the images, and provide strong support for subsequent applications of multi-angle images by applying post-processing steps such as dehazing, denoising, and super-resolution amplification to the multi-angle generation task of the images.
[0093] In some embodiments of the present application, the multi-angle image generation model is an MVDiffusion model.
[0094] The MVDiffusion model provides a new approach to solving the problem of multi-angle image generation. By jointly modeling the distribution of images from different perspectives, the model has significant advantages in multi-angle generation tasks. Compared with traditional GAN or VAE, the MVDiffusion model better captures the correlation between perspectives during the generation process and can generate more coherent and consistent multi-angle images.
[0095] The core of the MVDiffusion model is a UNet architecture consisting of multiple network modules on four feature pyramid levels on both sides of the encoder / decoder. Each UNet block processes the input features and learns 3D consistency through the following mechanism:
[0096] 1) Global self-attention mechanism: applied between UNet features of all images to learn global 3D structural information;
[0097] 2) Cross-attention mechanism: The CLIP embedding of the conditional image is injected into all other images through the CLIP model to enhance the correlation between the conditional image and the generated image;
[0098] 3) CNN layer: When processing the features of each image, it injects the time step frequency encoding Embedding and the image index learnable Embedding to capture the temporal information and the differences between images.
[0099] Based on this, the embodiment of the present application applies the MVDiffusion model to the task of multi-angle image generation, and inputs the image generation dataset and multi-angle category labels constructed in the previous embodiment into the UNet of the multi-angle image generation model for multi-angle image generation, making full use of the advantages of the MVDiffusion model to ensure that the image generation process of each angle can be coordinated with other angles, avoiding the problem of perspective inconsistency in traditional methods.
[0100] In summary, the key points of the multi-angle image generation method of the present application are:
[0101] 1) The MVDiffusion model is applied to multi-view image generation. A new solution is proposed to solve the problem of multi-view image generation consistency in current technology. It not only breaks through the bottleneck of existing technology, but also provides a more efficient and reliable way to achieve high-quality, multi-view consistent image generation.
[0102] 2) In the process design of the multi-angle generation method, this application covers the whole process design from data preprocessing to data set construction, to multi-angle image generation and image quality optimization. The data preprocessing part ensures that the input image has high quality, the data set construction part ensures the accuracy of the input data, and the image generation stage uses the advantages of the MVDiffusion model to ensure that the image generation process of each angle can be coordinated with other angles to avoid the problem of incoherent perspective in traditional methods; the final image quality optimization part further adjusts and improves the generated image to ensure that the final output image achieves the ideal visual effect.
[0103] 3) In terms of perspective consistency optimization, this application embeds the camera angle and inputs it into UNet as a category label, which greatly solves the problem of geometric and visual differences between images at different angles. In this way, the generated images not only maintain a high degree of consistency in details, but also show extremely high fidelity at different perspectives. In this way, whether observed from the front, side or back, the generated images can maintain the coherence and authenticity of the object structure, significantly improving the overall quality of multi-perspective generation.
[0104] The multi-angle image generation method of the present application has achieved at least the following technical effects:
[0105] 1) In terms of perspective consistency, by introducing the MVDiffusion model, we have successfully achieved high consistency in generating images from different perspectives. Compared with the prior art, this application significantly solves the common perspective inconsistency problem in traditional methods. The generated images can maintain coherence and consistency at all angles, both in terms of geometry and visual features. This breakthrough has greatly improved the performance of multi-perspective image generation in application scenarios, especially in applications that require the same object to be displayed from multiple angles, providing more accurate and consistent image generation capabilities.
[0106] 2) In terms of image quality, the images generated by this application show a very high level, especially in detail processing and texture expression compared with traditional methods. Whether it is image clarity, color reproduction, or the realism of detailed textures, there are great improvements. Through refined data construction and optimization, this application can generate more visually realistic and expressive images, meeting the needs of high-precision, high-quality image generation.
[0107] The present application also provides a multi-angle image generation device 500. Figure 5 As shown, a schematic diagram of the structure of a multi-angle image generation device in an embodiment of the present application is provided, wherein the multi-angle image generation device 500 includes: a pre-processing unit 510, a construction unit 520, a generation unit 530 and a post-processing unit 540, wherein:
[0108] A preprocessing unit 510 is used to obtain an input image and preprocess the input image to obtain a preprocessed image;
[0109] A construction unit 520, configured to construct an image generation data set based on the preprocessed image;
[0110] A generating unit 530, configured to generate initial multiple-angle images using a multi-angle image generating model according to the image generating data set and the multi-angle category information, wherein the multi-angle image generating model is a generating model based on a diffusion process;
[0111] The post-processing unit 540 is used to post-process the initial multiple angle images to obtain multiple angle images corresponding to the input image.
[0112] In some embodiments of the present application, the preprocessing unit 510 is specifically used to: perform background removal on the input image to obtain an image after background removal; perform subject extraction on the image after background removal to obtain a subject image; perform refinement processing on the subject image to obtain a refined image, wherein the refinement processing includes at least one of cropping, scaling, and adding margins.
[0113] In some embodiments of the present application, the generation unit 530 is specifically used to: digitize the preprocessed image to obtain image tensor data; copy the image tensor data into multiple copies according to multiple angles to be generated, and form the image generation data set with the multiple copies of image tensor data.
[0114] In some embodiments of the present application, the generation unit 530 is specifically used to: convert the preprocessed image into an image pixel array; normalize the image pixel values in the image pixel array to obtain a normalized image pixel array; add a white background to the image according to the normalized image pixel array, and combine the images using Alpha channel parameters to obtain a combined image; and convert the combined image into the image tensor data.
[0115] In some embodiments of the present application, the generation unit 530 is specifically used to: obtain the preset camera angle embedding information of the multi-angle image generation model during the training phase; transform the preset camera angle embedding information into a corresponding category label to obtain multi-angle category information; and generate initial multiple-angle images using the multi-angle image generation model based on the image generation data set and the multi-angle category information.
[0116] In some embodiments of the present application, the post-processing includes at least one of dehazing, denoising and super-resolution upscaling.
[0117] In some embodiments of the present application, the multi-angle image generation model is an MVDiffusion model.
[0118] It can be understood that the above-mentioned multi-angle image generating device can implement each step of the multi-angle image generating method provided in the aforementioned embodiment, and the relevant explanations about the multi-angle image generating method are applicable to the multi-angle image generating device, which will not be repeated here.
[0119] Figure 6 Schematic diagram of the structure of a device in the embodiment of the present application. Figure 6 As shown, the device includes one or more processors (or processing units), may further include one or more memories coupled to the processors, and may further include a communication module coupled to the processors.
[0120] The communication module can be used to communicate with other devices or apparatuses, such as the transmission or reception of data and / or signals. The communication module can have at least one communication module for communication. The communication module can include any interface necessary for communicating with other devices. Exemplarily, the communication module can be a transceiver, a circuit, a bus, a module, or other types of communication modules.
[0121] The processor may include, but is not limited to, at least one of the following: a general-purpose computer, a special-purpose computer, a microcontroller, a digital signal controller (DSP), or one or more of a controller-based multi-core controller architecture. The device may have multiple processors, such as application-specific integrated circuit chips, which are time-dependent and synchronized with a clock of a main processor.
[0122] The memory may include one or more non-volatile memories and one or more volatile memories. Examples of non-volatile memories include, but are not limited to, at least one of the following: read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, hard disk, compact disc (CD), digital video disc (DVD), or other magnetic storage and / or optical storage. Examples of volatile memories include, but are not limited to, at least one of the following: random access memory (RAM), or other volatile memories that do not persist during the duration of a power outage.
[0123] The computer program includes computer executable instructions executed by an associated processor. The program can be stored in ROM. The processor can perform any suitable actions and processes by loading the program into RAM.
[0124] The possible implementation of the present application can be implemented by means of a program, so that the communication device can perform any process discussed in the above embodiments. The possible implementation of the present application can also be implemented by hardware or by a combination of software and hardware.
[0125] In some embodiments, the program may be tangibly contained in a computer-readable storage medium, which may be included in the device (such as in a memory) or other storage device accessible by the device. The program may be loaded from the computer-readable storage medium to the RAM for execution. The computer-readable storage medium may include any type of tangible non-volatile memory, such as ROM, EPROM, flash memory, hard disk, CD, DVD, etc.
[0126] The present application embodiment also provides a computer-readable storage medium, on which computer instructions or program codes are stored, and when the processor runs the instructions or the program codes, the processor executes the methods and functions involved in any of the above embodiments. Computer-readable media can be any tangible medium containing or storing programs for or related to instruction execution systems, devices or equipment. Computer-readable media can be computer-readable signal media or computer-readable storage media. Computer-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or devices, or any suitable combination thereof. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrations. More detailed examples of computer-readable storage media include electrical connections with one or more wires, magnetic media (e.g., disks, floppy disks, hard disks, tapes, magnetic storage devices), optical media (e.g., optical storage devices, DVDs), semiconductor media (e.g., solid-state hard drives), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), or any suitable combination thereof, etc.
[0127] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The embodiment of the present application also provides at least one computer program product tangibly stored on a non-temporary computer-readable storage medium. The computer program product includes one or more computer executable instructions, such as instructions included in a program module, which are executed in a device on a real or virtual processor of the target to perform the process, method and function involved in any of the above embodiments. When the computer program instruction is loaded and executed on a computer, a process or function according to an embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instruction can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instruction can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center.
[0128] The present application embodiment also proposes a computer program product, including a computer program or instruction, when the computer program or instruction is run on a computer, the computer is made to perform the process, method and function in the above-mentioned embodiment. Usually, a program module includes routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or realize specific abstract data types. In various embodiments, the functions of program modules can be combined or divided between program modules as needed. Machine executable instructions for program modules can be executed in local or distributed devices. In distributed devices, program modules can be located in local and remote storage media.
[0129] In general, various embodiments of the present application may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software, which may be performed by a controller, microprocessor, or other computing device. Although various aspects of the embodiments of the present disclosure are shown and described as block diagrams, flow charts, or using some other graphical representations, it should be understood that the boxes, devices, systems, techniques, or methods described herein may be implemented as, for example, non-limiting examples, hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.
[0130] It should be noted that although the embodiments of the present application are described above in conjunction with the accompanying drawings, the above embodiments are not independent of each other, and they can also be combined to obtain other embodiments. The division of the modes, situations, categories and embodiments in the embodiments of the present application is only for the convenience of description and should not constitute a special limitation. The features in the various modes, categories, situations and embodiments can be combined with each other in a logical manner. The various implementation methods of the present application can be combined arbitrarily to achieve different technical effects. The embodiments of the present application no longer list various combinations.
[0131] In addition, although the operation of the method of the present disclosure is described in a particular order in the accompanying drawings, this does not require or imply that these operations must be performed in this particular order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flow chart can change the order of execution. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution. It should also be noted that the features and functions of two or more devices according to the present disclosure can be embodied in one device. Conversely, the features and functions of a device described above can be further divided into being embodied by multiple devices.
[0132] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0133] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A multi-angle image generation method, characterized in that: The multi-angle image generation method comprises: Acquire an input image and preprocess the input image to obtain a preprocessed image; constructing an image generation data set based on the preprocessed images; Generate initial multiple-angle images using a multi-angle image generation model according to the image generation data set and the multi-angle category information, wherein the multi-angle image generation model is a generation model based on a diffusion process; The initial multiple angle images are post-processed to obtain multiple angle images corresponding to the input image.
2. The multi-angle image generation method according to claim 1, characterized in that: The preprocessing of the input image to obtain the preprocessed image comprises: Performing background removal on the input image to obtain an image after background removal; Extracting the main body of the image after the background is removed to obtain a main body image; The main image is refined to obtain a refined image, wherein the refined processing includes at least one of cropping, scaling, and adding a margin.
3. The multi-angle image generation method according to claim 1, characterized in that: The constructing an image generation data set based on the preprocessed image comprises: Digitally processing the preprocessed image to obtain image tensor data; The image tensor data is copied into multiple copies according to multiple angles to be generated, and the multiple copies of image tensor data constitute the image generation data set.
4. The multi-angle image generation method according to claim 3, characterized in that: The digital processing of the preprocessed image to obtain image tensor data comprises: Converting the preprocessed image into an image pixel array; Normalizing the image pixel values in the image pixel array to obtain a normalized image pixel array; Add a white background to the image according to the normalized image pixel array, and combine the images using the Alpha channel parameter to obtain a combined image; The combined image is converted into the image tensor data.
5. The multi-angle image generation method according to claim 1, characterized in that: The step of generating the initial multiple-angle images using the multi-angle image generation model according to the image generation data set and the multi-angle category information includes: Obtaining preset camera angle embedding information of the multi-angle image generation model during the training phase; Transforming the preset camera angle embedding information into corresponding category labels to obtain multi-angle category information; The multi-angle image generation model is used to generate initial multi-angle images according to the image generation data set and the multi-angle category information.
6. The multi-angle image generation method according to claim 1, characterized in that: The post-processing includes at least one of defogging, denoising and super-resolution upscaling.
7. The multi-angle image generation method according to any one of claims 1 to 6, characterized in that: The multi-angle image generation model is an MVDiffusion model.
8. A multi-angle image generation device, characterized in that: The multi-angle image generating device comprises: A preprocessing unit, used to obtain an input image and preprocess the input image to obtain a preprocessed image; A construction unit, configured to construct an image generation data set based on the preprocessed image; A generating unit, configured to generate initial multiple-angle images using a multi-angle image generating model according to the image generating data set and the multi-angle category information, wherein the multi-angle image generating model is a generating model based on a diffusion process; The post-processing unit is used to post-process the initial multiple angle images to obtain multiple angle images corresponding to the input image.
9. A device comprising: processor; and a memory arranged to store computer executable instructions, wherein when the executable instructions are executed, the processor executes the multi-angle image generation method according to any one of claims 1 to 7.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the multi-angle image generation method according to any one of claims 1 to 7 is implemented.