Multi-view image generation method, apparatus, device, medium and product
Patent Information
- Application Number
- CN202611098222.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-08-18
AI Technical Summary
[0004]有鉴于此,本申请提供了一种多视角图像生成方法、装置、设备、介质及产品,以解决生成的多视角图像存在不一致的问题
[0010] The multi-view image generation method provided in this application reconstructs a 3D scene from original images from multiple perspectives, reconstructing the 3D scene corresponding to the original image. A 3D model is then added to the 3D scene and back-projected onto the image planes of each original perspective to form intermediate images. This editing method within the 3D scene ensures geometric consistency among the intermediate images obtained after back-projection, eliminating issues such as coordinate offset and pose drift, effectively avoiding image shaking problems found in related AIGC methods. Based on the intermediate images, texture replacement of the 3D model is performed using the image generation model, resulting in a final target image that not only has geometric consistency but also contains diverse visual content, achieving a visual realism close to that of a real photograph.
Smart Images

Figure CN122597684A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, specifically to multi-view image generation methods, apparatus, devices, media, and products. Background Technology
[0002] Multi-view datasets are an important data foundation when training vision-language-action (VLA) models. However, due to the limited amount of multi-view data, it is generally necessary to expand it.
[0003] Related solutions can augment data based on image generation models, but this approach is prone to breaking multi-view consistency constraints and cannot efficiently edit real-world scenes to augment the dataset while maintaining multi-view geometric consistency. This results in poor quality of the generated new multi-view images, which are unsuitable as training data. Summary of the Invention
[0004] In view of this, this application provides a method, apparatus, device, medium, and product for generating multi-view images to solve the problem of inconsistency in the generated multi-view images.
[0005] In a first aspect, this application provides a multi-view image generation method, the method comprising: Acquire raw images from multiple perspectives; Based on the original images from multiple perspectives, a 3D scene is reconstructed to determine the first scene point cloud of the 3D scene; A 3D model is added to the 3D scene, and the pose of the 3D model is optimized. Determine the second scene point cloud; the second scene point cloud includes the first scene point cloud and the model point cloud corresponding to the pose-optimized 3D model. The second scene point cloud is back-projected onto the image plane corresponding to each viewpoint to generate an intermediate image from multiple viewpoints; Based on the image generation model, texture replacement is performed on the portion of the image corresponding to the 3D model in the intermediate image from each viewpoint to generate a multi-view target image.
[0006] Secondly, this application provides a multi-view image generation apparatus, the apparatus comprising: The acquisition module is used to acquire raw images from multiple perspectives; The scene reconstruction module is used to reconstruct a three-dimensional scene based on the original images from multiple perspectives and determine the first scene point cloud of the three-dimensional scene. The pose optimization module is used to add a 3D model to the 3D scene and optimize the pose of the 3D model. The processing module is used to determine the second scene point cloud; the second scene point cloud includes the first scene point cloud and the model point cloud corresponding to the pose-optimized 3D model; the second scene point cloud is back-projected onto the image plane corresponding to each viewpoint to generate an intermediate image with multiple views; The image generation module is used to perform texture replacement on the portion of the image corresponding to the 3D model in the intermediate image from each viewpoint based on the image generation model, so as to generate a multi-view target image.
[0007] Thirdly, this application provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the multi-view image generation method described in the first aspect or any corresponding embodiment.
[0008] Fourthly, this application provides a computer-readable storage medium storing computer instructions for causing a computer to execute the multi-view image generation method described in the first aspect or any corresponding embodiment thereof.
[0009] Fifthly, this application provides a computer program product, including computer instructions for causing a computer to execute the multi-view image generation method described in the first aspect or any corresponding embodiment thereof.
[0010] The multi-view image generation method provided in this application reconstructs a 3D scene from original images from multiple perspectives, reconstructing the 3D scene corresponding to the original image. A 3D model is then added to the 3D scene and back-projected onto the image planes of each original perspective to form intermediate images. This editing method within the 3D scene ensures geometric consistency among the intermediate images obtained after back-projection, eliminating issues such as coordinate offset and pose drift, effectively avoiding image shaking problems found in related AIGC methods. Based on the intermediate images, texture replacement of the 3D model is performed using the image generation model, resulting in a final target image that not only has geometric consistency but also contains diverse visual content, achieving a visual realism close to that of a real photograph. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0012] Figure 1This is a schematic diagram illustrating an application scenario according to an embodiment of this application; Figure 2 This is a schematic flowchart of a first method for generating multi-view images according to an embodiment of this application; Figure 3 This is a schematic diagram of a second process for a multi-view image generation method according to an embodiment of this application; Figure 4 This is a structural block diagram of a multi-view image generation apparatus according to an embodiment of this application; Figure 5 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] It is understood that before using the technical solutions disclosed in the various embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0015] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0016] In research and model training for robot manipulation tasks, high-quality multi-view datasets are a crucial foundation for training vision-language-action models. Existing dataset augmentation methods mainly fall into the following categories: (1) Image generation method based on AIGC (Artificial Intelligence Generated Content): New scene images can be directly generated using diffusion models or generative adversarial networks. While these methods can generate diverse visual content, they suffer from significant image flickering, meaning that images generated between adjacent frames or different viewpoints lack consistency in texture, lighting, and geometric details. Furthermore, AIGC generation methods struggle to guarantee spatial alignment across multiple viewpoints; that is, AIGC-generated images cannot guarantee cross-view spatial consistency. This makes the generated data unsuitable for direct use in robot training tasks requiring multi-view consistency, thus failing to meet the needs of multi-view training data.
[0017] (2) Methods based on two-dimensional (2D) image editing: Techniques such as image inpainting and image transformation (e.g., image-to-image translation) are used to perform local editing on a single image (e.g., replacing objects or modifying textures). However, these methods can only modify a single image and cannot automatically synchronize the changes to other viewpoints within the same scene, thus violating multi-view consistency constraints. Furthermore, the editing results of a single image cannot be automatically propagated to other viewpoints; performing multi-view expansion processing image by image is inefficient and results in poor consistency.
[0018] (3) Data generation method based on pure three-dimensional (3D) simulation: Scenes are constructed and multi-view images are rendered in a virtual simulation environment. While this method can ensure consistency across multiple views, there is a significant domain gap between the rendered results and real-world data. That is, there is a difference between the data distribution generated by the virtual simulation (source domain) and the data distribution acquired by the camera in the real physical world (target domain). For example, there are obvious differences between simulated rendered images and real images in terms of texture, lighting, and materials. This leads to a performance degradation when the model trained on the simulation data is transferred to the real robot, affecting the model's generalization ability.
[0019] In addition, view synthesis can be performed based on 3D reconstruction. The basic process is as follows: 3D reconstruction is performed using multi-view images to obtain scene point clouds. After editing at the point cloud level, the images are projected back to each viewpoint to generate new images. This method usually only focuses on novel view synthesis. The multi-view images generated after point cloud editing may have problems such as physical distortion or poor visual realism, such as unnatural phenomena like objects floating or passing through objects.
[0020] As one optional application scenario in the embodiments of this application, such as Figure 1 As shown, application 101 is installed in terminal device 110, and user 130 can interact with application 101 through terminal device 110 and / or access device of terminal device 110.
[0021] For example, application 101 can be arbitrary. For instance, application 101 could be an application that supports 3D modeling. Figure 1 In the application scenario shown, if application 101 is active, the terminal device 110 can display the interface 102 of application 101. The interface 102 may include various pages that application 101 can provide, such as interactive pages, settings pages, query pages, etc.
[0022] In some embodiments, terminal device 110 is communicatively connected to server 120 to provide services to application 101. Terminal device 110 may be a mobile terminal, fixed terminal, or portable terminal, including but not limited to mobile phones, desktop computers, laptop computers, multimedia tablets, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 may also support any type of interface, and server 120 may be various types of computing systems or servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, and computing devices in cloud environments.
[0023] It should be noted that, Figure 1 This is merely an example of an application scenario and does not limit the scope of protection of this application.
[0024] This application provides a multi-view image generation method that reconstructs a 3D scene from original images from multiple perspectives, reconstructing the 3D scene corresponding to the original image. A 3D model is then added to the 3D scene and back-projected onto the image planes of each original perspective to form intermediate images. This editing method within the 3D scene ensures geometric consistency among the back-projected intermediate images, eliminating issues such as coordinate offset and pose drift, effectively avoiding image shaking problems found in related AIGC methods. Based on the intermediate images, texture replacement of the 3D model is performed using an image generation model, resulting in a final target image that not only has geometric consistency but also contains diverse visual content, achieving a visual realism close to that of a real photograph.
[0025] According to an embodiment of this application, a multi-view image generation method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0026] This embodiment provides a multi-view image generation method, which can be used in the aforementioned terminal device or server. Figure 2This is a flowchart of a multi-view image generation method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps.
[0027] Step S201: Obtain original images from multiple perspectives.
[0028] The multi-view image generation method provided in this embodiment can generate new multi-view images based on old multi-view images, thereby expanding the multi-view image; for ease of description, the old image is referred to as the original image.
[0029] A multi-view original image consists of multiple original images taken from different perspectives, with common areas shared between them; these original images can be, for example, RGB images. For instance, images of the same region can be acquired based on multiple images from different perspectives to generate a set of multi-view images corresponding to that region, i.e., multi-view original images. It can be understood that a set of multi-view original images contains multiple original images, for example, 3 or 5 original images.
[0030] For example, a robot contains three cameras: a head camera, a left-hand fisheye camera, and a right-hand fisheye camera. These three cameras can simultaneously capture images of a certain area, thus obtaining a raw image containing three perspectives.
[0031] Step S202: Reconstruct the 3D scene based on the original images from multiple perspectives to determine the first scene point cloud of the 3D scene.
[0032] In this embodiment, the multi-view original images are formed by capturing a three-dimensional space using cameras from different perspectives. Based on these multi-view original images, three-dimensional scene reconstruction can also be performed, reconstructing the captured three-dimensional space, i.e., the three-dimensional scene, from multiple two-dimensional original images. Furthermore, based on each original image from the multi-view perspective, the positional information (e.g., three-dimensional coordinates) of each point in the three-dimensional scene can be derived, thereby determining the scene point cloud in the three-dimensional scene, i.e., the first scene point cloud.
[0033] It can identify entities contained in original images from multiple perspectives, such as people, trees, and walls in the image. Through 3D scene reconstruction, it can determine the point cloud of the surface of each entity, thereby forming the first scene point cloud.
[0034] Optionally, in order to determine the point cloud of the first scene, it is generally necessary to use the camera parameters corresponding to each viewpoint. Therefore, before performing the three-dimensional scene reconstruction in step S202 above, the method further includes: determining the camera parameters corresponding to each viewpoint in the three-dimensional scene; the camera parameters are used to perform three-dimensional scene reconstruction on the original image and to perform back projection on the point cloud of the second scene.
[0035] In this embodiment, since the original images from multiple perspectives are acquired by each camera, the camera parameters corresponding to each perspective camera can be determined simultaneously when acquiring the original images. These camera parameters include camera intrinsic parameters and camera extrinsic parameters (such as camera pose). Therefore, the pre-recorded camera parameters can be directly obtained before 3D scene reconstruction.
[0036] Alternatively, camera parameters for each viewpoint can be derived from the original images from multiple perspectives.
[0037] For example, camera intrinsic and extrinsic parameters for each viewpoint can be determined through calibration. Alternatively, feature matching based on the Structure from Motion (SfM) method can be used to accurately estimate camera pose and determine camera intrinsic and extrinsic parameters for each viewpoint.
[0038] After determining the camera parameters for each viewpoint, the dense point cloud of the 3D scene can be reconstructed based on the Multiple View Stereo (MVS) and Neural Radiance Fields (NeRF) methods, forming the first scene point cloud.
[0039] Step S203: Add a 3D model to the 3D scene and optimize the pose of the 3D model.
[0040] In this embodiment, a pre-established 3D asset library is provided, containing multiple 3D models. These 3D models can be cup models, chair models, mobile phone models, vehicle models, etc., depending on the actual situation. After reconstructing the 3D scene, at least one 3D model is added to the 3D scene, and the pose of the added 3D model is optimized to ensure that the added 3D model matches the original 3D scene and avoids physical errors.
[0041] For example, if the original multi-view image is a captured image of an indoor table, then after 3D scene reconstruction, a 3D scene containing the table can be obtained, and the first scene point cloud includes the point cloud of the table. At this point, a cup model can be added to this 3D scene and placed on the upper surface of the table. It is necessary to adjust the pose of the cup model to ensure that the cup model fits snugly against the upper surface of the table and avoid physical errors such as clipping or floating.
[0042] Step S204: Determine the second scene point cloud; the second scene point cloud includes the first scene point cloud and the model point cloud corresponding to the pose-optimized 3D model.
[0043] For a 3D model with optimized pose, the 3D coordinates of its surface can also be determined, thus forming a point cloud. For example, the 3D coordinates of each vertex of the 3D model in the 3D scene can be determined, and each vertex can be mapped to a point in the point cloud to form the model point cloud.
[0044] Furthermore, by overlaying the original first scene point cloud with the model point cloud of the 3D model, a second scene point cloud can be formed. This second scene point cloud can represent the positions of existing objects in the 3D model (such as the table in the example above), as well as the positions of newly added 3D models (such as the cup model in the example above).
[0045] Step S205: Backproject the second scene point cloud onto the image plane corresponding to each viewpoint to generate an intermediate image with multiple views.
[0046] In this embodiment, after determining the second scene point cloud containing the newly added 3D model, the second scene point cloud is back-projected onto the image plane corresponding to each viewpoint, thereby generating a multi-view image, i.e., an intermediate image. It can be understood that the intermediate image contains the newly added 3D model.
[0047] This requires back-projection based on camera parameters from various viewpoints. Specifically, the second scene point cloud is a 3D point cloud in the world coordinate system. For any viewpoint, each 3D point cloud can be transformed to the camera coordinate system of that viewpoint based on the camera extrinsic parameters. Then, based on the camera intrinsic parameters, it can be projected from the camera coordinate system to the camera's two-dimensional image plane. Finally, based on the color information of the 3D point cloud, the corresponding pixels in the two-dimensional image plane are assigned values to form the intermediate image of that viewpoint.
[0048] In this embodiment, since the second scene point cloud has a new 3D model with optimized pose, after back-projecting it to each viewpoint, an intermediate image containing the 3D model can be formed. The 3D model in each intermediate image conforms to the pose under the corresponding viewpoint, and the 3D model under multiple viewpoints has geometric consistency. That is, the position and pose of the 3D model are accurate under different viewpoints, and there are no problems such as coordinate offset or pose drift. That is, there are no problems such as 3D model misalignment or shaking between intermediate images of each viewpoint.
[0049] Step S206: Based on the image generation model, perform texture replacement on the portion of the image corresponding to the 3D model in the intermediate image from each viewpoint to generate a multi-view target image.
[0050] As shown above, the generated intermediate image is a 2D view of the 3D model added to the original image; furthermore, the generated multi-view intermediate images maintain geometric consistency. Based on this intermediate image, the image generation model only performs texture replacement on the portion of the image corresponding to the 3D model. This operation does not change the geometric pose of the 3D model in the image, ensuring that the images generated by the image generation model maintain geometric consistency across different viewpoints. For ease of description, the image generated based on the image generation model is referred to as the target image.
[0051] Furthermore, by replacing the texture of the 3D model, such as changing the color and pattern of the 3D model, more target images with different textures can be generated based on the intermediate image, which can increase the number of multi-view images generated.
[0052] The image generation model can be a diffusion model or a large language model (LLM) with image generation capabilities; this embodiment does not limit the specific model.
[0053] It is understandable that the 3D model itself can also have a certain texture. Therefore, in the back projection process of step S205 above, the corresponding pixels of the 3D model in each intermediate image can also be assigned the corresponding color (each vertex of the 3D model can have color attributes). This makes the generated intermediate image itself a newly generated multi-view image, which can also be used for subsequent model training. Therefore, the multi-view intermediate image can also be used as a multi-view target image for subsequent use, but the realism may be slightly worse. For example, the multi-view intermediate image can also be added to the training set as a training sample for model training.
[0054] The multi-view image generation method provided in this embodiment reconstructs a 3D scene from the original images from multiple perspectives, reconstructing the 3D scene corresponding to the original image. A 3D model is then added to the 3D scene and back-projected onto the image planes of each original viewpoint to form intermediate images. This editing method within the 3D scene ensures geometric consistency among the intermediate images obtained after back-projection, eliminating issues such as coordinate offset and pose drift, effectively avoiding image shaking problems found in related AIGC methods. Based on the intermediate images, texture replacement of the 3D model is performed using the image generation model, resulting in a final target image that not only has geometric consistency but also contains diverse visual content, achieving a visual realism close to that of a real photograph.
[0055] This embodiment provides a multi-view image generation method, which can be used in the aforementioned terminal device or server. Figure 3 This is a flowchart of a multi-view image generation method according to an embodiment of this application, such as... Figure 3 As shown, the process includes the following steps.
[0056] Step S301: Obtain the original images from multiple perspectives.
[0057] Please see details Figure 2 Step S201 of the illustrated embodiment will not be described again here.
[0058] Step S302: Reconstruct the three-dimensional scene based on the original images from multiple perspectives to determine the first scene point cloud of the three-dimensional scene.
[0059] Please see details Figure 2 Step S202 of the illustrated embodiment will not be described again here.
[0060] Step S303: Add a 3D model to the 3D scene and optimize the pose of the 3D model.
[0061] Specifically, step S303, "Optimize the pose of the 3D model", includes steps S3031 to S3032.
[0062] Step S3031: Establish physical constraints between the 3D model and the point cloud of the first scene; the physical constraints include collision constraints.
[0063] In this embodiment, in order to ensure that the first scene point cloud in the three-dimensional scene and the newly added three-dimensional model conform to the physical laws, a collision constraint condition is established between the two. This collision constraint condition is used to constrain the three-dimensional model and the first scene point cloud to avoid the clipping problem between them.
[0064] In some optional implementations, the above step S3031 "establishing physical constraints between the 3D model and the point cloud of the first scene" may specifically include steps a1 to a3.
[0065] Step a1: Determine the directed distance field of the 3D model.
[0066] Step a2: Determine the collision distance between the first scene point cloud and the surface of the 3D model based on the directed distance field.
[0067] Step a3: Construct collision constraints to minimize the collision distance.
[0068] In this embodiment, based on the pose of the 3D model in the 3D scene, the directed distance field (SDF) of the 3D model can be determined. The directed distance field represents the distance from any position in the 3D scene (e.g., the position in the point cloud of the first scene) to the nearest surface of the 3D model, and the distance is given a sign to indicate whether it is inside or outside the 3D model, generally a positive or negative sign.
[0069] For example, for any location in a 3D scene The location can be determined based on the directed distance field. Signed distance to the nearest surface of the 3D model Among them, if This indicates that the position Outside the 3D model; if This indicates that the position On the surface of the three-dimensional model; if This indicates that the position Inside the 3D model, a collision problem exists.
[0070] After determining the directed distance field, the collision distance between any point in the first scene point cloud and the surface of the 3D model can be determined based on the directed distance field. The collision distance can be determined based on the symbolic distance mentioned above. If the symbolic distance is less than 0, it means that the point in the first scene point cloud has penetrated into the interior of the 3D model, that is, there is a collision problem between the two. In this case, the symbolic distance can be used as the collision distance. If the symbolic distance is greater than or equal to 0, it means that there is no collision problem between the two. In this case, the collision distance can be set to 0, or the symbolic distance can be set as the collision distance. There is no limitation here.
[0071] The collision distance between the first scene point cloud and the surface of the 3D model is minimized. That is, the collision constraint is used to minimize the penetration between the two, thereby penalizing all penetration points, so that there is no penetration or collision between the pose-optimized 3D model and the original objects in the 3D scene (i.e., the objects corresponding to the first scene point cloud, such as the desktop, walls, etc.).
[0072] Optionally, the aforementioned physical constraints may also include height constraints and / or gravity constraints. Furthermore, step S3031, "establishing physical constraints between the 3D model and the first scene point cloud," may also include step b1 and / or step b2.
[0073] Step b1: Determine the height difference between the bottom of the 3D model and the support surface point cloud used to place the 3D model in the first scene point cloud; construct a height constraint condition to minimize the height difference.
[0074] Step b2: Determine the directional deviation between the height direction of the 3D model and the normal direction of the point cloud of the supporting surface; construct gravity constraints to minimize the directional deviation.
[0075] In this embodiment, in order to ensure the pose optimization effect of the 3D model, in addition to setting collision constraints, height constraints and gravity constraints can also be set.
[0076] Specifically, the 3D model needs to be placed in a 3D scene, that is, the 3D scene has a supporting surface for placing the 3D model. The point cloud corresponding to the supporting surface can be determined based on the point cloud of the first scene. The supporting surface can be, for example, the ground, a tabletop, etc., depending on the actual situation.
[0077] It can calculate the height difference between the bottom of a 3D model and the point cloud of the supporting surface. Specifically, it can sample the bottom of the 3D model and the point cloud of the supporting surface separately. For any sampling point at the bottom of the 3D model, it determines the minimum distance between that sampling point and any sampling point in the point cloud of the supporting surface. The sum of the minimum distances (or the average of the minimum distances) of all sampling points at the bottom of the 3D model is taken as the height difference between the two.
[0078] In this embodiment, the height constraint condition is to minimize the height difference. By minimizing the height difference, the bottom of the 3D model after pose optimization can be in close contact with the support surface (such as a desktop), thus achieving the effect of placing the 3D model on the support surface.
[0079] Furthermore, the height direction of a 3D model can also be determined. For example, if the 3D model itself has a Z-axis, the orientation of this Z-axis in the 3D scene can be used as the height direction of the 3D model. Additionally, the supporting surface point cloud has normals, and the corresponding normal direction can be determined. Generally, this normal direction is basically parallel to the Z-axis of the 3D scene.
[0080] By minimizing the directional deviation between the height direction of the 3D model and the normal direction of the point cloud of the supporting surface, the corresponding gravity constraint conditions can be constructed, so that the 3D model after pose optimization conforms to the gravity constraint and will not violate physical laws, such as placing the 3D model on a vertical wall.
[0081] Step S3032: Solve the pose of the 3D model according to the physical constraints to determine the pose that meets the physical constraints.
[0082] In this embodiment, based on the above physical constraints, the pose (including position and orientation, a total of six dimensions) of the three-dimensional model can be solved, thereby optimizing the six-degree-of-freedom pose of the three-dimensional model so that the optimized pose of the three-dimensional model conforms to the physical constraints.
[0083] For example, physical constraints include collision constraints, height constraints, and gravity constraints, and corresponding loss functions can be set.
[0084] in, The loss value corresponding to the collision constraint can be determined based on the collision distance between the point cloud of the first scene and the surface of the 3D model; The loss value corresponding to the height constraint can be determined by the height difference between the bottom of the 3D model and the point cloud of the support surface; The loss value corresponding to the gravity constraint can be determined based on the directional deviation between the height direction of the 3D model and the normal direction of the point cloud of the support surface. , , The weighting coefficients are the corresponding losses.
[0085] The aforementioned loss function can be minimized using a global optimization algorithm to achieve pose optimization. For example, a double annealing optimization algorithm can be used to iteratively update the pose of the 3D model and calculate the total loss. The process continues until the total loss is minimized, for example, when the change in total loss is less than a preset value, or when the maximum number of iterations is reached. Optimization based on the double annealing algorithm is not only fast but also helps escape local optima, thus facilitating the finding of the globally optimal pose.
[0086] Step S304: Determine the second scene point cloud; the second scene point cloud includes the first scene point cloud and the model point cloud corresponding to the pose-optimized 3D model.
[0087] Please see details Figure 2 Step S204 of the illustrated embodiment will not be described again here.
[0088] Step S305: Backproject the second scene point cloud onto the image plane corresponding to each viewpoint to generate an intermediate image with multiple views.
[0089] Please see details Figure 2 Step S205 of the illustrated embodiment will not be described again here.
[0090] In some optional implementations, the above step S305, "backprojecting the second scene point cloud onto the image plane corresponding to each viewpoint to generate an intermediate image with multiple views", may include steps c1 to c2.
[0091] Step c1: For any viewpoint, backproject the second scene point cloud onto the image plane corresponding to the viewpoint, and determine the projection coordinates and depth values of each point cloud.
[0092] Step c2: Generate an intermediate image of the viewpoint based on the point cloud corresponding to the minimum depth value under the same projection coordinates.
[0093] When generating an image by back-projecting a scene point cloud, erroneous occlusion may occur. For example, objects further back relative to the camera may occlude objects further forward. In this embodiment, to correctly handle the occlusion relationships in the second scene point cloud, for any viewpoint, when back-projecting the second scene point cloud onto the image plane corresponding to the viewpoint, the projection coordinates (i.e., the two-dimensional coordinates of the projection points of the point cloud on the image plane in the corresponding two-dimensional coordinate system) and depth values (i.e., the distance between the point cloud and the camera at that viewpoint) of each point cloud in the second scene point cloud are determined.
[0094] If multiple point clouds have the same projection coordinates, it means that these point clouds have an occlusion relationship. In this case, the point cloud corresponding to the minimum depth value is determined, that is, the point cloud closest to the camera. Based on these point clouds, an intermediate image corresponding to this viewpoint is generated, which can ensure that the occlusion relationship of each object (including 3D models) in the generated intermediate image is correct.
[0095] For example, for a certain viewpoint, a depth buffer can be set up for its output intermediate image to record the minimum depth value corresponding to each pixel (each pixel corresponds to a projection coordinate). Initially, it can be empty or a maximum value. Iterate through each 3D point cloud to be projected in the second scene point cloud to determine the projection coordinates and depth value corresponding to each point cloud. Based on the projection coordinates of the point cloud, compare them with the currently recorded minimum depth value in the depth buffer. If it is less than the currently recorded minimum depth value, it means that the point cloud is a more foreground element, so update the minimum depth value in the depth buffer and cover the corresponding pixels in the intermediate image with the color corresponding to that point cloud; otherwise, discard the point cloud, that is, do not update the depth buffer or the rendered image.
[0096] Through the above processing, the depth sorting of each point cloud can be achieved, which can handle the occlusion relationship between the foreground (such as an inserted 3D model) and the background (such as the original objects in the 3D scene) and avoid visual errors such as clipping or inversion in the intermediate image.
[0097] Optionally, before the subsequent step S306 "replace the texture of the part of the image corresponding to the 3D model in the intermediate image from each viewpoint based on the image generation model", the method further includes: smoothing the edge region of the 3D model in the intermediate image to obtain an updated intermediate image; the updated intermediate image is used to generate the target image.
[0098] Due to the accuracy limitations of dense point clouds in 3D reconstruction, problems such as geometric discontinuities at the junction of the 3D model and the original scene may occur. After backprojection, artifacts such as jagged edges, color banding, and color abrupt changes may appear. In this embodiment, after backprojecting the second scene point cloud to generate an intermediate image, the edge regions of the 3D model in the intermediate image are smoothed to weaken or eliminate edge artifacts in the intermediate image. Then, step S306 is executed based on the updated intermediate image to generate the target image.
[0099] This process can involve color interpolation or linear blending of the edge regions of the 3D model in the intermediate image. For example, a weighted average of the RGB values of multiple adjacent pixels at the edge can be used to smooth out harsh color boundaries; alternatively, gradient weights can be applied to the edge regions to smoothly transition from the main object's color to the background color, reducing the sense of segmentation. Through smoothing, edge artifacts caused by insufficient reconstruction accuracy can be eliminated, improving the visual effect of the intermediate image without altering the underlying 3D geometric features.
[0100] Step S306: Based on the image generation model, perform texture replacement on the portion of the image corresponding to the 3D model in the intermediate image from each viewpoint to generate a multi-view target image.
[0101] Please see details Figure 2 Step S206 of the illustrated embodiment will not be described again here.
[0102] In some optional implementations, the above step S306, "replacing the texture of the part of the image corresponding to the 3D model in the intermediate image of each viewpoint based on the image generation model to generate a multi-view target image", may include steps d1 to d2.
[0103] Step d1: Input the object texture map and intermediate images from each viewpoint into the image generation model, and instruct the image generation model to perform texture replacement on the part of the image corresponding to the 3D model in the intermediate image of each viewpoint according to the object texture map.
[0104] Step d2: Generate multi-view target images based on the output of the image generation model.
[0105] In this embodiment, an object texture library can be pre-set, which contains multiple object texture maps. For the generated intermediate image (e.g., the updated intermediate image), it can be synchronously input into the image generation model along with the object texture map, so that the image generation model performs texture replacement on the part of the image corresponding to the 3D model in the intermediate image from each viewpoint according to the object texture map, that is, replaces the texture of the 3D model in the intermediate image with the texture corresponding to the object texture map.
[0106] For example, the image generation model could be a large language model, capable of replacing the texture of a 3D model based on texture replacement prompts. For instance, if the 3D model is a bottle, the prompt could include: "Replace the texture of the inserted bottle with the texture shown in the image."
[0107] In this embodiment, by combining 3D scene reconstruction and back projection to generate multi-view intermediate images with geometric consistency, and then using AI image generation technology, the realism of the final generated target image can be guaranteed, making up for the shortcomings of high cost and inflexibility of pure 3D modeling texture modification.
[0108] Optionally, if texture replacement is performed directly on each viewpoint based on the image generation model, there may be local visual inconsistencies. To further improve the visual effect of the image, the above step d2, "generating target images from multiple viewpoints based on the output of the image generation model", may include steps d21 to d25.
[0109] Step d21: Determine the multi-view images to be determined based on the output of the image generation model.
[0110] Step d22: For the first undetermined image corresponding to the first viewpoint, determine the undetermined region in the first undetermined image that corresponds to the 3D model.
[0111] Step d23: Determine the local scene point cloud corresponding to the region to be determined.
[0112] Step d24: Project the local scene point cloud onto the image plane corresponding to the second viewpoint to generate a reference image corresponding to the second viewpoint.
[0113] Step d25: Adjust the second undetermined image corresponding to the second viewpoint based on the reference image to obtain the target image corresponding to the second viewpoint.
[0114] In this embodiment, for ease of description, the image directly output by the image generation model is referred to as the undetermined image. For example, by inputting an object texture map and intermediate images from various viewpoints into the image generation model, the image generation model can output undetermined images from each viewpoint.
[0115] For a certain undetermined image from a specific viewpoint, namely the first undetermined image corresponding to the first viewpoint (e.g., the undetermined image corresponding to the viewpoint of the robot's head camera), the undetermined region corresponding to the 3D model can be located through image difference or segmentation model. This undetermined region serves as the subsequent editing region, which can be, for example, a rectangular region.
[0116] For the region to be determined, the associated point cloud, i.e., the local scene point cloud, can be determined. For example, a portion of the point cloud corresponding to the region to be determined can be determined from the second scene point cloud to form a local scene point cloud. The pixels of the region to be determined in the first image can be remapped back to the world coordinate system to generate a local incremental point cloud, which can also be used as a local scene point cloud.
[0117] It is understandable that, for this local scene point cloud, since it corresponds to the undetermined area of the first undetermined image, the attributes (such as color) of each point cloud in this local scene point cloud are consistent with the pixel attributes of the undetermined area.
[0118] The local scene point cloud is then projected onto the image plane corresponding to the second viewpoint. This second viewpoint differs from the first viewpoint, thus generating a reference image for the second viewpoint. This reference image is visually consistent with the region to be determined in the first image to be determined, and it mainly includes a 3D model. An image generation model can then be used to adjust the second image to be determined based on the reference image. This allows for local detail adjustments to the second image to ensure that the final target image from the second viewpoint is visually identical to the target image from the first viewpoint, conforming to the true imaging effect of the lens and improving multi-view visual consistency.
[0119] The generated target image can be directly added to the training set as a training sample; alternatively, the generated target image can be validated. Specifically, the validation process may include: Multi-view geometric consistency check: The reprojection error of corresponding points between different views is lower than the threshold. Visual consistency check: The photometric error between adjacent viewpoint images is below a threshold; Physical rationality verification: The inserted 3D model has no abnormal phenomena such as penetration or floating.
[0120] If the multi-view target images pass the above geometric consistency check, visual consistency check, and physical rationality check, the multi-view target images will be used as the multi-view dataset for training.
[0121] The multi-view image generation method provided in this embodiment, through a process of 3D reconstruction, adding 3D models, and back projection, fundamentally ensures geometric consistency across multiple perspectives, avoiding image jitter issues common in AIGC methods. Utilizing various physical constraints to optimize the pose of the 3D model ensures that the inserted 3D model adheres tightly to the support surface without penetration, exhibiting significantly better physical plausibility than methods such as random placement or simple gravity-based drop. Further local detail optimization of the images generated by the image generation model compensates for the shortcomings of pure geometric methods in terms of texture and lighting, ensuring that the visual realism of the generated images closely approximates that of real photographs. This method supports various editing operations such as adding and replacing 3D models, and modifying texture attributes, and can be driven by natural language commands, allowing for rapid expansion of multi-view image datasets without requiring specialized 3D modeling skills.
[0122] This embodiment also provides a multi-view image generation apparatus, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0123] This embodiment provides a multi-view image generation device, such as... Figure 4 As shown, the device includes: The acquisition module 401 is used to acquire original images from multiple perspectives; The scene reconstruction module 402 is used to reconstruct a three-dimensional scene based on the original images from multiple perspectives and determine the first scene point cloud of the three-dimensional scene. The pose optimization module 403 is used to add a three-dimensional model to the three-dimensional scene and optimize the pose of the three-dimensional model. The processing module 404 is used to determine the second scene point cloud; the second scene point cloud includes the first scene point cloud and the model point cloud corresponding to the pose-optimized 3D model; the second scene point cloud is back-projected onto the image plane corresponding to each viewpoint to generate an intermediate image with multiple views. The image generation module 405 is used to perform texture replacement on the portion of the image corresponding to the three-dimensional model in the intermediate image from each viewpoint based on the image generation model, so as to generate a multi-view target image.
[0124] In some optional implementations, the pose optimization of the 3D model includes: Establish physical constraints between the 3D model and the point cloud of the first scene; the physical constraints include collision constraints. The pose of the 3D model is solved based on the physical constraints to determine the pose that meets the physical constraints.
[0125] In some optional implementations, establishing the physical constraints between the 3D model and the first scene point cloud includes: Determine the directed distance field of the three-dimensional model; The collision distance between the first scene point cloud and the surface of the 3D model is determined based on the directed distance field; Construct collision constraints to minimize the collision distance.
[0126] In some optional implementations, the physical constraints may further include height constraints and / or gravity constraints; The physical constraints established between the 3D model and the first scene point cloud include: Determine the height difference between the bottom of the 3D model and the support surface point cloud in the first scene point cloud used to place the 3D model; construct a height constraint condition to minimize the height difference; And / or, Determine the directional deviation between the height direction of the 3D model and the normal direction of the point cloud of the supporting surface; construct a gravity constraint condition to minimize the directional deviation.
[0127] In some optional implementations, the step of back-projecting the second scene point cloud onto the image plane corresponding to each viewpoint to generate an intermediate image with multiple views includes: For any viewpoint, the second scene point cloud is back-projected onto the image plane corresponding to the viewpoint, and the projection coordinates and depth values corresponding to each point cloud are determined. An intermediate image of the viewpoint is generated based on the point cloud corresponding to the minimum depth value under the same projected coordinates.
[0128] In some optional implementations, before performing texture replacement on the portions of the intermediate images corresponding to the 3D model in each viewpoint based on the image generation model, the processing module 404 is further configured to: The edge regions of the 3D model in the intermediate image are smoothed to obtain an updated intermediate image; the updated intermediate image is used to generate the target image.
[0129] In some optional implementations, the step of performing texture replacement on the portion of the image corresponding to the 3D model in the intermediate images from each viewpoint based on the image generation model to generate a multi-view target image includes: The object texture map and the intermediate images from each viewpoint are input into the image generation model, and the image generation model is instructed to perform texture replacement on the portion of the image corresponding to the 3D model in the intermediate images from each viewpoint according to the object texture map; Generate multi-view target images based on the output of the image generation model.
[0130] In some optional implementations, generating a multi-view target image based on the output of the image generation model includes: The multi-view images to be determined are determined based on the output of the image generation model; For the first undetermined image corresponding to the first perspective, determine the undetermined region in the first undetermined image that corresponds to the three-dimensional model; Determine the local scene point cloud corresponding to the region to be determined; The local scene point cloud is projected onto the image plane corresponding to the second viewpoint to generate a reference image corresponding to the second viewpoint. The second undetermined image corresponding to the second viewpoint is adjusted based on the reference image to obtain the target image corresponding to the second viewpoint.
[0131] In some optional implementations, the scene reconstruction module 402 is further configured to: Determine the camera parameters corresponding to each viewpoint in the 3D scene; the camera parameters are used to reconstruct the 3D scene from the original image and to back-project the point cloud of the second scene.
[0132] The multi-view image generation apparatus provided in this disclosure can execute the multi-view image generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.
[0133] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0134] The following is a detailed reference. Figure 5The diagram illustrates a structural schematic suitable for implementing the electronic device described in the embodiments of this application. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from memory 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device. The processor 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0135] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0136] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a memory 508, or installed from a ROM 502. When the computer program is executed by the processor 501, it performs the functions defined in the multi-view image generation method of embodiments of this application.
[0137] Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0138] This application also provides a computer-readable storage medium. The methods described in this application can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the multi-view image generation method shown in the above embodiments is implemented.
[0139] A portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0140] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and all such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A method for generating multi-view images, characterized in that, The method includes: Acquire raw images from multiple perspectives; Based on the original images from multiple perspectives, a 3D scene is reconstructed to determine the first scene point cloud of the 3D scene; A 3D model is added to the 3D scene, and the pose of the 3D model is optimized. Determine the second scene point cloud; the second scene point cloud includes the first scene point cloud and the model point cloud corresponding to the pose-optimized 3D model. The second scene point cloud is back-projected onto the image plane corresponding to each viewpoint to generate an intermediate image from multiple viewpoints; Based on the image generation model, texture replacement is performed on the portion of the image corresponding to the 3D model in the intermediate image from each viewpoint to generate a multi-view target image.
2. The method according to claim 1, characterized in that, The pose optimization of the 3D model includes: Establish physical constraints between the 3D model and the point cloud of the first scene; the physical constraints include collision constraints. The pose of the 3D model is solved based on the physical constraints to determine the pose that meets the physical constraints.
3. The method according to claim 2, characterized in that, The physical constraints established between the 3D model and the first scene point cloud include: Determine the directed distance field of the three-dimensional model; The collision distance between the first scene point cloud and the surface of the 3D model is determined based on the directed distance field; Construct collision constraints to minimize the collision distance.
4. The method according to claim 2 or 3, characterized in that, The physical constraints also include height constraints and / or gravity constraints; The physical constraints established between the 3D model and the first scene point cloud include: Determine the height difference between the bottom of the 3D model and the point cloud of the support surface used to place the 3D model in the first scene point cloud; Construct height constraints to minimize the height difference; And / or, Determine the directional deviation between the height direction of the 3D model and the normal direction of the point cloud of the supporting surface; construct a gravity constraint condition to minimize the directional deviation.
5. The method according to claim 1, characterized in that, The step of back-projecting the second scene point cloud onto the image plane corresponding to each viewpoint to generate an intermediate image from multiple viewpoints includes: For any viewpoint, the second scene point cloud is back-projected onto the image plane corresponding to the viewpoint, and the projection coordinates and depth values corresponding to each point cloud are determined. An intermediate image of the viewpoint is generated based on the point cloud corresponding to the minimum depth value under the same projected coordinates.
6. The method according to claim 1, characterized in that, Before performing texture replacement on the portions of the intermediate images corresponding to the 3D model from each viewpoint based on the image generation model, the method further includes: The edge regions of the 3D model in the intermediate image are smoothed to obtain an updated intermediate image; the updated intermediate image is used to generate the target image.
7. The method according to claim 1, characterized in that, The step of performing texture replacement on the portion of the image corresponding to the 3D model in the intermediate images from each viewpoint based on the image generation model to generate a multi-view target image includes: The object texture map and the intermediate images from each viewpoint are input into the image generation model, and the image generation model is instructed to perform texture replacement on the portion of the image corresponding to the 3D model in the intermediate images from each viewpoint according to the object texture map; Generate multi-view target images based on the output of the image generation model.
8. The method according to claim 7, characterized in that, The step of generating a multi-view target image based on the output of the image generation model includes: The multi-view images to be determined are determined based on the output of the image generation model; For the first undetermined image corresponding to the first perspective, determine the undetermined region in the first undetermined image that corresponds to the three-dimensional model; Determine the local scene point cloud corresponding to the region to be determined; The local scene point cloud is projected onto the image plane corresponding to the second viewpoint to generate a reference image corresponding to the second viewpoint. The second undetermined image corresponding to the second viewpoint is adjusted based on the reference image to obtain the target image corresponding to the second viewpoint.
9. The method according to claim 1, characterized in that, The method further includes: Determine the camera parameters corresponding to each viewpoint in the 3D scene; the camera parameters are used to reconstruct the 3D scene from the original image and to back-project the point cloud of the second scene.
10. A multi-view image generation device, characterized in that, The device includes: The acquisition module is used to acquire raw images from multiple perspectives; The scene reconstruction module is used to reconstruct a three-dimensional scene based on the original images from multiple perspectives and determine the first scene point cloud of the three-dimensional scene. The pose optimization module is used to add a 3D model to the 3D scene and optimize the pose of the 3D model. The processing module is used to determine the second scene point cloud; the second scene point cloud includes the first scene point cloud and the model point cloud corresponding to the pose-optimized 3D model; the second scene point cloud is back-projected onto the image plane corresponding to each viewpoint to generate an intermediate image with multiple views; The image generation module is used to perform texture replacement on the portion of the image corresponding to the 3D model in the intermediate image from each viewpoint based on the image generation model, so as to generate a multi-view target image.
11. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the multi-view image generation method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the multi-view image generation method according to any one of claims 1 to 9.
13. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the multi-view image generation method according to any one of claims 1 to 9.