Image generation method and device
By obtaining images and camera parameters from different perspectives, fitting the parameterized mesh model and defining the local coordinate system and NeRF model, the problem of inaccurate local rendering effect of 3D virtual objects is solved, and high-quality image generation is achieved.
Patent Information
- Application Number
- CN202411006977.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-07-25
AI Technical Summary
The prior art is not realistic enough in the local rendering effect of 3D virtual objects, especially in the case of dynamic movement, resulting in poor quality of generated images.
By obtaining at least 2 images of different viewing angles, each image contains corresponding camera parameter information, fitting to generate a parameterized grid model, defining a local coordinate system and a NeRF model, and fusing the output results of multiple NeRF models to generate a target image.
The quality of the target generated images is effectively improved, especially in the case of dynamic movement, and the local rendering effect is significantly improved.
Smart Images

Figure CN118887340B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image generation, and in particular to an image generation method and device. Background Art
[0002] At present, 3D virtual object synthesis is becoming more and more common, but the existing technology is not realistic enough for local rendering, especially when the rendered local area moves dynamically, the local rendering effect is poor, so it needs to be improved. For example, in the process of generating the hand area of the human body, the relative position between the fingers can change at any time, resulting in poor display effect of the generated hand area. Summary of the invention
[0003] The present invention provides an image generation method and device, which can effectively improve the quality of target generated image generation.
[0004] According to one aspect of the present invention, there is provided an image generation method, comprising:
[0005] Obtain at least two images of different viewing angles corresponding to the target to be generated, wherein each image contains corresponding camera parameter information;
[0006] Generating a parameterized grid model based on the image fitting;
[0007] Defining N local coordinate systems based on the N vertices of the parameterized mesh model, where N is an integer greater than or equal to 2;
[0008] Determine the local position information of the position point in the three-dimensional space in each of the local coordinate systems based on the N local coordinate systems; wherein the three-dimensional space is a space containing the target to be generated;
[0009] Based on each of the N local coordinate systems, a NeRF model is defined respectively;
[0010] A target generation image is generated based on at least two of the NeRF models and the local position information corresponding to the local coordinate system.
[0011] According to another aspect of the present invention, there is provided an image generating device, comprising:
[0012] An image acquisition module is used to acquire at least two images of different viewing angles corresponding to the target to be generated, wherein each image contains corresponding camera parameter information;
[0013] A parameterized grid model generation module, used for generating a parameterized grid model based on the image fitting;
[0014] A local coordinate system definition module, used to define N local coordinate systems based on N vertices of the parameterized mesh model, where N is an integer greater than or equal to 2;
[0015] A local position information determination module, used to determine the local position information of a position point in the three-dimensional space in each of the local coordinate systems based on the N local coordinate systems; wherein the three-dimensional space is a space containing the target to be generated;
[0016] A NeRF model definition module, used for defining a NeRF model based on each of the N local coordinate systems;
[0017] The target generated image generation module is used to generate a target generated image based on at least two of the NeRF models and the local position information corresponding to the local coordinate system.
[0018] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0019] at least one processor; and
[0020] a memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the image generating method described in any embodiment of the present invention.
[0022] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the image generation method described in any embodiment of the present invention when executed.
[0023] The image generation scheme of the embodiment of the present invention obtains at least two images of different perspectives corresponding to the target to be generated, wherein each image contains corresponding camera parameter information; a parameterized mesh model is generated based on the image fitting; N local coordinate systems are defined based on the N vertices of the parameterized mesh model, wherein N is an integer greater than or equal to 2; local position information of position points in the three-dimensional space of the N local coordinate systems in each of the local coordinate systems; wherein the three-dimensional space is a space containing the target to be generated; a NeRF model is defined based on each local coordinate system in the N local coordinate systems; and a target generation image is generated based on at least two of the NeRF models and the local position information corresponding to the local coordinate systems. By fusing the output results of at least two NeRF models, the quality of target generation image generation can be effectively improved.
[0024] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0026] Figure 1 is a flowchart of an image generating method provided according to Embodiment 1 of the present invention;
[0027] Figure 2 It is a schematic diagram of the principle of collecting images in an image set provided by an embodiment of the present invention;
[0028] Figure 3 is a flowchart of an image generating method provided according to Embodiment 2 of the present invention;
[0029] Figure 4 is a flowchart of an image generating method provided according to Embodiment 3 of the present invention;
[0030] Figure 5 is a structural schematic diagram of an image generating device provided according to a fourth embodiment of the present invention;
[0031] Figure 6 It is a schematic diagram of the structure of an electronic device for implementing the image generating method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0032] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0034] Embodiment 1
[0035] Figure 1 A flowchart of an image generation method is provided for the first embodiment of the present invention. This embodiment is applicable to the case of generating an image. The method can be executed by an image generation device. The image generation device can be implemented in the form of hardware and / or software. The image generation device can be configured in an electronic device. Figure 1 As shown, the method includes:
[0036] S110, obtaining at least two images of different viewing angles corresponding to the target to be generated, wherein each image contains corresponding camera parameter information.
[0037] In an embodiment of the present invention, an image acquisition device is controlled to acquire images of a target to be generated at different shooting angles, so that at least two different visual images corresponding to the target to be generated are acquired through the image acquisition device. Among them, the image acquisition device may include various digital cameras, smart phones, portable wearable devices, drones and other devices with optical image information acquisition functions. Among them, each image contains corresponding camera parameter information, and the camera parameter information is the observation position coordinates and observation direction of the image acquisition device. The target to be generated can be a human body, a local area of a human body, other animals, or target objects such as vehicles. It should be noted that the embodiment of the present invention does not limit the target to be generated.
[0038] For example, Figure 2 The following is a schematic diagram of the principle of collecting images in an image set provided by an embodiment of the present invention. Figure 2 As shown in the figure, each image in the image set is taken in a global coordinate system (world coordinate system) with a camera at a fixed position and in a fixed field of view. The coordinates of the camera position can be called the viewpoint, and the RGB of each point in the image is the projection result of the RGB and transparency of each point on a ray with the viewpoint as the origin on the pixel point of the image.
[0039] S120: Generate a parameterized grid model based on the image fitting.
[0040] In an embodiment of the present invention, at least two images with different visuals acquired in S110 are fitted to generate a parameterized mesh model, wherein the parameterized mesh model is a mesh graph for representing the 3D topological structure of the target to be generated, and the parameterized mesh model may also be referred to as a 3D Mesh. The parameterized mesh model is composed of a plurality of vertex information and triangular face information, wherein a triangular face is a triangular face composed of three vertices, and the vertex information includes the coordinates of each vertex, and the coordinates of each vertex may be expressed as (x, y, z).
[0041] Exemplarily, a parametric model, such as MANO, is used to generate a parametric mesh model based on image fitting in an image set, wherein the parametric mesh model contains vertex information and connection information constituting the parametric mesh model. The vertex information is to use the viewpoint of each image as the origin and to make rays through the key points in the image. The points where the same key points in different images intersect are the position coordinates of the key points in three-dimensional space, that is, the vertices of the parametric mesh model. The connection information can be to connect multiple adjacent points according to rules to form a parametric mesh model, or to connect the key points to form a parametric mesh model after identifying the target to be generated according to image recognition technology.
[0042] Optionally, the generation of a parameterized mesh model based on the image fitting includes: determining a plurality of key points of the target to be generated in each of the images; and fitting and generating a parameterized mesh model of the target to be generated in a three-dimensional space according to the key points of each of the images. In an embodiment of the present invention, a plurality of key points of the target to be generated in each of at least two images of different viewing angles are determined, wherein the key points may be feature points of the target to be generated. For example, if the target to be generated is a human face region, the key points may be facial features, such as the mouth, left eye, nose, right eye, ear, eyebrow, and other feature points; and if the target to be generated is a human hand region, the key points may be feature points of the fingertips, knuckles, palms, and other feature points of each finger. Optionally, a plurality of key points of the target to be generated in each image may be determined by feature point recognition technology. According to the key points of each image, a parameterized mesh model of the target to be generated in a three-dimensional space is fitted and generated. Exemplarily, a parameterized mesh model (such as MANO) can be used to fit the captured 2D photos, and the fitting method is not limited to matching the detected 2D key points of the finger with each mesh vertex. Optionally, the parameterized mesh model is a triangular mesh.
[0043] S130. Define N local coordinate systems based on the N vertices of the parameterized mesh model, where N is an integer greater than or equal to 2.
[0044] Exemplarily, all vertices in the parameterized mesh model are determined, and then a local coordinate system is defined based on each vertex in the parameterized mesh model, wherein the number of local coordinate systems is the same as the number of vertices included in the parameterized mesh model. Another exemplary method is to determine all vertices in the parameterized mesh model, and then randomly select some vertices (at least two) from all vertices in the parameterized mesh model, and define a local coordinate system for each vertex in the selected part of the vertices, wherein the number of local coordinate systems is the same as the number of the selected part of the vertices.
[0045] S140. Determine local position information of a position point in the three-dimensional space in each of the local coordinate systems based on the N local coordinate systems; wherein the three-dimensional space is a space containing the target to be generated.
[0046] In an embodiment of the present invention, a three-dimensional space containing a target to be generated is determined, wherein the three-dimensional space is a space of a global coordinate system (world coordinate system), which is a space of the real world that has been defined when the images in the image set are captured. For example, the spatial region where the model is defined when the image is captured is a cubic meter space with the origin as the vertex, there are 8 cameras, and they are distributed at 8 vertices, the camera coordinates are (0,0,0), (1,0,0), (0,1,0), (0,0,1), (1,1,0), (1,0,1), (0,1,1), (1,1,1), and the viewing direction points to the center of the cube, that is, any point in the three-dimensional model, camera, or space outside the camera is a point that can be defined by the global coordinate system of the three-dimensional space. For each of all images of at least two different viewing angles, the local position information of the position point corresponding to the point in the image in the three-dimensional space in N local coordinate systems is determined respectively.
[0047] Optionally, determining the local position information of a position point in the three-dimensional space in each local coordinate system based on the N local coordinate systems includes: determining the global position information of the corresponding position point in the three-dimensional space in the global coordinate system according to the information of each point in the image; determining the linear transformation relationship between the global coordinate system and the local coordinate system for each local coordinate system in the N local coordinate systems; and converting the global position information of the position point in the three-dimensional space into the local position information of the position point in the local coordinate system based on the linear transformation relationship.
[0048] In an embodiment of the present invention, a global coordinate system is established, and the position information of the corresponding position point in the three-dimensional space in the global coordinate system is determined based on the information of each point in the image. For the convenience of description, the position information of the position point in the global coordinate system is referred to as global position information. It can be understood that the global position information of the position point in the three-dimensional space is the position information in the same global coordinate system. For each local coordinate system in the N local coordinate systems, the linear transformation relationship between the global coordinate system and the local coordinate system is determined, wherein the linear transformation relationship is used to reflect the change relationship between the two position points when the coordinate information of any position point in the global coordinate system is converted to the coordinate information of the corresponding position point in the local coordinate system. According to the linear transformation relationship between the global coordinate system and the local coordinate system, the global position information of the position point in the three-dimensional space is converted into the local position information of the position point in the corresponding local coordinate system. Through the above method, the local position information of the position point corresponding to each point in the image in the three-dimensional space in each local coordinate system can be obtained.
[0049] Exemplarily, the target to be generated is the hand area. During the movement of the human hand, the position and orientation of each vertex are dynamically changed, resulting in that even if the coordinates of a certain point in the local coordinate system remain unchanged, the global coordinates of the point are also transformed (the point in the local coordinate system changes with the vertex). However, assuming that the local coordinate system has a given coordinate point, no matter how the global coordinates of the point change, the features of the point will not change. Assuming that the vertex corresponding to the local coordinate system is the tip of the index finger, the relative position relationship between the index finger nail and the tip of the index finger is fixed. Therefore, the coordinates of the index finger nail in the local coordinate system of the tip of the index finger are unchanged. Even if the index finger moves, causing the global coordinates of the tip of the index finger to change, the features of the given local coordinates will not change. Therefore, in an embodiment of the present invention, N local coordinate systems are defined based on the N vertices of the parameterized mesh model, and the local position information of the position point corresponding to each point in the image in the three-dimensional space in each local coordinate system is determined based on the N local coordinate systems.
[0050] S150 . Based on each of the N local coordinate systems, define a NeRF model respectively.
[0051] In the embodiment of the present invention, a neural radiance field (NeRF) model is defined in each of the N local coordinate systems. It can be understood that the NeRF model corresponds to the local coordinate system one by one, and since there are N local coordinate systems, N corresponding NeRF models can be defined.
[0052] S160, generating a target generation image based on at least two of the NeRF models and the local position information corresponding to the local coordinate system.
[0053] In an embodiment of the present invention, a target generated image is generated according to N NeRF models and local position information of each position point in at least two images with different viewing angles in N local coordinate systems. Optionally, the target generated image is generated based on at least two NeRF models and the local position information corresponding to the local coordinate system, including: generating an initial generated image based on each NeRF model and the local position information corresponding to the local coordinate system; fusing the initial generated images generated by at least two NeRF models to generate a target generated image. For each local coordinate system in the N local coordinate systems, an initial generated image is generated according to the local position information of the position point corresponding to each point in the three-dimensional space in the images with at least two different viewing angles in the local coordinate system and the NeRF model corresponding to the local coordinate system. It can be understood that the local position information of the position point corresponding to each point in the three-dimensional space in the images with at least two different viewing angles in the local coordinate system is three-dimensionally reconstructed through the NeRF model corresponding to the local coordinate system, and the reconstructed image is used as the initial generated image. Exemplarily, the local position information of the position point corresponding to each point in the three-dimensional space in the local coordinate system of at least two images with different viewing angles is input into the NeRF model corresponding to the local coordinate system, and the image output by the NeRF model is used as the initial generated image corresponding to the local coordinate system. In the above manner, the initial generated images corresponding to N NeRF models can be obtained, and then the initial generated images corresponding to the N NeRF models are fused based on the image fusion technology, and the fused image is used as the target generated image. Exemplarily, the initial generated images output by the N NeRF models are linearly combined to generate the target generated image.
[0054] It should be emphasized that the NeRF model is essentially a function, which is related to the spatial position (x, y, z), viewing direction, color information (RGB) and volume density information (transparency). The first two (spatial position (x, y, z) and viewing direction) are input quantities, and the latter two (color information and volume density information) are output quantities. The process of model training is to sample points in the image, determine the coordinates, viewing direction, RGB and transparency of the sampling points, fit the model, and then calculate the loss value of the model through the loss function, and modify the model based on the loss value, and then sample again, and repeat the above process until the desired training model is achieved, for example, the loss value is within an acceptable range. The completion of model training indicates the completion of 3D modeling. After the model is trained, given the observation point and observation angle, the model will use the observation point as the origin to make a ray, and superimpose the RGB and transparency of each point (sampling point) on a ray to generate a pixel point on the initial generated image. Multiple pixels generate an initial generated image, and at least 2 initial generated images generate the target generated image.
[0055] Optionally, the initial generated image is generated based on each of the NeRF models and the local position information corresponding to the local coordinate system, including: for the position point in the three-dimensional space, respectively determining the vertex distance between the position point and the vertex of the parameterized mesh model corresponding to each local coordinate system; selecting from the local coordinate system a first target local coordinate system corresponding to the first target vertex whose vertex distance is less than a first preset distance threshold; and generating the initial generated image based on the NeRF model corresponding to each of the first target local coordinate systems and the local position information corresponding to the first target local coordinate system whose vertex distance is less than the first preset distance threshold.
[0056] In the embodiment of the present invention, since the local coordinate system is a corresponding coordinate system defined based on N vertices of the parameterized mesh model, and the local coordinate system corresponds to the N vertices of the parameterized mesh model one by one, the position points in the three-dimensional space are traversed, and the vertex distances between the position points and the vertices of the parameterized mesh model corresponding to each local coordinate system are calculated respectively. The local coordinate system corresponding to the first target vertex whose vertex distance is less than the first preset distance threshold is selected from the N local coordinate systems as the first target local coordinate system, wherein the first target local coordinate system can be one or more. The initial generated image is generated according to the NeRF model corresponding to each first target local coordinate system and the local position information whose vertex distance is less than the first preset distance threshold in the first target local coordinate system. It can be understood that the NeRF model corresponding to each first local position coordinate system only reconstructs the position points of the local position information whose vertex distance is less than the first preset distance threshold in the first local position coordinate system, thereby obtaining the initial generated position points output by the NeRF model corresponding to each first target local coordinate system, and the image composed of each initial generated position point is used as the initial generated image. It can be understood that the initial generation position points output by the NeRF model corresponding to all the first local coordinate systems are further fused to generate target generation position points, and the image formed by the target generation position points is the target generation image.
[0057] The image generation method of the embodiment of the present invention obtains at least two images of different perspectives corresponding to the target to be generated, wherein each image contains corresponding camera parameter information; a parameterized mesh model is generated based on the image fitting; N local coordinate systems are defined based on the N vertices of the parameterized mesh model, wherein N is an integer greater than or equal to 2; local position information of a position point in a three-dimensional space in each of the local coordinate systems is determined based on the N local coordinate systems; wherein the three-dimensional space is a space containing the target to be generated; a NeRF model is defined based on each of the N local coordinate systems; and a target generation image is generated based on at least two of the NeRF models and the local position information corresponding to the local coordinate systems. By fusing the output results of at least two NeRF models, the quality of the target generation image generation can be effectively improved.
[0058] Embodiment 2
[0059] Figure 3 A flowchart of an image generation method provided in Embodiment 2 of the present invention is shown in FIG. Figure 3 As shown, the method includes:
[0060] S310: Obtain at least two images of different viewing angles corresponding to the target to be generated, wherein each image contains corresponding camera parameter information.
[0061] S320: Generate a parameterized grid model based on the image fitting.
[0062] S330. Define N local coordinate systems based on the N vertices of the parameterized mesh model, where N is an integer greater than or equal to 2.
[0063] S340. Determine the local position information of a position point in the three-dimensional space in each of the local coordinate systems based on the N local coordinate systems; wherein the three-dimensional space is a space containing the target to be generated.
[0064] S350 . Based on each of the N local coordinate systems, define a NeRF model respectively.
[0065] S360: for the position point in the three-dimensional space, respectively determine the vertex distance between the position point and the vertex of the parameterized mesh model corresponding to each local coordinate system.
[0066] S370. Setting a model weight of the NeRF model corresponding to each local coordinate system according to the vertex distance.
[0067] In an embodiment of the present invention, for each position point in the three-dimensional space, a corresponding model weight is set for the NeRF model corresponding to each local coordinate system according to the vertex distance between the position point and the vertex of the parameterized mesh model corresponding to each local coordinate system. Among them, the model weight is negatively correlated with the corresponding vertex distance, that is, the larger the vertex distance, the smaller the corresponding model weight. The sum of the model weights corresponding to all NeRF model settings may be 1 or may not be 1.
[0068] S380. Based on at least two of the NeRF models and the corresponding model weights, fuse the initial generated images generated by at least two of the NeRF models to generate a target generated image.
[0069] In an embodiment of the present invention, at least two NeRF models are selected from N NeRF models according to the model weights. For the convenience of description, the selected at least two NeRF models may be referred to as target NeRF models, and the local position point information is respectively input into each target NeRF model to obtain the initial generated image output by each target NeRF model, and then the initial generated image is fused based on at least two target NeRF models and the corresponding model weights to generate the target generated image. Exemplarily, for a certain position point, the model weight of NeRF model 1 is 90, the model weight of NeRF model 2 is 80, and the model weights of other NeRF models are all less than 10, then NeRF model 1 and NeRF model 2 may be used as target NeRF models. The local position information corresponding to the position point is analyzed based on NeRF model 1 and NeRF model 2 respectively to obtain the initial generated position point corresponding to the position point, and then the initial generated position points output by NeRF model 1 and NeRF model 2 are fused based on the model weights of NeRF model 1 and NeRF model 2 to generate the target generated position point. Similarly, the target generation position point corresponding to each position point can be determined in the above manner, and a target generation image can be formed based on all target generation position points. Optionally, the local position point information can be input into N NeRF models respectively, and the initial generation images output by the N NeRF models are obtained respectively, and then the initial generation images are fused based on the N NeRF models and the corresponding model weights to generate the target generation image.
[0070] It is understandable that, in the embodiment of the present invention, although there are also 8 pictures, if fitting is performed in one model, first of all, the spatial area of fitting is large, and the sampling points are more. If the same accuracy is to be achieved, the amount of calculation is large and the training time is long; secondly, in the dynamic modeling process, each action moment needs to be sampled and trained from scratch (each action will cause the global coordinates to change), the amount of calculation is large and the time is long. The technical solution provided by the embodiment of the present invention, by defining N local coordinate systems and N NeRF models, for example, each finger joint defines a local coordinate system and a corresponding NeRF model, regardless of the finger joint in any position in the three-dimensional space, its position in the local coordinate system is basically unchanged, so when training the local model, the finger joint at the position can have a better reconstruction effect, especially in the dynamic modeling process, the finger joint at the position at each action moment has no significant change, and there is no need to start sampling training from 0, so the amount of training is greatly reduced and the training efficiency is improved. At the same time, each local model only trains points within a certain distance range around the origin, which greatly reduces the amount of training, and at the same time, the local rendering effect can be improved by increasing the sampling density and / or sampling times in the local space.
[0071] The image generation method of the embodiment of the present invention can ensure the locality of each NeRF model and reduce the total computational complexity of image generation, while improving the quality of target generated image generation by fusing the output results of at least two NeRF models.
[0072] Embodiment 3
[0073] Figure 4 A flowchart of an image generation method provided in Embodiment 3 of the present invention is shown in FIG. Figure 4 As shown, the method includes:
[0074] S410: Acquire at least two images of different viewing angles corresponding to the target to be generated, wherein each image contains corresponding camera parameter information.
[0075] S420: Generate a parameterized grid model based on the image fitting.
[0076] S430. Define N local coordinate systems based on the N vertices of the parameterized mesh model, where N is an integer greater than or equal to 2.
[0077] S440. Determine local position information of a position point in the three-dimensional space in each of the local coordinate systems based on the N local coordinate systems; wherein the three-dimensional space is a space containing the target to be generated.
[0078] S450 . Based on each of the N local coordinate systems, define a NeRF model respectively.
[0079] S460: for the position point in the three-dimensional space, respectively determine the vertex distance between the position point and the vertex of the parameterized mesh model corresponding to each local coordinate system.
[0080] In an embodiment of the present invention, since the local coordinate system is a corresponding coordinate system defined based on the N vertices of the parameterized mesh model, and the local coordinate system corresponds one-to-one to the N vertices of the parameterized mesh model, each position point in the three-dimensional space is traversed, and the vertex distance between the position point and the vertex of the parameterized mesh model corresponding to each local coordinate system is calculated respectively.
[0081] S470: Select, from the local coordinate system, a second target local coordinate system corresponding to a second target vertex whose vertex distance is less than a second preset distance threshold.
[0082] In the embodiment of the present invention, the local coordinate system corresponding to the second target vertex whose vertex distance is less than the second preset distance threshold is selected from the N local coordinate systems as the second target local coordinate system, wherein the second target local coordinate system may be one or more. It should be noted that the embodiment of the present invention does not limit the magnitude relationship between the first preset distance threshold and the second preset distance threshold.
[0083] S480: Generate a target generation image based on the NeRF models corresponding to at least two of the second target local coordinate systems and the local position information corresponding to the second target local coordinate system.
[0084] In an embodiment of the present invention, a target generation image is generated based on the NeRF models corresponding to at least two second target local coordinate systems and the local position information in the second target local coordinate system. It can be understood that the NeRF model corresponding to each second local position coordinate system reconstructs the position points of the local position information in the second local position coordinate system, thereby obtaining the initial generation position points output by the NeRF model corresponding to each second target local coordinate system, and then the initial generation position points output by the NeRF model corresponding to all second local coordinate systems are fused to generate the target generation position points, and then the image composed of each target generation position point is used as the target generation image.
[0085] The image generation method of the embodiment of the present invention can ensure the locality of each NeRF model and reduce the total computational complexity of image generation, while improving the quality of target generated image generation by fusing the output results of at least two NeRF models.
[0086] In some embodiments, the determining of the local position information of each position point in the three-dimensional space in each local coordinate system based on the N local coordinate systems also includes: determining the relative local pose information of at least two local coordinate systems based on the N local coordinate systems; or, determining the local pose information of the position point in the three-dimensional space in each local coordinate system based on the N local coordinate systems; the generating of the target image based on at least two NeRF models and the local position information in the corresponding local coordinate systems also includes: each NeRF model uses the local pose information as an additional input signal, and generates the target image based on at least two NeRF models and the local position information and the local pose information in the corresponding local coordinate systems. The advantage of such a setting is that each local NeRF uses the pose as an additional input signal, and can obtain the corresponding relationship between the pose of each position point and the surrounding joint activity angle, so that the subtle changes in the local surface brought about by the target to be generated in different poses can be reconstructed, which can further improve the quality of the target generated image.
[0087] In an embodiment of the present invention, for at least two of the N local coordinate systems, the relative local pose information between the at least two local coordinate systems is determined. For example, the relative pose information between the coordinate origins of the at least two local coordinate systems (including the three coordinate axis directions) can be used as the relative local pose information between the at least two local coordinate systems. Among them, the relative local pose information between the at least two local coordinate systems can be referred to as the first local pose information. Alternatively, for each of all images of at least two different viewing angles, the local pose information of the position point corresponding to each point in the image in the three-dimensional space in N local coordinate systems is determined respectively. Exemplarily, the global pose information of the position point corresponding to each point in the image in the three-dimensional space in the global coordinate system is determined, and according to the linear transformation relationship between the global coordinate system and each local coordinate system, the global pose information of the position point in the three-dimensional space is converted into the local pose information of the position point in the corresponding local coordinate system. Among them, the local pose information of the position point in the three-dimensional space in each local coordinate system can be referred to as the second local pose information. The local pose information (i.e., the first local pose information or the second local pose information) and the local position information are respectively input to at least two NeRF models, and the initial generated images output by the NeRF models are respectively obtained, and the initial generated images corresponding to the at least two NeRF models are fused to generate a target generated image. Exemplarily, taking the first local pose information as an additional input signal of each NeRF model as an example, the two local coordinate systems are the fingertip local coordinate system and the first finger joint local coordinate system. By determining the relative angle information between the fingertip local coordinate system and the first finger joint local coordinate system, the wrinkles of the skin at different joint angles can be obtained, and a training model can be generated to achieve the accurate generation of the local skin of the model at the unsampled joint angle. Another exemplary example is that the target to be generated is the hand area of the human body. By using the local pose information, such as the angle pose between a point on the local skin surface and the surrounding joints as an additional input signal of each NeRF model, the generated target generated image can contain the changes in the local skin surface of the hand caused by different poses, such as the local skin wrinkles when the finger joints are curled, etc., which can further improve the quality of the target generated image. It can be understood that the movement of the target to be generated is continuous, while the sampling is discrete. For example, the moving joints are photographed and sampled at 30 frames per second to generate images at each moment. At this time, the joint angles in the image at each moment are different and discontinuous. In the embodiment of the present invention, the training model can be better generated by additionally inputting local pose information to achieve image generation under continuous joint motion angles.
[0088] Embodiment 4
[0089] Figure 5 This is a schematic diagram of the structure of an image generating device provided by Embodiment 4 of the present invention. Figure 5As shown, the device comprises:
[0090] An image acquisition module 510 is used to acquire at least two images of different viewing angles corresponding to the target to be generated, wherein each image contains corresponding camera parameter information;
[0091] A parameterized grid model generation module 520, configured to generate a parameterized grid model based on the image fitting;
[0092] A local coordinate system definition module 530, configured to define N local coordinate systems based on N vertices of the parameterized mesh model, wherein N is an integer greater than or equal to 2;
[0093] A local position information determination module 540, configured to determine the local position information of a position point in the three-dimensional space in each of the local coordinate systems based on the N local coordinate systems;
[0094] A NeRF model definition module 550, configured to define a NeRF model based on each of the N local coordinate systems;
[0095] The target generated image generation module 560 is used to generate a target generated image based on at least two of the NeRF models and the local position information corresponding to the local coordinate system.
[0096] Optionally, the parameterized grid model generating module is used to:
[0097] Determine a plurality of key points of the target to be generated in each of the images;
[0098] According to the key points of each of the images, a parameterized grid model of the target to be generated in a three-dimensional space is fitted and generated.
[0099] Optionally, the parameterized mesh model is a triangular mesh.
[0100] Optionally, a local position information determination module is used to:
[0101] Determine the global position information of the corresponding position point in the three-dimensional space in the global coordinate system according to the information of each point in the image;
[0102] For each local coordinate system in the N local coordinate systems, determining a linear transformation relationship between the global coordinate system and the local coordinate system;
[0103] The global position information of the position points in the three-dimensional space is converted into the local position information of each position point in the local coordinate system based on the linear transformation relationship.
[0104] Optionally, the target generation image generation module includes:
[0105] An initial generated image generating unit, configured to generate an initial generated image based on each of the NeRF models and the local position information corresponding to the local coordinate system;
[0106] The target generated image generating unit is used to fuse the initial generated images generated by at least two of the NeRF models to generate a target generated image.
[0107] Optionally, the initial image generation unit is used to:
[0108] For the position point in the three-dimensional space, respectively determine vertex distances between the position point and vertices of the parameterized mesh model corresponding to each local coordinate system;
[0109] Selecting from the local coordinate system a first target local coordinate system corresponding to a first target vertex whose vertex distance is less than a first preset distance threshold;
[0110] The initial generated image is generated based on the NeRF model corresponding to each of the first target local coordinate systems and the local position information corresponding to the first target local coordinate system whose vertex distance is less than the first preset distance threshold.
[0111] Optionally, the target generation image generation module is used to:
[0112] For the position point in the three-dimensional space, respectively determine vertex distances between the position point and vertices of the parameterized mesh model corresponding to each local coordinate system;
[0113] Setting a model weight of the NeRF model corresponding to each local coordinate system according to the vertex distance;
[0114] Based on at least two of the NeRF models and the corresponding model weights, initial generated images generated by at least two of the NeRF models are fused to generate a target generated image.
[0115] Optionally, the target generation image generation module is used to:
[0116] For the position point in the three-dimensional space, respectively determine vertex distances between the position point and vertices of the parameterized mesh model corresponding to each local coordinate system;
[0117] Selecting from the local coordinate system a second target local coordinate system corresponding to a second target vertex whose vertex distance is less than a second preset distance threshold;
[0118] A target generation image is generated based on the NeRF models corresponding to at least two of the second target local coordinate systems and the local position information corresponding to the second target local coordinate system.
[0119] Optionally, the local position information determination module is further used to:
[0120] Determine relative local pose information of at least two of the local coordinate systems based on the N local coordinate systems; or,
[0121] Determine the local pose information of the position point in the three-dimensional space in each local coordinate system based on the N local coordinate systems;
[0122] The target generation image generation module is also used to:
[0123] Each of the NeRF models uses the local pose information as an additional input signal, and generates a target image based on at least two of the NeRF models and the local position information and the local pose information in the corresponding local coordinate system.
[0124] The image generating device provided in the embodiment of the present invention can execute the image generating method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0125] Embodiment 5
[0126] Figure 6 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0127] like Figure 6As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0128] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0129] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs the various methods and processes described above, such as an image generation method.
[0130] In some embodiments, the image generation method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the image generation method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform the image generation method in any other appropriate manner (e.g., by means of firmware).
[0131] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0132] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0133] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0134] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0135] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0136] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.
[0137] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.
[0138] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. An image generation method, characterized in that: include: Obtain at least two images of different viewing angles corresponding to the target to be generated, wherein each image contains corresponding camera parameter information; Generating a parameterized grid model based on the image fitting; Defining N local coordinate systems based on the N vertices of the parameterized mesh model, where N is an integer greater than or equal to 2; Determine the local position information of the position point in the three-dimensional space in each of the local coordinate systems based on the N local coordinate systems; wherein the three-dimensional space is a space containing the target to be generated; Based on each of the N local coordinate systems, a NeRF model is defined respectively; Generate a target generation image based on at least two of the NeRF models and the local position information corresponding to the local coordinate system; The step of generating a target image based on at least two of the NeRF models and the local position information corresponding to the local coordinate system includes: Generate an initial generated image based on each of the NeRF models and the local position information corresponding to the local coordinate system; The initial generated images generated by at least two of the NeRF models are fused to generate a target generated image.
2. The method according to claim 1, characterized in that The generating a parameterized grid model based on the image fitting comprises: Determine a plurality of key points of the target to be generated in each of the images; According to the key points of each of the images, a parameterized grid model of the target to be generated in a three-dimensional space is fitted and generated.
3. The method according to claim 2, characterized in that The parameterized mesh model is a triangular mesh.
4. The method according to claim 1, characterized in that: The determining, based on the N local coordinate systems, the local position information of the position point in the three-dimensional space in each local coordinate system comprises: Determine the global position information of the corresponding position point in the three-dimensional space in the global coordinate system according to the information of each point in the image; For each local coordinate system in the N local coordinate systems, determining a linear transformation relationship between the global coordinate system and the local coordinate system; The global position information of the position point in the three-dimensional space is converted into the local position information of the position point in the local coordinate system based on the linear transformation relationship.
5. The method according to claim 1, characterized in that The generating an initial generated image based on each of the NeRF models and the local position information corresponding to the local coordinate system includes: For the position point in the three-dimensional space, respectively determine vertex distances between the position point and vertices of the parameterized mesh model corresponding to each local coordinate system; Selecting from the local coordinate system a first target local coordinate system corresponding to a first target vertex whose vertex distance is less than a first preset distance threshold; The initial generated image is generated based on the NeRF model corresponding to each of the first target local coordinate systems and the local position information corresponding to the first target local coordinate system whose vertex distance is less than the first preset distance threshold.
6. The method according to claim 1, characterized in that The generating a target image based on at least two of the NeRF models and the local position information corresponding to the local coordinate system includes: For the position point in the three-dimensional space, respectively determine vertex distances between the position point and vertices of the parameterized mesh model corresponding to each local coordinate system; Setting a model weight of the NeRF model corresponding to each local coordinate system according to the vertex distance; Based on at least two of the NeRF models and the corresponding model weights, initial generated images generated by at least two of the NeRF models are fused to generate a target generated image.
7. The method according to claim 1, characterized in that The generating a target image based on at least two of the NeRF models and the local position information corresponding to the local coordinate system includes: For the position point in the three-dimensional space, respectively determine vertex distances between the position point and vertices of the parameterized mesh model corresponding to each local coordinate system; Selecting from the local coordinate system a second target local coordinate system corresponding to a second target vertex whose vertex distance is less than a second preset distance threshold; A target generation image is generated based on the NeRF models corresponding to at least two of the second target local coordinate systems and the local position information corresponding to the second target local coordinate system.
8. The method according to any one of claims 1 to 7, characterized in that: The determining of the local position information of the position point in the three-dimensional space in each local coordinate system based on the N local coordinate systems further includes: Determine relative local pose information of at least two of the local coordinate systems based on the N local coordinate systems; or, Determine the local pose information of the position point in the three-dimensional space in each local coordinate system based on the N local coordinate systems; The generating of the target image based on at least two of the NeRF models and the local position information corresponding to the local coordinate system further includes: Each of the NeRF models uses the local pose information as an additional input signal, and generates a target image based on at least two of the NeRF models and the local position information and the local pose information in the corresponding local coordinate system.
9. An image generating device, characterized in that: include: An image acquisition module is used to acquire at least two images of different viewing angles corresponding to the target to be generated, wherein each image contains corresponding camera parameter information; A parameterized grid model generation module, used for generating a parameterized grid model based on the image fitting; A local coordinate system definition module, used to define N local coordinate systems based on N vertices of the parameterized mesh model, where N is an integer greater than or equal to 2; A local position information determination module, used to determine the local position information of a position point in the three-dimensional space in each of the local coordinate systems based on the N local coordinate systems; wherein the three-dimensional space is a space containing the target to be generated; A NeRF model definition module, used for defining a NeRF model based on each of the N local coordinate systems; A target generation image generation module, used for generating a target generation image based on at least two of the NeRF models and the local position information corresponding to the local coordinate system; Wherein, the target generation image generation module includes: An initial generated image generating unit, configured to generate an initial generated image based on each of the NeRF models and the local position information corresponding to the local coordinate system; The target generated image generating unit is used to fuse the initial generated images generated by at least two of the NeRF models to generate a target generated image.
Citation Information
Patent Citations
Human hand image synthesis method and device, electronic equipment and storage medium
CN116758202A
Two-hand 3D reconstruction method based on monocular RGB image occlusion removal
CN117671133A