A design method, device and equipment for volumetric video acquisition system
By introducing shooting coverage evaluation index and secondary screening method of virtual three-dimensional object reconstruction in the volume video acquisition system, the problem of lack of theoretical guiding indicators in the hardware construction of volume video acquisition system in the prior art is solved, and the effect of quickly finding the optimal camera layout mode and saving costs is achieved.
Patent Information
- Application Number
- CN202411017896.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-07-29
AI Technical Summary
The hardware construction of the existing volume video acquisition system lacks theoretical guiding indicators and relies too much on actual reconstruction results for evaluation. The number of cameras and arrangements of solutions are too large, resulting in a significant increase in labor and time costs.
By introducing evaluation indicators for the shooting coverage of the target three-dimensional object, a collection of candidate camera arrangement modes is selected, and a secondary evaluation of the virtual three-dimensional object is performed to further filter out the target camera arrangement mode.
Quickly selecting the optimal camera arrangement mode from the huge solution space greatly saves labor and time costs, and is of guiding significance for the construction of hardware equipment.
Smart Images

Figure CN118827946B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical fields of computer graphics and computer vision, and in particular to a design method, device and equipment for a volumetric video acquisition system. Background Art
[0002] Volumetric video is used to capture the captured object in high-realism and high-quality three-dimensional space. The focus is on obtaining high-precision three-dimensional geometric reconstruction and high-resolution texture information. For example, human body volumetric video can record and reconstruct the dynamic form of a person in three dimensions. As a customized multi-view high-precision video, the acquisition system of volumetric video needs to be updated and deployed according to specific shooting requirements. Therefore, it has high construction and maintenance costs and is difficult to mass produce. The quality of volumetric video is strongly correlated with the number and arrangement of cameras in the acquisition system.
[0003] However, the hardware construction of the volumetric video acquisition system in the existing technology lacks theoretical guiding indicators and relies too much on actual reconstruction results as evaluation. In addition, the number and arrangement of cameras have too large a solution space, which requires a long debugging and experimental process during the hardware construction of the volumetric video acquisition system, greatly increasing the manpower and time costs. Summary of the invention
[0004] In view of this, the purpose of this application is to provide a design method, device and equipment for a volumetric video acquisition system, which first screens out a set of candidate camera arrangement patterns by introducing an evaluation index for the shooting coverage of a target three-dimensional object; then, a secondary evaluation is performed by reconstructing a virtual three-dimensional object to further screen out the target camera arrangement pattern. In this way, with the shooting coverage as a theoretical index, the combination of the two screenings can quickly screen out the optimal camera arrangement pattern from a huge solution space, greatly saving labor costs and time costs, and has guiding significance for the construction of hardware equipment.
[0005] The present application embodiment provides a design method for a volumetric video acquisition system, the design method comprising:
[0006] Generate a camera arrangement pattern set according to acquisition parameter constraints when the volumetric video acquisition system acquires image data of a target three-dimensional object at a target site;
[0007] For each camera arrangement mode in the camera arrangement mode set, simulate and obtain a shooting coverage parameter of the target three-dimensional object in the target site by the volumetric video acquisition system under the camera arrangement mode;
[0008] Filtering a candidate camera arrangement mode set from the camera arrangement mode set according to a shooting coverage parameter of each camera arrangement mode;
[0009] For each candidate camera arrangement mode in the candidate camera arrangement mode set, simulate and obtain a three-dimensional modeling effect of a volumetric video acquisition system on a virtual three-dimensional object under the candidate camera arrangement mode;
[0010] According to the three-dimensional modeling effect of the virtual three-dimensional object, a target camera arrangement pattern of the volumetric video acquisition system is screened out from the candidate camera arrangement pattern set.
[0011] Furthermore, for each candidate camera arrangement mode in the set of candidate camera arrangement modes, simulating a three-dimensional modeling effect of a volumetric video acquisition system on a virtual three-dimensional object under the candidate camera arrangement mode includes:
[0012] Configuring environmental parameters of a virtual rendering tool according to site parameters of the target site, and setting a pre-built three-dimensional model of a virtual three-dimensional object in the virtual rendering tool according to object parameters of the target three-dimensional object;
[0013] For each candidate camera arrangement mode, using the virtual rendering tool to simulate and obtain virtual image data captured by each camera in the volumetric video acquisition system for the virtual three-dimensional object under the camera arrangement mode;
[0014] According to the virtual image data captured by each camera for the virtual three-dimensional object, three-dimensional reconstruction is performed using a preselected three-dimensional reconstruction model to simulate the three-dimensional modeling effect after sampling and training of the virtual three-dimensional object through the volumetric video acquisition system under the candidate camera arrangement mode.
[0015] Furthermore, the three-dimensional reconstruction is performed using a preselected three-dimensional reconstruction model according to the virtual image data captured by each camera for the virtual three-dimensional object, simulating the three-dimensional modeling effect after sampling and training the virtual three-dimensional object through the volume video acquisition system in the camera arrangement mode, including:
[0016] Performing three-dimensional reconstruction using a preselected neural radiation field model according to virtual image data captured by each camera for the virtual three-dimensional object;
[0017] Obtain rendering images of reconstructed three-dimensional objects at different viewing angles based on the trained neural radiation field model;
[0018] Compare the rendered images of the reconstructed three-dimensional object at different viewing angles with the original images of the virtual three-dimensional object at different viewing angles to determine the three-dimensional modeling effect of the volumetric video acquisition system on the virtual three-dimensional object under the camera arrangement mode.
[0019] Furthermore, comparing the rendered images of the reconstructed three-dimensional object at different viewing angles and the original images of the virtual three-dimensional object at different viewing angles to determine the three-dimensional modeling effect of the volumetric video acquisition system on the virtual three-dimensional object under the camera arrangement mode includes:
[0020] For any viewing angle, comparing a rendered image of the reconstructed three-dimensional object at the viewing angle with an original image of the virtual three-dimensional object at the viewing angle, and determining a peak signal-to-noise ratio and a structural similarity of the rendered image relative to the original image;
[0021] According to the peak signal-to-noise ratio and structural similarity of the rendered image relative to the original image, the three-dimensional modeling effect of the volumetric video acquisition system on the virtual three-dimensional object under the camera arrangement mode is determined.
[0022] Further, the acquisition parameter constraint includes a parameter range corresponding to each acquisition parameter; the camera arrangement mode set is generated according to the acquisition parameter constraint when the volumetric video acquisition system acquires image data of the target three-dimensional object at the target site, including:
[0023] For each acquisition parameter, the parameter interval corresponding to the acquisition parameter is sampled at a predetermined interval to generate a value set of the acquisition parameter; wherein the acquisition parameter includes a camera parameter and a spatial range of a target site;
[0024] The value set of each acquisition parameter is traversed to generate the camera arrangement mode set in a permutation and combination manner; wherein each camera arrangement mode is used to indicate the spatial position and orientation of each camera included in the volumetric video acquisition system.
[0025] Furthermore, for each camera arrangement mode in the camera arrangement mode set, a shooting coverage parameter of the target three-dimensional object in the target site by the volumetric video acquisition system under the camera arrangement mode is simulated:
[0026] For each camera arrangement mode, determining a shooting coverage rate of each camera in the camera arrangement mode for the target three-dimensional object in the target site;
[0027] According to the shooting coverage of each camera, an average value and a variance of the shooting coverage of the camera arrangement pattern are determined as shooting coverage parameters of the camera arrangement pattern.
[0028] Further, according to the shooting coverage parameter of each camera arrangement mode, a candidate camera arrangement mode set is screened out from the camera arrangement mode set, including:
[0029] According to a preset rule, the average value and variance of the shooting coverage of each camera arrangement mode are mapped respectively to determine the average value interval and variance interval of the shooting coverage corresponding to each camera arrangement mode;
[0030] The candidate camera arrangement pattern set is screened out from the camera arrangement pattern set based on the average value interval and variance interval of the shooting coverage rate corresponding to each camera arrangement pattern.
[0031] Furthermore, the design method also includes:
[0032] Cameras are set at the target site according to the spatial position and orientation of each camera included in the volumetric video acquisition system indicated by the target camera arrangement pattern.
[0033] The present application also provides a design device for a volumetric video acquisition system, the design device comprising:
[0034] A generation module, for generating a camera arrangement pattern set according to acquisition parameter constraints when the volumetric video acquisition system acquires image data of a target three-dimensional object at a target site;
[0035] A first simulation module is used to simulate, for each camera arrangement mode in the camera arrangement mode set, to obtain a shooting coverage parameter of the target three-dimensional object in the target site by the volumetric video acquisition system under the camera arrangement mode;
[0036] A first screening module, configured to screen out a candidate camera arrangement mode set from the camera arrangement mode set according to a shooting coverage parameter of each camera arrangement mode;
[0037] A second simulation module is used to simulate, for each candidate camera arrangement mode in the candidate camera arrangement mode set, a three-dimensional modeling effect of a volumetric video acquisition system on a virtual three-dimensional object under the candidate camera arrangement mode;
[0038] The second screening module is used to screen out a target camera arrangement pattern of the volumetric video acquisition system from the candidate camera arrangement pattern set according to the three-dimensional modeling effect of the virtual three-dimensional object.
[0039] An embodiment of the present application also provides an electronic device, including: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, the steps of the design method of the volumetric video acquisition system as described above are performed.
[0040] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for designing a volumetric video acquisition system as described above are executed.
[0041] The embodiment of the present application provides a design method, device and equipment for a volumetric video acquisition system. In view of the problems in the prior art that the hardware construction of the volumetric video acquisition system lacks theoretical guiding indicators, relies too much on actual reconstruction results as evaluation, and the number and arrangement of cameras have too large a solution space, the evaluation index of the shooting coverage rate of the target three-dimensional object is introduced to first screen out a set of candidate camera arrangement modes, thereby providing a fast theoretical evaluation index, which is of guiding significance for the construction of hardware equipment; further, a secondary evaluation is performed through the reconstruction of the virtual three-dimensional object to further screen out the target camera arrangement mode. In this way, with the shooting coverage rate as a theoretical indicator, the combination of the two screenings can quickly screen out the optimal camera arrangement mode from the huge solution space, greatly saving labor costs and time costs.
[0042] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0044] Figure 1 A flow chart showing a method for designing a volumetric video acquisition system provided by an embodiment of the present application is shown;
[0045] Figure 2 One of the structural schematic diagrams of a design device for a volumetric video acquisition system provided in an embodiment of the present application is shown;
[0046] Figure 3 A second structural schematic diagram of another design device for a volumetric video acquisition system provided in an embodiment of the present application is shown;
[0047] Figure 4 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0048] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application usually described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, each other embodiment obtained by those skilled in the art without making creative work belongs to the scope of protection of the present application.
[0049] First, the application scenarios to which the present application is applicable are introduced. The present application can be applied to the fields of computer graphics and computer vision technology, specifically in the fields of virtual image technology, volumetric video technology, motion capture technology, etc.
[0050] Research has found that volumetric video is used to capture the captured object in high-realism and high-quality three-dimensional space, with the focus on obtaining high-precision three-dimensional geometric reconstruction and high-resolution texture information. For example, human body volumetric video can record and reconstruct the dynamic form of a person in three dimensions. As a customized multi-view high-precision video, the acquisition system of volumetric video needs to be updated and deployed according to specific shooting requirements. Therefore, it has high construction and maintenance costs and is difficult to mass produce. The quality of volumetric video is strongly correlated with the number and arrangement of cameras in the acquisition system.
[0051] However, the hardware construction of the volumetric video acquisition system in the existing technology lacks theoretical guiding indicators and relies too much on actual reconstruction results as evaluation. In addition, the number and arrangement of cameras have too large a solution space, which requires a long debugging and experimental process during the hardware construction of the volumetric video acquisition system, greatly increasing the manpower and time costs.
[0052] Based on this, an embodiment of the present application provides a design method for a volumetric video acquisition system to quickly design a camera arrangement mode of the volumetric video acquisition system, guide the construction of hardware equipment, and save manpower and time costs.
[0053] The following uses human body volume video as an example to specifically illustrate the implementation of the present application.
[0054] See also Figure 1 , Figure 1 This is a flow chart of a design method for a volumetric video acquisition system provided in an embodiment of the present application. Figure 1 As shown in , the design method provided by the embodiment of the present application includes:
[0055] S101. Generate a camera arrangement pattern set according to acquisition parameter constraints when a volumetric video acquisition system acquires image data of a target three-dimensional object at a target site.
[0056] It should be noted that the volumetric video acquisition system includes multiple cameras (camcorders), and a multi-view camera array formed by the multiple cameras is used to capture and shoot human body movements in three-dimensional space and generate a three-dimensional model.
[0057] Here, the acquisition parameter constraint refers to the constraint conditions set on some acquisition parameters when the volumetric video acquisition system acquires image data of the target three-dimensional object at the target site; specifically, it includes the parameter interval corresponding to each acquisition parameter, that is, the parameter range that each acquisition parameter can take.
[0058] During specific implementation, the acquisition parameter constraints of the human body volumetric video acquisition system can be determined according to project requirements and site conditions; then all possible camera arrangement modes under the acquisition parameter constraints can be determined to form a camera arrangement mode set.
[0059] In a possible implementation, step S101 may include:
[0060] S1011. For each acquisition parameter, perform interval sampling on a parameter interval corresponding to the acquisition parameter at a predetermined interval to generate a value set of the acquisition parameter.
[0061] In this step, for each acquisition parameter, the parameter interval corresponding to the acquisition parameter may be sampled at uniform predetermined intervals to obtain a discrete value set of the acquisition parameter.
[0062] The acquisition parameters include camera parameters and the spatial range of the target site; illustratively, the camera parameters may include the number of cameras, resolution, and field of view, etc.
[0063] S1012. Traverse the value set of each acquisition parameter, and generate the camera arrangement pattern set in a permutation and combination manner; wherein each camera arrangement pattern is used to indicate the spatial position and orientation of each camera included in the volumetric video acquisition system.
[0064] In a specific implementation, each parameter value in the value set of each acquisition parameter is traversed, and all combinations of parameter values, that is, all possible camera arrangement modes, are automatically generated by permutation and combination. The camera arrangement mode is used to indicate the spatial position and orientation of each camera included in the volumetric video acquisition system, which can be specifically represented by the three-dimensional spatial coordinates and rotation vector of each camera.
[0065] In one example, all cameras can be arranged and combined at all discrete points and in all directions based on the spatial range of the target site and the number of cameras. For example, a 1 cubic meter target site can be divided into 1000 cubes with a unit interval of 1 decimeter. One vertex of each cube can be used as a camera pre-selected point. There are 11*11*11=1331 pre-selected points. If each pre-selected point is spaced at 180 degrees, each pre-selected point has 2 pre-selected angles, with a total of 1331*2=2662 possibilities. If there are 2 cameras arranged, there are C 2662 2 =2662*2661 / 2=3541791 possible camera arrangements.
[0066] S102: For each camera arrangement mode in the camera arrangement mode set, simulate and obtain a shooting coverage parameter of the target three-dimensional object in the target site by the volumetric video acquisition system under the camera arrangement mode.
[0067] Here, the shooting coverage refers to the degree to which the data captured by the camera covers the target three-dimensional object.
[0068] In a specific implementation, step S102 may include:
[0069] First, for each camera arrangement mode, the shooting coverage of the target three-dimensional object in the target site by each camera in the camera arrangement mode is determined.
[0070] For volumetric video, the target three-dimensional object has a certain range of activity in the three-dimensional space of the target site. Therefore, based on the theory of multi-view geometry, the coverage rate of the target three-dimensional object in the target site for each camera in each camera arrangement mode can be calculated according to parameters such as the size of the device, the camera field of view angle range, the camera resolution, and the human body interaction range. Among them, the camera field of view refers to the field of view that the camera can capture, usually expressed as an angle, which determines the size and range of the scene that the camera can capture.
[0071] Secondly, the average value and variance of the shooting coverage of the camera arrangement pattern are determined according to the shooting coverage of each camera, as shooting coverage parameters of the camera arrangement pattern.
[0072] Here, the average value and variance of the shooting coverage of each camera in a certain camera arrangement mode are calculated as the shooting coverage parameter of the camera arrangement mode, so as to measure the overall shooting coverage effect of the camera arrangement mode on the target three-dimensional object.
[0073] S103: Filter out a candidate camera arrangement pattern set from the camera arrangement pattern set according to the shooting coverage parameter of each camera arrangement pattern.
[0074] In this step, according to the shooting coverage parameters of each camera arrangement mode, the camera arrangement mode with better overall shooting coverage effect can be preliminarily screened from the camera arrangement mode set to form a candidate camera arrangement mode set. That is, the camera arrangement mode with high mean and low variance is selected as the candidate camera arrangement mode with better indicators.
[0075] In specific implementation, step S103 may include:
[0076] According to preset rules, the mean value and variance of the shooting coverage of each camera arrangement mode are mapped respectively to determine the mean value interval and variance interval of the shooting coverage corresponding to each camera arrangement mode; based on the mean value interval and variance interval of the shooting coverage corresponding to each camera arrangement mode, the candidate camera arrangement mode set is screened out from the camera arrangement mode set.
[0077] As mentioned above, the solution space of the camera arrangement mode is very large, so the amount of calculation for comparing the specific values of the mean values and variances of all camera arrangement modes is also very large. In order to improve the operation speed and reduce the amount of calculation, the embodiment of the present application maps the mean value and variance of the shooting coverage of each camera arrangement mode from specific values to intervals, and then screens out a certain number of candidate camera arrangement modes through interval comparison.
[0078] As an example, for example: the mean of camera arrangement mode 1 is 1.2 and the variance is 0.34; the mean of camera arrangement mode 2 is 1.6 and the variance is 0.27; the mean of camera arrangement mode 3 is 1.65 and the variance is 0.19. The mean is discretized with a step size of 0.1 and the variance is discretized with a step size of 0.05. It is determined that the mean intervals of camera arrangement mode 2 and camera arrangement mode 3 are better than camera arrangement mode 1, and the variance interval of camera arrangement mode 3 is better than camera arrangement mode 2. The candidate camera arrangement modes finally selected can be 2 and 3, or 3; the number of candidate camera arrangement modes can be between 150 and 200.
[0079] In this way, this design method addresses the problem of lack of theoretical guiding indicators in the hardware construction process of the human body volumetric video acquisition system. Based on the multi-view geometry theory, the camera arrangement pattern's shooting coverage of the human body is calculated, thereby providing a fast theoretical evaluation indicator, which is of guiding significance for the construction of hardware equipment.
[0080] S104. For each candidate camera arrangement mode in the set of candidate camera arrangement modes, simulate and obtain a three-dimensional modeling effect of a volumetric video acquisition system on a virtual three-dimensional object under the candidate camera arrangement mode.
[0081] In this step, the embodiment of the present application further uses each candidate camera arrangement mode to simulate rendering of the virtual three-dimensional object, so as to perform secondary screening according to the rendering effect and select the optimal camera arrangement mode.
[0082] In a possible implementation, step S104 may include:
[0083] S1041: configuring environmental parameters of a virtual rendering tool according to site parameters of the target site, and setting a pre-built three-dimensional model of a virtual three-dimensional object in the virtual rendering tool according to object parameters of the target three-dimensional object.
[0084] In this step, in order to ensure the similarity between the simulated rendering and the actual acquisition scene, the environmental parameters of the virtual rendering tool can be configured according to the site parameters of the target site, such as configuring the three-dimensional space range of the virtual rendering tool according to the site size of the target site. After that, the pre-built three-dimensional model of the virtual three-dimensional object, such as a high-quality human body model, is imported into the virtual rendering tool; and the model is set according to the object parameters of the target three-dimensional object, such as setting the position of the three-dimensional model according to the interactive range of the human body in the target site.
[0085] In one example, the virtual rendering tool may be Open3d, which is an open source library for processing 3D data and supporting rapid development and applications such as 3D reconstruction, visualization, and analysis. For example, an existing high-quality human body model may be placed at the world coordinate origin in Open3d.
[0086] S1042: For each candidate camera arrangement mode, use the virtual rendering tool to simulate and obtain virtual image data captured by each camera in the volumetric video acquisition system for the virtual three-dimensional object under the camera arrangement mode.
[0087] In this step, a corresponding virtual camera arrangement is generated in the world coordinate system of the virtual rendering tool according to each candidate camera arrangement mode, and virtual image data of the virtual three-dimensional object captured by each camera in the volumetric video acquisition system under this virtual camera arrangement is rendered.
[0088] S1043. Perform three-dimensional reconstruction using a preselected three-dimensional reconstruction model according to the virtual image data captured by each camera for the virtual three-dimensional object, simulating the three-dimensional modeling effect after sampling and training the virtual three-dimensional object through the volumetric video acquisition system in the candidate camera arrangement mode.
[0089] In this step, by performing three-dimensional reconstruction using virtual image data, the quality of virtual image data obtained under different camera arrangement modes can be further determined based on the three-dimensional modeling effect, thereby measuring the advantages and disadvantages of different camera arrangement modes in order to select the optimal camera arrangement mode.
[0090] In specific implementation, step S1043 may include:
[0091] Step 1: Perform three-dimensional reconstruction using a preselected neural radiation field model based on virtual image data captured by each camera for the virtual three-dimensional object.
[0092] Exemplarily, the embodiment of the present application uses a neural radiance field model as a pre-selected 3D reconstruction model. The Neural Radiance Field (NeRF) model is a deep learning model for reconstructing specific targets or scenes. It uses a deep learning model to capture complex 3D scenes and generate high-quality view synthesis results.
[0093] In this step, the input of the neural radiation field model is the two-dimensional images of the same three-dimensional object taken from different perspectives, that is, the virtual image data taken by each camera for the virtual three-dimensional object; the neural radiation field model calculates the loss value by fitting the function of the neural radiation field model through the images from different perspectives to obtain the position, color and transparency parameter values of the midpoint in the three-dimensional space, so as to continuously adjust the model parameters to obtain a neural radiation field model (function) with acceptable accuracy that fits each spatial position point determined by the image. The process of three-dimensional reconstruction is the process of model sampling and training, and finally a model that fits the three-dimensional space captured by the image is generated.
[0094] Step 2: Obtain rendering images of the reconstructed three-dimensional object at different perspectives based on the trained neural radiation field model.
[0095] In this step, through three-dimensional reconstruction and virtual rendering, the output of the trained neural radiation field model is an image of the reconstructed three-dimensional object observed from different perspectives, that is, a rendered image of the reconstructed three-dimensional object at different perspectives.
[0096] Step 3: Compare the rendered images of the reconstructed three-dimensional object at different viewing angles with the original images of the virtual three-dimensional object at different viewing angles to determine the three-dimensional modeling effect of the volumetric video acquisition system on the virtual three-dimensional object under the camera arrangement mode.
[0097] In this step, the original images of the virtual three-dimensional object at different viewing angles are taken as true values, and compared with the rendered images of the reconstructed three-dimensional object at different viewing angles, so as to determine the three-dimensional modeling effect of the image data collected under the camera arrangement mode.
[0098] As an example, for any viewing angle, compare the rendered image of the reconstructed three-dimensional object at that viewing angle with the original image of the virtual three-dimensional object at that viewing angle to determine the peak signal-to-noise ratio and structural similarity of the rendered image relative to the original image; based on the peak signal-to-noise ratio and structural similarity of the rendered image relative to the original image, determine the three-dimensional modeling effect of the image data collected by the volumetric video acquisition system under this camera arrangement mode on the virtual three-dimensional object.
[0099] Among them, the Peak Signal to Noise Ratio (PSNR) is used to measure the pixel similarity between the NeRF rendered image and the original image. The larger the value, the better the 3D modeling effect. The Structural Similarity (SSIM) is used to measure the structural similarity between the NeRF rendered image and the original image. The larger the value, the better the 3D modeling effect.
[0100] S105 : Filtering a target camera arrangement pattern of a volumetric video acquisition system from the candidate camera arrangement pattern set according to the three-dimensional modeling effect of the virtual three-dimensional object.
[0101] In this step, considering the two reference indicators of PSNR and SSIM, the camera arrangement mode with the highest PSNR and SSIM values can be selected as the optimal target camera arrangement mode.
[0102] In this way, the design method solves the problem of too large a space for camera arrangement patterns. It uses virtual rendering tools and existing high-quality human body models to generate virtual image data. Based on theoretical evaluation indicators, NeRF is used to reconstruct the human body and conduct secondary evaluation and screening. It can find the optimal camera arrangement pattern for the human body volume video acquisition system, greatly saving manpower and time costs.
[0103] Furthermore, the design method also includes:
[0104] S106. Setting a camera at the target site according to the spatial position and orientation of each camera included in the volumetric video acquisition system indicated by the target camera arrangement pattern.
[0105] In this step, the actual hardware construction of the volumetric video acquisition system can be carried out based on the target camera arrangement mode determined in the above manner. Specifically, according to the spatial position and orientation of each camera included in the volumetric video acquisition system indicated by the target camera arrangement mode, the three-dimensional spatial coordinates and rotation vector of each camera are determined, and a high-precision ruler and protractor can be used to help determine the physical position and orientation of the camera, and the camera is set at the target site, thereby completing the hardware construction of the volumetric video acquisition system.
[0106] An embodiment of the present application provides a design method for a volumetric video acquisition system. According to acquisition parameter constraints when the volumetric video acquisition system acquires image data of a target three-dimensional object at a target site, a camera arrangement pattern set is generated; for each camera arrangement pattern in the camera arrangement pattern set, a shooting coverage parameter of the target three-dimensional object in the target site by the volumetric video acquisition system under the camera arrangement pattern is simulated; according to the shooting coverage parameter of each camera arrangement pattern, a candidate camera arrangement pattern set is screened out from the camera arrangement pattern set; for each candidate camera arrangement pattern in the candidate camera arrangement pattern set, a three-dimensional modeling effect of the volumetric video acquisition system under the camera arrangement pattern on a virtual three-dimensional object is simulated; according to the three-dimensional modeling effect on the virtual three-dimensional object, a target camera arrangement pattern of the volumetric video acquisition system is screened out from the candidate camera arrangement pattern set.
[0107] By introducing an evaluation index for the coverage of the target 3D object, we first screen out a set of candidate camera arrangement patterns; then, we reconstruct the virtual 3D object for a secondary evaluation to further screen out the target camera arrangement pattern. In this way, with the coverage as a theoretical index, the combination of the two screenings can quickly screen out the optimal camera arrangement pattern from the huge solution space, greatly saving labor and time costs, and providing guidance for the construction of hardware equipment.
[0108] See also Figure 2 , Figure 3 , Figure 2 This is one of the structural schematic diagrams of a design device for a volumetric video acquisition system provided in an embodiment of the present application. Figure 3 This is a second structural diagram of a design device for a volumetric video acquisition system provided in an embodiment of the present application. Figure 2 As shown in , the design device 200 includes:
[0109] A generation module 210, configured to generate a camera arrangement pattern set according to acquisition parameter constraints when the volumetric video acquisition system acquires image data of a target three-dimensional object at a target site;
[0110] A first simulation module 220 is used to simulate, for each camera arrangement mode in the camera arrangement mode set, to obtain a shooting coverage parameter of the target three-dimensional object in the target site by the volumetric video acquisition system under the camera arrangement mode;
[0111] A first screening module 230, configured to screen out a candidate camera arrangement pattern set from the camera arrangement pattern set according to a shooting coverage parameter of each camera arrangement pattern;
[0112] A second simulation module 240 is used to simulate, for each candidate camera arrangement mode in the candidate camera arrangement mode set, a three-dimensional modeling effect of the volumetric video acquisition system on the virtual three-dimensional object under the candidate camera arrangement mode;
[0113] The second screening module 250 is used to screen out a target camera arrangement pattern of the volumetric video acquisition system from the candidate camera arrangement pattern set according to the three-dimensional modeling effect of the virtual three-dimensional object.
[0114] Further, when the second simulation module 240 is used to simulate the three-dimensional modeling effect of the volumetric video acquisition system on the virtual three-dimensional object under each candidate camera arrangement mode in the candidate camera arrangement mode set, the second simulation module 240 is used to:
[0115] Configuring environmental parameters of a virtual rendering tool according to site parameters of the target site, and setting a pre-built three-dimensional model of a virtual three-dimensional object in the virtual rendering tool according to object parameters of the target three-dimensional object;
[0116] For each candidate camera arrangement mode, using the virtual rendering tool to simulate and obtain virtual image data captured by each camera in the volumetric video acquisition system for the virtual three-dimensional object under the camera arrangement mode;
[0117] According to the virtual image data captured by each camera for the virtual three-dimensional object, three-dimensional reconstruction is performed using a preselected three-dimensional reconstruction model to simulate the three-dimensional modeling effect after sampling and training of the virtual three-dimensional object through the volumetric video acquisition system under the candidate camera arrangement mode.
[0118] Further, when the second simulation module 240 is used to perform three-dimensional reconstruction using a preselected three-dimensional reconstruction model according to the virtual image data captured by each camera for the virtual three-dimensional object, and simulate the three-dimensional modeling effect after the virtual three-dimensional object is sampled and trained by the volumetric video acquisition system in the camera arrangement mode, the second simulation module 240 is used to:
[0119] Performing three-dimensional reconstruction using a preselected neural radiation field model according to virtual image data captured by each camera for the virtual three-dimensional object;
[0120] Obtain rendering images of reconstructed three-dimensional objects at different viewing angles based on the trained neural radiation field model;
[0121] Compare the rendered images of the reconstructed three-dimensional object at different viewing angles with the original images of the virtual three-dimensional object at different viewing angles to determine the three-dimensional modeling effect of the volumetric video acquisition system on the virtual three-dimensional object under the camera arrangement mode.
[0122] Further, when the second simulation module 240 is used to compare the rendered images of the reconstructed three-dimensional object at different viewing angles and the original images of the virtual three-dimensional object at different viewing angles to determine the three-dimensional modeling effect of the volumetric video acquisition system on the virtual three-dimensional object under the camera arrangement mode, the second simulation module 240 is used to:
[0123] For any viewing angle, comparing a rendered image of the reconstructed three-dimensional object at the viewing angle with an original image of the virtual three-dimensional object at the viewing angle, and determining a peak signal-to-noise ratio and a structural similarity of the rendered image relative to the original image;
[0124] According to the peak signal-to-noise ratio and structural similarity of the rendered image relative to the original image, the three-dimensional modeling effect of the volumetric video acquisition system on the virtual three-dimensional object under the camera arrangement mode is determined.
[0125] Further, the acquisition parameter constraint includes a parameter interval corresponding to each acquisition parameter; when the generation module 210 is used to generate a camera arrangement pattern set according to the acquisition parameter constraint when the volumetric video acquisition system acquires image data of a target three-dimensional object at a target site, the generation module 210 is used to:
[0126] For each acquisition parameter, the parameter interval corresponding to the acquisition parameter is sampled at a predetermined interval to generate a value set of the acquisition parameter; wherein the acquisition parameter includes a camera parameter and a spatial range of a target site;
[0127] The value set of each acquisition parameter is traversed to generate the camera arrangement mode set in a permutation and combination manner; wherein each camera arrangement mode is used to indicate the spatial position and orientation of each camera included in the volumetric video acquisition system.
[0128] Further, when the first simulation module 220 is used to simulate, for each camera arrangement mode in the camera arrangement mode set, obtaining a shooting coverage parameter of the target three-dimensional object in the target site by the volumetric video acquisition system under the camera arrangement mode, the first simulation module 220 is used to:
[0129] For each camera arrangement mode, determining a shooting coverage rate of each camera in the camera arrangement mode for the target three-dimensional object in the target site;
[0130] According to the shooting coverage of each camera, an average value and a variance of the shooting coverage of the camera arrangement pattern are determined as shooting coverage parameters of the camera arrangement pattern.
[0131] Further, when the first screening module 230 is used to screen out a candidate camera arrangement pattern set from the camera arrangement pattern set according to the shooting coverage parameter of each camera arrangement pattern, the first screening module 230 is used to:
[0132] According to a preset rule, the average value and variance of the shooting coverage of each camera arrangement mode are mapped respectively to determine the average value interval and variance interval of the shooting coverage corresponding to each camera arrangement mode;
[0133] The candidate camera arrangement pattern set is screened out from the camera arrangement pattern set based on the average value interval and variance interval of the shooting coverage rate corresponding to each camera arrangement pattern.
[0134] Further, such as Figure 3 As shown, the design device 200 also includes: a construction module 260; the construction module 260 is used to: set up cameras at the target site according to the spatial position and orientation of each camera included in the volumetric video acquisition system indicated by the target camera arrangement pattern.
[0135] See also Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 4 As shown in , the electronic device 400 includes a processor 410 , a memory 420 and a bus 430 .
[0136] The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 is running, the processor 410 communicates with the memory 420 via the bus 430. When the machine-readable instructions are executed by the processor 410, the above-mentioned Figure 1 The steps of the method for designing the volumetric video acquisition system in the method embodiment shown and the specific implementation thereof can be found in the method embodiment, which will not be described in detail here.
[0137] The present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the computer program can execute the above-mentioned Figure 1 The steps of the method for designing the volumetric video acquisition system in the method embodiment shown and the specific implementation thereof can be found in the method embodiment, which will not be described in detail here.
[0138] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0139] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0140] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0141] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0142] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present application can essentially be embodied in the form of a software product, or in other words, the part that contributes to the prior art or the part of the technical solution. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0143] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The protection scope of the present application is not limited thereto. Although the present application is described in detail with reference to the above-mentioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-mentioned embodiments within the technical scope disclosed in the present application, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A method for designing a volumetric video acquisition system, characterized in that: The design method comprises: Generate a camera arrangement pattern set according to acquisition parameter constraints when the volumetric video acquisition system acquires image data of a target three-dimensional object at a target site; For each camera arrangement mode in the camera arrangement mode set, a parameter of the shooting coverage rate of the target three-dimensional object in the target site by the volumetric video acquisition system under the camera arrangement mode is simulated; the shooting coverage rate refers to the degree to which the data captured by the camera covers the target three-dimensional object; Filtering a candidate camera arrangement mode set from the camera arrangement mode set according to a shooting coverage parameter of each camera arrangement mode; For each candidate camera arrangement mode in the candidate camera arrangement mode set, simulate and obtain a three-dimensional modeling effect of a volumetric video acquisition system on a virtual three-dimensional object under the candidate camera arrangement mode; According to the three-dimensional modeling effect of the virtual three-dimensional object, a target camera arrangement pattern of the volumetric video acquisition system is screened out from the candidate camera arrangement pattern set.
2. The design method according to claim 1, characterized in that: The simulating, for each candidate camera arrangement mode in the candidate camera arrangement mode set, a three-dimensional modeling effect of the volumetric video acquisition system on the virtual three-dimensional object under the candidate camera arrangement mode includes: Configuring environmental parameters of a virtual rendering tool according to site parameters of the target site, and setting a pre-built three-dimensional model of a virtual three-dimensional object in the virtual rendering tool according to object parameters of the target three-dimensional object; For each candidate camera arrangement mode, using the virtual rendering tool to simulate and obtain virtual image data captured by each camera in the volumetric video acquisition system for the virtual three-dimensional object under the camera arrangement mode; According to the virtual image data captured by each camera for the virtual three-dimensional object, three-dimensional reconstruction is performed using a preselected three-dimensional reconstruction model to simulate the three-dimensional modeling effect after sampling and training of the virtual three-dimensional object through the volumetric video acquisition system under the candidate camera arrangement mode.
3. The design method according to claim 2, characterized in that: The method of performing three-dimensional reconstruction using a preselected three-dimensional reconstruction model according to virtual image data captured by each camera for the virtual three-dimensional object, simulating the three-dimensional modeling effect after sampling and training the virtual three-dimensional object through the volume video acquisition system in the camera arrangement mode, includes: Performing three-dimensional reconstruction using a preselected neural radiation field model according to virtual image data captured by each camera for the virtual three-dimensional object; Obtain rendering images of reconstructed three-dimensional objects at different viewing angles based on the trained neural radiation field model; Compare the rendered images of the reconstructed three-dimensional object at different viewing angles with the original images of the virtual three-dimensional object at different viewing angles to determine the three-dimensional modeling effect of the volumetric video acquisition system on the virtual three-dimensional object under the camera arrangement mode.
4. The design method according to claim 3, characterized in that: The comparing the rendered images of the reconstructed three-dimensional object at different viewing angles and the original images of the virtual three-dimensional object at different viewing angles, and determining the three-dimensional modeling effect of the volumetric video acquisition system on the virtual three-dimensional object under the camera arrangement mode, includes: For any viewing angle, comparing a rendered image of the reconstructed three-dimensional object at the viewing angle with an original image of the virtual three-dimensional object at the viewing angle, and determining a peak signal-to-noise ratio and a structural similarity of the rendered image relative to the original image; According to the peak signal-to-noise ratio and structural similarity of the rendered image relative to the original image, the three-dimensional modeling effect of the volumetric video acquisition system on the virtual three-dimensional object under the camera arrangement mode is determined.
5. The design method according to claim 1, characterized in that: The acquisition parameter constraints include parameter intervals corresponding to each acquisition parameter; The generating of the camera arrangement pattern set according to the acquisition parameter constraints when the volumetric video acquisition system acquires image data of the target three-dimensional object at the target site comprises: For each acquisition parameter, the parameter interval corresponding to the acquisition parameter is sampled at a predetermined interval to generate a value set of the acquisition parameter; wherein the acquisition parameter includes a camera parameter and a spatial range of a target site; The value set of each acquisition parameter is traversed to generate the camera arrangement mode set in a permutation and combination manner; wherein each camera arrangement mode is used to indicate the spatial position and orientation of each camera included in the volumetric video acquisition system.
6. The design method according to claim 1, characterized in that: For each camera arrangement mode in the camera arrangement mode set, the shooting coverage parameter of the target three-dimensional object in the target site by the volumetric video acquisition system under the camera arrangement mode is simulated: For each camera arrangement mode, determining a shooting coverage rate of each camera in the camera arrangement mode for the target three-dimensional object in the target site; According to the shooting coverage of each camera, an average value and a variance of the shooting coverage of the camera arrangement pattern are determined as shooting coverage parameters of the camera arrangement pattern.
7. The design method according to claim 6, characterized in that: According to the shooting coverage parameter of each camera arrangement mode, a candidate camera arrangement mode set is screened out from the camera arrangement mode set, including: According to a preset rule, the average value and variance of the shooting coverage of each camera arrangement mode are mapped respectively to determine the average value interval and variance interval of the shooting coverage corresponding to each camera arrangement mode; The candidate camera arrangement pattern set is screened out from the camera arrangement pattern set based on the average value interval and variance interval of the shooting coverage rate corresponding to each camera arrangement pattern.
8. The design method according to claim 1, characterized in that: The design method further comprises: Cameras are set at the target site according to the spatial position and orientation of each camera included in the volumetric video acquisition system indicated by the target camera arrangement pattern.
9. A design device for a volumetric video acquisition system, characterized in that: The design device comprises: A generation module, for generating a camera arrangement pattern set according to acquisition parameter constraints when the volumetric video acquisition system acquires image data of a target three-dimensional object at a target site; A first simulation module is used to simulate, for each camera arrangement mode in the camera arrangement mode set, to obtain a shooting coverage parameter of the target three-dimensional object in the target site by the volumetric video acquisition system under the camera arrangement mode; the shooting coverage refers to the degree to which the data captured by the camera covers the target three-dimensional object; A first screening module, configured to screen out a candidate camera arrangement pattern set from the camera arrangement pattern set according to a shooting coverage parameter of each camera arrangement pattern; A second simulation module is used to simulate, for each candidate camera arrangement mode in the candidate camera arrangement mode set, a three-dimensional modeling effect of a volumetric video acquisition system on a virtual three-dimensional object under the candidate camera arrangement mode; The second screening module is used to screen out a target camera arrangement pattern of the volumetric video acquisition system from the candidate camera arrangement pattern set according to the three-dimensional modeling effect of the virtual three-dimensional object.
10. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and the machine-readable instructions are executed by the processor to execute the steps of the method for designing a volumetric video acquisition system as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Acquisition system and method for stereoscopic video
CN102572486A
Three-dimensional reconstruction method and device and electronic equipment
CN110910493A