Image data parsing device, scene estimation device, 3D fusion system
By designing an image data analysis device to perform off-site fusion of multi-terminal real scene data, the problem of difficult to achieve three-dimensional image fusion in the prior art is solved, and real-time, multi-terminal camera-coordinated video image fusion and realistic 3D visual fusion effects are realized.
Patent Information
- Application Number
- CN202210714758.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-06-23
AI Technical Summary
The prior art is difficult to achieve three-dimensional image fusion, and traditional image fusion methods cannot achieve realistic 3D fusion effects.
An image data analysis device is designed, including an image raw data analysis unit, a focal stack data analysis unit and a camera parameter analysis unit. Through these units, the multi-end real-life data is fused in a different location to realize three-dimensional visual fusion.
Real-time, multi-end camera-coordinated video images are realized, and the viewpoint can be freely changed while maintaining the consistency of the fused data, avoiding hollow phenomena, and realizing the realistic 3D visual fusion effect.
Smart Images

Figure CN115063561B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video technology, and more specifically, to an image data parsing device, a scene estimation device, and a 3D vision fusion system. Background Art
[0002] With the development of video technology, the applications of AR (Augmented Reality), VR (Virtual Reality), naked-eye 3D, MR (Mixed Reality), and XR (extended reality) have become increasingly perfect, triggering the rapid maturity of various 3D vision products and applications. However, the original design of these 3D vision products is based on the technical architecture of the virtual-real fusion solution. At the same time, image fusion in the academic stage still remains at the stage of two-dimensional data fusion such as the color and brightness of the scene.
[0003] Since the development of video technology to date, in order to meet the growing experience needs of users, the commonly used method in the fusion technology is to implant virtual objects into the actual captured image sequence to achieve the three-dimensional fusion effect, while traditional image fusion is difficult to achieve the technical presentation at the three-dimensional level. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and provide an image data parsing device, a scene estimation device, and a 3D vision fusion system, which can realize real-time, multi-terminal camera collaboration, and multi-application video image remote fusion, avoid the complex limitations in the screen end of the prior art solutions, break away from the three-dimensional rendering engine, can replace the virtual scene with the real scene data of multiple remote ends, and creatively complete the solution of real scene fusion with real scene, can realize the consistent expression of fusion data while freely changing the viewing point, allow the position of the main imaging end camera to change, and the content and scene of the fusion data change in real time following the change of the observation point, avoiding the phenomenon of emptiness, etc.
[0005] The purpose of the present invention is achieved through the following solutions:
[0006] An image data parsing device includes an image parsing unit for remote fusion of multi-terminal real scene data; specifically including:
[0007] An image raw data parsing unit;
[0008] A focus stack data parsing unit;
[0009] A camera parameter parsing unit.
[0010] Further, the image raw data parsing unit includes the following sub-units:
[0011] A calculation sub-unit for calculating the main body data in each frame extracted from the two-dimensional image and its scale in the current frame;
[0012] A quantization subunit, which is used to perform quantization operations on image data, perform inter-frame similarity judgments on video data of the same-end camera, and perform difference judgments on video data of different ends; the quantization operation methods include methods based on image color information, grayscale information, gradient information, and amplitude data in the frequency domain, and generate different intermediate data;
[0013] A metric subunit, including similarity measurement of same-end data and difference measurement of different-end data; the similarity measurement of same-end data includes non-equidistantly spaced marking of video frames to judge the position and scale information of dynamic objects in the video, so that the objects can maintain size and position stability during subsequent processing; the difference measurement of different-end data includes frame-by-frame estimation of the relationship factors between dynamic objects in multi-end videos, and after confirming the relationship factors, it is used to ensure local consistency of each frame of data after fusion;
[0014] A modeling subunit, which is used to jointly model and estimate the similarity measurement parameters of same-end data and the difference measurement parameters of different-end data to obtain a global metric factor, and use the global metric factor to ensure global consistency of the parsed video data.
[0015] Furthermore, the focus stack data parsing unit includes the following subunits:
[0016] A focus stack estimation subunit, which is used to normalize the focus stack data of multi-end videos to a common scale, and then process each frame of image data in the frequency domain to estimate the focal segment position where each frame of data is located;
[0017] A focus stack fusion subunit, which is used to complete the focus stack state conversion of the image data processed in the focus stack estimation subunit in the frequency domain, and then complete the fusion of this part of the image data.
[0018] Furthermore, the camera parameter parsing unit includes the following subunits:
[0019] An image-based camera parameter estimation subunit, which is used to establish the 3D relationship between multi-frame image data in the image raw data parsing unit and the focus stack data parsing unit, and estimate the CCD, FOV, and physical focal length of the camera through the reprojection process, so as to restore the frustum data of camera imaging;
[0020] A mapping solution subunit for the physical focal length of the camera and the focus stack data of the image, which is used to use the discrete focus stack range of each device obtained by the focus stack estimation subunit, combined with the results of the image-based camera parameter estimation subunit, to estimate the mapping relationship between the actual camera focal length range and the focus stack data, and fit the function change relationship between the data.
[0021] A scene estimation device, including a three-dimensional scene data reconstruction unit, configured to reconstruct three-dimensional scene data from the data parsed by the image data parsing device as described above; specifically including the following sub-units:
[0022] A screen parameterization estimation sub-unit, configured to display the dot matrix image on the screen, extract the coordinates of the dots from the captured screen image, and estimate the parametric function of the screen data in the Euclidean space;
[0023] A scene scale estimation sub-unit, configured to splice the camera imaging frustum data of different ends processed by the camera parameter parsing unit together, so that multiple-end cameras jointly form an equivalent visual imaging system, and obtain the final output scale of the scene;
[0024] A static scene reconstruction sub-unit, configured to, for a static scene, combine the scale data obtained by the image raw data parsing unit, and simultaneously simulate multiple planes based on the camera frustum structure to approximate the three-dimensional static scene space;
[0025] A dynamic scene reconstruction sub-unit, configured to, for a dynamic scene, calculate the motion trajectory and its geometric structure of the dynamic scene, combine the scale data obtained by the image raw data parsing unit, and restore the three-dimensional data of the dynamic scene to the real scale.
[0026] A 3D fusion system, including a fusion unit, configured to fuse the data estimated by the scene estimation device as described above and the data parsed by the image data parsing device as described in any one of the above; specifically including the following sub-units:
[0027] A geometric fusion sub-unit, configured to extract matching data using image information, establish a 3D geometric relationship, and convert the three-dimensional scene data of multiple ends into two-dimensional image data on an equivalent visual imaging system;
[0028] An image fusion sub-unit, configured to define the image data as different image blocks according to the 3D geometric relationship in the geometric fusion sub-unit, respectively establish the pixel data histograms of each block of image data, calculate the similarity degree between different image blocks, and then generate a corresponding mask image to assist the edge fusion between the image blocks;
[0029] A fusion consistency processing sub-unit, configured to calculate the image data obtained by converting the multi-end video data to the shooting end according to the geometric fusion sub-unit and the image fusion sub-unit, and then project it onto the display medium according to the parameters of the multi-end scene estimation device.
[0030] Furthermore, the fusion consistency processing subunit includes a geometric consistency processing subunit and an image data consistency processing subunit; the geometric consistency processing subunit is used to link the external parameter data of multiple-end cameras and the external camera parameters of the main imaging end, and calculate the relative pose relationship; the image data consistency processing subunit corrects the image data on the display medium by using a color mapping algorithm.
[0031] The beneficial effects of the present invention include:
[0032] The present invention proposes a real-time, multi-end camera collaborative, multi-application remote fusion technology. First, the present invention avoids the complex limitations in the screen end of the existing technical solutions, and can be applied to any number of screens, screens with any shape structure, and any type of ordinary screen (LED, LCD TV, curtain projection screen, etc.); second, the present invention is separated from the 3D rendering engine, and can use the real scene data of multiple remote ends to replace the virtual scene, and creatively completes a perfect solution for real scene fusion with real scene; third, the present invention can freely change the viewing point while maintaining the consistent expression of the fusion data, allows the position of the main imaging end camera to change, and the content and scene of the fusion data change in real time following the change of the observation point, avoiding the phenomenon of voids.
[0033] The embodiment of the present invention proposes a two-dimensional image data parsing process for the 3D effect fusion of multi-end real scene data, including a device, which solves the parsing of the image raw data, focus stack data, and camera parameters of any-end video. At the same time, the embodiment of the present invention diversifies the objects of two-dimensional image data parsing to meet the applications of different video products.
[0034] The embodiment of the present invention provides a scene estimation device, which can promote a higher degree of data fusion of the data fusion module, and at the same time, as completely as possible, restore the three-dimensional scene to reduce the limitations brought by the display medium (display screen).
[0035] The embodiment of the present invention provides a 3D vision fusion system. Different from the general image fusion process, the data fusion process of the embodiment of the present invention adds two processes of geometric fusion and fusion consistency processing at the beginning and end respectively. According to geometric fusion, image fusion, and fusion consistency processing, it can ensure the rationality and accuracy of the remote fusion output data. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0037] Figure 1 It is the system framework diagram of the embodiment of the present invention;
[0038] Figure 2 It is the schematic diagram of the equivalent frustum system and the multi-terminal imaging system in the embodiment of the present invention; (a) is the multi-terminal imaging system, and (b) is the equivalent frustum system. Specific implementation manners
[0039] All features disclosed in all embodiments in this specification, or steps in all methods or processes implicitly disclosed, except for mutually exclusive features and / or steps, can be combined and / or extended, replaced in any manner.
[0040] Regarding the technical problems to be solved: In the process of solving the problems existing in the existing virtual-real fusion scheme architecture described in the background in the embodiments of the present invention, the following technical problems are found: Image fusion in the academic stage still stays in the two-dimensional data fusion stage such as the color and brightness of the scene. Since the development of video technology to date, in order to meet the growing experience needs of users, the commonly used method in existing fusion technologies is to implant virtual objects into the actual captured image sequence to achieve the three-dimensional fusion effect, while traditional image fusion is difficult to achieve the technical presentation at the three-dimensional level.
[0041] One of the technical concepts of the present invention is to provide a 3D remote fusion system for video images. The remote fusion system of the present invention mainly includes the following three functional modules: image data parsing, multi-terminal scene estimation, and data fusion. The entire system framework is as Figure 1 shown.
[0042] I. Image data parsing
[0043] Different from the three-dimensional reconstruction of general scenes, the core of the embodiment of the present invention is to regard the main body (such as a person) in the video data as the fusion core, which greatly reduces the calculation amount in the subsequent data fusion stage. Whether it is traditional image fusion technology or AR, VR, XR and other technology products based on virtual-real fusion, their applications have very obvious functional limitations. Only when complete three-dimensional information is obtained can a more realistic 3D fusion effect be obtained. In order to get rid of the dilemmas caused by this phenomenon from both the product and technology perspectives, the embodiment of the present invention proposes a technology for remote fusion of multi-terminal real-scene data, aiming to solve the image raw data parsing, focus stack data parsing, and camera parameter parsing of arbitrary-end videos. At the same time, the embodiment of the present invention diversifies the objects of two-dimensional image data parsing to meet the applications of different video products.
[0044] 1) Image raw data parsing: Similar to traditional video signal processing, the input data of such products is a two-dimensional image sequence. Therefore, in this step, while parsing the image data according to traditional methods, we also designed corresponding methods to parse relevant intermediate data in the frequency domain to prepare for subsequent image focus stack estimation. In the specific implementation process, it includes the following steps:
[0045] Step a: First, we use a deep learning-based method for human body recognition. We collected human body recognition datasets under different conditions such as indoor, outdoor, strong illumination, and cloudy days. Through a large amount of training and optimization, we obtained a model with high accuracy, which can efficiently extract human body data from two-dimensional images and calculate the human data in each frame and its scale in the current frame.
[0046] Step b: After step a is completed, we need to perform quantization operations on the image data. The purpose is to judge the similarity between frames of video data from the same camera and the difference between video data from different cameras. The quantization operation method is based on image color information, grayscale information, gradient information, and amplitude data in the frequency domain. The generated different intermediate data act on subsequent different processes respectively.
[0047] Step c: After the image data goes through the quantization process in step b, it enters the measurement stage. This includes two parts: the similarity measurement of data from the same camera and the difference measurement of data from different cameras. The main process of the similarity measurement of data from the same camera is to mark video frames at non-equidistant intervals to judge information such as the position and scale of dynamic objects in the video, so that the object (person) can maintain the stability of size and position in subsequent processing; the main purpose of the difference measurement of data from different cameras is to estimate the relationship factor between dynamic objects in multi-camera videos frame by frame. After confirming the relationship factor, it is possible to ensure the local consistency of each frame of data after fusion.
[0048] Step d: After the data parsing in the above three steps, whether it is data from the same camera or data from different cameras, the data only completes the local consistency parsing. To obtain the global parsing result of all-camera video data, it is necessary to jointly model and estimate the similarity measurement parameters of data from the same camera and the difference measurement parameters of data from different cameras in step c to obtain a global measurement factor, so as to ensure the global consistency of the parsed video data.
[0049] 2) Parsing of focus stack data: For studio products, the realism of the fused data's 3D information perception is greatly affected by camera parameters. In the case of multi-end camera shooting, since the camera devices of videos at different ends may be different, due to the participation of internal parameter information such as focal length, the depth of field of the image formed at each end is different, and the blurred area and the focused area in the image are completely different. Therefore, in the image data parsing stage, it is necessary to estimate the focus stack information of each video frame to improve the richness of the parsed data at the 2D level.
[0050] Step a: Focus stack estimation: We need to estimate the focus stack information frame by frame for each end of the video data at the image level. As a reverse engineering of the imaging process, we have invented an image focus stack estimation method based on frequency domain data. First, normalize the focus stack data of multi-end videos to a common scale, and then process each frame of image data in the frequency domain to estimate the focal segment position where each frame of data is located.
[0051] Step b: Focus stack fusion: Since off-site fusion is a time-sequence-based processing flow, in the fusion stage of multi-end video frames, the focus stack data is usually different. Therefore, it is necessary to convert the off-site image data to the focus stack state of the main imaging end. Similar to the image refocusing operation, combining the frequency domain data in step a of the above step 2), this step will complete the focus stack state conversion of the image data in the frequency domain and complete the fusion of this part of the data.
[0052] 3) Parsing of camera parameters: The above steps 1) and 2) only parse and convert the input source data at the 2D image level to provide data support for subsequent fusion steps. To make the fusion effect reach a more perfect state, it is necessary to perform 3D parsing on the data. The main steps are as follows:
[0053] Step a: Estimation of camera parameters based on images: In the case where the camera parameters are unknown, we provide a more flexible solution to estimate its camera parameters. Combining the data in step a of the above step 1) and the above step 2), this step establishes the 3D relationship between multiple frames of image data, and estimates the CCD, FOV, and physical focal length of the camera through a not-too-complicated reprojection process, so as to restore the frustum data of camera imaging.
[0054] Step b: Mapping solution of the camera's physical focal length and the image focus stack data: Step a in the above step 2) obtains the discrete focus stack range of each device from the image level. Combining the mapping relationship between the actual camera focal length range and the focus stack data estimated in step a of the above step 3), and fitting the function change relationship between the data, it is convenient for data fusion at any focal length state in the subsequent fusion stage.
[0055] II. Multi-end scene estimation
[0056] The previous image data parsing module parses the image data and camera parameters and inputs them into this multi-terminal scene estimation module for the reconstruction of three-dimensional scene data. Its main purpose is to promote a higher degree of data fusion in the next data fusion module and to restore the three-dimensional scene as completely as possible to reduce the limitations brought by the display medium (display screen). The specific steps are as follows:
[0057] 1) Screen parameterization estimation: Although the final presentation effect of remote fusion does not depend on the type and structure of the display medium, if the presentation work needs to be completed on a specific display screen, it is necessary to perform parameterization estimation on the display medium and project the input data onto the screen with the correct geometric relationship. Common screens include single flat screens, L screens, triple-fold screens, and curved screens, etc. For a wide variety of display devices, the present invention designs a unified screen parameterization estimation method. By displaying a dot matrix image on the screen and extracting the coordinates of the dot matrix from the captured screen image, we can estimate the parameterization function of the screen data in the Euclidean space.
[0058] 2) Scene scale estimation: After the screen parameterization estimation in step 1), the scale problem at the imaging hardware device end is solved, and its output result directly affects the accuracy of the final projection data, ensuring that the data projected on the screen is geometrically consistent with the picture in the real scene. However, since remote fusion not only considers the data fusion of two ends, the input data may have more than one signal path. At the same time, combining the scale factor at the two-dimensional level obtained in the image data parsing module, we need to estimate the scale factor at the three-dimensional scene level on this basis to complete the scene scale estimation. Since the data from different ends is projected to different positions on the display device, there must be angle rotation and translation in the scene. This step combines the output of the camera parameter parsing in the image data parsing module, "stitches" the camera frustum data of different ends together, solves the visual drift caused by the combined action of different focal lengths and the scale factor at the image level, and finally enables multiple cameras to jointly form an equivalent visual imaging system, as Figure 2 shown. Among them, sub-figure (a) is the multi-terminal imaging system, and sub-figure (b) is the equivalent frustum system. After this step is completed, the final output scale of the scene is obtained, that is, the transformation standard for all scene data.
[0059] 3) Static scene reconstruction: Video data usually has relatively static background parts and moving main parts (usually people). To improve the efficiency of the entire project, we designed different methods to recover 3D data from 2D images for the two scenarios. For the static scene, this step combines the scale data obtained by the image data parsing module and, based on the camera frustum construction, simulates multiple planes to approximate the three-dimensional static scene space. Different from the 3D reconstruction methods in traditional SLAM or SFM technologies, the inventive method of the present invention does not need to introduce some reconstruction errors through triangulation algorithms and also avoids the efficiency problems caused by reprojection errors. This frustum-based multi-plane three-dimensional scene reconstruction is very suitable for the development of specific application scenarios in the embodiments of the present invention.
[0060] 4) Dynamic scene reconstruction: Considering the real-time requirements of the entire remote fusion system, the method for recovering the 3D data of the main body (person) in the video data is different from that of the static scene 3D scene recovery. In this step, it is not necessary to perform a complete 3D reconstruction of the dynamic scene. It is only necessary to estimate its motion trajectory and its geometric skeleton, and combine the scale data obtained by the image data parsing module to restore the 3D data of the dynamic scene to the real scale.
[0061] III. Data fusion
[0062] After the previous two modules, "image data parsing" and "multi-terminal scene estimation", all the input data required for fusion is obtained. Different from the general image fusion process, the data fusion process invented by us needs to add two processes, geometric fusion and fusion consistency processing, at the beginning and end respectively. Only by strictly following the process sequence of geometric fusion, image fusion, and fusion consistency processing can the rationality and accuracy of the output data of remote fusion be ensured.
[0063] 1) Geometric fusion: Through the scene scale estimation of the multi-terminal scene estimation module, an equivalent visual imaging system is obtained. Therefore, each terminal imaging system is transformed into a non-standard frustum system with an offset and an inclination angle on the basis of the standard visual imaging system. In this regard, in order to correctly perceive the three-dimensional relationship of the multi-terminal scene data on the equivalent visual imaging system, we first extract the corresponding matching data using the image information, establish the 3D geometric relationship of the fusion, and convert the three-dimensional scene data of multiple terminals into two-dimensional image data on the equivalent visual imaging system.
[0064] 2) 3D Visual Fusion: After the above step 1) of data fusion, the image data of multi-terminal video data fusion at the imaging end is obtained. However, due to the 3D geometric stitching relationship, there will inevitably be a hard segmentation phenomenon of the image at the stitching edge. This step mainly solves the fusion problem of the stitching edge of the image data at the image processing level. Define the image data as different image blocks according to the 3D geometric relationship in the above step 1) of data fusion, establish the pixel data histogram of each block of image data respectively, calculate the similarity degree between different blocks, and then generate the corresponding mask image to assist the edge fusion between image blocks.
[0065] 3) Fusion Consistency Processing: After obtaining the corrected image data, we calculate the image data of the multi-terminal video data converted to the shooting end according to steps 1) and 2), and then project it onto the display medium according to the relevant parameters of the multi-terminal scene estimation module. The work of this step mainly deals with the consistency problem between the projection data on the display medium and the real scene data, including geometric consistency and image data consistency. Since the geometric problem of this step has been initialized by the multi-terminal scene estimation, only the external parameter data of the multi-terminal cameras and the external parameters of the main imaging end camera need to be linked in the follow-up to calculate the relative pose relationship, so as to always maintain the geometric consistency of the remote fusion system in terms of data; the image consistency is mainly reflected in the consistent mapping of the screen color and the real scene color space. Correct the image data on the display medium through the color mapping algorithm to complete the image consistency processing.
[0066] Embodiment 1
[0067] An image data parsing device includes an image parsing unit for remote fusion of multi-terminal real scene data; specifically including:
[0068] Image raw data parsing unit;
[0069] Focus stack data parsing unit;
[0070] Camera parameter parsing unit.
[0071] Embodiment 2
[0072] On the basis of Embodiment 1, the image raw data parsing unit includes the following sub-units:
[0073] Calculation sub-unit, used to calculate the main body data in each frame extracted from the two-dimensional image and its scale in the current frame;
[0074] Quantization sub-unit, used to perform quantization operations on the image data, judge the similarity between frames of the video data of the same-end camera, and judge the difference of the video data of different-end cameras; the methods of quantization operations include methods based on image color information, gray information, gradient information, and amplitude data in the frequency domain, and generate different intermediate data;
[0075] The metric sub - unit includes the similarity metric of same - end data and the difference metric of different - end data. The similarity metric of same - end data includes non - equidistantly spacing and marking video frames to judge the position and scale information of dynamic subjects in the video, so that the subjects can maintain stability in size and position during subsequent processing. The difference metric of different - end data includes estimating the relationship factor between dynamic subjects in multi - end videos frame by frame, and after confirming the relationship factor, it is used to ensure local consistency for each frame of data after fusion.
[0076] The modeling sub - unit is used to jointly model and estimate the similarity metric parameters of same - end data and the difference metric parameters of different - end data to obtain a global metric factor, and use the global metric factor to ensure the global consistency of the parsed video data.
[0077] Embodiment 3
[0078] Based on Embodiment 1 or Embodiment 2, the focus - stack data parsing unit includes the following sub - units:
[0079] The focus - stack estimation sub - unit is used to normalize the focus - stack data of multi - end videos to a common scale, and then process each frame of image data in the frequency domain to estimate the focal - segment position where each frame of data is located.
[0080] The focus - stack fusion sub - unit is used to complete the focus - stack state conversion of the image data processed in the focus - stack estimation sub - unit in the frequency domain, and then complete the fusion of this part of the image data.
[0081] Embodiment 4
[0082] Based on Embodiment 3, the camera parameter parsing unit includes the following sub - units:
[0083] The image - based camera parameter estimation sub - unit is used to establish the 3D relationship between multi - frame image data in the image raw data parsing unit and the focus - stack data parsing unit, and estimate the CCD, FOV and physical focal length of the camera through the reprojection process, so as to restore the frustum data of camera imaging.
[0084] The mapping solution sub - unit of the camera physical focal length and the image focus - stack data is used to use the discrete focus - stack range of each device obtained by the focus - stack estimation sub - unit, combined with the results of the image - based camera parameter estimation sub - unit, to estimate the mapping relationship between the actual camera focal - length range and the focus - stack data, and fit the functional change relationship between the data.
[0085] Embodiment 5
[0086] A scene estimation device, including a three-dimensional scene data reconstruction unit, configured to reconstruct three-dimensional scene data from the data parsed by the image data parsing device described in Embodiment 1 or Embodiment 2; specifically including the following sub-units:
[0087] A screen parameterization estimation sub-unit, configured to display a dot matrix image on a screen, extract the coordinates of the dots from the captured screen image, and estimate the parameterized function of the screen data in the Euclidean space;
[0088] A scene scale estimation sub-unit, configured to splice the camera imaging frustum data of different ends processed by the camera parameter parsing unit together, so that multiple-end cameras jointly form an equivalent visual imaging system, and obtain the final output scale of the scene;
[0089] A static scene reconstruction sub-unit, configured to, for a static scene, combine the scale data obtained by the image raw data parsing unit, and simultaneously approximate the three-dimensional static scene space by simulating multiple planes based on the camera frustum structure;
[0090] A dynamic scene reconstruction sub-unit, configured to, for a dynamic scene, calculate the motion trajectory and its geometric structure of the dynamic scene, combine the scale data obtained by the image raw data parsing unit, and restore the three-dimensional data of the dynamic scene to the real scale.
[0091] Embodiment 6
[0092] A 3D fusion system, including a fusion unit, configured to fuse the data estimated by the scene estimation device described in Embodiment 5 and the data parsed by the image data parsing device described in any one of Embodiment 1 or Embodiment 2; specifically including the following sub-units:
[0093] A geometric fusion sub-unit, configured to extract matching data using image information, establish 3D geometric relationships, and convert the three-dimensional scene data of multiple ends into two-dimensional image data on an equivalent visual imaging system;
[0094] An image fusion sub-unit, configured to define the image data as different image blocks according to the 3D geometric relationships in the geometric fusion sub-unit, respectively establish the pixel data histograms of each block of image data, calculate the similarity between different image blocks, and then generate corresponding mask images to assist in edge fusion between image blocks;
[0095] A fusion consistency processing sub-unit, configured to convert the multi-end video data into image data at the shooting end according to the calculations of the geometric fusion sub-unit and the image fusion sub-unit, and then project it onto a display medium according to the parameters of the multi-end scene estimation device.
[0096] Embodiment 7
[0097] On the basis of Embodiment 6, the fusion consistency processing subunit includes a geometric consistency processing subunit and an image data consistency processing subunit; the geometric consistency processing subunit is used to link the external parameter data of multi-end cameras and the camera external parameter data of the main imaging end, and calculate the relative pose relationship; the image data consistency processing subunit corrects the image data on the display medium by using a color mapping algorithm.
[0098] Parts not involved in the present invention are the same as the prior art or can be implemented by using the prior art.
[0099] Except for the above examples, those skilled in the art can obtain inspiration according to the above disclosure or make modifications by using the knowledge or technology in related fields to obtain other embodiments. The features of each embodiment can be interchanged or replaced. As long as the modifications and changes made by those skilled in the art do not depart from the spirit and scope of the present invention, they should fall within the protection scope of the appended claims of the present invention.
Claims
1. An image data parsing device, characterized in that, An image parsing unit for cross - location fusion of multi - terminal real - scene data; specifically including: Image raw data parsing unit; Focus stack data parsing unit; Camera parameter parsing unit; The image raw data parsing unit includes the following sub - units: Calculation sub - unit, used to calculate the main data in each frame extracted from the two - dimensional image and its scale in the current frame; Quantization sub - unit, used to perform quantization operations on image data, judge the similarity between frames of video data of the same - end camera, and judge the difference of video data of different - end cameras; the quantization operation method is based on image color information, grayscale information, gradient information, and amplitude data in the frequency domain, and generates different intermediate data; Measurement sub - unit, including similarity measurement of same - end data and difference measurement of different - end data; similarity measurement of same - end data includes non - equidistantly spaced marking of video frames to judge the position and scale information of dynamic objects in the video, so that the object can maintain stability in size and position during subsequent processing; difference measurement of different - end data includes frame - by - frame estimation of the relationship factor between dynamic objects in multi - terminal videos, and after confirming the relationship factor, it is used to ensure local consistency of each frame of data after fusion; Modeling sub - unit, used to jointly model and estimate the similarity measurement parameters of same - end data and the difference measurement parameters of different - end data to obtain a global measurement factor, and use the global measurement factor to ensure global consistency of the parsed video data; the focus stack data parsing unit includes the following sub - units: Focus stack estimation sub - unit, used to normalize the focus stack data of multi - terminal videos to a common scale, and then process each frame of image data in the frequency domain to estimate the focus stack position where each frame of data is located; Focus stack fusion sub - unit, used to complete the focus stack state conversion of the image data processed in the focus stack estimation sub - unit in the frequency domain, and then complete the fusion of this part of the image data; The camera parameter parsing unit includes the following sub - units: Image - based camera parameter estimation sub - unit, used to establish the 3D relationship between multi - frame image data in the image raw data parsing unit and the focus stack data parsing unit, and estimate the CCD, FOV, and physical focal length of the camera through the reprojection process, so as to restore the frustum data of camera imaging; Mapping solution sub - unit for camera physical focal length and image focus stack data, used to use the discrete focus stack range of each device obtained by the focus stack estimation sub - unit, combined with the results of the image - based camera parameter estimation sub - unit, to estimate the mapping relationship between the actual camera focal length range and the focus stack data, and fit the functional change relationship between the data.
2. A scene estimation device, characterized in that, It includes a three - dimensional scene data reconstruction unit, used to reconstruct the three - dimensional scene data from the data parsed by the image data parsing device described in claim 1; specifically including the following sub - units: screen parameterization estimation sub - unit, used to display the dot - matrix image on the screen, extract the coordinates of the dots of the captured screen image, and estimate the parameterization function of the screen data in the Euclidean space; A scene scale estimation subunit, configured to splice the camera imaging frustum data of different ends processed by the camera parameter parsing unit together, so that multiple-end cameras jointly form an equivalent visual imaging system, and obtain the final output scale of the scene; A static scene reconstruction subunit, configured to, for a static scene, combine the scale data obtained by the image raw data parsing unit, and simultaneously approximate the three-dimensional static scene space by simulating multiple planes based on the camera frustum structure; A dynamic scene reconstruction subunit, configured to, for a dynamic scene, calculate the motion trajectory and its geometric structure of the dynamic scene, and combine the scale data obtained by the image raw data parsing unit to restore the three-dimensional data of the dynamic scene to the real scale.
3. A 3D fusion system, characterized in that, It includes a fusion unit, configured to fuse the data estimated by the scene estimation device according to Claim 2 and the data parsed by the image data parsing device according to Claim 1; specifically including the following subunits: A geometric fusion subunit, configured to extract matching data using image information, establish a 3D geometric relationship, and convert the three-dimensional scene data of multiple ends into two-dimensional image data on an equivalent visual imaging system; An image fusion subunit, configured to define the image data as different image blocks according to the 3D geometric relationship in the geometric fusion subunit, respectively establish a pixel data histogram of each block of image data, calculate the similarity degree between different image blocks, and then generate a corresponding mask image to assist the edge fusion between image blocks; A fusion consistency processing subunit, configured to calculate the image data obtained by converting the multi-end video data to the image data at the shooting end according to the geometric fusion subunit and the image fusion subunit, and then project it onto the display medium according to the parameters of the multi-end scene estimation device.
4. The 3D fusion system according to claim 3, characterized in that, The fusion consistency processing subunit includes a geometric consistency processing subunit and an image data consistency processing subunit; the geometric consistency processing subunit is used to link the external parameter data of the multi-end cameras and the external parameter data of the camera at the main imaging end, and calculate the relative pose relationship; the image data consistency processing subunit corrects the image data on the display medium using a color mapping algorithm.
Citation Information
Patent Citations
Multi-angle consistent plane detection and analysis method for monocular video scene three dimensional structure
CN106570507A
3D scene engineering simulation and real-life scene fusion system
US10769843B1