A video fusion method and device based on video foreground extraction
By generating 3D models using Gaussian mixture models and constrained delaunay triangulation algorithms, the problem of unreasonable foreground display in free-viewpoint and multi-video overlapping areas in existing technologies is solved, achieving reasonable display and information integration under any viewpoint.
Patent Information
- Application Number
- CN202210252167.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-15
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2042-03-15
AI Technical Summary
Existing technologies cannot adapt to foreground display in free-viewpoint and multi-video overlapping areas, resulting in unreasonable foreground display in 3D scenes.
The foreground and background are separated by Gaussian mixture model, the foreground contour is extracted and the bounding box is calculated using morphological methods, and a 3D model is generated by constrained delaunay triangulation algorithm. The foreground pose is adjusted according to the viewpoint to display it reasonably from any viewpoint.
It enables reasonable display of the foreground from any angle, solves the problem of foreground display in overlapping areas of multiple videos, and provides a good user experience and information integration capabilities.
Smart Images

Figure CN114648476B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a video fusion method and device based on video foreground extraction. BACKGROUND
[0002] Taking pictures of a monitoring scene by means of a drone or a digital camera can be used to reconstruct a three-dimensional scene. Before reconstruction, the pictures are taken with certain standards. Firstly, the pictures taken should completely cover the scene to be reconstructed; secondly, each picture should have more than 30% overlap with at least one other picture; and finally, each picture should not contain a close-up shot. Once all the pictures are taken, the three-dimensional reconstruction is performed by means of SfM and MVS algorithms to obtain the camera pose of the pictures taken, the sparse point cloud and the dense point cloud of the scene. With the dense point cloud, the three-dimensional model of the monitoring scene can be obtained by means of Poisson reconstruction and mesh simplification algorithms.
[0003] For a monitoring video with a fixed view angle, the moving objects in the video are generally referred to as foreground, and the remaining part after the foreground is removed is referred to as background. The prior art has detected and extracted the moving objects (pedestrians, vehicles, etc.) in a single monitoring video, expressed the objects in a three-dimensional scene by mapping after estimating the depth, and the result is shown in Figure 1 . Figure 1 The foreground is shown in a three-dimensional virtual scene by means of a plane. There is also a method based on a single video, which expresses the foreground in a simple three-dimensional scene by means of a plane and a prefabricated body, and the result is shown in Figure 2 . Figure 2 The foreground is shown in a three-dimensional virtual scene by means of a plane and a simple prefabricated body. However, the prior art has the following disadvantages:
[0004] 1. The prior art can only reasonably display the foreground in a three-dimensional scene at a certain specific angle, and cannot adapt to a free view angle.
[0005] 2. The prior art does not consider the foreground display in the overlapping area of multiple videos in a three-dimensional scene. SUMMARY
[0006] The embodiments of the present application provide a video fusion method and device based on video foreground extraction, which at least solve the technical problem that the prior art cannot adapt to a free view angle.
[0007] According to an embodiment of the present application, a video fusion method based on video foreground extraction is provided, which comprises the following steps:
[0008] extracting the foreground and its contour in multiple videos;
[0009] calibrating the pose of the extracted video foreground in a three-dimensional virtual scene;
[0010] combining the video foregrounds after pose calibration with the video foreground contours to generate corresponding three-dimensional models;
[0011] fusing the foregrounds of multiple videos in the three-dimensional model.
[0012] Further, the extracting the foregrounds and their contours in the multiple videos comprises:
[0013] using a Gaussian mixture model to separate and extract the foregrounds from the backgrounds in the multiple videos, and then extracting the contours of the foregrounds and calculating the bounding boxes by a morphological method.
[0014] Further, the calibrating the poses of the extracted video foregrounds in the three-dimensional virtual scene comprises:
[0015] assuming that the plane where the foreground model is located is parallel to the imaging plane, extracting the contours of multiple foregrounds in each monitoring video and calculating the bounding boxes, and sending a ray from the optical center of the monitoring camera to the center of the lower edge of the bounding box on the imaging plane, and the intersection of the ray and the scene is the position of the foreground in the virtual three-dimensional scene.
[0016] Further, the combining the video foregrounds after pose calibration with the video foreground contours to generate corresponding three-dimensional models comprises:
[0017] generating the corresponding three-dimensional models by a constrained delaunay triangulation algorithm after combining the video foregrounds after pose calibration with the video foreground contours.
[0018] Further, the fusing the foregrounds of multiple videos in the three-dimensional model comprises:
[0019] when multiple monitoring cameras simultaneously shoot the same object, selecting the optimal foreground for display according to the viewpoint of the user;
[0020] calculating the distance between the foregrounds from different videos, if the distances between multiple foregrounds are close to each other, regarding these foregrounds as belonging to the same object and putting them into a candidate set, and then calculating the unit direction vector v f , from the foreground coordinates to the observation camera, and calculating the unit normal vector n f of the plane where each foreground is located according to the normal vector of the imaging plane of the corresponding monitoring camera. f ·n f The foreground that can maximize v
[0021] Further, the method further comprises:
[0022] adjusting the pose of the foreground according to the observation camera to ensure that the foreground always faces the observation camera.
[0023] According to another embodiment of the present application, a video fusion device based on video foreground extraction is provided, comprising:
[0024] a foreground extraction unit configured to extract foregrounds and their contours from a plurality of videos;
[0025] a pose calibration unit configured to calibrate poses of the extracted foregrounds in a three-dimensional virtual scene;
[0026] a three-dimensional model generation unit configured to generate corresponding three-dimensional models by combining the foregrounds after pose calibration with the foreground contours;
[0027] a fusion unit configured to fuse the foregrounds of the plurality of videos in the three-dimensional models.
[0028] Further, the device further comprises:
[0029] a pose adjustment unit configured to adjust the foreground poses according to an observation camera to ensure that the foregrounds always face the observation camera.
[0030] A storage medium storing a program file capable of implementing any one of the above video fusion methods based on video foreground extraction.
[0031] A processor configured to run a program, wherein the program, when running, performs any one of the above video fusion methods based on video foreground extraction.
[0032] The video fusion method and device based on video foreground extraction in the embodiments of the present application can extract foregrounds and their contours from a plurality of videos, calibrate poses of the extracted foregrounds in a three-dimensional virtual scene, generate corresponding three-dimensional models by combining the foregrounds after pose calibration with the foreground contours, and fuse the foregrounds of the plurality of videos in the three-dimensional models. The present application can extract foregrounds from multiple videos and display the foregrounds in a virtual three-dimensional scene, can calculate poses of the foregrounds according to a viewing angle, and thus can obtain reasonable observation results at any viewing angle and provide users with good user experience.
[0033] The present application has at least the following advantages:
[0034] 1. Unlike the conventional foreground display method, the present application can transform foregrounds according to a viewing angle, so that the foregrounds always face an observation camera.
[0035] 2. The present application can fuse foregrounds of multiple monitoring videos to display the foregrounds and solve the problem of repeated display of foregrounds in multiple videos.
[0036] 3. The present application supports users to observe foregrounds at any viewing angle and provides users with good user experience.
[0037] 4. The present application can fuse foregrounds of multiple videos and has strong information integration capability.
[0038] 5. The application can calculate the pose of the foreground according to the visual angle, so that reasonable observation results can be obtained at any visual angle.
[0039] 6. The application can extract the foreground in multiple videos and display the fusion in a virtual three-dimensional scene. BRIEF DESCRIPTION OF DRAWINGS
[0040] The drawings described herein are used to provide further understanding of the present application, and form a part of the present application. The illustrative embodiments of the present application and their description serve to explain the present application. They do not constitute an improper limitation on the present application. In the drawings:
[0041] Figure 1 A foreground image displayed by a plane in a three-dimensional virtual scene;
[0042] Figure 2 A foreground image displayed by a plane and a simple prefabricated body in a three-dimensional virtual scene;
[0043] Figure 3 A foreground calibration diagram in a video fusion method based on video foreground extraction of the present application;
[0044] Figure 4 A foreground extraction process diagram in a video fusion method based on video foreground extraction of the present application;
[0045] Figure 5 A two-video foreground fusion result diagram in a video fusion method based on video foreground extraction of the present application. DETAILED DESCRIPTION
[0046] In order to make the person in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the person in the art without creative labor should belong to the protection scope of the present application.
[0047] It is to be understood that the terminology "first", "second" and the like used in the specification and the claims of the application as well as the foregoing drawings is merely intended to distinguish between similar objects and not necessarily for describing a particular sequential order. It is to be understood that the use of such terms can be interchanged insofar as is appropriate to the subject matter at hand. Furthermore, the terms "comprising", "having", "including", and the like, when used in the specification and in the following claims, are intended to mean the open-ended inclusion of the stated steps or elements, that the processes, methods, systems, products, or apparatuses, comprising such steps or elements do not necessarily consist of those steps or elements, but can include other steps or elements not stated or inherent to these processes, methods, products, or apparatuses.
[0048] Embodiment 1
[0049] According to an embodiment of the application, there is provided a video fusion method based on video foreground extraction, comprising the following steps:
[0050] extracting foreground and its contour in multiple videos;
[0051] calibrating pose of the extracted video foreground in a three-dimensional virtual scene;
[0052] generating a corresponding three-dimensional model by combining the video foreground after pose calibration with the video foreground contour;
[0053] fusing foreground of multiple videos in the three-dimensional model.
[0054] The video fusion method based on video foreground extraction in the embodiment of the application extracts foreground and its contour in multiple videos, calibrates pose of the extracted video foreground in a three-dimensional virtual scene, generates a corresponding three-dimensional model by combining the video foreground after pose calibration with the video foreground contour, and fuses foreground of multiple videos in the three-dimensional model. The application can extract foreground in multiple videos and fuse and display in a virtual three-dimensional scene, and can calculate pose of the foreground according to a view angle, so that a reasonable observation result is obtained at any view angle, providing a good user experience.
[0055] The extracting foreground and its contour in multiple videos comprises:
[0056] Gaussian mixture model is used to separate and extract foreground from background in multiple videos, and then a morphological method is used to extract contour of the foreground and calculate a bounding box.
[0057] The calibrating pose of the extracted video foreground in a three-dimensional virtual scene comprises:
[0058] Assuming the plane containing the foreground model is parallel to the imaging plane, the contours of multiple foregrounds in each surveillance video are extracted and the bounding box is calculated. A ray is emitted from the optical center of the surveillance camera toward the center of the lower edge of the bounding box on the imaging plane. The intersection of this ray with the scene is the position of the foreground in the virtual 3D scene.
[0059] The process of generating a corresponding 3D model by combining the pose-calibrated video foreground with the video foreground contour includes:
[0060] The video foreground, after pose calibration, is combined with the video foreground contour to generate the corresponding 3D model using the constrained delaunaytriangulation algorithm.
[0061] The foreground that integrates multiple videos into the 3D model includes:
[0062] When multiple surveillance cameras capture images of the same object simultaneously, the optimal foreground is selected for display based on the user's viewpoint.
[0063] Calculate the distance between foregrounds from different videos. If multiple foregrounds are close to each other, they are considered to belong to the same object and added to the candidate set. Then, calculate the unit direction vector v from the foreground coordinates to the viewing camera. f Simultaneously, based on the imaging plane normal vector of the corresponding monitoring camera, the unit normal vector n of the plane containing each foreground is calculated. f In the candidate set, v can be f ·n f The maximized prospect is selected as the optimal prospect.
[0064] The methods also include:
[0065] Adjust the foreground pose according to the observation camera to ensure that the foreground is always facing the observation camera.
[0066] The video fusion method based on video foreground extraction of the present invention will be described in detail below with specific embodiments:
[0067] This invention can integrate the foreground from multiple surveillance videos and display it reasonably in a three-dimensional virtual scene, so that users can obtain reasonable observation results of the foreground in all videos from any viewpoint in the three-dimensional virtual scene.
[0068] Traditional multi-channel monitoring systems present fragmented images, making it difficult for users to intuitively understand the monitored content. In some cases, users may only be interested in certain moving objects in the video. To address these moving objects in the video, this invention proposes a video fusion method based on video foreground extraction, which allows users to more easily understand and track the movement of objects in the scene.
[0069] The technical solution of the present application can be briefly summarized as: extracting foreground and its contour -> calibrating foreground pose -> generating three-dimensional model according to foreground contour -> fusing foreground of multiple videos -> adjusting foreground pose according to observation camera to ensure that foreground always faces observation camera.
[0070] Firstly, the moving object (i.e. foreground) in the video can be extracted by Gaussian Mixture Model (GMM) algorithm. In the present application, Gaussian Mixture Model is used to separate and extract foreground from background. Then the contour of the foreground can be extracted and the bounding box can be calculated by morphological method. The extraction result is shown in Figure 4 After obtaining the foreground, the pose needs to be calibrated in the three-dimensional virtual scene. After the pose calibration, the corresponding three-dimensional model can be generated according to the contour by constrained delaunay triangulation algorithm. Then the foreground in multiple videos is fused. Finally, the billboard technique can be used to ensure that the model always faces the observation camera, so that the user can easily track the moving object in the video.
[0071] The calibration process of the foreground is as follows: firstly, the contour of the foreground is extracted and the bounding box is calculated for each monitoring video. A ray is emitted from the optical center of the monitoring camera to the center of the lower edge of the bounding box on the imaging plane. The intersection point of the ray and the scene is the position of the foreground in the virtual three-dimensional scene. The present application assumes that the plane where the foreground model is located is parallel to the imaging plane. Therefore, the intersection point of the ray and the three-dimensional virtual environment and the normal vector of the imaging plane are obtained, and the plane where the foreground mesh model is located can be uniquely determined. The calibration process is shown in Figure 3 .
[0072] The multi-video foreground fusion process: there is a problem to be solved in the foreground generation process. When multiple monitoring cameras simultaneously capture the same object, the optimal foreground needs to be selected according to the user's viewpoint for display. The present application first calculates the distance between the foregrounds from different videos. If the distances between multiple foregrounds are very close to each other, these foregrounds are considered to belong to the same object and are put into the candidate set. Then the unit direction vector v f from the foreground coordinates to the observation camera is calculated, and the unit normal vector n f of the plane where each foreground is located is calculated according to the imaging plane normal vector of the corresponding monitoring camera. Then the foreground in the candidate set that can maximize v f ·n f is selected as the optimal foreground. The multi-video foreground fusion result is shown in Figure 5 . Figure 5 The left image is the fusion result under the video shooting angle, and the middle and right images are the fusion results after rotating a certain angle.
[0073] Embodiment 2
[0074] According to another embodiment of the present application, a video fusion device based on video foreground extraction is provided, comprising:
[0075] a foreground extraction unit configured to extract foregrounds and their contours from a plurality of videos;
[0076] a pose calibration unit configured to calibrate poses of the extracted foregrounds in a three-dimensional virtual scene;
[0077] a three-dimensional model generation unit configured to generate corresponding three-dimensional models by combining the foregrounds after pose calibration with the foreground contours;
[0078] a fusion unit configured to fuse the foregrounds of the plurality of videos in the three-dimensional models.
[0079] The video fusion device based on video foreground extraction in the embodiment of the present application extracts foregrounds and their contours from a plurality of videos, calibrates poses of the extracted foregrounds in a three-dimensional virtual scene, generates corresponding three-dimensional models by combining the foregrounds after pose calibration with the foreground contours, and fuses the foregrounds of the plurality of videos in the three-dimensional models. The present application can extract foregrounds from a plurality of videos and display them in a virtual three-dimensional scene, and can calculate poses of the foregrounds according to a viewing angle, so that reasonable observation results can be obtained at any viewing angle, providing a good user experience.
[0080] The device further comprises:
[0081] a pose adjustment unit configured to adjust the poses of the foregrounds according to a viewing camera to ensure that the foregrounds always face the viewing camera.
[0082] The video fusion device based on video foreground extraction of the present application will be described in detail below with specific embodiments:
[0083] The present application can fuse foregrounds in a plurality of monitoring videos and display them reasonably in a three-dimensional virtual scene, so that a user can obtain reasonable observation results of the foregrounds in all videos at any viewing point in the three-dimensional virtual scene.
[0084] Traditional multi-channel monitoring systems are mutually isolated between different pictures, which is not conducive to users to intuitively understand the monitoring content. In some cases, users only care about some moving objects in the videos. For these moving objects in the videos, the present application proposes a video fusion device based on video foreground extraction, which can make it easier for users to understand and track the movement of objects in the videos in the scene.
[0085] The technical solution of the present application can be briefly summarized as follows: an extraction unit extracts foreground and its contour; a pose calibration unit calibrates foreground pose; a three-dimensional model generation unit generates a three-dimensional model according to the foreground contour; a fusion unit fuses foregrounds of multiple videos; and a pose adjustment unit adjusts foreground pose according to an observation camera to ensure that the foreground always faces the observation camera.
[0086] First, moving objects (i.e. foregrounds) in the video can be extracted by a Gaussian Mixture Model (GMM) algorithm. In the present application, the Gaussian Mixture Model is used to separate and extract the foreground from the background. Then, the contour of the foreground can be extracted by a morphological method and a bounding box can be calculated. The extraction result is shown in FIG. 1. Figure 4 After obtaining the foreground, the pose needs to be calibrated in a three-dimensional virtual scene. After the pose is calibrated, a corresponding three-dimensional model can be generated according to the contour by a constrained delaunay triangulation algorithm. Then, the foregrounds in multiple videos are fused. Finally, the billboard technique can be used to ensure that the model always faces the observation camera, so that the user can easily track the moving object in the video.
[0087] The calibration process of the foreground of the pose calibration unit is as follows: first, the contour of each foreground is extracted and a bounding box is calculated. A ray is emitted from the optical center of the monitoring camera to the center of the lower edge of the bounding box on the imaging plane. The intersection point of the ray and the scene is the position of the foreground in the virtual three-dimensional scene. The present application assumes that the plane where the foreground model is located is parallel to the imaging plane. Therefore, the intersection point of the ray and the three-dimensional virtual environment and the normal vector of the imaging plane can uniquely determine the plane where the foreground mesh model is located. The calibration process is shown in FIG. 2. Figure 3
[0088] The multi-video foreground fusion process of the fusion unit: there is a problem to be solved in the foreground generation process. When multiple monitoring cameras simultaneously capture the same object, the optimal foreground needs to be selected according to the user's viewpoint for display. The present application first calculates the distance between the foregrounds from different videos. If the distances between multiple foregrounds are very close to each other, these foregrounds are considered to belong to the same object and are put into a candidate set. Then, the unit direction vector v f from the foreground coordinates to the observation camera is calculated, as well as the unit normal vector n f of the plane where each foreground is located. The present application can calculate the unit normal vector of the imaging plane of the corresponding monitoring camera, so that the foreground in the candidate set that can maximize v f · n f is selected as the optimal foreground. The multi-video foreground fusion result is shown in FIG. 3. Figure 5 Figure 5 The left image is a fusion result image in a video shooting perspective, and the middle and right images are fusion result images after rotating by a certain angle.
[0089] Embodiment 3
[0090] A storage medium, the storage medium stores a program file capable of implementing any one of the above video fusion methods based on video foreground extraction.
[0091] Embodiment 4
[0092] A processor, the processor is used to run a program, wherein the program executes the video fusion method based on video foreground extraction of any one of the above when running.
[0093] The above-mentioned embodiment numbers of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0094] In the above-mentioned embodiments of the present application, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0095] In several embodiments provided in the present application, it should be understood that the disclosed technical contents can be implemented by other ways. Among them, the system embodiments described above are only schematic, for example, the division of units can be a logical function division, and actual implementation can have another division way, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, units or modules, which can be electrical or other forms.
[0096] The units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment scheme.
[0097] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of software functional unit.
[0098] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server or a network device, etc.) execute all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0099] The above is only the preferred embodiment of the present application, it should be pointed out that, for those skilled in the art, without departing from the principles of the present application, can make a number of improvements and refinements, these improvements and refinements should also be considered as the protection scope of the present application.
Claims
1. A video fusion method based on video foreground extraction, characterized in that, Includes the following steps: Extract the foreground and its outline from multiple videos; Position the extracted video foreground in a 3D virtual scene; The video foreground after pose calibration is combined with the video foreground outline to generate the corresponding 3D model; Integrate the foreground of multiple videos into a 3D model; The foreground of fusing multiple videos in a 3D model includes: When multiple surveillance cameras capture images of the same object simultaneously, the optimal foreground is selected for display based on the user's viewpoint. Calculate the distance between foregrounds from different videos. If multiple foregrounds are close to each other, they are considered to belong to the same object and added to the candidate set. Then, calculate the unit direction vector v from the foreground coordinates to the viewing camera. f Simultaneously, based on the imaging plane normal vector of the corresponding monitoring camera, the unit normal vector n of the plane containing each foreground is calculated. f In the candidate set, v can be f ·n f The maximized prospect is selected as the optimal prospect.
2. The video fusion method based on video foreground extraction according to claim 1, characterized in that, The extraction of the foreground and its outline from multiple videos includes: A Gaussian mixture model is used to separate and extract the foreground and background in multiple videos. Then, morphological methods are used to extract the outline of the foreground and calculate the bounding box.
3. The video fusion method based on video foreground extraction according to claim 2, characterized in that, The extraction of video foreground pose in the 3D virtual scene includes: Assuming the plane containing the foreground model is parallel to the imaging plane, the contours of multiple foregrounds in each surveillance video are extracted and the bounding box is calculated. A ray is emitted from the optical center of the surveillance camera toward the center of the lower edge of the bounding box on the imaging plane. The intersection of this ray with the scene is the position of the foreground in the virtual 3D scene.
4. The video fusion method based on video foreground extraction according to claim 3, characterized in that, The step of combining the pose-calibrated video foreground with the video foreground contour to generate a corresponding 3D model includes: The video foreground, after pose calibration, is combined with the video foreground contour to generate the corresponding 3D model using the constrained delaunaytriangulation algorithm.
5. The video fusion method based on video foreground extraction according to claim 1, characterized in that, The method further includes: Adjust the foreground pose according to the observation camera to ensure that the foreground is always facing the observation camera.
6. A video fusion apparatus based on video foreground extraction, utilizing the video fusion method based on video foreground extraction as described in claim 1, characterized in that, include: The extraction unit is used to extract the foreground and its outline from multiple videos; The pose calibration unit is used to calibrate the pose of the extracted video foreground in a 3D virtual scene; The 3D model generation unit is used to combine the video foreground after pose calibration with the video foreground outline to generate the corresponding 3D model. The fusion unit is used to fuse the foreground of multiple videos in a 3D model.
7. The video fusion apparatus based on video foreground extraction according to claim 6, characterized in that, The device further includes: The pose adjustment unit is used to adjust the pose of the foreground according to the observation camera to ensure that the foreground is always facing the observation camera.
8. A storage medium, characterized in that, The storage medium stores program files capable of implementing the video fusion method based on video foreground extraction as described in any one of claims 1 to 5.
9. A processor, characterized in that, The processor is used to run a program, wherein the program executes the video fusion method based on video foreground extraction as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Virtual and actual reality integration method of multiple video streams and three-dimensional scene
CN104599243A
Three-dimensional reconstruction method for dynamic target in static scene
CN111524233A