Three-dimensional video fusion method and application thereof in video security system

Through the three-dimensional video fusion method, the multi-scale optical flow and feature point matching technology are used to build a static and dynamic three-dimensional scene model, solving the problem of single viewing angle and environmental interference of two-dimensional video surveillance, and achieving high-precision and high-reliability security monitoring.

CN120279093AActive Publication Date: 2025-07-08YUNTU DATA TECH (ZHENGZHOU) CO LTD

Patent Information

Application Number
CN202510365577.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-08
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

In the existing video security system, two-dimensional video surveillance has a single viewing angle and is susceptible to complex environment interference, resulting in blind spots and missing information, making it difficult to meet the security requirements of high reliability and high accuracy. The existing video fusion method cannot build a depth and three-dimensional monitoring scenario.

Method used

The three-dimensional video fusion method is adopted to obtain multi-angle video sequences through high-precision three-dimensional fusion of multi-scale optical flow, and static and dynamic three-dimensional scene models are constructed, and weighted fusion is performed to generate high-quality three-dimensional videos.

Benefits of technology

It improves the accuracy and real-time nature of security monitoring, provides more comprehensive and accurate scene information, and meets the requirements of high reliability of video security systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279093A_ABST
    Figure CN120279093A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video processing, and discloses a three-dimensional video fusion method and application thereof in a video security system. The method comprises the following steps: processing a static scene and a dynamic scene, constructing a three-dimensional scene model, obtaining an optical flow field corresponding to each layer of image of an image pyramid, obtaining multi-scale optical flow information, and correspondingly matching image feature points of a video frame with image feature points of the three-dimensional scene model to generate a feature point matching result, the corresponding relation between the video frame and the three-dimensional scene model is found according to the feature point matching result, accurate position information is provided for subsequent fusion, the fusion weight is calculated according to the feature point matching result, the fused three-dimensional video is obtained through processing according to the fusion weight, and the fused three-dimensional video can be used in a video security and protection system. More comprehensive and accurate scene information can be provided, the precision and real-time performance of security and protection monitoring can be improved, and the requirement of a video security and protection system for high reliability is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video processing, and in particular, to a three-dimensional video fusion method and its application in a video security system. Background Art

[0002] In today's digital age, the demand for security is increasing day by day. As a key defense line for protecting people and property, the importance of the video security system is self-evident. With the continuous expansion of the city scale, the increasing complexity of public places, and the gradual diversification of criminal means, in the prior art, a video security system integrating multiple key technologies such as video surveillance, alarm, access control management, IP intercom, power environment acquisition, patrol management, and fire protection has become an important defense line for ensuring security. However, the traditional video security system faces severe challenges in the data processing link because the data sources at the bottom layer are extensive, making the data extremely complex. At the same time, these data show significant heterogeneity, resulting in great difficulties in data integration and analysis. A large amount of valuable data is submerged in the complex information and cannot be efficiently utilized. In addition, the two-dimensional video surveillance used in the existing video security system has problems such as a single perspective and being easily interfered by complex environments. When facing situations such as light changes and obstacles, it is extremely easy to have monitoring blind spots and information loss, and it is impossible to accurately monitor the target area, making it difficult to meet the current requirements for high reliability and high accuracy of the security system.

[0003] In response to these problems, the existing common video fusion methods mainly focus on the splicing and overlay of two-dimensional video images. Although it can expand the monitoring perspective to a certain extent, it cannot construct a monitoring scene with depth and three-dimensionality, and the judgment of the spatial position and movement trajectory of objects is not accurate enough. It lacks in-depth consideration of scene details and environmental factors, and is difficult to play an effective role in complex security scenarios, and cannot provide continuous and reliable support for the security system. Summary of the Invention

[0004] This application provides a three-dimensional video fusion method and its application in a video security system. Through multi-scale optical flow high-precision three-dimensional fusion, the requirements for high reliability of security are achieved, and the high precision and high real-time performance of security monitoring are improved.

[0005] In a first aspect, the present application provides a three-dimensional video fusion method, which includes: obtaining video sequences of multiple angles of a target scene, where the video sequences include a number of video frames; when the target scene is a static scene, analyzing the video sequences through a structured light three-dimensional reconstruction module to obtain a static three-dimensional mesh model, and when the target scene is a dynamic scene, analyzing the video sequences through structure from motion to obtain a dynamic three-dimensional mesh model, and constructing a three-dimensional scene model based on the static three-dimensional mesh model and the dynamic three-dimensional mesh model; constructing an image pyramid for each of the video frames, obtaining the optical flow field corresponding to each layer of the image of the image pyramid through an optical flow algorithm, and fusing the optical flow fields through a multi-scale convolutional neural network to obtain multi-scale optical flow information; extracting corresponding image feature points from the video frames and the three-dimensional scene model and combining the multi-scale optical flow information, and matching the image feature points of the video frames with the image feature points of the three-dimensional scene model to generate a feature point matching result; calculating a fusion weight according to the feature point matching result, and performing weighted fusion on the video frames and the three-dimensional scene model according to the fusion weight to obtain a fused three-dimensional video.

[0006] In a second aspect, the present application provides an application of the three-dimensional video fusion method in a video security system, where the video security system includes: an acquisition module, which is used to obtain video sequences of multiple angles of a target scene, and the video sequences include a number of video frames; a construction module, when the target scene is a static scene, analyzing the video sequences through a structured light three-dimensional reconstruction module to obtain a static three-dimensional mesh model, and when the target scene is a dynamic scene, analyzing the video sequences through structure from motion to obtain a dynamic three-dimensional mesh model, and the construction module is used to construct a three-dimensional scene model based on the static three-dimensional mesh model and the dynamic three-dimensional mesh model; a multi-scale optical flow module, which is used to construct an image pyramid for each of the video frames, obtain the optical flow field corresponding to each layer of the image of the image pyramid through an optical flow algorithm, and fuse the optical flow fields through a multi-scale convolutional neural network to obtain multi-scale optical flow information; a feature point matching module, which is used to extract corresponding image feature points from the video frames and the three-dimensional scene model and combine the multi-scale optical flow information, and match the image feature points of the video frames with the image feature points of the three-dimensional scene model to generate a feature point matching result; a fusion module, which is used to calculate a fusion weight according to the feature point matching result, and perform weighted fusion on the video frames and the three-dimensional scene model according to the fusion weight to obtain a fused three-dimensional video.

[0007] In the technical solution provided by this application, after obtaining video sequences from multiple angles of a target scene, a three-dimensional scene model is constructed by separately processing and combining the static scene and the dynamic scene in the target scene. The static scene in the target scene is analyzed by a structured light three-dimensional reconstruction module on the video sequences to obtain a static three-dimensional mesh model. The dynamic scene in the target scene is analyzed by a structure from motion on the video sequences to obtain a dynamic three-dimensional mesh model. Then, a three-dimensional scene model is constructed based on the static three-dimensional mesh model and the dynamic three-dimensional mesh model, and a complete three-dimensional scene model containing static and dynamic elements can be obtained, providing a unified three-dimensional space framework for subsequent video fusion. In addition, an image pyramid is constructed for each video frame, the optical flow field corresponding to each layer of the image in the image pyramid is obtained through an optical flow algorithm, and the optical flow fields are fused through a multi-scale convolutional neural network to obtain multi-scale optical flow information. Images of different layers can capture objects and motion details of different sizes, which helps to improve the accuracy of optical flow calculation. The fused multi-scale optical flow information can provide a more reliable basis for subsequent feature point matching. Then, corresponding image feature points are extracted from the video frames and the three-dimensional scene model and combined with the multi-scale optical flow information, and the image feature points of the video frames are matched with the image feature points of the three-dimensional scene model to generate a feature point matching result. The extraction of feature points can reduce the data volume and help to improve the efficiency of subsequent matching. Through the feature point matching result, the corresponding relationship between the video frames and the three-dimensional scene model can be found, providing accurate position information for subsequent fusion. Finally, the fusion weight is calculated according to the feature point matching result, and the video frames and the three-dimensional scene model are weighted and fused according to the fusion weight to obtain a fused three-dimensional video. Among them, the fusion weight calculated according to the feature point matching result not only reflects the importance of the video frames and the three-dimensional scene model at different positions, but also more reasonably distributes the contributions of the video frames and the three-dimensional scene model, making the fused image more natural and realistic. The fused three-dimensional video presents a richer and more vivid visual effect. In a video security system, it can provide more comprehensive and accurate scene information, helping to improve the accuracy and real-time performance of security monitoring and meeting the requirements of the video security system for high reliability. Description of the Drawings

[0008] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0009] Figure 1 Schematic diagram of an embodiment of the three-dimensional video fusion method in the embodiment of this application;

[0010] Figure 2 This is a schematic diagram of an embodiment of a video security system for the application of the three-dimensional video fusion method in the video security system in the embodiment of the present application. Detailed implementation manners

[0011] The embodiment of the present application provides a three-dimensional video fusion method and its application in a video security system. Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0012] For ease of understanding, the specific process of the embodiment of the present application is described below. Please refer to Figure 1 An embodiment of the three-dimensional video fusion method in the embodiment of the present application includes:

[0013] Step S101, obtaining video sequences of multiple angles of a target scene, where the video sequences include a plurality of video frames;

[0014] Step S102, when the target scene is a static scene, analyzing the video sequence through a structured light three-dimensional reconstruction module to obtain a static three-dimensional mesh model, and when the target scene is a dynamic scene, analyzing the video sequence through structure from motion to obtain a dynamic three-dimensional mesh model, and constructing a three-dimensional scene model based on the static three-dimensional mesh model and the dynamic three-dimensional mesh model;

[0015] Step S103, constructing an image pyramid for each of the video frames, obtaining the optical flow field corresponding to each layer of the image of the image pyramid through an optical flow algorithm, and fusing the optical flow fields through a multi-scale convolutional neural network to obtain multi-scale optical flow information;

[0016] Step S104, extracting corresponding image feature points from the video frames and the three-dimensional scene model and combining the multi-scale optical flow information, and matching the image feature points of the video frames with the image feature points of the three-dimensional scene model to generate a feature point matching result;

[0017] Step S105: Calculate the fusion weights according to the feature point matching results, and perform weighted fusion on the video frame and the three-dimensional scene model according to the fusion weights to obtain a fused three-dimensional video.

[0018] It can be understood that the execution subject of this application can be a video security system, or a terminal or a server. Specifically, it is not limited here.

[0019] In a specific embodiment, obtaining video sequences of a target scene from multiple angles is the starting step of the entire three-dimensional video fusion method, providing basic data for subsequent three-dimensional reconstruction and fusion. To obtain multi-angle information of the target scene, it is necessary to reasonably arrange video acquisition devices. According to the shape and characteristics of the scene, cameras are distributed at different positions and angles. According to factors such as the lighting conditions and motion speed of the target scene, appropriate acquisition parameters are set, such as frame rate, exposure time, gain, etc. For a fast-moving scene, the frame rate needs to be increased to capture clear motion trajectories. For a scene with relatively dim lighting, the exposure time or gain needs to be appropriately increased, but attention should be paid to avoiding problems such as overexposure or excessive noise in the image.

[0020] During the acquisition process, it is necessary to ensure the stable operation of the camera to avoid video frame jitter caused by vibration or other interferences. At the same time, information such as the acquisition time, location, and scene description should be recorded for subsequent processing and analysis. The acquired video data is stored in a suitable storage device, such as a hard disk array, cloud storage, etc. It is necessary to ensure that the storage device has sufficient capacity and reliable performance to prevent data loss or damage.

[0021] In a specific embodiment, the structured light three-dimensional reconstruction module projects a structured light pattern with encoded information, such as a sine stripe, onto a static target scene through a specific projection device. In each frame of the video frame, the camera captures the reflected structured light pattern. Through image preprocessing steps, including denoising and enhancing contrast, it prepares for accurately extracting the structured light encoded information. Finally, the structured light encoded information is extracted from the preprocessed image through a decoding algorithm. After extracting the structured light encoded information, the absolute phase value of each pixel in the image is calculated through a phase unwrapping algorithm. Taking the sine stripe projection as an example, the initially obtained is the wrapped phase value. Through the phase unwrapping algorithm, the wrapped phase is unwrapped into the absolute phase to obtain continuous and accurate phase information.

[0022] According to the calculated absolute phase value of each pixel, combined with the calibration parameters of the structured light, the three-dimensional spatial coordinates corresponding to each pixel are calculated using the principle of triangulation. Triangulation is based on the propagation and projection relationship of light rays. By using the known position information of the camera and the projector, the pixel points on the two-dimensional image are mapped into the three-dimensional space. The three-dimensional spatial coordinates of all pixels are combined to form point cloud data. Point cloud data is a set composed of a large number of discrete three-dimensional points, and each point represents the information of a position in the scene.

[0023] Use the Poisson surface reconstruction algorithm to process the generated point cloud data to generate a static three-dimensional mesh model. The Poisson surface reconstruction algorithm regards the point cloud data as a set of directed sampled points, estimates an implicit surface function by solving the Poisson equation, and then generates a continuous three-dimensional mesh model, which can well handle the noise and incompleteness in the point cloud data and generate a smooth and continuous surface.

[0024] For a dynamic scene, use the structure from motion algorithm to process the video frames of the dynamic scene. First, use the feature extraction algorithm to extract the feature points of each frame of the image. These feature points have unique local features and can be stably detected in different image frames. Then, use the K-nearest neighbor algorithm to match the feature points of adjacent video frame images. The K-nearest neighbor algorithm finds the K nearest neighbor points of each feature point in the adjacent frame by calculating the distance between the feature points, and filters out reliable matching point pairs according to the distance threshold. Based on the matched feature point pairs, combined with the calibration parameters of the camera and the principle of multi-view geometry, calculate the three-dimensional coordinates of the feature points by the triangulation method. The triangulation method uses the projection relationship of the feature points under different perspectives to solve their positions in the three-dimensional space.

[0025] Combine the calculated three-dimensional coordinates of the feature points into a sparse point cloud. Since the sparse point cloud only contains the three-dimensional information of some points with obvious features in the image, in order to obtain a more complete three-dimensional model, use the multi-view stereo network algorithm to further generate a dense point cloud from the sparse point cloud. The multi-view stereo network algorithm uses the image information from multiple perspectives and, through techniques such as stereo matching, supplements more three-dimensional points on the basis of the sparse point cloud to obtain a denser and more complete point cloud data. Also use the Poisson surface reconstruction algorithm to process the dense point cloud to generate a dynamic three-dimensional mesh model. Poisson surface reconstruction can convert the discrete point cloud data into a continuous three-dimensional mesh, making the model have better visualization effects and convenience for subsequent processing.

[0026] Align the static three-dimensional mesh model and the dynamic three-dimensional mesh model to the same coordinate system through the iterative closest point algorithm (ICP algorithm). The ICP algorithm continuously iterates to find the optimal rotation and translation transformation between the two models, making the sum of the distances of the corresponding points between them the smallest. The specific steps are as follows:

[0027] Select an initial rotation matrix and translation vector;

[0028] For each vertex in the static 3D mesh model, find the vertex in the dynamic 3D mesh model that is closest to it, forming corresponding point pairs;

[0029] According to the current corresponding point pairs, use the least squares method to calculate the rotation matrix and translation vector that can make the two models closer;

[0030] Use the calculated rotation matrix and translation vector to transform one of the models;

[0031] Repeat the above steps until the convergence condition is met.

[0032] Use the Iterative Closest Point (ICP) algorithm again to match the boundary vertices of the meshes of the static 3D mesh model and the dynamic 3D mesh model. Find the corresponding vertices on the boundaries of the two models through the ICP algorithm, providing a basis for subsequent triangulation and merging operations.

[0033] Perform Delaunay triangulation on the matched boundary vertices to generate triangular patches in the stitching area. Delaunay triangulation can ensure that the generated triangular patches are as close to equilateral triangles as possible, avoiding the appearance of long and narrow triangles, thereby improving the quality of the mesh. Merge the boundary vertices of the meshes of the static 3D mesh model and the dynamic 3D mesh model, removing duplicate vertices to ensure the consistency of the mesh.

[0034] Unify the texture coordinates of the meshes of the static 3D mesh model and the dynamic 3D mesh model after merging the boundary vertices. Since the two models may have different texture mapping methods, the texture coordinates need to be adjusted so that they can correctly map the texture after merging. Use Poisson image editing technology to transition the textures of the static 3D mesh model and the dynamic 3D mesh model. Poisson image editing solves the Poisson equation to smoothly transition the textures of the two models, eliminating the obvious boundaries at the texture stitching points and making the textures look more natural.

[0035] Connect the meshes of the static 3D mesh model and the dynamic 3D mesh model after transitioning the textures according to the generated triangular patches to construct a complete 3D scene model. During the connection process, ensure that the topological structure of the mesh is correct and the connection relationship between vertices and edges is reasonable, thereby obtaining a high-quality 3D scene model containing static and dynamic elements.

[0036] An image pyramid is a method for representing an image at multiple scales. By performing Gaussian blur and downsampling operations on the original image, a series of image layers with different resolutions are generated. Just like a pyramid, the resolution gets lower as you go up. This multi-scale representation helps capture image features and motion information at different scales, improving the accuracy and robustness of subsequent optical flow calculation and feature fusion.

[0037] Apply a Gaussian filter to the input video frame image. The Gaussian filter replaces each pixel value in the image with the weighted average of its neighboring pixels through a convolution operation, and the weights are determined by the Gaussian distribution. The kernel size and standard deviation of the Gaussian filter are important parameters, usually selected according to the characteristics of the image and application requirements. For example, for an image with more noise, a larger kernel size and standard deviation can be chosen to enhance the smoothing effect. Perform a downsampling operation on the Gaussian-blurred image, usually by sampling every other row and every other column, that is, only retain the pixels in the image that are every other row and every other column, thus reducing the resolution of the image by half. Repeat the above Gaussian blur and downsampling operations on the first-layer image to obtain the second-layer image, and perform Gaussian blur and downsampling operations on the second-layer image again to obtain the third-layer image.

[0038] Optical flow refers to the trend of pixel gray value changes caused by the movement of objects in an image. The optical flow field describes the movement speed and direction of each pixel in the image. By calculating the optical flow fields of image layers with different resolutions, the movement information of objects can be captured at different scales, providing a basis for subsequent multi-scale feature fusion.

[0039] Calculate the optical flow fields corresponding to the first-layer image, the second-layer image, and the third-layer image respectively for the first-layer image, the second-layer image, and the third-layer image through an optical flow algorithm. The optical flow fields with different resolutions contain movement information at different scales, but their resolutions are different and cannot be directly fused. They need to be upsampled to the same resolution first, then use convolution kernels of different sizes to extract their respective features, and finally perform dynamic weighted fusion through the squeeze-and-excitation (SE) module of a multi-scale convolutional neural network to generate multi-scale optical flow information with consistent image features. Use an interpolation method to upsample the optical flow fields corresponding to the second-layer image and the third-layer image so that their resolutions are the same as the optical flow field corresponding to the first-layer image. For the aligned optical flow fields, perform convolution operations using convolution kernels of different sizes to extract image features. Use the squeeze-and-excitation (SE) module to perform dynamic weighted fusion on the three feature maps. The SE module adaptively learns the importance weights of each channel through global average pooling, fully connected layers, and the Sigmoid activation function, and then applies these weights to the feature maps.

[0040] By processing video frames using the Scale-Invariant Feature Transform (SIFT) algorithm, a set of feature points is obtained through scale-space extreme value detection, feature point localization, orientation assignment, and feature descriptor generation. First, the three-dimensional scene model is projected into a two-dimensional image according to the camera parameters, and then the SIFT algorithm is used to extract the set of feature points of the projected image. Combining multi-scale optical flow information, the positions of the feature points in the video frames are predicted in the projected image, narrowing the matching search range. The KNN algorithm is used to calculate the distances of the feature descriptors, and reliable matching pairs are screened out. The RANSAC algorithm is used to calculate the homography matrix, which is optimized through multiple iterations. Inliers are screened out based on the projection error, and the matching pairs that pass the verification are the final feature point matching results. The accurate and reliable feature point matching results provide the key corresponding relationships for subsequent three-dimensional video fusion. Through these matching pairs, the spatial position relationship between the video frames and the three-dimensional scene model can be determined, thus achieving their accurate fusion and generating high-quality three-dimensional videos. For example, in a security monitoring system, real-time videos can be fused with the three-dimensional scene model to provide more comprehensive and accurate scene information, improving the accuracy and real-time performance of monitoring.

[0041] Perform Gaussian difference operations on the video frame images to construct a scale space. Search for local extreme points in the scale space. These points may be candidate feature points. The positions and scales of the feature points are accurately determined by fitting a three-dimensional quadratic function. At the same time, points with low contrast and edge responses are removed to improve the stability of the feature points. One or more main directions are assigned to each feature point. Based on the gradient direction histogram within the neighborhood of the feature point, a neighborhood is selected around the feature point, which is divided into multiple sub-regions. The gradient direction histogram within each sub-region is calculated, and these histograms are combined into a high-dimensional vector as the feature descriptor of the feature point.

[0042] Project the three-dimensional scene model onto the perspective of the video frame to generate a projected image, and then use the SIFT algorithm to extract the feature points in the projected image, realizing feature matching between the video frame and the three-dimensional scene model from the same perspective. Specifically, according to the camera parameters of the video frame, the three-dimensional scene model is projected onto a two-dimensional plane to generate a projected image. The SIFT algorithm is used to extract the feature points of the projected image. The steps are the same as those for extracting the feature points of the video frame and will not be elaborated here.

[0043] Multi-scale optical flow information can predict the motion trajectories of feature points between different frames, narrow down the search range for feature point matching, and improve the accuracy and efficiency of matching. The K-Nearest Neighbor (KNN) algorithm finds the nearest K neighbors by calculating the distances between feature point descriptors, thereby determining the matching pairs. Specifically, based on the multi-scale optical flow information, predict the possible positions of feature points in the next frame of the video frame. Use the KNN algorithm to match the set of feature points in the video frame and the set of feature points in the 3D scene model. Usually, take K = 2, that is, find the two nearest neighbors for each feature point. By comparing the distance ratio between the nearest neighbor and the second nearest neighbor, filter out reliable matching pairs. If the distance ratio between the nearest neighbor and the second nearest neighbor is less than a certain threshold (such as 0.8), then this matching pair is considered reliable.

[0044] The homography matrix describes the projective transformation relationship between two planes. If the matching pairs are correct, then they satisfy the homography transformation. By calculating the homography matrix and verifying the consistency of the matching pairs, mis-matched points can be removed. Specifically, use the Random Sample Consensus (RANSAC) algorithm to calculate the homography matrix from the matching pairs. The RANSAC algorithm randomly selects a set of matching pairs, calculates the homography matrix, and then verifies the consistency of other matching pairs according to this matrix. Iterate continuously until the optimal homography matrix is found. According to the calculated homography matrix, verify each matching pair. If the projection error of the matching pair is less than a certain threshold, then this matching pair is considered consistent.

[0045] The confidence score of the matching pair can be calculated based on factors such as the matching distance of the feature point descriptors and the projection error of the homography matrix. The smaller the matching distance and the projection error, the more reliable the matching pair, and the higher the confidence score. For each matching pair, calculate the distance between its feature point descriptors, normalize the distance, and then take the reciprocal as the matching distance score. Calculate the projection error of each matching pair according to the homography matrix, normalize the error, and then take the reciprocal as the projection error score. Perform a weighted average on the matching distance score and the projection error score to obtain the confidence score of the matching pair.

[0046] The local weight considers the distribution of the matching pairs in the local area, and the global weight considers the overall situation of all matching pairs. By calculating the local weight and the global weight, the importance of the matching pairs can be more comprehensively reflected. Specifically, divide the image into multiple local areas, count the sum of the confidence scores of the matching pairs in each local area, and then perform normalization to obtain the local weight. Calculate the sum of the confidence scores of all matching pairs, and then perform normalization to obtain the global weight. Obtain the fusion weight by assigning different coefficients to the local weight and the global weight. The selection of the coefficients can be adjusted according to the specific application scenario. The specific implementation formula is:

[0047] W = α * w local + β * w glocal

[0048] Wherein, W is the fusion weight, w local is the local weight, w glocal is the global weight, α is the coefficient of the local weight, β is the coefficient of the global weight, and α + β = 1.

[0049] At the pixel position of each image, the pixel values of the video frame and the pixel values of the image in the three-dimensional scene model are weighted and fused according to the fusion weight to obtain the fused pixel values, and then these fused pixel values are sorted into a sequence to form a fused three-dimensional video. According to the pixel values of the video frame and the pixel values of the image in the three-dimensional scene model, the fused pixel values are calculated. The specific implementation formula is as follows:

[0050] I fused = W * I model + (1 - W) * I video

[0051] Wherein, I fused is the fused pixel value, I video is the pixel value of the video frame, I model is the pixel value of the image in the three-dimensional scene model, and W is the fusion weight. By traversing each pixel of the video frame and the image in the three-dimensional scene model, the fused pixel values are calculated according to the fusion weight, and the fused pixel values are sorted into a sequence to form the three-dimensional video, and are saved as a video file through a video processing library.

[0052] In this embodiment, specifically, a three-dimensional scene model is constructed by separately processing and combining the static scene and the dynamic scene in the target scene. The static scene in the target scene is analyzed by a structured light three-dimensional reconstruction module on the video sequence to obtain a static three-dimensional mesh model. The dynamic scene in the target scene is analyzed by structure from motion on the video sequence to obtain a dynamic three-dimensional mesh model. Then, a three-dimensional scene model is constructed based on the static three-dimensional mesh model and the dynamic three-dimensional mesh model, and a complete three-dimensional scene model containing static and dynamic elements can be obtained, providing a unified three-dimensional space framework for subsequent video fusion. In addition, an image pyramid is constructed for each video frame, the optical flow field corresponding to each layer of the image of the image pyramid is obtained through the optical flow algorithm, and the multi-scale optical flow information is obtained by fusing the optical flow fields through a multi-scale convolutional neural network. Images of different layers can capture objects and motion details of different sizes, which helps to improve the accuracy of optical flow calculation. The fused multi-scale optical flow information can provide a more reliable basis for subsequent feature point matching. Then, corresponding image feature points are extracted from the video frame and the three-dimensional scene model and combined with the multi-scale optical flow information, and the image feature points of the video frame are matched with the image feature points of the three-dimensional scene model to generate a feature point matching result. The extraction of feature points can reduce the data volume and help to improve the efficiency of subsequent matching. The corresponding relationship between the video frame and the three-dimensional scene model can be found through the feature point matching result, providing accurate position information for subsequent fusion. Finally, the fusion weight is calculated according to the feature point matching result, and the video frame and the three-dimensional scene model are weighted and fused according to the fusion weight to obtain a fused three-dimensional video. Among them, the fusion weight calculated according to the feature point matching result not only reflects the importance of the video frame and the three-dimensional scene model at different positions, but also more reasonably distributes the contributions of the video frame and the three-dimensional scene model, making the fused image more natural and real. The fused three-dimensional video presents a more rich and vivid visual effect. In a video security system, it can provide more comprehensive and accurate scene information, helping to improve the accuracy and real-time performance of security monitoring and meeting the requirements of the video security system for high reliability.

[0053] In a specific embodiment, obtaining the video sequences of the target scene from multiple perspectives is the starting step of the entire 3D video fusion method, providing the basic data for subsequent 3D reconstruction and fusion. To obtain the multi-angle information of the target scene, it is necessary to reasonably arrange the video acquisition devices. According to the shape and characteristics of the scene, the cameras are distributed at different positions and angles. According to factors such as the lighting conditions and motion speed of the target scene, appropriate acquisition parameters are set, such as frame rate, exposure time, gain, etc. For a fast-moving scene, it is necessary to increase the frame rate to capture clear motion trajectories. For a scene with relatively dim lighting, it is necessary to appropriately increase the exposure time or gain, but attention should be paid to avoiding problems such as overexposure or excessive noise in the image.

[0054] During the acquisition process, it is necessary to ensure the stable operation of the cameras to avoid video frame jitter caused by vibration or other interferences. At the same time, information such as the acquisition time, location, and scene description should be recorded for subsequent processing and analysis. The acquired video data is stored in a suitable storage device, such as a hard disk array, cloud storage, etc. It is necessary to ensure that the storage device has sufficient capacity and reliable performance to prevent data loss or damage.

[0055] In a specific embodiment, the process of executing step S102 may specifically include the following steps:

[0056] (1) When the target scene is a static scene, the structured light 3D reconstruction module is used to extract the structured light coding information in each frame image of the video frame, and the absolute phase value of each pixel in the image is calculated through the phase unwrapping algorithm;

[0057] (2) According to the absolute phase value, the three-dimensional space coordinates corresponding to each pixel are calculated, the three-dimensional space coordinates of the pixels are combined into point cloud data, and the Poisson surface reconstruction is used to generate a static 3D mesh model based on the point cloud data;

[0058] (3) When the target scene is a dynamic scene, the structure from motion is used to extract the feature points in each frame image of the video frame and match the feature points in the images of adjacent video frames through the K-nearest neighbor algorithm;

[0059] (4) According to the feature points, the three-dimensional coordinates of the feature points are calculated, the three-dimensional coordinates of the feature points are combined into a sparse point cloud, the multi-view stereo network algorithm is used to generate a dense point cloud from the sparse point cloud, and the Poisson surface reconstruction is used to generate a dynamic 3D mesh model from the dense point cloud;

[0060] (5) The static 3D mesh model and the dynamic 3D mesh model are aligned to the same coordinate system through the iterative closest point algorithm, and the static 3D mesh model and the dynamic 3D mesh model are merged using the mesh stitching algorithm to construct the 3D scene model.

[0061] Specifically, merging the static three-dimensional mesh model and the dynamic three-dimensional mesh model using the mesh stitching algorithm to construct the three-dimensional scene model includes:

[0062] (1) Matching the boundary vertices of the meshes of the static three-dimensional mesh model and the dynamic three-dimensional mesh model through the iterative closest point algorithm;

[0063] (2) Performing Delaunay triangulation on the matched boundary vertices to generate triangular patches in the stitching area, and merging the boundary vertices of the meshes of the static three-dimensional mesh model and the dynamic three-dimensional mesh model;

[0064] (3) Unifying the texture coordinates of the meshes of the static three-dimensional mesh model and the dynamic three-dimensional mesh model after merging the boundary vertices, and using Poisson image editing to transition the textures of the static three-dimensional mesh model and the dynamic three-dimensional mesh model;

[0065] (4) Constructing the three-dimensional scene model according to the meshes of the static three-dimensional mesh model and the dynamic three-dimensional mesh model that are connected and transitioned with the textures according to the triangular patches.

[0066] In a specific embodiment, the structured light three-dimensional reconstruction module projects a structured light pattern with encoded information, such as a sine stripe, onto a static target scene through a specific projection device. In each frame image of the video frame, the camera captures the reflected structured light pattern. Through image preprocessing steps, including denoising and enhancing contrast, it prepares for accurately extracting the structured light encoded information in the subsequent steps. Finally, through a decoding algorithm, the structured light encoded information is extracted from the preprocessed image. After extracting the structured light encoded information, through a phase unwrapping algorithm, the absolute phase value of each pixel in the image is calculated. Taking the sine stripe projection as an example, initially the wrapped phase value is obtained, and through the phase unwrapping algorithm, the wrapped phase is unwrapped into the absolute phase to obtain continuous and accurate phase information.

[0067] According to the absolute phase value of each pixel calculated, combined with the calibration parameters of the structured light, the three-dimensional space coordinates corresponding to each pixel are calculated using the principle of triangulation. Triangulation is based on the propagation and projection relationship of light rays. Through the known position information of the camera and the projector, the pixel points on the two-dimensional image are mapped into the three-dimensional space. Combining the three-dimensional space coordinates of all pixels forms point cloud data. Point cloud data is a set composed of a large number of discrete three-dimensional points, and each point represents the information of a position in the scene.

[0068] The generated point cloud data is processed using the Poisson surface reconstruction algorithm to generate a static three-dimensional mesh model. The Poisson surface reconstruction algorithm regards the point cloud data as a directed set of sampled points, estimates an implicit surface function by solving the Poisson equation, and then generates a continuous three-dimensional mesh model, which can well handle the noise and incompleteness in the point cloud data and generate a smooth and continuous surface.

[0069] For dynamic scenes, the Structure from Motion algorithm is used to process the video frames of the dynamic scene. First, the feature extraction algorithm is used to extract the feature points of each frame of the image. These feature points have unique local features and can be stably detected in different image frames. Then, the K-Nearest Neighbor algorithm is used to match the feature points of adjacent video frame images. The K-Nearest Neighbor algorithm finds the K nearest neighbor points of each feature point in the adjacent frame by calculating the distance between the feature points, and filters out reliable matching point pairs according to the distance threshold. Based on the matched feature point pairs, combined with the calibration parameters of the camera and the principles of multi-view geometry, the three-dimensional coordinates of the feature points are calculated by the triangulation method. The triangulation method uses the projection relationship of the feature points under different perspectives to solve their positions in the three-dimensional space.

[0070] The three-dimensional coordinates of the calculated feature points are combined into a sparse point cloud. Since the sparse point cloud only contains the three-dimensional information of some points with obvious features in the image, in order to obtain a more complete three-dimensional model, the Multi-View Stereo Network algorithm is used to further generate a dense point cloud from the sparse point cloud. The Multi-View Stereo Network algorithm uses the image information from multiple perspectives and, through techniques such as stereo matching, supplements more three-dimensional points on the basis of the sparse point cloud to obtain a denser and more complete point cloud data. Similarly, the Poisson surface reconstruction algorithm is used to process the dense point cloud to generate a dynamic three-dimensional mesh model. Poisson surface reconstruction can convert the discrete point cloud data into a continuous three-dimensional mesh, making the model have better visualization effects and convenience for subsequent processing.

[0071] The static three-dimensional mesh model and the dynamic three-dimensional mesh model are aligned to the same coordinate system through the Iterative Closest Point algorithm (ICP algorithm). The ICP algorithm continuously iterates to find the optimal rotation and translation transformation between the two models, making the sum of the distances between their corresponding points the smallest. The specific steps are as follows:

[0072] Select an initial rotation matrix and translation vector;

[0073] For each vertex in the static three-dimensional mesh model, find the vertex in the dynamic three-dimensional mesh model that is closest to it to form a corresponding point pair;

[0074] According to the current corresponding point pairs, use the least squares method to calculate the rotation matrix and translation vector that can make the two models closer;

[0075] Transform one of the models using the calculated rotation matrix and translation vector;

[0076] Repeat the above steps until the convergence condition is met.

[0077] Use the iterative closest point algorithm to match the boundary vertices of the static 3D mesh model and the dynamic 3D mesh model again. Find the corresponding vertices on the boundaries of the two models through the iterative closest point algorithm, providing a basis for subsequent triangulation and merging operations.

[0078] Perform Delaunay triangulation on the matched boundary vertices to generate triangular patches in the stitching area. Delaunay triangulation can ensure that the generated triangular patches are as close to equilateral triangles as possible, avoiding the appearance of long and narrow triangles, thereby improving the quality of the mesh. Merge the boundary vertices of the static 3D mesh model and the dynamic 3D mesh model, removing duplicate vertices to ensure the consistency of the mesh.

[0079] Unify the texture coordinates of the meshes of the static 3D mesh model and the dynamic 3D mesh model after merging the boundary vertices. Since the two models may have different texture mapping methods, the texture coordinates need to be adjusted so that they can correctly map the texture after merging. Use Poisson image editing technology to transition the textures of the static 3D mesh model and the dynamic 3D mesh model. Poisson image editing solves the Poisson equation to smoothly transition the textures of the two models, eliminating the obvious boundaries at the texture seams and making the textures look more natural.

[0080] Connect the meshes of the static 3D mesh model and the dynamic 3D mesh model with the transitioned texture according to the generated triangular patches to construct a complete 3D scene model. During the connection process, ensure that the topological structure of the mesh is correct and the connection relationship between vertices and edges is reasonable, thereby obtaining a high-quality 3D scene model containing static and dynamic elements.

[0081] In a specific embodiment, the process of executing step S103 may specifically include the following steps:

[0082] (1) Perform Gaussian blur and downsampling on the image of the video frame to generate a multi-layer image, and the multi-layer image forms the image pyramid. The image pyramid includes at least the first-layer image, the second-layer image, and the third-layer image. The resolution of the first-layer image is greater than that of the second-layer image, and the resolution of the second-layer image is greater than that of the third-layer image;

[0083] (2) Calculate the optical flow fields corresponding to the first-layer image, the second-layer image, and the third-layer image respectively for the first-layer image, the second-layer image, and the third-layer image through the optical flow algorithm;

[0084] (3) Generate multi-scale optical flow information with consistent image features by fusing the optical flow fields corresponding to the first-layer image, the optical flow fields corresponding to the second-layer image, and the optical flow fields corresponding to the third-layer image through a multi-scale convolutional neural network.

[0085] Specifically, performing Gaussian blur and downsampling on the images of the video frames to generate multiple layers of images, the multiple layers of images forming the image pyramid, the image pyramid at least including a first-layer image, a second-layer image, and a third-layer image, the resolution of the first-layer image being greater than the resolution of the second-layer image, and the resolution of the second-layer image being greater than the resolution of the third-layer image, includes:

[0086] Performing Gaussian blur and downsampling on the images of the video frames to obtain the first-layer image;

[0087] Performing Gaussian blur and downsampling on the first-layer image to obtain the second-layer image, the resolution of the second-layer image being lower than the resolution of the first-layer image;

[0088] Performing Gaussian blur and downsampling on the second-layer image to obtain the third-layer image, the resolution of the third-layer image being lower than the resolution of the second-layer image.

[0089] Specifically, the generating of multi-scale optical flow information with consistent image features by fusing the optical flow fields corresponding to the first-layer image, the optical flow fields corresponding to the second-layer image, and the optical flow fields corresponding to the third-layer image through a multi-scale convolutional neural network includes:

[0090] Aligning the optical flow fields corresponding to the first-layer image, the optical flow fields corresponding to the second-layer image, and the optical flow fields corresponding to the third-layer image to the same resolution through upsampling;

[0091] Extracting image features from the aligned optical flow fields corresponding to the first-layer image through 3×3 convolution to obtain a first feature map, extracting image features from the aligned optical flow fields corresponding to the second-layer image through 5×5 convolution to obtain a second feature map, and extracting image features from the aligned optical flow fields corresponding to the third-layer image through 7×7 convolution to obtain a third feature map;

[0092] Dynamically weighting and fusing the first feature map, the second feature map, and the third feature map through the squeeze-and-excitation module in the multi-scale convolutional neural network and then using 3×3 convolution to generate multi-scale optical flow information with consistent image features.

[0093] An image pyramid is a method for representing an image at multiple scales. By performing Gaussian blur and downsampling operations on the original image, a series of image layers with different resolutions are generated. Just like a pyramid, the resolution decreases as you go up. This multi-scale representation helps capture image features and motion information at different scales, improving the accuracy and robustness of subsequent optical flow calculation and feature fusion.

[0094] Apply a Gaussian filter to the input video frame image. The Gaussian filter replaces each pixel value in the image with the weighted average of its neighboring pixels through a convolution operation, and the weights are determined by the Gaussian distribution. The kernel size and standard deviation of the Gaussian filter are important parameters, usually selected according to the characteristics of the image and application requirements. For example, for an image with more noise, a larger kernel size and standard deviation can be selected to enhance the smoothing effect. Perform a downsampling operation on the Gaussian-blurred image, usually using an interlaced sampling method, that is, only retain the pixels in every other row and every other column of the image, thereby reducing the resolution of the image by half. Repeat the above Gaussian blur and downsampling operations on the first-layer image to obtain the second-layer image, and perform Gaussian blur and downsampling operations on the second-layer image again to obtain the third-layer image.

[0095] Optical flow refers to the trend of pixel gray value changes caused by the movement of objects in an image. The optical flow field describes the movement speed and direction of each pixel in the image. By calculating the optical flow fields of image layers with different resolutions, the movement information of objects can be captured at different scales, providing a basis for subsequent multi-scale feature fusion.

[0096] Calculate the optical flow fields corresponding to the first-layer image, the second-layer image, and the third-layer image respectively through an optical flow algorithm for the first-layer image, the second-layer image, and the third-layer image. The optical flow fields with different resolutions contain movement information at different scales, but their resolutions are different and cannot be directly fused. They need to be upsampled to the same resolution first, then use convolution kernels of different sizes to extract their respective features, and finally perform dynamic weighted fusion through the squeeze-and-excitation (SE) module of a multi-scale convolutional neural network to generate multi-scale optical flow information with consistent image features. Use an interpolation method to upsample the optical flow fields corresponding to the second-layer image and the third-layer image so that their resolutions are the same as the optical flow field corresponding to the first-layer image. For the aligned optical flow fields, perform convolution operations using convolution kernels of different sizes to extract image features. Use the squeeze-and-excitation (SE) module to perform dynamic weighted fusion on the three feature maps. The SE module adaptively learns the importance weights of each channel through global average pooling, a fully connected layer, and a Sigmoid activation function, and then applies these weights to the feature maps.

[0097] In a specific embodiment, the process of performing step S104 may specifically include the following steps:

[0098] (1) Extract feature points in the video frame using the Scale-Invariant Feature Transform (SIFT) algorithm to obtain the feature point set of the video frame;

[0099] (2) Project the 3D scene model onto the perspective of the video frame to generate a projection image, and use the SIFT algorithm to extract feature points in the projection image and generate the feature point set of the 3D scene model;

[0100] (3) Combine the multi-scale optical flow information to predict the feature point positions, and use the K-Nearest Neighbor (KNN) algorithm to match the feature point set of the video frame with the feature point set of the 3D scene model to obtain a set of matching pairs;

[0101] (4) Verify the consistency of the set of matching pairs through a homography matrix, and the set of matching pairs that pass the verification is the feature point matching result.

[0102] By processing the video frame using the Scale-Invariant Feature Transform (SIFT) algorithm, a feature point set is obtained through scale-space extreme value detection, feature point localization, orientation assignment, and feature descriptor generation. First, project the 3D scene model into a 2D image according to the camera parameters, and then use the SIFT algorithm to extract the feature point set of the projection image. Combine the multi-scale optical flow information to predict the positions of the video frame feature points in the projection image, narrow the matching search range, calculate the feature descriptor distances using the KNN algorithm, filter out reliable matching pairs, calculate the homography matrix using the RANSAC algorithm, perform multiple iterations for optimization, filter out inliers based on the projection error, and the matching pairs that pass the verification are the final feature point matching results. The accurate and reliable feature point matching results provide the key corresponding relationships for subsequent 3D video fusion. Through these matching pairs, the spatial position relationship between the video frame and the 3D scene model can be determined, thereby achieving the accurate fusion of the two and generating high-quality 3D videos. For example, in a security monitoring system, it is possible to fuse real-time video with a 3D scene model to provide more comprehensive and accurate scene information, improving the accuracy and real-time performance of monitoring.

[0103] In a specific embodiment, the process of executing step S105 may specifically include the following steps:

[0104] (1) Calculate the confidence scores of the matching pairs for the feature point matching result;

[0105] (2) Calculate the local weights and global weights for the confidence scores of the matching pairs;

[0106] (3) By assigning different coefficients to the local weight and the global weight, and then linearly adding them to obtain the fusion weight, the calculation formula is:

[0107] W = α * wlocal + β * w glocal

[0108] where W is the fusion weight, w local is the local weight, and w glocal is the global weight, α is the coefficient of the local weight, β is the coefficient of the global weight, and α + β = 1;

[0109] (4) For the video frame and the image of the three - dimensional scene model, weighted fusion is performed at each pixel position of each image according to the fusion weight to obtain the fused three - dimensional video.

[0110] Specifically, the weighted fusion at each pixel position of each image according to the fusion weight to obtain the fused three - dimensional video includes:

[0111] Based on the pixel value of the video frame and the pixel value of the image in the three - dimensional scene model, the fused pixel value is calculated, and the calculation formula is:

[0112] I fused = W * I model + (1 - W) * I video

[0113] where I fused is the fused pixel value, I video is the pixel value of the video frame, I model is the pixel value of the image in the three - dimensional scene model, and W is the fusion weight;

[0114] The calculated fused pixel values are sorted into a sequence to form the three - dimensional video.

[0115] Perform a Difference of Gaussians operation on the video frame image to construct a scale space. Search for local extreme points in the scale space. These points may be candidate feature points. The position and scale of the feature points are accurately determined by fitting a three - dimensional quadratic function. At the same time, points with low contrast and edge responses are removed to improve the stability of the feature points. Assign one or more main directions to each feature point. Based on the gradient direction histogram within the neighborhood of the feature point, select a neighborhood around the feature point, divide the neighborhood into multiple sub - regions, calculate the gradient direction histogram within each sub - region, and combine these histograms into a high - dimensional vector as the feature descriptor of the feature point.

[0116] Project the 3D scene model onto the perspective of the video frame to generate a projected image, and then use the SIFT algorithm to extract the feature points in the projected image, so as to achieve feature matching between the video frame and the 3D scene model from the same perspective. Specifically, according to the camera parameters of the video frame, project the 3D scene model onto a two-dimensional plane to generate a projected image, and use the SIFT algorithm to extract the feature points from the projected image. The steps are the same as those for extracting the feature points of the video frame and will not be elaborated here.

[0117] The multi-scale optical flow information can predict the motion trajectories of feature points between different frames, narrow down the search range of feature point matching, and improve the accuracy and efficiency of matching. The K-Nearest Neighbor (KNN) algorithm finds the nearest K neighbors by calculating the distances between feature point descriptors, thereby determining the matching pairs. Specifically, according to the multi-scale optical flow information, predict the possible positions of the feature points in the video frame in the next frame, and use the KNN algorithm to match the set of feature points in the video frame and the set of feature points in the 3D scene model. Usually, K = 2, that is, find the two nearest neighbors for each feature point. By comparing the distance ratio between the nearest neighbor and the second-nearest neighbor, filter out the reliable matching pairs. If the distance ratio between the nearest neighbor and the second-nearest neighbor is less than a certain threshold (such as 0.8), then this matching pair is considered reliable.

[0118] The homography matrix describes the projective transformation relationship between two planes. If the matching pairs are correct, then they satisfy the homography transformation. By calculating the homography matrix and verifying the consistency of the matching pairs, the mismatched points can be removed. Specifically, use the Random Sample Consensus (RANSAC) algorithm to calculate the homography matrix from the matching pairs. The RANSAC algorithm randomly selects a set of matching pairs, calculates the homography matrix, and then verifies the consistency of other matching pairs according to this matrix. Iterate continuously until the optimal homography matrix is found. According to the calculated homography matrix, verify each matching pair. If the projection error of the matching pair is less than a certain threshold, then this matching pair is considered consistent.

[0119] The confidence score of the matching pair can be calculated based on factors such as the matching distance of the feature point descriptors and the projection error of the homography matrix. The smaller the matching distance and the projection error, the more reliable the matching pair, and the higher the confidence score. For each matching pair, calculate the distance between its feature point descriptors, normalize the distance, and then take the reciprocal as the matching distance score. Calculate the projection error of each matching pair according to the homography matrix, normalize the error, and then take the reciprocal as the projection error score. Perform a weighted average on the matching distance score and the projection error score to obtain the confidence score of the matching pair.

[0120] The local weight considers the distribution of matching pairs in the local area, and the global weight considers the overall situation of all matching pairs. By calculating the local weight and the global weight, the importance of the matching pairs can be more comprehensively reflected. Specifically, the image is divided into multiple local areas, the sum of the confidence scores of the matching pairs in each local area is statistically calculated, and then normalized to obtain the local weight. The sum of the confidence scores of all matching pairs is calculated and then normalized to obtain the global weight. By assigning different coefficients to the local weight and the global weight, and then linearly adding them to obtain the fusion weight, the selection of the coefficients can be adjusted according to the specific application scenario. The calculation formula is:

[0121] W = α * w local + β * w glocal

[0122] Where, W is the fusion weight, w local is the local weight, w glocal is the global weight, α is the coefficient of the local weight, β is the coefficient of the global weight, and α + β = 1.

[0123] At the pixel position of each image, the pixel values of the video frame and the pixel values of the image in the three-dimensional scene model are weighted and fused according to the fusion weight to obtain the fused pixel values. Then, these fused pixel values are sorted into a sequence to form the fused three-dimensional video. Based on the pixel values of the video frame and the pixel values of the image in the three-dimensional scene model, the fused pixel values are calculated. The calculation formula is:

[0124] I fused = W * I model +(1 - W) * I video

[0125] Where, I fused is the fused pixel value, I video is the pixel value of the video frame, I model is the pixel value of the image in the three-dimensional scene model, and W is the fusion weight. By traversing each pixel of the video frame and the image in the three-dimensional scene model, the fused pixel values are calculated according to the fusion weight, and the fused pixel values are sorted into a sequence to form the three-dimensional video, which is saved as a video file through a video processing library.

[0126] The three-dimensional video fusion method in the embodiments of the present application is described above. Next, the application of the three-dimensional video fusion method in the embodiments of the present application in a video security system will be described. Please refer to Figure 2 , the application of the three-dimensional video fusion method in the embodiments of the present application in a video security system. The video security system includes:

[0127] An acquisition module 201, which is used to acquire a video sequence of multiple angles of a target scene, and the video sequence includes a plurality of video frames;

[0128] A construction module 202, when the target scene is a static scene, analyzes the video sequence through a structured light three-dimensional reconstruction module to obtain a static three-dimensional mesh model, and when the target scene is a dynamic scene, analyzes the video sequence through structure from motion to obtain a dynamic three-dimensional mesh model. The construction module 202 is used to construct a three-dimensional scene model based on the static three-dimensional mesh model and the dynamic three-dimensional mesh model;

[0129] A multi-scale optical flow module 203, which is used to construct an image pyramid for each video frame, obtain the optical flow field corresponding to each layer of the image of the image pyramid through an optical flow algorithm, and fuse the optical flow fields through a multi-scale convolutional neural network to obtain multi-scale optical flow information;

[0130] A feature point matching module 204, which is used to extract corresponding image feature points from the video frame and the three-dimensional scene model and combine the multi-scale optical flow information, and match the image feature points of the video frame with the image feature points of the three-dimensional scene model to generate a feature point matching result;

[0131] A fusion module 205, which is used to calculate a fusion weight according to the feature point matching result, and perform weighted fusion on the video frame and the three-dimensional scene model according to the fusion weight to obtain a fused three-dimensional video.

[0132] In the embodiments of the present application, a three-dimensional scene model is constructed by separately processing and combining the static scene and the dynamic scene in the target scene. The static scene in the target scene is analyzed by a structured light three-dimensional reconstruction module for the video sequence to obtain a static three-dimensional mesh model. The dynamic scene in the target scene is analyzed by structure from motion for the video sequence to obtain a dynamic three-dimensional mesh model. Then, a three-dimensional scene model is constructed based on the static three-dimensional mesh model and the dynamic three-dimensional mesh model, and a complete three-dimensional scene model including static and dynamic elements can be obtained, providing a unified three-dimensional space framework for subsequent video fusion. In addition, an image pyramid is constructed for each video frame, the optical flow field corresponding to each layer of the image of the image pyramid is obtained by an optical flow algorithm, and the optical flow fields are fused by a multi-scale convolutional neural network to obtain multi-scale optical flow information. Images of different layers can capture objects and motion details of different sizes, which helps to improve the accuracy of optical flow calculation. The fused multi-scale optical flow information can provide a more reliable basis for subsequent feature point matching. Then, corresponding image feature points are extracted from the video frame and the three-dimensional scene model and combined with the multi-scale optical flow information, and the image feature points of the video frame are matched with the image feature points of the three-dimensional scene model to generate a feature point matching result. The extraction of feature points can reduce the data volume and help to improve the efficiency of subsequent matching. The corresponding relationship between the video frame and the three-dimensional scene model can be found through the feature point matching result, providing accurate position information for subsequent fusion. Finally, the fusion weight is calculated according to the feature point matching result, and the video frame and the three-dimensional scene model are weighted and fused according to the fusion weight to obtain a fused three-dimensional video. Among them, the fusion weight calculated according to the feature point matching result not only reflects the importance of the video frame and the three-dimensional scene model at different positions, but also more reasonably distributes the contributions of the video frame and the three-dimensional scene model, making the fused image more natural and real. The fused three-dimensional video presents a more rich and vivid visual effect. In a video security system, it can provide more comprehensive and accurate scene information, helping to improve the accuracy and real-time performance of security monitoring and meeting the requirements of the video security system for high reliability.

[0133] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.

[0134] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A three-dimensional video fusion method, characterized in that The three-dimensional video fusion method includes: Obtaining video sequences of multiple angles of a target scene, where the video sequences include a number of video frames; When the target scene is a static scene, analyzing the video sequences through a structured light three-dimensional reconstruction module to obtain a static three-dimensional mesh model. When the target scene is a dynamic scene, analyzing the video sequences through structure from motion to obtain a dynamic three-dimensional mesh model, and constructing a three-dimensional scene model based on the static three-dimensional mesh model and the dynamic three-dimensional mesh model; Constructing an image pyramid for each of the video frames, obtaining the optical flow field corresponding to each layer of the image of the image pyramid through an optical flow algorithm, and fusing the optical flow fields through a multi-scale convolutional neural network to obtain multi-scale optical flow information; Extracting corresponding image feature points from the video frames and the three-dimensional scene model and combining the multi-scale optical flow information, and matching the image feature points of the video frames with the image feature points of the three-dimensional scene model to generate a feature point matching result; Calculating a fusion weight according to the feature point matching result, and performing weighted fusion on the video frames and the three-dimensional scene model according to the fusion weight to obtain a fused three-dimensional video.

2. The three-dimensional video fusion method according to claim 1, characterized in that The step of, when the target scene is a static scene, analyzing the video sequences through a structured light three-dimensional reconstruction module to obtain the three-dimensional geometric data of the target scene, and when the target scene is a dynamic scene, analyzing the video sequences through structure from motion to obtain the three-dimensional geometric data of the target scene, and constructing a three-dimensional scene model based on the three-dimensional geometric data, includes: When the target scene is a static scene, using the structured light three-dimensional reconstruction module to extract the structured light coding information in each frame of the video frames, and calculating the absolute phase value of each pixel of the image through a phase unwrapping algorithm; Calculating the three-dimensional spatial coordinates corresponding to each pixel according to the absolute phase value, combining the three-dimensional spatial coordinates of the pixels into point cloud data, and using Poisson surface reconstruction to generate a static three-dimensional mesh model based on the point cloud data; When the target scene is a dynamic scene, using structure from motion to extract the feature points of each frame of the video frames and matching the feature points of the images of adjacent video frames through the K-nearest neighbor algorithm; Calculating the three-dimensional coordinates of the feature points according to the feature points, combining the three-dimensional coordinates of the feature points into sparse point clouds, using a multi-view stereo network algorithm to generate dense point clouds from the sparse point clouds, and using Poisson surface reconstruction to generate a dynamic three-dimensional mesh model from the dense point clouds; Aligning the static three-dimensional mesh model and the dynamic three-dimensional mesh model to the same coordinate system through the iterative closest point algorithm, and using a mesh stitching algorithm to merge the static three-dimensional mesh model and the dynamic three-dimensional mesh model to construct the three-dimensional scene model.

3. The three-dimensional video fusion method according to claim 2, wherein The step of using a mesh stitching algorithm to merge the static three-dimensional mesh model and the dynamic three-dimensional mesh model to construct the three-dimensional scene model includes: Matching the boundary vertices of the meshes of the static three-dimensional mesh model and the dynamic three-dimensional mesh model through the iterative closest point algorithm; Perform Delaunay triangulation on the matched boundary vertices to generate triangular patches of the stitching region, and merge the boundary vertices of the meshes of the static three-dimensional mesh model and the dynamic three-dimensional mesh model; Unify the texture coordinates of the meshes of the static three-dimensional mesh model and the dynamic three-dimensional mesh model after merging the boundary vertices, and use Poisson image editing to transition the textures of the static three-dimensional mesh model and the dynamic three-dimensional mesh model; Construct the three-dimensional scene model according to the meshes of the static three-dimensional mesh model and the dynamic three-dimensional mesh model that are connected with the transitional texture through the triangular patches.

4. The three-dimensional video fusion method according to claim 1, wherein For each of the video frames, construct an image pyramid, obtain the optical flow field corresponding to each layer of the image of the image pyramid through the optical flow algorithm, and fuse the optical flow fields through a multi-scale convolutional neural network to obtain multi-scale optical flow information, including: Perform Gaussian blur and downsampling on the image of the video frame to generate multiple layers of images, and the multiple layers of images form the image pyramid. The image pyramid includes at least a first-layer image, a second-layer image, and a third-layer image. The resolution of the first-layer image is greater than that of the second-layer image, and the resolution of the second-layer image is greater than that of the third-layer image; Calculate the optical flow field corresponding to the first-layer image, the optical flow field corresponding to the second-layer image, and the optical flow field corresponding to the third-layer image respectively for the first-layer image, the second-layer image, and the third-layer image through the optical flow algorithm; Fuse the optical flow field corresponding to the first-layer image, the optical flow field corresponding to the second-layer image, and the optical flow field corresponding to the third-layer image through a multi-scale convolutional neural network to generate multi-scale optical flow information with consistent image features.

5. The three-dimensional video fusion method according to claim 4, wherein The performing Gaussian blur and downsampling on the image of the video frame to generate multiple layers of images, and the multiple layers of images form the image pyramid. The image pyramid includes at least a first-layer image, a second-layer image, and a third-layer image. The resolution of the first-layer image is greater than that of the second-layer image, and the resolution of the second-layer image is greater than that of the third-layer image, includes: Perform Gaussian blur and downsampling on the image of the video frame to obtain the first-layer image; Perform Gaussian blur and downsampling on the first-layer image to obtain the second-layer image, and the resolution of the second-layer image is lower than that of the first-layer image; Perform Gaussian blur and downsampling on the second-layer image to obtain the third-layer image, and the resolution of the third-layer image is lower than that of the second-layer image.

6. The three-dimensional video fusion method according to claim 4, characterized in that The fusing the optical flow field corresponding to the first-layer image, the optical flow field corresponding to the second-layer image, and the optical flow field corresponding to the third-layer image through a multi-scale convolutional neural network to generate multi-scale optical flow information with consistent image features, includes: Align the optical flow field corresponding to the first-layer image, the optical flow field corresponding to the second-layer image, and the optical flow field corresponding to the third-layer image to the same resolution through upsampling; Extract image features from the optical flow field corresponding to the aligned first-layer image through 3×3 convolution to obtain a first feature map, extract image features from the optical flow field corresponding to the aligned second-layer image through 5×5 convolution to obtain a second feature map, and extract image features from the optical flow field corresponding to the aligned third-layer image through 7×7 convolution to obtain a third feature map; Dynamically weight and fuse the first feature map, the second feature map, and the third feature map through the squeeze-and-excitation module in the multi-scale convolutional neural network, and then use 3×3 convolution to generate multi-scale optical flow information with consistent image features.

7. The three-dimensional video fusion method according to claim 1, wherein Extract corresponding image feature points from the video frame and the three-dimensional scene model and combine the multi-scale optical flow information. Matching the image feature points of the video frame with the image feature points of the three-dimensional scene model to generate a feature point matching result includes: Use the scale-invariant feature transform algorithm to extract feature points in the video frame to obtain a set of feature points of the video frame; Project the three-dimensional scene model onto the perspective of the video frame to generate a projection image, and use the scale-invariant feature transform algorithm to extract feature points in the projection image and generate a set of feature points of the three-dimensional scene model; Predict the feature point positions in combination with the multi-scale optical flow information, and use the K-nearest neighbor algorithm to match the set of feature points of the video frame with the set of feature points of the three-dimensional scene model to obtain a set of matching pairs; Verify the consistency of the set of matching pairs through the homography matrix, and the set of matching pairs that pass the verification is the feature point matching result.

8. The three-dimensional video fusion method according to claim 1, wherein Calculate the fusion weights according to the feature point matching result, and perform weighted fusion on the video frame and the three-dimensional scene model according to the fusion weights to obtain a fused three-dimensional video, including: Calculate the confidence scores of the matching pairs from the feature point matching result; Calculate the local weights and global weights from the confidence scores of the matching pairs; Obtain the fusion weights by assigning different coefficients to the local weights and the global weights; For the images of the video frame and the three-dimensional scene model, perform weighted fusion at each pixel position of each image according to the fusion weights to obtain the fused three-dimensional video.

9. The three-dimensional video fusion method according to claim 8, wherein Performing weighted fusion at each pixel position of each image according to the fusion weights to obtain the fused three-dimensional video includes: Calculate the fused pixel values based on the pixel values of the video frame and the pixel values of the image in the three-dimensional scene model; Organize the calculated fused pixel values into a sequence to form the three-dimensional video.

10. Application of a three-dimensional video fusion method in a video security system for implementing the three-dimensional video fusion method according to any one of claims 1-9, characterized in that, The video security system includes: An acquisition module, where the acquisition module is used to acquire video sequences of a target scene from multiple angles, and the video sequences include several video frames; A construction module, when the target scene is a static scene, analyzes the video sequence through a structured light three-dimensional reconstruction module to obtain a static three-dimensional mesh model, and when the target scene is a dynamic scene, analyzes the video sequence through structure from motion to obtain a dynamic three-dimensional mesh model. The construction module is used to construct a three-dimensional scene model based on the static three-dimensional mesh model and the dynamic three-dimensional mesh model; A multi-scale optical flow module, which is used to construct an image pyramid for each video frame, obtain the optical flow field corresponding to each layer of the image of the image pyramid through an optical flow algorithm, and fuse the optical flow fields through a multi-scale convolutional neural network to obtain multi-scale optical flow information; A feature point matching module, which is used to extract corresponding image feature points from the video frame and the three-dimensional scene model and combine the multi-scale optical flow information, and match the image feature points of the video frame with the image feature points of the three-dimensional scene model to generate a feature point matching result; A fusion module, which is used to calculate a fusion weight according to the feature point matching result, and perform weighted fusion on the video frame and the three-dimensional scene model according to the fusion weight to obtain a fused three-dimensional video.

Citation Information

Patent Citations

  • Time-space jointed multi-view video interpolation and three-dimensional modeling method

    CN102446366A

  • Dynamic scene HDR reconstruction method based on deep learning

    CN111242883A

  • Video super-resolution reconstruction method based on multi-frame fusion optical flow

    CN111311490A

  • Video fusion method

    CN112584120A

  • Scene video processing method and device, equipment and storage medium

    CN117274446A

Cited By

  • Three-dimensional model video fusion method and device based on machine learning and medium

    CN122156555A

  • A machine learning-based three-dimensional model video fusion method, device and medium

    CN122156555B