Real-time video stream distortion correction and splicing method for live-action fusion

By generating anchor point sets through multi-sensor unit self-organizing networks and a five-modal fusion algorithm, and constructing an adaptive element calibration matrix, the problems of high latency and poor seamlessness of dynamic scenes in existing video distortion correction and stitching methods are solved, and efficient and accurate mapping of multiple real-time video streams to virtual space is achieved.

CN121961865APending Publication Date: 2026-05-01广东际合科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
广东际合科技有限公司
Filing Date
2026-01-16
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, video distortion correction and stitching methods suffer from problems such as poor anchor point generation stability, high processing latency, error accumulation, and dynamic target misalignment, making it difficult to achieve dynamic and seamless mapping of multiple real-time videos to virtual space.

Method used

A multi-sensor unit self-organizing network is used to synchronously collect video streams, inertial measurement data, millimeter-wave radar distance data, and thermal imaging temperature data. A multi-modal spatial anchor point set is generated through a five-modal fusion algorithm, an anchor point association index is established, and an adaptive element calibration matrix is ​​constructed to realize the dynamic mapping of video streams to a unified virtual space.

Benefits of technology

It achieves dynamic, accurate, and seamless mapping of multiple real-time video streams to a unified virtual space, reduces splicing latency, improves the spatial uniqueness and stability of anchor points, and ensures the continuity of dynamic targets and splicing accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121961865A_ABST
    Figure CN121961865A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video processing and live-action fusion, and discloses a real-time video stream distortion correction and splicing method for live-action fusion. According to the method, multi-source data are synchronously collected through a camera array ad hoc network integrating multiple sensing units, effective spatial anchor points are generated and screened through five-mode fusion, video elements are decomposed, association with the anchor points is established, distortion correction and spatial calibration are synchronously completed based on an adaptive element calibration matrix, and the distortion correction and spatial calibration precision is improved. And finally, dynamic seamless fusion of multiple paths of videos is realized through virtual space adaptive grid splicing. According to the method, a traditional linear processing flow is thoroughly overturned, the core problems of high splicing delay, poor dynamic scene adaptability and insufficient multi-element cooperative calibration precision in the prior art are solved, the method is suitable for various live-action fusion scenes such as urban traffic monitoring and large-scale venue security and protection, and the real-time performance and the seamless performance of video splicing are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Real-time video stream distortion correction and stitching method for real-scene fusion Technical Field

[0001] This invention belongs to the field of video processing and real-scene fusion technology, specifically relating to a real-time video stream distortion correction and stitching method for real-scene fusion. Background Technology

[0002] With the rapid development of real-scene fusion technology, seamlessly stitching and mapping multiple real-time video streams onto a unified virtual space to achieve global real-time perception of the real world has become a core requirement in fields such as intelligent monitoring and smart transportation. In existing technologies, video distortion correction and stitching mostly follow a linear process of "distortion correction - geometric alignment - dynamic fusion," which has significant limitations: First, it relies on single image sensor data and lacks multi-dimensional sensor information collaboration, resulting in poor anchor point generation stability and susceptibility to environmental interference; second, distortion correction and spatial alignment are performed step-by-step, increasing processing latency and easily leading to error accumulation, making it difficult to meet real-time requirements; third, the fixed-density design of grid stitching cannot adapt to different scene complexities, and splicing breaks or texture distortion are prone to occur in dynamic target areas; fourth, dynamic target calibration only focuses on the spatial position of a single frame, ignoring the temporal continuity of motion, making target misalignment prone to occur during cross-frame stitching.

[0003] Furthermore, existing methods rely heavily on external synchronization devices for sensor data acquisition, resulting in high latency in data interaction between cameras and insufficient timestamp alignment accuracy. The lack of a clear correlation mechanism between element decomposition and calibration makes it difficult to guarantee the spatial consistency of the three types of elements (geometry, texture, and dynamic targets). These problems collectively make it difficult for existing technologies to achieve dynamic and seamless mapping of multiple real-time video streams to virtual space, limiting the application of reality fusion technology in complex scenarios. Summary of the Invention

[0004] To address the aforementioned shortcomings in existing technologies, this invention provides a real-time video stream distortion correction and stitching method for real-scene fusion, which solves the technical problems of high stitching delay, poor seamlessness of dynamic scenes, and insufficient accuracy of multi-element collaborative calibration caused by the linear process of "distortion correction-geometric alignment-dynamic fusion" in the background technology, thereby achieving dynamic, accurate, and seamless mapping of multiple real-time video streams to a unified virtual space.

[0005] To address the aforementioned technical problems, this invention employs the following technical solution: a real-time video stream distortion correction and stitching method for real-scene fusion, comprising the following steps: S1: Constructing a self-organizing network using a camera array integrating multiple sensor units, simultaneously acquiring multiple real-time video streams, inertial measurement data, spatial location data, millimeter-wave radar distance data, and thermal imaging temperature data, and achieving timestamp alignment of multi-source data based on millimeter-wave radar communication; S2: Generating a multimodal spatial anchor point set using a five-modal fusion algorithm of vision-inertial-lidar-millimeter-wave radar-thermal imaging, filtering effective anchor points through cross-validation of multi-sensor data, and establishing a cross-camera anchor point association index; S3: Decomposing each video stream into three categories of elements: geometry, texture, and dynamic targets, clarifying the corresponding association between each type of element and the anchor point; S4: Constructing an adaptive element calibration matrix based on the anchor point association index, and synchronously generating a standardized element set through matrix operations; S5: Mapping the standardized element set to a unified virtual space, dynamically adjusting the virtual space grid density based on scene complexity, and completing the stitching.

[0006] Furthermore, the camera integrates a six-axis IMU, a GPS module, a LiDAR sensor, a millimeter-wave radar sensor, and a thermal imaging sensor. The self-organizing network enables direct data interaction between cameras through the communication function of the millimeter-wave radar, and compensates for time delays in distance calculations between cameras to align timestamps of multi-source data. The multi-source data includes video streams, IMU attitude data, GPS position coordinates, LiDAR point cloud data, millimeter-wave radar distance data, thermal imaging temperature data, and camera intrinsic parameter data.

[0007] Furthermore, the cross-validation of multi-sensor data in S2 specifically includes: verifying the accuracy of anchor point spatial coordinates by using the spatial distance difference between lidar point cloud data and millimeter-wave radar distance data; filtering static anchor points by using thermal imaging temperature data and excluding anchor points corresponding to dynamic heat sources; the effective anchor points include virtual spatial coordinates, cross-camera association identifiers, and stability tags.

[0008] Furthermore, in S3, video element decomposition is performed using a lightweight Transformer network. Geometric elements include contour features, keypoint coordinates, and plane equation parameters; texture elements include color matrix, texture gradient, and color histogram; and dynamic target elements include motion region mask, target bounding box, and motion vector. Each type of element establishes a unique mapping relationship with its corresponding valid anchor point through anchor point association identifiers.

[0009] Furthermore, the adaptive element calibration matrix construction in S4 includes using anchor point association indexes as rows and element types as columns, incorporating five-modal sensing calibration parameters; synchronous calibration specifically includes: synchronously correcting radial distortion, tangential distortion deviation and spatial coordinate offset of geometric elements through matrix operations, and texture elements and dynamic target elements achieve pixel-level collaborative deformation based on the calibration results of geometric elements, ensuring the consistency of spatial position and shape of the three types of elements.

[0010] Furthermore, the virtual space adaptive mesh stitching in S5 specifically includes: pre-setting a basic virtual space mesh, dynamically adjusting the mesh density based on the quantified values ​​of scene complexity; maintaining the basic mesh density in static areas, and achieving seamless stitching through coordinate matching of mesh nodes and standardized elements.

[0011] Furthermore, the virtual space adaptive mesh stitching in S5 also includes an element-mesh bidirectional adaptation mechanism: for irregular geometric elements, a non-uniform rational B-spline mesh is used to replace the traditional rectangular mesh, and the mesh nodes automatically fit the outline curve of the element. For texture elements in standardized elements, texture block sampling is performed based on the spatial distribution of mesh nodes to avoid texture stretching or compression distortion.

[0012] Furthermore, the synchronous calibration in S4 also includes spatiotemporal joint calibration of dynamic targets: in the time dimension, based on the motion trajectory of the dynamic target in the previous 5 frames, the theoretical position and shape of the current frame are predicted by Kalman filtering as calibration constraints; in the spatial dimension, the accuracy of distortion correction parameters of geometric elements within the bounding box of the dynamic target is improved.

[0013] Furthermore, it also includes quantifying the dynamic target density by the area ratio of the dynamic target bounding box, and quantifying the texture change rate by the inter-frame difference of the image gradient histogram; based on the quantization results, a linear interpolation algorithm is used to calculate the target grid density, and the grid node spacing is negatively correlated with the scene complexity quantification value to ensure the stitching accuracy of dynamic complex regions.

[0014] Furthermore, the adaptive element calibration matrix described in S4 also includes a parameter dynamic iterative optimization mechanism, which uses the least squares method to update the sensor calibration parameters in the matrix in real time based on the splicing quality monitoring results in S5.

[0015] Compared with existing technologies, this invention has the following advantages: 1. It achieves simultaneous execution of distortion correction and spatial calibration through an adaptive element calibration matrix. Combined with multi-source data self-organizing network acquisition and adaptive mesh stitching, the overall processing flow is coordinated and linked, significantly reducing stitching latency and meeting the needs of real-time scene fusion; 2. It integrates five-modal sensor data and selects effective anchor points through cross-validation to ensure the spatial uniqueness and stability of anchor points. The calibration matrix incorporates five-modal parameters to achieve element-level precise calibration, significantly improving the spatial consistency of the three types of elements and resulting in better distortion correction; 3. It dynamically adjusts the mesh density based on scene complexity and combines an element-mesh bidirectional adaptation mechanism to ensure stitching accuracy in dynamic complex areas while avoiding computational redundancy in static areas. It also solves the problems of irregular geometric element stitching and texture distortion, greatly improving seamlessness; 4. Through a spatiotemporal joint calibration mechanism, it integrates trajectory prediction in the time dimension and accuracy improvement in the spatial dimension, effectively avoiding cross-frame stitching misalignment of dynamic targets and ensuring the continuity of target trajectories in dynamic scenes. Attached Figure Description

[0016] Figure 1 shows a schematic diagram of the real-time video stream distortion correction and stitching method for real-scene fusion proposed in an embodiment of this application; Figure 2 shows a schematic diagram of the anchor point set of the real-time video stream distortion correction and stitching method for real-scene fusion proposed in an embodiment of this application; Figure 3 shows a schematic diagram of the element set mapping of the real-time video stream distortion correction and stitching method for real-scene fusion proposed in an embodiment of this application. Detailed Implementation

[0017] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0018] Example 1: As shown in Figure 1, this invention uses real-scene fusion monitoring of a large convention center as an application scenario. This scenario requires seamlessly stitching and mapping real-time video streams collected by multiple cameras distributed in key areas such as exhibition halls, passages, and entrances / exits of the convention center onto a 3D virtual model of the convention center. This enables global real-time monitoring of exhibitors, exhibits, and their movement, supporting on-site scheduling and security management. Specifically, it involves a real-time video stream distortion correction and stitching method for real-scene fusion. This method includes: S1: Collecting multi-source data and constructing a multi-source data self-organizing network: This step is to construct a synchronous acquisition system to provide data support for subsequent processing. In specific implementation, cameras deployed in key areas of the convention center construct a self-organizing network through the communication function of millimeter-wave radar. Direct data interaction between cameras can be achieved without the need for additional external synchronization equipment. Each camera has a built-in six-axis IMU, GPS module, lidar sensor, millimeter-wave radar sensor, and thermal imaging sensor. The working principles and synergistic effects of each component are as follows: The six-axis IMU consists of a three-axis accelerometer and a three-axis gyroscope. By sensing the acceleration and angular velocity of objects, it outputs the camera's attitude angles in real time, such as pitch angle, roll angle, yaw angle, and motion state data; The GPS module receives positioning signals from multiple satellites and calculates the camera's precise geographic coordinates, providing a benchmark for multi-device spatial position calibration; The lidar sensor emits a laser beam and receives reflected signals, calculates the laser propagation time and distance, and generates high-density point cloud data to reconstruct the three-dimensional spatial structure of the exhibition scene; The millimeter-wave radar sensor has both communication and detection functions. On the one hand, it builds a self-organizing network through millimeter-wave signals to achieve data interaction between devices. On the other hand, it emits millimeter waves and receives target reflected waves to obtain the target's distance and speed information; The thermal imaging sensor detects the infrared radiation emitted by objects and converts it into a temperature distribution image to identify exhibitors and heat-generating equipment targets.

[0019] After each component is activated, it synchronously collects real-time video streams, IMU attitude data, GPS position coordinates, LiDAR point cloud data, millimeter-wave radar distance data, and thermal imaging temperature data for the corresponding area. Their synergistic effect is reflected in: using the millimeter-wave radar self-organizing network as the communication core to ensure the real-time transmission of data collected by each component; achieving spatial alignment between cameras through the fusion of GPS position data and IMU attitude data; LiDAR point cloud data providing three-dimensional spatial supplementation for the video stream; and thermal imaging data compensating for the recognition shortcomings of visible light video in low-light or occluded scenes. The multi-dimensional data corroborate and complement each other, while simultaneously collecting camera intrinsic parameter data to lay the foundation for subsequent data calibration and fusion.

[0020] Simultaneously, by calculating the physical distance between cameras to determine the data transmission time delay, time delay compensation is applied to all collected data to ensure that the timestamps of the exhibition center scene data collected by different cameras at the same moment are completely synchronized. This provides a reliable data foundation for subsequent multimodal fusion and anchor point generation. This step achieves direct synchronization between nodes through a millimeter-wave radar self-organizing network, eliminating the need for external clock sources such as GPS or dedicated synchronization servers. This fundamentally avoids the risk of signal loss or delay due to obstruction by the exhibition center's building structure, significantly improving synchronization stability. Furthermore, in terms of anti-interference and robustness, the self-organizing network supports multi-node mutual synchronization; the failure of a single node does not affect the synchronization of the entire network. Millimeter-wave communication also exhibits strong resistance to obstruction and electromagnetic interference, making it more reliable in the densely populated and complex equipment environment of an exhibition compared to traditional single-source synchronization protocols.

[0021] S2: Multimodal Spatial Anchor Point Generation and Screening: This step aims to establish a highly reliable spatial benchmark, providing accurate correlation basis for element calibration. Based on the multi-source data collected in the previous step, a five-modal fusion algorithm combining vision, inertial, lidar, millimeter-wave radar, and thermal imaging is run to fully integrate the advantages of various sensor data, generating a multimodal spatial anchor point set covering static key locations such as exhibition center pillars, exhibit booth corners, and entrance / exit signs. The fusion algorithm adopts a three-level architecture of "preprocessing - hierarchical fusion - output calibration" to achieve accurate fusion of multimodal data.

[0022] Data preprocessing stage: The acquired five-modal data are standardized and calibrated. Keyframes are extracted from the video stream using frame differencing, and the SIFT algorithm is used to detect feature points and remove redundant noise, providing visual recognition for anchor points. Inertial measurement unit (IMU) data undergoes Kalman filtering to eliminate zero-bias errors, outputting a smooth attitude angle sequence. LiDAR data is filtered through pass-through to remove ground interference points, offsetting the impact of installation jitter on anchor point positioning and generating a simplified point cloud. Millimeter-wave radar data is processed using the clutter suppression algorithm to remove clutter while retaining effective target distance information. Thermal imaging data is segmented using adaptive thresholding to extract temperature regions, providing a direct basis for subsequent selection of static anchor points.

[0023] Layered fusion stage: A progressive fusion strategy of "data layer - feature layer - decision layer" is adopted. Data layer fusion achieves spatial coordinate unification, mapping visual, laser, and millimeter-wave data to the same geographic coordinate system; feature layer fusion allocates modal weights through an attention mechanism to match and associate feature points; decision layer fusion is based on Bayesian inference, integrates the judgment results of each modality, and outputs accurate spatial target information.

[0024] The coordinate system calculation method is as follows: Using the GPS geographic coordinate system as a reference, coordinate transformation is performed on each modal data. The formula for transforming the lidar point cloud coordinates (X_L, Y_L, Z_L) to the geographic coordinate system (X_G, Y_G, Z_G) is: X_G = X_L × × -Y_L× +Z_L× +X0;Y_G=X_L× -Y_L× × +Z_L× +Y0;Z_G=-X_L× +Z_L× +Z0; where θ is the yaw angle output by the IMU, φ is the pitch angle, and (X0, Y0, Z0) is the installation offset of the LiDAR relative to the GPS module. The visual feature point coordinates are transformed from pixel coordinates (u, v) to world coordinates through the camera intrinsic parameter matrix K: [X_W, Y_W, Z_W]^T = Z_C·K⁻¹·[u, v, 1]^T (Z_C is the feature point depth, supplemented by LiDAR data).

[0025] Feature layer weight allocation calculation: An adaptive weighted fusion algorithm is adopted, dynamically allocating weights based on the confidence scores of each modality. The confidence score of the i-th modality is defined as C_i = 1 - σ_i / σ_max (σ_i is the variance of the data for that modality, and σ_max is the maximum variance across all modalities). Then, the fusion weight W_i = C_i / ΣC_i (i = 1 to 5, corresponding to visual, inertial, laser, millimeter-wave, and thermal imaging, respectively). The fusion feature vector F = ΣW_i·F_i (F_i is the feature vector for each modality). Then, based on the judgment results of each modality for static targets, the post-verification probability is calculated, and finally, the coordinates are corrected using the difference between the distance data from the laser radar and the millimeter-wave radar.

[0026] Secondly, anchor points are screened using a rigorous dual verification mechanism to ensure both static attributes and spatial accuracy. The verification process involves two parts: 1) Spatial coordinate accuracy verification: Utilizing the 3D spatial positioning advantage of LiDAR point cloud data and the distance measurement advantage of millimeter-wave radar, the spatial distance difference for the same physical location in both types of data is calculated. A preset distance difference threshold is established; if the spatial distance difference of an anchor point exceeds this threshold, it is deemed to have unacceptable spatial coordinate deviation and is directly removed, preventing inaccurate anchor point positioning due to errors in a single sensor. 2) Static attribute verification: Through temporal analysis of multiple consecutive frames of thermal imaging temperature data, the temperature fluctuation range of the area corresponding to the anchor point is calculated. If the fluctuation range is below a preset stability threshold, it is determined to be a static anchor point with stable temperature; otherwise, it is determined to be a dynamic anchor point corresponding to dynamic heat sources such as exhibitors or mobile devices and is excluded.

[0027] Output calibration stage: Deviation correction is applied to the fused spatial data. Pixel coordinates are converted to world coordinates using camera intrinsic parameters, generating an anchor point set. The final effective anchor points include virtual space coordinates, cross-camera association identifiers, and stability labels. The virtual space coordinates are calibrated based on a unified virtual space coordinate system, ensuring all anchor points are on the same spatial reference, providing a unified reference for subsequent element mapping. The cross-camera association identifier uses a composite coding rule of "camera ID - physical location code - anchor point sequence number," clearly identifying the original camera to which the anchor point belongs and directly associating it with its corresponding physical scene location, providing a unique identifier for cross-camera anchor point matching. Stability... Based on indicators such as the consistency of five-modal data fusion, the magnitude of spatial distance difference, and the range of temperature fluctuation, the stability labels divide anchor points into three stability levels: high, medium, and low. High-stability anchor points are prioritized for core area element calibration, while medium- and low-stability anchor points are used as supplements or for non-critical areas, achieving differentiated utilization of anchor point resources. A full-camera anchor point association index can be established through association identifiers to clarify the cross-device anchor point correspondence at the same physical location. For example, if anchor point "1-005" (the 5th anchor point of node 1) corresponds to anchor point "2-012" in the data of node 2, then the association relationship between the two is recorded in the index table to facilitate subsequent cross-camera data calibration and stitching.

[0028] As shown in the table below: S3: Decompose video stream elements and clarify their corresponding relationships with anchor points: This step focuses on "precise decomposition and unique mapping," employing a five-stage closed-loop implementation process: "video preprocessing → lightweight Transformer network inference → three types of element extraction → element-anchor point mapping verification → structured output." Specifically, video preprocessing involves standardizing the real-time video stream to adapt to edge camera computing power and improve decomposition accuracy. First, frame normalization: all video frames are uniformly scaled to 640×480 pixels, with pixel values ​​normalized to the [0,1] range to eliminate the impact of size differences on network inference. Second, noise suppression: a 3×3 kernel Gaussian filter is used to smooth the image, filtering interference from ambient light changes and sensor noise. Third, keyframe sampling: combining the timestamp synchronization information output from S1, one keyframe is extracted every three frames, reducing redundant computation while ensuring real-time performance. The final output is a standardized keyframe sequence, serving as the network input data source.

[0029] Lightweight Transformer Network Deployment and Inference: MobileViT was selected as the core decomposition network. This network, through a lightweight design of "local image patch embedding + global attention mechanism," reduces the number of parameters by 75% compared to the standard ViT, and can be directly deployed on cameras. First, based on a dataset specifically for exhibition scenarios, including samples of exhibition hall walls, pillars, exhibits, and attendees, the network was pre-trained and fine-tuned to enable it to accurately recognize elements in exhibition scenarios. Then, the pre-processed keyframes were input into the network. Through local feature extraction, global attention modeling, and multi-scale feature fusion, the network outputs a feature map containing "element category, pixel position, and preliminary attributes," providing a foundation for subsequent element extraction.

[0030] Secondly, three types of core elements are accurately extracted. Based on the feature map output by the network, the detailed extraction of the three types of elements is carried out in a targeted manner to ensure that the element information is complete and can be used for subsequent calibration: Geometric element extraction: Focusing on the static structural region in the feature map, the Canny edge detection algorithm is used to extract continuous contour features and output contour coordinate sequence; Harris corner detection is used to locate key points such as column vertices and wall corners, and combined with S2 coordinate transformation, the pixel coordinates are converted into three-dimensional world coordinates; the least squares method is used to fit the plane equation ax+by+cz+d=0 to quantify the spatial posture of the planar structure such as the exhibition hall walls and exhibition stands.

[0031] Texture element extraction: Within the image region corresponding to the geometric element, the RGB color matrix is ​​extracted in units of 3×3 neighborhood windows to represent the local color distribution; the Sobel operator is used to calculate the horizontal and vertical texture gradients to capture the details of wall texture and exhibit surface texture; the RGB three-channel 256 bins color histogram is statistically analyzed to form global texture features, providing a basis for the uniqueness of element differentiation.

[0032] Dynamic target element extraction: The motion region is accurately segmented by combining the motion region mask output by the network with the background subtraction method; the lightweight YOLOv8 model is called to generate target bounding boxes (format [x1,y1,x2,y2]) to locate dynamic targets such as exhibitors and moving exhibits; the Lucas-Kanade optical flow method is used to calculate motion vectors to represent the direction and speed of target motion, providing support for subsequent dynamic / static element differentiation.

[0033] Element-Anchor Point Unique Mapping: Relying on the effective anchor point set output by S2 and the cross-camera association index, precise binding of "element-anchor point" is achieved. Specifically, this involves first constructing a mapping index library, importing effective anchor point information from S2, including anchor point association identifiers, 3D spatial coordinates, and stability labels, and establishing an index table with "anchor point association identifiers" as the core; secondly, performing spatial coordinate matching, converting the extracted pixel positions of various elements into 3D spatial coordinates through S2's pixel-to-world coordinate transformation, and matching the nearest effective anchor point; finally, performing uniqueness verification, setting an IOU threshold, if a single element matches multiple anchor points, prioritizing the anchor point with the highest stability label and the closest spatial distance to complete the binding, ensuring mapping uniqueness; secondly, removing elements that cannot be matched with effective anchor points, such as noise or temporary interference targets; and finally, performing structured output, integrating the extracted element information with the bound anchor point information to generate a structured data list of "element ID - element type - core attribute - anchor point association identifier - 3D spatial coordinates", which is output to the subsequent calibration module, providing a clear "element-space" association object for targeted calibration. This step, together with S2, forms a progressive "support-connection" relationship. S2, through five-modal fusion and dual filtering, generates a highly accurate and stable set of effective anchor points and cross-camera association identifiers, thereby providing a precise timestamp benchmark for synchronization and a spatial basis for achieving "unique mapping between elements and anchor points." For example, if the region containing the contour features of a geometric element contains the anchor point "1-005," then the association identifier of that geometric element is set to "1-005," thus achieving precise binding between elements and anchor points. Together, this constitutes a complete and feasible technical solution chain from "spatial benchmark construction" to "video element decomposition and spatial association."

[0034] S4: Construct an adaptive element calibration matrix based on the anchor point association index. Simultaneously complete element-level distortion correction and spatial coordinate calibration through matrix operations to generate a standardized element set. The overall workflow of this step is based on the core link of "benchmark import - matrix construction - collaborative calibration - dynamic optimization". Specifically, firstly, core benchmark data is imported based on the anchor point association index established in S3. At the same time, the original parameters and calibration experience values ​​of the five-modal sensors from the previous multi-source data acquisition stage are retrieved. Based on this, an adaptive element calibration matrix is ​​constructed. The working principle of this matrix is ​​to use the anchor point association index as the row vector and the three types of elements, namely geometric elements, texture elements, and dynamic target elements, as the column vectors. The calibration parameters of the five-modal sensors (including IMU attitude calibration coefficients, LiDAR point cloud accuracy correction parameters, millimeter-wave radar range compensation parameters, camera intrinsic parameter distortion coefficients, and thermal imaging temperature association calibration parameters) are normalized and quantized and integrated into the corresponding positions of the matrix, forming a three-in-one structured calibration framework of "anchor point - element - sensor parameter", which provides a unified computing carrier for subsequent collaborative calibration. Once the matrix is ​​constructed, the core synchronous distortion correction and spatial calibration process is immediately initiated. This step works by simultaneously completing multi-dimensional calibration requirements through a single matrix operation of the calibration matrix. For geometric elements, the Brown-Conrady distortion model is used to calculate the correction amounts for radial distortion (k1, k2, k3) and tangential distortion (p1, p2) based on the pre-stored camera intrinsic distortion coefficients in the matrix. This simultaneously corrects distortion deviations caused by lens optical characteristics, such as bending of exhibition hall walls and deformation of columns. At the same time, combined with the IMU attitude parameters and GPS position references integrated into the matrix, the global coordinate offset caused by differences in camera installation positions is corrected through spatial coordinate transformation formulas, ensuring that the virtual spatial coordinates of geometric elements are completely aligned with the actual exhibition center scene. Texture elements and dynamic target elements are mapped in real time to the calibration results of geometric elements through preset element association weights in the matrix. Pixel-level collaborative deformation is achieved based on the bilinear interpolation principle in image registration, ensuring consistency in spatial position and shape among the three types of elements.

[0035] Simultaneously with the implementation of collaborative calibration, a dynamic target spatiotemporal joint calibration mechanism is launched. Its working principle is as follows: for dynamic targets such as exhibitors and moving exhibits, the motion trajectory data of the previous 5 frames is extracted from historical data. The state equation and observation equation are constructed through the Kalman filter algorithm to predict the theoretical position and morphological parameters of the dynamic target in the current frame. These are then incorporated into the calibration matrix operation as calibration constraints to avoid calibration deviations caused by motion blur of the dynamic target. In the spatial dimension, the area covered by the bounding box of the dynamic target is identified in real time through the calibration matrix. The accuracy of the distortion correction parameters of the geometric elements in this area is automatically improved. Higher resolution grid sampling is used for local optimization to ensure smooth connection between the edge of the dynamic target and the surrounding static elements. This deep synergy between the predictive constraints in the temporal dimension and the local optimization in the spatial dimension significantly improves the calibration accuracy of the dynamic target.

[0036] Furthermore, the adaptive element calibration matrix also possesses a dynamic iterative optimization mechanism for parameters. Its working principle is to use the quality monitoring results of subsequent virtual space splicing (including indicators such as anchor point matching rate, element overlap, and pixel error) as feedback signals. When the monitored indicators are lower than the preset threshold, the least squares optimization process is automatically triggered. With the goal of minimizing calibration error, the five-modal sensor calibration parameters in the matrix are updated in real time. At the same time, the element association weights are adjusted through the gradient descent algorithm to achieve continuous optimization of calibration accuracy. This iterative mechanism forms a closed loop with the collaborative calibration process mentioned above, enabling the calibration system to adapt to dynamic environmental changes such as personnel flow and lighting changes within the convention center, and always maintain a high-precision calibration state. The entire process, through the core carrier of the calibration matrix, deeply integrates multi-element calibration requirements, multi-modal sensing parameters, dynamic constraints, and iterative optimization logic. Their synergistic effect is reflected in the matrix construction providing a unified framework for calibration, synchronous calibration achieving efficient alignment of multiple elements, spatiotemporal joint calibration specifically addressing dynamic target deviation issues, and iterative optimization ensuring long-term calibration stability. Ultimately, it abandons the traditional step-by-step logic of "correcting first and then aligning," achieving precise multi-element collaborative calibration, significantly reducing processing latency, and providing high-quality calibration data support for subsequent virtual space stitching.

[0037] S5: Virtual Space Stitching: First, the calibrated and standardized element set output from S4 is mapped to the unified virtual space corresponding to the 3D virtual model of the exhibition center. This mapping process uses the anchor point spatial coordinates generated in step S2 as an intermediary to ensure the accuracy of the initial mapping position of the elements. Then, a basic mesh is preset, using a rectangular mesh aligned with the virtual space coordinate system as the basic mesh. The mesh node spacing is set to 10cm. Its working principle is to uniformly divide the virtual space to construct an initial spatial framework. Each mesh node is bound to a unique virtual space coordinate, providing a basic positioning reference for subsequent element adaptation. Based on this, the mesh density is dynamically adjusted according to the scene complexity quantization value. The quantization rules are as follows: Dynamic target density quantization: The real-time dynamic degree of the scene is reflected by statistically analyzing the proportion of dynamic targets. Specifically, the dynamic target bounding box data decomposed in step S3 is used as input, and the ratio of the total area of ​​all dynamic target bounding boxes in each frame to the total area of ​​the image is calculated (quantization formula: dynamic target density = Σ dynamic target bounding box area / total image area × 100%); Texture change rate quantization: The scene detail richness is reflected by characterizing the inter-frame differences in texture. First, the texture elements decomposed in step S3 are extracted, and an image gradient histogram is generated. Next, the difference between histograms of adjacent frames is calculated using Bach distance. The core logic of Bach distance is to quantify the change in texture features by calculating the similarity between two probability distributions. Target mesh density calculation: Based on the comprehensive weight allocation of dynamic and static features, a linear interpolation algorithm is used to fuse the results of the first two quantizations to calculate the target mesh density (calculation formula: mesh density = base density × (1 + 0.6 × dynamic target density + 0.4 × texture change rate)). In static areas such as open passages and walls (dynamic target density < 5%, difference < 0.2), the base mesh density is maintained to balance stitching accuracy and computational efficiency.

[0038] For irregular geometric elements such as curved passages and irregularly shaped exhibit stands in the exhibition center, the core technology adopts non-uniform rational B-spline (NURBS) mesh for adaptation. A smooth and continuous curved mesh is generated by interpolation of controllable control points. The specific process is as follows: First, extract the key points of the contour curve of the irregular geometric element after calibration in step S4. Use these key points as control points of the NURBS mesh. By adjusting the weight coefficients of the control points, the mesh curve can accurately fit the contour edge of the geometric element. Finally, a NURBS mesh that perfectly fits the irregular contour is generated. Compared with traditional rectangular mesh, it can effectively avoid splicing gaps caused by contour fitting deviation.

[0039] For texture elements in the standardized elements, a two-way adaptation mechanism of texture block sampling + bilinear interpolation is adopted. The working principle is to divide the texture sampling area based on the spatial distribution of grid nodes, so that the texture sampling range is accurately matched with the grid carrying range: First, the sampling accuracy is determined according to the grid node spacing (5cm spacing corresponds to 5×5 pixel sampling, 10cm spacing corresponds to 10×10 pixel sampling). The corresponding texture pixel block is extracted in each grid area. Then, the sampled texture pixels are processed by the bilinear interpolation algorithm to smooth the transition. The core logic of this algorithm is to generate the gray value of the middle pixel by weighted averaging of the gray values ​​of four adjacent pixels, thereby avoiding distortion caused by texture stretching or compression.

[0040] Next, seamless stitching and quality verification are performed: Based on the precise coordinate mapping of anchor point association, using the anchor point association index in step S2 as a bridge, a one-to-one correspondence between the coordinates of the calibrated elements and the coordinates of the grid nodes is established. First, the anchor point association identifier of each standardized element is extracted, and the virtual space coordinates corresponding to the identifier are queried. Then, the key feature points of the elements (such as the vertices of geometric elements and the sampling center points of texture elements) are bound to the grid nodes of the corresponding coordinates. Through the global association of grid nodes, the elements collected by multiple cameras are integrated into a unified virtual space to complete the global stitching. Quality verification is achieved by quantifying the stitching effect through dual indicators and constructing a closed loop of "stitching-verification-adjustment".

[0041] After stitching, the process also includes calculating element overlap: using the intersection-union ratio (IoU) algorithm, the ratio of the area of ​​the overlapping region of adjacent elements to the area of ​​the union of the two elements is calculated. If the overlap is ≤2%, it is determined that there is no excessive overlap. Then, the gap width is detected: by calculating the shortest distance between the edge pixels of adjacent elements in the virtual space, if the distance is ≤1 pixel, it is determined that there is no obvious gap. If both indicators meet the standards, the stitching is qualified. If there is overlap or gap, the grid density (such as densifying the grid in the overlapping area) and the element mapping coordinates (fine-tuning based on the anchor point coordinates) are readjusted until the qualified standard is met.

[0042] The advantages of this invention are: it generates high-precision and high-stability spatial anchor points through five-modal fusion and cross-validation, and on this basis, it achieves accurate decomposition and unique mapping of three types of elements in the video stream: geometry, texture, and dynamic targets.

[0043] By using an adaptive calibration matrix to achieve element-level, millisecond-level spatiotemporal alignment and distortion correction, the problem of "data silos" is effectively solved, laying a solid foundation for high-quality data fusion and analysis.

[0044] The above are merely embodiments of the present invention. The circuits, electronic components, and modules involved are all prior art, fully achievable by those skilled in the art, and require no further explanation. The scope of protection in this application does not involve improvements to the software and methods. Commonly known structures and characteristics in the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all prior art in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the application.

Claims

1. A real-time video stream distortion correction and stitching method for real-scene fusion, characterized in that, Includes the following steps: S1: A self-organizing network is built by integrating a camera array with multiple sensing units to simultaneously collect multiple real-time video streams, inertial measurement data, spatial position data, millimeter-wave radar distance data and thermal imaging temperature data, and to achieve multi-source data timestamp alignment based on millimeter-wave radar communication; S2: A five-modal fusion algorithm combining vision, inertial, lidar, millimeter-wave radar, and thermal imaging is used to generate a multimodal spatial anchor point set. Valid anchor points are selected through cross-validation of multi-sensor data, and an anchor point association index is established. S3: Each video stream is decomposed into three types of elements: geometry, texture, and dynamic targets, clarifying the corresponding association between each type of element and the anchor point. S4: An adaptive element calibration matrix is ​​constructed based on the anchor point association index, and a standardized element set is generated through matrix operations. S5: The standardized element set is mapped to a unified virtual space, and the virtual space grid density is dynamically adjusted based on scene complexity and then stitched together.

2. The real-time video stream distortion correction and stitching method for real-scene fusion as described in claim 1, characterized in that: The camera integrates a six-axis IMU, GPS module, LiDAR sensor, millimeter-wave radar sensor, and thermal imaging sensor. The self-organizing network enables direct data interaction between cameras through the communication function of the millimeter-wave radar, and compensates for time delays in distance calculation between cameras to align the timestamps of multi-source data. The multi-source data includes video stream, IMU attitude data, GPS position coordinates, LiDAR point cloud data, millimeter-wave radar distance data, thermal imaging temperature data, and camera intrinsic parameter data.

3. The real-time video stream distortion correction and stitching method for real-scene fusion as described in claim 1, characterized in that: The cross-validation of multi-sensor data in S2 specifically includes: verifying the accuracy of anchor point spatial coordinates by using the spatial distance difference between lidar point cloud data and millimeter-wave radar distance data; filtering static anchor points by using thermal imaging temperature data and excluding anchor points corresponding to dynamic heat sources; the effective anchor points include virtual spatial coordinates, cross-camera association identifiers, and stability tags.

4. The real-time video stream distortion correction and stitching method for real-scene fusion as described in claim 1, characterized in that: In S3, video element decomposition is performed using a lightweight Transformer network. Geometric elements include contour features, keypoint coordinates, and plane equation parameters; texture elements include color matrix, texture gradient, and color histogram; and dynamic target elements include motion region mask, target bounding box, and motion vector. Each type of element establishes a unique mapping relationship with its corresponding valid anchor point through anchor point association identifiers.

5. The real-time video stream distortion correction and stitching method for real-scene fusion as described in claim 1, characterized in that: The adaptive element calibration matrix construction in S4 includes using anchor point association indexes as rows and element types as columns, incorporating five-modal sensing calibration parameters; synchronous calibration specifically includes: synchronously correcting radial distortion, tangential distortion deviation and spatial coordinate offset of geometric elements through matrix operations; texture elements and dynamic target elements achieve pixel-level collaborative deformation based on the calibration results of geometric elements, ensuring the consistency of spatial position and shape of the three types of elements.

6. The real-time video stream distortion correction and stitching method for real-scene fusion as described in claim 1, characterized in that: The virtual space adaptive mesh stitching in S5 specifically includes: a preset virtual space base mesh, dynamically adjusting the mesh density based on the quantified values ​​of scene complexity; maintaining the base mesh density in static areas, and achieving seamless stitching through coordinate matching of mesh nodes and standardized elements.

7. The real-time video stream distortion correction and stitching method for real-scene fusion as described in claim 6, characterized in that: The virtual space adaptive mesh stitching in S5 also includes an element-mesh bidirectional adaptation mechanism: for irregular geometric elements, a non-uniform rational B-spline mesh is used to replace the traditional rectangular mesh, and the mesh nodes fit the outline curve of the element. For texture elements in the standardized elements, texture block sampling is performed based on the spatial distribution of the mesh nodes to avoid texture stretching or compression distortion.

8. The real-time video stream distortion correction and stitching method for real-scene fusion as described in claim 1, characterized in that: Synchronous calibration in S4 also includes spatiotemporal joint calibration of dynamic targets: in the time dimension, based on the motion trajectory of the dynamic target in the previous 5 frames, the theoretical position and shape of the current frame are predicted by Kalman filtering as calibration constraints; in the spatial dimension, the accuracy of distortion correction parameters of geometric elements within the bounding box of the dynamic target is improved.

9. The real-time video stream distortion correction and stitching method for real-scene fusion as described in claim 1, characterized in that: It also includes quantifying dynamic target density by the area ratio of dynamic target bounding boxes, quantifying texture change rate by the inter-frame difference of image gradient histograms, and calculating target grid density using a linear interpolation algorithm based on the quantization results. The grid node spacing is negatively correlated with the scene complexity quantization value to ensure the stitching accuracy of dynamic complex regions.

10. The real-time video stream distortion correction and stitching method for real-scene fusion as described in claim 1, characterized in that: The adaptive element calibration matrix described in S4 also includes a parameter dynamic iterative optimization mechanism. Based on the splicing quality monitoring results in S5, the least squares method is used to update the sensor calibration parameters in the matrix in real time.