Dynamic splicing method and system for different-source security and protection video streams
By deploying PTP across the entire link and mapping the physical location of cameras, combined with lens distortion correction and photometric correction, and dynamically switching splicing modes, the problem of time synchronization and spatial mapping of heterogeneous video streams in smart cities and large park security is solved, achieving high-precision video stream splicing and intelligent interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIV OF TECH
- Filing Date
- 2026-01-19
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies are difficult to be compatible with multiple heterogeneous video inputs, and cannot achieve sub-microsecond time synchronization, accurate spatial mapping, and dynamic adaptive stitching. This results in problems such as frame errors, tearing, color difference, sudden changes in brightness and darkness, and geometric distortion in video streams during cross-camera target tracking and panoramic situational awareness, which cannot meet the advanced application requirements of smart cities and large park security.
By deploying PTP across the entire chain, sub-microsecond time synchronization is achieved. By combining the physical location of the camera with the GIS platform, a two-way mapping between image pixels and geographic coordinates is established. Lens distortion correction, angle correction, and light correction are performed. The stitching mode is dynamically switched, multiple protocol pushes are supported, and AI analysis results are superimposed on the stitched image.
It achieves strict alignment of multiple video frames in the time dimension and precise matching in the spatial dimension, eliminating splicing problems caused by device differences and uneven lighting, enhancing the visualization and traceability of the security system, and supporting various terminal requirements and intelligent interaction.
Smart Images

Figure CN121907973A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of dynamic video stream splicing technology, specifically a method and system for dynamic splicing of heterogeneous security video streams. Background Technology
[0002] With the rapid growth in demand for security in smart cities, intelligent transportation, and large parks, the deployment of multi-camera collaborative monitoring has become a mainstream trend. However, in practical applications, heterogeneous video devices (such as network cameras, analog cameras, HD PTZ cameras, and mobile terminals) from different manufacturers, using different protocols, and with different imaging characteristics are common, resulting in significant differences in video streams in terms of time synchronization, spatial alignment, color consistency, and geometric structure.
[0003] Traditional video stitching methods typically assume that all video sources have the same time reference, consistent imaging parameters, and fixed relative positions. This makes it difficult to handle complex conditions in real-world security scenarios, such as heterogeneous equipment, drastic changes in lighting, and dynamic shifts in overlapping viewpoints. Especially in advanced applications such as cross-camera target tracking and panoramic situational awareness, the lack of high-precision time synchronization (usually relying solely on NTP with millisecond-level errors) and geographic information-based spatial mapping mechanisms will lead to problems such as frame errors, tearing, color differences, abrupt changes in brightness, or geometric distortion in the stitched images. This severely impacts the efficiency of event analysis and backtracking. Furthermore, existing stitching systems mostly employ static fusion strategies, which cannot adaptively adjust stitching weights based on scene content (such as occlusion, moving objects, and abrupt changes in lighting) and lack effective correction methods for non-linear imaging modes such as fisheye and wide-angle lenses.
[0004] Therefore, there is an urgent need for a video stream splicing method and system that can be compatible with multiple heterogeneous video inputs, achieve sub-microsecond time synchronization, accurate spatial mapping, dynamic adaptive splicing, and support intelligent interaction and multi-protocol distribution, so as to improve the overall visualization capability and intelligence level of large-scale security monitoring systems. Summary of the Invention
[0005] The purpose of this invention is to provide a dynamic splicing method and system for heterogeneous security video streams to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method and system for dynamic splicing of heterogeneous security video streams, comprising the following steps:
[0007] S1: Users add camera information in the management interface. The system calls the corresponding adapter according to the protocol type, automatically scans video sources in the network, and establishes a mapping between physical location and logical coordinates.
[0008] S2: Video stream frames enter the time synchronization buffer. The system uses the master clock source as a reference to correct the timestamps of each stream frame.
[0009] S3: Perform geometric correction, photometric correction, and de-editing on each video stream;
[0010] S4: Extract feature points of overlapping areas of adjacent video streams, dynamically switch splicing modes according to scene changes, and simultaneously fuse and render the video streams;
[0011] S5: Push the stitched video stream to the client or platform via RTMP, WebRTC, or HLS, and support on-demand cropping, scaling, and view switching. At the same time, AI analysis results are overlaid on the stitched screen, and the original camera can be viewed by clicking on the stitched screen.
[0012] As a further preferred embodiment of this technical solution: the video source in S1 refers to a network camera, analog camera, high-definition camera, video recording device, streaming media server, or mobile device;
[0013] As a further preferred embodiment of this technical solution, the specific method for establishing the mapping between physical location and logical coordinates in S1 is as follows:
[0014] Given the camera's location (GPS) and orientation, and combining the focal length and sensor size, calculate the projected polygon of the view frustum on the ground, mark the camera's field of view in the GIS platform, establish the correspondence between the four corner points of the image and their geographic coordinates, and use affine transformation, polynomial correction, or grid interpolation to establish a nonlinear mapping, supporting the mapping of any latitude and longitude point to the pixel position in the video frame.
[0015] As a further preferred embodiment of this technical solution: the specific method for correcting the video stream using timestamps in S2 is as follows:
[0016] The network deploys a PTP master clock, the switch supports transparent clocks, and the camera acts as a PTP slave clock. After synchronization, the internal clock error is <1μs. Each frame is accompanied by a PTP timestamp, and all frames directly use PTP time as the absolute time without conversion.
[0017] As a further preferred embodiment of this technical solution: the geometric correction and photometric correction methods in S3 are further divided into:
[0018] Geometric correction: lens distortion correction, angle correction, fisheye / panoramic image unfolding;
[0019] Photometric correction: white balance adjustment, brightness and contrast equalization adjustment, histogram matching;
[0020] As a further preferred embodiment of this technical solution, the specific steps for lens distortion correction, angle correction, and fisheye / panoramic image unfolding are as follows:
[0021] Lens distortion correction: The camera is calibrated in advance using a calibration board to record its "deformation rules". During real-time processing, the curved image is "straightened" according to these rules to restore the true geometric structure, making the originally curved corridor edges, door frames, and driveway lines straight.
[0022] Viewpoint correction: Manually or automatically select a ground area in the image, and the system will "flatten" this slanted area to make it look like it is viewed from directly above, so that the positions of vehicles and pedestrians in the image are more consistent with the real map relationship, which is convenient for cross-camera tracking;
[0023] Fisheye / panoramic image unfolding: The entire spherical image is "unfolded" into a wide planar image, and the objects that were originally squeezed together are reasonably separated, making them easier to identify and stitch together;
[0024] As a further preferred embodiment of this technical solution, the specific steps for white balance adjustment, brightness and contrast equalization adjustment, and histogram matching are as follows:
[0025] White balance adjustment: Automatically analyzes the neutral colors in the image, adjusts the intensity of the red, green and blue channels, and makes the white adjustment whiter, so that the color style of multiple images is consistent and no obvious color blocks appear when stitching.
[0026] Brightness and contrast balance adjustment: enhances details in dark areas, suppresses overexposed areas, makes the overall brightness and darkness more balanced, and eliminates the abrupt feeling of "one side bright and one side dark" at the seams.
[0027] Histogram matching: Select one path as the "reference image", and the other paths will automatically adjust their color distribution to be as close as possible to the reference, so that the whole stitched image looks like it was taken by the same device;
[0028] As a further preferred embodiment of this technical solution: the specific method for extracting feature points of overlapping regions between adjacent video streams in S4 is as follows:
[0029] For two adjacent video frames from different cameras, their potential overlapping areas are first delineated by camera orientation and field of view. Feature points are detected in these two areas and their descriptors are calculated. By matching the descriptors, pairs of points with the same name are found, that is, pairs of pixel coordinates representing the same physical point.
[0030] Compared with the prior art, the beneficial effects of the present invention are:
[0031] 1. In this invention, by adopting PTP full-link deployment, sub-microsecond (<1μs) time synchronization is achieved, ensuring that multiple video frames are strictly aligned in the time dimension; at the same time, by combining parameters such as camera physical location (GPS), orientation, and focal length, a two-way mapping relationship between image pixels and geographic coordinates is established on the GIS platform to achieve accurate matching in the spatial dimension, providing a reliable foundation for cross-camera target tracking and event correlation.
[0032] 2. In this invention, the geometric structure of the real scene is restored by means of geometric correction such as lens distortion correction, viewing angle correction, and fisheye expansion; combined with photometric correction strategies such as white balance adjustment, brightness / contrast equalization, and histogram matching, the splicing color difference and brightness abrupt change caused by equipment differences, uneven lighting, or color temperature deviation are effectively eliminated, so that the spliced picture is visually coherent and has a unified style.
[0033] 3. In this invention, the stitched panoramic video supports multiple protocols such as RTMP, WebRTC, and HLS for push, meeting the needs of different terminals. It also provides on-demand cropping, scaling, and perspective switching functions, and can overlay AI analysis results on the screen. Users can click on the stitched screen to check the original camera source, greatly enhancing the visualization and traceability of the security system. Attached Figure Description
[0034] Figure 1 This is a flowchart of a dynamic splicing method and system for heterogeneous security video streams according to the present invention. Detailed Implementation
[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0036] Example:
[0037] Please see Figure 1 As shown, this invention provides a technical solution: a method and system for dynamically splicing heterogeneous security video streams, comprising the following steps:
[0038] S1: Users add camera information in the management interface. The system calls the corresponding adapter according to the protocol type, automatically scans video sources in the network, and establishes a mapping between physical location and logical coordinates.
[0039] S2: Video stream frames enter the time synchronization buffer. The system uses the master clock source as a reference to correct the timestamps of each stream frame.
[0040] S3: Perform geometric correction, photometric correction, and de-editing on each video stream;
[0041] S4: Extract feature points of overlapping areas of adjacent video streams, dynamically switch splicing modes according to scene changes, and simultaneously fuse and render the video streams;
[0042] S5: Pushes the stitched video stream to the client or platform via RTMP, WebRTC, or HLS, and supports on-demand cropping, scaling, and view switching. It also overlays AI analysis results on the stitched image and supports clicking on the stitched image to check the original camera footage.
[0043] Further: Dynamically switch splicing modes according to changes in the scenario, including: fixed splicing for static deployment, adaptive splicing for changes in content in overlapping areas, AI-assisted splicing, and semantic segmentation to avoid mis-splitting;
[0044] Secondly, the specific method of fixed stitching is as follows: read the coordinates of the pre-calibrated homography matrix, use fixed fusion weights, for example, the closer to the image of camera A, the higher the weight of A, the closer to the image of camera B, the higher the weight of B, the overlapping areas overlap, the lower the weight, and the overlapping areas of the images become lighter when fused.
[0045] Adaptive stitching: SIFT / ORB feature points are re-detected in the overlapping area of the current frame to obtain the corner points or edge intersections of the image. RANSAC is used to remove mismatches, and new homography matrix image coordinates are calculated. The fusion weights are dynamically adjusted according to the content of the overlapping area. For example, if one side image A is blurry, its weight is reduced; if one side image B has strong light overexposure, the other side is used first for image fusion.
[0046] AI-assisted stitching: Real-time semantic segmentation is performed on each video frame to output pixel-level label maps: distinguishing: static background (ground, buildings), dynamic foreground (people, vehicles), and invalid areas (occlusion, blur). In images with static backgrounds, the images are fused according to matrix coordinates. Dynamic images retain their original viewpoints without forced stitching, or are filled using temporal interpolation. At the same time, invalid areas are stitched together. For example: a tree occludes an intersection in camera A, but is visible in camera B → only the image from camera B is used to fill the area.
[0047] Secondly, the specific workflow for dynamically switching splicing modes is as follows:
[0048] S11: When the system starts, it first loads the basic parameters of all cameras, including position, orientation, focal length, field of view, etc., and pre-calculates the potential overlap area between each pair of adjacent cameras based on this information. For fixed-installation and stable camera combinations, the system will generate a fixed homography matrix in advance through the calibration process. The homography matrix specifically refers to generating several sets of coordinates for multiple objects in the image based on the space in the area, which serve as the position code of the objects.
[0049] S22: Perform real-time analysis on the images, analyzing whether the images within the matrix overlap area of multiple images taken by multiple cameras are clear, whether there are overly bright or dark areas, whether the lighting has changed significantly, whether a large number of moving targets have entered the overlap area, whether there are occlusions causing some areas to be invisible, and whether the feature point matching is stable (feature point matching refers to the matching between repeated objects in multiple sets of images).
[0050] S33: Based on the above analysis, the system determines which stitching mode should be used according to priority: if the scene is stable—that is, the image is clear, the lighting is uniform, there is no obstruction, and there are no dense moving targets—then the fixed stitching mode is enabled.
[0051] The system directly uses the pre-stored homography matrix to align the two images and fuses them according to preset fixed weights. This method has low computational cost, high efficiency, and is suitable for long-term operation.
[0052] If the scene experiences changes in image quality—for example, one side of the image becomes blurry, the other side experiences overexposure due to strong light, or the feature matching error increases due to slight camera displacement caused by temperature changes—then switch to adaptive stitching mode.
[0053] At this point, the system will re-detect stable feature points (such as corner points and edge intersections) in the overlapping area of the current frame. The SIFT or ORB algorithm is commonly used. The SIFT algorithm performs multiple Gaussian blurs on the original image to generate a set of "increasingly blurry" images. Local extrema are found in these images. An extrema is a pixel that is brighter or darker than its 26 neighbors. Then, the extrema in the image are removed.
[0054] Next, the RANSAC algorithm is used to remove incorrect matches and the feature points applicable to the current frame are recalculated to achieve more accurate geometric alignment. At the same time, the system will dynamically adjust the fusion weights according to the actual quality of the two images: the side with poor quality reduces its participation and the side with good quality dominates the output, thereby ensuring the visual consistency of the fusion region.
[0055] If the scene is complex and there is semantic interference, such as a large number of pedestrians and vehicles passing through the overlapping area, or a road being blocked by leaves, pillars, etc., the system will activate the AI-assisted splicing mode.
[0056] In this mode, the system performs real-time semantic segmentation on each image and video frame, generating pixel-level label maps that clearly distinguish three types of regions:
[0057] Static backgrounds, including the ground, walls, and fixed structures: This part can be safely geometrically aligned and blended without the need for excessive algorithms to re-integrate features, resulting in rapid video stream splicing.
[0058] Dynamic foreground, including people and vehicles: To avoid the target being "torn" or "half-cut", the system retains its complete form in the original camera view and does not perform cross-frame averaging. If the target happens to be located at the stitching boundary, the system will combine the motion trajectory of the previous and next frames and fill in the missing part through temporal interpolation.
[0059] Specifically, in multi-camera video stitching, if a pedestrian or vehicle happens to be located in the overlapping boundary area of two camera images, the traditional pixel-weighted fusion method will "split" the target in two: the left side comes from camera A, and the right side comes from camera B. When the target crosses the stitching boundary, and part of it is cropped or occluded in the main view of the current frame, for example, only half of the car is visible, the system will not leave it blank or fill it with background, but will instead activate the temporal context reasoning mechanism:
[0060] 1. Utilize the complete appearance and motion trajectory of the target in the first and last few frames;
[0061] 2. Predict the location and shape of the missing part in the current frame by optical flow estimation or motion vector extrapolation;
[0062] Third, use temporal interpolation (such as inter-frame copying based on motion compensation or generative completion) to synthesize reasonable pixel content and fill in the gaps.
[0063] Invalid areas, such as those that are obstructed, severely blurred, or have lens smudges: The system checks whether another camera can see the area clearly. If it can (for example, if A is blocked by a tree at the intersection, but B can see it clearly), the corresponding position is filled directly with B's image to achieve intelligent content completion.
[0064] In this embodiment, specifically: the video source in S1 refers to a network camera, analog camera, high-definition camera, video recording device, streaming media server, and mobile device.
[0065] In this embodiment, specifically, the method for establishing the mapping between physical location and logical coordinates in S1 is as follows:
[0066] Given the camera's location (GPS) and orientation, and combining the focal length and sensor size, calculate the projected polygon of the view frustum on the ground. Mark the camera's field of view in the GIS platform, establish the correspondence between the four corner points of the image and their geographic coordinates, and use affine transformation, polynomial correction, or grid interpolation to establish a nonlinear mapping. This supports mapping any latitude and longitude point to the pixel position in the video frame.
[0067] Further: Forward mapping (geography → pixel): Given the latitude and longitude of a ground point, the pixel coordinates of that point in a camera's view can be calculated;
[0068] Reverse mapping (pixel → geography): Given a pixel in the image, its corresponding latitude and longitude can be deduced;
[0069] In this embodiment, specifically: the method for correcting the video stream using timestamps in S2 is as follows:
[0070] The network deploys a PTP master clock, the switch supports transparent clocks, and the camera acts as a PTP slave clock. After synchronization, the internal clock error is less than 1μs. Each frame is accompanied by a PTP timestamp, and all frames directly use PTP time as the absolute time without conversion.
[0071] Further: The correction of video streams by timestamps relies on the full-link deployment of the high-precision time synchronization protocol PTP: The system deploys a PTP master clock in the network core. Switches that support PTP transparent clock function automatically record and compensate for dwell time and link delay when forwarding PTP packets. Each camera acts as a PTP slave clock and locks the master clock through the best master clock algorithm (BMC). Hardware timestamp units are used to achieve sub-microsecond clock discipline, so that the deviation between the real-time clock (RTC) inside all cameras and the master clock is controlled within 1 microsecond.
[0072] In this embodiment, specifically, the geometric correction and photometric correction methods in S3 are further divided into:
[0073] Geometric correction: lens distortion correction, angle correction, fisheye / panoramic image unfolding;
[0074] Photometric correction: white balance adjustment, brightness and contrast equalization adjustment, histogram matching.
[0075] In this embodiment, the specific steps for lens distortion correction, angle correction, and fisheye / panoramic image unfolding are as follows:
[0076] Lens distortion correction: The camera is calibrated in advance using a calibration board to record its "deformation rules". During real-time processing, the curved image is "straightened" according to these rules to restore the true geometric structure, making the originally curved corridor edges, door frames, and driveway lines straight.
[0077] Viewpoint correction: Manually or automatically select a ground area in the image, and the system will "flatten" this slanted area to make it look like it is viewed from directly above, so that the positions of vehicles and pedestrians in the image are more consistent with the real map relationship, which is convenient for cross-camera tracking;
[0078] Fisheye / panoramic image unfolding: The entire spherical image is "unfolded" into a wide planar image, and the objects that were originally squeezed together are reasonably separated, making them easier to identify and stitch together.
[0079] Further: Lens distortion correction is based on obtaining the intrinsic parameter matrix and distortion coefficients through camera calibration. It uses reverse mapping to infer the position of each output pixel in the original distorted image, and reconstructs the distortion-free image through interpolation to restore the straightness of the line.
[0080] Viewpoint correction involves selecting four corresponding points on the ground plane in the image, calculating the corresponding matrix H from the current tilted viewpoint to the virtual orthogonal overhead viewpoint, and projecting the region of interest onto a unified bird's-eye view plane, thereby eliminating perspective distortion, making the object's trajectory linearly correspond to the real geographic coordinates, and supporting cross-camera target association.
[0081] Fisheye / panoramic images are unfolded based on lens projection models, such as equidistant projection and equal solid angle projection, by unfolding spherical or hemispherical imaging data into a rectangular plane according to latitude and longitude grids.
[0082] Further: White balance adjustment is based on the gray world assumption or reference white point detection. By statistically analyzing the pixel areas in the image that are close to neutral gray or white, the gain coefficients of the red (R), green (G), and blue (B) channels are calculated, and each channel is linearly scaled to eliminate color cast caused by differences in light source color temperature, ensuring that white objects present a consistent neutral color on different cameras.
[0083] Brightness and contrast balance adjustment usually employs adaptive gamma correction, CLAHE, or dynamic range compression methods based on local mean-variance. This method enhances the visibility of dark areas and suppresses overexposure of highlights while preserving details, making the illumination transition between adjacent video areas at the splicing boundary natural and avoiding abrupt changes in brightness.
[0084] Histogram matching uses a high-quality or center-view video stream as a reference and aligns it using the cumulative distribution function (CDF). This forces the statistical distribution of brightness or color channels in other video streams to be consistent with the reference image, thereby achieving stylistic uniformity in overall tone, saturation, and brightness levels.
[0085] In this embodiment, specifically, the method for extracting feature points in the overlapping region of adjacent video streams in S4 is as follows:
[0086] For two adjacent video frames from different cameras, their potential overlapping areas are first delineated by camera orientation and field of view. Feature points are detected in these two areas and their descriptors are calculated. By matching the descriptors, pairs of points with the same name are found, that is, pairs of pixel coordinates representing the same physical point.
[0087] Further: Based on the known extrinsic and intrinsic parameters of the camera, the common visible area of the two videos in physical space is estimated on their respective image planes by ray backprojection or frustum intersection calculation, thereby narrowing the feature detection range from the entire image to this local area, improving efficiency and matching robustness;
[0088] Key points are extracted in two overlapping regions using scale-invariant or affine-invariant feature detection algorithms, and descriptors with robustness to rotation, scale, and illumination are generated.
[0089] Next, the two sets of descriptors are matched by an approximate nearest neighbor search, and erroneous matches are eliminated by combining geometric constraints. Finally, the pairs of corresponding points that satisfy the epipolar geometry or homography model are retained.
[0090] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method and system for dynamically stitching heterogeneous security video streams, characterized in that, Includes the following steps: S1: Users add camera information in the management interface. The system calls the corresponding adapter according to the protocol type, automatically scans video sources in the network, and establishes a mapping between physical location and logical coordinates. S2: Video stream frames enter the time synchronization buffer. The system uses the master clock source as a reference to correct the timestamps of each stream frame. S3: Perform geometric correction, photometric correction, and de-editing on each video stream; S4: Extract feature points of overlapping areas of adjacent video streams, dynamically switch splicing modes according to scene changes, and simultaneously fuse and render the video streams; S5: Pushes the stitched video stream to the client or platform via RTMP, WebRTC, or HLS, and supports on-demand cropping, scaling, and view switching. It also overlays AI analysis results on the stitched image and supports clicking on the stitched image to check the original camera footage.
2. The dynamic splicing method and system for heterogeneous security video streams according to claim 1, characterized in that: In S1, the video source refers to network cameras, analog cameras, high-definition cameras, video recording equipment, streaming media servers, and mobile devices.
3. The method and system for dynamic splicing of heterogeneous security video streams according to claim 2, characterized in that: The specific method for establishing the mapping between physical location and logical coordinates in S1 is as follows: Given the camera's location (GPS) and orientation, and combining the focal length and sensor size, calculate the projected polygon of the view frustum on the ground. Mark the camera's field of view in the GIS platform, establish the correspondence between the four corner points of the image and their geographic coordinates, and use affine transformation, polynomial correction, or grid interpolation to establish a nonlinear mapping. This supports mapping any latitude and longitude point to the pixel position in the video frame.
4. The dynamic splicing method and system for heterogeneous security video streams according to claim 3, characterized in that: The specific method for correcting the video stream using timestamps in S2 is as follows: The network deploys a PTP master clock, the switch supports transparent clocks, and the camera acts as a PTP slave clock. After synchronization, the internal clock error is less than 1μs. Each frame is accompanied by a PTP timestamp, and all frames directly use PTP time as the absolute time without conversion.
5. The dynamic splicing method and system for heterogeneous security video streams according to claim 4, characterized in that: The geometric correction and photometric correction methods in S3 are further divided into: Geometric correction: lens distortion correction, angle correction, fisheye / panoramic image unfolding; Photometric correction: white balance adjustment, brightness and contrast equalization adjustment, histogram matching.
6. The dynamic splicing method and system for heterogeneous security video streams according to claim 5, characterized in that: The specific steps for lens distortion correction, angle correction, and fisheye / panoramic image unfolding are as follows: Lens distortion correction: The camera is calibrated in advance using a calibration board to record its "deformation rules". During real-time processing, the curved image is "straightened" according to these rules to restore the true geometric structure, making the originally curved corridor edges, door frames, and driveway lines straight. Viewpoint correction: Manually or automatically select a ground area in the image, and the system will "flatten" this slanted area to make it look like it is viewed from directly above, so that the positions of vehicles and pedestrians in the image are more consistent with the real map relationship, which facilitates cross-camera tracking; Fisheye / panoramic image unfolding: The entire spherical image is "unfolded" into a wide planar image, and the objects that were originally squeezed together are reasonably separated, making them easier to identify and stitch together.
7. The dynamic splicing method and system for heterogeneous security video streams according to claim 6, characterized in that: The specific steps for white balance adjustment, brightness and contrast equalization adjustment, and histogram matching are as follows: White balance adjustment: Automatically analyzes the neutral colors in the image, adjusts the intensity of the red, green and blue channels, and makes the white adjustment whiter, so that the color style of multiple images is consistent and no obvious color blocks appear when stitching. Brightness and contrast balance adjustment: enhances details in dark areas, suppresses overexposed areas, makes the overall brightness and darkness more balanced, and eliminates the abrupt feeling of "one side bright and one side dark" at the seams. Histogram matching: Select one channel as the "reference image", and the other channels will automatically adjust their color distribution to be as close as possible to the reference, so that the whole stitched image looks like it was taken by the same device.
8. The dynamic splicing method and system for heterogeneous security video streams according to claim 7, characterized in that: The specific method for extracting feature points in the overlapping region of adjacent video streams in S4 is as follows: For two adjacent video frames from different cameras, their potential overlapping areas are first delineated by camera orientation and field of view. Feature points are detected in these two areas and their descriptors are calculated. By matching the descriptors, pairs of points with the same name are found, that is, pairs of pixel coordinates representing the same physical point.