Smoke and fire target identification method based on multi-channel video stitching

By generating high-quality panoramic images through multi-channel video stitching technology and combining them with intelligent fire analysis, the problems of large blind spots, isolated data, difficulty in coordination, and high false alarm and missed alarm rates in forest fire monitoring have been solved, realizing full-area, high-precision, and real-time forest fire monitoring.

CN121415136APending Publication Date: 2026-01-27HARBIN XINGUANG OPTIC-ELECTRONICS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511567155.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing forest fire monitoring technologies suffer from problems such as large monitoring blind spots, isolated data, difficulty in coordination, and high false alarm and missed alarm rates, making it difficult to meet the needs of smart forestry for high-precision, wide-coverage, and all-weather monitoring.

Method used

Using multi-channel video stitching technology, high-quality panoramic stitched images are generated through image preprocessing, feature detection and matching, scale scaling and rotation transformation, mesh optimization and texture mapping synthesis, and combined with the fire intelligent analysis module to realize fire detection.

Benefits of technology

It enables continuous panoramic forest fire monitoring across the entire area, with low cost, high accuracy, minimal environmental impact, and wide applicability, making it suitable for real-time monitoring of large-scale fire spread trends.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121415136A_ABST
    Figure CN121415136A_ABST
Patent Text Reader

Abstract

The invention provides a smoke and fire target identification method based on multi-channel video stitching, and belongs to the field of forest fire monitoring. The problems of large monitoring blind area, data isolation, difficult collaboration and high false report and missing report rate in the prior art are solved. The method comprises the following steps: preprocessing an image acquired by a multi-camera hardware system; image feature detection and matching are carried out, correlation information between sequences is found out, mismatching points are eliminated through image registration verification so as to improve robustness, and high-precision image matching points are generated; necessary scale zooming and rotation transformation adjustment are carried out on the images, and the camera focal length and the three-dimensional rotation attitude of each image are accurately estimated; a grid optimization technology is adopted, alignment precision between images is optimized in a global level, seams are minimized, optimized image data are seamlessly fused through a texture mapping synthesis technology, and a high-quality panoramic spliced image is output; and a fire behavior intelligent analysis module is constructed to realize fire behavior detection. The system is mainly used for forest fire monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of forest fire monitoring, and in particular relates to a method for identifying smoke and fire targets based on multi-channel video stitching. Background Technology

[0002] Currently, forest fire monitoring mainly relies on manual patrols, lookout towers, drone swarms, and fixed PTZ cameras. Manual patrols and lookout towers are limited by terrain obstructions and manpower, making it difficult to achieve continuous 24 / 7 coverage, resulting in large blind spots and high missed detection rates. While drones can carry multiple sensors, image processing is still primarily based on single frames or scattered perspectives, lacking spatial continuity and unable to support the judgment of large-scale fire spread trends. Fixed PTZ cameras are limited by physical angle constraints, limiting their monitoring range and making it difficult to simultaneously capture scattered fire sources, let alone visually present the global relationship between fire sources and surrounding vegetation and terrain. Furthermore, handheld radar patrols are inefficient, have short battery life, and high labor costs, resulting in severely insufficient continuous monitoring capabilities. Existing technologies generally suffer from large monitoring blind spots, isolated data, difficulty in coordination, and high false alarm and missed detection rates, failing to meet the urgent needs of smart forestry for high-precision, wide-coverage, and all-weather forest fire monitoring.

[0003] In summary, a method for identifying fireworks targets based on multi-channel video stitching is proposed. Summary of the Invention

[0004] In view of this, the present invention aims to propose a fireworks target recognition method based on multi-channel video stitching, so as to solve the problems of large monitoring blind spots, isolated data, difficulty in coordination, and high false alarm and false negative rates in the existing technology.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a method for identifying fireworks targets based on multi-channel video stitching, comprising the following steps:

[0006] S1. Preprocess the images acquired by the multi-camera hardware system;

[0007] S2. Perform image feature detection and matching to find the correlation information between sequences. Eliminate mismatched points through image registration verification to improve robustness and generate high-precision image matching points.

[0008] S3. Perform necessary scaling and rotation transformations on the images, and accurately estimate the camera focal length and three-dimensional rotation attitude of each image.

[0009] S4. Employing mesh optimization technology, the alignment accuracy between images is optimized and seams are minimized at the global level. The optimized image data is seamlessly integrated through texture mapping synthesis technology to output a high-quality panoramic stitched image.

[0010] S5. Construct a fire situation intelligent analysis module to realize fire situation detection.

[0011] Furthermore, the image preprocessing in S1 includes:

[0012] S11. Distortion Correction: The Zhang Zhengyou calibration method is used to calibrate the intrinsic parameters of each camera, and the coordinates are corrected using the correction formula.

[0013] S12. The multi-scale Retinex algorithm is used to decompose the image into illuminance and reflectance components. The illuminance component reflects global illumination changes, while the reflectance component reflects the details of vegetation, smoke, and fire. Logarithmic transformation and Gaussian filtering are performed on the illuminance component to suppress strong light areas and enhance details in shadow areas. The image is reconstructed by multiplying the reflectance component by the adjusted illuminance component to ensure that the brightness deviation of images captured by different cameras is ≤8%.

[0014] S13. For salt and pepper noise in forest scenes, median filtering is used; for blurred Gaussian noise caused by fog, wavelet transform is used for noise reduction; the effect is verified by peak signal-to-noise ratio (PSNR) after noise reduction.

[0015] Furthermore, the modified formula is as follows:

[0016]

[0017] Where: k1, k2, k3 are radial distortion coefficients; p1, p2 are tangential distortion coefficients; (x dist ,y dist (x) represents the coordinates of the distorted pixel; undist ,y undist () represents the ideal coordinates;

[0018] Furthermore, S21 and SURF feature point detection generate high-quality feature points, and the homography matrix optimizes matching through global constraints to ensure accurate fitting between virtual objects and physical planes. By utilizing the initially estimated relative pose information, the quality and accuracy of matching points are improved, and a Hessian matrix is ​​constructed.

[0019] S22. Guiding feature point matching based on the initial relative pose between images;

[0020] S23. Use the Random Sample Consensus Algorithm (RANSAC) to further filter out erroneous feature matching points.

[0021] Furthermore, the Hessian matrix is:

[0022]

[0023] c(x,y,σ)=D xx D yy -(0.9D xy) 2

[0024] Where: D xx D xy and D yy The approximate convolution value obtained using the box filter; the value of the determinant of c(x,y,σ).

[0025] Furthermore, S31, estimate the fundamental matrix F and the essential matrix E;

[0026] S32, Camera focal length f estimation;

[0027] S33. Minimize reprojection error through bundle adjustment optimization;

[0028] S34. Output data that meets the reprojection error requirements.

[0029] Furthermore, S4 includes;

[0030] S41. Using a cylindrical projection algorithm, the image is projected from a two-dimensional plane to a cylindrical coordinate system. The focal length 3D transformation simulates the transformation of the camera lens. That is, by adjusting the focal length, the viewing distance and depth information in the image are controlled, so as to maintain a consistent perspective effect and reduce horizontal distortion when images from different perspectives are stitched together.

[0031] S42. Divide the image into a grid, calculate the distance weight between the center point of each grid and the interior point, reconstruct the DLT matrix and solve for local homography;

[0032] S43. The local homography matrix in each small square is Gaussian weighted using the Moving DLT algorithm, and the image stitching is completed by transforming the relationship.

[0033] S44. Using transparency-weighted fusion: In overlapping areas, a transparency-based weighting method controls the contribution ratio of each image by dynamically adjusting the transparency of the images, thereby smoothly transitioning the stitching areas.

[0034] S45, outputs seamless panoramic images.

[0035] Furthermore, the weighted fusion algorithm formula in S44 is as follows:

[0036]

[0037] Where: I(x,y) is the gray value of the fused image, I1(x,y) is the gray value of the left image, S1 is the region of I1(x,y) after removing the overlapping area on the image, I2(x,y) is the gray value of the right image, S2 is the region of I2(x,y) after removing the overlapping area on the image, the overlapping area after stitching is S, w1 and w2 are the weights of the overlapping area, and w1+w2=1, w1∈(0,1).

[0038] Furthermore, the construction of the fire situation intelligent analysis module in S5 includes the following steps:

[0039] S51. Construct a forest fire dataset. Expand the sample dataset by random rotation, scaling, flipping, brightness adjustment and adding smoke blurring strategy. Use LabelImg tool to label fire point areas. The label format is VOC format.

[0040] S52. Model Training and Optimization: The basic model is YOLOv8-Tiny, and the initial weights are the pre-trained weights from the COCO dataset; Forest Scene Adaptation: Modify the feature fusion module in the neck of the model to increase cross-layer connections between shallow and deep features and enhance the extraction of small fire point features.

[0041] S53. Real-time identification and reasoning: When a fire point is identified and an alarm is triggered, it can be pushed to the visualization terminal via the cloud;

[0042] S54. Coordinate Mapping Relationship Construction: Based on Camera GPS Coordinates (X... cam ,Y cam Z cam Using the intrinsic parameter K, establish the pixel coordinates (u,v) and geographic coordinates (X) of the panoramic image. gco ,Y gco Mapping model of )

[0043] S55. Ground marker calibration: Select one type of ground marker from the panoramic image.

[0044] Furthermore, the mapping model is as follows:

[0045]

[0046] Where: s is the depth factor, estimated based on forest terrain DEM data, R is the camera pose matrix, (c x ,c y ) represents the coordinates of the camera's principal point, and K represents the camera's intrinsic parameter matrix.

[0047] Beneficial effects:

[0048] 1. By applying multi-source image fusion and intelligent analysis technology, high-definition cameras are deployed and spatiotemporally synchronized through the "hexagonal coverage rule" to obtain a continuous panoramic view of the entire area, thereby providing a more complete spatial representation of forest fire characteristics;

[0049] 2. This method integrates SURF feature detection, Moving DLT mesh optimization, and cylindrical projection image stitching techniques. The core process includes: first, image feature detection and matching to identify correlations between sequences; then, image registration verification to eliminate mismatches and improve robustness; next, generating high-precision image matching points; then, accurately estimating the camera focal length and 3D rotation attitude of each image; accordingly, adjusting the images through necessary scaling and rotation transformations; further, employing mesh optimization techniques to optimize the alignment accuracy between images globally and minimize seams; finally, seamlessly fusing the optimized image data using texture mapping synthesis techniques to output a high-quality panoramic stitched image, which, combined with a fire intelligent analysis module, enables comprehensive, high-precision, and real-time monitoring of forest areas. Compared to previous methods, this method is lower in cost, higher in accuracy, less affected by environmental conditions, and applicable to a wider range of scenarios. Attached Figure Description

[0050] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0051] Figure 1 This is a flowchart illustrating the overall workflow of a fireworks target recognition method based on multi-channel video stitching as described in this invention.

[0052] Figure 2 This is a flowchart of a multi-channel camera panoramic stitching process for a fireworks target recognition method based on multi-channel video stitching, as described in this invention.

[0053] Figure 3 The flowchart of the RANSAC algorithm for a fireworks target recognition method based on multi-channel video stitching described in this invention is shown below.

[0054] Figure 4 Figure (a) shows the cylindrical projection principle of the firework target recognition method based on multi-channel video stitching described in this invention.

[0055] Figure 5 Figure (b) shows the cylindrical projection principle of the fireworks target recognition method based on multi-channel video stitching described in this invention.

[0056] Figure 6 Figure (c) illustrates the cylindrical projection principle of the firework target recognition method based on multi-channel video stitching described in this invention.

[0057] Figure 7 This is a flowchart of the focal length 3D transformation process for a fireworks target recognition method based on multi-channel video stitching as described in this invention.

[0058] Figure 8 This is a schematic diagram of the cylindrical projection effect of the fireworks target recognition method based on multi-channel video stitching described in this invention;

[0059] Figure 9 This is a flowchart of the intelligent fire analysis process for a smoke and fire target recognition method based on multi-channel video stitching as described in this invention.

[0060] Figure 10 This is a system architecture diagram of a fireworks target recognition method based on multi-channel video stitching as described in this invention. Detailed Implementation

[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present invention can be combined with each other, and the described embodiments are only some embodiments of the present invention, not all embodiments.

[0062] It should be noted that the descriptions of "left," "right," "left side," "right side," "upper part," "lower part," "top," and "bottom" in this invention are defined based on the orientation or positional relationships shown in the accompanying drawings. They are merely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the described structure must be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0063] In the description of this invention, unless otherwise expressly specified and limited, the terms "installation," "connection," and "joining" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0064] See appendix Figures 1 to 10 The present invention describes a method for identifying fireworks targets based on multi-channel video stitching. The overall architecture of the invention is as follows: Figure 10As shown, the system adopts a three-layer distributed structure: the perception layer collects data synchronously through multiple cameras deployed in mountainous areas; the processing layer utilizes edge computing servers to sequentially complete image preprocessing, panoramic stitching, and AI fire analysis to accurately locate fire points; finally, the application layer realizes panoramic visualization, electronic map positioning, and multi-terminal early warning. This system achieves an automated closed loop from fire detection and location to early warning, greatly improving the efficiency and response speed of forest fire prevention.

[0065] like Figure 1 As shown: Preprocessing of images acquired by a multi-camera hardware system.

[0066] To address image interference in forest scenes (such as sudden changes in lighting, foliage obstruction, and fog), the preprocessing workflow is as follows:

[0067] Distortion correction: The Zhang Zhengyou calibration method was used to calibrate the intrinsic parameters of each camera to obtain the focal length f. x / f y Principal point coordinates (c x ,c y ) and distortion coefficients (radial distortion k1, k2, k3, tangential distortion p1, p2); based on the calibration parameters, the distorted pixel coordinates (x) are corrected using the following formula. dist ,y dist to ideal coordinates (x) undist ,y undist ):

[0068]

[0069] in, Ensure that the edges of the corrected image are not stretched (e.g., the tree trunk is not bent).

[0070] Illumination equalization: The multi-scale Retinex algorithm is used to decompose the image into "illuminance component" and "reflection component". The illuminance component reflects global illumination changes, while the reflection component reflects detailed features such as vegetation, smoke, and fire. Logarithmic transformation and Gaussian filtering are applied to the illuminance component to suppress strong light areas and enhance details in shadow areas. Image reconstruction: Reflection component × adjusted illuminance component to ensure that the brightness deviation of images captured by different cameras is ≤8%.

[0071] Noise suppression: Median filtering was used for salt and pepper noise in the forest scene; wavelet transform was used for denoising the blurred Gaussian noise caused by fog; the effect was verified by peak signal-to-noise ratio (PSNR) after denoising.

[0072] like Figure 2 As shown: Image feature detection and matching are performed to find the correlation information between sequences. Image registration verification is used to eliminate mismatched points to improve robustness and generate high-precision image matching points.

[0073] Further feature detection and matching are performed on the preprocessed image: The SURF feature point detection method is a very practical trade-off, achieving a good balance between speed, robustness, and accuracy. In terms of speed, SURF uses an integral image to quickly calculate the sum of pixels in the image region, and uses a box filter instead of the Gaussian second derivative to avoid time-consuming convolution operations. Furthermore, it uses the Haar wavelet response instead of the gradient histogram to reduce computational complexity. The integral image II(x,y) is calculated using the following formula to achieve fast calculation of the pixel sum in any rectangular region:

[0074] II(x,y)=I(x,y)+II(x-1,y)+II(x,y-1)-II(x-1,y-1)

[0075] In terms of robustness, SURF simulates scale space through changes in the size of box filters and calculates the principal direction based on Haar wavelet responses, ensuring feature rotation invariance and exhibiting good stability against noise and illumination variations. Regarding accuracy and stability, SURF generates 64 / 128-dimensional floating-point descriptors, demonstrating high discriminative power for complex textures and superior scale-space continuity.

[0076] Homography matrices can accurately describe the mapping relationship between two views, with clear transformation relationships and controllable computational complexity. Homography matrices are chosen for global geometric constraint optimization matching results. SURF quickly generates high-quality feature points, and the homography matrix, through global constraint optimization matching, ensures accurate alignment of virtual objects with the physical plane, achieving a balance between speed and accuracy.

[0077] This method improves the quality and accuracy of matching points by utilizing initially estimated relative pose information, and is applied to multi-view geometry problems, especially when dealing with large viewpoint differences or mismatches between images. The key idea is to introduce pose priors during the matching process, optimizing the geometric consistency of matching point pairs to reduce mismatches and improve the accuracy of the final reconstruction. First, feature points in the image are extracted using the SURF feature detection algorithm. The SURF algorithm uses an approximate Haar wavelet method to extract feature points, based on a Hessian determinant-based blob feature detection method. This method uses a box filter on the integral image to simplify the second-order differential template and construct the Hessian matrix elements, shortening the feature extraction time and improving efficiency. The approximate Hessian matrix and its determinant values ​​are as follows:

[0078]

[0079] c(x,y,σ)=D xx D yy -(0.9D xy ) 2

[0080] Where D xx D xy and D yy This is an approximate convolution value obtained using a box filter. If c(x,y,σ) is greater than the set threshold, the pixel is determined to be a keyword. Non-maximum suppression is performed in a 3×3×3 pixel neighborhood centered on the keyword. Finally, the precise localization of the SURF feature point is achieved by interpolating the blob features.

[0081] Then, feature point matching is guided by the initial relative pose between images (e.g., based on coarse extrinsic parameters or camera pose estimation). For example, multiple visible light cameras are placed horizontally, and the acquired images are stitched together. The slope of the line segment connecting matching feature points in multiple images should be close to zero. If the slope is greater than a certain threshold, it can be determined as an incorrect match. Based on this prior information, the matching process can search for matching points within a more reasonable range, thereby avoiding incorrect matches that do not conform to geometric constraints.

[0082] Furthermore, to further filter out erroneous feature matching points, the Random Sample Consensus (RANSAC) algorithm is used to optimize the initially screened matching points to obtain better matching point pairs. RANSAC is a highly robust model fitting algorithm used to estimate mathematical model parameters from data containing a large number of outliers. This algorithm assumes that the data consists of inliers (correct data that conforms to the model) and outliers (noise or erroneous data). It finds the largest consistent set (inlier set) that supports the model by randomly sampling the smallest subset to fit the model. Furthermore, based on the Monte Carlo method, multiple random samplings are used to cover potentially correct models, avoiding local optima. The RANSAC algorithm process is as follows: Figure 3 As shown:

[0083] For each data point p i Calculate its error e to the model. i :

[0084] e i =dist(M,p i )

[0085] If e i If p < δ (δ is a preset threshold), then p is determined. i For interior points, RANSAC requires iterative sampling to ensure that at least one sample does not contain exterior points.

[0086]

[0087] Where, p success The expected success probability is set to 0.99 by default, ε is the proportion of outliers, and k is the minimum sample set size.

[0088] The application of combining cylindrical projection algorithms with focal length 3D transformation can better handle wide-angle or panoramic image stitching, especially in scenes with a wide dynamic range and large changes in viewing angle. Cylindrical projection provides an effective geometric transformation for image stitching, while focal length 3D transformation simulates the focal length effect of a camera, making the image projection more realistic, especially when depth and three-dimensional spatial position are involved.

[0089] The basic principle of cylindrical projection is illustrated as follows: Figure 4 , Figure 5 , Figure 6 This allows us to obtain the pixel space transformation relationship between the projected image plane ABC and the projected cylinder DBE. Through this transformation relationship, the pixels on the plane can be compressed onto the cylinder.

[0090] like Figure 5 The quadrilateral GHEF represents the original image to be processed. After projection, it becomes the curved surface JDILCK (marked by yellow dots).

[0091] The top view is as follows Figure 6 DCE is the image plane to be processed, and FCG is the surface obtained by projection.

[0092] Let the original image have width W, height H, and angle FOG be the camera's field of view angle α (typically 45°, i.e., PI / 4). The radius (focal length) f of the circle is:

[0093]

[0094] The width (length of curve FCG) W′ of the target image is calculated sequentially:

[0095] W′=f×α

[0096] The height H′ of the target image remains unchanged, H′=H, that is:

[0097]

[0098] Focal length 3D transformation refers to explicitly considering the camera's position and internal parameters in space within an image stitching framework. This projects points from the image back into three-dimensional space, and then, based on the relative positions of adjacent cameras, reprojects them onto the final stitching plane. Focal length estimation is achieved through... Figure 7 Process for building a realistic 3D projection model:

[0099] Based on the camera projection model, the projection relationship between the image point coordinate x and the 3D point X is as follows:

[0100] λx=K[R|t]X

[0101] in:

[0102]

[0103] Where R,t are camera extrinsic parameters and λ is the depth scaling factor. Matching point pairs is constrained according to the fundamental matrix. satisfy:

[0104]

[0105] The fundamental matrix F can be decomposed into:

[0106] F = K -T [t] × RK -1

[0107] Where [t] × It is the antisymmetric matrix of the translation vector.

[0108] The images are scaled and rotated as necessary, and the camera focal length and 3D rotation attitude of each image are accurately estimated.

[0109] Based on geometric constraints, F is estimated using the 8-point method, and then the intrinsic parameter constraints are substituted in:

[0110] F = K -T EK -1

[0111] The essential matrix E satisfies:

[0112] E = [t] × R

[0113] Let the image width be w, height be h, and principal point be h. Derivation of focal length:

[0114]

[0115] Minimize reprojection error through bundle adjustment optimization:

[0116]

[0117] Where π(·) is the projection function, x ij Let j be the observation of the j-th point in the i-th image.

[0118] By employing mesh optimization technology, the alignment accuracy between images is optimized and seams are minimized at the global level. The optimized image data is then seamlessly fused using texture mapping synthesis technology to output a high-quality panoramic stitched image.

[0119] Cylindrical projection algorithms, by projecting images from a two-dimensional plane to a cylindrical coordinate system, can better maintain the consistency of the image along the horizontal axis and reduce horizontal distortion, making them particularly suitable for panoramic image stitching. Focal length 3D transformation can simulate the changes in a camera lens, taking into account the impact of focal length on image depth and perspective. By adjusting the focal length, the viewing distance and depth information in the image can be controlled, thus ensuring a consistent perspective effect when stitching images from different viewpoints and avoiding unnatural distortion.

[0120] Through the previous steps of matching point detection, generation, and camera parameter estimation, an initial image transformation relationship is obtained. Image stitching mesh optimization is a technique to improve stitching quality and accuracy, playing a crucial role, especially in processing large-scale image stitching. The core idea is to reduce distortion and unnatural transitions in overlapping areas between images by optimizing the image transformation parameters. This algorithm divides the image into a mesh, calculates the distance weights between the center point of each mesh and its interior points, reconstructs the DLT matrix, solves for local homography, and finally generates the stitching result through weighted fusion, thereby improving the stitching effect.

[0121] Suppose images A and B are a set of images that need to be stitched together, then the detected feature point pairs between them are p = [x, y, 1]. T Given q = [x′, y′, 1], the correspondence between feature point pairs is as follows:

[0122]

[0123] The transformation relationship between homogeneous coordinates p and q is expressed as follows:

[0124] q~Hp

[0125] In the above formula, matrix H ∈ R 3×3 H = [h1, h2, h3, h4, h5, h6, h7, h8, h9], since p and q are in the same direction, o 3×1 = q × Hq, therefore, substituting the feature points p and q into the matrix H yields:

[0126]

[0127] The above equation can be transformed into A·H = 0, that is:

[0128]

[0129] The calculated h can be expressed as:

[0130]

[0131] After obtaining the homography matrix H, the local homography matrix in each small square is Gaussian-weighted using the Moving DLT algorithm. The local homography matrix within the grid can then be expressed as:

[0132]

[0133] Where j is the number of n×n image blocks into which the image is divided, W j After obtaining the local homography matrix of each grid by assigning Gaussian weights to each grid, the image stitching is completed through transformation relationships.

[0134] Due to factors such as parallax, stitching artifacts exist after image registration. Image fusion refers to combining image information from different sources or perspectives during the image stitching process to generate a seamless composite image. The core task of image fusion is to ensure a natural transition between stitched images, reduce seams and color differences between images, and optimize visual effects, making the stitched image as close as possible to the real scene in terms of quality and detail. Transparency-weighted fusion is used: in overlapping areas, a transparency-based weighting method dynamically adjusts the transparency of each image to control the contribution ratio of each image, thereby smoothly transitioning the stitched area.

[0135] The formula for the weighted fusion algorithm is shown below:

[0136]

[0137] Where I(x,y) is the gray value of the merged image, I1(x,y) is the gray value of the left image, S1 is the region of I1(x,y) after removing the overlapping area, I2(x,y) is the gray value of the right image, S2 is the region of I2(x,y) after removing the overlapping area, the overlapping area after stitching is S, w1 and w2 are the weights of the overlapping area, and w1+w2=1, w1∈(0,1). The weights are calculated using the fade-in / fade-out method, and the weight calculation formula is as follows:

[0138]

[0139] Where D is the length of the overlapping region in the two images, and d is the distance between the point to be merged in the overlapping region and the left boundary.

[0140] like Figure 9 The fire situation intelligent analysis module shown is constructed to realize fire detection.

[0141] (I) Panoramic Fire Identification Algorithm:

[0142] (1) Forest fire dataset construction: Collect samples from different scenarios: covering sunny / rainy / foggy days, daytime / nighttime, coniferous / broadleaf forest / shrubland scenarios, including sample data of flames and smoke at different scales;

[0143] (2) Data augmentation: The sample dataset is expanded by using strategies such as random rotation, scaling, flipping, brightness adjustment, and adding smoke blur.

[0144] (3) Labeling: The fire point area was labeled using the LabelImg tool, and the label format was VOC.

[0145] (4) Model training and optimization: The basic model is YOLOv8-Tiny, and the initial weights are the pre-trained weights from the COCO dataset; Forest scene adaptation: Modify the feature fusion module in the neck of the model, increase the cross-layer connection between shallow features and deep features, and strengthen the extraction of small fire point features.

[0146] (5) Real-time identification and reasoning: When a fire point is identified, an early warning is triggered and can be pushed to the visualization terminal via the cloud.

[0147] (II) Precise location of the fire source

[0148] (1) Coordinate mapping relationship construction: based on camera GPS coordinates (X cam ,Y cam Z cam Using the intrinsic parameter K, establish the pixel coordinates (u,v) and geographic coordinates (X) of the panoramic image. gco ,Y gco Mapping model of )

[0149]

[0150] Where s is the depth factor, estimated based on forest terrain DEM data, and R is the camera pose matrix (c x ,c y ) represents the coordinates of the camera's principal point, and K represents the camera's intrinsic parameter matrix;

[0151] Ground landmark calibration: Select a ground landmark from the panoramic image, such as a fire tower, signboard, or building with known geographic coordinates, and calculate the correction coefficients δX and δY of the mapping model; then perform positioning modifications.

[0152] X gco,corr =X gco +δX

[0153] Y gco,corr =Y gco +δY

[0154] Positioning accuracy is ensured through positioning correction operations.

[0155] The embodiments of the present invention disclosed above are merely illustrative of the invention. These embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention.

Claims

1. A method for identifying fireworks targets based on multi-channel video stitching, characterized in that, Includes the following steps: S1. Preprocess the images acquired by the multi-camera hardware system; S2. Perform image feature detection and matching to find the correlation information between sequences. Eliminate mismatched points through image registration verification to improve robustness and generate high-precision image matching points. S3. Perform necessary scaling and rotation transformations on the images, and accurately estimate the camera focal length and three-dimensional rotation attitude of each image. S4. Employing mesh optimization technology, the alignment accuracy between images is optimized and seams are minimized at the global level. The optimized image data is seamlessly integrated through texture mapping synthesis technology to output a high-quality panoramic stitched image. S5. Construct an intelligent fire analysis module to detect fires.

2. The method for identifying fireworks targets based on multi-channel video stitching according to claim 1, characterized in that: The image preprocessing in S1 includes: S11. Distortion Correction: The Zhang Zhengyou calibration method is used to calibrate the intrinsic parameters of each camera, and the coordinates are corrected using the correction formula. S12. The multi-scale Retinex algorithm is used to decompose the image into illuminance and reflectance components. The illuminance component reflects global illumination changes, while the reflectance component reflects the details of vegetation, smoke, and fire. Logarithmic transformation and Gaussian filtering are performed on the illuminance component to suppress strong light areas and enhance details in shadow areas. The image is reconstructed by multiplying the reflectance component by the adjusted illuminance component to ensure that the brightness deviation of images captured by different cameras is ≤8%. S13. For salt and pepper noise in forest scenes, median filtering is used; for blurred Gaussian noise caused by fog, wavelet transform is used for noise reduction; the effect is verified by peak signal-to-noise ratio (PSNR) after noise reduction.

3. The method for identifying fireworks targets based on multi-channel video stitching according to claim 2, characterized in that: The corrected formula is: Where: k1, k2, k3 are radial distortion coefficients; p1 and p2 are tangential distortion coefficients; (x dist ,y dist () represents the coordinates of the distorted pixel; (x undist ,y undist () represents the ideal coordinates; 4. The method for identifying fireworks targets based on multi-channel video stitching according to claim 1, characterized in that: S2 includes the following steps: S21. SURF feature point detection generates high-quality feature points. The homography matrix optimizes the matching through global constraints to ensure accurate fitting between the virtual object and the physical plane. The quality and accuracy of the matching points are improved by utilizing the relative pose information estimated in the initial stage. A Hessian matrix is ​​constructed. S22. Guiding feature point matching based on the initial relative pose between images; S23. Use the Random Sample Consensus Algorithm (RANSAC) to further filter out erroneous feature matching points.

5. The method for identifying fireworks targets based on multi-channel video stitching according to claim 4, characterized in that: The Hessian matrix is: c(x,y,σ)=D xx D yy -(0.9D xy ) 2 Where: D xx D xy and D yy This is an approximate convolution value obtained using a box filter; The value of the determinant of c(x,y,σ).

6. The method for identifying fireworks targets based on multi-channel video stitching according to claim 5, characterized in that: S3 includes: S31. Estimate the fundamental matrix F and the essential matrix E; S32, Camera focal length f estimation; S33. Minimize reprojection error through bundle adjustment optimization; S34. Output data that meets the reprojection error requirements.

7. The method for identifying fireworks targets based on multi-channel video stitching according to claim 6, characterized in that: S4 includes; S41. Using a cylindrical projection algorithm, the image is projected from a two-dimensional plane to a cylindrical coordinate system. The focal length 3D transformation simulates the transformation of the camera lens. That is, by adjusting the focal length, the viewing distance and depth information in the image are controlled, so as to maintain a consistent perspective effect and reduce horizontal distortion when images from different perspectives are stitched together. S42. Divide the image into a grid, calculate the distance weight between the center point of each grid and the interior point, reconstruct the DLT matrix and solve for local homography; S43. The local homography matrix in each small square is Gaussian weighted using the Moving DLT algorithm, and the image stitching is completed by transforming the relationship. S44. Using transparency-weighted fusion: In overlapping areas, a transparency-based weighting method controls the contribution ratio of each image by dynamically adjusting the transparency of the images, thereby smoothly transitioning the stitching areas. S45, outputs seamless panoramic images.

8. The method for identifying fireworks targets based on multi-channel video stitching according to claim 7, characterized in that: The weighted fusion algorithm formula in S44 is as follows: Where: I(x,y) is the gray value of the fused image, I1(x,y) is the gray value of the left image, S1 is the region of I1(x,y) after removing the overlapping area on the image, I2(x,y) is the gray value of the right image, S2 is the region of I2(x,y) after removing the overlapping area on the image, the overlapping area after stitching is S, w1 and w2 are the weights of the overlapping area, and w1+w2=1, w1∈(0,1).

9. The method for identifying fireworks targets based on multi-channel video stitching according to claim 1, characterized in that: The construction of the intelligent fire analysis module in S5 includes the following steps: S51. Construct a forest fire dataset. Expand the sample dataset by random rotation, scaling, flipping, brightness adjustment and adding smoke blurring strategy. Use LabelImg tool to label fire point areas. The label format is VOC format. S52. Model Training and Optimization: The basic model is YOLOv8-Tiny, and the initial weights are the pre-trained weights from the COCO dataset; Forest Scene Adaptation: Modify the feature fusion module in the neck of the model to increase cross-layer connections between shallow and deep features and enhance the extraction of small fire point features. S53. Real-time identification and reasoning: When a fire point is identified and an alarm is triggered, it can be pushed to the visualization terminal via the cloud; S54. Coordinate Mapping Relationship Construction: Based on Camera GPS Coordinates (X... cam ,Y cam Z cam Using the intrinsic parameter K, establish the pixel coordinates (u,v) and geographic coordinates (X) of the panoramic image. gco ,Y gco Mapping model of ) S55. Ground marker calibration: Select one type of ground marker from the panoramic image.

10. A method for identifying fireworks targets based on multi-channel video stitching according to claim 9, characterized in that: The mapping model is as follows: Where: s is the depth factor, estimated based on forest terrain DEM data, R is the camera pose matrix, (c x ,c y ) represents the coordinates of the camera's principal point, and K represents the camera's intrinsic parameter matrix.