Image acquisition and processing system and method based on three-light pod
Through the image acquisition and processing system based on the three-light pod, the problem of low recognition accuracy and low integration in the existing technology of image acquisition system in complex environments is solved, real-time processing and intelligent analysis of multi-source data is realized, and the system's perception ability and scheduling efficiency are improved to adapt to the needs of multiple scenarios.
Patent Information
- Application Number
- CN202510633101.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing image acquisition systems have low recognition accuracy and poor adaptability in complex environments, and lack efficient multi-source image fusion algorithms, resulting in data redundancy, low information utilization, delays in image acquisition and processing, low system integration, inconsistent control and data links, increasing deployment complexity.
A three-optical pod-based image acquisition and processing system is designed, including a three-optical pod module, an image acquisition synchronization module, an image processing and fusion module, an edge computing and data transmission module, and a human-computer interaction and control terminal. It adopts a multi-channel synchronization control interface, an embedded processing unit and an algorithm engine, combining edge computing and protocol unified mechanisms to realize real-time processing and intelligent analysis of multi-source data.
It significantly enhances the system's comprehensive perception and real-time response capabilities for emergencies, improves the intelligence level of command and dispatch, has the ability to automatically identify events, predict risks and schedule optimization, reduces response time and resource losses, improves the visualization and interactivity of scheduling execution, supports the collaborative needs of multiple departments, and realizes the system's self-learning and continuous optimization.
Smart Images

Figure CN120499501A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) remote sensing and image processing, and in particular to an image acquisition and processing system and method based on a three-light pod. Background Art
[0002] Existing image acquisition systems often use a single sensor type (such as visible light or infrared), resulting in low recognition accuracy and poor adaptability in complex environments. With the development of triple-light pods (integrating visible light, infrared, and lidar), multi-source information image systems with fused perception capabilities have become a research hotspot. However, current triple-light pod-based systems generally suffer from the following shortcomings: a lack of efficient multi-source image fusion algorithms leads to data redundancy and low information utilization; image acquisition and processing delays hinder real-time applications; and low system integration and inconsistent control and data links increase deployment complexity. Therefore, there is an urgent need to design an image acquisition and processing system and method with a high degree of integration, strong processing capabilities, and intelligent analysis capabilities. Summary of the Invention
[0003] (1) Technical problems solved
[0004] In response to the shortcomings of existing technologies, the present invention provides an image acquisition and processing system based on a three-light pod, which solves the following problems: the existing technology lacks an efficient multi-source image fusion algorithm, resulting in data redundancy and low information utilization; image acquisition and processing have delays, making real-time applications difficult to achieve; and the system has low system integration and inconsistent control and data links, which increases deployment complexity.
[0005] (2) Technical solution
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: an image acquisition and processing system based on a three-light pod, comprising a three-light pod module, an image acquisition synchronization module, an image processing and fusion module, an edge computing and data transmission module, and a human-computer interaction and control terminal;
[0007] The three-light pod module is the core front-end acquisition component;
[0008] The image acquisition synchronization module is used to synchronously control the exposure and sampling time of the three light sensors to achieve time-space aligned data acquisition;
[0009] The image processing and fusion module is the core part of data processing and consists of an embedded processing unit and an algorithm engine;
[0010] The edge computing and data transmission module is used to improve system processing capabilities, reduce backhaul pressure, and implement front-end intelligent judgment;
[0011] The human-computer interaction and control terminal can be a PC software or a mobile tablet App.
[0012] As a further preferred embodiment of the present invention, the three-light pod module integrates the following sensors and control interfaces, including a visible light camera unit: using a low-light high-definition CMOS image sensor with a resolution of up to 1920×1080, supporting automatic exposure adjustment, electronic image stabilization and optical image stabilization; an infrared thermal imaging unit: using an uncooled microbolometer with a detection band of 8-14μm and a thermal sensitivity of less than 50mK, supporting quantitative temperature measurement and high and low temperature alarms; a lidar unit: using a rotating or MEMS scanning lidar with a ranging accuracy of ±2cm, a field of view of 360°×30°, a sampling frequency of not less than 10Hz, and outputting high-density point cloud data.
[0013] As a further preferred embodiment of the present invention, the image acquisition synchronization module is designed with a multi-channel synchronization control interface to ensure that the spatial positions and time frames of the three types of images are completely consistent, including: a hardware trigger synchronization circuit: an FPGA controller is used to generate a unified frame trigger signal, and simultaneously control the three types of sensors to start sampling; a timestamp calibration unit: through GPS time reference, the acquisition time is unified; an attitude compensation mechanism: in conjunction with the IMU measurement unit, the pod attitude changes are dynamically compensated to avoid image offset.
[0014] As a further preferred embodiment of the present invention, the image processing and fusion module includes an image preprocessing unit for visible light image preprocessing: including brightness equalization, white balance correction, and noise suppression; for infrared image processing: temperature map generation, color band pseudo-color mapping, and heat source extraction; for LiDAR point cloud preliminary screening: noise filtering, ground plane identification, and coordinate conversion;
[0015] The image registration and fusion unit includes geometric correction: using external parameter calibration and IMU attitude solution to perform spatial correction on the image; image registration algorithm: based on multi-scale SURF feature point extraction and affine transformation model; fusion algorithm: visible light and infrared image fusion: using multi-channel wavelet fusion algorithm to enhance edge and thermal features; image and point cloud fusion: projecting the LiDAR depth map into the image coordinate system through a spatial projection model to achieve three-dimensional information superposition; the final fused image retains brightness, thermal features and spatial distance information, supporting subsequent intelligent analysis;
[0016] Intelligent recognition and tracking unit, real-time target detection, heat source anomaly judgment, based on temperature threshold and area growth; spatial ranging and size estimation, fusion of depth map to achieve 3D modeling; multi-frame time series tracking, dynamic target path analysis based on Kalman filtering and IoU matching.
[0017] As a further preferred embodiment of the present invention, the hardware platform of the edge computing and data transmission module: integrates ARM Cortex-A72+GPU / NPU; software environment: builds an image processing container based on Docker, supports dynamic loading of AI models; data return: video stream compression adopts H.265 encoding; structured data is pushed in real time using the MQTT protocol; supports asynchronous uploading of local SD storage data to ensure data security.
[0018] As a further preferred embodiment of the present invention, the functions of the human-computer interaction and control terminal include: real-time multi-channel image browsing, visual display of target detection results; automatic alarm and voice broadcast;
[0019] Data query and export; parameter setting and pod control.
[0020] As a further preferred embodiment of the present invention, the specific method steps include the following:
[0021] Step S1: Start the system and initialize. After the flight platform starts, the image system self-checks; the GPS positioning and attitude sensor are initialized; the synchronization control module is activated and begins to coordinate the work of the three types of sensors;
[0022] Step S2: Image data is collected synchronously, and the control module sends a unified trigger signal; visible light images, infrared images, and lidar data are collected at the same time; each set of images is attached with position information, posture information, and a timestamp to form an image frame set;
[0023] Step S3: Image preprocessing and calibration: image enhancement and denoising are performed immediately after image input; image angle and distortion are corrected using IMU / GPS data; and preliminary image alignment preparation is completed;
[0024] Step 4: Image registration and fusion: Use a multi-scale feature matching algorithm to register the visible light and infrared images; convert the LiDAR point cloud data into a depth map and map it to image coordinates; and fuse the three-source image data in a weighted adaptive manner to obtain a fused image.
[0025] Step 5: Intelligent analysis and target recognition: The fused image is input into the AI recognition model; the model outputs the location, category, and confidence level of the recognition box; if a high-temperature target or abnormal behavior is found, the event is recorded and an alarm is sent;
[0026] Step 6: Display and transmit the results. The detection results are superimposed on the real-time image; the fused video stream and recognition results are transmitted back to the ground end; and the data is saved for later analysis and model optimization.
[0027] (3) Beneficial effects
[0028] This invention provides an image acquisition and processing system and method based on a three-light pod. It has the following beneficial effects: By accessing multi-source heterogeneous data, combined with edge computing and a unified protocol mechanism, the system significantly enhances its comprehensive perception and real-time response capabilities for emergencies. It improves the intelligence level of command and dispatch, integrating multimodal data and intelligent algorithms. The system is capable of automatic event identification, risk prediction, and dispatch optimization, eliminating the lag of manual judgment and reliance on experience, and improving decision-making quality. It significantly optimizes resource allocation efficiency. By introducing heuristic optimization algorithms and strategy simulation and deduction mechanisms, it achieves dynamic optimal dispatch of emergency resources, effectively reducing response time and resource loss. It enhances the visualization and interactivity of dispatch execution. Three-dimensional maps, video streams, and task flow visualization make the dispatch process clear and controllable. Support for multi-screen interaction and voice command issuance improves command efficiency and user experience. It enables self-learning and continuous optimization of the dispatch system. Through task feedback, model retraining, and knowledge graph construction, the system has the ability to continuously evolve, achieving a qualitative upgrade from "manual dispatch" to "intelligent decision-making." To adapt to the needs of multi-scenario and multi-department collaboration, the system supports docking with multi-department platforms such as fire protection, transportation, and medical care to improve the systematization, coordination, and intelligence of urban governance. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 This is a schematic diagram of the system principle framework of the present invention;
[0030] Figure 2 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION
[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0032] See also Figure 1-2 , an embodiment of the present invention provides a technical solution: an image acquisition and processing system based on a three-light pod, comprising a three-light pod module, an image acquisition synchronization module, an image processing and fusion module, an edge computing and data transmission module, and a human-computer interaction and control terminal;
[0033] The three-light pod module is the core front-end acquisition component;
[0034] Image acquisition synchronization module, used to synchronously control the exposure and sampling time of the three light sensors to achieve spatiotemporal aligned data acquisition;
[0035] The image processing and fusion module is the core part of data processing and consists of an embedded processing unit and an algorithm engine;
[0036] Edge computing and data transmission modules are used to improve system processing capabilities, reduce backhaul pressure, and enable front-end intelligent judgment;
[0037] Human-computer interaction and control terminal, the terminal can be PC software or mobile tablet App.
[0038] The three-light pod module integrates the following sensors and control interfaces: a visible light camera unit: using a low-light high-definition CMOS image sensor with a resolution of up to 1920×1080, supporting automatic exposure adjustment, electronic image stabilization, and optical image stabilization; an infrared thermal imaging unit: using an uncooled microbolometer with a detection band of 8-14μm and a thermal sensitivity of less than 50mK, supporting quantitative temperature measurement and high and low temperature alarms; a lidar unit: using a rotating or MEMS scanning lidar with a ranging accuracy of ±2cm and a field of view of 360°
[0039] ×30°, sampling frequency not less than 10Hz, output high-density point cloud data.
[0040] To ensure that the spatial positions and time frames of the three types of images are completely consistent, the image acquisition synchronization module is designed with a multi-channel synchronization control interface, including: a hardware trigger synchronization circuit: an FPGA controller is used to generate a unified frame trigger signal, which simultaneously controls the three types of sensors to start sampling; a timestamp calibration unit: a GPS time reference is used to achieve unified acquisition time; and an attitude compensation mechanism: in conjunction with the IMU measurement unit, dynamic compensation is performed for pod attitude changes to avoid image offset.
[0041] The image processing and fusion module includes an image preprocessing unit for visible light image preprocessing:
[0042] Including brightness balance, white balance correction, and noise suppression;
[0043] Used for infrared image processing: temperature map generation, color band pseudo-color mapping, heat source extraction; used for LiDAR point cloud initial screening: noise filtering, ground plane identification, coordinate conversion;
[0044] The image registration and fusion unit includes geometric correction: using external parameter calibration and IMU attitude solution to perform spatial correction on the image; image registration algorithm: based on multi-scale SURF feature point extraction and affine transformation model; fusion algorithm: visible light and infrared image fusion: using multi-channel wavelet fusion algorithm to enhance edge and thermal features; image and point cloud fusion: projecting the LiDAR depth map into the image coordinate system through a spatial projection model to achieve three-dimensional information superposition; the final fused image retains brightness, thermal features and spatial distance information, supporting subsequent intelligent analysis;
[0045] Intelligent recognition and tracking unit, real-time target detection, heat source anomaly judgment, based on temperature threshold and area growth; spatial ranging and size estimation, fusion of depth map to achieve 3D modeling; multi-frame time series tracking, dynamic target path analysis based on Kalman filtering and IoU matching.
[0046] The hardware platform of the edge computing and data transmission module integrates ARM Cortex-A72+GPU / NPU. The software environment uses Docker to build image processing containers and supports dynamic loading of AI models. Data return uses H.265 encoding for video stream compression. Structured data is pushed in real time using the MQTT protocol. Asynchronous upload of local SD storage data is supported to ensure data security.
[0047] The functions of the human-computer interaction and control terminal include: real-time multi-channel image browsing, visual display of target detection results; automatic alarm and voice broadcast;
[0048] Data query and export; parameter setting and pod control.
[0049] The method steps of the image acquisition and processing system based on the three-light pod include the following:
[0050] Start the system and initialize it. After the flight platform starts, the image system self-checks; the GPS positioning and attitude sensors are initialized; the synchronization control module is activated and begins to coordinate the work of the three types of sensors;
[0051] Image data is collected synchronously, and the control module sends a unified trigger signal; visible light images, infrared images, and lidar data are collected at the same time; each set of images is attached with position information, posture information, and timestamp to form an image frame set;
[0052] Image preprocessing and calibration: enhance and denoise the image immediately after input; use IMU / GPS data to correct image angle and distortion; complete preliminary image alignment preparation;
[0053] Image registration and fusion: Use a multi-scale feature matching algorithm to register visible light and infrared images; convert LiDAR point cloud data into a depth map and map it to image coordinates; fuse the three-source image data in a weighted adaptive manner to obtain a fused image;
[0054] Intelligent analysis and target recognition: Fusion image input AI recognition model; the model outputs the location, category, and confidence level of the recognition box; if a high-temperature target or abnormal behavior is found, the event is recorded and an alarm is sent;
[0055] The results are displayed and transmitted back, and the detection results are superimposed on the real-time image; the fused video stream and recognition results are transmitted back to the ground end; and the data is saved for later analysis and model optimization.
[0056] Example 1: Application in forest fire prevention intelligent patrol scenario
[0057] This embodiment provides a three-light pod image acquisition and processing system. The system is mounted on a multi-rotor drone platform, has a system configuration consistent with the above claims, and has good field test performance.
[0058] 1. System Hardware Deployment
[0059] Three-light pod module:
[0060] The pod is installed under the belly of the drone and is fixed with a shock-absorbing structure;
[0061] The pod integrates:
[0062] Visible light camera (resolution 1920×1080, support 20x optical zoom);
[0063] Infrared thermal imager (wavelength 8-14 μm, temperature sensitivity ≤ 45 mK, field of view 40°);
[0064] Rotating LiDAR module (30 lines, 360°×30° field of view, 100,000 points / second);
[0065] All sensors are integrated into the same optical axis structure, equipped with a gimbal stabilization system that supports ±90° pitch adjustment.
[0066] Image acquisition synchronization module:
[0067] The FPGA timing control board is responsible for triggering the three types of image acquisition;
[0068] The synchronization signal cycle is 5 frames per second (5Hz), ensuring low latency and high consistency;
[0069] The GPS module outputs timestamps, and the IMU sensor provides 6-degree-of-freedom attitude information to assist in data correction.
[0070] Image processing and fusion module:
[0071] Using the Jetson Xavier NX edge computing platform;
[0072] Perform image preprocessing, registration fusion and target analysis during flight;
[0073] The wavelet fusion algorithm is used to fuse the infrared thermal image and the visible light image for display, and the LiDAR point cloud depth is mapped to the image to form a three-dimensional temperature map.
[0074] Edge computing and transmission module:
[0075] The processed video stream is uploaded to the ground station in real time via the 5G module;
[0076] The image stream is compressed using H.265 encoding, and structured target identification information is pushed via the MQTT protocol;
[0077] At the same time, high-definition source images and AI recognition results are stored locally in TF cards for subsequent analysis.
[0078] Human-computer interaction and control terminal:
[0079] The ground station uses Windows system control software and Android tablet app;
[0080] The software interface supports real-time browsing of multi-channel images, heat source location marking, and target trajectory drawing;
[0081] You can set the high temperature threshold, detection target type, automatic photo taking frequency, data saving path, etc.
[0082] 2. System Usage Process
[0083] According to the method process of claim 7, the specific operations are as follows:
[0084] Step S1: System startup and initialization
[0085] Ground operators activate the drone and pod systems;
[0086] After the pod's internal self-test is complete, connect the GPS positioning and IMU modules;
[0087] The pod's attitude stabilizes and the FPGA control system begins to send a unified trigger signal.
[0088] Step S2: Synchronous image acquisition
[0089] The visible light camera collects RGB images;
[0090] Thermal imager captures scene heat map (temperature coverage range -20℃-300℃);
[0091] The laser radar simultaneously collects environmental point cloud data;
[0092] All data are packaged as image frame sets with timestamps, coordinates and pose information.
[0093] Step S3: Image preprocessing and calibration
[0094] The preprocessing module performs brightness equalization and sharpening on RGB images;
[0095] Heatmaps are pseudo-colored for easy visualization;
[0096] The point cloud completes ground filtering and orientation alignment;
[0097] The system performs spatial correction on images based on GPS / IMU data to ensure that the data is fused in a unified space.
[0098] Step S4: Image registration and fusion
[0099] Use the multi-scale SURF feature point algorithm to accurately align infrared and visible light images;
[0100] The point cloud depth map is projected onto the image plane to generate a three-dimensional heat map;
[0101] The fused image is displayed in the form of visual + thermal imaging + depth map, enhancing target recognition in forest scenes.
[0102] Step S5: Intelligent analysis and target recognition
[0103] Use the YOLOv7 model to identify targets such as "people", "smoke", and "open flames" in images;
[0104] Combined with the thermal map, abnormally high temperature areas are identified and marked as "suspected fire sources";
[0105] Simultaneously track the trajectory of moving targets and identify the dynamics of suspicious individuals;
[0106] If the temperature is found to be higher than the set threshold (such as 60°C) or the image characteristics of an open flame are detected, the alarm mechanism will be automatically triggered.
[0107] Step S6: Result display and data transmission
[0108] The ground terminal displays the fused image, superimposed with identification frames and temperature labels;
[0109] The system records the alarm location, time and screenshots;
[0110] Low latency (less than 500ms) is maintained during video transmission, and structured data is pushed to the backend system in real time.
[0111] The actual operation test of this embodiment shows that:
[0112] Multi-source image fusion significantly improves the detectability of hidden fire sources in forests;
[0113] Real-time thermal imaging + spatial depth enhances the analysis capability of complex terrain areas;
[0114] The system has low latency and high data integrity, making it suitable for a variety of emergency response scenarios;
[0115] Edge AI intelligent judgment reduces the background recognition burden and significantly improves response speed;
[0116] The entire system is portable, easy to deploy, has strong environmental adaptability, and is capable of mass promotion.
[0117] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, from all perspectives, the embodiments should be regarded as illustrative and non-restrictive. The scope of the present invention is defined by the appended claims, not the foregoing description, and it is intended that all variations that come within the meaning and range of equivalents of the claims be included within the present invention. Any reference signs in the claims should not be construed as limiting the claim to which they relate.
[0118] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
Claims
1. An image acquisition and processing system based on a three-light pod, characterized by: It includes three-light pod module, image acquisition and synchronization module, image processing and fusion module, edge computing and data transmission module, and human-computer interaction and control terminal; The three-light pod module is the core front-end acquisition component; The image acquisition synchronization module is used to synchronously control the exposure and sampling time of the three light sensors to achieve time-space aligned data acquisition; The image processing and fusion module is the core part of data processing and consists of an embedded processing unit and an algorithm engine; The edge computing and data transmission module is used to improve system processing capabilities, reduce backhaul pressure, and implement front-end intelligent judgment; The human-computer interaction and control terminal can be a PC software or a mobile tablet App.
2. The image acquisition and processing system based on the three-light pod according to claim 1 is characterized by: The three-light pod module integrates the following sensors and control interfaces, including a visible light camera unit: using a low-light high-definition CMOS image sensor with a resolution of up to 1920×1080, supporting automatic exposure adjustment, electronic image stabilization and optical image stabilization; an infrared thermal imaging unit: using an uncooled microbolometer with a detection band of 8-14μm and a thermal sensitivity of less than 50mK, supporting quantitative temperature measurement and high and low temperature alarms; a lidar unit: using a rotating or MEMS scanning lidar with a ranging accuracy of ±2cm, a field of view of 360°×30°, a sampling frequency of not less than 10Hz, and outputting high-density point cloud data.
3. The image acquisition and processing system based on the three-light pod according to claim 1 is characterized in that: To ensure that the spatial positions and time frames of the three types of images are completely consistent, the image acquisition synchronization module is designed with a multi-channel synchronization control interface, including: a hardware trigger synchronization circuit: using an FPGA controller to generate a unified frame trigger signal, simultaneously controlling the three types of sensors to start sampling; a timestamp calibration unit: using GPS time reference to achieve unified acquisition time; and an attitude compensation mechanism: in conjunction with the IMU measurement unit, dynamically compensates for pod attitude changes to avoid image offset.
4. The image acquisition and processing system based on the three-light pod according to claim 1 is characterized in that: The image processing and fusion module includes an image preprocessing unit for visible light image preprocessing: including brightness equalization, white balance correction, and noise suppression; for infrared image processing: temperature map generation, color band pseudo-color mapping, and heat source extraction; for LiDAR point cloud initial screening: noise filtering, ground plane identification, and coordinate conversion; The image registration and fusion unit includes geometric correction: using external parameter calibration and IMU attitude solution to perform spatial correction on the image; image registration algorithm: based on multi-scale SURF feature point extraction and affine transformation model; fusion algorithm: visible light and infrared image fusion: using multi-channel wavelet fusion algorithm to enhance edge and thermal features; Image and point cloud fusion: The LiDAR depth map is projected into the image coordinate system through a spatial projection model to achieve 3D information overlay. The final fused image retains brightness, thermal characteristics, and spatial distance information, providing support for subsequent intelligent analysis. Intelligent recognition and tracking unit, real-time target detection, heat source anomaly judgment, based on temperature threshold and area growth; Spatial ranging and size estimation, fusion of depth maps to achieve 3D modeling; multi-frame time series tracking, dynamic target path analysis based on Kalman filtering and IoU matching.
5. The image acquisition and processing system based on the three-light pod according to claim 1 is characterized in that: The hardware platform of the edge computing and data transmission module: integrates ARM Cortex-A72+GPU / NPU; the software environment: builds an image processing container based on Docker, supports dynamic loading of AI models; data return: video stream compression uses H.265 encoding; structured data is pushed in real time using the MQTT protocol; supports asynchronous uploading of local SD storage data to ensure data security.
6. The image acquisition and processing system based on the three-light pod according to claim 1 is characterized in that: The functions of the human-computer interaction and control terminal include: real-time multi-channel image browsing, visual display of target detection results; automatic alarm and voice broadcast; data query and export; parameter setting and pod control.
7. The method for using the image acquisition and processing system based on the three-light pod according to claim 1 is characterized by: The specific steps include the following: Step S1: Start the system and initialize. After the flight platform starts, the image system self-checks; the GPS positioning and attitude sensors are initialized; the synchronization control module is activated and begins to coordinate the work of the three types of sensors; Step S2: Image data is collected synchronously, and the control module sends a unified trigger signal; visible light images, infrared images, and lidar data are collected at the same time; each set of images is attached with position information, posture information, and a timestamp to form an image frame set; Step S3: Image preprocessing and calibration: image enhancement and denoising are performed immediately after image input; image angle and distortion are corrected using IMU / GPS data; and preliminary image alignment preparation is completed; Step 4: Image registration and fusion: Use a multi-scale feature matching algorithm to register the visible light and infrared images; convert the LiDAR point cloud data into a depth map and map it to image coordinates; and fuse the three-source image data in a weighted adaptive manner to obtain a fused image. Step 5: Intelligent analysis and target recognition: the fused image is input into the AI recognition model; the model outputs the recognition box location, category, and confidence level; If a high-temperature target or abnormal behavior is detected, the event will be recorded and an alarm will be sent; Step 6: Display and transmit the results, and superimpose the test results on the real-time image; The fused video stream and recognition results are transmitted back to the ground end; the data is saved for later analysis and model optimization.