Pedestrian danger level detection method and device based on image fusion
Through the improvement of the multi-scale fusion algorithm of the Laplace pyramid and the YOLOv5 model, combined with the model pruning technology, the accuracy of detection of pedestrian hazards in low visible light conditions is solved, the accurate identification and positioning of pedestrians is achieved, and the intelligent capabilities of optoelectronic equipment are improved.
Patent Information
- Application Number
- CN202510461094.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
AI Technical Summary
Existing optoelectronic equipment cannot effectively fuse infrared and low light images for pedestrian hazard detection under low visible light conditions, and cannot accurately identify and locate dangerous targets.
The multi-scale fusion algorithm of Laplace pyramid is used to fuse the low light and infrared images, and the SPDConv module is introduced based on the YOLOv5 model. Combined with model pruning technology, a pedestrian hazard detection model is built and deployed to the AI image processing board.
In the low-visible light complex urban combat scenario, the precise detection and positioning of pedestrian danger levels has been achieved, which has improved detection efficiency and accuracy and enhanced the intelligence level of optoelectronic equipment.
Smart Images

Figure CN120299088A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of image detection, and in particular, to a method and device for detecting the danger level of pedestrians based on image fusion. Background Art
[0002] High-precision detection of the danger level of pedestrians can obtain more accurate target information, thereby improving the accuracy of tactical decisions. In addition, it can improve the strike accuracy and effect of weapon systems, avoid ineffective attacks and defenses, and thus improve the efficiency of war. However, there are currently not many works on detecting the danger level of pedestrians. Most tasks focus on general target detection tasks. The difference between the two is that target detection focuses on more general object categories, while the detection of the danger level of pedestrians can be regarded as a subtask of target detection, and its goal is to identify the danger level of pedestrians and locate pedestrians with a higher danger level.
[0003] In addition, existing optoelectronic devices generally use detection images of a single band (visible light or infrared) to detect the danger level of pedestrians, and do not have the ability to accurately detect the danger level of pedestrians in the case of fusing low visible light and infrared images. At the same time, they also cannot perform the functions of locating and reporting danger for targets with a higher danger level. Summary of the Invention
[0004] Embodiments of the present application provide a method and device for detecting the danger level of pedestrians based on image fusion, so as to solve the problems of danger level detection and locking target reporting in a fused light scene.
[0005] On the one hand, the present application provides a method for detecting the danger level of pedestrians based on image fusion, and the method includes:
[0006] Obtain low-light images and infrared image datasets of targets with different danger levels, and construct augmented samples;
[0007] Use the Laplacian pyramid multiscale fusion algorithm to fuse the low-light image and the infrared image, and set labels for the fused image to divide the training set, validation set, and test set;
[0008] Construct a pedestrian danger level detection model, perform iterative training according to the training set and the validation set, and perform pruning operations on the model network structure after training is completed;
[0009] Convert the weights of the pruned pedestrian danger level detection model, perform functional tests based on the test set, and deploy it to the AI image processing board.
[0010] Specifically, the use of the Laplacian pyramid multiscale fusion algorithm to fuse the low-light image and the infrared image includes:
[0011] Obtain the original images of each type respectively, perform Gaussian filtering and downsampling operations on the original images to obtain multi-scale intermediate images; different-scale intermediate images correspond to different image resolutions.
[0012] Interpolate and magnify the intermediate image G at the target level i and perform differential processing in combination with the intermediate image G at the previous level i-1 to obtain the Laplacian interpolation image LP at the target scale i ;
[0013] Fuse the Laplacian interpolation images of low-light and infrared light at the same level to obtain a fused image; and perform step-by-step difference magnification processing on the fused images of different scales to restore and superimpose the fusion or obtain a fused reconstruction image.
[0014] Specifically, adopt three-level Gaussian filtering processing. The original image G0 is filtered by Gaussian and downsampled to obtain the first-level intermediate image G1; after the first-level intermediate image is interpolated and magnified, the first-level intermediate interpolation image G 1_E is obtained; the first-level intermediate image G1 is filtered by Gaussian and downsampled to obtain the second-level intermediate image G2, and after interpolation and magnification, the second-level intermediate interpolation image G 2_E .
[0015] Specifically, in the differential processing process, the intermediate image G at the previous level i-1 is subjected to differential operation with the intermediate interpolation image G at the target level i_E to output the Laplacian interpolation image LP at the target scale i , and the formula is as follows:
[0016] LP i =G i-1 -G i_E .
[0017] Specifically, adopt three-level image fusion, fuse the Laplacian interpolation images of low-light and infrared light at the same scale to obtain the fused image at the corresponding scale; the i-th level Laplacian interpolation image of low-light and the i-th level Laplacian interpolation image of infrared light are fused to obtain the i-th level fused image.
[0018] Specifically, the step-by-step difference magnification processing of the fused images of different scales to restore and superimpose the fusion or obtain a fused reconstruction image includes:
[0019] Perform difference magnification and downsampling on the third-level fused image to form the third-level intermediate fused image, and superimpose and fuse the third-level intermediate fused image with the second-level fused image to obtain the second-level intermediate fused image;
[0020] The second-level intermediate fusion image is subjected to difference magnification and downsampling to form the first-level intermediate fusion image, and the first-level intermediate fusion image is superimposed and fused with the first-level fusion image to form a fusion reconstruction image.
[0021] Specifically, the pruning operation on the model network structure after training includes:
[0022] Randomly search and generate a large number of candidate pruning schemes that meet the target constraints;
[0023] Quickly update the mean and variance parameters of the calibration normalization layer, and test the accuracy of all subnetworks after pruning calibration on the validation set;
[0024] Select the part with the highest accuracy of the subnetworks after calibration, determine the corresponding pruning scheme as the target pruning strategy, and finely adjust the accuracy of the pruned model to restore it to the target accuracy.
[0025] Specifically, the weight conversion of the pruned pedestrian danger level detection model includes:
[0026] First, convert the pt weight after the above pruning to the ONNX format, configure the model on the RKNN-Toolkit2 platform and import the model, then build the RKNN model through the rknn.build() interface, and finally export the RKNN model into an RKNN format file through the rknn.export_rknn() interface for subsequent deployment.
[0027] On the other hand, the present application provides a pedestrian danger level detection device based on image fusion. The device is used for deploying a pedestrian danger level detection model based on image fusion, and the device includes:
[0028] A control box and a photoelectric ball assembly connected to the control box through a cable;
[0029] The photoelectric ball assembly includes a detection component and a two-dimensional turntable. The detection component adopts an integrated design form of an optoelectronic payload, integrating a visible light continuous zoom camera, an uncooled infrared thermal imager, a low-light camera, and a laser rangefinder; the control box includes a housing, a power main control board, a servo control board, and an AI image processing board. The servo control board controls the two-dimensional turntable to perform pitch angle and attitude adjustment through a cable, and the AI image processing board obtains image data through a cable for automatic recognition and target positioning.
[0030] Specifically, the two-dimensional turntable includes an azimuth mechanism and a U-shaped pitch mechanism. The pitch mechanism is installed on the azimuth mechanism, and the spherical housing is installed on two rotating shafts in the U-shaped groove; the azimuth mechanism rotates azimuthally according to instructions, and the U-shaped pitch mechanism controls the spherical housing to perform pitch angle adjustment according to instructions.
[0031] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include: The method and optoelectronic device provided by this solution can accurately detect the danger level of pedestrians while ensuring the inference speed in low visible light and complex urban combat scenarios, and can be used for panoramic recognition and positioning of dangerous targets on the battlefield. The importance of this optoelectronic device lies in its ability to provide more efficient, accurate and real-time information acquisition and processing services by integrating advanced detection and analysis technologies, thereby helping soldiers better complete tasks in complex environments. To improve the efficiency and accuracy of information acquisition of optoelectronic detection devices, in this application, infrared images and low-light images obtained in low visible light and complex urban combat scenarios are calibrated and fused, and then an improved yolov5 model is used to complete the classification and positioning of different dangerous pedestrians. In addition, in order to balance the model detection efficiency, the improved model is pruned to achieve fast and accurate recognition and positioning of pedestrian danger. Finally, the algorithm is integrated into a new optoelectronic device to further improve the intelligent level of the optoelectronic device. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a flowchart of the method for detecting and deploying the danger level of pedestrians based on image fusion provided by the embodiments of the present application;
[0033] Figure 2 is a schematic diagram of the algorithm for extracting Laplacian images in the embodiments of the present application;
[0034] Figure 3 is a schematic diagram of the algorithm for generating a fused reconstruction image based on the extracted Laplacian image in the embodiments of the present application;
[0035] Figure 4 is a schematic diagram of the structure of the device for detecting the danger level of pedestrians based on image fusion provided by the present application;
[0036] Figure 5 is Figure 4 the schematic diagram of the structure of the optoelectronic sphere device in
[0037] Figure 6 is Figure 4 the schematic diagram of the structure of the two-dimensional turntable in DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0039] Figure 1 is a flowchart of the method for detecting and deploying the danger level of pedestrians based on image fusion provided by the embodiments of the present application, including the following steps:
[0040] S1. Obtain the low-light image and infrared image datasets of targets with different danger levels, and construct augmented samples;
[0041] Infrared thermal imaging systems have a large dynamic range, high signal-to-noise ratio, and high temperature resolution. They have excellent target recognition characteristics and the ability to penetrate fog, haze, and conventional smoke screens. It is easy for observers to detect clues of the target, but the image grayscale is determined by the temperature difference between the background and the target and cannot reflect the real scene. Compared with infrared images, low-light images have richer scene details and stronger expression ability for local textures and details. Therefore, infrared thermal images and low-light images have good complementarity, and the fused image of the two has better performance in extracting image features.
[0042] In this step, data augmentation can be achieved by using data enhancement methods such as random flipping, random cropping, and random brightness to expand the samples.
[0043] Of course, image calibration is also required before obtaining the images. To ensure the quality and accuracy of detection, when designing the optical system, it is necessary to ensure that the fields of view and magnifications of the visible light objective lens and the infrared objective lens are matched, so as to achieve strict image spatial domain matching. At the same time, the distortions of the two lenses need to be matched with each other. Therefore, when designing the visible light objective lens and the infrared objective lens, the distortion is strictly controlled so that the distortion errors of the two lenses are within one pixel.
[0044] Due to the alignment errors of the sight, the requirement for the consistency of image spatial domain matching is relatively high. Therefore, during the design process of the sight, not only optical interval adjustment mechanisms, elevation and azimuth adjustment mechanisms for double optical axes are designed in the physical space, but also electronic stepless zoom technology, image displacement technology, etc. are configured to ensure that the sizes and positions of the infrared images and visible light images are completely consistent, so as to meet the requirements of spatial domain matching.
[0045] S2. Use the Laplacian pyramid multi-scale fusion algorithm to fuse the low-light image and the infrared image, and set labels for the fused image to divide the training set, validation set, and test set;
[0046] The Laplacian pyramid algorithm is developed based on the Gaussian pyramid algorithm. Its main function is to separate the high-frequency details and low-frequency background of the image, and different fusion rules are selected for different frequency layers of the image during the fusion process to achieve the purpose of image detail enhancement.
[0047] When performing image fusion, the Laplacian pyramid multi-scale fusion algorithm is adopted. The advantages of this method are: different fusion strategies can be adopted for the high-frequency part and the low-frequency part of the image, which can better highlight the key points. In real-time systems, this algorithm can also adopt pipeline processing without causing an extension of the delay. In terms of color processing, in addition to using the traditional global mean-variance transfer method, local color transfer methods, background-target separate transfer methods, multi-feature domain color space transfer, etc. are tried to improve the color discrimination and color richness between the key target and the background.
[0048] S3. Construct a pedestrian danger level detection model, perform iterative training based on the training set and the validation set, and prune the model network structure after the training is completed;
[0049] Because in the low visible light and complex urban combat scenarios, it may be faced with the situation that the electro-optical detection device is far away from pedestrians, resulting in a small image of the target presented in the field of view of the device. At the same time, in order to enhance the feature expression ability after the fusion of low-light images and infrared images, the SPDConv module is introduced based on the YOLOv5 model in this application. The result verifies that the addition of this module has a positive help for improving the model performance. It can not only improve the detection ability of the model for small samples, but also improve the detection performance of the model in different scenarios.
[0050] In order to capture the real-time situation on the battlefield faster, it is necessary to improve the detection speed of the model as much as possible. In this application, pruning operation is performed on the improved YOLOv5 model to reduce the number of model parameters on the basis of ensuring the detection performance of the model, so as to improve the detection efficiency of the model for the target.
[0051] S4. Convert the weights of the pruned pedestrian danger level detection model, perform functional testing based on the test set and deploy it to the AI image processing board.
[0052] For the pruned model, the conversion of the model weights is then carried out. First, convert the pt weights to the onnx format, and then convert from onnx to the rknn format that can be run by the NPU. At the same time, develop a program for target detection and danger level monitoring on the development board for identification and positioning and reporting of targets with a higher danger level. Finally, assemble and connect the components of the development board and the electro-optical detection device to form an electro-optical equipment that can identify the danger level of pedestrians in urban combat scenarios, which has manual sector scanning and automatic sector scanning functions, various detection and positioning and status reporting functions, as well as day and night detection and ranging functions.
[0053] Figure 2 It is a schematic diagram of the algorithm for extracting the Laplacian image in the embodiment of this application. The process of fusing low-light images and infrared images using the Laplacian pyramid multi-scale fusion algorithm can be summarized as follows:
[0054] S21. Obtain the original images of each type respectively, perform Gaussian filtering processing and downsampling operation on the original images to obtain multi-scale intermediate images;
[0055] Each type mentioned here refers to two types, namely low-light images and infrared images. This application is introduced with a three-layer Laplacian algorithm. After the original image G0 is subjected to Gaussian filtering and downsampling processing, the first-level intermediate image G1 is obtained, and the first-level intermediate image G1 continues to be subjected to Gaussian filtering and downsampling to obtain the second-level intermediate image G2. Different-scale intermediate images correspond to different image resolutions.
[0056] S22. Interpolate and magnify the target-level intermediate image G i and perform differential processing in combination with the upper-level intermediate image G i-1 to obtain the Laplacian interpolation image LP at the target scale i ;
[0057] Gaussian filtering is an image processing technique mainly used to remove noise and details in images, making the images smoother. However, a simple magnification operation will cause jagged edges and blurring effects in the images. Therefore, interpolating and magnifying after Gaussian filtering can effectively reduce these side effects and make the magnified images more natural and clear. Interpolating and magnifying is the process of estimating unknown pixel values based on known pixel values. Common interpolation methods include nearest-neighbor interpolation, bilinear interpolation, and bicubic interpolation, etc. These methods calculate the values of new pixels by considering the values of surrounding pixels, thereby reducing jagged edges and blurring effects when magnifying images. In the embodiments of the present application, methods such as nearest-neighbor interpolation, bilinear interpolation, and bicubic interpolation can be used to complete this.
[0058] The first-level intermediate image G1 is interpolated and magnified to obtain the first-level intermediate image G 1_E , and the second-level intermediate image G2 is interpolated and magnified to obtain the second-level intermediate image G 2_E . The first-level intermediate image G 1_E is combined with the original image G0 for differential processing to obtain the first-level Laplacian interpolation image LP1; the second-level intermediate image G 2_E is combined with the first-level intermediate image G1 for differential processing to obtain the second-level Laplacian interpolation image LP2. Since there are a total of three levels of structures, the second-level intermediate image G2 directly outputs the third-level Laplacian interpolation image LP3.
[0059] During the differential processing, the upper-level intermediate image G i-1 is subjected to differential operation with the target-level intermediate interpolation image G i_E to output the Laplacian interpolation image LP at the target scale i , and the formula is expressed as follows:
[0060] LP i = G i-1 - G i_E .
[0061] S23. Fuse the Laplacian interpolation images of the same-level low-light and infrared light to obtain a fused image; and perform step-by-step difference magnification processing on the fused images of different scales to restore, superimpose and fuse or obtain a fused reconstructed image.
[0062] Figure 3It is a schematic diagram of an algorithm for generating a fused reconstructed image based on the extracted Laplacian image in an embodiment of the present application. Corresponding to Figure 2 logically, the fusion is also a three-level image fusion. The low-light Laplacian interpolation image and the infrared Laplacian interpolation image of the same scale are fused to obtain the fusion image of the corresponding scale. The i-th level low-light Laplacian interpolation image IL_LPi and the i-th level infrared Laplacian interpolation image IR_LPi are fused to obtain the i-th level fusion image RHi.
[0063] Specifically, the third-level fusion image RH3 is interpolated and enlarged and downsampled to form the third-level intermediate fusion image LP3_E. The third-level intermediate fusion image LP3_E is superimposed and fused with the second-level fusion image RH2 to obtain the second-level intermediate fusion image LP2_E. Further, the second-level intermediate fusion image LP2_E is interpolated and enlarged and downsampled to form the first-level intermediate fusion image LP1_E. The first-level intermediate fusion image LP1_E is superimposed and fused with the first-level fusion image RH1 to form the fused reconstructed image.
[0064] In the low visible light and complex urban combat scenario, dangerous elements may hide in positions far from the optoelectronic detection device. The result presented in the device is that the pixel area occupied by the target in the image is very small. To enable the optoelectronic detection device to have better long-distance capture performance, it is necessary to improve the small target detection ability of the model. In addition, the image after fusing the low-light image and the infrared image may have a noisy background. In this solution, the SPD Conv module is added to the YOLOv5 model to improve the model's perception ability of small objects, the ability to extract target features in different scenarios, and the adaptability of target detection.
[0065] In the traditional convolutional architecture network, as the network depth increases, the stride convolution and pooling layers will gradually reduce the spatial resolution of the image, resulting in the loss of detailed information of small objects, making it difficult for the network to accurately identify these small objects. The SPD Conv module reduces each spatial dimension of the input feature map to the channel dimension while retaining the information within the channel, compressing and retaining the detailed information scattered in a larger spatial dimension in deeper channels. In this process, the spatial dimension will decrease, while the size of the channel dimension will increase. This can reduce the size of the spatial dimension without losing information and retain the information within the channel, helping the model to more effectively extract key features, thereby improving the detection performance for low-resolution images and small objects.
[0066] There are a large number of redundant parameters in deep networks, and the activation values of many neurons approach 0. Inactivating these neurons or directly removing them can make the model exhibit similar expressive capabilities. Model pruning can effectively alleviate the problem of over-parameterization of the model. By removing the "unimportant" weights in the model, the model reduces the number of parameters and the amount of computation, while trying to ensure that the accuracy of the model is not affected. This can reduce the storage, bandwidth, and computational requirements of the model for hardware. Considering that optoelectronic devices generally have limited hardware resources and the algorithm needs to maintain a fast inference speed, the model is pruned to accelerate the model inference and implementation.
[0067] Model pruning includes unstructured pruning and structured pruning. The former has a higher model compression rate and accuracy. However, due to the "irregularity" of its computational characteristics, it requires specific hardware support to achieve the acceleration effect, so the acceleration effect on general-purpose hardware is not good. While structured pruning sacrifices the model compression rate or accuracy, it makes the weight matrix more regular and structured, so the acceleration effect on general-purpose hardware is good and it is more conducive to hardware acceleration. Based on this, this application adopts search-based channel pruning, a type of structured pruning, to reduce the number of parameters and the amount of computation of the model while minimizing the impact on model performance, so as to meet the hardware limitations. The pruning process can be summarized as follows:
[0068] S31. Randomly search to generate a large number of candidate pruning schemes that meet the target constraints
[0069] S32. Update the mean and variance parameters of the calibration normalization layer, and test the accuracy of all subnets after pruning calibration on the validation set;
[0070] S33. Select the part with the highest accuracy of the subnets after calibration, determine the corresponding pruning scheme as the target pruning strategy, and fine-tune the accuracy of the pruned model to restore it to the target accuracy.
[0071] In a feasible implementation, EagleEye uses adaptive batch normalization technology for pruning. Its overall workflow is as follows: First, a large number of pruning schemes are generated using a random strategy; then, for different pruning strategies, the parameters of their BN layers are corrected; then, for different pruning strategies, the accuracy of their pruning strategies is measured, and the optimal pruning strategy is selected; finally, the accuracy of the optimal pruning strategy is fine-tuned for restoration. EagleEye evaluates the potential of candidate subnets based on adaptive BN, which is more accurate than other evaluation methods. It can not only accelerate the pruning efficiency but also ensure the model accuracy.
[0072] The AI image board adopted in this solution is the Rockchip RK3588 image processing board equipped with an NPU. The NPU is very suitable for embedded neural networks and edge computing. However, to use the NPU unit of RK3588 for inference acceleration, the model needs to be converted into an RKNN format model first, otherwise it cannot be used. Therefore, the pt weights after pruning in the previous step need to be converted into the ONNX format first and then further into the RKNN format.
[0073] ONNX is an open ecosystem that supports developers in easily converting models between different deep learning frameworks. Therefore, ONNX can be used as an intermediate representation format to describe the architecture and its parameters of the model, enhancing the portability and flexibility of the model. The pt format is a format unique to the PyTorch framework and can only be used in PyTorch, while the ONNX weight format is a cross-platform model format that can run in a variety of hardware and software environments, and the file will contain more metadata and descriptive information. When converting the pt weights, the torch.onnx.export() function is used to export the model. The key to the conversion process is to ensure that the operations and layers of the model have corresponding representations in the ONNX format to maintain the functionality and performance of the model. After the conversion is completed, it is necessary to verify that the converted model is functionally the same as the original model to ensure the same inference results.
[0074] The RKNN model format is a model format designed specifically for the RKNN inference framework, which can effectively improve the inference efficiency. In the stage of converting ONNX to RKNN, the deep learning model will be converted into the RKNN format for efficient inference on the RKNPU platform. This process generally includes 5 stages: first, obtain the original model; then make necessary configurations in RKNN-Toolkit2, such as normalization parameters, quantization parameters, and target platforms, etc.; then use the appropriate loading interface to import the model into RKNN-Toolkit2 and select the correct loading interface according to the model framework; then build the RKNN model through the rknn.build() interface, and quantization can be selected to improve the inference performance of the model on the hardware; finally, export the RKNN model into a file (.rknn format) through the rknn.export_rknn() interface for subsequent deployment.
[0075] After obtaining the model that can accelerate inference on the development board, a program needs to be developed to identify and locate the danger level of pedestrians. When the danger level of the detected target is greater than the set limit, it needs to be reported to assist combat personnel in judging the safety of the current scene.
[0076] Based on the above detection method, this application also provides a corresponding pedestrian danger level detection device based on image fusion, such as Figure 4As shown in the figure, the device includes:
[0077] A control box and an optoelectronic sphere assembly connected to the control box through a cable;
[0078] The optoelectronic sphere assembly includes a detection component and a two-dimensional turntable. The detection component adopts an optoelectronic payload integrated design form, integrating a visible light continuous zoom camera, an uncooled infrared thermal imager, a low-light-level camera, and a laser rangefinder.
[0079] Figure 5 is Figure 4 a schematic structural diagram of the optoelectronic sphere device. As can be seen from Figure 4 it, various cameras and display devices are located in the cavity between the front cover and the rear cover of the spherical shell. A laser window and various imaging windows are provided on the shell. The visible light continuous zoom camera, the uncooled infrared thermal imager, the low-light-level camera, and the laser rangefinder are integrated and fixed on the optical bench in the optical window. After the components are individually adjusted, they are installed in the pitch frame, which can effectively isolate and reduce the influence of the turntable environmental stress on the performance of the optical payload.
[0080] The control box includes parts such as a shell, a power main control board, a servo control board, and an AI image processing board. The servo control board is the hardware platform for the operation of the servo software. Its main functions include implementing a digital control loop, reading information such as angle data, angular rate data, and target deviation amount, calculating the comprehensive control amount, and driving the actuator to operate. The two-dimensional turntable is controlled through a cable for pitch angle and attitude adjustment. The AI image processing board obtains image data through a cable for automatic recognition and target positioning.
[0081] Figure 6 is a schematic structural diagram of the two-dimensional turntable. The two-dimensional turntable adopts a two-axis and two-frame structure form, including a azimuth mechanism and a U-shaped pitch mechanism. It uses a two-axis gyro inertial element to sense the carrier deflection and realizes the stable control function of the sight line through the turntable servo control system. The azimuth mechanism consists of an azimuth shell, an azimuth rotating shaft, an azimuth bearing, an azimuth DC torque motor, an azimuth encoder, a slip ring, etc. The pitch mechanism consists of a bracket, a pitch motor, a pitch encoder, a pitch rotating shaft, a pitch bearing, etc. The bracket is of U-shaped structure, and the left and right sides of the spherical shell are fixedly connected to the pitch rotating shaft to realize the pitch adjustment function of the product. The azimuth mechanism rotates azimuthally according to the instruction, and the U-shaped pitch mechanism controls the spherical shell to adjust the pitch angle according to the instruction.
[0082] In summary, the method and optoelectronic device provided by this solution can accurately detect the danger level of pedestrians while ensuring the inference speed in low visible light and complex urban combat scenarios. It can be used for panoramic recognition and positioning of dangerous targets on the battlefield. The importance of this optoelectronic device lies in its ability to provide more efficient, accurate, and real-time information acquisition and processing services by integrating advanced detection and analysis technologies, thereby helping soldiers better complete tasks in complex environments. To improve the efficiency and accuracy of information acquisition of optoelectronic detection devices, in this application, infrared images and low-light images obtained in low visible light and complex urban combat scenarios are calibrated and fused, and then an improved yolov5 model is used to complete the classification and positioning of different dangerous pedestrians. In addition, to balance the model detection efficiency, the improved model is pruned to achieve fast and accurate recognition and positioning of pedestrian danger. Finally, the algorithm is integrated into a new optoelectronic device to further improve the intelligence level of the optoelectronic device.
[0083] This specific embodiment is only an interpretation of the present invention and does not limit the present invention. After reading this specification, those skilled in the art can make modifications to this embodiment without creative contributions as needed, but as long as they are within the scope of the claims of the present invention, they are protected by the patent law.
Claims
1. A pedestrian danger level detection method based on image fusion, characterized in that, The method includes: Obtaining the low-light image and infrared image datasets of targets with different levels of danger, and constructing augmented samples; Using the Laplacian pyramid multi-scale fusion algorithm to fuse the low-light image and the infrared image, setting labels for the fused image, and dividing it into a training set, a validation set, and a test set; Constructing a pedestrian danger level detection model, performing iterative training based on the training set and the validation set, and pruning the model network structure after the training is completed; Converting the weights of the pruned pedestrian danger level detection model, performing functional testing based on the test set, and deploying it to the AI image processing board.
2. The method according to claim 1, wherein The use of the Laplacian pyramid multi-scale fusion algorithm to fuse the low-light image and the infrared image includes: Respectively obtaining the original images of each type, performing Gaussian filtering and downsampling operations on the original images to obtain multi-scale intermediate images; different-scale intermediate images correspond to different image resolutions; The target level intermediate image G i Perform interpolation and magnification processing, combined with the previous level intermediate image G i-1 Perform differential processing to obtain the Laplace interpolation image LP of the target scale i ; Fusing the Laplacian interpolation images of the same-level low-light and infrared light to obtain a fused image; and performing step-by-step difference amplification processing on the fused images of different scales to restore and superimpose the fusion or obtain a fused reconstruction image.
3. The method according to claim 2, characterized in that Using three - level Gaussian filtering processing, the original image G0 obtains the first - level intermediate image G1 after Gaussian filtering and downsampling; after the first - level intermediate image is interpolated and magnified, the first - level intermediate interpolated image G 1_E is obtained; the first - level intermediate image G1 obtains the second - level intermediate image G2 after Gaussian filtering and downsampling, and the second - level intermediate interpolated image G 2_E is obtained.
4. The method according to claim 3, wherein In the differential processing, the intermediate image G of the previous level i-1 is differentially operated with the intermediate interpolation image G of the target level i_E to output the Laplacian interpolation image LP at the target scale i , and the formula is expressed as follows: LP i = G i-1 - G i_E .
5. The method according to claim 3, wherein Adopting three-level image fusion, fusing the Laplacian interpolation image of the low-light and the Laplacian interpolation image of the infrared light at the same scale to obtain a fused image corresponding to that scale; the fused image of the i-th level Laplacian interpolation image of the low-light and the i-th level Laplacian interpolation image of the infrared light is fused to obtain the i-th level fused image.
6. The method according to claim 5, wherein The step-by-step difference amplification processing of the fused images of different scales to restore and superimpose the fusion or obtain a fused reconstruction image includes: Performing difference amplification and downsampling on the third-level fused image to form a third-level intermediate fused image, and superimposing and fusing the third-level intermediate fused image with the second-level fused image to obtain a second-level intermediate fused image; Performing difference amplification and downsampling on the second-level intermediate fused image to form a first-level intermediate fused image, and superimposing and fusing the first-level intermediate fused image with the first-level fused image to form a fused reconstruction image.
7. The method according to claim 1, characterized in that, The pruning operation of the model network structure after the training is completed includes: Randomly searching and generating a large number of candidate pruning schemes that meet the target constraints; Updating the mean and variance parameters of the calibration normalization layer, and testing the accuracy of all subnetworks after pruning calibration on the validation set; Selecting the part with the highest accuracy of the subnetworks after calibration, determining the corresponding pruning scheme as the target pruning strategy, and fine-tuning the accuracy of the pruned model to restore it to the target accuracy.
8. The method according to claim 1, wherein The conversion of the weights of the pruned pedestrian danger level detection model includes: First converting the pt weights after the previous pruning step to the ONNX format, configuring the model on the RKNN-Toolkit2 platform and importing the model, then constructing the RKNN model through the rknn.build() interface, and finally exporting the RKNN model into an RKNN format file through the rknn.export_rknn() interface for subsequent deployment.
9. A pedestrian danger level detection device based on image fusion, characterized in that, The device is used for pedestrian danger level detection based on image fusion, and the device includes: A control box and an optoelectronic ball assembly connected to the control box by a cable; The optoelectronic sphere assembly includes a detection component and a two-dimensional turntable. The detection component adopts an integrated design form of optoelectronic payload, integrating a visible light continuous zoom camera, an uncooled infrared thermal imager, a low-light-level camera, and a laser rangefinder. The control box includes a housing, a power main control board, a servo control board, and an AI image processing board. The servo control board controls the two-dimensional turntable to perform pitch angle and attitude adjustment through a cable, and the AI image processing board obtains image data through a cable for automatic recognition and target positioning.
10. The pedestrian danger level detection device based on image fusion according to claim 8, characterized in that, The two-dimensional turntable includes an azimuth mechanism and a U-shaped pitch mechanism. The pitch mechanism is installed on the azimuth mechanism, and the spherical housing is installed on two rotating shafts in the U-shaped groove. The azimuth mechanism rotates azimuthally according to instructions, and the U-shaped pitch mechanism controls the spherical housing to perform pitch angle adjustment according to instructions.