Intrusion detection and identification method, system and medium based on infrared and visible light fusion

Through geometric correction, pixel-level weighted fusion and bilateral filtering processing in infrared and visible image fusion, combined with adaptive ELMAN neural network and background differential processing, the problem of poor image fusion quality in the prior art is solved, and high-quality image fusion and intrusion detection effects are achieved.

CN119107596BActive Publication Date: 2025-05-23SHENZHEN SED ELECTRONICS EQUIP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411136099.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2025-05-23
Estimated Expiration
2044-08-19

AI Technical Summary

Technical Problem

In the fusion process of visible light and infrared images, the prior art leads to the loss of details and color distortion of pixel fusion, edge model and high-frequency information, and the image quality is not ideal, which is manifested as a decrease in clarity and contrast, and the noise is amplified.

Method used

By acquiring the images of infrared and visible cameras at the same time, geometric correction and pixel-level weighted fusion are performed, the weighting coefficient is determined using an adaptive ELMAN neural network, and bilateral filtering and background differential processing are performed after the fusion to achieve intrusion detection and recognition.

Benefits of technology

The quality and clarity of the fusion image are improved, the thermal information of infrared images and the detailed information of visible light images are retained, the brightness, contrast and detail retention capabilities of the image are enhanced, and the complex and dynamic features in the image data are effectively processed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119107596B_ABST
    Figure CN119107596B_ABST
Patent Text Reader

Abstract

The present invention provides an intrusion detection and identification method, system and medium based on the fusion of infrared and visible light. The method includes: obtaining images detected by an infrared camera and a visible light camera at the same time, which are infrared images and visible light images respectively; geometrically correcting the infrared image and the visible light image to make the viewing angle and scale of the two consistent; performing pixel-level weighted fusion of the infrared and visible light images, determining the weighting coefficient based on an adaptive ELMAN neural network and calculating the pixel value after fusion to obtain an infrared / visible light fusion image; performing bilateral filtering on the fusion image to achieve global smoothing post-processing; performing differential processing on the globally smoothed fusion image and the background image to achieve intrusion detection and identification. The present invention can make the fused image have better performance in terms of brightness, contrast and detail retention, effectively process the complex and dynamic features in the image data, and improve the quality of the fusion image and the overall performance of the detection system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of security and monitoring technology, in particular to intrusion detection technology, and specifically to an intrusion detection and identification method, system and medium based on the fusion of infrared and visible light. Background Art

[0002] In security and monitoring technology, both infrared imaging and visible light imaging play a vital role. They each have unique advantages and applications, and can play an important role in different environments and needs. Visible light cameras can capture clear images, provide richer color and texture information, and are suitable for well-lit environments, and are suitable for tasks such as face recognition, pedestrian detection and tracking, and behavior analysis. Infrared cameras rely on the thermal radiation of target objects rather than the reflection of light, and can work in low-light and no-light environments. By capturing the thermal radiation emitted by objects, they can identify targets at night or in low visibility conditions. For example, the human body emits radiation of a specific temperature in the infrared band, and infrared cameras can detect this radiation, especially in low-light or dark environments, and are used for night security, border protection, night security monitoring of warehouses and industrial areas, and monitoring of wild animals or non-invasive monitoring in nature reserves.

[0003] In modern security systems, infrared and visible light imaging technologies are often used in combination to provide all-weather and all-day monitoring capabilities. By utilizing image data in two different bands, the detection accuracy and reliability under different environmental conditions can be improved. For example, a multimodal monitoring system: combining infrared night vision with visible light monitoring during the day to ensure that the monitoring system can operate effectively during both daytime and nighttime, combined with further algorithm analysis of visible light and infrared fusion images, for more complex detection and alarms, such as intrusion detection, abnormal behavior warnings, etc.

[0004] At present, the fusion of visible light and infrared images is achieved through image processing methods such as pixel fusion, feature fusion, and decision fusion. For example, weighted pixel processing or wavelet transform is used to directly fuse visible light and infrared images at the pixel level of the image to obtain a fused image. Among them, pixel-level weighted fusion is a commonly used method, which calculates the fused pixel value by weighted average of the pixels of the two images. However, in the existing calculation, the threshold value of the weighted calculation is fixed, which leads to the problem of detail and color distortion, edge model and high-frequency information loss in pixel fusion, resulting in unsatisfactory quality of the fused image, which is manifested as reduced clarity and contrast, amplified noise, and reduced image quality, which affects the subsequent image processing steps. Summary of the invention

[0005] In view of the defects of the prior art, according to the first aspect of the present invention, an intrusion detection and identification method based on the fusion of infrared and visible light is proposed, comprising the following steps:

[0006] Step S101, acquiring images detected by an infrared camera and a visible light camera at the same time, which are an infrared image and a visible light image respectively;

[0007] Step S102, geometrically correcting the infrared image and the visible light image so that the viewing angles and scales of the two images are consistent;

[0008] Step S103, performing pixel-level weighted fusion of the infrared image and the visible light image, determining the weighting coefficient based on the adaptive ELMAN neural network and calculating the fused pixel value according to the grayscale values ​​of the infrared image and the visible light image, to obtain an infrared / visible light fused image;

[0009] Step S104, performing bilateral filtering on the infrared / visible light fusion image to achieve global smoothing post-processing; and

[0010] Step S105: performing differential processing on the infrared / visible light fusion image after global smoothing and the background image to achieve intrusion detection and identification.

[0011] As an optional implementation, the step S101 further includes graying the acquired visible light image to obtain a grayscale image corresponding to the visible light image.

[0012] As an optional implementation, in step S102, geometric correction is performed on the infrared image and the visible light image so that the viewing angle and scale of the two images are consistent, including:

[0013] Interpolate and resample the infrared image and the visible light image to make them have the same size; and

[0014] Perform an affine transformation on the infrared image so that it is at the same viewing angle as the visible light image.

[0015] As an optional implementation, in step S103, the infrared image and the visible light image are subjected to pixel-level weighted fusion, the weighting coefficient is determined based on the adaptive ELMAN neural network, and the fused pixel value is calculated according to the grayscale value of the infrared image and the visible light image to obtain the infrared / visible light fused image, including:

[0016] Pixel-level weighted fusion is performed based on the following pixel-level weighted calculation:

[0017] IF(x,y)=k 1 *RGB(x,y)+k 2 *IR(x,y)

[0018] Where IF(x,y) represents the value of the pixel point (x,y) in the infrared / visible light fusion image, RGB(x,y) and IR(x,y) represent the values ​​of the corresponding pixel point (x,y) in the visible light image and infrared image respectively; k 1 , k 2 The weighting coefficients representing the visible light image and the infrared image respectively are set to determine the coefficient value of each infrared / visible light fusion image through an adaptive ELMAN neural network.

[0019] As an optional implementation, in step S103, the process of determining the weighting coefficient of each infrared / visible light fusion image by using the adaptive ELMAN neural network includes:

[0020] The adaptive ELMAN neural network is a dynamic recursive neural network structure, which consists of an input layer, a hidden layer, a context layer and an output layer. The input layer and the output layer each include two neuron nodes, and the number of neuron nodes in the hidden layer and the context layer is the same. The input layer inputs a signal, and the hidden layer uses a nonlinear function as an excitation function. The context layer is used to memorize the output value of the hidden layer at the previous moment, so that the output of the hidden layer is self-connected back to the input of the hidden layer after being delayed by the context layer, and the output weight coefficient is output through weighted processing in the output layer;

[0021] The grayscale mean Gray of the visible light image and infrared image at time t RGB With Gray IR As the input layer input of the adaptive ELMAN neural network, the number of neuron nodes h of the hidden layer of the adaptive ELMAN neural network is determined by the following method:

[0022]

[0023] Among them, m and n represent the number of neuron nodes in the input layer and the output layer respectively, and a represents a constant with a value range of 1 to 10;

[0024] The hidden layer outputs the brightness mean, standard deviation and coefficient of variation of the fused image at the current time t. The context layer receives the output of the hidden layer and memorizes the output value of the hidden layer at the previous time t-1, so that the output of the hidden layer is self-connected back to the input of the hidden layer after being delayed by the context layer. Finally, the two outputs after weighted processing by the two neurons in the output layer are the weighted coefficients k of the visible light image and the infrared image. 1 , k 2 ;

[0025] The brightness mean of the fused image refers to the mean of the grayscale values ​​of the fused image, reflecting the overall brightness of the image;

[0026] The standard deviation represents the degree of dispersion between the grayscale value of the pixel points of the fused image and the mean value, which is expressed as:

[0027] Standard deviation = square root of variance, where variance is the sum of the squares of the grayscale value of each pixel in the fused image and the mean of the grayscale value;

[0028] The coefficient of variation represents the contrast and clarity of the fused image, and is expressed as:

[0029] Coefficient of variation = standard deviation / mean brightness.

[0030] As an optional implementation, in step S103, the adaptive ELMAN neural network is a pre-trained network model, and the training process includes:

[0031] Initialize network parameters, including weights from input to hidden layer, from hidden layer to output layer, from hidden layer to context layer, and randomly initialize bias;

[0032] Forward propagation: From the input layer to the hidden layer, the weighted sum of the input signal and the weight is calculated, the bias is added, and then the activation function is applied; the context layer is updated, the hidden layer output of the previous time step is stored in the context layer and used as an additional input for the current time step, which helps the network remember the previous state; from the hidden layer to the output layer, the final output signal is calculated using the output of the hidden layer and including feedback from the context layer;

[0033] Evaluation loss: Use the mean squared error loss function to evaluate the difference between the network output and the true label;

[0034] Backpropagation: Compute the gradient of each weight with respect to the loss function via the chain rule, then update the weights and biases via the gradient descent optimization algorithm; and

[0035] Parameter adaptive adjustment: Use learning rate decay to adjust the learning rate and use L1 / L2 regularization to prevent overfitting.

[0036] As an optional implementation, in step S104, bilateral filtering is performed on the infrared / visible light fusion image to achieve global smoothing post-processing, including:

[0037] Traverse each pixel: Apply bilateral filtering to each pixel of the infrared / visible light fusion image, and for each pixel, calculate the weighted average based on the dynamically adjusted spatial kernel and intensity kernel in its neighborhood;

[0038] Weight calculation: Calculate filter weights, including spatial weights based on the spatial distance between pixels and a dynamically adjusted spatial kernel, and intensity weights based on the difference in pixel values ​​and a dynamically adjusted intensity kernel;

[0039] Output to generate an updated image: Update the current pixel value of each pixel position to the weighted average of its neighboring pixels, and output the overall updated image;

[0040] The dynamically adjusted spatial kernel and intensity kernel are filter kernels adjusted based on the gradient and variance statistical characteristics of the fused image, including dynamically adjusting the size of the spatial kernel according to the gradient size and adjusting the range of the intensity kernel according to the variance.

[0041] As an optional implementation, in step S105, the infrared / visible light fusion image processed after global smoothing is subjected to differential processing with the background image to achieve intrusion detection and identification, including:

[0042] Acquire background infrared images and background visible light images of the monitoring area collected in multiple time periods according to a preset cycle, and obtain an infrared / visible light fused background image after image fusion;

[0043] According to the moment of the infrared / visible light fusion image after global smoothing, the time period is determined;

[0044] Perform background difference processing on the infrared / visible light fusion image after global smoothing and the infrared / visible light fusion background image of the same time period to achieve foreground target intrusion detection;

[0045] Perform pedestrian target detection on foreground targets, identify pedestrians and issue early warnings as well as dynamic trajectory tracking.

[0046] According to the second aspect of the purpose of the present invention, a computer-readable medium storing software is also proposed, wherein the software includes instructions that can be executed by one or more computers, and when the instructions are executed by the one or more computers, the process of the intrusion detection and identification method based on the fusion of infrared and visible light as mentioned above is performed.

[0047] According to a third aspect of the present invention, a computer system is also provided, comprising:

[0048] one or more processors;

[0049] Memory, storing instructions that can be operated;

[0050] Among them, when the instruction is executed by one or more processors, the aforementioned one or more processors perform operations, and the operations include the process of executing the aforementioned intrusion detection and identification method based on the fusion of infrared and visible light.

[0051] In combination with the intrusion detection and identification method based on the fusion of infrared and visible light of the present invention, before pixel-level weighted fusion, geometric correction is performed on the visible light and infrared images to ensure the consistency of the two images in perspective and scale, reduce the registration error caused by the perspective difference, and thus improve the quality of the fused image; pixel-level weighted fusion combines the advantages of infrared and visible light images, so that the fused image retains both the thermal information of the infrared image and the detail information of the visible light image, and the weighting coefficient is determined by the adaptive ELMAN neural network, and the fusion effect can be dynamically adjusted according to the actual situation of the image; after fusion, bilateral filtering is further used to retain the edge information of the image while globally smoothing, which helps to remove noise without blurring the edges and high-frequency details, thereby improving the quality and clarity of the fused image.

[0052] Compared with the prior art, the visible light and infrared image fusion of the present invention performs pixel calculation based on the weighting system determined by the adaptive ELMAN neural network. On the one hand, through the delayed self-connection mechanism of the context layer, the adaptive ELMAN neural network can better capture the time dependence of the image data and improve the dynamic adjustment of the weighting coefficients, so that the fusion effect can better adapt to the dynamic changes in the image and fit the characteristics of the actual image, so that the fused image has better performance in brightness, contrast and detail retention, effectively processes the complex and dynamic features in the image data, and improves the quality of the fused image and the overall performance of the detection system. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a flowchart of an intrusion detection and identification method based on infrared and visible light fusion according to an embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to better understand the technical content of the present invention, specific embodiments are given and described as follows in conjunction with the accompanying drawings.

[0055] Example 1

[0056] Combination Figure 1 As shown, the intrusion detection and identification method based on infrared and visible light fusion according to an embodiment of the present invention includes the following steps:

[0057] Step S101, acquiring images detected by an infrared camera and a visible light camera at the same time, which are an infrared image and a visible light image respectively;

[0058] Step S102, geometrically correcting the infrared image and the visible light image so that the viewing angles and scales of the two images are consistent;

[0059] Step S103, performing pixel-level weighted fusion of the infrared image and the visible light image, determining the weighting coefficient based on the adaptive ELMAN neural network and calculating the fused pixel value according to the grayscale values ​​of the infrared image and the visible light image, to obtain an infrared / visible light fused image;

[0060] Step S104, performing bilateral filtering on the infrared / visible light fusion image to achieve global smoothing post-processing; and

[0061] Step S105: performing differential processing on the infrared / visible light fusion image after global smoothing and the background image to achieve intrusion detection and identification.

[0062] Wherein, in the step S101, the obtained visible light image is also converted into grayscale to obtain a grayscale image corresponding to the visible light image.

[0063] In step S102, the infrared image and the visible light image are geometrically corrected so that the viewing angles and scales of the two images are consistent, including:

[0064] Interpolate and resample the infrared image and the visible light image to make them have the same size; and

[0065] Perform an affine transformation on the infrared image so that it is at the same viewing angle as the visible light image.

[0066] As an optional embodiment, the aforementioned interpolation resampling uses a bilinear interpolation algorithm to resample the low-resolution infrared image so that it has the same size as the high-resolution visible light image, that is, the same width and height.

[0067] It should be understood that in the embodiments of the present invention, whether it is geometric correction of images or subsequent fusion, it is necessary to ensure that the infrared and visible light cameras capture images of the same monitoring area scene at the same time to ensure synchronous shooting.

[0068] As an optional embodiment, in step S103, the infrared image and the visible light image are subjected to pixel-level weighted fusion, a weighting coefficient is determined based on an adaptive ELMAN neural network, and a fused pixel value is calculated according to the grayscale values ​​of the infrared image and the visible light image to obtain an infrared / visible light fused image, including:

[0069] Pixel-level weighted fusion is performed based on the following pixel-level weighted calculation:

[0070] IF(x,y)=k 1 *RGB(x,y)+k 2 *IR(x,y)

[0071] Where IF(x,y) represents the value of the pixel point (x,y) in the infrared / visible light fusion image, RGB(x,y) and IR(x,y) represent the values ​​of the corresponding pixel point (x,y) in the visible light image and infrared image respectively; k 1 , k 2 The weighting coefficients representing the visible light image and the infrared image respectively are set to determine the coefficient value of each infrared / visible light fusion image through an adaptive ELMAN neural network.

[0072] As an optional embodiment, in step S103, the process of determining the weighting coefficient of each infrared / visible light fusion image by using the adaptive ELMAN neural network includes:

[0073] The adaptive ELMAN neural network is a dynamic recursive neural network structure, and the network parameters are adjusted through an adaptive mechanism. It consists of an input layer, a hidden layer, a context layer and an output layer. The input layer and the output layer each include two neuron nodes, and the number of neuron nodes in the hidden layer and the context layer is the same. The signal is input through the input layer, and the hidden layer uses a nonlinear function as an excitation function. The context layer is used to memorize the output value of the hidden layer at the previous moment, so that the output of the hidden layer is self-connected back to the input of the hidden layer after being delayed by the context layer, and the output weight coefficient is output through weighted processing in the output layer;

[0074] The grayscale mean Gray of the visible light image and the infrared image at time t RGB With Gray IR As the input layer input of the adaptive ELMAN neural network, the number of neuron nodes h of the hidden layer of the adaptive ELMAN neural network is determined by the following method:

[0075]

[0076] Among them, m and n represent the number of neuron nodes in the input layer and the output layer respectively, and a represents a constant with a value range of 1 to 10;

[0077] The hidden layer outputs the brightness mean, standard deviation and coefficient of variation of the fused image at the current time t. The context layer receives the output of the hidden layer and memorizes the output value of the hidden layer at the previous time t-1, so that the output of the hidden layer is self-connected back to the input of the hidden layer after being delayed by the context layer. Finally, the two outputs after weighted processing by the two neurons in the output layer are the weighted coefficients k of the visible light image and the infrared image. 1 , k 2 .

[0078] The brightness mean of the fused image refers to the mean of the grayscale values ​​of the fused image, reflecting the overall brightness of the image. A larger mean of the grayscale value indicates a greater overall brightness of the image, and vice versa.

[0079] The standard deviation indicates the degree of dispersion between the grayscale value of the pixel points of the fused image and the mean value. The larger the standard deviation, the higher the image quality. In this example, the image standard deviation is expressed as:

[0080] Standard deviation = square root of variance, where variance is the sum of the squares of the grayscale value of each pixel in the fused image and the mean of the grayscale value.

[0081] The aforementioned coefficient of variation indicates the contrast and clarity of the fused image. The larger the coefficient of variation, the better the contrast and clarity of the image. In this example, the coefficient of variation of the image is expressed as:

[0082] Coefficient of variation = standard deviation / mean brightness.

[0083] In the embodiment of the present invention, it can be calculated based on the OpenCV public library.

[0084] As an optional implementation, in step S103, the adaptive ELMAN neural network is a pre-trained network model, and the training process includes:

[0085] Initialize network parameters, including weights from input to hidden layer, from hidden layer to output layer, from hidden layer to context layer, and randomly initialize bias;

[0086] Forward propagation: From the input layer to the hidden layer, the weighted sum of the input signal and the weight is calculated, the bias is added, and then the activation function is applied; the context layer is updated, the hidden layer output of the previous time step is stored in the context layer and used as an additional input for the current time step, which helps the network remember the previous state; from the hidden layer to the output layer, the final output signal is calculated using the output of the hidden layer and including feedback from the context layer;

[0087] Evaluation loss: Use the mean squared error loss function to evaluate the difference between the network output and the true label;

[0088] Backpropagation: Compute the gradient of each weight with respect to the loss function via the chain rule, then update the weights and biases via the gradient descent optimization algorithm; and

[0089] Parameter adaptive adjustment: Use learning rate decay to adjust the learning rate and use L1 / L2 regularization to prevent overfitting.

[0090] In the embodiment of the present invention, in the aforementioned adaptive ELMAN neural network, the number of neuron nodes in the input layer and the output layer are configured to be 2 respectively, and the value of the constant a is configured to be in the range of 1-5.

[0091] As an optional embodiment, in step S104, bilateral filtering is performed on the infrared / visible light fusion image to implement global smoothing post-processing, including:

[0092] Traverse each pixel: Apply bilateral filtering to each pixel of the infrared / visible light fusion image, and for each pixel, calculate the weighted average based on the dynamically adjusted spatial kernel and intensity kernel in its neighborhood;

[0093] Weight calculation: Calculate filter weights, including spatial weights based on the spatial distance between pixels and a dynamically adjusted spatial kernel, and intensity weights based on the difference in pixel values ​​and a dynamically adjusted intensity kernel;

[0094] Output to generate an updated image: Update the current pixel value of each pixel position to the weighted average of its neighboring pixels, and output the overall updated image;

[0095] The dynamically adjusted spatial kernel and intensity kernel are filter kernels adjusted based on the gradient and variance statistical characteristics of the fused image, including dynamically adjusting the size of the spatial kernel according to the gradient size and adjusting the range of the intensity kernel according to the variance.

[0096] It should be understood that in this embodiment, adaptive bilateral filtering is used for global post-processing, and the spatial kernel and the intensity kernel are dynamically adjusted according to the local characteristics of the image, that is, the standard deviation σ_s of the spatial kernel and the standard deviation σ_r of the intensity kernel are adjusted to automatically reduce the degree of filtering in areas with edges or rich details, and enhance the filtering effect in smoother areas.

[0097] In another embodiment, a classic bilateral filtering algorithm may be used to define a spatial neighbor function (such as a Gaussian function, used to reduce the influence of pixels farther from the center pixel) and an intensity similarity function (such as a Gaussian function, used to reduce the influence of pixels that are greatly different from the center pixel), and to perform weighted averaging of two Gaussian functions to update each pixel to implement bilateral filtering, thereby protecting the edges while removing noise and preventing edge blurring.

[0098] As an optional embodiment, in step S105, the infrared / visible light fusion image processed after global smoothing is subjected to differential processing with the background image to implement intrusion detection and identification, including:

[0099] Acquire background infrared images and background visible light images of the monitoring area collected in multiple time periods according to a preset cycle, and obtain an infrared / visible light fused background image after image fusion;

[0100] According to the moment of the infrared / visible light fusion image after global smoothing, the time period is determined;

[0101] Perform background difference processing on the infrared / visible light fusion image after global smoothing and the infrared / visible light fusion background image of the same time period to achieve foreground target intrusion detection;

[0102] Perform pedestrian target detection on foreground targets, identify pedestrians and issue early warnings, as well as dynamic pedestrian trajectory tracking.

[0103] It should be understood that the aforementioned background image, also called a background image model, is used to represent a scene when there is no intruder. The background image can be obtained in the following manner: when there is no intruder, images are taken according to a preset time period, and the infrared and visible light images obtained are directly used as the background image after being fused.

[0104] Here, we divide a day into 24 hours. The aforementioned time period division standard is limited to 24 hours a day, for example, the background image is captured every hour, every 30 minutes, or every 15 minutes.

[0105] Therefore, the globally smoothed fusion image is subtracted pixel by pixel from the background image to generate a difference image, thereby obtaining the changed areas in the highlighted image, i.e., possible intruders.

[0106] Then, by applying threshold segmentation to the difference image, the areas with significant differences (possible intruders) are distinguished from the background. According to the set threshold, the pixels above the threshold are marked as foreground (possible intruders) and the pixels below the threshold are marked as background.

[0107] Furthermore, pedestrian detection algorithms (such as classification and recognition algorithms based on neural networks) are used to trigger an alarm or perform corresponding early warning responses once a suspected intruder is detected. At the same time, the detected intruder is further processed, such as tracking, recording, or more sophisticated behavior, motion analysis, and face recognition.

[0108] Example 2

[0109] In combination with the above embodiments of the intrusion detection and identification method based on the fusion of infrared and visible light, according to the present invention, a computer system is also proposed, including: one or more processors; and a memory.

[0110] The memory is configured to store operable instructions.

[0111] Among them, when the instruction is executed by one or more processors, it causes the aforementioned one or more processors to perform operations, and the operations include the process of executing the intrusion detection and identification method based on the fusion of infrared and visible light of the aforementioned embodiment.

[0112] Example 3

[0113] In combination with the embodiments of the intrusion detection and identification method based on the fusion of infrared and visible light in the above embodiments, a computer-readable medium storing software is also proposed according to the present invention, wherein the software includes instructions that can be executed by one or more computers, and the instructions, when executed by the one or more computers, perform the process of the intrusion detection and identification method based on the fusion of infrared and visible light in the aforementioned embodiments.

[0114] Although the present invention has been disclosed as above with preferred embodiments, it is not intended to limit the present invention. A person with ordinary knowledge in the technical field to which the present invention belongs may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be determined by the definition of the claims.

Claims

1. An intrusion detection and identification method based on the fusion of infrared and visible light, characterized in that: The following steps are involved: Step S101, acquiring images detected by an infrared camera and a visible light camera at the same time, which are an infrared image and a visible light image respectively; Step S102, geometrically correcting the infrared image and the visible light image so that the viewing angles and scales of the two images are consistent; Step S103, performing pixel-level weighted fusion of the infrared image and the visible light image, determining the weighting coefficient based on the adaptive ELMAN neural network and calculating the fused pixel value according to the grayscale values ​​of the infrared image and the visible light image, to obtain an infrared / visible light fused image; Step S104, performing bilateral filtering on the infrared / visible light fusion image to achieve global smoothing post-processing; as well as Step S105, performing differential processing on the infrared / visible light fusion image after global smoothing and the background image to achieve intrusion detection and identification; Wherein, in the step S103, the infrared image and the visible light image are subjected to pixel-level weighted fusion, the weighting coefficient is determined based on the adaptive ELMAN neural network, and the fused pixel value is calculated according to the grayscale value of the infrared image and the visible light image to obtain the infrared / visible light fused image, including: Pixel-level weighted fusion is performed based on the following pixel-level weighted calculation: IF(x,y)=k1*RGB(x,y)+ k2*IR(x,y) Wherein, IF(x,y) represents the value of the pixel point (x,y) in the infrared / visible light fusion image, RGB(x,y) and IR(x,y) represent the values ​​of the corresponding pixel point (x,y) in the visible light image and the infrared image respectively; k1 and k2 represent the weighting coefficients of the visible light image and the infrared image respectively, which are set to determine the coefficient value of each infrared / visible light fusion image through the adaptive ELMAN neural network; And, the process of determining the weighting coefficient of each infrared / visible light fusion image by the adaptive ELMAN neural network includes: The adaptive ELMAN neural network is a dynamic recursive neural network structure, which consists of an input layer, a hidden layer, a context layer and an output layer. The input layer and the output layer each include two neuron nodes, and the number of neuron nodes in the hidden layer and the context layer is the same. The input layer inputs a signal, and the hidden layer uses a nonlinear function as an excitation function. The context layer is used to memorize the output value of the hidden layer at the previous moment, so that the output of the hidden layer is self-connected back to the input of the hidden layer after being delayed by the context layer, and the output weight coefficient is output through weighted processing in the output layer; The grayscale mean Gray of the visible light image and infrared image at time t RGB With Gray IR As the input layer input of the adaptive ELMAN neural network, the number of neuron nodes in the hidden layer of the adaptive ELMAN neural network h Determined by: ; Among them, m and n represent the number of neuron nodes in the input layer and the output layer respectively, and a represents a constant with a value range of 1 to 10; The hidden layer outputs the brightness mean, standard deviation and coefficient of variation of the fused image at the current time t. The context layer receives the output of the hidden layer and memorizes the output value of the hidden layer at the previous time t-1, so that the output of the hidden layer is self-connected back to the input of the hidden layer after being delayed by the context layer. Finally, the two outputs after weighted processing by the two neurons in the output layer are the weighted coefficients k1 and k2 of the visible light image and the infrared image. The brightness mean of the fused image refers to the mean of the grayscale values ​​of the fused image, reflecting the overall brightness of the image; The standard deviation represents the degree of dispersion between the grayscale value of the pixel points of the fused image and the mean value, which is expressed as: Standard deviation = square root of variance, where variance is the sum of the squares of the grayscale value of each pixel in the fused image and the mean of the grayscale value; The coefficient of variation represents the contrast and clarity of the fused image, and is expressed as: Coefficient of variation = standard deviation / mean brightness.

2. The intrusion detection and identification method based on infrared and visible light fusion according to claim 1 is characterized in that: The step S101 also includes graying the acquired visible light image to obtain a grayscale image corresponding to the visible light image.

3. The intrusion detection and identification method based on infrared and visible light fusion according to claim 1 is characterized in that: In step S102, geometric correction is performed on the infrared image and the visible light image so that the viewing angle and scale of the two images are consistent, including: Interpolate and resample the infrared image and the visible light image to make them have the same size; and Perform an affine transformation on the infrared image so that it is at the same viewing angle as the visible light image.

4. The intrusion detection and identification method based on infrared and visible light fusion according to claim 1 is characterized in that: In step S103, the adaptive ELMAN neural network is a pre-trained network model, and the training process includes: Initialize network parameters, including weights from input to hidden layer, from hidden layer to output layer, from hidden layer to context layer, and randomly initialize bias; Forward propagation: From the input layer to the hidden layer, the weighted sum of the input signal and the weight is calculated, the bias is added, and then the activation function is applied; the context layer is updated, the hidden layer output of the previous time step is stored in the context layer and used as an additional input for the current time step, which helps the network remember the previous state; from the hidden layer to the output layer, the final output signal is calculated using the output of the hidden layer and including feedback from the context layer; Evaluation loss: Use the mean squared error loss function to evaluate the difference between the network output and the true label; Backpropagation: Compute the gradient of each weight with respect to the loss function via the chain rule, then update the weights and biases via the gradient descent optimization algorithm; and Parameter adaptive adjustment: Use learning rate decay to adjust the learning rate and use L1 / L2 regularization to prevent overfitting.

5. The intrusion detection and identification method based on infrared and visible light fusion according to claim 4 is characterized in that: In step S104, bilateral filtering is performed on the infrared / visible light fusion image to achieve global smoothing post-processing, including: Traverse each pixel: Apply bilateral filtering to each pixel of the infrared / visible light fusion image, and for each pixel, calculate the weighted average based on the dynamically adjusted spatial kernel and intensity kernel in its neighborhood; Weight calculation: Calculate filter weights, including spatial weights based on the spatial distance between pixels and a dynamically adjusted spatial kernel, and intensity weights based on the difference in pixel values ​​and a dynamically adjusted intensity kernel; Output to generate an updated image: Update the current pixel value of each pixel position to the weighted average of its neighboring pixels, and output the overall updated image; The dynamically adjusted spatial kernel and intensity kernel are filter kernels adjusted based on the gradient and variance statistical characteristics of the fused image, including dynamically adjusting the size of the spatial kernel according to the gradient size and adjusting the range of the intensity kernel according to the variance.

6. The intrusion detection and identification method based on infrared and visible light fusion according to claim 1 is characterized in that: In step S105, the infrared / visible light fusion image after global smoothing is subjected to differential processing with the background image to achieve intrusion detection and identification, including: Acquire background infrared images and background visible light images of the monitoring area collected in multiple time periods according to a preset cycle, and obtain an infrared / visible light fused background image after image fusion; According to the moment of the infrared / visible light fusion image after global smoothing, the time period is determined; Perform background difference processing on the infrared / visible light fusion image after global smoothing and the infrared / visible light fusion background image of the same time period to achieve foreground target intrusion detection; Perform pedestrian target detection on foreground targets, identify pedestrians and issue early warnings as well as dynamic trajectory tracking.

7. A computer-readable medium storing software, characterized in that: The software includes instructions that can be executed by one or more computers, and when the instructions are executed by the one or more computers, the process of the intrusion detection and identification method based on the fusion of infrared and visible light as described in any one of claims 1-6 is performed.

8. A computer system, characterized in that: include: one or more processors; Memory, storing instructions that can be operated; Wherein, when the instruction is executed by one or more processors, the aforementioned one or more processors perform operations, and the operations include the process of executing the intrusion detection and identification method based on the fusion of infrared and visible light as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Lightweight infrared and visible light image fusion method based on convolutional neural network

    CN116681636A