Method and device for removing external interference features from video surveillance images

By establishing a migration conversion model between optical surveillance images and infrared surveillance images, and using the generative adversarial network to generate virtual images, the problem of removing external interference features in traditional methods is solved, the accuracy of recognition of video surveillance images is improved, and parking difficulties and parking disorders are alleviated.

CN114519703BActive Publication Date: 2025-08-22AI SUPER EYE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210127008.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-11
Publication Date
2025-08-22
Estimated Expiration
2042-02-11

AI Technical Summary

Technical Problem

Traditional image processing methods cannot effectively remove external environmental interference characteristics, resulting in poor accuracy in video surveillance image detection and recognition, and cannot effectively alleviate the problems of difficulty in parking and parking disorder.

Method used

By establishing a migration conversion model between the optical monitoring image and the infrared monitoring image, a virtual image is generated using the generative adversarial network to form a cyclic link, and the image processing model is optimized through the loss function to remove external interference characteristics.

Benefits of technology

It improves the recognition accuracy of video surveillance images and effectively alleviates the problems of difficulty in parking and parking mess on the roadside.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114519703B_ABST
    Figure CN114519703B_ABST
Patent Text Reader

Abstract

The present application discloses a method and apparatus for removing external interference features from video surveillance images. The method includes: an optical surveillance image and an infrared surveillance image have overlapping areas, the optical surveillance image being an image obtained by an optical camera without external interference features, and the infrared surveillance image being an image obtained by an infrared camera; inputting the optical surveillance image and the infrared surveillance image into an optical-infrared transformation model to generate a first virtual infrared surveillance image; inputting the first virtual infrared surveillance image and the optical surveillance image into the infrared optical transformation model to generate a first virtual optical surveillance image, and forming a first loop link; inputting the infrared surveillance image and the optical surveillance image into the infrared optical transformation model to generate a second virtual optical surveillance image; and inputting the second virtual optical surveillance image and the infrared surveillance image into the optical-infrared transformation model to generate a second virtual infrared surveillance image, and forming a second loop link.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method and device for removing external interference features from video surveillance images. Background Art

[0002] As the number of cars continues to increase, parking difficulties and irregularities have become a major issue that needs to be addressed in my country's urban transportation development. Intelligent parking management using high-positioned video cameras is currently a mainstream solution, effectively alleviating these issues.

[0003] Due to the influence of external factors such as sunlight, vehicle high and low beam headlights, rain, snow, or fog, some areas of the captured image may lose color, texture, and other features. However, traditional image processing methods use scale-invariant feature transformation algorithms to process the captured images, which cannot accurately extract feature points, has poor real-time performance, and cannot remove interfering features caused by the external environment, resulting in inaccurate detection and recognition of images captured under the influence of external factors.

[0004] Application Contents

[0005] The purpose of this application is to solve the technical problem that traditional image processing methods cannot remove interference features caused by the external environment. To achieve the above purpose, this application provides a method and device for removing external interference features from video surveillance images.

[0006] This application provides a method for removing external interference features from video surveillance images, including:

[0007] Acquire an optical monitoring image set and an infrared monitoring image set, wherein the optical monitoring images in the optical monitoring image set and the infrared monitoring images in the infrared monitoring image set have overlapping areas, the optical monitoring images are feature images obtained by an optical camera without external interference, and the infrared monitoring images are images obtained by an infrared camera;

[0008] Inputting the optical monitoring image and the infrared monitoring image into an optical-infrared conversion model to generate a first virtual infrared monitoring image;

[0009] Inputting the first virtual infrared monitoring image and the optical monitoring image into an infrared optical transformation model to generate a first virtual optical monitoring image;

[0010] Inputting the infrared monitoring image and the optical monitoring image into the infrared optical transformation model to generate a second virtual optical monitoring image;

[0011] inputting the second virtual optical monitoring image and the infrared monitoring image into the optical-infrared conversion model to generate a second virtual infrared monitoring image;

[0012] The optical monitoring image, the infrared monitoring image, the first virtual infrared monitoring image, and the first virtual optical monitoring image form a first loop link, and the infrared monitoring image, the optical monitoring image, the second virtual optical monitoring image, and the second virtual infrared monitoring image form a second loop link;

[0013] Constructing a loss function of an image processing model according to the loss function of the optical infrared transformation model, the loss function of the infrared optical transformation model, the loss function of the first loop link, and the loss function of the second loop link;

[0014] The image processing model is trained and optimized according to the total loss function to obtain a trained image processing model, and the infrared monitoring image corresponding to the optical monitoring image to be measured is processed according to the trained image processing model to obtain a feature image without external interference.

[0015] In one embodiment, the present application provides a device for removing external interference features from a video surveillance image, comprising:

[0016] An image acquisition module, configured to acquire an optical monitoring image set and an infrared monitoring image set, wherein the optical monitoring images in the optical monitoring image set and the infrared monitoring images in the infrared monitoring image set have overlapping areas, the optical monitoring images are characteristic images obtained by an optical camera without external interference, and the infrared monitoring images are images obtained by an infrared camera;

[0017] a first virtual infrared monitoring image generation module, configured to input the optical monitoring image and the infrared monitoring image into an optical infrared conversion model to generate a first virtual infrared monitoring image;

[0018] A first virtual optical monitoring image generation module, configured to input the first virtual infrared monitoring image and the optical monitoring image into an infrared optical transformation model to generate a first virtual optical monitoring image;

[0019] a second virtual optical monitoring image generating module, configured to input the infrared monitoring image and the optical monitoring image into the infrared optical transformation model to generate a second virtual optical monitoring image;

[0020] a second virtual infrared monitoring image generating module, configured to input the second virtual optical monitoring image and the infrared monitoring image into the optical-infrared conversion model to generate a second virtual infrared monitoring image;

[0021] a loop link forming module, configured to form a first loop link with the optical surveillance image, the infrared surveillance image, the first virtual infrared surveillance image, and the first virtual optical surveillance image, and to form a second loop link with the infrared surveillance image, the optical surveillance image, the second virtual optical surveillance image, and the second virtual infrared surveillance image;

[0022] a loss function construction module, configured to construct a loss function of an image processing model according to the loss function of the optical infrared transformation model, the loss function of the infrared optical transformation model, the loss function of the first loop link, and the loss function of the second loop link;

[0023] The module for acquiring characteristic images without external interference is used to train and optimize the image processing model according to the total loss function to obtain a trained image processing model, and to process the infrared monitoring image corresponding to the optical monitoring image to be measured according to the trained image processing model to obtain a characteristic image without external interference.

[0024] In the aforementioned method and apparatus for removing external interference features from video surveillance images, this application establishes a migration and conversion model between optical surveillance images and infrared surveillance images that do not contain external interference features. When the optical surveillance image to be measured contains external interference features, a migration and conversion is performed on the infrared surveillance image corresponding to the optical surveillance image to obtain a non-interference feature image that does not contain external interference features, thereby achieving the purpose of removing external interference features. Consequently, when detecting and identifying the non-interference feature image with the removed external interference features, recognition accuracy can be improved, effectively alleviating the problems of difficult and chaotic roadside parking. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a flowchart of the steps of the method for removing external interference features from video surveillance images provided by this application.

[0026] Figure 2 This is a schematic diagram of the structure of the device for removing external interference features of video surveillance images provided by this application. DETAILED DESCRIPTION

[0027] The technical solution of the present application is further described in detail below through the accompanying drawings and examples.

[0028] See Figure 1 , this application provides a method for removing external interference features from video surveillance images, including:

[0029] S10, acquiring an optical monitoring image set and an infrared monitoring image set, wherein the optical monitoring images in the optical monitoring image set and the infrared monitoring images in the infrared monitoring image set have overlapping areas, the optical monitoring images are feature images obtained by an optical camera without external interference, and the infrared monitoring images are images obtained by an infrared camera;

[0030] S20, inputting the optical monitoring image and the infrared monitoring image into an optical-infrared conversion model to generate a first virtual infrared monitoring image;

[0031] S30, inputting the first virtual infrared monitoring image and the optical monitoring image into an infrared optical transformation model to generate a first virtual optical monitoring image;

[0032] S40, inputting the infrared monitoring image and the optical monitoring image into an infrared optical transformation model to generate a second virtual optical monitoring image;

[0033] S50, inputting the second virtual optical monitoring image and the infrared monitoring image into the optical-infrared conversion model to generate a second virtual infrared monitoring image;

[0034] S60, forming a first loop link with the optical monitoring image, the infrared monitoring image, the first virtual infrared monitoring image, and the first virtual optical monitoring image, and forming a second loop link with the infrared monitoring image, the optical monitoring image, the second virtual optical monitoring image, and the second virtual infrared monitoring image;

[0035] S70, constructing a loss function of the image processing model according to the loss function of the optical infrared transformation model, the loss function of the infrared optical transformation model, the loss function of the first loop link, and the loss function of the second loop link;

[0036] S80, training and optimizing the image processing model according to the total loss function to obtain a trained image processing model, and processing the infrared monitoring image corresponding to the optical monitoring image to be measured according to the trained image processing model to obtain a feature image without external interference.

[0037] In S10, an optical camera and a thermal infrared camera are installed on a roadside pole as two different types of devices. By adjusting the installation position of the optical camera or the thermal infrared camera, the optical monitoring image and the infrared monitoring image acquired respectively have overlapping parts. The image data acquired by the optical camera and the thermal infrared camera have the same frame rate (Frame Per Second, FPS). Alternatively, by appropriately rotating or translating the images acquired by the optical camera or the thermal infrared camera, the optical monitoring image and the infrared monitoring image have overlapping images. The overlapping area of ​​the optical monitoring image and the infrared monitoring image can make the optical monitoring image and the infrared monitoring image have the same target area. An image without external interference features can be understood as an image without external interference features. External interference features include interference features caused by factors such as reflections, fog, raindrops, shadows, etc. in the acquired image due to external environmental factors such as sunlight, vehicle high and low beam headlights, rain, snow, or fog. The infrared monitoring image is an image formed by the infrared camera based on the amount of infrared thermal radiation energy of the target object. Infrared cameras are intuitive and efficient, effectively identifying targets even in darkness, rain, snow, fog, and other weather conditions. Consequently, infrared surveillance images are less susceptible to environmental influences. Optical and infrared surveillance images are captured at different times, such as morning, noon, and evening. They are also captured at camera positions corresponding to different camera pole positions. These images clearly capture objects such as vehicles, pedestrians, and license plates, and also show vehicles entering, parked, and exiting parking spaces.

[0038] In S20, the optical-infrared conversion model is an image-video synthesis model. The optical-infrared conversion model may be a synthesis algorithm capable of mapping the optical surveillance image to the infrared surveillance image to form a first virtual infrared surveillance image. In one embodiment, the optical-infrared conversion model may be a Generative Adversarial Network (GAN).

[0039] In S30, the infrared optical transformation model is an image and video synthesis model. The infrared optical transformation model can be a synthesis algorithm capable of mapping the infrared surveillance image to the optical surveillance image to form a first virtual optical surveillance image. In one embodiment, the infrared optical transformation model can be a generative adversarial network (GAN). Through S20 to S30, the process of mapping the optical surveillance image to the infrared surveillance image and then back to the optical surveillance image is implemented, forming a first loop link.

[0040] In S40 , the infrared monitoring image can be mapped to the optical monitoring image through the infrared optical transformation model to form a second virtual optical monitoring image.

[0041] In S50, the second virtual optical monitoring image is mapped to the infrared monitoring image using the optical-infrared transformation model to form a second virtual infrared monitoring image. The process of mapping the infrared monitoring image to the optical monitoring image and then back to the infrared monitoring image is implemented through S40 to S50, forming a second loop.

[0042] In S60, the optical monitoring image and the first virtual optical monitoring image in the first loop link have a high degree of similarity. By comparing the two to establish the loss function of the first loop link, the feasibility of both the optical infrared transformation model and the infrared optical transformation model can be demonstrated. The infrared monitoring image and the second virtual infrared monitoring image in the second loop link have a high degree of similarity. By comparing the two to establish the loss function of the second loop link, the feasibility of both the optical infrared transformation model and the infrared optical transformation model can be demonstrated. By establishing two loop links, model training of both the optical infrared transformation model and the infrared optical transformation model can be performed from multiple perspectives.

[0043] In S70 , the optical-infrared conversion model, the infrared-optical conversion model, the first loop link, and the second loop link together form the entire image processing model. Parameters of the image processing model as a whole are adjusted and optimized to obtain trained and optimized parameters. By constructing loss functions for each step to form an overall loss function, parameter adjustment and optimization are performed as a whole, optimizing parameters from multiple perspectives to achieve a balance between each link and obtain a stable model.

[0044] In S80, the image processing model is optimized and trained to obtain optimized model parameters, thereby obtaining a trained image processing model. The trained image processing model is used to process an infrared monitoring image corresponding to the optical monitoring image to be tested, thereby obtaining an image free of external interference features. The optical monitoring image to be tested contains external interference features. By removing the external interference features from the optical monitoring image to be tested, an infrared monitoring image corresponding to the optical monitoring image to be tested is obtained, and the infrared monitoring image is converted to obtain an image free of external interference features corresponding to the optical monitoring image to be tested.

[0045] Therefore, this application utilizes a method for removing external interference features from video surveillance images to establish a migration and conversion model between optical surveillance images and infrared surveillance images that do not contain external interference features. When the optical surveillance image to be tested contains external interference features, a migration and conversion is performed on the corresponding infrared surveillance image to obtain a non-interference feature image that does not contain external interference features, thereby achieving the purpose of removing external interference features. Consequently, when detecting and recognizing the non-interference feature image after removing external interference features, recognition accuracy can be improved, effectively alleviating the problems of difficult and chaotic roadside parking.

[0046] In one embodiment, in S10, the overlap between the optical surveillance image and the infrared surveillance image is 95% to 100%. By setting the range of the overlap, the optical surveillance image and the infrared surveillance image can share the same target area, resulting in consistent imagery in the overlapped area. The optical camera and the thermal infrared camera are synchronized in time, and video de-framing is performed to obtain images captured by the two cameras that are synchronized in time and have similar images. In one embodiment, the optical surveillance image and the infrared surveillance image are pre-processed, such as by unifying the image size.

[0047] In one embodiment, S20, inputting the optical monitoring image and the infrared monitoring image into an optical-infrared conversion model to generate a first virtual infrared monitoring image includes:

[0048] S210, inputting the optical monitoring image and the infrared monitoring image into a generator network of an optical-infrared conversion model, and outputting a first virtual infrared monitoring image;

[0049] S220: Input the first virtual infrared monitoring image and the infrared monitoring image into a discriminator network of an optical infrared transformation model to discriminate the first virtual infrared monitoring image and the infrared monitoring image.

[0050] In this embodiment, the optical infrared transformation model is a generative adversarial network, comprising a generator network and a discriminator network. By iteratively training the generator and discriminator networks, the two networks compete with each other to produce a better optical infrared transformation model. The generator network is used to generate a first virtual infrared surveillance image. The discriminator network is used to determine whether the generated first virtual infrared surveillance image is realistic.

[0051] In one embodiment, the input image feature size of the generator network of the optical infrared transformation model is 256*256*3 (H*W*C, Height*Width*Channel). The input image passes through multiple convolutional layers, normalization layers, and nonlinear activation layers, reducing the feature map size to 64*64*256. For further feature extraction, the feature map size remains unchanged after passing through a residual convolution module. Finally, the feature map size is restored to the original image size of 256*256*3 after passing through multiple deconvolution layers.

[0052] The residual convolution module consists of multiple connections of residual convolutional neural networks. Normalization layers include, but are not limited to, instance normalization layers and adaptive instance normalization layers. Non-linear activation layers include, but are not limited to, non-linear activation functions such as rectified linear units (ReLUs).

[0053] In one embodiment, the discriminator network of the optical infrared transformation model is a binary classification network, and the classification goal is to distinguish the original input real image data from the image data generated by the generator. The discriminator network of the optical infrared transformation model has a relatively shallow number of layers, and the input image feature size is 256*256*3 (H*W*C, Height*Width*Channel, height*width*number of channels). The feature map is downsampled through multiple convolution operations in sequence, and then the features are classified using a binary classification loss function. Binary classification loss functions include but are not limited to cross entropy loss function, logarithmic loss function, etc.

[0054] In one embodiment, after inputting the optical monitoring image and the infrared monitoring image into the optical-infrared transformation model to generate a first virtual infrared monitoring image at step S20, the method further comprises: inputting the first virtual infrared monitoring image and the optical monitoring image into the infrared-optical transformation model at step S30 before generating the first virtual optical monitoring image.

[0055] S201, inputting the first two adjacent optical monitoring images of the optical monitoring image and the first two adjacent infrared monitoring images of the infrared monitoring image into an optical-infrared conversion model to obtain the first two adjacent virtual infrared monitoring images of a first virtual infrared monitoring image;

[0056] S202, inputting the first two adjacent virtual infrared surveillance images into an infrared surveillance prediction model for prediction to obtain a predicted infrared surveillance image;

[0057] S203: Perform an average calculation on the predicted infrared surveillance image and the first virtual infrared surveillance image to obtain an average virtual infrared surveillance image.

[0058] In S201, the first two adjacent optical surveillance images can be understood as the two frames preceding the optical surveillance image corresponding to the current moment. For example, if the optical surveillance image corresponding to the current moment is the i-th optical surveillance image, then the first two adjacent optical surveillance images are the i-1th optical surveillance image and the i-2th optical surveillance image, respectively. Similarly, the first two adjacent infrared surveillance images of an infrared surveillance image can be understood as the two frames preceding the infrared surveillance image corresponding to the current moment. For example, if the infrared surveillance image corresponding to the current moment is the i-th infrared surveillance image, then the first two adjacent infrared surveillance images are the i-1th infrared surveillance image and the i-2th infrared surveillance image, respectively. The i-1th optical surveillance image and the i-1th infrared surveillance image are input into the optical-infrared transformation model to obtain the i-1th virtual infrared surveillance image. The i-2th optical surveillance image and the i-2th infrared surveillance image are input into the optical-infrared transformation model to obtain the i-2th virtual infrared surveillance image.

[0059] In S202, the i-1th frame virtual infrared monitoring image and the i-2th frame virtual infrared monitoring image are input into the infrared monitoring prediction model for prediction to obtain the i-th frame predicted infrared monitoring image, which can also be understood as the current predicted infrared monitoring image. The infrared monitoring prediction model can be a time predictor. In one embodiment, the time predictor is implemented by an optical flow neural network (FlowNet). There is a correlation between adjacent optical monitoring images, adjacent infrared monitoring images, and adjacent virtual infrared monitoring images. The optical flow neural network is used to find the corresponding relationship between the previous frame image and the current frame image, and the motion information of the object between the adjacent frame images is calculated. There is object motion information between the previous frame image and the current frame image, and between the current frame image and the next frame image, and a corresponding relationship can be established between them. Based on the motion information of the object between the adjacent frame images, the data of the future frame image is predicted, that is, the data of the next frame image can be predicted.

[0060] In S203, the average calculation can be understood as performing an average process on the two images. By combining the predicted image with the virtual image and taking an average, a more accurate average virtual infrared monitoring image can be obtained.

[0061] The infrared surveillance prediction model can predict the current virtual infrared surveillance image based on the two previous adjacent virtual infrared surveillance images, enabling prediction of future data from past data. By predicting adjacent frames, the training model is constrained by time series constraints, further improving the accuracy and realism of the current virtual infrared surveillance image.

[0062] In one embodiment, the first two adjacent virtual infrared surveillance images are input into the contraction layer of the convolutional layer of an optical flow neural network. Feature maps of each of the first two adjacent virtual infrared surveillance images are extracted to obtain two adjacent feature maps. The two adjacent feature maps are then input into the correlation layer of the optical flow neural network to calculate the correlation features of the two adjacent feature maps. The correlation features of the two adjacent feature maps are then input into the expansion layer of the deconvolutional layer of the optical flow neural network to perform optical flow prediction, thereby achieving a prediction of the next image frame. Thus, the optical flow neural network and the first two adjacent virtual infrared surveillance images can be used to predict the next image frame, obtaining a predicted infrared surveillance image.

[0063] In one embodiment, S30, inputting the first virtual infrared monitoring image and the optical monitoring image into the infrared optical transformation model to generate the first virtual optical monitoring image includes:

[0064] S310: Input the average virtual infrared monitoring image and the optical monitoring image into an infrared optical transformation model to generate a first virtual optical monitoring image.

[0065] In this embodiment, the averaged virtual infrared surveillance image is sequentially processed by the predictor and averaging to obtain a more accurate and realistic virtual infrared surveillance image. Furthermore, the averaged virtual infrared surveillance image and the optical surveillance image are input into the infrared optical transformation model to obtain a more accurate and realistic first virtual optical surveillance image. This improves the model's conversion accuracy, preserves more realistic and accurate image features, and facilitates subsequent detection and recognition.

[0066] In one embodiment, S310, inputting the average virtual infrared monitoring image and the optical monitoring image into an infrared optical transformation model to generate a first virtual optical monitoring image includes:

[0067] S311, inputting the average virtual infrared monitoring image and the optical monitoring image into a generator network of an infrared optical transformation model, and outputting a first virtual optical monitoring image;

[0068] S312: Input the first virtual optical monitoring image and the optical monitoring image into a discriminator network of an infrared optical transformation model to discriminate the first virtual optical monitoring image and the optical monitoring image.

[0069] In this embodiment, the infrared optical transformation model is a generative adversarial network, comprising a generator network and a discriminator network. By iteratively training the generator and discriminator networks, the two networks compete with each other to produce a better infrared optical transformation model. The generator network is used to generate a first virtual optical monitoring image. The discriminator network is used to determine whether the generated first virtual optical monitoring image is realistic.

[0070] The infrared-optical transformation model and the optical-infrared transformation model form two mirror-symmetric generative adversarial networks, forming a recurrent generative adversarial network. Compared to Cycle-GAN, which focuses solely on spatial information, the recurrent generative adversarial network adds temporal constraints, enabling more realistic and accurate transfer results. The recurrent generative adversarial network utilizes unpaired training data and two mirror-symmetric generative adversarial networks. During model training, it learns to map optical surveillance images to infrared surveillance images and back again. The combined game between the generator and discriminator networks makes the recurrent generative adversarial network more stable.

[0071] In one embodiment, the input image feature size of the generator network of the infrared optical transformation model is 256*256*3 (H*W*C, Height*Width*Channel). The input image passes through multiple convolutional layers, normalization layers, and nonlinear activation layers, and the feature map size is reduced to 64*64*256. For further feature extraction, it passes through the residual convolution module, and the feature map size remains unchanged. Finally, it passes through multiple deconvolution layers to restore the feature map size to the original image size of 256*256*3.

[0072] The residual convolution module consists of multiple connections of residual convolutional neural networks. Normalization layers include, but are not limited to, instance normalization layers and adaptive instance normalization layers. Non-linear activation layers include, but are not limited to, non-linear activation functions such as rectified linear units (ReLUs).

[0073] In one embodiment, the discriminator network of the infrared optical transformation model is a binary classification network, and the classification goal is to distinguish the original input real image data from the image data generated by the generator. The discriminator network layer of the infrared optical transformation model is relatively shallow, and the input image feature size is 256*256*3 (H*W*C, Height*Width*Channel, height*width*number of channels). The feature map is downsampled through multiple convolution operations in sequence, and then the features are classified using a binary classification loss function. The binary classification loss function includes but is not limited to the cross entropy loss function, the logarithmic loss function, etc.

[0074] In one embodiment, S40, inputting the infrared monitoring image and the optical monitoring image into the infrared optical transformation model to generate a second virtual optical monitoring image includes:

[0075] S410, inputting the infrared monitoring image and the optical monitoring image into a generator network of an infrared optical transformation model, and outputting a second virtual optical monitoring image;

[0076] S420: Input the infrared monitoring image and the optical monitoring image into a discriminator network of the infrared optical transformation model to discriminate the second virtual optical monitoring image and the optical monitoring image.

[0077] In this embodiment, the infrared optical transformation model is the same as that in S30, and the corresponding descriptions of the generator network and discriminator network are also the same. The infrared optical transformation model is a generative adversarial network, including a generator network and a discriminator network. By iteratively training the generator network and the discriminator network, the two networks compete with each other to generate a better infrared optical transformation model. The generator network is used to generate a second virtual optical monitoring image. The discriminator network is used to determine whether the generated second virtual optical monitoring image is realistic.

[0078] In one embodiment, after inputting the infrared monitoring image and the optical monitoring image into the infrared optical transformation model to generate a second virtual optical monitoring image at step S40, and before inputting the second virtual optical monitoring image and the infrared monitoring image into the optical infrared transformation model to generate the second virtual infrared monitoring image at step S50, the method further includes:

[0079] S401, inputting the first two adjacent infrared monitoring images of the infrared monitoring image and the first two adjacent optical monitoring images of the optical monitoring image into an infrared optical transformation model to obtain the first two adjacent virtual optical monitoring images of the second virtual optical monitoring image;

[0080] S402, inputting the first two adjacent virtual optical monitoring images into an optical monitoring prediction model for prediction to obtain a predicted optical monitoring image;

[0081] S403 : averaging the predicted optical monitoring image and the second virtual optical monitoring image to obtain an average virtual optical monitoring image.

[0082] In S401, the first two adjacent optical surveillance images can be understood as the two frames preceding the optical surveillance image corresponding to the current moment. For example, if the optical surveillance image corresponding to the current moment is the i-th optical surveillance image, then the first two adjacent optical surveillance images are the i-1th optical surveillance image and the i-2th optical surveillance image, respectively. Similarly, the first two adjacent infrared surveillance images of an infrared surveillance image can be understood as the two frames preceding the infrared surveillance image corresponding to the current moment. For example, if the infrared surveillance image corresponding to the current moment is the i-th infrared surveillance image, then the first two adjacent infrared surveillance images are the i-1th infrared surveillance image and the i-2th infrared surveillance image, respectively. The i-1th infrared surveillance image and the i-1th optical surveillance image are input into the infrared optical transformation model to obtain the i-1th virtual optical surveillance image. The i-2th infrared surveillance image and the i-2th optical surveillance image are input into the infrared optical transformation model to obtain the i-2th virtual optical surveillance image.

[0083] In S402, the i-1th frame virtual optical monitoring image and the i-2th frame virtual optical monitoring image are input into the optical monitoring prediction model for prediction to obtain the i-th frame predicted optical monitoring image, which can also be understood as the current predicted optical monitoring image. The optical monitoring prediction model can be a time predictor. In one embodiment, the time predictor is implemented by an optical flow neural network (FlowNet). There is a correlation between adjacent infrared monitoring images, adjacent optical monitoring images, and adjacent virtual optical monitoring images. The optical flow neural network is used to find the corresponding relationship between the previous frame image and the current frame image, and the motion information of the object between the adjacent frame images is calculated. There is object motion information between the previous frame image and the current frame image, and between the current frame image and the next frame image, and a corresponding relationship can be established between them. Based on the motion information of the object between the adjacent frame images, the data of the future frame image is predicted, that is, the data of the next frame image can be predicted.

[0084] In S403, the average calculation can be understood as performing an average process on the two images. By combining the predicted image and the virtual image and taking an average, a more accurate average virtual optical monitoring image can be obtained.

[0085] The optical monitoring prediction model can predict the current virtual optical monitoring image based on the two previous adjacent virtual optical monitoring images, enabling prediction of future data from past data. By predicting adjacent frames, the training model is constrained by time series, further improving the accuracy and realism of the current virtual optical monitoring image.

[0086] In one embodiment, the first two adjacent virtual optical monitoring images are input into the contraction portion of the convolutional layer of an optical flow neural network. Feature maps of each of the first two adjacent virtual optical monitoring images are extracted to obtain two adjacent feature maps. The two adjacent feature maps are then input into the correlation layer of the optical flow neural network to calculate the correlation features of the two adjacent feature maps. The correlation features of the two adjacent feature maps are then input into the expansion portion of the deconvolutional layer of the optical flow neural network to perform optical flow prediction, thereby achieving a prediction of the next image frame. Thus, the optical flow neural network and the first two adjacent virtual optical monitoring images can be used to predict the next image frame, obtaining a predicted optical monitoring image.

[0087] In one embodiment, S50, inputting the second virtual optical monitoring image and the infrared monitoring image into the optical-infrared conversion model to generate the second virtual infrared monitoring image includes:

[0088] S510 , inputting the average virtual optical monitoring image and the infrared monitoring image into an optical-infrared conversion model to generate a second virtual infrared monitoring image.

[0089] In this embodiment, the averaged virtual optical surveillance image is sequentially processed by the predictor and averaging to obtain a more accurate and realistic virtual optical surveillance image. Furthermore, the averaged virtual optical surveillance image and the infrared surveillance image are input into the optical-infrared conversion model to obtain a more accurate and realistic second virtual infrared surveillance image. This improves the model's conversion accuracy, preserves more realistic and accurate image features, and facilitates subsequent detection and recognition.

[0090] In one embodiment, S510, inputting the average virtual optical monitoring image and the infrared monitoring image into an optical-infrared conversion model to generate a second virtual infrared monitoring image includes:

[0091] S511, inputting the average virtual optical monitoring image and the infrared monitoring image into a generator network of an optical-infrared conversion model, and outputting a second virtual infrared monitoring image;

[0092] S512: Input the average virtual optical monitoring image and the infrared monitoring image into a discriminator network of the optical infrared conversion model to discriminate the second virtual infrared monitoring image and the infrared monitoring image.

[0093] In this embodiment, the optical infrared transformation model is the same as the description of the optical infrared transformation model in S20. The corresponding descriptions of the generator network and the discriminator network are also the same. The optical infrared transformation model is a generative adversarial network, including a generator network and a discriminator network. By iteratively training the generator network and the discriminator network, the two networks compete with each other to generate a better optical infrared transformation model. The generator network is used to generate a second virtual infrared surveillance image. The discriminator network is used to determine whether the generated second virtual infrared surveillance image is realistic.

[0094] The optical-infrared transformation model and the infrared-optical transformation model form two mirror-symmetric generative adversarial networks, forming a recurrent generative adversarial network. Compared to Cycle-GAN, which focuses solely on spatial information, the recurrent generative adversarial network adds temporal constraints, enabling a more natural and smooth transfer effect. The recurrent generative adversarial network utilizes unpaired training data and two mirror-symmetric generative adversarial networks to learn how to map infrared surveillance images to optical surveillance images and back again when training the infrared-optical transformation model. The interplay between the generator and discriminator networks makes the recurrent generative adversarial network more stable.

[0095] In one embodiment, S70, a loss function of an image processing model is constructed based on the loss function of the optical infrared transformation model, the loss function of the infrared optical transformation model, the loss function of the first loop link, and the loss function of the second loop link. The loss function of the image processing model is:

[0096]

[0097] Among them, X represents the infrared monitoring image, Y represents the optical monitoring image, G represents the generator network of the infrared optical transformation model, F represents the generator network of the optical infrared transformation model, and D Y Represents the discriminator network of the infrared optical transformation model, D X a discriminator network representing the optical infrared transformation model;

[0098] L GAN (G, D Y , X, Y) represents the loss function of the infrared optical transformation model, L GAN (F, D X , Y, X) represents the loss function of the optical infrared transformation model, represents the loss function of the second loop link, represents the loss function of the first loop link, and λ1 and λ2 represent weight parameters.

[0099] In this embodiment, two mirror-symmetric generative adversarial networks are formed using an infrared-optical transformation model and an optical-infrared transformation model, forming a recurrent generative adversarial network. This network can simultaneously learn the process of mapping an infrared surveillance image to an optical surveillance image and then back to an infrared surveillance image, and the process of mapping an optical surveillance image to an infrared surveillance image and then back to an optical surveillance image. The optical-infrared transformation model, the infrared-optical transformation model, the first recurrent link, and the second recurrent link together constitute the image processing model. By constructing a loss function for the image processing model, it is possible to adjust various parameters in the model, thereby obtaining optimized parameters and forming a stable image processing model.

[0100] In one embodiment, the loss function of the first recurrent link and the loss function of the second recurrent link may be L1 norm loss functions.

[0101] In one embodiment, according to the infrared monitoring prediction model and the optical monitoring prediction model in S201 to S203 and S401 to S403, S70 constructs a loss function of the image processing model based on the loss function of the optical infrared transformation model, the loss function of the infrared optical transformation model, the loss function of the first loop link, and the loss function of the second loop link. The loss function of the image processing model is:

[0102]

[0103] in, represents the loss function of the infrared monitoring prediction model, represents the loss function of the optical monitoring prediction model, and λ3 and λ4 represent weight parameters.

[0104] In this embodiment, the optical-infrared conversion model, the infrared-optical conversion model, the infrared monitoring prediction model, the optical monitoring prediction model, the first loop link, and the second loop link collectively constitute the image processing model. By constructing a loss function for the image processing model, it is possible to adjust various parameters in the model, thereby obtaining optimized parameters and forming a stable image processing model.

[0105] In one embodiment, S80, training and optimizing the image processing model according to the total loss function to obtain a trained image processing model, and processing the infrared monitoring image corresponding to the optical monitoring image to be measured according to the trained image processing model to obtain a feature image free of external interference, including:

[0106] S810, training and optimizing the optical infrared transformation model, the infrared optical transformation model, the first loop link, and the second loop link according to the total loss function to obtain a trained infrared optical transformation model;

[0107] S820 , obtaining an infrared monitoring image to be measured corresponding to the optical monitoring image to be measured, and inputting the infrared monitoring image to be measured into the trained infrared optical transformation model to obtain a feature image without external interference.

[0108] In S810, the total loss function includes the loss function of the optical infrared transformation model, the loss function of the infrared optical transformation model, the loss function of the first loop link, and the loss function of the second loop link. By adjusting the total loss function to the target balance point, the optical infrared transformation model, the infrared optical transformation model, the first loop link, and the second loop link can be made to influence each other, constraining the infrared optical transformation model from multiple angles. The infrared optical transformation model is adjusted as a whole through the total loss function, and finally a trained infrared optical transformation model is obtained.

[0109] At S820, the optical surveillance image to be tested contains external interference features. The optical surveillance image to be tested also overlaps with the infrared surveillance image to be tested. The infrared surveillance image to be tested is input into the trained infrared optical transformation model to obtain an image free of external interference features. This image free of external interference features shares the same target technical features as the optical surveillance image to be tested, but lacks the external interference features, enabling accurate and reliable application for target detection and recognition.

[0110] In one embodiment, in S80, the image processing model is trained and optimized according to the total loss function to obtain a trained image processing model, and the infrared monitoring image corresponding to the optical monitoring image to be measured is processed according to the trained image processing model to obtain a feature image without external interference, further comprising:

[0111] S830, adjusting the loss function of the image processing model to the target equilibrium point, and obtaining a trained image processing model. The target equilibrium point is:

[0112]

[0113] Among them, G * represents the weight of the generator network of the infrared optical transformation model, F * represents the weight of the generator network of the optical infrared transformation model, and the min-max function represents finding the discriminator network of the infrared optical transformation model and the discriminator network of the optical infrared transformation model, so that the difference between the feature distribution of the input image and the generation distribution of the generator network of the infrared optical transformation model and the generator network of the optical infrared transformation model is maximized, and the difference between the feature distribution of the input image and the generation distribution of the generator network of the infrared optical transformation model and the generator network of the optical infrared transformation model is minimized, and the distribution difference between adjacent frame virtual optical monitoring images is minimized.

[0114] In this embodiment, the min-max function can be understood as finding the discriminator networks corresponding to the two models to maximize the difference between the feature distribution of the input data and the generated distribution of the generator. Simultaneously, the generator networks corresponding to the two models are adjusted to minimize the difference between the distribution of the input data and the generated distribution of the generator network. By adding a temporal infrared monitoring prediction model and an optical monitoring prediction model, constraints on future frame data based on past frame data are implemented, minimizing the distribution difference between adjacent frame data and making the previous and next frame images generated by the generator more accurate and realistic.

[0115] In one embodiment, when calculating the difference between the distribution of input data and the distribution of data generated by the generator network, the difference can be described by functions such as KL divergence (Kullback-Leibler Divergence) or JS divergence (Jensen–Shannon Divergence).

[0116] See Figure 2 The present application provides a device 100 for removing external interference features from video surveillance images, including an image acquisition module 10, a first virtual infrared surveillance image generation module 20, a first virtual optical surveillance image generation module 30, a second virtual optical surveillance image generation module 40, a second virtual infrared surveillance image generation module 50, a cyclic link formation module 60, a loss function construction module 70, and a feature image acquisition module 80 without external interference.

[0117] The image acquisition module 10 is used to acquire a set of optical surveillance images and an infrared surveillance image set. The optical surveillance images in the optical surveillance image set and the infrared surveillance images in the infrared surveillance image set have overlapping areas. The optical surveillance images are feature images captured by an optical camera without external interference, and the infrared surveillance images are images captured by an infrared camera. The first virtual infrared surveillance image generation module 20 is used to input the optical surveillance images and the infrared surveillance images into an optical-infrared transformation model to generate a first virtual infrared surveillance image. The first virtual optical surveillance image generation module 30 is used to input the first virtual infrared surveillance image and the optical surveillance image into the infrared-optical transformation model to generate a first virtual optical surveillance image.

[0118] The second virtual optical monitoring image generation module 40 is configured to input the infrared monitoring image and the optical monitoring image into the infrared optical transformation model to generate a second virtual optical monitoring image. The second virtual infrared monitoring image generation module 50 is configured to input the second virtual optical monitoring image and the infrared monitoring image into the optical-infrared transformation model to generate a second virtual infrared monitoring image. The loop link formation module 60 is configured to form a first loop link with the optical monitoring image, the infrared monitoring image, the first virtual infrared monitoring image, and the first virtual optical monitoring image, and to form a second loop link with the infrared monitoring image, the optical monitoring image, the second virtual optical monitoring image, and the second virtual infrared monitoring image.

[0119] The loss function construction module 70 is used to construct the loss function of the image processing model based on the loss function of the optical infrared transformation model, the loss function of the infrared optical transformation model, the loss function of the first loop link, and the loss function of the second loop link. The interference-free feature image acquisition module 80 is used to train and optimize the image processing model based on the total loss function to obtain a trained image processing model, and then process the infrared monitoring image corresponding to the optical monitoring image to be measured based on the trained image processing model to obtain a interference-free feature image.

[0120] In this embodiment, the relevant description of the image acquisition module 10 may refer to the relevant description of S10 in the above embodiment. The relevant description of the first virtual infrared monitoring image generation module 20 may refer to the relevant description of S20 in the above embodiment. The relevant description of the first virtual optical monitoring image generation module 30 may refer to the relevant description of S30 in the above embodiment. The relevant description of the second virtual optical monitoring image generation module 40 may refer to the relevant description of S40 in the above embodiment. The relevant description of the second virtual infrared monitoring image generation module 50 may refer to the relevant description of S50 in the above embodiment. The relevant description of the loop link formation module 60 may refer to the relevant description of S60 in the above embodiment. The relevant description of the loss function construction module 70 may refer to the relevant description of S70 in the above embodiment. The relevant description of the feature image acquisition module 80 without external interference may refer to the relevant description of S80 in the above embodiment.

[0121] In one embodiment, the device 100 for removing external interference features of video surveillance images also includes an adjacent virtual infrared surveillance image acquisition module (not marked in the figure), a predicted infrared surveillance image acquisition module (not marked in the figure) and an average virtual infrared surveillance image acquisition module (not marked in the figure).

[0122] The adjacent virtual infrared surveillance image acquisition module is used to input the first two adjacent optical surveillance images of the optical surveillance image and the first two adjacent infrared surveillance images of the infrared surveillance image into the optical infrared transformation model to obtain the first two adjacent virtual infrared surveillance images of the first virtual infrared surveillance image. The predicted infrared surveillance image acquisition module is used to input the first two adjacent virtual infrared surveillance images into the infrared surveillance prediction model for prediction to obtain a predicted infrared surveillance image. The average virtual infrared surveillance image acquisition module is used to average the predicted infrared surveillance image with the first virtual infrared surveillance image to obtain an average virtual infrared surveillance image.

[0123] In this embodiment, the description of the adjacent virtual infrared surveillance image acquisition module can refer to the description of S201 in the above embodiment. The description of the predicted infrared surveillance image acquisition module can refer to the description of S202 in the above embodiment. The description of the averaged virtual infrared surveillance image acquisition module can refer to the description of S203 in the above embodiment.

[0124] In one embodiment, the first virtual optical monitoring image generation module 30 includes a first acquisition module (not shown). The first acquisition module is configured to input the average virtual infrared monitoring image and the optical monitoring image into an infrared optical transformation model to generate a first virtual optical monitoring image. For a description of the first acquisition module in this embodiment, reference can be made to the description of S310, S311, and S312 in the above embodiment.

[0125] In one embodiment, the apparatus 100 for removing external interference features of a video surveillance image further includes an adjacent virtual optical surveillance image acquisition module, a predicted optical surveillance image acquisition module, and an average virtual optical surveillance image acquisition module.

[0126] The adjacent virtual optical monitoring image acquisition module is used to input the first two adjacent infrared monitoring images of the infrared monitoring image and the first two adjacent optical monitoring images of the optical monitoring image into the infrared optical transformation model to obtain the first two adjacent virtual optical monitoring images of the second virtual optical monitoring image. The predicted optical monitoring image acquisition module is used to input the first two adjacent virtual optical monitoring images into the optical monitoring prediction model for prediction to obtain a predicted optical monitoring image. The average virtual optical monitoring image acquisition module is used to average the predicted optical monitoring image and the second virtual optical monitoring image to obtain an average virtual optical monitoring image.

[0127] In this embodiment, the description of the adjacent virtual optical monitoring image acquisition module can refer to the description of S401 in the above embodiment. The description of the predicted optical monitoring image acquisition module can refer to the description of S402 in the above embodiment. The description of the averaged virtual optical monitoring image acquisition module can refer to the description of S403 in the above embodiment.

[0128] In one embodiment, the second virtual infrared surveillance image generation module 50 includes a second acquisition module (not shown). The second acquisition module is configured to input the averaged virtual optical surveillance image and the infrared surveillance image into the optical-infrared transformation model to generate the second virtual infrared surveillance image. For a description of the second acquisition module in this embodiment, reference can be made to the description of S50, S511, and S512 in the above embodiment.

[0129] In one embodiment, the interference-free characteristic image acquisition module 80 includes a training and optimization module (not shown) and a third acquisition module (not shown). The training and optimization module is configured to train and optimize the optical infrared transformation model, the infrared optical transformation model, the first loop link, and the second loop link based on a total loss function to obtain a trained infrared optical transformation model. The third acquisition module is configured to obtain a target infrared surveillance image corresponding to the interference-free characteristic image and input the target infrared surveillance image into the trained infrared optical transformation model to obtain the interference-free characteristic image.

[0130] In this embodiment, the description of the training optimization module can refer to the description of S810 in the above embodiment. The description of the third acquisition module can refer to the description of S820 in the above embodiment.

[0131] In one embodiment, in the image acquisition module 10, the overlap area between the optical monitoring image and the infrared monitoring image is 95% to 100%. In this embodiment, the description of the overlap area can refer to the relevant description in the above embodiment.

[0132] In the various embodiments described above, the specific order or hierarchy of steps in the disclosed processes is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged without departing from the scope of protection of this disclosure. The accompanying method claims present the elements of the various steps in an exemplary order and are not intended to be limited to the specific order or hierarchy described.

[0133] Those skilled in the art will also appreciate that the various illustrative logical blocks, units, and steps listed in the embodiments of the present application can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly demonstrate the interchangeability of hardware and software, the various illustrative components, units, and steps described above have generally described their functions. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art may use various methods to implement the described functions for each specific application, but such implementation should not be understood as exceeding the scope of protection of the embodiments of the present application.

[0134] The various illustrative logic blocks described in the embodiments of the present application, or units can be implemented or operated by the design of a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, a discrete gate or transistor logic, a discrete hardware component, or any combination thereof. The general-purpose processor can be a microprocessor, alternatively, the general-purpose processor can also be any traditional processor, controller, microcontroller or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a digital signal processor core, or any other similar configuration to implement.

[0135] The steps of the methods or algorithms described in the embodiments of the present application can be directly embedded in hardware, a software module executed by a processor, or a combination of the two. The software module can be stored in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Alternatively, the storage medium can also be integrated into the processor. The processor and the storage medium can be provided in an ASIC, which can be provided in a user terminal. Alternatively, the processor and the storage medium can also be provided in different components in the user terminal.

[0136] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of this application. It should be understood that the above description is only the specific implementation methods of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.

Claims

1. A method for removing external interference features from video surveillance images, characterized in that: include: Acquire an optical monitoring image set and an infrared monitoring image set, wherein the optical monitoring images in the optical monitoring image set and the infrared monitoring images in the infrared monitoring image set have overlapping areas, the optical monitoring images are feature images obtained by an optical camera without external interference, and the infrared monitoring images are images obtained by an infrared camera; Inputting the optical monitoring image and the infrared monitoring image into an optical-infrared conversion model to generate a first virtual infrared monitoring image; Inputting the first virtual infrared monitoring image and the optical monitoring image into an infrared optical transformation model to generate a first virtual optical monitoring image; Inputting the infrared monitoring image and the optical monitoring image into the infrared optical transformation model to generate a second virtual optical monitoring image; inputting the second virtual optical monitoring image and the infrared monitoring image into the optical-infrared conversion model to generate a second virtual infrared monitoring image; The optical monitoring image, the infrared monitoring image, the first virtual infrared monitoring image, and the first virtual optical monitoring image form a first loop link, and the infrared monitoring image, the optical monitoring image, the second virtual optical monitoring image, and the second virtual infrared monitoring image form a second loop link; Constructing a total loss function of an image processing model according to the loss function of the optical infrared transformation model, the loss function of the infrared optical transformation model, the loss function of the first loop link, and the loss function of the second loop link; Training and optimizing the image processing model according to the total loss function to obtain a trained image processing model, and processing the infrared monitoring image corresponding to the optical monitoring image to be measured according to the trained image processing model to obtain a feature image without external interference; After inputting the optical monitoring image and the infrared monitoring image into the optical-infrared transformation model to generate the first virtual infrared monitoring image, and before inputting the first virtual infrared monitoring image and the optical monitoring image into the infrared-optical transformation model to generate the first virtual optical monitoring image, the method further includes: Inputting the first two adjacent optical monitoring images of the optical monitoring image and the first two adjacent infrared monitoring images of the infrared monitoring image into the optical-infrared transformation model to obtain the first two adjacent virtual infrared monitoring images of the first virtual infrared monitoring image; Inputting the first two adjacent virtual infrared monitoring images into an infrared monitoring prediction model for prediction to obtain a predicted infrared monitoring image; averaging the predicted infrared surveillance image and the first virtual infrared surveillance image to obtain an average virtual infrared surveillance image; The step of inputting the first virtual infrared monitoring image and the optical monitoring image into an infrared optical transformation model to generate a first virtual optical monitoring image includes: The average virtual infrared monitoring image and the optical monitoring image are input into the infrared optical transformation model to generate the first virtual optical monitoring image.

2. The method for removing external interference features from video surveillance images according to claim 1, characterized in that: After inputting the infrared monitoring image and the optical monitoring image into the infrared optical transformation model to generate a second virtual optical monitoring image, and before inputting the second virtual optical monitoring image and the infrared monitoring image into the optical infrared transformation model to generate a second virtual infrared monitoring image, the method further includes: Inputting the first two adjacent infrared monitoring images of the infrared monitoring image and the first two adjacent optical monitoring images of the optical monitoring image into the infrared optical transformation model to obtain the first two adjacent virtual optical monitoring images of the second virtual optical monitoring image; Inputting the first two adjacent virtual optical monitoring images into an optical monitoring prediction model for prediction to obtain a predicted optical monitoring image; The predicted optical monitoring image and the second virtual optical monitoring image are averaged to obtain an average virtual optical monitoring image.

3. The method for removing external interference features from video surveillance images according to claim 2, characterized in that: The step of inputting the second virtual optical monitoring image and the infrared monitoring image into the optical-infrared conversion model to generate the second virtual infrared monitoring image includes: The average virtual optical monitoring image and the infrared monitoring image are input into the optical infrared conversion model to generate the second virtual infrared monitoring image.

4. The method for removing external interference features from video surveillance images according to claim 1, characterized in that: The image processing model is trained and optimized according to the total loss function to obtain a trained image processing model, and the image to be tested is processed according to the trained image processing model to obtain a feature image without external interference, including: Performing training and optimization on the optical infrared transformation model, the infrared optical transformation model, the first loop link, and the second loop link according to the total loss function to obtain a trained infrared optical transformation model; An infrared monitoring image to be measured corresponding to the image containing external interference characteristics is obtained, and the infrared monitoring image to be measured is input into the trained infrared optical transformation model to obtain the image without external interference characteristics.

5. The method for removing external interference features from video surveillance images according to claim 1, characterized in that: The optical monitoring image in the optical monitoring image set and the infrared monitoring image in the infrared monitoring image set have an overlapping area, and the overlapping area between the optical monitoring image and the infrared monitoring image is 95% to 100%.

6. A device for removing external interference features from video surveillance images, characterized in that: include: An image acquisition module, configured to acquire an optical monitoring image set and an infrared monitoring image set, wherein the optical monitoring images in the optical monitoring image set and the infrared monitoring images in the infrared monitoring image set have overlapping areas, the optical monitoring images are characteristic images obtained by an optical camera without external interference, and the infrared monitoring images are images obtained by an infrared camera; a first virtual infrared monitoring image generation module, configured to input the optical monitoring image and the infrared monitoring image into an optical infrared conversion model to generate a first virtual infrared monitoring image; A first virtual optical monitoring image generation module, configured to input the first virtual infrared monitoring image and the optical monitoring image into an infrared optical transformation model to generate a first virtual optical monitoring image; a second virtual optical monitoring image generating module, configured to input the infrared monitoring image and the optical monitoring image into the infrared optical transformation model to generate a second virtual optical monitoring image; a second virtual infrared monitoring image generating module, configured to input the second virtual optical monitoring image and the infrared monitoring image into the optical-infrared conversion model to generate a second virtual infrared monitoring image; a loop link forming module, configured to form a first loop link with the optical surveillance image, the infrared surveillance image, the first virtual infrared surveillance image, and the first virtual optical surveillance image, and to form a second loop link with the infrared surveillance image, the optical surveillance image, the second virtual optical surveillance image, and the second virtual infrared surveillance image; a loss function construction module, configured to construct a total loss function of an image processing model according to the loss function of the optical infrared transformation model, the loss function of the infrared optical transformation model, the loss function of the first loop link, and the loss function of the second loop link; a feature image acquisition module free of external interference, configured to train and optimize the image processing model according to the total loss function to obtain a trained image processing model, and to process the infrared monitoring image corresponding to the optical monitoring image to be measured according to the trained image processing model to obtain a feature image free of external interference; The device further comprises: an adjacent virtual infrared monitoring image acquisition module, configured to input the first two adjacent optical monitoring images of the optical monitoring image and the first two adjacent infrared monitoring images of the infrared monitoring image into the optical-infrared transformation model to obtain the first two adjacent virtual infrared monitoring images of the first virtual infrared monitoring image; A predicted infrared monitoring image acquisition module is used to input the first two adjacent virtual infrared monitoring images into the infrared monitoring prediction model for prediction to obtain a predicted infrared monitoring image; an average virtual infrared surveillance image acquisition module, configured to average the predicted infrared surveillance image and the first virtual infrared surveillance image to obtain an average virtual infrared surveillance image; The first virtual optical monitoring image generation module includes: The first acquisition module is configured to input the average virtual infrared monitoring image and the optical monitoring image into the infrared optical transformation model to generate the first virtual optical monitoring image.

7. The device for removing external interference features of video surveillance images according to claim 6, characterized in that: Also includes: an adjacent virtual optical monitoring image acquisition module, configured to input the first two adjacent infrared monitoring images of the infrared monitoring image and the first two adjacent optical monitoring images of the optical monitoring image into the infrared optical transformation model to obtain the first two adjacent virtual optical monitoring images of the second virtual optical monitoring image; A predicted optical monitoring image acquisition module is used to input the first two adjacent virtual optical monitoring images into the optical monitoring prediction model for prediction, thereby obtaining a predicted optical monitoring image; The average virtual optical monitoring image acquisition module is configured to average the predicted optical monitoring image and the second virtual optical monitoring image to obtain an average virtual optical monitoring image.

8. The device for removing external interference features of video surveillance images according to claim 7, characterized in that: The second virtual infrared monitoring image generation module includes: The second acquisition module is configured to input the average virtual optical monitoring image and the infrared monitoring image into the optical-infrared conversion model to generate the second virtual infrared monitoring image.

9. The device for removing external interference features of video surveillance images according to claim 6, characterized in that: The feature image acquisition module without external interference includes: a training optimization module, configured to perform training optimization on the optical-infrared transformation model, the infrared optical transformation model, the first loop link, and the second loop link according to the total loss function to obtain a trained infrared optical transformation model; The third acquisition module is used to obtain the infrared monitoring image to be tested corresponding to the feature image containing external interference, and input the infrared monitoring image to be tested into the trained infrared optical transformation model to obtain the feature image without external interference.

10. The device for removing external interference features of video surveillance images according to claim 6, characterized in that: In the image acquisition module, the overlap area between the optical monitoring image and the infrared monitoring image is 95% to 100%.