Image defogging method and system based on optical flow dynamic feature enhanced attention mechanism network
By building a network of enhanced attention mechanisms for optical flow dynamic characteristics, the problem of poor fog removal effect of surveillance video is solved, and the efficient and clear fog removal effect is achieved, the quality and safety of surveillance video are improved. It is suitable for monitoring scenarios in complex environments such as railways, airports and construction sites.
Patent Information
- Application Number
- CN202510332407.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-20
AI Technical Summary
The existing surveillance video defog removal algorithm is poor, resulting in impairment of the clarity and accuracy of the surveillance video, and the inability to effectively extract key scene features.
A network of enhanced attention mechanisms based on optical flow dynamic features is constructed, including background feature extraction module, color feature extraction module and sliding window attention mechanism module. Combined with composite loss function and gradual optimization algorithm, the monitoring video is preprocessed, atomized and trained, and high-quality defogging images are output.
It improves the fog removal effect of surveillance video in a haze environment, enhances the ability to extract key information, improves the clarity and security of surveillance video, can process in real time or near real-time, and improves the safety efficiency and accuracy of surveillance scenes.
Smart Images

Figure CN120339116A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer vision technology, and particularly relates to an image dehazing method and system based on an optical flow dynamic feature enhanced attention mechanism network. Background Art
[0002] Outdoor monitoring scenarios usually have complex environmental conditions, such as atmospheric haze, rain and snow weather, etc. These factors will greatly affect the clarity and accuracy of monitoring videos. In order to improve the working efficiency and security of the monitoring system, dehazing technology has become an important research field; dehazing technology can not only improve the video quality, but also effectively extract key scene features to assist in automatic monitoring and judgment.
[0003] Currently, the monitoring videos of existing monitoring scenarios are generally dehazed by simple dehazing algorithms. Dehazing with simple dehazing algorithms will cause damage to the monitoring videos and the dehazing effect is not good; therefore, how to achieve efficient dehazing of the monitoring videos of monitoring scenarios is a technical problem to be solved urgently. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide an image dehazing method based on an optical flow dynamic feature enhanced attention mechanism network, which can solve the technical problem that the simple dehazing algorithm in the prior art has a poor dehazing effect on monitoring videos.
[0005] To solve the above technical problem, this application is implemented as follows:
[0006] In the first aspect, the embodiments of this application provide an image dehazing method based on an optical flow dynamic feature enhanced attention mechanism network, and the method includes:
[0007] Obtain multiple segments of monitoring videos of a monitoring scenario, perform a first preprocessing on the multiple segments of monitoring videos to obtain corresponding multiple segments of preprocessed monitoring videos;
[0008] Obtain a training set and a test set according to the multiple segments of preprocessed monitoring videos, and perform a second preprocessing on the training set to obtain a preprocessed training set;
[0009] Perform fogging processing on the multiple segments of preprocessed monitoring videos to construct a fogging image flow set according to the fogged multiple segments of monitoring videos;
[0010] Construct an optical flow dynamic feature enhanced attention mechanism network, and the optical flow dynamic feature enhanced attention mechanism network includes a background feature extraction module, a color feature extraction module, and a sliding window attention mechanism module;
[0011] Construct a composite loss function, and train the optical flow dynamic feature enhanced attention mechanism network according to the preprocessed training set, the set of foggy images, and the composite loss function;
[0012] Test the trained optical flow dynamic feature enhanced attention mechanism network according to the test set, input the image to be defogged into the tested optical flow dynamic feature enhanced attention mechanism network, and output the defogged image;
[0013] Process the defogged image according to the progressive optimization algorithm to obtain the final target defogged image.
[0014] As an optional implementation manner of the first aspect of the present application, obtain multiple segments of surveillance videos of a surveillance scene, and perform first preprocessing on the multiple segments of surveillance videos to obtain corresponding multiple segments of preprocessed surveillance videos; specifically:
[0015] Obtain multiple segments of surveillance videos of a surveillance scene, and set the same resolution and frame rate for each surveillance video in the multiple segments of surveillance videos;
[0016] Perform normalization processing on the multiple segments of surveillance videos so that the time length and video format of each surveillance video in the multiple segments of surveillance videos are the same, and obtain the multiple segments of preprocessed surveillance videos.
[0017] As an optional implementation manner of the first aspect of the present application, obtain a training set and a test set according to the multiple segments of preprocessed surveillance videos, and perform second preprocessing on the training set to obtain a preprocessed training set; specifically:
[0018] Extract images from each preprocessed surveillance video in the multiple segments of preprocessed surveillance videos to obtain each real image set corresponding to each preprocessed surveillance video;
[0019] Perform random central fogging processing on each real image in each real image set to obtain each foggy image set corresponding to each real image set
[0020] Divide each foggy image set according to a preset ratio to obtain each training subset and each test subset corresponding to each foggy image set;
[0021] Extract any real image from each real image set to construct a real reference image set, construct a training set according to each training subset and the real reference image set, and construct a test set according to each test subset and the real reference image set;
[0022] Perform image scaling processing on the training set to adjust the sizes of the images in the test set to be the same, and perform random cropping processing and horizontal flipping processing on the training set after the image scaling processing to obtain the preprocessed training set.
[0023] As an alternative implementation of the first aspect of the present application, atomize the multi-segment preprocessed monitoring videos to construct an atomized image stream set according to the atomized multi-segment monitoring videos; specifically:
[0024] Atomize the multi-segment preprocessed monitoring videos according to the fog synthesis algorithm based on depth image information to obtain multi-segment atomized videos;
[0025] Randomly sample N atomized videos from the multi-segment atomized videos, and perform continuous image extraction on each atomized video in the N atomized videos to obtain each atomized image stream corresponding to each atomized video;
[0026] Construct the atomized image stream set according to each atomized image stream corresponding to each atomized video.
[0027] As an alternative implementation of the first aspect of the present application, construct a composite loss function, and train the optical flow dynamic feature enhanced attention mechanism network according to the preprocessing training set, the atomized image stream set, and the composite loss function; specifically:
[0028] Construct a composite loss function, and input the preprocessing training set and the atomized image stream set into the optical flow dynamic feature enhanced attention mechanism network for training;
[0029] Obtain the loss difference between the training result of the optical flow dynamic feature enhanced attention mechanism network for the preprocessing training set and the corresponding real image
[0030] The composite loss function optimizes the parameters of the optical flow dynamic feature enhanced attention mechanism network according to the loss difference until the optical flow dynamic feature enhanced attention mechanism network converges, and completes the training of the composite loss function for the optical flow dynamic feature enhanced attention mechanism network.
[0031] As an alternative implementation of the first aspect of the present application, input the preprocessing training set and the atomized image stream set into the optical flow dynamic feature enhanced attention mechanism network for training; specifically:
[0032] Input the atomized images in the preprocessing training set into a 3×3 convolutional layer, and after passing through a 3×3 convolutional layer and a sliding window attention mechanism module in sequence, obtain the first attention feature;
[0033] Input the same atomization image into the color feature extraction module, and obtain color features after being processed by the color feature extraction module; input the real reference image from the same preprocessed monitoring video as the atomization image in the preprocessed training set into the background feature extraction module, and obtain background features after being processed by the background feature extraction module;
[0034] Input the first attention feature and the background feature into a downsampling layer for feature fusion processing to obtain a first fusion feature, and input the first fusion feature into a sliding window attention mechanism module for processing to obtain a second attention feature;
[0035] Input the second attention feature and the color feature into a downsampling layer for feature fusion processing to obtain a second fusion feature; input the second fusion feature into a sliding window attention mechanism module and an upsampling layer in sequence for processing to obtain a third attention feature;
[0036] Input the third attention feature and the second attention feature into an SK channel attention fusion module for feature fusion processing to obtain a third fusion feature; input the third fusion feature into a sliding window attention mechanism module for processing to obtain a fourth attention feature;
[0037] Input the atomization image stream set into an optical flow network for feature extraction processing to obtain optical flow features, and input the optical flow features and the fourth attention feature into an upsampling layer for feature fusion processing to obtain a fourth fusion feature;
[0038] Input the fourth fusion feature and the first attention feature into an SK channel attention fusion module for feature fusion to obtain a fifth fusion feature, and input the fifth fusion feature into a sliding window attention mechanism module and a 3×3 convolutional layer in sequence for processing to obtain a defogged image.
[0039] As an optional implementation manner of the first aspect of the present application, the composite loss function is represented by the following formula:
[0040]
[0041] L = L1 + SSIM(x, y),
[0042] where L represents the composite loss function, x(p) represents the pixel value of the defogged image at pixel point p, y(p) represents the pixel value of the real image corresponding to the defogged image at pixel point p, N represents the total number of pixels in the image, and ∑ p∈P represents the summation over all pixel points p in the image, μ x represents the pixel mean of the defogged image, μ ydenotes the pixel mean of the ground truth image corresponding to the dehazed image, σ x denotes the standard deviation of the pixel values of the dehazed image, σ y denotes the standard deviation of the pixel values of the ground truth image corresponding to the dehazed image, σ xy denotes the pixel value covariance between the dehazed image and the corresponding ground truth image, and C1 and C2 denote constants;
[0043] The data processing process of the sliding window attention mechanism module is represented by the following formula:
[0044]
[0045] where B represents the relative position bias term, Q represents the query matrix, K represents the key matrix, V represents the value matrix, ||Q|| represents the L2 norm of Q, ||K|| represents the L2 norm of K, and β represents the learnable scaling factor.
[0046] where B represents the relative position bias term, Q represents the query matrix, K represents the key matrix, V represents the value matrix, ||Q|| represents the L2 norm of Q, ||K|| represents the L2 norm of K, and λ represents the learnable scaling factor.
[0047] In a second aspect, an image dehazing system based on an optical flow dynamic feature enhanced attention mechanism network is provided in an embodiment of the present application. The system includes:
[0048] The first acquisition module: acquires multiple segments of monitoring videos of a monitoring scene, performs first preprocessing on the multiple segments of monitoring videos, and obtains corresponding multiple segments of preprocessed monitoring videos;
[0049] The second acquisition module: obtains a training set and a test set according to the multiple segments of preprocessed monitoring videos, and performs second preprocessing on the training set to obtain a preprocessed training set;
[0050] The first processing module: performs fogging processing on the multiple segments of preprocessed monitoring videos to construct a fogged image flow set according to the fogged multiple segments of monitoring videos;
[0051] The construction module: constructs an optical flow dynamic feature enhanced attention mechanism network, and the optical flow dynamic feature enhanced attention mechanism network includes a background feature extraction module, a color feature extraction module, and a sliding window attention mechanism module;
[0052] The training module: constructs a composite loss function, and trains the optical flow dynamic feature enhanced attention mechanism network according to the preprocessed training set, the fogged image flow set, and the composite loss function;
[0053] Testing module: Test the trained optical flow dynamic feature enhanced attention mechanism network according to the test set, input the image to be dehazed into the tested optical flow dynamic feature enhanced attention mechanism network, and output the dehazed image;
[0054] Second processing module: Process the dehazed image according to the progressive optimization algorithm to obtain the final target dehazed image.
[0055] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0056] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0057] In the embodiment of the present application, compared with the prior art, the following beneficial effects are achieved:
[0058] (1) By constructing an optical flow dynamic feature enhanced attention mechanism network, key information in the scene monitoring video, especially optical flow features and color features, can be more effectively extracted and utilized, thereby enhancing the adaptability to the foggy environment; the introduction of the composite loss function makes the network training more accurate, can minimize the difference between the dehazed result and the real image, and improve the dehazing effect.
[0059] (2) Performing the first preprocessing on the monitoring video and the second preprocessing on the training set. The first preprocessing includes setting the same resolution and frame rate, and standardization processing. The second preprocessing includes image scaling, random cropping, and horizontal flipping, etc., which helps to improve the network's adaptability to different scenes and monitoring conditions; by random center fogging processing and constructing a set of fogged image streams, the diversity and complexity of the training data are increased, further enhancing the generalization performance of the network.
[0060] (3) The sliding window attention mechanism module in the optical flow dynamic feature enhanced attention mechanism network can pay more attention to the feature regions with obvious fogging in the input features, thereby achieving a more refined dehazing effect; the collaborative effect of the background feature extraction module and the color feature extraction module enables the network to more accurately identify and process the feature elements in the monitoring scene.
[0061] (4) The application of the progressive optimization algorithm further improves the quality of the dehazed image, making the dehazed image clearer and more natural; the introduction of the SK channel attention fusion module enhances the network's attention allocation ability for different feature channels, which helps to improve the detail performance and overall visual effect of the dehazed image.
[0062] (5) This method can perform defogging processing on the monitoring video of the monitoring scene in real time or near real time, which helps to improve the efficiency and accuracy of scene monitoring security; the defogged monitoring video can provide a clearer field of view, which helps the monitoring personnel to detect and handle potential security hazards in a timely manner. Description of the Drawings
[0063] Figure 1 is a flowchart of an image defogging method based on an optical flow dynamic feature enhanced attention mechanism network provided by some embodiments of the present application;
[0064] Figure 2 is a structural diagram of an optical flow dynamic feature enhanced attention mechanism network of an image defogging method based on an optical flow dynamic feature enhanced attention mechanism network provided by some embodiments of the present application;
[0065] Figure 3 is a structural diagram of a background feature extraction module of an image defogging method based on an optical flow dynamic feature enhanced attention mechanism network provided by some embodiments of the present application;
[0066] Figure 4 is a structural diagram of a color feature extraction module of an image defogging method based on an optical flow dynamic feature enhanced attention mechanism network provided by some embodiments of the present application;
[0067] Figure 5 is a defogging effect diagram of an image defogging method based on an optical flow dynamic feature enhanced attention mechanism network provided by some embodiments of the present application;
[0068] Figure 6 is a schematic diagram of a monitoring device of an image defogging method based on an optical flow dynamic feature enhanced attention mechanism network provided by some embodiments of the present application. Detailed Embodiments
[0069] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0070] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0071] The following will combine the accompanying drawings and through specific embodiments and their application scenarios, a method and system for image defogging based on an optical flow dynamic feature enhanced attention mechanism network provided by the embodiments of this application will be described in detail.
[0072] Refer to Figure 1 As shown, a method for image defogging based on an optical flow dynamic feature enhanced attention mechanism network includes the following steps:
[0073] S100: Obtain multiple segments of surveillance videos of the surveillance scene, perform a first preprocessing on the multiple segments of surveillance videos to obtain corresponding multiple segments of preprocessed surveillance videos;
[0074] It should be noted that the scene surveillance includes railway operation scenes, airport operation scenes, construction site work scenes, etc. In this embodiment, the surveillance scene is described by taking the railway operation scene as an example;
[0075] It should be noted that S100 is specifically:
[0076] S110: Obtain multiple segments of surveillance videos of the railway operation scene according to railway surveillance, and set the same resolution and frame rate for each surveillance video in the multiple segments of surveillance videos;
[0077] S120: Perform a normalization process on the multiple segments of surveillance videos to make the time length and video format of each surveillance video in the multiple segments of surveillance videos the same, and obtain multiple segments of preprocessed surveillance videos.
[0078] Furthermore, in S110, the multiple segments of surveillance videos obtained are processed to make them into a unified resolution of 1080*720, with 30 frames per second; then in S120, the time length of each segment of surveillance video is adjusted to 40 minutes, and the format is set to the mp4 format to obtain multiple segments of preprocessed surveillance videos, and the multiple segments of preprocessed surveillance videos are named video(1), video(2),..., video(n) in sequence;
[0079] S200: Obtain a training set and a test set according to the multiple segments of preprocessed surveillance videos, and perform a second preprocessing on the training set to obtain a preprocessed training set;
[0080] It should be noted that S200 is specifically:
[0081] S210: Extract images from each preprocessed monitoring video in multiple segments of preprocessed monitoring videos to obtain each set of real images corresponding to each preprocessed monitoring video;
[0082] S220: Perform random central fogging processing on each real image in each set of real images to obtain each set of fogged images corresponding to each set of real images;
[0083] S230: Divide each set of fogged images according to a preset ratio to obtain each training subset and each test subset corresponding to each set of fogged images;
[0084] S240: Extract any real image from each set of real images to construct a real reference image set, construct a training set according to each training subset and the real reference image set, and construct a test set according to each test subset and the real reference image set;
[0085] S250: Perform image scaling processing on the training set to adjust the sizes of the images in the test set to be the same, and perform random cropping processing and horizontal flipping processing on the training set after image scaling processing to obtain a preprocessed training set.
[0086] Further, in S210, image extraction processing is performed on each of the preprocessed surveillance videos Video(1), Video(2), …, Video(n). One real image is extracted every 2 seconds, and 1,200 real images are extracted from each preprocessed surveillance video. Specifically, the real images corresponding to the preprocessed surveillance video Video(1) are Video(1-1), Video(1-2), …, Video(1-1200), the real images corresponding to the preprocessed surveillance video Video(2) are Video(2-1), Video(2-2), …, Video(2-1200), and similarly, the real images corresponding to the preprocessed surveillance video Video(n) are Video(n-1), Video(n-2), …, Video(n-1200); for each preprocessed surveillance video, the 1,200 images extracted from the preprocessed surveillance video constitute the corresponding real image sets GT(1), GT(2), …, GT(n) of the preprocessed surveillance video; specifically, the real image set GT(1) is composed of Video(1-1), Video(1-2), …, Video(1-1200), the real image set GT(2) is composed of Video(2-1), Video(2-2), …, Video(2-1200), and similarly, the real image set GT(n) is composed of Video(n-1), Video(n-2), …, Video(n-1200); in S220, random central fogging processing is performed on each real image in GT(1), GT(2), …, GT(n) to obtain the corresponding fogged image sets Hazy(1), Hazy(2), …, Hazy(n); where Hazy(1) includes Hazy(1-1), Hazy(1-2), …, Hazy(1-1200), Hazy(2) includes Hazy(2-1), Hazy(2-2), …, Hazy(2-1200), and similarly, Hazy(n) includes Hazy(n-1), Hazy(n-2), …, Hazy(n-1200);The Hazy(1-1), Hazy(1-2), …, Hazy(1-1200) in Hazy(1) correspond one-to-one with the Video(1-1), Video(1-2), …, Video(1-1200) in GT(1), the Hazy(2-1), Hazy(2-2), …, Hazy(2-1200) in Hazy(2) correspond one-to-one with the Video(2-1), Video(2-2), …, Video(2-1200) in GT(2), and similarly, the Hazy(n-1), Hazy(n-2), …, Hazy(n-1200) in Hazy(n) correspond one-to-one with the Video(n-1), Video(n-2), …, Video(n-1200) in GT(n); in S230, according to a preset ratio, which is 8:2 in this embodiment, the foggy image sets Hazy(1), Hazy(2), …, Hazy(n) are sequentially divided to obtain each training subset and each test subset corresponding to each foggy image set. Specifically, the foggy image set Hazy(1) corresponds to the training subset Train(1) and the test subset Test(1), the foggy image set Hazy(2) corresponds to the training subset Train(2) and the test subset Test(2), and similarly, the foggy image set Hazy(n) corresponds to the training subset Train(n) and the test subset Test(n). Among them, the ratio of foggy images in Train(1) and Test(1) is 8:2, the ratio of foggy images in Train(2) and Test(2) is 8:2, and similarly, the ratio of foggy images in Train(2) and Test(2) is 8:2; in S240, in the real image sets GT(1), GT(2), …, GT(n), any one picture is randomly extracted as the real reference images Re(1), Re(2), …, Re(n) corresponding to the real image sets. All the extracted real reference images Re(1), Re(2), …, Re(n) form the real reference image set Re; the training set Train is formed by the training subsets Train(1), Train(2), …, Train(n) and the real reference image set Re, and the test set Test is formed by the test subsets Test(1), Test(2), …, Test(n) and the real reference image set Re; in S250, all the images in the training set Train are scaled to make the sizes of all the images the same, and then all the pictures are randomly cropped and horizontally flipped to prevent overfitting, obtaining the preprocessed training set, where the preprocessed training set is composed of the training subsets Train(1), Train(2), …, Train(n) and the reference image set Re after the second preprocessing (scaling, random cropping and horizontal flipping).;
[0087] S300: Perform fogging processing on multiple segments of pre - processed surveillance videos to construct a set of fogged image streams based on the fogged multiple segments of surveillance videos;
[0088] It should be noted that S300 specifically is:
[0089] S310: Perform fogging processing on multiple segments of pre - processed surveillance videos according to the fog synthesis algorithm based on depth image information to obtain multiple segments of fogged videos;
[0090] S320: Randomly sample N segments of fogged videos from the multiple segments of fogged videos, and perform continuous image extraction on each fogged video among the N segments of fogged videos to obtain each fogged image stream corresponding to each fogged video;
[0091] S330: Construct a set of fogged image streams according to each fogged image stream corresponding to each fogged video.
[0092] Furthermore, in S310, perform fogging processing on multiple segments of pre - processed surveillance videos Video(1), Video(2), …, Video(n) according to the fog synthesis algorithm based on depth image information to obtain the corresponding multiple segments of fogged videos WH(1), WH(2), …, WH(n); in S320, randomly sample N segments of fogged videos from the n segments of fogged videos WH(1), WH(2), …, WH(n), where N < n; perform continuous image extraction on each fogged video among the N segments of fogged videos to obtain each fogged image stream corresponding to each fogged video among the N segments of fogged videos; in S330, construct a set of fogged image streams according to the N fogged image streams corresponding to the N segments of fogged videos obtained in S320.
[0093] S400: Construct an optical flow dynamic feature enhanced attention mechanism network, and the optical flow dynamic feature enhanced attention mechanism network includes a background feature extraction module, a color feature extraction module, and a sliding window attention mechanism module;
[0094] It should be noted that, referring to Figure 2 、 Figure 3 and Figure 4 as shown, the background feature extraction module is used to extract background features in the railway scene, and these features are crucial for distinguishing the foreground (such as trains, personnel) and the background (such as tracks, sky); the color feature extraction module is used to extract color information in the image, and fog usually causes the loss or blurring of color information, so the extraction of color features helps to restore the image details affected by fog; the sliding window attention mechanism module uses the optical flow dynamic feature enhanced attention mechanism to apply the attention mechanism to the image in a sliding window manner, enabling the network to focus on important local regions in the image, thereby improving the defogging effect.
[0095] S500: Construct a composite loss function and train the optical flow dynamic feature enhanced attention mechanism network based on the preprocessed training set, the set of atomized image streams, and the composite loss function;
[0096] It should be noted that S500 is specifically as follows:
[0097] S510: Construct a composite loss function and input the preprocessed training set and the set of atomized image streams into the optical flow dynamic feature enhanced attention mechanism network for training;
[0098] S520: Obtain the loss difference between the training result of the optical flow dynamic feature enhanced attention mechanism network for the preprocessed training set and the corresponding real image;
[0099] S530: The composite loss function optimizes the parameters of the optical flow dynamic feature enhanced attention mechanism network according to the loss difference until the optical flow dynamic feature enhanced attention mechanism network converges, completing the training of the optical flow dynamic feature enhanced attention mechanism network by the composite loss function.
[0100] It should be noted that in S510, inputting the preprocessed training set and the set of atomized image streams into the optical flow dynamic feature enhanced attention mechanism network for training is specifically as follows:
[0101] S511: Input the atomized image in the preprocessed training set into a 3×3 convolutional layer, and after passing through a 3×3 convolutional layer and a sliding window attention mechanism module in sequence, obtain the first attention feature;
[0102] S512: Input the same atomized image into the color feature extraction module, and obtain the color feature after being processed by the color feature extraction module; input the real reference image from the same preprocessed monitoring video as the atomized image in the preprocessed training set into the background feature extraction module, and obtain the background feature after being processed by the background feature extraction module;
[0103] S513: Input the first attention feature and the background feature into a downsampling layer for feature fusion processing to obtain the first fusion feature, and input the first fusion feature into a sliding window attention mechanism module for processing to obtain the second attention feature;
[0104] S514: Input the second attention feature and the color feature into a downsampling layer for feature fusion processing to obtain the second fusion feature; input the second fusion feature into a sliding window attention mechanism module and an upsampling layer in sequence for processing to obtain the third attention feature;
[0105] S515: Input the third attention feature and the second attention feature into an SK channel attention fusion module for feature fusion processing to obtain a third fusion feature; input the third fusion feature into a sliding window attention mechanism module for processing to obtain a fourth attention feature;
[0106] S516: Input the atomization image stream set into an optical flow network for feature extraction processing to obtain an optical flow feature, and input the optical flow feature and the fourth attention feature into an upsampling layer for feature fusion processing to obtain a fourth fusion feature;
[0107] S517: Input the fourth fusion feature and the first attention feature into an SK channel attention fusion module for feature fusion to obtain a fifth fusion feature, and after sequentially passing the fifth fusion feature through a sliding window attention mechanism module and a 3×3 convolutional layer, obtain a defogged image;
[0108] Further, assume that the hazy image Hazy(1-2) belongs to the training subset Train(1); in S511, the hazy image Hazy(1-2) after the second preprocessing in the preprocessed training set is input into a 3×3 convolutional layer. After passing through a 3×3 convolutional layer and a sliding window attention mechanism module in sequence, a first attention feature is obtained. The depth of the first attention feature is 24, the length is 720, and the width is 1080; in S512, the hazy image Hazy(1-2) after the second preprocessing is input into a color feature extraction module. After being processed by the color feature extraction module, a color feature is obtained. The network depth of the color feature is 12, the length is 180, and the width is 270; the real reference image from the same preprocessed monitoring video as the hazy image Hazy(1-2) in the preprocessed training set is input into the background feature extraction module. Since the hazy image Hazy(1-2) comes from the preprocessed monitoring video Video(1), and the real reference image corresponding to the preprocessed monitoring video Video(1) is Re(1), the real reference image Re(1) is input into the background feature extraction module. After being processed by the background feature extraction module, a background feature is obtained. The depth of the background feature is 48, the length is 360, and the width is 540; in S513, the first attention feature and the background feature are input into a downsampling layer for feature fusion processing to obtain a first fusion feature. The depth of the first fusion feature is 48, the length is 360, and the width is 540; the first fusion feature is input into a sliding window attention mechanism module for processing to obtain a second attention feature. The depth of the second attention feature is 48, the length is 360, and the width is 540; in S514, the second attention feature and the color feature are input into a downsampling layer for feature fusion processing to obtain a second fusion feature. The depth of the second fusion feature is 96, the length is 180, and the width is 270; the second fusion feature is input into a sliding window attention mechanism module and an upsampling layer in sequence for processing to obtain a third attention feature. The depth of the third attention feature is 48, the length is 360, and the width is 540; in S515, the third attention feature and the second attention feature are input into an SK channel attention fusion module for feature fusion processing to obtain a third fusion feature. The depth of the third fusion feature is 48, the length is 360, and the width is 540; the third fusion feature is input into a sliding window attention mechanism module for processing to obtain a fourth attention feature. The depth of the fourth attention feature is 48, the length is 360, and the width is 540; in S516, the hazy image stream set is input into an optical flow network for feature extraction processing to obtain an optical flow feature. The depth of the optical flow feature is 48, the length is 360, and the width is 540; the optical flow feature and the fourth attention feature are input into an upsampling layer for feature fusion processing to obtain a fourth fusion feature. The depth of the fourth fusion feature is 24, the length is 720, and the width is 1080;In S517, the fourth fusion feature and the first attention feature are input into an SK channel attention fusion module for feature fusion to obtain a fifth fusion feature. The depth of the fifth fusion feature is 24, the length is 720, and the width is 1080. After passing the fifth fusion feature through a sliding window attention mechanism module and a 3×3 convolutional layer in sequence, a defogged image is obtained.
[0109] It should be noted that before inputting the foggy image stream set into the optical flow network for feature extraction processing in S516, the optical flow network has already been trained. That is, when constructing the optical flow dynamic feature enhanced attention mechanism network, the optical flow network used is an already trained optical flow network with the best weights.
[0110] It should be noted that in S520, the loss difference between the training result of the optical flow dynamic feature enhanced attention mechanism network on the preprocessing training set and the corresponding real image is obtained. Specifically:
[0111] Suppose the optical flow dynamic feature enhanced attention mechanism network is trained by the foggy image Hazy(1-2) in S510, and the defogged image QHazy(1-2) is output. The difference between the defogged image QHazy(1-2) and the corresponding real image value is obtained. Since Hazy(1-2) comes from the random central fogging of the real image Video(1-2), the real image corresponding to the defogged image QHazy(1-2) is Video(1-2). Therefore, that is, the loss difference between the defogged image QHazy(1-2) and the real image Video(1-2) is obtained.
[0112] It should be noted that in S530, the composite loss function optimizes the parameters of the optical flow dynamic feature enhanced attention mechanism network according to the loss difference until the optical flow dynamic feature enhanced attention mechanism network converges, completing the training of the composite loss function on the optical flow dynamic feature enhanced attention mechanism network. Specifically:
[0113] Suppose the optical flow dynamic feature enhanced attention mechanism network is trained by the foggy image Hazy(1-2) in S510 and the defogged image is output. Therefore, in S530, the composite loss function optimizes the parameters of the optical flow dynamic feature enhanced attention mechanism network according to the loss difference between the defogged image QHazy(1-2) and the real image Video(1-2) obtained in S520. Through iterative optimization of the training process, until the optical flow dynamic feature enhanced attention mechanism network converges, completing the training of the composite loss function on the optical flow dynamic feature enhanced attention mechanism network.
[0114] It should be noted that the composite loss function is represented by the following formula:
[0115]
[0116] L = L1 + SSIM(x, y),
[0117] where L represents the composite loss function, x(p) represents the pixel value of the defogged image at pixel point p, y(p) represents the pixel value of the corresponding real image of the defogged image at pixel point p, N represents the total number of pixels in the image, and ∑ p∈P denotes the summation over all pixel points p in the image, μ x represents the mean pixel value of the defogged image, μ y represents the mean pixel value of the corresponding real image of the defogged image, σ x represents the standard deviation of the pixel values of the defogged image, σ y represents the standard deviation of the pixel values of the corresponding real image of the defogged image, σ xy represents the covariance of the pixel values of the defogged image and the corresponding real image, and C1 and C2 represent constants.
[0118] The data processing process of the sliding window attention mechanism module is represented by the following formula:
[0119]
[0120] where B represents the relative position bias term, Q represents the query matrix, K represents the key matrix, V represents the value matrix, ||Q|| represents the L2 norm of Q, ||K|| represents the L2 norm of K, and β represents the learnable scaling factor.
[0121] S600: Test the trained optical flow dynamic feature enhanced attention mechanism network according to the test set. Input the image to be defogged into the optical flow dynamic feature enhanced attention mechanism network that passes the test, and output the defogged image;
[0122] It should be noted that the trained network is tested using the test set to verify its defogging effect; after passing the test; the railway operation scene monitoring video to be defogged is used to extract defogged images at 15 frames per second, obtaining a set of consecutive images to be defogged, and this set of consecutive images to be defogged is sequentially input into the optical flow dynamic feature enhanced attention mechanism network that passes the test, and a set of consecutive defogged images is output. Refer to Figure 5 As shown, it is a defogged image obtained after processing one of the images to be defogged.
[0123] S700: Process the defogged image according to the progressive optimization algorithm to obtain the final target defogged image;
[0124] It should be noted that in S700, refer to Figure 6As shown, considering the characteristics of surveillance cameras, most camera directions are slightly tilted downward. In the real-time video image, the pixel values of the y-axis increase from top to bottom, and the pixel values of the x-axis increase from left to right. The smaller the y-axis pixel, the farther the real-time detection distance. The greater the distance, the greater the thickness of the fog, and the greater the impact of the fog.
[0125] Therefore, a progressive optimization algorithm is designed to optimize the obtained defogged image to obtain a more accurate target defogged image. The progressive optimization algorithm is expressed by the following formula:
[0126]
[0127] T(x) = t(x) + δ · f(t(x)),
[0128] where δ min represents the minimum value of the position change factor, δ represents the position change factor, H represents the height of the defogged image, y represents the position of the current pixel point, T(x) represents the target defogged image, t(x) represents the defogged image, and f(t(x)) represents image enhancement processing of the defogged image.
[0129] An image dehazing method based on an optical flow dynamic feature enhanced attention mechanism network according to this embodiment can more effectively extract and utilize key information in railway scene monitoring videos, especially optical flow features and color features, thereby enhancing the adaptability to foggy environments through constructing an optical flow dynamic feature enhanced attention mechanism network; the introduction of a composite loss function makes network training more accurate, capable of minimizing the difference between the dehazing result and the real image and improving the dehazing effect; the first preprocessing of the monitoring video and the second preprocessing of the training set, where the first preprocessing includes setting the same resolution and frame rate and normalization processing, and the second preprocessing includes image scaling, random cropping, and horizontal flipping, etc., helps to improve the network's adaptability to different railway scenes and monitoring conditions; through random central fogging processing and constructing a set of fogged images, the diversity and complexity of training data are increased, further enhancing the network's generalization performance; the sliding window attention mechanism module in the optical flow dynamic feature enhanced attention mechanism network can pay more attention to the feature regions with obvious fogging in the input features, thus achieving a more refined dehazing effect; the collaborative effect of the background feature extraction module and the color feature extraction module enables the network to more accurately identify and process feature elements in railway scenes; the application of the progressive optimization algorithm further improves the quality of the dehazed image, making the dehazed image clearer and more natural; the introduction of the SK channel attention fusion module enhances the network's attention distribution ability for different feature channels, helping to improve the detail performance and overall visual effect of the dehazed image; this method can dehaze the monitoring video of railway scenes in real-time or near real-time, helping to improve the efficiency and accuracy of railway safety monitoring; the dehazed monitoring video can provide a clearer view, helping the monitoring personnel to discover and handle potential safety hazards in time.
[0130] It should be noted that for an image dehazing method based on an optical flow dynamic feature enhanced attention mechanism network provided in the embodiment of this application, the execution subject can be an image dehazing system based on an optical flow dynamic feature enhanced attention mechanism network, or a control module in the image dehazing system based on an optical flow dynamic feature enhanced attention mechanism network for executing and loading an image dehazing method based on an optical flow dynamic feature enhanced attention mechanism network. In the embodiment of this application, taking an image dehazing system based on an optical flow dynamic feature enhanced attention mechanism network for executing and loading an image dehazing method based on an optical flow dynamic feature enhanced attention mechanism network as an example, an image dehazing method based on an optical flow dynamic feature enhanced attention mechanism network provided in the embodiment of this application is described.
[0131] An image dehazing system based on an optical flow dynamic feature enhanced attention mechanism network includes:
[0132] The first acquisition module: acquires multiple segments of surveillance videos of a surveillance scenario, performs first preprocessing on the multiple segments of surveillance videos to obtain corresponding multiple segments of preprocessed surveillance videos;
[0133] The second acquisition module: obtains a training set and a test set according to the multiple segments of preprocessed surveillance videos, and performs second preprocessing on the training set to obtain a preprocessed training set;
[0134] The first processing module: performs fogging processing on the multiple segments of preprocessed surveillance videos to construct a fogging image stream set according to the fogged multiple segments of surveillance videos;
[0135] The construction module: constructs an optical flow dynamic feature enhanced attention mechanism network, and the optical flow dynamic feature enhanced attention mechanism network includes a background feature extraction module, a color feature extraction module, and a sliding window attention mechanism module;
[0136] The training module: constructs a composite loss function, and trains the optical flow dynamic feature enhanced attention mechanism network according to the preprocessed training set, the fogging image stream set, and the composite loss function;
[0137] The testing module: tests the trained optical flow dynamic feature enhanced attention mechanism network according to the test set, inputs the image to be defogged into the qualified optical flow dynamic feature enhanced attention mechanism network after testing, and outputs a defogged image;
[0138] The second processing module: processes the defogged image according to the progressive optimization algorithm to obtain a final target defogged image.
[0139] An image defogging system based on an optical flow dynamic feature enhanced attention mechanism network in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a notebook computer, a handheld computer, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), etc. The embodiments of the present application do not make specific limitations.
[0140] An image defogging system based on an optical flow dynamic feature enhanced attention mechanism network in an embodiment of the present application can be a device with an operating system. The operating system can be the Android operating system, the iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.
[0141] An image defogging system based on an optical flow dynamic feature enhanced attention mechanism network provided in an embodiment of the present application can implement Figures 1 to 6 each process implemented by an image defogging method based on an optical flow dynamic feature enhanced attention mechanism network in the method embodiment. To avoid repetition, it will not be elaborated here.
[0142] According to an image defogging system based on an optical flow dynamic feature enhanced attention mechanism network in this embodiment, by preprocessing multi-surveillance videos of a railway scene through a first acquisition module, the quality and consistency of video data can be improved, laying a good foundation for subsequent processing steps. This helps reduce noise, enhance contrast, or perform other necessary image enhancement operations, thereby improving the defogging effect. The second acquisition module obtains a training set and a test set according to multiple preprocessed surveillance videos, ensuring that the training and testing of the model are carried out on independent data sets, which helps improve the generalization ability of the model and avoid overfitting. The first processing module performs fogging processing on the preprocessed surveillance videos to construct a set of fogged image flows. This method of simulating a real foggy environment helps the model learn defogging strategies under different fogging degrees and improves the adaptability of the model in practical applications. The optical flow dynamic feature enhanced attention mechanism network in the construction module can more accurately capture key information in the railway scene through a background feature extraction module, a color feature extraction module, and a sliding window attention mechanism module. This mechanism helps the model pay more attention to important visual cues when processing complex scenes, thereby improving the accuracy and efficiency of defogging. The training module trains the optical flow dynamic feature enhanced attention mechanism network by constructing a composite loss function; this method of multi-objective optimization helps improve the overall performance of the model and enables it to achieve the best defogging effect while meeting multiple requirements. The testing module tests the trained model to ensure the stability and reliability of the model in practical applications. By inputting the image to be defogged into the network that passes the test, a high-quality defogged image can be output to meet the requirements of practical applications. The second processing module further processes the defogged image using a progressive optimization algorithm to obtain the final target defogged image. This algorithm can further improve the quality of the image through step-by-step iteration and optimization, making the defogging effect more natural and clear.
[0143] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-mentioned embodiment of the image dehazing method based on the optical flow dynamic feature enhanced attention mechanism network, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0144] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above-mentioned embodiment of the image dehazing method based on the optical flow dynamic feature enhanced attention mechanism network, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0145] Wherein, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0146] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described method may be executed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0147] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0148] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. An image defogging method based on an attention mechanism network that enhances optical flow dynamic features, characterized in that The method includes: Obtain multiple surveillance videos of a surveillance scenario, perform a first preprocessing on the multiple surveillance videos to obtain corresponding multiple preprocessed surveillance videos; Obtain a training set and a test set according to the multiple preprocessed surveillance videos, and perform a second preprocessing on the training set to obtain a preprocessed training set; Perform fogging processing on the multiple preprocessed surveillance videos to construct a fogged image stream set according to the fogged multiple surveillance videos; Construct an optical flow dynamic feature enhanced attention mechanism network, which includes a background feature extraction module, a color feature extraction module, and a sliding window attention mechanism module; Construct a composite loss function, and train the optical flow dynamic feature enhanced attention mechanism network according to the preprocessed training set, the fogged image stream set, and the composite loss function; Test the trained optical flow dynamic feature enhanced attention mechanism network according to the test set, input the image to be de-fogged into the optical flow dynamic feature enhanced attention mechanism network that passes the test, and output a de-fogged image; Process the de-fogged image according to the progressive optimization algorithm to obtain the final target de-fogged image.
2. The image defogging method based on an optical flow dynamic feature enhanced attention mechanism network according to claim 1, wherein, The obtaining multiple surveillance videos of a surveillance scenario, performing a first preprocessing on the multiple surveillance videos to obtain corresponding multiple preprocessed surveillance videos; Specifically: Obtain multiple surveillance videos of a surveillance scenario, and set the same resolution and frame rate for each surveillance video in the multiple surveillance videos; Perform a normalization process on the multiple surveillance videos so that the time length and video format of each surveillance video in the multiple surveillance videos are the same, to obtain the multiple preprocessed surveillance videos.
3. The image dehazing method based on an optical flow dynamic feature enhanced attention mechanism network according to claim 2, characterized in that, The obtaining a training set and a test set according to the multiple preprocessed surveillance videos, and performing a second preprocessing on the training set to obtain a preprocessed training set; Specifically: Perform image extraction on each preprocessed surveillance video in the multiple preprocessed surveillance videos to obtain each corresponding true image set for each preprocessed surveillance video; Perform random central fogging processing on each true image in each true image set to obtain each corresponding fogged image set for each true image set; Divide each fogged image set according to a preset ratio to obtain each corresponding training subset and each test subset for each fogged image set; Extract any one true image from each true image set to construct a true reference image set, construct a training set according to each training subset and the true reference image set, and construct a test set according to each test subset and the true reference image set; Perform image scaling processing on the training set to adjust the images in the test set to have the same size, and perform random cropping processing and horizontal flipping processing on the training set after the image scaling processing to obtain the preprocessed training set.
4. The image defogging method based on the optical flow dynamic feature enhanced attention mechanism network according to claim 3, characterized in that, The performing fogging processing on the multiple preprocessed surveillance videos to construct a fogged image stream set according to the fogged multiple surveillance videos; Specifically: Perform fogging processing on the multiple preprocessed surveillance videos according to the depth image information synthesis fog algorithm to obtain multiple fogged videos; Randomly sample N atomization videos from the multiple atomization videos, and perform continuous image extraction on each atomization video in the N atomization videos to obtain each atomization image stream corresponding to each atomization video; Construct the atomization image stream set according to each atomization image stream corresponding to each atomization video.
5. A method for image dehazing based on an optical flow dynamic feature enhanced attention mechanism network according to claim 4, characterized in that Construct the composite loss function, and train the optical flow dynamic feature enhanced attention mechanism network according to the preprocessed training set, the atomization image stream set and the composite loss function; specifically: Construct the composite loss function, and input the preprocessed training set and the atomization image stream set into the optical flow dynamic feature enhanced attention mechanism network for training; Obtain the loss difference between the training result of the optical flow dynamic feature enhanced attention mechanism network on the preprocessed training set and the corresponding real image The composite loss function optimizes the parameters of the optical flow dynamic feature enhanced attention mechanism network according to the loss difference until the optical flow dynamic feature enhanced attention mechanism network converges, and completes the training of the composite loss function on the optical flow dynamic feature enhanced attention mechanism network.
6. The image defogging method based on an optical flow dynamic feature enhanced attention mechanism network according to claim 5, characterized in that The inputting the preprocessed training set and the atomization image stream set into the optical flow dynamic feature enhanced attention mechanism network for training; specifically: Input the atomization image in the preprocessed training set into a 3×3 convolutional layer, and after passing through a 3×3 convolutional layer and a sliding window attention mechanism module in sequence, obtain the first attention feature; Input the same atomization image into the color feature extraction module, and obtain the color feature after being processed by the color feature extraction module; Input the real reference image in the preprocessed training set that comes from the same preprocessed monitoring video as the atomization image into the background feature extraction module, and obtain the background feature after being processed by the background feature extraction module; Input the first attention feature and the background feature into a downsampling layer for feature fusion processing to obtain the first fusion feature, and input the first fusion feature into a sliding window attention mechanism module for processing to obtain the second attention feature; Input the second attention feature and the color feature into a downsampling layer for feature fusion processing to obtain the second fusion feature; input the second fusion feature into a sliding window attention mechanism module and an upsampling layer in sequence for processing to obtain the third attention feature; Input the third attention feature and the second attention feature into an SK channel attention fusion module for feature fusion processing to obtain the third fusion feature; input the third fusion feature into a sliding window attention mechanism module for processing to obtain the fourth attention feature; Input the atomization image stream set into the optical flow network for feature extraction processing to obtain the optical flow feature, and input the optical flow feature and the fourth attention feature into an upsampling layer for feature fusion processing to obtain the fourth fusion feature; Input the fourth fusion feature and the first attention feature into an SK channel attention fusion module for feature fusion to obtain a fifth fusion feature. After sequentially processing the fifth fusion feature through a sliding window attention mechanism module and a 3×3 convolutional layer, a defogged image is obtained.
7. An image defogging method based on an optical flow dynamic feature enhanced attention mechanism network according to claim 6, characterized in that, The composite loss function is represented by the following formula: L = L1 + SSIM(x, y), Among them, L represents the composite loss function, x(p) represents the pixel value of the dehazed image at pixel point p, y(p) represents the pixel value of the corresponding real image of the dehazed image at pixel point p, N represents the total number of pixels in the image, and ∑ p∈P denotes the summation over all pixel points p in the image, μ x represents the mean pixel value of the dehazed image, μ y represents the mean pixel value of the corresponding real image of the dehazed image, σ x represents the standard deviation of the pixel values of the dehazed image, σ y represents the standard deviation of the pixel values of the corresponding real image of the dehazed image, σ xy denotes the covariance of the pixel values of the dehazed image and the corresponding real image, and C1 and C2 represent constants; The data processing process of the sliding window attention mechanism module is represented by the following formula: where B represents the relative position bias term, Q represents the query matrix, K represents the key matrix, V represents the value matrix, ||Q|| represents the L2 norm of Q, ||K|| represents the L2 norm of K, and β represents a learnable scaling factor.
8. An image dehazing system based on an optical flow dynamic feature enhanced attention mechanism network, which can implement an image dehazing method based on an optical flow dynamic feature enhanced attention mechanism network according to any one of claims 1-7, characterized in that, The system includes: The first acquisition module: Acquire multiple segments of surveillance videos of a surveillance scene, perform first preprocessing on the multiple segments of surveillance videos to obtain corresponding multiple segments of preprocessed surveillance videos; The second acquisition module: Obtain a training set and a test set according to the multiple segments of preprocessed surveillance videos, and perform second preprocessing on the training set to obtain a preprocessed training set; The first processing module: Perform fogging processing on the multiple segments of preprocessed surveillance videos to construct a fogged image stream set according to the fogged multiple segments of surveillance videos; The construction module: Construct an optical flow dynamic feature enhanced attention mechanism network, and the optical flow dynamic feature enhanced attention mechanism network includes a background feature extraction module, a color feature extraction module, and a sliding window attention mechanism module; The training module: Construct a composite loss function, and train the optical flow dynamic feature enhanced attention mechanism network according to the preprocessed training set, the fogged image stream set, and the composite loss function; The testing module: Test the trained optical flow dynamic feature enhanced attention mechanism network according to the test set, input the image to be defogged into the optical flow dynamic feature enhanced attention mechanism network that passes the test, and output a defogged image; The second processing module: Process the defogged image according to the progressive optimization algorithm to obtain a final target defogged image.
9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of an image defogging method based on an optical flow dynamic feature enhanced attention mechanism network according to any one of claims 1-7 are implemented.
10. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of an image defogging method based on an optical flow dynamic feature enhanced attention mechanism network according to any one of claims 1-7 are implemented.
Citation Information
Patent Citations
De-blurring convolutional neural network training method and device, equipment and storage medium
CN114240764A
Visual defogging method and system based on three-branch neural network
CN118521508A
Video smoke saliency target detection method based on weak supervised learning
CN119206160A
Infrared image deblurring method and system for unmanned aerial vehicle platform
CN119359592A
Video defogging method and system based on enhanced alignment
CN119477749A