A linkage control system based on video surveillance and its control method
By improving the YOLOv8 model and multi-scale feature extraction technology, the abnormal detection problem of traditional video surveillance systems under noise interference is solved, and efficient linkage control and intelligent response are achieved.
Patent Information
- Application Number
- CN202510634683.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-16
AI Technical Summary
When traditional video surveillance systems face high-frequency noise and burst pulse noise, it is difficult to accurately capture the operating details of water conservancy facilities, resulting in missed or false alarms of abnormal detection algorithms, and the system linkage intelligent control efficiency is inefficient.
The improved YOLOv8 model is used to replace the Conv modules of the first, fifth and seventh layers as the downsampling fusion module, combining Laplace pyramid and multi-scale feature extraction, and linkage control is carried out through the target recognition and abnormal detection model, and tracking tags and alarm information are generated for linkage control.
It improves the accuracy of abnormal detection and the efficiency of system linkage control, can identify suspicious behaviors in real time and automatically generate alarms, reduce manual misjudgment, and improves response speed and intelligent control capabilities.
Smart Images

Figure CN120151488B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of water conservancy video monitoring, and particularly relates to a linkage control system based on video monitoring and its control method. Background Art
[0002] With the rapid development of technology, the demand for efficient and intelligent monitoring and linkage control in the field of water conservancy video monitoring is becoming increasingly urgent. In the field of video monitoring, application programs use computer vision technology to observe people, objects, abnormal states and other activities in a large number of video sequences, which plays an important role in water conservancy video monitoring work. The effective elements in video data include river channels, river surface hydrological elements, floating objects on the river surface, ships, special species, drowning persons, abnormal water quality states, riverbanks, intruding persons, vehicles, etc.; however, most traditional video monitoring systems operate independently, and there is a lack of an intelligent linkage mechanism among video monitoring subsystems such as management station houses, river channels and river surfaces, and riverbanks. Once an abnormal event occurs, it relies on manual judgment and manual operation of each subsystem, which not only has a slow response, but also has a very low coordination efficiency among subsystems. At the same time, there is a lack of continuous recording and pre-judgment and early warning of abnormal events in terms of time and space.
[0003] Patent CN116320292B discloses a water conservancy monitoring and control system based on big data, including a water conservancy monitoring real-time acquisition system, a water conservancy monitoring video reading and analysis system, a water conservancy monitoring data decision system, a grid-connected alarm system and a data elimination and encryption system. By compressing video frames of video frame data and then reorganizing the compressed video frames, it can effectively reduce the number of video frames and prevent instability caused by too large number of frames, ensuring that the videos in the water conservancy monitoring areas collected by multiple collection devices in each water conservancy monitoring collection area can well ensure synchronization. By matching audio data with video frames and then processing the noise in the video frames, it can effectively analyze the operation conditions of water conservancy in the video. By comparing the water flow sound and the video picture synchronously, the accuracy of the analyzed data is improved. However, in terms of video noise, although the system processes the noise, some high-frequency noises and burst pulse noises still interfere with the capture of the operation details of water conservancy facilities in the video, resulting in the masking of key features, affecting the judgment of normal and abnormal states by the abnormal detection algorithm, and causing missed reports or false reports in abnormal detection, resulting in low efficiency of intelligent linkage control among systems. Summary of the Invention
[0004] The object of the present invention is to solve the problem that high-frequency noises and burst pulse noises still interfere with the capture of the operation details of water conservancy facilities in the video, resulting in the masking of key features, affecting the judgment of normal and abnormal states by the abnormal detection algorithm, causing missed reports or false reports in abnormal detection, and resulting in low efficiency of intelligent linkage control among systems, and to propose a linkage control system based on video monitoring and its control method.
[0005] In the first aspect of the implementation of the present invention, a linkage control method based on video surveillance is first proposed, and the method includes:
[0006] Obtain a continuous image sequence uploaded by cameras in the target area; the continuous image sequence is obtained by frame extraction after the cameras collect video data;
[0007] Perform target recognition on each image in the continuous image sequence to obtain a set of valid bounding boxes;
[0008] Substitute all the valid bounding boxes into a preset feature extraction model to obtain the detection features corresponding to each bounding box, and check whether the detection features exist by searching a preset database;
[0009] If the detection features exist, obtain the corresponding label ID, and obtain the movement trajectory according to the coordinate sequence of the label ID in each camera;
[0010] Substitute the movement trajectory into an anomaly detection model to obtain an anomaly result. If the anomaly result indicates an anomaly, generate a tracking label for the bounding box, issue an alarm to generate alarm information, and perform linkage control on security devices according to the tracking label and the alarm information.
[0011] Optionally, performing target recognition on each image in the continuous image sequence to obtain a set of valid bounding boxes includes:
[0012] For each image in the continuous image sequence, preprocess the image to obtain a target detection image;
[0013] Substitute the target detection image into a preset target recognition model to obtain a set of valid bounding boxes;
[0014] The preset target recognition model is an improved YOLOv8 model; replace the Conv modules in the first layer, fifth layer, and seventh layer of the YOLOv8 model with downsampling fusion modules;
[0015] The working principle of the downsampling fusion module is as follows:
[0016] The downsampling fusion module includes a first module and a second module; the first module includes obtaining an input feature map, performing average pooling and convolution operations on the feature map to obtain a first output feature map; performing channel dimension splitting on the first output feature map to obtain 4 second output feature maps; applying the Softmax function to all the second output feature maps to obtain 4 sets of feature weights;
[0017] The second module includes performing convolution operations on the input feature map to obtain a third output feature map; performing channel dimension splitting on the third output feature map to obtain 4 fourth output feature maps;
[0018] Weighted fusion operations are respectively performed on the four fourth output feature maps according to the four groups of feature weights to obtain the final output feature map.
[0019] Optionally, substituting all valid bounding boxes into the preset feature extraction model to obtain the detection features corresponding to each bounding box includes:
[0020] For each valid bounding box, extracting the image region corresponding to the valid bounding box to obtain the initial input image;
[0021] Substituting the initial input image into the Laplacian pyramid to obtain the first-scale feature, the second-scale feature, and the third-scale feature; the first-scale feature is the high-scale feature; the second-scale feature is the medium-scale feature; the third-scale feature is the low-scale feature;
[0022] Substituting the first-scale feature map into the first enhanced feature extraction module to obtain the first enhanced feature; substituting the first enhanced feature, the second-scale feature, and the third-scale feature into the second enhanced feature extraction module to obtain the second enhanced feature;
[0023] Performing feature fusion on the second-scale feature and the third-scale feature respectively with the second enhanced feature to obtain the first fusion feature and the second fusion feature;
[0024] Performing scale normalization on the first enhanced feature, the first fusion feature, and the second fusion feature and then performing feature fusion to obtain the detection features.
[0025] Optionally, the first enhanced feature extraction module includes:
[0026] Performing channel decomposition on the first-scale feature map to obtain the first channel feature and the second channel feature;
[0027] Substituting the first channel feature and the second channel feature into the full-channel extraction module respectively to obtain the first full-channel fusion feature and the second full-channel fusion feature;
[0028] Performing splicing on the first full-channel fusion feature and the second full-channel fusion feature to obtain the first enhanced feature.
[0029] Optionally, the second enhanced feature extraction module includes:
[0030] Performing splicing on the first enhanced feature, the second-scale feature, and the third-scale feature to obtain the first spliced feature;
[0031] Passing the first spliced feature through the multi-scale dilation attention module and the multi-head attention module respectively to obtain the first attention feature and the second attention feature;
[0032] The first attention feature and the second attention feature are superimposed and fused to obtain a second enhanced feature.
[0033] In the second aspect of the implementation of the present invention, a linkage control system based on video surveillance is proposed, including:
[0034] A frame extraction processing module, configured to obtain a continuous image sequence uploaded by a camera within a target area; the continuous image sequence is obtained by performing frame extraction processing on the video data collected by the camera;
[0035] A target recognition module, configured to perform target recognition on each image in the continuous image sequence to obtain a set of effective bounding boxes;
[0036] A detection feature generation module, configured to substitute all effective bounding boxes into a preset feature extraction model to obtain a detection feature corresponding to each bounding box, and determine whether the detection feature exists by searching a preset database;
[0037] A motion trajectory determination module, configured to, if the detection feature exists, obtain the corresponding label ID, and obtain a motion trajectory according to the coordinate sequence of the label ID in each camera;
[0038] A linkage control module, configured to substitute the motion trajectory into an anomaly detection model to obtain an anomaly result. If the anomaly result indicates an anomaly, generate a tracking label for the bounding box, issue an alarm to generate an alarm message, and perform linkage control on security devices according to the tracking label and the alarm message.
[0039] Optionally, the target recognition module includes:
[0040] A preprocessing module, configured to, for each image in the continuous image sequence, perform preprocessing on the image to obtain a target detection image;
[0041] A bounding box determination module, configured to substitute the target detection image into a preset target recognition model to obtain a set of effective bounding boxes;
[0042] A model improvement module, where the preset target recognition model is an improved YOLOv8 model; replace the Conv modules in the first layer, fifth layer, and seventh layer of the YOLOv8 model with downsampling fusion modules;
[0043] The working principle of the downsampling fusion module is:
[0044] The downsampling fusion module includes a first module and a second module; the first module includes obtaining an input feature map, performing average pooling and convolution operations on the feature map to obtain a first output feature map; performing channel dimension splitting on the first output feature map to obtain 4 second output feature maps; applying the Softmax function to all the second output feature maps to obtain 4 sets of feature weights;
[0045] The second module includes: performing a convolution operation on the input feature map to obtain a third output feature map; performing channel dimension splitting on the third output feature map to obtain four fourth output feature maps;
[0046] Performing weighted fusion operations on the four fourth output feature maps according to four groups of feature weights to obtain a final output feature map.
[0047] Optionally, the detection feature generation module includes:
[0048] A region extraction module, configured to extract the image region corresponding to each valid bounding box to obtain an initial input image;
[0049] A multi-scale feature extraction module, configured to substitute the initial input image into a Laplacian pyramid to obtain a first-scale feature, a second-scale feature, and a third-scale feature; the first-scale feature is a high-scale feature; the second-scale feature is a medium-scale feature; the third-scale feature is a low-scale feature;
[0050] A feature enhancement module, configured to substitute the first-scale feature map into a first enhanced feature extraction module to obtain a first enhanced feature; substituting the first enhanced feature, the second-scale feature, and the third-scale feature into a second enhanced feature extraction module to obtain a second enhanced feature;
[0051] A first feature fusion module, configured to perform feature fusion on the second-scale feature and the third-scale feature respectively with the second enhanced feature to obtain a first fusion feature and a second fusion feature;
[0052] A second feature fusion module, configured to perform scale normalization on the first enhanced feature, the first fusion feature, and the second fusion feature and then perform feature fusion to obtain a detection feature.
[0053] Optionally, the first enhanced feature extraction module includes:
[0054] A channel decomposition module, configured to perform channel decomposition on the first-scale feature map to obtain a first channel feature and a second channel feature;
[0055] A channel fusion feature extraction module, configured to substitute the first channel feature and the second channel feature into a full-channel extraction module respectively to obtain a first full-channel fusion feature and a second full-channel fusion feature;
[0056] A full-channel fusion module, configured to splice the first full-channel fusion feature and the second full-channel fusion feature to obtain a first enhanced feature.
[0057] Optionally, the second enhanced feature extraction module includes:
[0058] A scale feature splicing module, configured to splice the first enhanced feature, the second scale feature, and the third scale feature to obtain a first spliced feature;
[0059] An attention feature extraction module, configured to respectively pass the first spliced feature through a multi-scale dilation attention module and a multi-head attention module to obtain a first attention feature and a second attention feature;
[0060] An attention feature superposition module, configured to superpose and fuse the first attention feature and the second attention feature to obtain a second enhanced feature.
[0061] Advantages of the present invention:
[0062] The present invention proposes a linkage control method based on video surveillance, which acquires a continuous image sequence uploaded by a camera in a target area, performs target recognition on each image in the continuous image sequence to obtain a set of effective bounding boxes; substitutes all effective bounding boxes into a preset feature extraction model to obtain a detection feature corresponding to each bounding box, and determines whether the detection feature exists by searching a preset database; if the detection feature exists, the corresponding label ID is obtained, and a motion trajectory is obtained according to the coordinate sequence of the label ID in each camera; the motion trajectory is substituted into an anomaly detection model to obtain an anomaly result, if the anomaly result shows an anomaly, a tracking label is generated for the bounding box, and an alarm is issued to generate alarm information, and the security equipment is controlled in a linkage manner according to the tracking label and the alarm information. The continuous image sequence is monitored and analyzed in real time, and the feature detection efficiency is improved by first detecting the effective bounding box and then performing feature extraction. Combined with the anomaly detection model, suspicious behaviors can be effectively identified. Once an anomaly is detected, the system automatically generates a tracking label and alarm information, thereby performing linkage control and improving the linkage intelligent control efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] The present invention will be further described below with reference to the accompanying drawings.
[0064] Figure 1 is a flowchart of a linkage control method based on video surveillance provided by an embodiment of the present invention;
[0065] Figure 2 is a schematic diagram of camera linkage control provided by an embodiment of the present invention;
[0066] Figure 3 is a schematic diagram of a downsampling fusion module provided by an embodiment of the present invention;
[0067] Figure 4 is a framework diagram of a linkage control system based on video surveillance provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0069] All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0070] The embodiments of the present invention provide a linkage control method based on video monitoring. Refer to Figure 1 , Figure 1 which is a flowchart of a linkage control method based on video monitoring provided by the embodiments of the present invention. The method includes the following steps:
[0071] S101, obtaining a continuous image sequence uploaded by a camera in a target area;
[0072] S102, performing target recognition on each image in the continuous image sequence to obtain a set of effective bounding boxes;
[0073] S103, substituting all the effective bounding boxes into a preset feature extraction model to obtain detection features corresponding to each bounding box, and checking whether the detection features exist by querying a preset database according to the detection features;
[0074] S104, if the detection features exist, obtaining the corresponding label ID, and obtaining a motion trajectory according to the coordinate sequence of the label ID in each camera;
[0075] S105, substituting the motion trajectory into an anomaly detection model to obtain an anomaly result. If the anomaly result indicates an anomaly, generating a tracking label for the bounding box, issuing an alarm to generate alarm information, and performing linkage control on security devices according to the tracking label and the alarm information;
[0076] Among them, the continuous image sequence is obtained by frame extraction after the camera collects video data.
[0077] Based on the linkage control method based on video monitoring provided by the embodiments of the present invention, real-time monitoring and analysis are performed on the continuous image sequence, and feature detection efficiency is improved by first detecting effective bounding boxes and then performing feature extraction. Combining with the anomaly detection model, suspicious behaviors can be effectively identified. Once an anomaly is detected, the system automatically generates a tracking label and alarm information, thereby performing linkage control and improving the efficiency of linkage intelligent control.
[0078] In one implementation, the effective elements in the video data include river channels, river surface hydrological elements, floating objects on the river surface, boats, special species, persons falling into the water, abnormal water quality states, river embankments, intruding persons, vehicles, etc.
[0079] In one implementation, the security devices can be cameras, warning lights, loudspeakers, etc.; the associated control of the security devices can be the movement control of the cameras, the flashing of the warning lights, the alarm of the loudspeakers, etc.; for example, see Figure 2 , Figure 2 which is a schematic diagram of the associated control of the cameras, where circles 1, 2, 3, and 4 respectively represent the cameras, the dashed arrows represent the water flow direction, d is the floating object, such as Figure 2 when the A camera 4 in detects the floating object, at this time the system will start to track the movement trajectory of the target in real time according to its position and tracking label, the camera 4 will continuously track the floating object d, and control the monitoring view of the adjacent camera 3 to move in the direction of the floating object d. As the floating object d moves along the water flow direction, as shown in Figure 2 B in, so that the detection object always appears within the monitoring range of the camera. Similarly, the cameras 1 and 2 will also perform tracking operations; when the camera 4 cannot detect the floating object d, the view of the camera 4 will return to the default angle; the default angle is determined by the technical personnel; if the abnormal result is that no abnormality occurs, no operation will be performed.
[0080] In one implementation, through real-time monitoring and analysis of the continuous image sequence, combined with the abnormal detection model, it can effectively identify suspicious behaviors (persons falling into the water, abnormal floating objects, etc.), automatically identify abnormal situations, improve the response speed, and reduce manual misjudgment; the target area is the monitoring area composed of all cameras.
[0081] In one implementation, frame extraction processes the video stream at a preset fixed frame rate (such as 25fps), the built-in edge computing node extracts frames from the video, and generates a continuous image sequence in chronological order; the edge computing node performs lightweight preprocessing on the images (such as normalizing the size to 640×480 pixels), and uploads the key frames to the central control cloud in real time through the 5G network, and the ordinary frames are transmitted and backed up through the optical fiber during the idle period.
[0082] In one implementation, the target area is the surveillance area composed of all cameras. The preset database is constructed as follows: based on the locations of all cameras, the surveillance area is constructed, and a two-dimensional coordinate system is constructed according to the surveillance area. The location of each camera and the monitoring area are displayed in the two-dimensional coordinate system. When an object that first appears in the surveillance area is detected, its features are collected and recorded in the database, and a corresponding label ID is generated. The feature record includes the coordinates in the two-dimensional coordinate system, the recording time, the recorded features, and the label ID. When it is recognized that the object leaves the detection area, it will be deleted from the database. The anomaly detection model is a commonly used model in the market for anomaly judgment based on running estimation. The preset database stores the features corresponding to different label IDs.
[0083] In one implementation, through target recognition and feature comparison, the label ID can be confirmed in the preset database, coordinate association can be performed among multiple cameras, a complete motion trajectory can be generated, and cross-camera personnel tracking can be achieved. Once an anomaly is detected, the system automatically generates a tracking label and alarm information, and immediately triggers security equipment (such as warning lights, broadcasts, etc.).
[0084] In one embodiment, the effective bounding box set obtained by performing target recognition on each image in the continuous image sequence includes:
[0085] For each image in the continuous image sequence, the image is preprocessed to obtain a target detection image;
[0086] The target detection image is substituted into a preset target recognition model to obtain an effective bounding box set;
[0087] The preset target recognition model is an improved YOLOv8 model; the Conv modules in the first layer, fifth layer, and seventh layer of the YOLOv8 model are replaced with downsampling fusion modules;
[0088] The working principle of the downsampling fusion module is as follows:
[0089] The downsampling fusion module includes a first module and a second module; the first module includes obtaining an input feature map, performing average pooling and convolution operations on the feature map to obtain a first output feature map; performing channel dimension segmentation on the first output feature map to obtain 4 second output feature maps; applying the Softmax function to all the second output feature maps to obtain 4 sets of feature weights;
[0090] The second module includes performing a convolution operation on the input feature map to obtain a third output feature map; performing channel dimension segmentation on the third output feature map to obtain 4 fourth output feature maps;
[0091] Weighted fusion operations are respectively performed on the 4 fourth output feature maps according to the 4 sets of feature weights to obtain a final output feature map.
[0092] In one implementation, refer to Figure 3 , Figure 3 which is a schematic diagram of the downsampling fusion module. Among them, the four second output feature maps are F1, F2, F3, and F4 respectively. The four sets of feature weights obtained by passing the four second output feature maps through the Softmax function are w1, w2, w3, and w4 respectively. The four fourth output feature maps are F11, F12, F13, and F14 respectively. The channels corresponding to F1 and F11 are the same, the channels corresponding to F2 and F12 are the same, the channels corresponding to F3 and F13 are the same, and the channels corresponding to F4 and F14 are the same. The convolution module is a convolution kernel with a size of 3×3 and a stride of 2. The pooling module has a pooling kernel size of 2×2, a stride of 1, and a padding of 0. Then the final output feature map F = w1 * F11 + w2 * F12 + w3 * F13 + w4 * F14.
[0093] In one implementation, the preprocessing is common technical means such as size adjustment, equal-proportion scaling and padding, color space conversion, and normalization, which are used to make the input image more suitable for the target detection model.
[0094] In one implementation, the overall spatial semantic features are extracted through average pooling and convolution to enhance the model's perception of the context. Channel-by-channel processing and weighted fusion enable the model to focus on more distinguishable regions and improve the accuracy of target perception. The preprocessing is the same as the preprocessing method in "performing preprocessing on the image to obtain a target detection image".
[0095] In one implementation, the downsampling fusion module adopts a dual-branch structure. The first module can capture the weight relationship between different channels of the feature map through average pooling and convolution operations, as well as subsequent channel splitting and Softmax processing, and extract the global information and inter-channel correlation features of the feature map. The convolution and channel splitting operations of the second module can extract features from different scales and perspectives. The feature fusion of the two modules enables the model to obtain richer and more diverse feature information and improve the feature extraction ability for targets.
[0096] In one implementation, traditional convolutional downsampling is prone to losing fine-grained feature information. The replaced module, through a fusion mechanism, weights the feature map of the second module according to the feature weights generated by the first module, can dynamically adjust according to the feature importance during the fusion process, reduce information loss during downsampling, and is more beneficial for the detection of small targets or targets with rich details.
[0097] In one embodiment, substituting all valid bounding boxes into the preset feature extraction model to obtain the detection features corresponding to each bounding box includes:
[0098] For each valid bounding box, extracting the image region corresponding to the valid bounding box to obtain an initial input image;
[0099] Substitute the initial input image into the Laplacian pyramid to obtain the first-scale feature, the second-scale feature, and the third-scale feature; the first-scale feature is a high-scale feature; the second-scale feature is a medium-scale feature; the third-scale feature is a low-scale feature;
[0100] Substitute the first-scale feature map into the first enhanced feature extraction module to obtain the first enhanced feature; substitute the first enhanced feature, the second-scale feature, and the third-scale feature into the second enhanced feature extraction module to obtain the second enhanced feature;
[0101] Use the second enhanced feature to perform feature fusion on the second-scale feature and the third-scale feature respectively to obtain the first fusion feature and the second fusion feature;
[0102] Perform scale normalization on the first enhanced feature, the first fusion feature, and the second fusion feature, and then perform feature fusion to obtain the detection feature.
[0103] In one implementation, different-scale features are obtained through the Laplacian pyramid (the first-scale feature is the L0 layer and belongs to the high-scale feature; the second-scale feature is the L1 layer and belongs to the medium-scale feature; the third-scale feature is the L2 layer and belongs to the high-scale feature), which can enable the model to capture information about targets of different sizes in the image; small-scale feature maps are helpful for detecting small targets because they retain more detailed information; large-scale feature maps are more helpful for detecting large targets because they have a broader field of view and semantic information, which can improve the model's detection ability for targets of different sizes and enhance the model's generalization ability.
[0104] In one implementation, feature fusion is to add all the features together and take the average; scale normalization is to normalize the scales of the first enhanced feature, the first fusion feature, and the second fusion feature to be the same as the scale of the input initial input image.
[0105] In one implementation, the first enhanced feature extraction module processes the first-scale feature map to obtain the first enhanced feature, and the second enhanced feature extraction module combines the first enhanced feature, the second-scale feature, and the third-scale feature to obtain the second enhanced feature. This hierarchical enhanced feature extraction method can further excavate and strengthen the useful features in the image, remove noise and irrelevant information, improve the quality and expression ability of the features, and thus improve the accuracy of target detection.
[0106] In one implementation, use the second enhanced feature to perform feature fusion on the second-scale feature and the third-scale feature respectively to obtain the first fusion feature and the second fusion feature. Feature fusion can integrate the advantages of different-scale features, combine high-level semantic information and low-level detailed information, enable the model to better understand the image content, and more accurately locate and identify targets.
[0107] In one implementation, the first enhanced feature, the first fusion feature, and the second fusion feature are subjected to scale normalization and then feature fusion is performed to obtain the detection feature. Scale normalization ensures that features of different scales are comparable during fusion, avoiding feature fusion bias caused by scale differences. The final feature fusion can integrate the feature information extracted and processed at each stage, forming a more comprehensive and discriminative detection feature, improving the accuracy and robustness of object detection.
[0108] In one embodiment, the first enhanced feature extraction module includes:
[0109] Channel decomposition is performed on the first-scale feature map to obtain a first channel feature and a second channel feature;
[0110] The first channel feature and the second channel feature are respectively input into the full-channel extraction module to obtain a first full-channel fusion feature and a second full-channel fusion feature;
[0111] The first full-channel fusion feature and the second full-channel fusion feature are concatenated to obtain the first enhanced feature.
[0112] In one implementation, channel decomposition is performed on the first enhanced feature to obtain a first channel feature and a second channel feature, which can separate different channel information originally mixed in one feature representation. Different channels often encode different aspects of the image information, such as color, texture, shape, etc. Through decomposition, these information can be processed and analyzed separately more carefully, so as to more deeply mine the feature information of the image, which helps to improve the model's understanding ability of the image.
[0113] In one implementation, the full-channel extraction module includes 4 convolutional layers (convolutional layer 1 (convolution kernel size is 3x3, stride 1, padding 1), convolutional layer 2 (convolution kernel size is 3x3, stride 1, padding 1), convolutional layer 3 (convolution kernel size is 7x7, stride 1, padding 1), and convolutional layer 4 (convolution kernel size is 7x7, stride 1, padding 1)). The input feature passes through 4 convolutional layers respectively to obtain a set of convolutional images (convolutional feature 1, convolutional feature 2, convolutional feature 3, and convolutional feature 4); convolutional feature 1 passes through the activation function ReLU to obtain weight 1, and according to weight 1, feature fusion is performed on convolutional feature 1 and convolutional feature 2 to obtain fusion feature 1; fusion feature 1 passes through the activation function ReLU to obtain weight 2, and according to weight 2, feature fusion is performed on fusion feature 1 and convolutional feature 3 to obtain fusion feature 2; fusion feature 2 passes through the activation function ReLU to obtain weight 3, and according to weight 3, feature fusion is performed on fusion feature 2 and convolutional feature 4 to obtain the full-channel fusion feature, where feature fusion of two features according to the weight is to multiply the weight by the previous feature, the value of 1 minus the weight, multiply the latter feature, and finally add them to obtain the full-channel fusion feature.
[0114] In one implementation, after different channel features are processed by the full-channel extraction module, they may have advantages in different aspects. The splicing operation can combine these advantages to form a more comprehensive and powerful feature representation, which helps the model to detect and identify targets more accurately. Through this operation of channel decomposition and fusion, the model can process and learn features from multiple perspectives, enhancing its adaptability to different image variations (such as illumination changes, noise interference, etc.), thereby improving the robustness of the model. Even in a complex image environment, it can more stably extract and utilize effective feature information, reducing the occurrence of false detections and missed detections.
[0115] In one embodiment, the second enhanced feature extraction module includes:
[0116] Splice the first enhanced feature, the second scale feature, and the third scale feature to obtain a first spliced feature;
[0117] Pass the first spliced feature through a multi-scale dilated attention module and a multi-head attention module respectively to obtain a first attention feature and a second attention feature;
[0118] Overlay and fuse the first attention feature and the second attention feature to obtain a second enhanced feature.
[0119] In one implementation, feature information from different sources and at different scales is integrated together. Different features may contain semantic information and detailed information at different levels of the image. The splicing operation enables the model to utilize these diverse information simultaneously, enriching the expression content of the features and providing a more comprehensive input for subsequent processing.
[0120] In one implementation, the first spliced feature will first undergo a 1×1 convolution and then enter the multi-scale dilated attention module. The multi-scale dilated attention module constructs attention calculation branches using dilated convolutions with different dilation rates, and each branch independently calculates attention weights. For the output of each dilated convolution branch, an attention weight matrix is calculated, where the dot-product attention mechanism is adopted. The multi-scale dilated attention calculation result is added element-wise to the module input (i.e., the feature after Convs) to obtain a residual feature. The result after residual connection is layer-normalized and then fed into a feed-forward neural network layer to obtain the output feature, where the feed-forward neural network layer consists of two fully-connected layers, with a non-linear activation function (such as ReLU) used in the middle. The first fully-connected layer increases the input feature dimension (such as increasing it to 4 times the original), and after being activated by ReLU, the second fully-connected layer restores the dimension. The output of the feed-forward neural network is residually connected to the input of this layer, and layer normalization is performed again to output the final feature.
[0121] In one implementation, the first concatenated feature is first subjected to a 1×1 convolution and then feature extraction and channel compression are performed to obtain Q, K, and V. Q, K, and V are divided into h heads (h = 8) along the channel dimension. Each head processes partial channel features, reducing the computational load and capturing feature relationships in different subspaces. Attention calculation is independently performed for each head, and the dot product attention mechanism is adopted. The results of the multi-head attention calculation are added element-wise to the module input (the feature after being processed by Conv and PE), preserving the original input information. After layer normalization of the result of the element-wise addition, multi-head features are obtained. The multi-head features are substituted into the second feed-forward neural network layer. The second feed-forward neural network layer is similar to the feed-forward neural network layer and consists of two fully connected layers with a ReLU activation function in the middle. The first layer increases the dimension, and the second layer reduces the dimension to achieve non-linear transformation and enhanced expression of features. The output of the feed-forward neural network is connected to the input of this layer through a residual connection, and layer normalization is performed again to output the final features.
[0122] In one implementation, features are weighted at different scales to focus on important information in different-sized regions of the image. In this way, the model can, according to the size of the target and the distribution of features, specifically highlight the parts that are more critical for the detection or recognition task, suppress the interference of irrelevant information, improve the effectiveness and discriminability of features, and is particularly significant for the detection and recognition of multi-scale targets.
[0123] In one implementation, the first concatenated feature passes through the multi-head attention module to obtain the second attention feature. The multi-head attention mechanism can capture the relationships between features from multiple different representation subspaces, and each head can learn feature interaction information in different aspects. This helps the model to more comprehensively understand the associations between features, discover more complex semantic relationships, and thus enhance the expressive ability of features.
[0124] Based on the same inventive concept, the embodiments of the present invention also provide a linkage control system based on video surveillance. Refer to Figure 4 , Figure 4 which is the framework diagram of a linkage control system based on video surveillance provided by the embodiments of the present invention, including:
[0125] A frame extraction processing module, configured to obtain a continuous image sequence uploaded by a camera within a target area; the continuous image sequence is obtained by performing frame extraction on the video data collected by the camera;
[0126] A target recognition module, configured to perform target recognition on each image in the continuous image sequence to obtain a set of valid bounding boxes;
[0127] A detection feature generation module, configured to substitute all valid bounding boxes into a preset feature extraction model to obtain the detection feature corresponding to each bounding box, and determine whether the detection feature exists by querying a preset database according to the detection feature;
[0128] A motion trajectory determination module, configured to, if the detection feature exists, obtain the corresponding label ID, and obtain the motion trajectory according to the coordinate sequence of the label ID in each camera;
[0129] A linkage control module, configured to substitute the motion trajectory into an anomaly detection model to obtain an anomaly result. If the anomaly result indicates an anomaly, generate a tracking label for the bounding box, send out an alarm to generate an alarm message, and perform linkage control on the security devices according to the tracking label and the alarm message.
[0130] Based on the linkage control system based on video surveillance provided by an embodiment of the present invention, real-time monitoring and analysis are performed on a continuous image sequence, and feature extraction is performed after first detecting an effective bounding box, improving the detection efficiency of features. Then, combined with an anomaly detection model, suspicious behaviors can be effectively identified. Once an anomaly is detected, the system automatically generates a tracking label and an alarm message, thereby performing linkage control and improving the efficiency of linkage intelligent control.
[0131] In one embodiment, the target recognition module includes:
[0132] A preprocessing module, configured to, for each image in the continuous image sequence, perform preprocessing on the image to obtain a target detection image;
[0133] A bounding box determination module, configured to substitute the target detection image into a preset target recognition model to obtain a set of effective bounding boxes;
[0134] A model improvement module, configured to set the preset target recognition model as an improved YOLOv8 model; replace the Conv modules in the first layer, fifth layer, and seventh layer of the YOLOv8 model with downsampling fusion modules;
[0135] The working principle of the downsampling fusion module is as follows:
[0136] The downsampling fusion module includes a first module and a second module; the first module includes obtaining an input feature map, performing average pooling and convolution operations on the feature map to obtain a first output feature map; performing channel dimension splitting on the first output feature map to obtain 4 second output feature maps; applying the Softmax function to all the second output feature maps to obtain 4 sets of feature weights;
[0137] The second module includes performing convolution operations on the input feature map to obtain a third output feature map; performing channel dimension splitting on the third output feature map to obtain 4 fourth output feature maps;
[0138] Performing weighted fusion operations on the 4 fourth output feature maps according to the 4 sets of feature weights to obtain a final output feature map.
[0139] In one embodiment, the detection feature generation module includes:
[0140] A region extraction module, configured to extract the image region corresponding to each valid bounding box to obtain an initial input image;
[0141] A multi-scale feature extraction module, configured to substitute the initial input image into a Laplacian pyramid to obtain a first-scale feature, a second-scale feature, and a third-scale feature; the first-scale feature is a high-scale feature; the second-scale feature is a medium-scale feature; the third-scale feature is a low-scale feature;
[0142] A feature enhancement module, configured to substitute the first-scale feature map into a first enhanced feature extraction module to obtain a first enhanced feature; substitute the first enhanced feature, the second-scale feature, and the third-scale feature into a second enhanced feature extraction module to obtain a second enhanced feature;
[0143] A first feature fusion module, configured to perform feature fusion on the second-scale feature and the third-scale feature respectively with the second enhanced feature to obtain a first fusion feature and a second fusion feature;
[0144] A second feature fusion module, configured to perform scale normalization on the first enhanced feature, the first fusion feature, and the second fusion feature and then perform feature fusion to obtain a detection feature.
[0145] In one embodiment, the first enhanced feature extraction module includes:
[0146] A channel decomposition module, configured to perform channel decomposition on the first-scale feature map to obtain a first channel feature and a second channel feature;
[0147] A channel fusion feature extraction module, configured to substitute the first channel feature and the second channel feature into a full-channel extraction module respectively to obtain a first full-channel fusion feature and a second full-channel fusion feature;
[0148] A full-channel fusion module, configured to splice the first full-channel fusion feature and the second full-channel fusion feature to obtain a first enhanced feature.
[0149] In one embodiment, the second enhanced feature extraction module includes:
[0150] A scale feature splicing module, configured to splice the first enhanced feature, the second-scale feature, and the third-scale feature to obtain a first spliced feature;
[0151] An attention feature extraction module, configured to respectively pass the first spliced feature through a multi-scale dilated attention module and a multi-head attention module to obtain a first attention feature and a second attention feature;
[0152] An attention feature superposition module is used to superpose and fuse the first attention feature and the second attention feature to obtain a second enhanced feature.
[0153] The above has described an embodiment of the present invention in detail, but the content is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the present invention application shall still fall within the scope covered by the patent of the present invention.
Claims
1. A linkage and co-control method based on video surveillance, characterized in that, The method includes: Obtain a continuous image sequence uploaded by a camera within a target area; the continuous image sequence is obtained by frame extraction after the camera collects video data; Perform object recognition on each image in the continuous image sequence to obtain a set of valid bounding boxes; Substitute all valid bounding boxes into a preset feature extraction model to obtain detection features corresponding to each bounding box, and check whether the detection features exist by searching a preset database according to the detection features; If the detection features exist, obtain the corresponding label ID, and obtain the motion trajectory according to the coordinate sequence of the label ID in each camera; Substitute the motion trajectory into an anomaly detection model to obtain an anomaly result. If the anomaly result indicates an anomaly, generate a tracking label for the bounding box, issue an alarm to generate alarm information, and perform linkage control on security devices according to the tracking label and the alarm information; Among them, substituting all valid bounding boxes into a preset feature extraction model to obtain detection features corresponding to each bounding box includes: For each valid bounding box, extract the image area corresponding to the valid bounding box to obtain an initial input image; Substitute the initial input image into a Laplacian pyramid to obtain a first-scale feature, a second-scale feature, and a third-scale feature; the first-scale feature is a high-scale feature; the second-scale feature is a medium-scale feature; the third-scale feature is a low-scale feature; Substitute the first-scale feature map into a first enhanced feature extraction module to obtain a first enhanced feature; substitute the first enhanced feature, the second-scale feature, and the third-scale feature into a second enhanced feature extraction module to obtain a second enhanced feature; Perform feature fusion on the second-scale feature and the third-scale feature respectively with the second enhanced feature to obtain a first fusion feature and a second fusion feature; Perform scale normalization on the first enhanced feature, the first fusion feature, and the second fusion feature, and then perform feature fusion to obtain detection features.
2. The linkage control method based on video monitoring according to claim 1, wherein, Performing object recognition on each image in the continuous image sequence to obtain a set of valid bounding boxes includes: For each image in the continuous image sequence, preprocess the image to obtain a target detection image; Substitute the target detection image into a preset object recognition model to obtain a set of valid bounding boxes; The preset object recognition model is an improved YOLOv8 model; replace the Conv modules in the first layer, the fifth layer, and the seventh layer of the YOLOv8 model with downsampling fusion modules; The working principle of the downsampling fusion module is: The downsampling fusion module includes a first module and a second module; the first module includes obtaining an input feature map, performing average pooling and convolution operations on the feature map to obtain a first output feature map; performing channel dimension division on the first output feature map to obtain 4 second output feature maps; applying the Softmax function to all second output feature maps to obtain 4 sets of feature weights; The second module includes performing a convolution operation on the input feature map to obtain a third output feature map; performing channel dimension division on the third output feature map to obtain 4 fourth output feature maps; Weighted fusion operations are respectively performed on the four fourth output feature maps according to four groups of feature weights to obtain the final output feature map.
3. The linkage control method based on video surveillance according to claim 1, wherein The first enhanced feature extraction module includes: Performing channel decomposition on the first-scale feature map to obtain a first channel feature and a second channel feature; Substituting the first channel feature and the second channel feature into the full-channel extraction module respectively to obtain a first full-channel fusion feature and a second full-channel fusion feature; Performing splicing on the first full-channel fusion feature and the second full-channel fusion feature to obtain a first enhanced feature.
4. A linkage control method based on video surveillance according to claim 1, characterized in that, The second enhanced feature extraction module includes: Performing splicing on the first enhanced feature, the second-scale feature, and the third-scale feature to obtain a first spliced feature; Passing the first spliced feature through a multi-scale dilation attention module and a multi-head attention module respectively to obtain a first attention feature and a second attention feature; Performing superposition fusion on the first attention feature and the second attention feature to obtain a second enhanced feature.
5. A linkage control system based on video surveillance, characterized in that, The system includes: A frame extraction processing module, configured to obtain a continuous image sequence uploaded by a camera within a target area; the continuous image sequence is obtained by performing frame extraction on video data collected by the camera; A target recognition module, configured to perform target recognition on each image in the continuous image sequence to obtain a set of valid bounding boxes; A detection feature generation module, configured to substitute all valid bounding boxes into a preset feature extraction model to obtain a detection feature corresponding to each bounding box, and determine whether the detection feature exists by searching a preset database; A motion trajectory determination module, configured to, if the detection feature exists, obtain the corresponding label ID, and obtain a motion trajectory according to the coordinate sequence of the label ID in each camera; A linkage control module, configured to substitute the motion trajectory into an anomaly detection model to obtain an anomaly result. If the anomaly result indicates an anomaly, generate a tracking label for the bounding box, issue an alarm to generate alarm information, and perform linkage control on security devices according to the tracking label and the alarm information; Wherein, the detection feature generation module includes: A region extraction module, configured to, for each valid bounding box, extract the image region corresponding to the valid bounding box to obtain an initial input image; A multi-scale feature extraction module, configured to substitute the initial input image into a Laplacian pyramid to obtain a first-scale feature, a second-scale feature, and a third-scale feature; the first-scale feature is a high-scale feature; the second-scale feature is a medium-scale feature; the third-scale feature is a low-scale feature; A feature enhancement module, configured to substitute the first-scale feature map into the first enhanced feature extraction module to obtain a first enhanced feature; substitute the first enhanced feature, the second-scale feature, and the third-scale feature into the second enhanced feature extraction module to obtain a second enhanced feature; A first feature fusion module, configured to perform feature fusion on the second enhanced feature with the second-scale feature and the third-scale feature respectively to obtain a first fusion feature and a second fusion feature; A second feature fusion module, configured to perform scale normalization on the first enhanced feature, the first fusion feature, and the second fusion feature and then perform feature fusion to obtain a detection feature.
6. The linkage control system based on video surveillance according to claim 5, characterized in that, The target recognition module includes: A preprocessing module, which is used to preprocess each image in the continuous image sequence to obtain a target detection image; A bounding box determination module, which is used to substitute the target detection image into a preset target recognition model to obtain a set of effective bounding boxes; A model improvement module, which is used to set the preset target recognition model as an improved YOLOv8 model; replace the Conv modules in the first, fifth, and seventh layers of the YOLOv8 model with downsampling fusion modules; The working principle of the downsampling fusion module is as follows: The downsampling fusion module includes a first module and a second module; the first module includes obtaining an input feature map, performing average pooling and convolution operations on the feature map to obtain a first output feature map; performing channel dimension splitting on the first output feature map to obtain 4 second output feature maps; applying the Softmax function to all the second output feature maps to obtain 4 sets of feature weights; The second module includes performing a convolution operation on the input feature map to obtain a third output feature map; performing channel dimension splitting on the third output feature map to obtain 4 fourth output feature maps; Performing weighted fusion operations on the 4 fourth output feature maps according to the 4 sets of feature weights to obtain a final output feature map.
7. The linkage control system based on video surveillance according to claim 5, characterized in that, The first enhanced feature extraction module includes: A channel decomposition module, which is used to decompose the first-scale feature map into a first channel feature and a second channel feature; A channel fusion feature extraction module, which is used to substitute the first channel feature and the second channel feature into a full-channel extraction module to obtain a first full-channel fusion feature and a second full-channel fusion feature; A full-channel fusion module, which is used to splice the first full-channel fusion feature and the second full-channel fusion feature to obtain a first enhanced feature.
8. The linkage control system based on video surveillance according to claim 5, wherein The second enhanced feature extraction module includes: A scale feature splicing module, which is used to splice the first enhanced feature, the second-scale feature, and the third-scale feature to obtain a first spliced feature; An attention feature extraction module, which is used to pass the first spliced feature through a multi-scale dilation attention module and a multi-head attention module respectively to obtain a first attention feature and a second attention feature; An attention feature superposition module, which is used to superpose and fuse the first attention feature and the second attention feature to obtain a second enhanced feature.
Citation Information
Patent Citations
Strip steel surface defect detection method based on multi-scale feature fusion and related device
CN118735878A
Image processing method and system for intelligent security and protection monitoring
CN118887622A