Method, apparatus, device and storage medium for splicing surveillance videos

Through the improved YOLOv5 model and SE-ResNet-50 network, humanoid detection and feature extraction are performed, combined with Kalman filter and distributed computing framework, the problem of insufficient target continuity and real-time processing capabilities in multi-camera systems is solved, and the quality and stability of video stitching are improved.

CN119359537BActive Publication Date: 2025-07-18SHENZHEN ANKED SHITONG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411396281.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-08
Publication Date
2025-07-18
Estimated Expiration
2044-10-08

AI Technical Summary

Technical Problem

The existing video stitching methods have insufficient dynamic target continuity in multi-camera systems, resulting in the problem of target breakage, repetition or loss. At the same time, the real-time processing capabilities of large-scale video surveillance systems are insufficient, making it difficult to meet the needs of low latency and high throughput. The stability and accuracy of object detection and tracking in complex scenarios are affected by factors such as occlusion and lighting changes.

Method used

The improved YOLOv5 model is used for humanoid detection, combined with spatial attention and channel attention mechanisms, and the face features are extracted using SE-ResNet-50 backbone network and self-attention mechanism, motion prediction and trajectory extrapolation are performed through Kalman filters, distributed computing framework and dynamic task allocation mechanism are introduced, and video stitching is realized in combination with dual-stream network processing method.

Benefits of technology

It improves the accuracy and robustness of object detection in complex scenarios, enhances the discrimination of face features, improves the accuracy and stability of multi-objective tracking, realizes efficient processing of large-scale video data, and improves the quality and visual effects of spliced videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119359537B_ABST
    Figure CN119359537B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of video splicing, and discloses a method, device, equipment and storage medium for splicing surveillance videos. The method includes: performing video frame division on the surveillance videos collected by a surveillance terminal to obtain a plurality of target video frames; performing human detection to obtain a human detection result; aligning and cropping the face regions in the human detection result to obtain standardized face images, and extracting face feature vectors; performing motion prediction and trajectory extrapolation to obtain a multi-target human tracking result; inputting the multi-target human tracking result into a distributed computing framework for task allocation and cross-node feature matching to obtain a global human tracking result; performing two-stream network processing on the plurality of target video frames according to the global human tracking result, calculating video splicing parameters, and performing a splicing operation on the plurality of target video frames to obtain a spliced video. The present invention effectively improves the quality and visual effect of the spliced video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video splicing, and particularly to a method, device, equipment and storage medium for splicing surveillance videos. Background Art

[0002] With the wide application of video surveillance systems in fields such as public security, traffic management, and smart cities, higher requirements are put forward for high-quality and large-scale video surveillance. Traditional single cameras are difficult to meet the needs of large-scene surveillance, and it has become an inevitable trend for multiple cameras to work together. However, multi-camera systems face a series of technical challenges such as video splicing, target tracking, and real-time processing.

[0003] Existing video splicing methods often ignore the continuity of dynamic targets, resulting in problems such as target breaks, repetitions, or losses in the splicing results. In addition, the massive data generated by large-scale video surveillance systems poses a severe test to real-time processing capabilities, and traditional centralized computing architectures are difficult to meet the requirements of low latency and high throughput. At the same time, target detection and tracking in complex scenarios still face many difficulties, such as factors like occlusion, illumination changes, and pose diversity affecting the stability and accuracy of algorithms. Summary of the Invention

[0004] The present invention provides a method, device, equipment and storage medium for splicing surveillance videos, and effectively improves the quality and visual effect of the spliced videos.

[0005] In a first aspect, the present invention provides a method for splicing surveillance videos, and the method for splicing surveillance videos includes:

[0006] Performing video frame division on the surveillance videos collected by a surveillance terminal to obtain a plurality of target video frames;

[0007] Inputting the plurality of target video frames into an improved YOLOv5 model for human detection to obtain human detection results;

[0008] Aligning and cropping the face regions in the human detection results to obtain standardized face images, and extracting face feature vectors;

[0009] Performing motion prediction and trajectory extrapolation based on the face feature vectors and the human detection results to obtain multi-target human tracking results;

[0010] Inputting the multi-target human tracking results into a distributed computing framework for task allocation and cross-node feature matching to obtain global human tracking results;

[0011] Perform two-stream network processing on the multiple target video frames according to the global human form tracking result, calculate video stitching parameters, and perform a stitching operation on the multiple target video frames to obtain a stitched video.

[0012] In a second aspect, the present invention provides a monitoring video stitching device, which includes:

[0013] A frame splitting module, configured to split a monitoring video collected by a monitoring terminal into video frames to obtain multiple target video frames;

[0014] A detection module, configured to input the multiple target video frames into an improved YOLOv5 model for human form detection to obtain a human form detection result;

[0015] An extraction module, configured to align and crop the face regions in the human form detection result to obtain a standardized face image, and extract face feature vectors;

[0016] A prediction module, configured to perform motion prediction and trajectory extrapolation based on the face feature vectors and the human form detection result to obtain a multi-target human form tracking result;

[0017] A tracking module, configured to input the multi-target human form tracking result into a distributed computing framework for task allocation and cross-node feature matching to obtain a global human form tracking result;

[0018] A calculation module, configured to perform two-stream network processing on the multiple target video frames according to the global human form tracking result, calculate video stitching parameters, and perform a stitching operation on the multiple target video frames to obtain a stitched video.

[0019] In a third aspect of the present invention, there is provided a computer device, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor invokes the instructions in the memory so that the computer device executes the above-mentioned monitoring video stitching method.

[0020] In a fourth aspect of the present invention, there is provided a computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium, and when the instructions are run on a computer, the computer is made to execute the above-mentioned monitoring video stitching method.

[0021] In the technical solution provided by the present invention, human detection is performed through an improved YOLOv5 model, and spatial attention and channel attention mechanisms are introduced, which improves the accuracy and robustness of object detection in complex scenarios. The improved FaceNet network with the SE-ResNet-50 backbone network and self-attention mechanism is used to extract face features, enhancing the discriminability of face features and facilitating subsequent object tracking and identity recognition. The Kalman filter with an adaptive noise covariance matrix is used for motion prediction and trajectory extrapolation, improving the accuracy and stability of multi-object tracking, especially in the case of short-term occlusion of objects. A distributed computing framework and a dynamic task allocation mechanism are introduced to achieve efficient processing of large-scale video data, significantly improving the real-time performance and scalability of the system. Through the two-stream network processing method, spatial semantic features and temporal residual noise information are comprehensively considered, improving the accuracy and coherence of video stitching and effectively solving the problems of breakage and repetition of dynamic objects during the stitching process. The present invention effectively improves the quality and visual effect of the stitched video. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 It is a schematic diagram of the steps of the method for stitching surveillance videos in the embodiments of the present invention;

[0024] Figure 2 It is a schematic diagram of the structure of the device for stitching surveillance videos in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] The embodiments of the present invention provide a method, device, equipment and storage medium for stitching surveillance videos. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0026] For ease of understanding, the specific process of the embodiments of the present invention will be described below. Please refer to Figure 1 , an embodiment of the method for splicing monitoring videos in the embodiments of the present invention includes:

[0027] Step S1: Perform video frame segmentation on the monitoring videos collected by the monitoring terminal to obtain a plurality of target video frames;

[0028] It can be understood that the execution subject of the present invention can be a splicing device for monitoring videos, or a terminal or a server. Specifically, no limitation is made here. In the embodiments of the present invention, the server is taken as the execution subject for illustration.

[0029] Specifically, monitor terminal is used to collect monitoring video data, including continuous video streams in different scenarios. The collected video stream is sampled at a preset frame rate, and multiple initial video frames are extracted from the video. The setting of the frame rate is determined according to actual needs. For example, 30 frames or 60 frames are extracted per second to ensure that the extracted initial video frames can fully reflect the dynamic information in the monitoring video. After sampling, a series of initial video frames obtained may have different resolutions or quality problems. Each initial video frame is subjected to resolution normalization processing to unify the resolutions of all frames to a predetermined standard resolution, such as 1920x1080 or 1280x720, so that each frame image has consistency during subsequent processing, thereby reducing possible calculation deviations or processing errors caused by different resolutions. Bilateral filtering is performed on the normalized initial video frames. Bilateral filtering is a smoothing filtering method that preserves the edge features of the image. By considering the information in both the spatial domain and the pixel value domain simultaneously, it effectively removes the noise in the image while retaining the edge details. In bilateral filtering, two key parameters are set: the radius of the spatial domain kernel function and the standard deviation of the value domain kernel function. The radius of the spatial domain kernel function is set to 5, indicating that the influence range of a pixel during the filtering process is a radius area of 5 pixels; while the standard deviation of the value domain kernel function is set to 50, indicating the smoothness of the pixel value changes considered during the filtering process. After processing, multiple preliminarily denoised video frames are obtained. Edge-preserving filtering is performed on the preliminarily denoised video frames to enhance the edge features in the image. Edge-preserving filtering is a technique that can effectively enhance the edges of the image and can highlight the edge structure in the image while smoothing the image noise. The edge-enhanced video frames are divided into non-overlapping 8x8 blocks. For each small block, its cumulative distribution function (CDF) is calculated to obtain a local CDF matrix. The cumulative distribution function can describe the distribution of pixel values in the image and helps to enhance the local contrast of the image. Based on the local CDF matrix, the mapping function of each pixel is calculated using bilinear interpolation. Bilinear interpolation is an image processing technique that can calculate the pixel values of unknown points on a given pixel grid, making the image processing result smoother and more refined. Through the interpolation method, the local cumulative distribution function information is mapped to the entire image to obtain a global mapping matrix, which describes the mapping relationship of each pixel in the entire image. Based on the global mapping matrix, pixel-level gray value remapping is performed on each edge-enhanced video frame to enhance the contrast of the entire image, and multiple contrast-enhanced video frames are obtained. Color space conversion is performed on the contrast-enhanced video frames. The image is converted from the RGB color space to the YUV color space. The YUV color space can better separate the luminance information and color information of the image. In the YUV color space, only the Y channel (luminance channel) is subjected to histogram equalization. Histogram equalization adjusts the luminance distribution of the image to make the contrast of the image more uniform and clear.After processing, the YUV image is then converted back to the RGB color space, preserving the original color information of the image while improving the brightness distribution of the image. The resulting multiple video frames have better contrast and can also maintain the original color effect.

[0030] Step S2: Input multiple target video frames into the improved YOLOv5 model for human detection to obtain the human detection results;

[0031] Specifically, multi-scale transformation is performed on the input video frames to enhance the model's detection ability for targets of different sizes. Three different-scale feature maps are generated through bilinear interpolation, corresponding to 1 / 8, 1 / 16, and 1 / 32 of the original input size respectively. The multi-scale input feature maps are sequentially input into the backbone network of the improved YOLOv5 model. The improved backbone network is different from the original YOLOv5 model, and a spatial attention module is added behind each convolutional block. The spatial attention module emphasizes important spatial regions in the image by generating a two-dimensional attention weight map, ignoring irrelevant or unimportant background information. The spatial attention module calculates the weights of each pixel on the feature map and adjusts the feature values at different positions, so that the features in the important regions are more emphasized in the network, and a preliminary feature map is obtained. Channel attention processing is performed on the preliminary feature map containing spatial attention information. The channel attention mechanism is realized by calculating the global average pooling value of each channel, and the pooling value represents the importance of each channel on the entire feature map. The pooling result is input into a multi-layer perceptron composed of two fully connected layers. In the multi-layer perceptron, the number of neurons in the first layer is set to 1 / 16 of the number of channels, effectively reducing the number of parameters and increasing the model's non-linear representation ability. The number of neurons in the second layer is restored to the original number of channels to match the initial feature map. The output channel weights are mapped to between 0 and 1 through the Sigmoid activation function, thus effectively controlling the importance of each channel. The channel weights are multiplied element-wise with the preliminary feature map to obtain an enhanced feature map containing both spatial and channel attention information. The enhanced feature map containing spatial and channel attention information is input into the Feature Pyramid Network for processing. The Feature Pyramid Network performs feature fusion through a top-down path and lateral connections. The top-down path uses 1x1 convolutions to adjust the number of channels, and each layer performs a 2x upsampling operation, gradually upsampling the lower-resolution feature maps to higher resolutions. At the same time, the lateral connections use 1x1 convolutions to process the features of the same layer, enabling the fusion of feature information from different levels. The design of the Feature Pyramid Network enables the network to utilize feature maps of different scales simultaneously, making the model perform better in detecting humanoid targets of different sizes. Through feature fusion, a multi-scale fused feature map is obtained. A 3x3 convolutional kernel is applied to the multi-scale fused feature map for feature extraction, effectively capturing local spatial information and extracting richer local features while keeping the size of the feature map unchanged. The number of channels is compressed to the preset number of anchor boxes multiplied by (5 + the number of classes) using a 1x1 convolutional kernel. 5 represents the 4 coordinate values and 1 confidence value of the bounding box, and the number of classes represents the possible different types in humanoid detection, such as target objects like people and vehicles. After processing, the original prediction results contain data such as the position information, confidence, and class probability of each anchor box.Decode the original prediction results, convert the predicted bounding box coordinates from the offsets relative to the grid to the absolute coordinates of the image. Multiply the predicted relative coordinates by the stride of the feature map and add the upper left coordinates of the corresponding grid. Convert the relative positions of the anchor boxes to the absolute positions on the image to accurately locate the positions of the targets. At the same time, apply the Sigmoid function to process the confidence and class probabilities and map them between 0 and 1. The prediction results obtained after decoding clearly reflect the absolute positions and class probabilities of the targets. Based on the decoded prediction results, use the non-maximum suppression algorithm to filter out overlapping detection boxes. The non-maximum suppression algorithm sorts all the detection boxes in descending order of confidence and preferentially retains the detection result with the highest confidence. For each detection box, compare its overlap with other detection boxes one by one, and judge by calculating the intersection over union of the overlapping boxes. When the intersection over union is greater than the preset threshold (set to 0.5), only retain the box with the highest confidence and delete all other overlapping detection boxes, thereby effectively removing duplicate detections or overlapping detections and avoiding false detections and missed detections caused by the overlap of multiple boxes, and finally obtaining the human detection results.

[0032] Step S3: Align and crop the face regions in the human detection results to obtain normalized face images and extract face feature vectors;

[0033] Specifically, the MTCNN algorithm is applied to the face region in the human detection result for key point detection. MTCNN is a multi-task learning model that can perform face detection and simultaneously locate the key points of the face. The MTCNN algorithm consists of three cascaded networks, namely P-Net (Proposal Network), R-Net (Refine Network), and O-Net (Output Network). During the detection process, P-Net generates candidate face bounding boxes, R-Net filters and refines the candidate boxes, and O-Net further adjusts and outputs the final face bounding box and the coordinates of five key points. Through the MTCNN algorithm, the coordinates of the left eye center, right eye center, nose tip, left mouth corner, and right mouth corner in each face image are obtained. Based on the coordinates of the five key points, an affine transformation matrix is calculated. Affine transformation is a geometric transformation that can perform operations such as translation, rotation, scaling, and shearing on the image to transform the face image to a standard position. The face is aligned. The face image is adjusted to a predetermined standard position through affine transformation, so that the centers of the two eyes are on the horizontal line, and the relative positions of the eyes and the mouth conform to the conventional distribution of face geometry. The calculation of the affine transformation matrix is based on the relationship between the coordinates of the five key points and the coordinates of the key points in the standard position, and is solved by the least squares method. The original face image is multiplied by this affine transformation matrix to obtain the aligned face image. The aligned face image is cropped and scaled. The face image is cropped to only contain the face region, and the cropped face image is scaled to the standard size of 112x112 pixels. The standardized size can ensure that different face images have the same scale when input into the neural network, avoiding the problem of poor feature extraction effect caused by inconsistent face sizes. The pixel values are normalized, and the pixel values are scaled to the interval [-1,1] to eliminate the influence of external factors such as illumination and contrast on the image, improve the robustness and stability of the model, and obtain the standardized face image. The standardized face image is input into the SE-ResNet-50 backbone network for feature extraction. SE-ResNet-50 is a deep convolutional neural network that adds a Squeeze-and-Excitation (SE) module on the basis of the traditional ResNet-50 network. ResNet-50 contains 50 convolutional layers. By introducing the residual structure, the problem of gradient disappearance in the deep network is effectively alleviated, enabling the network to better learn deeper features. The introduction of the SE module enhances the expression ability of the network. Through global average pooling, the features of each convolutional layer are compressed in terms of global information in space, and then the importance weights of each channel are learned through two fully connected layers.The number of neurons in the first fully connected layer is 1 / 16 of the number of channels, which is used to reduce the computational complexity and introduce non-linear relationships. The second layer restores the number of neurons to the original number of channels and maps the weights to the interval [0, 1] through the Sigmoid function, so as to assign a weight value to each channel. Apply the self-attention mechanism to the face feature map. By calculating the correlation between each position in the feature map and all other positions, an attention weight matrix is generated to capture the long-range dependencies between different regions in the image. The self-attention mechanism obtains the attention weight matrix by calculating the correlation between each position in the face feature map and all other positions, and then multiplies this matrix by the original feature map to obtain the enhanced feature map. Input the enhanced feature map into the global average pooling layer for processing. Global average pooling reduces the spatial dimension of the enhanced feature map to 1x1, only retaining the channel dimension. Compress the spatial information of each channel into a global information, eliminating the influence of spatial position on the features, while reducing the number of parameters and avoiding overfitting. After global average pooling, a 2048-dimensional feature vector is obtained. Input the 2048-dimensional feature vector into the fully connected layer to map it to a 512-dimensional space to obtain the face feature vector.

[0034] Step S4: Based on the face feature vector and the humanoid detection result, perform motion prediction and trajectory extrapolation to obtain the multi-target humanoid tracking result;

[0035] Specifically, a state vector is established for each target in the human detection result, including information such as the center coordinates, width, height, speed, and acceleration of the target, which describes the position, size, and motion state of the target in space. At the same time, the state vector and covariance matrix of the Kalman filter are initialized. The Kalman filter is a recursive filtering algorithm that estimates the system state through prediction and update, and can provide relatively accurate state estimation results in the presence of noise. The state vector at initialization is set by the target position and size information extracted from the human detection result, while the covariance matrix is used to describe the uncertainty of the state vector. After initialization, an initial state estimate is obtained. Based on the initial state estimate, the state transition matrix is used to predict the position of the target in the next frame. The state transition matrix is a linear model that describes how the system transitions from the current state to the next state, taking into account factors such as the speed and acceleration of the target to make the prediction more accurate. By using the state transition matrix to predict the target state, a prior state estimate is obtained. The prior state estimate is an estimated result of the possible position, size, and motion state of the target in the next frame based on the information in the current frame. According to the prior state estimate and the human detection result in the current frame, a matching cost matrix is calculated. The matching cost matrix is a two-dimensional matrix, where each element represents the matching cost between a certain predicted target and a certain detected target. The matching cost consists of two parts. One is the difference in position, i.e., the Euclidean distance of the target position, and the other is the difference in size, i.e., the difference in the width and height of the target. In this way, each predicted target is matched with the detected targets in the current frame. The Hungarian algorithm is used for data association. The Hungarian algorithm is a method for solving the bipartite graph matching problem, which can find a globally optimal matching scheme in the case of many-to-many. Through the Hungarian algorithm, the matching pairs of the target and the detection result are obtained. For each target in the matching pairs, the cosine similarity between the face feature vectors is calculated. The cosine similarity is an index that measures the similarity degree between two vectors, and its value range is between [-1, 1]. The closer it is to 1, the more similar the two vectors are. The cosine similarity effectively measures whether the faces in two different detection boxes belong to the same person. Combining the cosine similarity with the position information can more accurately determine the identity of the target, and the target identity confirmation result is obtained. Based on the target identity confirmation result and the current observation value, the state estimate of the Kalman filter is updated to obtain the posterior state estimate. The observation noise and process noise in the posterior state estimate are adaptively adjusted. The observation noise and process noise respectively describe the uncertainties in the measurement process and the system motion process. By calculating the sample covariance of the observation residual sequence in the recent A frames and using the exponentially weighted average method to update the noise covariance matrix, an adaptive noise covariance matrix is obtained.The sample covariance can reflect the fluctuations of the observation residuals. The exponentially weighted moving average method can respond more quickly to the latest change trends while retaining historical information, enabling the Kalman filter to better adapt to the changes in the actual scenario. Substitute the adaptive noise covariance matrix into the prediction and update equations of the Kalman filter to recalculate the Kalman gain and state estimation. The Kalman gain describes the weight allocation of the observation information in the state update. The adaptively adjusted Kalman gain can better balance the relationship between prediction and observation, resulting in a more optimized state estimation result. For the unmatched detection results, initialize new trajectories. For each new detection result, initialize a new trajectory based on its position, velocity, and other information and track it. At the same time, for the trajectories that have not been matched for consecutive A frames, use the long short-term memory network (LSTM) for trajectory extrapolation. The LSTM predicts the positions in the next A frames based on the historical motion trajectory of the target. Through trajectory extrapolation, the blank areas in the trajectory are effectively filled, obtaining a more continuous and complete multi-target humanoid tracking result.

[0036] Step S5: Input the multi-target humanoid tracking result into the distributed computing framework for task allocation and cross-node feature matching to obtain the global humanoid tracking result;

[0037] Specifically, spatiotemporal segmentation is performed on the multi-target human tracking results, and the continuous B-frame tracking results are divided into individual task units, which are convenient for parallel processing and scheduling management in a distributed computing framework. Each task unit represents the tracking data and its corresponding spatial information within a certain period of time (i.e., B frames), so that it can be better allocated to different computing nodes for processing. Predict the future load of each node according to the CPU usage rate, memory occupancy rate, and network bandwidth of each computing node. Analyze and model the resource usage of each node through historical data and trend analysis models to obtain the node load prediction value. The load prediction value reflects the possible resource occupancy of each node at future moments. According to the node load prediction value and the complexity of the task unit, perform a preliminary allocation of the task unit. The complexity of the task unit is determined by factors such as the number of tracking targets, the complexity of the movement trajectory, and the data volume. For task units with higher complexity, try to allocate them to computing nodes with lower load and sufficient resources to ensure their processing efficiency and stability. Optimize the load balance of the initially allocated task plan. Through task migration and splitting operations, control the load difference of each node within a preset threshold. Task migration refers to transferring task units on some nodes to other nodes with lighter load, while task splitting is to decompose task units with higher complexity into multiple smaller task units for parallel processing on different nodes. Through optimization operations, obtain a more reasonable and balanced task allocation plan to ensure the load balance of all nodes, thereby improving the processing efficiency and stability of the entire distributed computing framework. Send the optimized task allocation plan to each computing node and construct a local feature index on each node. The local feature index refers to using an efficient data structure such as a hash table to store face feature vectors and corresponding trajectory IDs on each node. The hash table can quickly perform access operations on feature vectors, thereby improving the matching efficiency in subsequent feature matching processes. The local feature index structure on each node ensures the rapid retrieval and comparison of face features within the node. Construct a global feature index on the central control node. The global feature index adopts a hierarchical tree structure to organize the feature information of each node, and each leaf node corresponds to the local index of a computing node. This structure can effectively manage and organize the feature data in the entire distributed system, facilitating quick positioning to the target node when performing cross-node matching. For cross-node face matching requests, first perform a rough match in the global feature index. Quickly screen out the K leaf nodes that are most likely to contain the target feature through the global feature index. The rough match quickly locates the node range where the target feature is located through methods such as hash mapping and KD tree. Perform an exact match in the local indexes of the K leaf nodes. The exact match calculates the similarity between the target feature and all feature vectors in the local feature index to find the most matching result and obtain the cross-node matching result.Based on the cross-node matching results, the trajectory segments on different nodes are associated and fused to eliminate duplicate trajectory records. The process of trajectory fusion includes comparing the start and end times of the trajectory segments, merging the overlapping trajectory segments, and removing duplicate trajectory data. Global consistency processing is performed on the trajectory IDs, and a unique trajectory ID is assigned to each target globally, ensuring that the same target has one and only one unique identifier throughout the system, and finally obtaining the global human tracking result.

[0038] Step S6: Perform dual-stream network processing on multiple target video frames according to the global human tracking result, calculate the video stitching parameters, and perform a stitching operation on the multiple target video frames to obtain a stitched video.

[0039] Specifically, multiple target video frames are reasonably grouped according to the global human tracking results. The same human target continuously tracked is divided into the same video segment, and the frame sequence in each video segment is regarded as an independent analysis unit. The entire video stream is sliced into multiple video segments, and each video segment contains E frame images. ResNext-101 network is applied to each frame in the segmented multiple video segments for feature extraction. ResNext-101 is a deep convolutional neural network that contains 32 residual blocks with a cardinality of 4, and these residual blocks can extract multi-level and rich image features. Each residual block outputs a 256-dimensional feature representation, and after stacking, it is reduced to a 2048-dimensional spatial feature vector through a global average pooling layer. All frames of each video segment are represented as an E×2048-dimensional spatial feature sequence, where E represents the number of frames contained in the segment. The self-attention mechanism is applied to the spatial feature sequence, and the mutual correlation between video frames is measured by calculating an E×E attention weight matrix. Each element in the attention weight matrix represents the correlation magnitude between two frames in the sequence. After multiplying the attention weight matrix by the original spatial feature sequence, a weighted spatial feature sequence is obtained, making the originally relatively independent frame sequence more coherent in terms of spatial features. At the same time, the difference image is calculated between each video frame in the video segment and its previous frame, and E - 1 residual noise images are obtained using the inter-frame difference method. Inter-frame difference can effectively reflect the motion changes and differences between two consecutive frames, thereby extracting the dynamic information in the video. The residual noise images are stacked into an E - 1×H×W×3 residual noise sequence, where H and W are the height and width of the image respectively. The residual noise sequence is used to capture the motion features and change patterns between frames in the video. The residual noise sequence is input into a 3D convolutional neural network. This 3D convolutional neural network contains 4 3D convolutional layers, with the convolutional kernel size of each layer being 3×3×3, the stride set to 1×2×2, and the number of channels being 64, 128, 256, and 512 in sequence. The 3D convolutional network can extract features simultaneously in both the spatial and temporal dimensions, capturing the dynamic change information in the video. After being processed by 4 3D convolutional layers, through 3D global average pooling, the feature map is reduced to a 512-dimensional temporal feature vector. Each video segment is represented by an E×512-dimensional temporal feature sequence. The temporal feature sequence describes the dynamic pattern that changes over time in the video segment. The weighted spatial feature sequence and the temporal feature sequence are subjected to a bilinear pooling operation to obtain an E×d-dimensional fused feature sequence, where d is the dimension of the fused feature. Based on the fused feature sequence, a two-layer fully connected neural network is used to perform regression calculations for the relative position offset, rotation angle, and scaling factor between adjacent video segments. The first layer of the fully connected network is used to perform a preliminary linear transformation on the fused features to extract features that contribute to the regression of the stitching parameters; the second layer of the fully connected network then maps the features to the specific stitching parameters.Through the trained regression model, calculate the relative position relationship between adjacent video segments to obtain video stitching parameters, including information such as translation, rotation, and scaling between video segments. According to the video stitching parameters, perform affine transformation on multiple video segments. Affine transformation is a geometric transformation method that can perform translation, rotation, and scaling simultaneously, aligning different video segments into the global coordinate system. The transformed images are stitched according to their relative positions, connecting all segments into a complete stitched video, resulting in a spatio-temporally coherent and visually appealing stitched video.

[0040] In the embodiments of the present invention, a modified YOLOv5 model is used for human detection, introducing spatial attention and channel attention mechanisms, which improves the accuracy and robustness of object detection in complex scenarios. An improved FaceNet network with an SE-ResNet-50 backbone network and self-attention mechanism is used to extract face features, enhancing the discriminability of face features and facilitating subsequent object tracking and identity recognition. A Kalman filter with an adaptive noise covariance matrix is used for motion prediction and trajectory extrapolation, improving the accuracy and stability of multi-object tracking, especially in the case of short-term occlusion of objects. A distributed computing framework and a dynamic task allocation mechanism are introduced to achieve efficient processing of large-scale video data, significantly improving the real-time performance and scalability of the system. Through the two-stream network processing method, spatial semantic features and temporal residual noise information are comprehensively considered, improving the accuracy and coherence of video stitching and effectively solving the problems of breakage and repetition of dynamic objects during the stitching process. The present invention effectively improves the quality and visual effect of the stitched video.

[0041] In a specific embodiment, the process of executing step S1 may specifically include the following steps:

[0042] The monitoring terminal collects monitoring videos, samples the monitoring videos at a preset frame rate to obtain multiple initial video frames, and performs resolution normalization processing on each of the multiple initial video frames to obtain multiple normalized video frames; performs bilateral filtering on the multiple normalized video frames, where the radius of the spatial domain kernel function is set to 5 and the standard deviation of the range domain kernel function is set to 50, to obtain multiple preliminarily denoised video frames; performs edge-preserving filtering on the multiple preliminarily denoised video frames to obtain multiple edge-enhanced video frames, divides the edge-enhanced video frames into non-overlapping 8x8 blocks, calculates the cumulative distribution function for each block to obtain a local CDF matrix; based on the local CDF matrix, uses bilinear interpolation to calculate the mapping function for each pixel to obtain a global mapping matrix, and remaps the gray values of the edge-enhanced video frames at the pixel level according to the global mapping matrix to obtain multiple contrast-enhanced video frames; performs color space conversion on the multiple contrast-enhanced video frames, converts the RGB color space to the YUV color space, applies histogram equalization only to the Y channel, and then converts the processed YUV image back to the RGB color space to obtain multiple target video frames.

[0043] Specifically, the collected original monitoring video is sampled at a preset frame rate to obtain multiple initial video frames. Assume that the frame rate of the original monitoring video is R frames per second, and the preset frame rate is R s frames per second, and perform downsampling on the video. Extract one video frame every frame from the original video to obtain multiple initial video frames. Perform resolution normalization processing on the initial video frames. Assume that the original resolution of each video frame is H orig ×W orig′ and the target resolution is H target ×W target . The purpose of resolution normalization is to adjust all the initial video frames to the same resolution to ensure consistent spatial characteristics between frames during subsequent processing. The normalization process is implemented using bilinear interpolation, and the interpolation formula is:

[0044]

[0045] where I target (x,y) represents the pixel value at coordinates (x,y) in the normalized frame, and I orig (x′,y′) represents the pixel value at the corresponding position in the original frame. W orig and H orig are the width and height of the original frame respectively, and W target and H targetis the width and height of the target frame. The pixel values in the original frame are mapped to the target resolution frame through this formula to obtain multiple video frames with standardized resolution. Bilateral filtering is performed on the standardized video frames. Bilateral filtering is a filtering method that can remove noise while retaining image edge details. The image is smoothed by considering the weights of both the spatial domain and the value domain. The mathematical expression of bilateral filtering is:

[0046]

[0047] Among them, I filtered (x, y) represents the pixel value at the coordinate (x, y) after bilateral filtering, I norm (x,y represents the pixel value at (x,y) in the frame after resolution normalization, Ω is the filter window range, usually a square area. d The standard deviation of the spatial domain kernel function determines the size of the filter window and is set to 5; while σ r is the standard deviation of the range kernel function, which controls the weight of color similarity and is set to 50. W(x,y) is a normalization factor used to ensure that the pixel values of the filtering results are within the correct range. By calculating with this formula, it is possible to remove noise while maintaining the integrity of the edge area of the image, and obtain multiple video frames after preliminary denoising. The video frames after preliminary denoising are processed by edge-preserving filtering to enhance the edge features in the image. Edge-preserving filtering suppresses the filtering effect in flat areas by assigning higher weights to areas with large gradient changes. The edge-enhanced video frame is divided into 8x8 non-overlapping blocks, and the cumulative distribution function (CDF) of each block is calculated to capture the pixel distribution characteristics of the local area in the image. The cumulative distribution function is defined as:

[0048]

[0049] Among them, CDF(x,y,k) represents the cumulative probability of the pixel point with the pixel value less than or equal to k at the coordinate (x,y) in block B. is the indicator function, when I enhanced When (x′, y′)≤k, the indicator function is 1, otherwise it is 0. By calculating the cumulative distribution function of each non-overlapping block, the local CDF matrix can be obtained to describe the distribution of pixel values in each block. Based on the local CDF matrix, the mapping function of each pixel is calculated using bilinear interpolation to generate a global mapping matrix. The bilinear interpolation method can smoothly extend the local CDF information to the entire image, so that each pixel value can be adjusted according to the cumulative distribution characteristics of the block in which it is located. The formula for calculating the mapping function of each pixel is:

[0050] I mapped (x,y)=I enhanced(x, y) · (1 + α · (CDF(x, y, I enhanced (x, y)) - 0.5));

[0051] wherein, I mapped (x, y) is the pixel value after mapping, α is an adjustment coefficient used to control the mapping intensity, and CDF(x, y, I enhanced (x, y)) is the cumulative probability of the original pixel value in its local block. By remapping the gray value of each pixel through this formula, the contrast of the image is effectively enhanced, and multiple video frames with enhanced contrast are obtained. Perform color space conversion on the video frames with enhanced contrast, converting the RGB color space to the YUV color space. The three channels in the RGB color space respectively represent the red, green, and blue color components, while in the YUV color space, the Y channel represents the luminance information, and the U and V channels represent the chrominance information. The color space conversion formula is:

[0052]

[0053] wherein, Y, U, and V are respectively the luminance and chrominance components, and R, G, and B are respectively the red, green, and blue components. After converting to the YUV color space, only apply histogram equalization to the Y channel (luminance channel) to uniformize the luminance distribution and enhance the luminance contrast of the image. Histogram equalization stretches the luminance value distribution of the original image to the entire luminance range, making the luminance distribution more uniform, thereby improving the visual effect of the image. After the processing is completed, convert the YUV image back to the RGB color space to restore the color information and obtain multiple target video frames.

[0054] In a specific embodiment, the process of executing step S2 may specifically include the following steps:

[0055] Perform multi-scale transformation on multiple target video frames, generate three feature maps of different scales through bilinear interpolation method, which are 1 / 8, 1 / 16, and 1 / 32 of the original input size respectively, to obtain multi-scale input feature maps; input the multi-scale input feature maps into the backbone network of the improved YOLOv5 model in sequence. A spatial attention module is added after each convolutional block in this backbone network. The spatial attention module emphasizes important spatial regions by generating a two-dimensional attention weight map to obtain a preliminary feature map containing spatial attention information; perform channel attention processing on the preliminary feature map. First, calculate the global average pooling value of each channel, then input the pooling result into a multi-layer perceptron composed of two fully connected layers, where the number of neurons in the first layer is 1 / 16 of the number of channels, and the second layer restores to the original number of channels. Finally, obtain the channel weights through the Sigmoid activation function, and multiply the channel weights element-wise with the preliminary feature map to obtain an enhanced feature map containing both spatial and channel attention information; input the enhanced feature map into the Feature Pyramid Network, and perform feature fusion through the top-down path and lateral connections. Among them, the top-down path uses 1x1 convolution to adjust the number of channels and 2x upsampling, and the lateral connections use 1x1 convolution to process the features of the same layer. Finally, obtain a multi-scale fusion feature map; apply a 3x3 convolutional kernel to the multi-scale fusion feature map for feature extraction while keeping the size of the feature map unchanged, and then use a 1x1 convolutional kernel to compress the number of channels to the preset number of anchor boxes multiplied by (5 + the number of classes), where 5 represents the 4 coordinate values and 1 confidence value of the bounding box, to obtain the original prediction result; decode the original prediction result, convert the predicted bounding box coordinates from the offset relative to the grid to the absolute coordinates of the image, specifically by multiplying the predicted relative coordinates by the stride of the feature map and adding the upper left corner coordinates of the corresponding grid. At the same time, apply the Sigmoid function to process the confidence and class probabilities, and map them to between 0 and 1 to obtain the decoded prediction result; based on the decoded prediction result, use the non-maximum suppression algorithm to filter out overlapping detection boxes. First, sort all detection boxes in descending order of confidence, then compare them one by one, calculate the intersection over union (IoU) of the overlapping boxes. When the IoU is greater than the preset threshold of 0.5, keep the box with the highest confidence and delete other overlapping boxes to obtain the human detection result.

[0056] Specifically, perform multi-scale transformation on the target video frame. Assume the size of the original input image is H×W, and use the bilinear interpolation method to generate feature maps of 1 / 8, 1 / 16, and 1 / 32 of the original size respectively, to obtain three feature maps of different scales. For the generation of each feature map, the formula of the bilinear interpolation method is:

[0057]

[0058] where, I scaled (x,y) represents the scale after transformation as W scaled ×Hscaled The pixel value at coordinates (x, y) in the feature map, I orig (x′, y′) represents the pixel value at coordinates (x′, y′) in the original input image, W scaled and H scaled respectively represent the width and height of the scaled image. Through this formula, the original input image is scaled to different scales to capture different features of the target at different resolutions. The feature maps at these three different scales respectively correspond to 1 / 8, 1 / 16, and 1 / 32 of the original input size, aiming to extract multi-level features of the target at different scales to handle complex scenarios such as target size and perspective changes. The generated multi-scale input feature maps are sequentially input into the backbone network of the improved YOLOv5 model. A spatial attention module is added after each convolutional block in this backbone network. The spatial attention module emphasizes important spatial regions in the image by generating a two-dimensional attention weight map. The spatial attention module generates weights according to the importance of different positions in the input feature map, and the calculation formula is:

[0059]

[0060] where, Attention spatial (x, y) represents the weight at (x, y) in the spatial attention weight map, Conv 3×3 represents a 3x3 convolution operation, I input is the input feature map, and σ is the Sigmoid activation function. This module captures the local information of the input feature map through convolution operations and maps the weights to between 0 and 1 through the Sigmoid activation function to adjust the importance of each position. Channel attention processing is performed on the preliminary feature map containing spatial attention information, and different weights are assigned to each channel according to its importance. The average value of each channel is calculated through global average pooling, and the pooling formula is:

[0061]

[0062] where, Pool avg (c) represents the global average pooling value of channel c, I input (x, y, c) is the pixel value of the input feature map at channel c and coordinates (x, y). The pooling result is input into a multi-layer perceptron composed of two fully connected layers. The number of neurons in the first layer is 1 / 16 of the original number of channels, effectively reducing the number of parameters and increasing the non-linear expression ability of the model; the number of neurons in the second layer is restored to the original number of channels. The channel weights are calculated through the Sigmoid activation function:

[0063] Attention channel (c) = σ9MLP(Pool avg (c)));

[0064] Among them, Attention channel (c) is the weight of channel c, and MLP represents the multi-layer perceptron operation. Multiply the weight with the preliminary feature map element-wise to obtain an enhanced feature map that contains both spatial and channel attention information. Input the enhanced feature map into the feature pyramid network for feature fusion. The feature pyramid network fuses feature information at different scales through a top-down path and lateral connections. In the top-down path, 1x1 convolutions are used to adjust the number of channels and perform 2x upsampling; in the lateral connections, 1x1 convolutions are used to process the feature maps from the backbone network. The role of the feature pyramid network is to fuse high-resolution detail information with low-resolution semantic information to obtain a multi-scale fused feature map. Apply a 3x3 convolutional kernel to the multi-scale fused feature map for feature extraction to keep the spatial size of the feature map unchanged. The 3x3 convolutional kernel can effectively capture local spatial information and extract richer features while keeping the size of the feature map unchanged. Use a 1x1 convolutional kernel to compress the number of channels to the preset number of anchor boxes multiplied by 5 + the number of classes, where 5 represents the 4 coordinate values of the bounding box and 1 confidence value. The 4 coordinate values of the bounding box represent the center coordinates, width, and height of the target box, and the confidence value represents the confidence of the model in the existence of the target. After the operation, the original prediction result is obtained. Decode the original prediction result, and the predicted bounding box coordinates are converted from the offsets relative to the grid to the absolute coordinates of the image. The decoding calculation formula is:

[0065]

[0066] Among them, (x,y) are the absolute coordinates of the center of the bounding box, w and h are the width and height of the bounding box respectively, (t x ,t y ) are the relative offsets predicted by the model, (c x ,c y ) are the coordinates of the upper left corner of the grid, s is the stride of the feature map, (p w ,p h ) are the width and height of the anchor box, and σ(·) is the Sigmoid function. Through this formula, the relative coordinates are converted into the absolute coordinates in the image, and through and Convert the width and height to real values, and at the same time apply the Sigmoid function to map the confidence and class probabilities to between 0 and 1 to obtain the decoded prediction results. Based on the decoded prediction results, use the non-maximum suppression algorithm to filter out overlapping detection boxes. The non-maximum suppression sorts all detection boxes in descending order of confidence, then compares the overlap between them one by one, and calculates the intersection over union (IoU). When the IoU is greater than a preset threshold (e.g., 0.5), keep the box with the highest confidence and delete other overlapping boxes. Thus, duplicate detections are effectively removed to obtain the final human detection results.

[0067] In a specific embodiment, the process of performing step S3 may specifically include the following steps:

[0068] Apply the face key point detection algorithm based on the multi-task cascaded convolutional network to the face region in the human detection result to locate 5 key points, including the center of the left eye, the center of the right eye, the tip of the nose, the left corner of the mouth, and the right corner of the mouth, to obtain the face key point coordinates; based on the face key point coordinates, calculate the affine transformation matrix to transform the face image to the standard position, so that the centers of the two eyes are on the horizontal line and at a fixed position in the image, to obtain the aligned face image; crop and scale the aligned face image, adjust the aligned face image to a size of 112x112 pixels, and perform pixel value normalization to scale the pixel values to the interval [-1,1] to obtain the standardized face image; input the standardized face image into the SE-ResNet-50 backbone network, the SE-ResNet-50 backbone network contains 50 convolutional layers, and a Squeeze-and-Excitation module is added after each residual block to learn the correlation between channels through global average pooling and two fully connected layers to obtain the face feature map; apply the self-attention mechanism to the face feature map, calculate the correlation between each position in the face feature map and all other positions, generate the attention weight matrix, multiply the attention weight matrix by the face feature map to obtain the enhanced feature map; input the enhanced feature map into the global average pooling layer to reduce the spatial dimension of the enhanced feature map to 1x1, retain the channel dimension, obtain a 2048-dimensional feature vector, and apply a fully connected layer to the 2048-dimensional feature vector to map the 2048-dimensional feature vector to the 512-dimensional space to obtain the face feature vector.

[0069] Specifically, crop and preprocess the face region in the human detection result. Assume that the bounding box coordinates (x1, y1, x2, y2) of the face region are obtained from the human detection result, where (x1, y1) represents the upper left corner coordinates of the face bounding box, and (x2, y2) represents the lower right corner coordinates. Crop the face region from the original image to obtain a sub-image I containing the face. face. To avoid the boundary being uneven or the head part being cropped, the bounding box is appropriately enlarged during cropping. For example, the width and height are each increased by 10%. The face sub-image I face is input into the face key point detection model based on the Multi-task Cascaded Convolutional Networks (MTCNN). The MTCNN is a model with a three-level network structure, namely the Proposal Network, the Refinement Network, and the Output Network. They successively perform step-by-step refined detection and key point localization on the face region. The Proposal Network performs an initial scan of the input image to generate face candidate boxes and rough coordinates of five key points; the Refinement Network further corrects and filters the candidate boxes and refines the coordinates of the five key points; the Output Network performs accurate regression of the final face bounding box and key point coordinates on the filtered candidate boxes. Through the step-by-step refinement process of these three-level networks, the MTCNN can accurately locate the five key points of the face in complex scenes: the center of the left eye, the center of the right eye, the tip of the nose, the left corner of the mouth, and the right corner of the mouth. The face image is aligned by transforming it to a standardized position and pose. The alignment process is achieved by calculating the affine transformation matrix. The affine transformation is a linear transformation that preserves the parallelism of lines and includes geometric operations such as translation, rotation, scaling, and shearing. To calculate the affine transformation matrix, a standard key point template is defined. Assume that the coordinates of the centers of the two eyes in the standard template are (x le , y le ) and (x re , y re ), and the coordinates of the tip of the nose, the left corner of the mouth, and the right corner of the mouth are (x′ nose , y′ nose ), (x′ lm , y′ lm ), and (x′ rm , y′ rm ) respectively. According to the actual key point coordinates in the target image and the key point coordinates in the standard template, the affine transformation matrix M is calculated. The calculation of the affine transformation matrix uses the least squares method to solve the linear equations. Assume that the key point coordinates in the target image are (x i , y i ), and the corresponding coordinates in the standard template are (x′ i , y′ i ). The transformation matrix M can be expressed as:

[0070]

[0071] where m 11 , m 12 , m 13 , m 21 , m 22 , m 23They are the parameters of the affine transformation matrix. By substituting the actual coordinates and standard coordinates into this system of equations and solving using the least squares method, the affine transformation matrix M is obtained. According to the transformation matrix, the face image I face is subjected to an affine transformation to align the face image to the position of the center of the two eyes horizontally in the standard template. The aligned face image is cropped and scaled. It is cropped to only contain the face region and adjusted to the standard size of 112x112 pixels. The pixel values are normalized by scaling the pixel values to the range [-1, 1], eliminating the influence of lighting conditions on the image, and at the same time improving the numerical stability of the model during training and inference. The formula for pixel value normalization is:

[0072]

[0073] where I norm (x, y) is the normalized pixel value, and I resized (x, y) is the pixel value at (x, y) of the resized standardized face image. Through the normalization formula, the image with original pixel values in the range [0, 255] is mapped to the interval [-1, 1] to obtain the standardized face image. The standardized face image is input into the SE-ResNet-50 backbone network for feature extraction. SE-ResNet-50 is an improved residual neural network structure with a Squeeze-and-Excitation (SE) module added after each residual block. The SE module models the global information of each channel and learns the correlation between different channels. The implementation of the SE module is divided into two steps: Squeeze and Excitation. The Squeeze operation compresses the spatial information of each channel into a scalar through global average pooling, that is:

[0074]

[0075] where z c represents the pooling result of the c-th channel, and I feat (i, j, c) is the pixel value at (i, j) of the c-th channel in the feature map, and H and W are the height and width of the feature map respectively. The Excitation operation then learns the importance weights of each channel through two fully connected layers:

[0076] s c = σ(W2·ReLU(W1·z c ));

[0077] where s cDenote the weight of the c-th channel, where W1 and W2 are the weight matrices of the fully connected layers respectively, and σ represents the Sigmoid activation function. By multiplying the weights with the original feature map channel by channel, the enhanced feature map is obtained. The SE module can improve the attention distribution among channels in the model, thereby enhancing the feature expression ability. Apply the self-attention mechanism to the face feature map. By calculating the correlation between each position in the feature map and all other positions, an attention weight matrix is generated, and an enhanced feature map containing global context information is obtained. Input the enhanced feature map into the global average pooling layer to reduce its spatial dimension to 1x1, only retaining the channel dimension. Through the global average pooling operation, the spatial information of the feature map is compressed into global information, while retaining the feature information in the channel dimension. Assume that after global average pooling, a 2048-dimensional feature vector is obtained. Next, apply a fully connected layer to it to map the 2048-dimensional feature vector to a 512-dimensional space, obtaining a face feature vector.

[0078] In a specific embodiment, the process of executing step S4 may specifically include the following steps:

[0079] Establish a state vector for each target in the humanoid detection result, including the center coordinates, width, height, speed, and acceleration of the target. Initialize the state vector and covariance matrix of the Kalman filter to obtain an initial state estimate; based on the initial state estimate, use the state transition matrix to predict the position of the target in the next frame to obtain a prior state estimate, and calculate the matching cost matrix according to the prior state estimate and the humanoid detection result of the current frame. Use the Hungarian algorithm for data association to obtain the matching pairs of the target and the detection result; for each target in the matching pairs, calculate the cosine similarity between the face feature vectors, and combine the position information to obtain the target identity confirmation result, and update the state estimate of the Kalman filter based on the target identity confirmation result and the current observation value to obtain a posterior state estimate; adaptively adjust the observation noise and process noise in the posterior state estimate. By calculating the sample covariance of the observation residual sequence of the last A frames and using the exponentially weighted average method to update the noise covariance matrix, an adaptive noise covariance matrix is obtained; substitute the adaptive noise covariance matrix into the prediction and update equations of the Kalman filter to recalculate the Kalman gain and state estimate to obtain an optimized state estimate; for the unmatched detection results, initialize new trajectories; for the trajectories that have not been matched for A consecutive frames, use a long short-term memory network for trajectory extrapolation to predict the positions of the next A frames to obtain the multi-target humanoid tracking result.

[0080] Specifically, extract the current detection information of each target. Assume that the detection result of a certain target in the t-th frame includes the center coordinates (x t , y t ), width w t and height h t, and its state vector is represented as:

[0081]

[0082] Among them, and respectively represent the velocity components of the target center coordinates, and respectively represent the change rates of the target width and height. The initial state estimate of the Kalman filter includes the state vector x t and the covariance matrix P t . The state vector describes the motion state of the target, while the covariance matrix is used to describe the uncertainty of the state. The initial covariance matrix P0 is set as a diagonal matrix, and its diagonal elements represent the initial uncertainties of each state component:

[0083]

[0084] Among them, respectively represent the uncertainties of position and size, and respectively represent the uncertainties of the velocity components. Through the initialization method, the motion state of each target is modeled. Based on the initial state estimate, the state transition matrix is used to predict the position of the target in the next frame. The state transition model of the Kalman filter is expressed as:

[0085] x t+1 = Fx t + Bu t + w t ;

[0086] Among them, F is the state transition matrix, u t is the control input vector, B is the control input matrix, w t is the process noise, and it is usually assumed that w t ~ N(0, Q t ). In the case of no control input, the control input matrix B can be ignored. The state transition matrix F defines the dynamic evolution law of the state vector. For example, in a simple uniform motion model, the state transition matrix is set as:

[0087]

[0088] Among them, Δt represents the time interval between two adjacent frames. The prior state estimate of the next frame is obtained through the state transition matrix:

[0089]

[0090] At the same time, the predicted covariance matrix is:

[0091]

[0092] Among them, Q t is the process noise covariance matrix. The prior state estimate describes the predicted state of the target without the current observation. Based on the prior state estimate and the human detection results of the current frame, the matching cost matrix is calculated. The element d ij of the matching cost matrix D represents the matching cost between the i-th predicted target and the j-th detected target, which is jointly determined by the position deviation and the size difference:

[0093]

[0094] Among them, (x i , y i ) and (x j , y j ) represent the center coordinates of the predicted target and the detected target respectively, w i , h i and w j , h j represent the width and height of the predicted target and the detected target respectively, and α is the importance weight used to balance the position and size differences. Through the formula, the prior state estimate and the current observation are combined to calculate the matching cost between each pair of targets. The Hungarian algorithm is used for data association of the matching cost matrix. The Hungarian algorithm is an optimization algorithm for solving the bipartite graph matching problem, which can find the combination with the minimum cost among all possible matching pairs. The matching pairs of the target and the detection result are obtained through the Hungarian algorithm, that is, it is determined which detection results correspond to which tracking targets. For each target in the matching pair, the cosine similarity between its face feature vectors is calculated. Suppose the feature vector of the i-th target is f i , and the feature vector of the j-th detection result is f j , the cosine similarity is expressed as:

[0095]

[0096] Among them, cos(θ ij ) represents the cosine similarity between the two feature vectors, f i ·f j represents their inner product, ||f i || and ||f j || represent their norms respectively. The value range of the cosine similarity is between [-1, 1], and the closer it is to 1, the more similar their directions are and the more similar the features are. Combining the position information and the cosine similarity, the identity confirmation result of the target is obtained. Based on the target identity confirmation result and the current observation, the state estimate of the Kalman filter is updated to obtain the posterior state estimate. The update calculation formula is:

[0097]

[0098] Among them, z t+1 is the current observation value, H is the observation matrix, usually the identity matrix, and K t+1 is the Kalman gain, and its calculation formula is:

[0099]

[0100] Among them, R t+1 is the observation noise covariance matrix. The updated covariance matrix is:

[0101]

[0102] The observation value is fused into the predicted state through the update formula to obtain a more accurate posterior state estimate. The observation noise and process noise in the posterior state estimate are adaptively adjusted. By calculating the sample covariance of the observation residual sequence of the most recent A frames and using the exponential weighted average method to update the noise covariance matrix. Let the observation residual sequence of the most recent A frames be r t , and its sample covariance is:

[0103]

[0104] Among them, is the mean of the observation residual sequence. The observation noise covariance matrix R t is updated using the exponential weighted average method:

[0105] R t = βR t-1 +(1 - β)S t ;

[0106] Among them, β is the smoothing coefficient, usually taking values between 0.8 and 0.95. This method enables the noise covariance matrix to be continuously adaptively adjusted during the tracking process, so as to better adapt to the noise changes in different environments. Substitute the adaptive noise covariance matrix into the prediction and update equations of the Kalman filter, recalculate the Kalman gain and state estimate, and obtain the optimized state estimate. The method of dynamically adjusting the noise covariance can effectively improve the stability and tracking accuracy of the Kalman filter in complex environments. For the unmatched detection results, new trajectories are initialized; for those trajectories that have not been matched for consecutive A frames, it may be that the target has disappeared briefly or been occluded. In order to continue tracking when the target reappears, a long short-term memory network (LSTM) is used for trajectory extrapolation. LSTM is a neural network model that can process time series data. By memorizing the historical information of the previous time steps, it predicts the positions of future time steps. Predict the target positions of the future A frames through the trained LSTM model to obtain a more complete multi-target humanoid tracking result.

[0107] In a specific embodiment, the process of executing step S5 may specifically include the following steps:

[0108] Perform spatio-temporal segmentation on the multi-target humanoid tracking results, divide the tracking results of consecutive B frames into a task unit to obtain multiple task units, and predict the future load of each node based on the CPU usage rate, memory occupancy rate, and network bandwidth of each computing node to obtain node load prediction values; according to the node load prediction values and the complexity of the task units, perform preliminary allocation of the task units to obtain an initial task allocation plan, and perform load balancing optimization on the initial task allocation plan. Through task migration and splitting operations, control the load difference of each node within a preset threshold to obtain an optimized task allocation plan; send the optimized task allocation plan to each computing node, and build a local feature index on each node. Use a hash table to store the face feature vectors and the corresponding trajectory IDs to obtain a local feature index structure; build a global feature index on the central control node, and use a hierarchical tree structure to organize the feature information of each node. Each leaf node corresponds to the local index of a computing node to obtain a global feature index structure; for cross-node face matching requests, first perform a rough match in the global feature index to locate the most likely K leaf nodes, and then perform an exact match in the local indexes of these K leaf nodes to obtain cross-node matching results; based on the cross-node matching results, associate and fuse the trajectory segments on different nodes, eliminate duplicate trajectories, and perform global consistency processing on the trajectory IDs to obtain global humanoid tracking results.

[0109] Specifically, perform block processing on the multi-target humanoid tracking data according to the time dimension. Assume that in the entire tracking data, the tracking result of each target contains T frames, and the tracking information of each frame is represented as t i ={x i , y i , w i , h i , id i}, where x i and y i are the center coordinates of the target respectively, w i and h i are the width and height of the target respectively, and id i is the unique identifier of the target. In order to divide the tracking data into multiple task units, each task unit should contain consecutive B-frame tracking results to ensure good continuity of the task units in time. The divided task unit is represented as T k ={t kB , t kB+1 ,…, t( k+1 ) B-1}, where k represents the number of the task unit, and the value range of k is For example, for a segment of humanoid tracking data containing 300 frames, if B = 50 is selected, it is divided into 6 task units, and each task unit contains the tracking results of 50 frames. Based on the resource information such as the CPU usage rate, memory occupancy rate, and network bandwidth of each computing node, predict the future load of each node to obtain the load prediction value of the node. Let the CPU usage rate, memory occupancy rate, and network bandwidth of the j-th node at the current moment t be CPU j (t), Mem j (t), and Net j (t). Use a linear regression model to predict the load situation at the future moment t + Δt. The expression of the linear regression model is:

[0110]

[0111] where represents the load prediction value of the j-th node at the future moment, α1, α2, α3 are regression coefficients, and β is the bias term. By fitting the historical data, estimate the values of the regression coefficients and the bias term to obtain the load prediction value of each node, reflecting the resource occupancy of each node at the future moment. According to the node load prediction value and the complexity of the task unit, make a preliminary allocation of the task unit. The complexity of the task unit is determined by factors such as the number of targets it contains, the motion complexity of the targets, and the data volume. Assume that the complexity of the k-th task unit is C k , and it is calculated by the following formula:

[0112]

[0113] where N k is the number of targets in the task unit, (x i , y i ) is the center coordinate of the target in each frame, representing the motion complexity of the target, D k is the data volume of the task unit, which is related to factors such as the image resolution and the number of frames, and γ1, γ2, γ3 are weight coefficients. Measure the complexity of each task unit through the formula to provide a basis for task allocation. The preliminary allocation scheme is implemented through a greedy algorithm, and preferentially allocate the task units with greater complexity to the nodes with lighter loads. Let the preliminary allocation scheme be M = {(k, j)}, indicating that the k-th task unit is allocated to the j-th computing node. To optimize the preliminary allocation scheme, perform load balancing optimization. Through task migration and task splitting operations, control the load difference of each node within a preset threshold. Let the load of each node be L j , and the goal is to make the loads of all nodes satisfy:

[0114]

[0115] Among them, δ is a preset load difference threshold. Task migration is to transfer some task units from the node with a heavier load to the node with a lighter load, and task splitting is to decompose the task units with higher complexity into multiple sub-task units with lower complexity for parallel execution on different nodes. After the task migration and splitting operations, a task allocation scheme M with optimized load balancing is obtained. opt After the optimized task allocation scheme is sent to each computing node, a local feature index is constructed on each node. The local feature index is used to store face feature vectors and corresponding track IDs for efficient face matching and retrieval within the node. Assume that there are N different tracks on each node, and each track contains M feature vectors. A hash table is used to store the mapping relationship between feature vectors and track IDs. Through this storage method of the hash table, fast retrieval and matching of feature vectors are realized. On the central control node, a global feature index structure is constructed. The global feature index adopts a hierarchical tree structure to organize the feature information of each node. Each leaf node corresponds to the local index of a computing node. The root node of the tree structure represents the feature information of the entire distributed system, the internal nodes represent the feature information of sub-regions or sub-nodes, and the leaf nodes store the actual feature data. The construction of the tree structure is achieved through a bottom-up clustering algorithm. The feature information of each computing node is aggregated into a leaf node, and then adjacent leaf nodes are merged into an internal node, and so on, until the entire tree structure is constructed. When processing a face matching request across nodes, a rough match is performed in the global feature index to quickly locate the K leaf nodes that are most likely to contain the target feature. Let the input feature vector be f, and the similarity between each leaf node and this feature vector is calculated by the following formula:

[0116]

[0117] Among them, S j represents the maximum cosine similarity between the input feature vector f and all feature vectors in the j-th leaf node, and θ ij represents the angle between the input feature vector and the i-th feature vector in the leaf node. According to the size of the similarity S j , the top K leaf nodes are selected for exact matching. When performing exact matching in the local indexes of the selected K leaf nodes, the cosine similarity between the input feature vector and each local feature vector is calculated one by one to find the matching result with the highest similarity, effectively reducing the search space and improving the matching efficiency. Based on the cross-node matching results, the track segments on different nodes are associated and fused. The association of track segments is achieved by comparing their spatio-temporal overlap and feature similarity. Assume that the time overlap of two track segments is Δt, the space overlap is Δs, and the feature similarity is cos(θ). The association degree of the track segments is expressed as:

[0118] R = α·Δt + β·Δs + γ·cos(θ);

[0119] Among them, α, β, and γ are the weight coefficients of the time overlap degree, space overlap degree, and feature similarity degree respectively. According to the magnitude of the correlation degree R, it is determined whether two trajectory segments belong to the same target. For the trajectory segments determined to be the same target, trajectory fusion is performed, and the trajectory points of different segments are merged into a complete trajectory. Global consistency processing is performed on the trajectory ID, and a unique global ID is assigned to each target to ensure that the same target on different nodes has a unique identifier in the entire system, and the global human tracking result is obtained.

[0120] In a specific embodiment, the process of executing step S6 may specifically include the following steps:

[0121] Group multiple target video frames according to the global human tracking result, divide the continuously tracked human targets into the same video segment to obtain multiple video segments, and each video segment contains E frames of images; apply the ResNext-101 network to each video frame in the multiple video segments. The ResNext-101 network contains 32 residual blocks with a cardinality of 4, and each residual block outputs 256-dimensional features. Finally, a 2048-dimensional spatial feature vector is obtained through global average pooling, and then an E×2048-dimensional spatial feature sequence is obtained; apply the self-attention mechanism to the spatial feature sequence to calculate the E×E attention weight matrix, where each element represents the correlation between frames, and then multiply the attention weight matrix by the original feature sequence to obtain the weighted spatial feature sequence; calculate the difference image between each video frame and the previous frame in the multiple video segments, use the inter-frame difference method to obtain E - 1 residual noise images, and stack these images into an E - 1×H×W×3 residual noise sequence, where H and W are the height and width of the image respectively; input the residual noise sequence into a 3D convolutional neural network. The 3D convolutional neural network contains 4 3D convolutional layers, the convolutional kernel size of each layer is 3×3×3, the stride is 1×2×2, and the number of channels is 64, 128, 256, and 512 in sequence. Finally, a 512-dimensional temporal feature vector is obtained through 3D global average pooling, and then an E×512-dimensional temporal feature sequence is obtained; perform a bilinear pooling operation on the weighted spatial feature sequence and the temporal feature sequence to obtain an E×d-dimensional fused feature sequence; based on the fused feature sequence, use a two-layer fully connected neural network to regressively calculate the relative position offset, rotation angle, and scaling factor between adjacent video segments to obtain video stitching parameters; according to the video stitching parameters, perform an affine transformation on the multiple video segments, use bilinear interpolation for image resampling, and then stitch the transformed images according to the relative positions to obtain the stitched video.

[0122] Specifically, analyze the global human tracking results. The global human tracking results usually include the trajectory data of the target, where each trajectory consists of the target detection boxes of a series of frames and the corresponding timestamps. Divide the consecutive frames into video segments, set a fixed number of frames E as the length of each video segment, and obtain multiple consecutive video segments containing E frames. To extract the spatial features of each frame in each video segment, use the ResNext-101 network. ResNext-101 is a deep convolutional neural network that contains 32 residual blocks with a cardinality of 4, and each residual block outputs 256-dimensional features. The expression form of each residual block can be represented as:

[0123]

[0124] where y is the output of the residual block, x is the input feature, F(x) represents the convolutional operation, and each residual block contains multiple parallel branches, each branch using a different convolutional kernel. Through the stacking of residual blocks, ResNext-101 can extract richer image features. Through the global average pooling operation, map the features of each frame into a 2048-dimensional feature vector. The formula for global average pooling is:

[0125]

[0126] where, f i is the 2048-dimensional feature vector of the i-th frame, f ijk represents the feature value of the feature map at the position (j,k) in the i-th frame, and H and W are the height and width of the feature map respectively. Through this formula, compress the feature map of each frame into a global feature vector, and obtain a spatial feature sequence of E×2048 dimensions. To capture the temporal correlation between the frames in the video segment, apply the self-attention mechanism to the spatial feature sequence. The self-attention mechanism can capture the dependence relationship between the elements in the sequence, and represent the inter-frame correlation by calculating the attention weight matrix of E×E. The calculation formula of the attention weight matrix is:

[0127]

[0128] where, α ij represents the attention weight between the i-th frame and the j-th frame, q i =W q f i and k j =W k f j represent the query vector and key vector of the i-th frame and the j-th frame respectively, W q and W kis a learnable linear transformation matrix. Through this formula, each frame in the original feature sequence is associated with other frames to obtain a weighted spatial feature sequence. To capture the dynamic information in the video clip, the difference image between each video frame and the previous frame is calculated. The difference image reflects the motion changes between frames in the video. Using the inter-frame difference method, E - 1 residual noise images are obtained. The formula for the inter-frame difference method is:

[0129] ΔI t =I t -I t-1 ;

[0130] where, ΔI t represents the difference image between the t-th frame and the (t - 1)-th frame, and I t and I t-1 represent the images of the t-th frame and the (t - 1)-th frame respectively. The difference images are stacked into a residual noise sequence of (E - 1)×H×W×3, where H and W are the height and width of the image respectively. The residual noise sequence is input into a 3D convolutional neural network for temporal feature extraction. The 3D convolutional neural network can extract features simultaneously in the spatial and temporal dimensions and is suitable for processing spatio-temporal information in videos. Suppose the 3D convolutional neural network contains 4 3D convolutional layers, the size of the convolutional kernel in each layer is 3×3×3, the stride is 1×2×2, and the number of channels is 64, 128, 256, 512 in sequence. The formula for 3D convolution is:

[0131]

[0132] where, f i,j,k represents the feature value at the position (i, j, k) of the output feature map of the 3D convolutional layer, W d,m,n represents the weight of the 3D convolutional kernel at the depth d, height m, and width n, and I i+d-1,j+m-1,k+n-1 represents the value at the corresponding position of the input feature map. D, M, N represent the sizes of the convolutional kernel in the depth, height, and width respectively. Through the stacking of 3D convolutional layers, rich temporal features in the video clip are extracted. Finally, through 3D global average pooling, the output feature map is reduced to a 512-dimensional temporal feature vector, obtaining a temporal feature sequence of E×512 dimensions. To fuse the spatial features and temporal features of the video clip, a bilinear pooling operation is performed on the weighted spatial feature sequence and the temporal feature sequence. Bilinear pooling is a feature fusion method that can capture the interaction information between two different feature vectors. The formula for bilinear pooling is:

[0133]

[0134] where, f fusion,t represents the fused feature of the t-th frame, f space,t and f time,trespectively represent the spatial feature and the temporal feature of the t-th frame, represents the element-wise multiplication operation. After the bilinear pooling operation, an E×d-dimensional fused feature sequence is obtained, where d is the dimension of the fused feature. Based on the fused feature sequence, a two-layer fully connected neural network is used to perform regression calculations on the relative position offset, rotation angle, and scaling factor between adjacent video segments. Let the relative position offset between adjacent video segments P1 and P2 be (Δx,Δy), the rotation angle be θ, and the scaling factor be s. The input of the fully connected neural network is the fused feature vector, and the output is the stitching parameters (Δx,Δy,θ,s). The calculation formula of the regression network is:

[0135]

[0136] where, W fcl and W fc2 respectively represent the weight matrices of the first and second fully connected layers, b fc1 and b fc2 respectively represent the bias terms, and σ(·) represents the non-linear activation function. Through the two-layer fully connected network, the fused feature is mapped to the stitching parameter space to obtain the stitching parameters of adjacent video segments. According to the video stitching parameters, affine transformation is performed on multiple video segments. Affine transformation is a linear transformation that performs translation, rotation, and scaling simultaneously. Let the coordinates of the input image be (x,y), and the transformed coordinates be (x′,y′). The formula for affine transformation is:

[0137]

[0138] where, s represents the scaling factor, θ represents the rotation angle, and Δx and Δy respectively represent the position offset. Through this formula, each pixel point in the image is mapped to the new coordinates after transformation. In order to maintain the details of the image during the transformation process, bilinear interpolation is used to resample the image. The formula for bilinear interpolation is: I′(x′,y′) = (1 - Δx)(1 - Δy)I(x1,y1) + Δx(1 - Δy)I(x2,y1) + (1 - Δx)ΔyI(x1,y2) + ΔxΔyI(x2,y2). Where, I′(x′,y′) represents the pixel value after resampling, I(x1,y1), I(x2,y1), I(x1,y2), and I(x2,y2) respectively represent the four adjacent pixel values, and Δx and Δy respectively represent the coordinate offsets of the fractional parts. The image is smoothed through bilinear interpolation, and then the transformed images are stitched according to the relative positions to obtain the final stitched video.

[0139] The stitching method of the surveillance video in the embodiment of the present invention is described above. Next, the stitching device of the surveillance video in the embodiment of the present invention will be described. Please refer to Figure 2, an embodiment of the monitoring video stitching device in the embodiments of the present invention includes:

[0140] A frame division module for dividing the monitoring video collected by the monitoring terminal into video frames to obtain a plurality of target video frames;

[0141] A detection module for inputting the plurality of target video frames into an improved YOLOv5 model for human detection to obtain a human detection result;

[0142] An extraction module for aligning and cropping the face regions in the human detection result to obtain a standardized face image and extracting face feature vectors;

[0143] A prediction module for performing motion prediction and trajectory extrapolation based on the face feature vectors and the human detection result to obtain a multi-target human tracking result;

[0144] A tracking module for inputting the multi-target human tracking result into a distributed computing framework for task allocation and cross-node feature matching to obtain a global human tracking result;

[0145] A calculation module for performing two-stream network processing on the plurality of target video frames according to the global human tracking result, calculating video stitching parameters, and performing a stitching operation on the plurality of target video frames to obtain a stitched video.

[0146] Through the collaborative cooperation of the above-mentioned various components, human detection is performed by an improved YOLOv5 model, and the spatial attention and channel attention mechanisms are introduced to improve the target detection accuracy and robustness in complex scenarios. The improved FaceNet network with the SE-ResNet-50 backbone network and the self-attention mechanism is used to extract face features, enhancing the discriminability of face features, which is beneficial to subsequent target tracking and identity recognition. The Kalman filter with an adaptive noise covariance matrix is used for motion prediction and trajectory extrapolation, improving the accuracy and stability of multi-target tracking, especially in the case of short-term target occlusion. The introduction of a distributed computing framework and a dynamic task allocation mechanism realizes the efficient processing of large-scale video data, significantly improving the real-time performance and scalability of the system. Through the two-stream network processing method, the spatial semantic features and temporal residual noise information are comprehensively considered, improving the accuracy and coherence of video stitching, and effectively solving the problems of breakage and repetition of dynamic targets during the stitching process. The present invention effectively improves the quality and visual effect of the stitched video.

[0147] The present invention also provides a computer device, which includes a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the processor executes the steps of the monitoring video stitching method in the above-mentioned embodiments.

[0148] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the method for stitching surveillance videos.

[0149] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0150] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0151] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for splicing surveillance videos, characterized in that, The method includes: Performing video frame segmentation on the monitoring video collected by the monitoring terminal to obtain a plurality of target video frames; Inputting the plurality of target video frames into an improved YOLOv5 model for human detection to obtain human detection results. Specifically, it includes: performing multi-scale transformation on the plurality of target video frames, generating three different-scale feature maps through bilinear interpolation method, which are 1 / 8, 1 / 16, and 1 / 32 of the original input size respectively, to obtain multi-scale input feature maps; sequentially inputting the multi-scale input feature maps into the backbone network of the improved YOLOv5 model, and a spatial attention module is added after each convolutional block in this backbone network. The spatial attention module emphasizes important spatial regions by generating a two-dimensional attention weight map to obtain a preliminary feature map containing spatial attention information; performing channel attention processing on the preliminary feature map. First, calculate the global average pooling value of each channel, then input the pooling result into a multi-layer perceptron composed of two fully connected layers, where the number of neurons in the first layer is 1 / 16 of the number of channels, and the second layer restores to the original number of channels. Finally, obtain the channel weights through the Sigmoid activation function, multiply the channel weights element-wise with the preliminary feature map to obtain an enhanced feature map containing both spatial and channel attention information; inputting the enhanced feature map into a feature pyramid network, and performing feature fusion through a top-down path and lateral connections. The top-down path uses 1×1 convolution to adjust the number of channels and 2x upsampling, and the lateral connections use 1×1 convolution to process the features of the same layer, finally obtaining a multi-scale fusion feature map; applying a 3×3 convolution kernel to the multi-scale fusion feature map for feature extraction while keeping the size of the feature map unchanged, and then using a 1×1 convolution kernel to compress the number of channels to the sum of the preset number of anchor boxes multiplied by 5 and the number of classes, where 5 represents the 4 coordinate values and 1 confidence value of the bounding box, to obtain the original prediction result; decoding the original prediction result, converting the predicted bounding box coordinates from the offset relative to the grid to the absolute coordinates of the image, specifically by multiplying the predicted relative coordinates by the stride of the feature map and adding the upper left coordinate of the corresponding grid, and at the same time applying the Sigmoid function to process the confidence and class probabilities, mapping them to between 0 and 1 to obtain the decoded prediction result; based on the decoded prediction result, using the non-maximum suppression algorithm to filter overlapping detection boxes. First, sort all detection boxes in descending order of confidence, then compare them one by one, calculate the intersection over union of the overlapping boxes, and when the intersection over union is greater than the preset threshold of 0.5, retain the box with the highest confidence and delete other overlapping boxes to obtain the human detection results; Aligning and cropping the face region in the human detection results to obtain a standardized face image, and extracting a face feature vector; Performing motion prediction and trajectory extrapolation based on the face feature vector and the human detection results to obtain multi-target human tracking results; Inputting the multi-target human tracking results into a distributed computing framework for task assignment and cross-node feature matching to obtain global human tracking results; Perform two-stream network processing on the multiple target video frames according to the global human form tracking result, calculate video stitching parameters, and perform a stitching operation on the multiple target video frames to obtain a stitched video.

2. The splicing method of the monitoring video according to claim 1, characterized in that, Performing video frame splitting on the surveillance video collected by the surveillance terminal to obtain multiple target video frames, including: Collecting a surveillance video through a surveillance terminal, sampling the surveillance video at a preset frame rate to obtain multiple initial video frames, and performing resolution normalization processing on each frame of the multiple initial video frames to obtain multiple normalized video frames; Performing bilateral filtering on the multiple normalized video frames, where the spatial domain kernel function radius is set to 5 and the range domain kernel function standard deviation is set to 50, to obtain multiple preliminarily denoised video frames; Performing edge-preserving filtering on the multiple preliminarily denoised video frames to obtain multiple edge-enhanced video frames, dividing the edge-enhanced video frames into non-overlapping 8×8 blocks, and calculating the cumulative distribution function for each block to obtain a local CDF matrix; Based on the local CDF matrix, calculating the mapping function for each pixel using bilinear interpolation to obtain a global mapping matrix, and remapping the gray values of the edge-enhanced video frames at the pixel level according to the global mapping matrix to obtain multiple contrast-enhanced video frames; Performing color space conversion on the multiple contrast-enhanced video frames, converting the RGB color space to the YUV color space, only applying histogram equalization to the Y channel, and then converting the processed YUV image back to the RGB color space to obtain multiple target video frames.

3. The method for splicing surveillance videos according to claim 1, wherein Aligning and cropping the face region in the human form detection result to obtain a normalized face image, and extracting a face feature vector, including: Applying a face key point detection algorithm based on a multi-task cascaded convolutional network to the face region in the human form detection result to locate 5 key points, including the left eye center, right eye center, nose tip, left mouth corner, and right mouth corner, to obtain face key point coordinates; Calculating an affine transformation matrix based on the face key point coordinates to transform the face image to a standard position, such that the two eye centers are on a horizontal line and at a fixed position in the image, to obtain an aligned face image; Cropping and scaling the aligned face image, adjusting the aligned face image to a size of 112×112 pixels, and performing pixel value normalization, scaling the pixel values to the range [-1, 1], to obtain a normalized face image; Inputting the normalized face image into a SE-ResNet-50 backbone network, the SE-ResNet-50 backbone network contains 50 convolutional layers, adding a Squeeze-and-Excitation module after each residual block, and learning the correlation between channels through global average pooling and two fully connected layers to obtain a face feature map; Applying a self-attention mechanism to the face feature map, calculating the correlation between each position in the face feature map and all other positions, generating an attention weight matrix, and multiplying the attention weight matrix by the face feature map to obtain an enhanced feature map; Input the enhanced feature map into the global average pooling layer, reduce the spatial dimension of the enhanced feature map to 1×1 while retaining the channel dimension, obtain a 2048-dimensional feature vector, and apply a fully connected layer to the 2048-dimensional feature vector to map the 2048-dimensional feature vector to a 512-dimensional space, obtaining a face feature vector.

4. The method for splicing surveillance videos according to claim 1, characterized in that, Based on the face feature vector and the humanoid detection result, perform motion prediction and trajectory extrapolation to obtain a multi-object humanoid tracking result, including: Establish a state vector for each target in the humanoid detection result, including the center coordinates, width, height, speed, and acceleration of the target, and initialize the state vector and covariance matrix of the Kalman filter to obtain an initial state estimate; Based on the initial state estimate, use the state transition matrix to predict the position of the target in the next frame to obtain a prior state estimate, and calculate a matching cost matrix according to the prior state estimate and the humanoid detection result in the current frame. Use the Hungarian algorithm for data association to obtain the matching pairs of the target and the detection result; For each target in the matching pairs, calculate the cosine similarity between the face feature vectors, and combine the position information to obtain the target identity confirmation result. Based on the target identity confirmation result and the current observation value, update the state estimate of the Kalman filter to obtain a posterior state estimate; Adaptively adjust the observation noise and process noise in the posterior state estimate. By calculating the sample covariance of the observation residual sequence in the most recent A frames and using the exponential weighted average method to update the noise covariance matrix, obtain an adaptive noise covariance matrix; Substitute the adaptive noise covariance matrix into the prediction and update equations of the Kalman filter, recalculate the Kalman gain and state estimate, and obtain an optimized state estimate; For the unmatched detection results, initialize new trajectories; for the trajectories that have not been matched for consecutive A frames, use a long short-term memory network for trajectory extrapolation to predict the positions in the next A frames, obtaining a multi-object humanoid tracking result.

5. The splicing method of the monitoring video according to claim 1, characterized in that Input the multi-object humanoid tracking result into a distributed computing framework for task assignment and cross-node feature matching to obtain a global humanoid tracking result, including: Perform spatio-temporal segmentation on the multi-object humanoid tracking result, divide the tracking results of consecutive B frames into a task unit to obtain multiple task units, and predict the future load of each node based on the CPU usage rate, memory occupancy rate, and network bandwidth of each computing node to obtain a node load prediction value; According to the node load prediction value and the complexity of the task unit, perform a preliminary assignment of the task unit to obtain an initial task assignment scheme, and perform load balancing optimization on the initial task assignment scheme. Through task migration and splitting operations, control the load difference of each node within a preset threshold to obtain an optimized task assignment scheme; Send the optimized task assignment scheme to each computing node, and construct a local feature index on each node. Use a hash table to store the face feature vector and the corresponding trajectory ID to obtain a local feature index structure; Build a global feature index at the central control node, organize the feature information of each node using a hierarchical tree structure, where each leaf node corresponds to the local index of a computing node, and obtain the global feature index structure; For a face matching request across nodes, first perform a rough match in the global feature index to locate the most likely K leaf nodes, and then perform an exact match in the local indexes of these K leaf nodes to obtain the cross-node matching result; Based on the cross-node matching result, associate and fuse the trajectory segments on different nodes, eliminate duplicate trajectories, and perform global consistency processing on the trajectory IDs to obtain the global humanoid tracking result.

6. The splicing method of the surveillance video according to claim 1, wherein The double-stream network processing of the multiple target video frames according to the global humanoid tracking result to calculate the video stitching parameters and perform a stitching operation on the multiple target video frames to obtain a stitched video includes: Group the multiple target video frames according to the global humanoid tracking result, divide the continuously tracked humanoid targets into the same video segment to obtain multiple video segments, and each video segment contains E frame images; Apply the ResNext-101 network to each video frame in the multiple video segments. The ResNext-101 network includes 32 residual blocks with a cardinality of 4, each residual block outputs 256-dimensional features, and finally obtain a 2048-dimensional spatial feature vector through global average pooling, and then obtain an E×2048-dimensional spatial feature sequence; Apply the self-attention mechanism to the spatial feature sequence to calculate an E×E attention weight matrix, where each element represents the correlation between frames, and then multiply the attention weight matrix by the original feature sequence to obtain a weighted spatial feature sequence; Calculate the difference image between each video frame and the previous frame in the multiple video segments, use the inter-frame difference method to obtain E - 1 residual noise images, and stack these images into an E - 1×H×W×3 residual noise sequence, where H and W are the height and width of the image respectively; Input the residual noise sequence into a 3D convolutional neural network. The 3D convolutional neural network includes 4 3D convolutional layers, the convolutional kernel size of each layer is 3×3×3, the stride is 1×2×2, and the number of channels is 64, 128, 256, 512 in sequence. Finally, obtain a 512-dimensional temporal feature vector through 3D global average pooling, and then obtain an E×512-dimensional temporal feature sequence; Perform a bilinear pooling operation on the weighted spatial feature sequence and the temporal feature sequence to obtain an E×d-dimensional fused feature sequence; Based on the fused feature sequence, use a two-layer fully connected neural network to regressively calculate the relative position offset, rotation angle, and scaling factor between adjacent video segments to obtain the video stitching parameters; According to the video stitching parameters, perform an affine transformation on the multiple video segments, use bilinear interpolation for image resampling, and then stitch the transformed images according to the relative positions to obtain the stitched video.

7. A splicing device for monitoring videos, characterized in that, A monitoring video stitching device for executing the monitoring video stitching method according to any one of claims 1 - 6, the monitoring video stitching device includes: A frame division module, configured to perform video frame division on the monitoring video collected by the monitoring terminal to obtain a plurality of target video frames; A detection module, configured to input the plurality of target video frames into an improved YOLOv5 model for human detection to obtain a human detection result; An extraction module, configured to align and crop the face region in the human detection result to obtain a standardized face image, and extract a face feature vector; A prediction module, configured to perform motion prediction and trajectory extrapolation based on the face feature vector and the human detection result to obtain a multi-target human tracking result; A tracking module, configured to input the multi-target human tracking result into a distributed computing framework for task allocation and cross-node feature matching to obtain a global human tracking result; A calculation module, configured to perform two-stream network processing on the plurality of target video frames according to the global human tracking result, calculate video stitching parameters, and perform a stitching operation on the plurality of target video frames to obtain a stitched video.

8. A computer device, characterized in that, It includes a memory and a processor, the memory stores a computer program that can run on the processor, and is characterized in that when the processor executes the computer program, the stitching method of the monitoring video according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, the processor is caused to execute the stitching method of the monitoring video according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Hidden danger violation identification method and device for explosion-related video

    CN114898181A

  • Pedestrian multi-target tracking method based on deep learning

    CN117237411A