Movement detection method and system based on frame difference algorithm
By using a motion detection method based on frame difference algorithm, which utilizes Y component data dimensionality reduction and dynamic threshold judgment, the accuracy and cost problems of motion detection in existing technologies are solved, and efficient and low-cost motion target detection is achieved in complex environments.
Patent Information
- Application Number
- CN202511118766.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing motion detection methods lack accuracy and stability in complex environments, and have high hardware costs, making them difficult to run in real time on embedded devices. In particular, they have high false detection and false negative rates in scenarios with changing day and night lighting.
A motion detection method based on frame difference algorithm is adopted. By extracting the Y component data of adjacent frame images for data dimensionality reduction processing, a frame difference information matrix is established. Then, the moving target is judged by enabling mask and dynamic threshold, which reduces the computational complexity and hardware cost.
It improves the accuracy and stability of motion detection, reduces computational complexity and hardware costs, is suitable for cost-sensitive monitoring scenarios, and achieves optimized control of equipment costs.
Smart Images

Figure CN120976269A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of motion detection technology, and more specifically, to a motion detection method and system based on frame difference algorithm. Background Technology
[0002] With the rapid development of intelligent security, traffic monitoring, and other fields, real-time detection of moving targets in various scenarios has become a critical requirement. In complex environments, such as urban roads, public places, or nighttime scenes, it is essential to accurately identify moving targets such as vehicles, pedestrians, and animals, while simultaneously achieving low false alarm and low false negative rates. However, while pursuing high detection accuracy, equipment cost is also a crucial factor that must be considered. Therefore, how to reduce hardware investment while ensuring detection effectiveness has become an important direction for the development of current mobile detection technologies.
[0003] Existing motion detection solutions largely rely on complex hardware computing resources, employing high-performance processors or dedicated chips to implement image algorithm processing, such as optical flow and background modeling. Specifically, optical flow detects motion by calculating pixel motion vectors, but its high computational complexity makes it difficult to run in real-time on embedded devices. Background modeling relies on static background assumptions, making it sensitive to changes in lighting and dynamic backgrounds (such as swaying leaves), which can easily lead to false detections, limiting its application in large-scale or cost-sensitive scenarios. Furthermore, most motion detection methods do not adequately consider data biases caused by factors such as changes in day and night lighting when processing image data, affecting the stability and accuracy of the detection results.
[0004] Therefore, an optimized motion detection method and system based on the frame difference algorithm is desired. Summary of the Invention
[0005] To address the aforementioned technical problems, this application is proposed. Embodiments of this application provide a motion detection method and system based on a frame difference algorithm. For adjacent first and second frame images, the Y component data of both are extracted to avoid color channel interference. Data dimensionality reduction is then performed to map the high-dimensional image matrix into a low-dimensional matrix, reducing computational complexity. Subsequently, frame difference information is calculated on the Y component data of the two adjacent frames after dimensionality reduction to establish a frame difference information matrix. Enabled regions within the frame difference information matrix are delineated based on an enable mask. The frame difference values of these regions are compared with dynamically adjusted preset thresholds to determine the presence of a moving target. This method leverages the sensitivity of the Y component to brightness changes, effectively improving detection accuracy. Simultaneously, data dimensionality reduction reduces computational complexity and hardware costs, making it suitable for cost-sensitive monitoring scenarios requiring precise motion detection. It can optimize equipment costs while ensuring detection performance.
[0006] Accordingly, according to one aspect of this application, a motion detection method based on a frame difference algorithm is provided, comprising:
[0007] Acquire adjacent first and second frame images;
[0008] Y component data of the first frame image and the second frame image are extracted respectively to obtain Y component data of the first frame and Y component data of the second frame;
[0009] Data mapping and dimensionality reduction are performed on the first frame Y component data and the second frame Y component data to obtain the first frame Y component dimensionality reduction data and the second frame Y component dimensionality reduction data.
[0010] Calculate the frame difference information matrix between the first frame Y component dimensionality-reduced data and the second frame Y component dimensionality-reduced data;
[0011] Moving target detection is performed based on the frame difference information matrix to obtain moving target detection results.
[0012] According to another aspect of this application, a motion detection system based on a frame difference algorithm is provided, comprising:
[0013] The adjacent image frame acquisition module is used to acquire adjacent first and second frame images;
[0014] The Y component extraction module is used to extract the Y component data of the first frame image and the second frame image respectively to obtain the Y component data of the first frame and the Y component data of the second frame.
[0015] The data dimensionality reduction module is used to perform data mapping and dimensionality reduction on the first frame Y component data and the second frame Y component data to obtain the first frame Y component dimensionality reduction data and the second frame Y component dimensionality reduction data.
[0016] The frame difference calculation module is used to calculate the frame difference information matrix between the first frame Y component dimensionality reduction data and the second frame Y component dimensionality reduction data;
[0017] The moving target detection module is used to perform moving target detection based on the frame difference information matrix to obtain the moving target detection result.
[0018] Compared with existing technologies, the motion detection method and system based on frame difference algorithm provided in this application extracts the Y component data of adjacent first and second frame images to avoid color channel interference, and maps the high-dimensional image matrix to a low-dimensional matrix through data dimensionality reduction to reduce computational complexity. Subsequently, frame difference information is calculated on the Y component data of the two adjacent frames after dimensionality reduction to establish a frame difference information matrix. Enabled regions in the frame difference information matrix are delineated based on an enable mask. The presence of a moving target is determined by comparing the frame difference values of these regions with a dynamically adjusted preset threshold. This method utilizes the sensitivity of the Y component to brightness changes, effectively improving detection accuracy. Simultaneously, data dimensionality reduction reduces computational complexity and hardware costs, making it suitable for cost-sensitive monitoring scenarios requiring precise motion detection. It can optimize equipment costs while ensuring detection performance. Attached Figure Description
[0019] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0020] Figure 1 This is a flowchart of a motion detection method based on a frame difference algorithm according to an embodiment of this application.
[0021] Figure 2 This is a schematic diagram of the data flow of a motion detection method based on the frame difference algorithm according to an embodiment of this application.
[0022] Figure 3 This is a flowchart of step S2 in the motion detection method based on the frame difference algorithm according to an embodiment of this application.
[0023] Figure 4 This is a flowchart of step S5 in the motion detection method based on the frame difference algorithm according to an embodiment of this application.
[0024] Figure 5 This is a flowchart of step S52 in the motion detection method based on the frame difference algorithm according to an embodiment of this application.
[0025] Figure 6 This is a flowchart of step S523 in the motion detection method based on the frame difference algorithm according to an embodiment of this application.
[0026] Figure 7 This is a block diagram of a motion detection system based on a frame difference algorithm according to an embodiment of this application. Detailed Implementation
[0027] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein. It is worth noting that in this application, all data acquisition actions are performed in accordance with the relevant data protection laws and policies of the country where the application is located, and with authorization from the owner of the relevant device.
[0028] Figure 1 This is a flowchart of a motion detection method based on a frame difference algorithm according to an embodiment of this application. Figure 2 This is a schematic diagram of the data flow of a motion detection method based on a frame difference algorithm according to an embodiment of this application. Figure 1 and Figure 2 As shown, the motion detection method based on the frame difference algorithm according to an embodiment of this application includes the following steps: S1, acquiring adjacent first frame images and second frame images; S2, extracting Y component data from the first frame image and the second frame image respectively to obtain first frame Y component data and second frame Y component data; S3, performing data mapping and dimensionality reduction on the first frame Y component data and the second frame Y component data to obtain first frame Y component dimensionality-reduced data and second frame Y component dimensionality-reduced data; S4, calculating the frame difference information matrix between the first frame Y component dimensionality-reduced data and the second frame Y component dimensionality-reduced data; S5, performing motion target detection based on the frame difference information matrix to obtain a motion target detection result.
[0029] In the aforementioned motion detection method based on frame difference algorithm, step S1 involves acquiring adjacent first and second frame images. It should be understood that the temporal correlation between consecutive frames in a video stream is the foundation of motion detection; that is, motion can be identified by comparing the differences between two consecutive frames. Therefore, to establish the input source for inter-frame difference analysis, this application, based on the principle of temporal continuity in video streams, uses a sensor to acquire images in real time and processes them through an ISP (Image Signal Processor) to generate a continuous sequence of video frames. Adjacent first and second frame images are then extracted from this sequence for motion detection. In specific implementation, after initialization, the sensor captures images at a preset frame rate (e.g., 30fps). The ISP performs denoising, white balance, and color correction, outputting a standardized RGB format video stream to ensure the stability and consistency of subsequent data processing. This method enables the real-time acquisition of adjacent frames, providing high-quality raw data for subsequent motion detection based on frame difference analysis and avoiding detection errors caused by image distortion or delay.
[0030] In the aforementioned motion detection method based on the frame difference algorithm, step S2 involves extracting the Y component data from the first frame image and the second frame image to obtain the first frame Y component data and the second frame Y component data. It should be understood that color information in RGB image data is easily affected by changes in illumination and noise, especially in day-night transition scenes where the stability of chromaticity information is poor. Therefore, to avoid the negative impact of color channels on motion detection, this application leverages the robustness of the luminance component (Y component) under illumination changes. By extracting the Y component data from the first frame image and the second frame image as the core processing object, it focuses on luminance changes, simplifies the amount of data processed subsequently, and highlights key motion-related information.
[0031] Figure 3 This is a flowchart of step S2 in the motion detection method based on the frame difference algorithm according to an embodiment of this application. Figure 3 As shown, step S2 includes: S21, extracting YUV data from the first frame image and the second frame image to obtain first frame YUV data and second frame YUV data; S22, performing noise reduction processing on the first frame YUV data and the second frame YUV data to obtain first frame noise-reduced YUV data and second frame noise-reduced YUV data; S23, extracting the Y component from the first frame noise-reduced YUV data and the second frame noise-reduced YUV data to obtain first frame Y component data and second frame Y component data.
[0032] Specifically, in step S21, the YUV data of the first frame image and the second frame image are extracted to obtain first frame YUV data and second frame YUV data. It should be understood that the YUV color space can separate the luminance component (Y) from the chrominance component (UV), making it more suitable for focusing on luminance information sensitive to illumination changes in motion detection. Therefore, in order to obtain standardized input to meet the algorithm's requirements, this application converts the first frame image and the second frame image from the RGB color space to the YUV color space to obtain first frame YUV data and second frame YUV data. In the YUV color space, the luminance component Y can be processed independently of the chrominance components U and V, which not only reduces processing complexity but also enhances adaptability to illumination changes.
[0033] Specifically, in step S22, the first frame YUV data and the second frame YUV data are denoised to obtain the first frame denoised YUV data and the second frame denoised YUV data. It should be understood that noise in the YUV data (such as sensor thermal noise and random noise under low light) and color temperature abrupt changes in day-night switching scenes can significantly affect the stability of the luminance component. Therefore, in order to eliminate the negative impact of noise on motion detection and improve data reliability, this application uses a joint temporal and spatial denoising technique, employing multi-frame cumulative filtering and spatial smoothing algorithms to preprocess the first frame YUV data and the second frame YUV data to reduce noise interference. Specifically, for the first frame YUV data and the second frame YUV data, temporal denoising techniques (such as 3D denoising) are first used, employing a weighted average of UV component information from multiple frames to suppress random noise; simultaneously, spatial Gaussian filtering or median filtering is applied to the UV channel of the current frame to eliminate isolated noise points. For the Y channel, lightweight spatial denoising (such as edge-preserving filtering) is used to avoid excessive smoothing that could lead to blurred motion edges. In this way, the first and second frames of denoised YUV data, after noise reduction processing, retain luminance details while significantly reducing noise interference in the chroma channel, providing a high signal-to-noise ratio input for subsequent Y component extraction.
[0034] Specifically, in step S23, the Y component is extracted from the first frame's denoised YUV data and the second frame's denoised YUV data to obtain the first frame's Y component data and the second frame's Y component data. It should be understood that since motion detection relies heavily on brightness changes, and the UV component is prone to introducing irrelevant chromaticity fluctuations in day / night lighting differences or complex backgrounds, this application, based on the principle that the luminance component (Y) dominates motion representation, separates the Y channel and discards the UV component, compressing multi-channel data into a single-channel grayscale matrix to obtain the first frame's Y component data and the second frame's Y component data. In this way, the amount of data processed per frame is reduced from three channels (YUV) to a single channel (Y), reducing computational complexity while ensuring the integrity of motion-sensitive information, providing efficient and stable input for subsequent dimensionality reduction and frame difference analysis.
[0035] In the aforementioned motion detection method based on frame difference algorithm, step S3 involves data mapping and dimensionality reduction of the first frame Y component data and the second frame Y component data to obtain dimensionality-reduced first frame Y component data and dimensionality-reduced second frame Y component data. In a specific example of this application, step S3 includes: dividing the dimensionality-reduced first frame Y component data into regions to obtain a predetermined number of sub-regions; calculating the average value of multiple sub-elements within each sub-region in the predetermined number of sub-regions to obtain the dimensionality-reduced first frame Y component data. Specifically, directly calculating frame difference on the high-dimensional matrix of the original image (such as 1920×1080 pixels in a 1080p image) requires a large amount of computing power and is difficult to process in real time in embedded devices. Therefore, in order to reduce the computational burden while retaining key motion information, this application further performs spatial dimensionality reduction and local feature aggregation on the first frame Y component data and the second frame Y component data. By dividing the high-dimensional Y component matrix into a predetermined number of sub-regions and calculating the average value of all elements within each sub-region, data compression is achieved to obtain the dimensionality-reduced first frame Y component data and the dimensionality-reduced second frame Y component data. In one specific embodiment of this application, the predetermined number of sub-regions is 5×5 sub-regions, and the first frame Y component dimensionality-reduced data is a 5×5 matrix. That is, the first frame Y component data and the second frame Y component data are divided from the original M×N matrix into 5×5 sub-regions. Within each sub-region, a single value representing that region is generated through mean calculation, thereby mapping the original matrix into a 5×5 low-dimensional matrix. In this way, while preserving the characteristics of motion-sensitive regions, the computational load is reduced to 1 / ((M / 5)×(N / 5)) of the original size, significantly improving algorithm efficiency.
[0036] In the aforementioned motion detection method based on the frame difference algorithm, step S4 involves calculating the frame difference information matrix between the first frame's Y component dimensionality-reduced data and the second frame's Y component dimensionality-reduced data. In a specific example of this application, step S4 includes: calculating the positional difference between the first frame's Y component dimensionality-reduced data and the second frame's Y component dimensionality-reduced data to obtain the frame difference information matrix, wherein the size of the frame difference information matrix is 5×5. It should be understood that the essence of moving target detection is the difference in brightness distribution between adjacent frames. By comparing the pixel value differences at corresponding positions in two adjacent frames, the movement of objects in the scene can be captured, thereby reflecting information such as the position and contour of the moving target. Therefore, this application further constructs the frame difference information matrix based on the frame difference algorithm by calculating the positional difference between the first frame's Y component dimensionality-reduced data and the second frame's Y component dimensionality-reduced data. The frame difference information matrix has the same dimension as the first frame Y component dimensionality reduction data and the second frame Y component dimensionality reduction data, with a size of 5×5. Each element in the matrix represents the average brightness change in the corresponding sub-region between two adjacent frames, which can clearly show the changes between the two frames, highlight the position and general outline of the moving target in the image, and provide core data support for subsequent accurate judgment of the moving target.
[0037] In the above-described motion detection method based on the frame difference algorithm, step S5 involves performing motion target detection based on the frame difference information matrix to obtain the motion target detection result. Wherein, Figure 4 This is a flowchart of step S5 in the motion detection method based on the frame difference algorithm according to an embodiment of this application. Figure 4 As shown, step S5 includes: S51, determining the enabled region in the frame difference information matrix based on the enabled mask; S52, determining whether there is a moving target in the enabled region based on the comparison between the frame difference value of the enabled region and a preset threshold.
[0038] Specifically, step S51 determines the enabled region in the frame difference information matrix based on the enabled mask. It should be understood that, considering the potential interference variations in non-critical regions (such as fixed backgrounds or edge noise) within the frame difference information matrix, directly analyzing the global matrix would increase computational load and easily lead to misjudgments (such as leaf swaying in a dynamic scene). Therefore, to focus on the effective area that actually needs monitoring, improve detection efficiency, and eliminate interference from irrelevant areas, this application introduces an enabled mask (a pre-defined region of interest) to delineate the effective detection range within the frame difference information matrix. In specific implementation, an enabled mask (i.e., a 5×5 matrix) matching the size of the frame difference information matrix is generated in advance according to the actual needs of the monitoring scene (such as the location of key monitoring areas). Regions with an element value of 1 in the mask are defined as enabled regions (i.e., the target area to be detected), and regions with a value of 0 are non-enabled regions (i.e., ignored background areas). By performing element-wise multiplication of the enable mask with the frame difference information matrix, only the frame difference values of the enabled regions are retained, and the data of the non-enabled regions are filtered out. This avoids redundant calculations of the entire matrix and allows analysis to be performed only on preset key regions, significantly reducing the amount of data processing. At the same time, it eliminates interference from fixed backgrounds or non-monitored areas, making the detection process more targeted.
[0039] Specifically, step S52 determines whether a moving target exists in the enabled region based on a comparison between the frame difference value of the enabled region and a preset threshold. More specifically, if the frame difference value of the enabled region is greater than the preset threshold, a moving target is determined to exist in the enabled region; if the frame difference value of the enabled region is less than or equal to the preset threshold, a moving target is determined not to exist in the enabled region. That is, if the frame difference value in the enabled region exceeds the preset threshold, it indicates that there is a significant brightness change in the enabled region, and the region is determined to have a moving target, and the corresponding location is marked as a potential moving target region; conversely, if the frame difference value is less than or equal to the preset threshold, the region is considered to have no obvious motion change and belongs to static background, and is excluded from the list of moving targets. In this way, moving targets in the monitoring scene can be accurately and efficiently identified, providing a reliable foundation for subsequent advanced functions such as target tracking and behavior analysis.
[0040] Specifically, considering that when determining moving targets in the enabled area based on a preset threshold, there may be significant differences in lighting conditions, background complexity, and moving target characteristics under different monitoring scenarios, a fixed preset threshold may be difficult to adapt to all situations. Therefore, to further improve the accuracy and robustness of moving target detection, this application proposes an adaptive threshold adjustment strategy. First, an initial preset threshold is set based on a typical monitoring scenario. Then, by performing contextual correlation analysis on the frame difference information matrix, the current scene features and the dynamic characteristics of the moving target are mined. Based on this, the initial preset threshold is dynamically adjusted to better match actual monitoring needs.
[0041] Figure 5 This is a flowchart of step S52 in the motion detection method based on the frame difference algorithm according to an embodiment of this application. Figure 5 As shown, in step S52, the setting of the preset threshold includes: S521, extracting an initial preset threshold; S522, performing convolutional encoding on the frame difference information matrix to obtain a regional brightness change correlation feature map; S523, performing cross-regional brightness-related semantic enhancement on the regional brightness change correlation feature map to obtain an enhanced regional brightness change correlation feature map; S524, performing feature decoding on the enhanced regional brightness change correlation feature map to obtain a threshold adjustment coefficient; S525, dynamically adjusting the initial preset threshold based on the threshold adjustment coefficient to obtain the preset threshold.
[0042] Specifically, step S521 involves extracting an initial preset threshold. That is, by setting an initial preset threshold as a basic judgment threshold, a benchmark reference is provided for subsequent dynamic adjustments. Here, the initial preset threshold is set based on experience. It is determined by pre-analyzing the frame difference data distribution of typical monitoring scenarios (such as indoor / outdoor scenes with uniform lighting) to initially distinguish moving targets from static backgrounds. This provides a quantifiable basic judgment standard for motion detection, avoiding blind adjustments without reference values.
[0043] Specifically, in step S522, the frame difference information matrix is convolutionally encoded to obtain a regional brightness change correlation feature map. It should be understood that the local difference distribution of the frame difference information matrix implies the spatiotemporal correlation of moving targets, but its original numerical information only reflects pixel-level differences, lacking an abstract representation of regional brightness change patterns, making it difficult to distinguish between real movement and local noise or sudden changes in illumination. Therefore, to deeply explore the correlation of brightness changes between regions, this application, based on the local perception characteristics of convolutional neural networks, transforms the pixel-level data of the frame difference information matrix into regional correlation features through convolutional encoding operations. Specifically, multiple convolutional kernels are used to perform sliding window convolution operations on the frame difference information matrix, with each kernel responsible for capturing brightness change patterns in different directions and scales, and generating corresponding feature maps. By stacking multiple convolutional layers, higher-level regional brightness change correlation features are gradually abstracted, resulting in a multi-channel regional brightness change correlation feature map to reflect the spatiotemporal distribution pattern of moving targets. In this way, the original numerical matrix is transformed into a feature representation that includes spatial contextual relationships, which helps to suppress the interference of noise from individual pixels, highlight the brightness variation trend of continuous regions, and provide a more representative input for subsequent threshold adjustment.
[0044] Specifically, step S523 involves performing cross-regional brightness-related semantic enhancement on the region brightness change correlation feature map to obtain an enhanced region brightness change correlation feature map. It should be understood that, limited by the local characteristics of convolutional operations, although convolutional encoding can capture the correlation of brightness changes between regions, it is still difficult to fully capture global contextual information, potentially misjudging isolated high-frame-difference regions as moving targets. Therefore, to enhance the global semantic understanding of the region brightness change correlation feature map, this application further proposes a multi-directional context-aware cross-regional brightness-related semantic enhancement method. This method aims to enhance the overall understanding of each local region in the region brightness change correlation feature map by fusing contextual information from multiple directions, reducing misjudgments caused by insufficient local information.
[0045] Figure 6 This is a flowchart of step S523 in the motion detection method based on the frame difference algorithm according to an embodiment of this application. Figure 6As shown, step S523 includes: S5231, extracting the channel feature vector at the (h,w)th pixel position from the region brightness change correlation feature map as the region brightness correlation enhancement feature vector; S5232, performing multi-directional context awareness on the region brightness correlation enhancement feature vector to obtain a first-direction context awareness encoding vector, a second-direction context awareness encoding vector, and a third-direction context awareness encoding vector; S5233, performing attention aggregation on the first-direction context awareness encoding vector, the second-direction context awareness encoding vector, and the third-direction context awareness encoding vector to obtain a region brightness change correlation feature enhancement encoding vector, wherein the region brightness change correlation feature enhancement encoding vector is the channel feature vector at the (h,w)th pixel position of the enhanced region brightness change correlation feature map.
[0046] More specifically, step S5231 can be expressed by the formula:
[0047]
[0048]
[0049] in, The brightness variation of the region is associated with a feature map. , , These are the height, width, and number of channels of the feature map associated with the brightness change in the region, respectively. This represents the feature vector to be enhanced associated with the brightness of the region, that is, the channel feature vector at the (h,w) pixel position in the feature map associated with the brightness change of the region.
[0050] It is understandable that, considering the limitations of standard convolutional operations due to the fixed size and isotropic nature of the local receptive field, it is difficult to capture long-distance or specific directional spatial dependencies (such as the continuity of road edges or the trajectory direction of vehicle movement), resulting in a lack of integration of global structured information in feature representations. Therefore, in order to effectively integrate multi-directional contextual information for each pixel position in the region brightness change associated feature map, this application uses local features as semantic anchors and extracts the channel-dimensional feature vectors of each pixel position (h,w) in the region brightness change associated feature map as the objects to be enhanced, thereby achieving pixel-by-pixel contextual enhancement of the region brightness change associated feature map.
[0051] More specifically, step S5232 includes: First, in the region brightness change correlation feature map, taking the region brightness correlation feature vector to be enhanced as the center, feature sampling is performed along a first direction, a second direction, and a third direction to obtain a set of region brightness correlation feature context vectors to be enhanced in the first direction, a set of region brightness correlation feature context vectors to be enhanced in the second direction, and a set of region brightness correlation feature context vectors to be enhanced in the third direction. The region brightness correlation feature vector to be enhanced is located at the center of the set of region brightness correlation feature context vectors to be enhanced in the first direction, the set of region brightness correlation feature context vectors to be enhanced in the second direction, and the set of region brightness correlation feature context vectors to be enhanced in the third direction, as expressed by the formula:
[0052]
[0053]
[0054] in, Indicates the first One direction, In direction The step size vector used when sampling. This represents the set of context vectors for the brightness-related features to be enhanced in the first direction region, the set of context vectors for the brightness-related features to be enhanced in the second direction region, and the set of context vectors for the brightness-related features to be enhanced in the third direction region. This represents the sampling radius, which is the channel feature vector sampled from the center pixel outwards to both sides in each direction, at k pixel positions. If the coordinates exceed the limits, zero values are used for padding. Represents a set The first in Context vectors, For set The number of context vectors in the middle.
[0055] Here, the region brightness-related feature vector to be enhanced encapsulates the brightness change patterns (such as gradients and texture intensity) learned by the convolutional network in a local region, but has not yet incorporated cross-regional contextual association. Considering that moving targets in natural scenes often extend along specific directions (such as vehicles moving horizontally and pedestrians moving vertically), in order to explicitly capture the directional contextual information in the region brightness change-related feature map, this application performs contextual association modeling based on spatial anisotropy. By sequentially sampling the region brightness change-related feature map along a preset direction (such as horizontal, vertical, or diagonal), a multi-directional contextual information set of the region brightness-related feature vector to be enhanced is constructed. Specifically, centered on the region brightness-associated feature vector (h, w), with a predetermined step size, K pixel positions of channel feature vectors are sampled on both sides of the center position along a first direction (e.g., the horizontal direction), forming a first-direction context information set including 2k+1 channel feature vectors (including the center position), i.e., the set of the first-direction region brightness-associated feature context vectors to be enhanced. Similarly, the same operation is performed along the second direction (vertical direction) and the third direction (45° diagonal) to generate three sets of context vectors. Each set contains a center vector and its context vectors extending along the direction, capturing the brightness change trend in different directions (such as the continuity of motion trajectory or the repetition pattern of background texture), which can provide a structured sequence input for subsequent context-aware enhancement.
[0056] Next, the sets of context vectors for the first-direction region brightness-related features to be enhanced, the sets of context vectors for the second-direction region brightness-related features to be enhanced, and the sets of context vectors for the third-direction region brightness-related features to be enhanced are respectively input into the cross-context semantic association perception module based on the converter structure to obtain the first-direction context-aware encoding vector, the second-direction context-aware encoding vector, and the third-direction context-aware encoding vector for the region brightness-related features to be enhanced, which can be expressed by the formula:
[0057]
[0058]
[0059] in, This represents a direction context perceptron based on a converter structure. For the set of orientation context perceptrons based on the converter structure The set of contextual latent state features obtained after context-aware encoding. This represents the first-direction context-aware encoding vector of the region brightness correlation feature to be enhanced, the second-direction context-aware encoding vector of the region brightness correlation feature to be enhanced, and the third-direction context-aware encoding vector of the region brightness correlation feature to be enhanced. Represents a set The first in Each contextual hidden state feature, that is, extracting the set The contextual latent state features at the center position are used as the first direction context-aware encoding vector of the region brightness association enhancement feature, the second direction context-aware encoding vector of the region brightness association enhancement feature, and the third direction context-aware encoding vector of the region brightness association enhancement feature.
[0060] Here, traditional recurrent networks (such as RNNs) struggle to process long sequences in parallel. Therefore, to dynamically model long-range dependencies within directional sequences, this application utilizes the feature interaction principle of self-attention mechanism and employs a Transformer architecture to perform context-aware encoding on the context set for each direction. Specifically, for the set of context vectors representing the brightness-related features to be enhanced in the first direction, the Transformer encoder calculates the correlation weights between all vectors within the set through its internal self-attention layer. Based on these correlation weights, it dynamically aggregates context information, suppresses irrelevant or noisy information, and enhances the perception of the surrounding environment at each pixel location. In the encoded output sequence, the hidden state corresponding to the center position (h, w) is extracted as the first-direction context-aware encoding vector representing the brightness-related features to be enhanced. Similarly, the context sets for the second direction (vertical direction) and the third direction (45° diagonal) are also processed by the Transformer encoder to obtain context-aware encoding vectors for their respective directions. This generates independent context-aware encoding vectors in the three directions. In this way, we can adaptively focus on contextual information that is strongly related to the central features in the focusing sequence (such as the associated brightness changes of distant moving targets), break through the limitations of the local receptive field, and improve the modeling accuracy of directional semantics.
[0061] More specifically, step S5233 includes, firstly, performing multi-dimensional sensitive multilayer feature modulation on the first-direction context-aware encoding vector of the region brightness-related feature to be enhanced, the second-direction context-aware encoding vector of the region brightness-related feature to be enhanced, and the third-direction context-aware encoding vector of the region brightness-related feature to be enhanced, respectively, to obtain the multi-directional modulated first-direction context-aware encoding vector of the region brightness-related feature to be enhanced, the multi-directional modulated second-direction context-aware encoding vector of the region brightness-related feature to be enhanced, and the multi-directional modulated third-direction context-aware encoding vector of the region brightness-related feature to be enhanced, expressed by the formula:
[0062]
[0063]
[0064]
[0065] in, and Directions and direction The embedding representation vector, For direction Relative to direction Attention score This indicates exponentiation with base e. Indicates direction Compared to Attention weights, and , Indicates multi-directional modulation That is, the brightness correlation of the multi-directional modulation region to be enhanced feature. Orientation Context-Aware Encoding Vector
[0066] It is understandable that contextual information from a single direction may be insufficient to fully describe complex motion patterns (e.g., a vehicle moving diagonally needs to combine horizontal and vertical features), and simple concatenation or averaging would introduce information redundancy. Furthermore, considering the correlation between contextual information from each direction and contextual information from other directions, this application first models the correlation between contextual information from each direction and the contextual information from each direction. Then, employing a cross-directional attention mechanism, it modulates the contextual information from each direction separately to enhance the specific features of each direction and suppress redundant information. This yields a first-direction context-aware encoding vector for the brightness-correlation feature to be enhanced in the multi-directional modulation region, a second-direction context-aware encoding vector for the brightness-correlation feature to be enhanced in the multi-directional modulation region, and a third-direction context-aware encoding vector for the brightness-correlation feature to be enhanced in the multi-directional modulation region. In this way, the discriminative components in each direction encoding vector (e.g., the horizontal motion trajectory features of the vehicle) can be adaptively enhanced while suppressing interference from noise or irrelevant background, thus improving the semantic discriminativeness of the directional features.
[0067] Next, the first-direction context-aware encoding vector of the multi-directional modulation region brightness-related features to be enhanced, the second-direction context-aware encoding vector of the multi-directional modulation region brightness-related features to be enhanced, and the third-direction context-aware encoding vector of the multi-directional modulation region brightness-related features to be enhanced are fused to obtain the region brightness change-related feature enhancement encoding vector. In a specific example of this application, it can be expressed by the formula:
[0068]
[0069] in, This represents the enhanced encoding vector of the brightness change associated features in the region.
[0070] In other words, by fusing the first-direction context-aware encoding vector of the multi-directional modulation region brightness-related features to be enhanced, the second-direction context-aware encoding vector of the multi-directional modulation region brightness-related features to be enhanced, and the third-direction context-aware encoding vector of the multi-directional modulation region brightness-related features to be enhanced, multi-directional information is integrated to generate a region brightness change-related feature enhancement encoding vector corresponding to the (h,w) position. This region brightness change-related feature enhancement encoding vector not only includes the brightness change pattern of the local region but also incorporates cross-regional, multi-directional contextual information, possessing both global semantics and local details, thereby significantly improving the modeling capability for complex moving targets and background changes.
[0071] Specifically, in a preferred embodiment of this application, the first-direction context-aware encoding vector of the multi-directional modulation region brightness-related enhanced feature, the second-direction context-aware encoding vector of the multi-directional modulation region brightness-related enhanced feature, and the third-direction context-aware encoding vector of the multi-directional modulation region brightness-related enhanced feature are weighted and fused based on non-isotropic response collaborative optimization to obtain the enhanced encoding vector of the region brightness change-related feature. Here, considering that the context-aware encoding vectors of different directions have directional difference field polarization effects (i.e., significant imbalance in the amplitude of multi-dimensional feature responses) during fusion, resulting in signal distortion during feature interaction (e.g., excessive suppression of effective information in the diagonal direction by horizontal features), thereby reducing the discriminative ability of the enhanced encoding vector of the region brightness change-related feature. Therefore, in order to balance the contribution of multi-directional features and suppress polarization phenomena, this application proposes a non-isotropic response collaborative optimization mechanism. By modeling the dynamic constraint relationship between features in each direction and dynamically adjusting the aggregation weight coefficient, the final enhanced feature can retain dimensional uniqueness and reflect cross-dimensional topological correlation.
[0072] Specifically, before performing position-wise summation on the first-direction context-aware encoding vector of the brightness-related features to be enhanced in the multi-directional modulation region, the second-direction context-aware encoding vector of the brightness-related features to be enhanced in the multi-directional modulation region, and the third-direction context-aware encoding vector of the brightness-related features to be enhanced in the multi-directional modulation region, the context-aware encoding vector for each direction after multi-directional modulation is... Calculate its standardized response relative to the other two directions:
[0073]
[0074] in, The first feature to be enhanced is the brightness correlation of the multi-directional modulation region. Orientation context-aware encoding vector, and , express eigenvectors of the bilinear response;
[0075] Then, under the covariant differential representation of multidimensional analysis, respectively... right Find the partial derivatives:
[0076]
[0077] in, express The directional polarization response adjustment factor. Here, due to the context-aware encoding vectors of the other two dimensions in the denominator position. They are essentially symmetrical, so their partial derivatives are the same. This means that by calculating the partial derivatives, the overall field polarization response exhibits directional difference response calibration.
[0078] In this way, then with and After applying dot product weighting correction, the calculation is as follows:
[0079]
[0080] Based on this, the multi-dimensional field interaction effect under dimension-dependent polarization intensity is effectively improved, enhancing the coupling and fusion efficiency of the first-direction context-aware encoding vector of the brightness-related features to be enhanced in the multi-directional modulation region, the second-direction context-aware encoding vector of the brightness-related features to be enhanced in the multi-directional modulation region, and the third-direction context-aware encoding vector of the brightness-related features to be enhanced in the multi-directional modulation region, thereby improving the enhancement encoding vector of the region brightness change-related features. The accuracy of the expression.
[0081] Specifically, step S524 involves performing feature decoding on the enhanced region brightness change correlation feature map to obtain a threshold adjustment coefficient. That is, to map the high-level semantic features in the enhanced region brightness change correlation feature map into numerical threshold adjustment coefficients. Based on the feature decoding regression principle, this application uses a fully connected layer to perform decoding operations on the enhanced region brightness change correlation feature map, converting the high-level semantic information in the feature map into specific numerical outputs. Specifically, firstly, global average pooling is performed on the enhanced region brightness change correlation feature map to compress the spatial dimension and generate a one-dimensional feature vector, extracting its global features; then, a multi-layer fully connected network (including a non-linear activation function) is used to perform a non-linear transformation on the global features, learning the complex characteristics of the current scene, and mapping it into a real-valued threshold adjustment coefficient (range 0.5~1.5, used for multiplicative adjustment of the initial preset threshold). The threshold adjustment coefficient reflects the complexity of the current scene. When there is a lot of dynamic background interference in the scene, the threshold adjustment coefficient will increase accordingly to raise the detection threshold and reduce false alarms. Conversely, when the scene is relatively simple and the moving target features are obvious, the threshold adjustment coefficient will decrease accordingly to lower the detection threshold and improve detection sensitivity. In this way, a quantitative mapping relationship between scene features and threshold adjustment is established, enabling the moving target detection algorithm to adapt to different monitoring environments and effectively cope with complex and ever-changing scene challenges.
[0082] Specifically, in step S525, the initial preset threshold is dynamically adjusted based on the threshold adjustment coefficient to obtain the preset threshold. That is, the initial preset threshold is combined with the threshold adjustment coefficient through multiplication to generate a real-time preset threshold. This dynamic adjustment mechanism allows the preset threshold to automatically adapt to the characteristics of the current scene (such as light intensity and background dynamism): in scenes with drastic light changes or high background dynamism, the threshold automatically increases to filter interference; in low-light or static background scenes, the threshold automatically decreases to capture weak motion signals. This effectively solves the limitations of fixed thresholds in complex environments, improves the adaptability of the motion detection algorithm to diverse scenes, and thus significantly reduces the false detection rate and false negative rate, making the detection results more reliable.
[0083] In summary, the motion detection method based on the frame difference algorithm according to the embodiments of this application is explained. For adjacent first and second frame images, the Y component data of both are extracted to avoid color channel interference. Data dimensionality reduction is then performed to map the high-dimensional image matrix into a low-dimensional matrix to reduce computational complexity. Subsequently, frame difference information is calculated on the Y component data of the two adjacent frames after dimensionality reduction to establish a frame difference information matrix. Enabled regions in the frame difference information matrix are delineated based on an enable mask. The frame difference values of these regions are compared with dynamically adjusted preset thresholds to determine the presence of a moving target. This method utilizes the sensitivity of the Y component to brightness changes, effectively improving detection accuracy. Simultaneously, data dimensionality reduction reduces computational complexity and hardware costs, making it suitable for cost-sensitive monitoring scenarios requiring precise motion detection. It can optimize equipment costs while ensuring detection performance.
[0084] Furthermore, this application also provides a motion detection system based on a frame difference algorithm.
[0085] Figure 7 This is a block diagram of a motion detection system based on a frame difference algorithm according to an embodiment of this application. Figure 7 As shown, the motion detection system 100 based on the frame difference algorithm according to an embodiment of this application includes: an adjacent image frame acquisition module 110, used to acquire adjacent first frame images and second frame images; a Y component extraction module 120, used to extract the Y component data of the first frame image and the second frame image respectively to obtain first frame Y component data and second frame Y component data; a data dimensionality reduction module 130, used to perform data mapping and dimensionality reduction on the first frame Y component data and the second frame Y component data to obtain first frame Y component dimensionality reduction data and second frame Y component dimensionality reduction data; a frame difference calculation module 140, used to calculate the frame difference information matrix between the first frame Y component dimensionality reduction data and the second frame Y component dimensionality reduction data; and a moving target detection module 150, used to perform moving target detection based on the frame difference information matrix to obtain a moving target detection result.
[0086] Here, those skilled in the art will understand that the specific operations of each module in the above-described motion detection system based on the frame difference algorithm have been described above. Figures 1 to 6 The motion detection method based on the frame difference algorithm has been described in detail, and therefore, its repeated description will be omitted.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A motion detection method based on frame difference algorithm, characterized in that, include: Acquire adjacent first and second frame images; Y component data of the first frame image and the second frame image are extracted respectively to obtain Y component data of the first frame and Y component data of the second frame; Data mapping and dimensionality reduction are performed on the first frame Y component data and the second frame Y component data to obtain the first frame Y component dimensionality reduction data and the second frame Y component dimensionality reduction data. Calculate the frame difference information matrix between the first frame Y component dimensionality-reduced data and the second frame Y component dimensionality-reduced data; Moving target detection is performed based on the frame difference information matrix to obtain moving target detection results.
2. The motion detection method based on frame difference algorithm according to claim 1, characterized in that, Extracting the Y component data of the first frame image and the second frame image respectively to obtain the first frame Y component data and the second frame Y component data includes: Extract the YUV data of the first frame image and the second frame image to obtain the first frame YUV data and the second frame YUV data; Denoising is performed on the first frame of YUV data and the second frame of YUV data to obtain the first frame of denoised YUV data and the second frame of denoised YUV data. Y components are extracted from the first frame of denoised YUV data and the second frame of denoised YUV data to obtain the first frame Y component data and the second frame Y component data.
3. The motion detection method based on frame difference algorithm according to claim 1, characterized in that, Data mapping and dimensionality reduction are performed on the first frame Y component data and the second frame Y component data to obtain the first frame Y component dimensionality-reduced data and the second frame Y component dimensionality-reduced data, including: The first frame of Y component dimensionality-reduced data is divided into regions to obtain a predetermined number of sub-regions; The average value of multiple sub-elements within each of the predetermined number of sub-regions is calculated to obtain the dimensionality reduction data of the Y component of the first frame.
4. The motion detection method based on frame difference algorithm according to claim 3, characterized in that, The predetermined number of sub-regions is 5×5 sub-regions, and the dimensionality reduction data of the Y component of the first frame is a 5×5 matrix.
5. The motion detection method based on frame difference algorithm according to claim 4, characterized in that, Calculating the frame difference information matrix between the first frame Y component dimensionality reduction data and the second frame Y component dimensionality reduction data includes: calculating the positional difference between the first frame Y component dimensionality reduction data and the second frame Y component dimensionality reduction data to obtain the frame difference information matrix, wherein the frame difference information matrix has a size of 5×5.
6. The motion detection method based on frame difference algorithm according to claim 1, characterized in that, Moving target detection is performed based on the frame difference information matrix to obtain moving target detection results, including: The enabled regions in the frame difference information matrix are determined based on the enable mask; Based on the comparison between the frame difference value of the enabled region and a preset threshold, it is determined whether there is a moving target in the enabled region.
7. The motion detection method based on frame difference algorithm according to claim 6, characterized in that, Based on a comparison between the frame difference value of the enabled region and a preset threshold, it is determined whether a moving target exists in the enabled region, including: If the frame difference value of the enabled region is greater than the preset threshold, it is determined that there is a moving target in the enabled region; If the frame difference value of the enabled region is less than or equal to the preset threshold, it is determined that there is no moving target in the enabled region.
8. The motion detection method based on frame difference algorithm according to claim 7, characterized in that, The setting of the preset threshold includes: Extract the initial preset threshold; The frame difference information matrix is convolutionally encoded to obtain a region brightness change correlation feature map; Cross-regional brightness-related semantic enhancement is performed on the aforementioned regional brightness change correlation feature map to obtain an enhanced regional brightness change correlation feature map; The feature map associated with the brightness change in the enhanced region is decoded to obtain the threshold adjustment coefficient; The initial preset threshold is dynamically adjusted based on the threshold adjustment coefficient to obtain the preset threshold.
9. The motion detection method based on frame difference algorithm according to claim 8, characterized in that, Perform cross-regional brightness-related semantic enhancement on the aforementioned regional brightness change correlation feature map to obtain an enhanced regional brightness change correlation feature map, including: The channel feature vector at the (h,w)th pixel position is extracted from the region brightness change correlation feature map as the region brightness correlation feature vector to be enhanced; Multi-directional context awareness is performed on the region brightness-related feature vector to be enhanced to obtain the first-direction context-aware encoding vector, the second-direction context-aware encoding vector, and the third-direction context-aware encoding vector of the region brightness-related feature to be enhanced. Attention aggregation is performed on the first direction context-aware encoding vector of the region brightness-related feature to be enhanced, the second direction context-aware encoding vector of the region brightness-related feature to be enhanced, and the third direction context-aware encoding vector of the region brightness-related feature to be enhanced to obtain the region brightness change-related feature enhancement encoding vector, wherein the region brightness change-related feature enhancement encoding vector is the channel feature vector at the (h,w) pixel position of the enhanced region brightness change-related feature map.
10. A motion detection system based on a frame difference algorithm, characterized in that, include: The adjacent image frame acquisition module is used to acquire adjacent first and second frame images; The Y component extraction module is used to extract the Y component data of the first frame image and the second frame image respectively to obtain the Y component data of the first frame and the Y component data of the second frame. The data dimensionality reduction module is used to perform data mapping and dimensionality reduction on the first frame Y component data and the second frame Y component data to obtain the first frame Y component dimensionality reduction data and the second frame Y component dimensionality reduction data. The frame difference calculation module is used to calculate the frame difference information matrix between the first frame Y component dimensionality reduction data and the second frame Y component dimensionality reduction data; The moving target detection module is used to perform moving target detection based on the frame difference information matrix to obtain the moving target detection result.