Movement detection method and system based on deep learning

By combining deep learning algorithms with a post-processing mechanism based on physical constraints, the calibration threshold is dynamically adjusted, which solves the problem of false detection in motion detection in complex scenarios and improves the accuracy and practicality of detection.

CN120976267AInactive Publication Date: 2025-11-18HANGZHOU CLOSELI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511118765.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing deep learning-based motion detection methods are prone to false detections and misjudgments in complex scenarios, especially under conditions such as changes in lighting and target occlusion, which can lead to false alarms or missed detections, affecting detection efficiency and accuracy.

Method used

By initializing device parameters and configuring deep learning computing power, multi-scale feature analysis is performed to extract the bounding box coordinates and area information of potential targets. The calibration threshold is dynamically fine-tuned in combination with the confidence score to eliminate small-area noise and instantaneous interference, and to determine the effective moving targets.

Benefits of technology

This improves the practicality of motion detection algorithms in security monitoring and intelligent transportation scenarios, reduces false detection rates, and enhances detection accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976267A_ABST
    Figure CN120976267A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of movement detection, and particularly discloses a movement detection method and system based on deep learning, and the method comprises the steps: inputting the obtained to-be-detected image information into a movement detection model for multi-scale feature analysis through initializing equipment parameters and configuring the deep learning computing power; extracting bounding box coordinates and area information of the potential target; and then, based on the confidence score of the potential target belonging to each target category, the calibration threshold is dynamically and finely adjusted, and through comparison between the calibration threshold and the potential target area information, an effective moving target is judged, and small-area noise and instantaneous interference are eliminated. According to the method, the semantic understanding advantage of a deep learning algorithm on complex scenes is reserved, tiny false detection targets are filtered out through a physical constraint post-processing mechanism, the false detection rate is reduced, and therefore the practicability of a mobile detection algorithm in the scenes such as security monitoring and intelligent traffic is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of movement detection, and more particularly, to a movement detection method and system based on deep learning. BACKGROUND

[0002] With the rapid development of intelligent security, traffic monitoring and smart home, etc., the demand for real-time detection and accurate identification of moving targets in the scene is increasingly urgent. For example, in the security scene, it is necessary to efficiently distinguish between pedestrians, vehicles, animals and other moving targets to avoid safety hazards caused by misjudgment or missed detection; in the intelligent traffic system, accurate identification of the dynamic behavior of vehicles and pedestrians is crucial to optimizing traffic management. Traditional movement detection methods rely on image pixel-level change analysis, such as inter-frame difference, background subtraction, etc., and have poor robustness to environmental light changes, camera jitter and background complexity, which can easily cause false positives or false negatives.

[0003] In recent years, with the rapid development of computer vision technology, especially the rise of deep learning technology, deep learning-based movement detection methods have shown great potential in the field of movement detection due to their powerful feature extraction and pattern recognition capabilities. However, existing deep learning-based movement detection methods have made some progress, but they still have false detection and misjudgment problems in the face of multi-scale targets and dynamic background interference in complex scenes, especially in complex conditions such as light changes and target occlusion. For example, non-target objects such as swaying leaves and light refraction artifacts are incorrectly identified as moving targets, resulting in a large number of invalid alarms, which not only increases the cost of manual verification, but also may mask real threats due to interference information, affecting detection efficiency and accuracy.

[0004] Therefore, an optimized deep learning-based movement detection method and system are expected. SUMMARY

[0005] To solve the above technical problems, the present application is proposed. The embodiments of the present application provide a deep learning-based movement detection method and system, which initializes device parameters and configures deep learning computing power, inputs the obtained image information to be detected into a movement detection model for multi-scale feature analysis to extract the bounding box coordinates and area information of potential targets; then, based on the confidence scores of the potential targets belonging to each target category, the calibration threshold is dynamically fine-tuned, and by comparing the calibration threshold with the area information of the potential targets, the effective moving targets are determined, and small-area noise and transient interference are eliminated. This method not only retains the semantic understanding advantage of deep learning algorithms for complex scenes, but also filters out small false detection targets through physical constraint post-processing mechanism, reduces the false detection rate, and thus improves the practicality of the movement detection algorithm in security monitoring, intelligent transportation and other scenes.

[0006] Accordingly, according to an aspect of the present application, there is provided a deep learning-based moving object detection method, comprising:

[0007] S1: device start-up and initialization;

[0008] S2: obtaining first frame image information;

[0009] S3: inputting the first frame image information into a moving object detection algorithm model to obtain a detection result;

[0010] S4: in response to the detection result being a potential moving object, extracting a bounding box coordinate of the potential moving object and calculating an area of the potential moving object based on the bounding box coordinate of the potential moving object;

[0011] S5: if the area of the potential moving object exceeds a designated threshold, marking the potential moving object as a valid moving object; and if the area of the potential moving object does not exceed the designated threshold, marking the potential moving object as invalid information.

[0012] According to another aspect of the present application, there is provided a deep learning-based moving object detection system, comprising:

[0013] an initialization module for device start-up and initialization;

[0014] a first frame image obtaining module for obtaining first frame image information;

[0015] a moving object detection module for inputting the first frame image information into a moving object detection algorithm model to obtain a detection result;

[0016] an area calculation module for, in response to the detection result being a potential moving object, extracting a bounding box coordinate of the potential moving object and calculating an area of the potential moving object based on the bounding box coordinate of the potential moving object;

[0017] a moving object validity determination module for, if the area of the potential moving object exceeds a designated threshold, marking the potential moving object as a valid moving object; and if the area of the potential moving object does not exceed the designated threshold, marking the potential moving object as invalid information.

[0018] Compared with the prior art, the deep learning-based mobile detection method and system provided by the application can initialize device parameters and configure deep learning computing power, input acquired image information to be detected into a mobile detection model for multi-scale feature analysis to extract boundary box coordinates and area information of potential targets, then dynamically fine-tune a calibration threshold based on confidence scores of the potential targets belonging to various target categories, and determine effective mobile targets by comparing the calibration threshold with the area information of the potential targets to eliminate small-area noise and transient interference. The method retains the semantic understanding advantage of deep learning algorithms for complex scenes, filters out small false detection targets through a physical constraint post-processing mechanism, reduces the false detection rate, and thus improves the practicability of the mobile detection algorithm in security monitoring, intelligent transportation and other scenes. BRIEF DESCRIPTION OF DRAWINGS

[0019] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description of embodiments of the present application, when taken in conjunction with the accompanying drawings. The drawings provided in the present application are used to provide further understanding of embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with embodiments of the present application, and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0020] Figure 1 A flowchart of the deep learning-based mobile detection method according to the embodiments of the present application.

[0021] Figure 2 A data flow diagram of the deep learning-based mobile detection method according to the embodiments of the present application.

[0022] Figure 3 A flowchart of step S3 in the deep learning-based mobile detection method according to the embodiments of the present application.

[0023] Figure 4 A flowchart of step S5 in the deep learning-based mobile detection method according to the embodiments of the present application.

[0024] Figure 5 A flowchart of step S53 in the deep learning-based mobile detection method according to the embodiments of the present application.

[0025] Figure 6 A flowchart of step S533 in the deep learning-based mobile detection method according to the embodiments of the present application.

[0026] Figure 7 A block diagram of the deep learning-based mobile detection system according to the embodiments of the present application. DETAILED DESCRIPTION

[0027] Hereinafter, the example embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and the present application is not limited to the described example embodiments. It is worth noting that in the present application, all actions of obtaining data are carried out in accordance with the corresponding data protection regulations and policies of the country where the device is located, and with the authorization given by the owner of the corresponding device.

[0028] Figure 1 Flowchart of the deep learning-based mobile detection method according to the embodiments of the present application. Figure 2 Data flow diagram of the deep learning-based mobile detection method according to the embodiments of the present application. As shown in Figure 1 and Figure 2 The deep learning-based mobile detection method according to the embodiments of the present application includes the following steps: S1, starting and initializing the device; S2, obtaining first frame image information; S3, inputting the first frame image information into a mobile detection algorithm model to obtain a detection result; S4, in response to the detection result being a potential moving target, extracting the bounding box coordinates of the potential moving target and calculating the area of the potential moving target based on the bounding box coordinates of the potential moving target; S5, if the area of the potential moving target exceeds a designated threshold, marking the potential moving target as a valid moving target; and if the area of the potential moving target does not exceed the designated threshold, marking the potential moving target as invalid information.

[0029] In the above deep learning-based mobile detection method, the step S1 is starting and initializing the device. It can be understood that the mobile detection system needs to rely on the collaborative configuration of hardware computing power and software parameters to ensure real-time and accuracy. Therefore, in order to build a stable and efficient detection environment, the present application is based on the principles of hardware resource optimization and parameter preloading, and the initialization process is automatically executed after the device is started to realize the reasonable allocation of system resources. In one specific example of the present application, the step S1 includes: starting the device; initializing parameters; configuring deep learning hardware computing power; and configuring software parameters of the mobile detection algorithm model.

[0030] Specifically, the initialization process includes starting the camera or sensor device, performing system self-check and basic service startup; performing parameter initialization, setting basic running parameters such as image acquisition interface, data buffer size, network communication configuration, etc.; configuring deep learning hardware computing power, such as identifying and specifying the use of specific computing acceleration units (such as GPU, NPU, DSP), allocating the required memory or video memory resources, and loading and optimizing the computation graph; and configuring the software parameters of the mobile detection algorithm model, such as loading the pre-trained model weight file, setting the pre-processing parameters required for model inference (such as image input size, pixel normalization range), and preliminary post-processing parameters (such as area calibration threshold). In this way, resource integration and state calibration can be completed before detection, providing a low-latency, high-throughput running basis for subsequent image processing and model inference, avoiding performance fluctuations caused by hardware resource competition or parameter conflicts.

[0031] In the above deep learning-based mobile detection method, the step S2, the first frame image information is acquired. It should be understood that mobile detection needs to be based on visual data of the actual scene for analysis, and image is the most direct information carrier in computer vision. Therefore, in order to obtain the basic visual information of the scene to be detected, the present application based on image sensing and data acquisition technology captures the visual data of the scene at the current time through calling the image acquisition interface as the first frame image for detection, providing the original input for subsequent target recognition. Specifically, first, a complete digital image data is read from the sensor according to the preset image acquisition parameters (such as resolution, frame rate, exposure time, etc.), stored in the system memory, and subjected to necessary basic format conversion to form the first frame image information for subsequent steps. In continuous detection applications, this step will be repeated continuously to obtain continuous image frames, but here it refers specifically to the initial data input in a single or first detection cycle.

[0032] In the above deep learning-based mobile detection method, the step S3, the first frame image information is input into the mobile detection algorithm model to obtain a detection result. Specifically, considering that deep learning models have significant advantages in image feature extraction and semantic segmentation. Therefore, in order to efficiently identify potential moving targets from the first frame image information, the present application constructs a mobile detection algorithm model based on deep learning algorithm, converts the first frame image information into standardized data recognizable by the model through scaling, normalization and format conversion, and then uses the convolutional neural network structure of the model to extract multi-scale image features to generate potential target information.

[0033] Figure 3 The flowchart of step S3 in the deep learning-based mobile detection method according to the embodiment of the present application is shown in FIG. 3. As shown in FIG. 3, the step S3 includes steps S31-S33. Figure 3As shown, the step S3 comprises: S31, scaling the first frame image information to obtain scaled first frame image information; S32, normalizing the scaled first frame image information to obtain normalized first frame image information; S33, format converting the normalized first frame image information to obtain standardized first frame image information; and S34, inputting the standardized first frame image information into the mobile detection algorithm model to obtain the detection result, the detection result being a list of potential target moving target information, the potential target moving target information including boundary box coordinates, target type and confidence score of the potential target moving target.

[0034] Specifically, the step S31 comprises scaling the first frame image information to obtain scaled first frame image information. It should be understood that a deep learning model usually requires a fixed size of input image, while the resolution of an actually collected image may vary due to different devices or scenes. Therefore, in order to adapt to the input requirements of the model and improve the calculation efficiency, the present application scales the first frame image information by a bilinear interpolation or a region interpolation algorithm based on the principle of image resampling and feature reservation to realize size standardization. Specifically, according to the preset input size of the model (such as 640x640 pixels), the ratio of the original image width and height to the target size is calculated, and the image is scaled to the target resolution according to the ratio, and the pixel transition is smoothed by the interpolation algorithm to avoid loss of details. For example, if the original image is 1920x1080 pixels, it is scaled to 640x360 pixels, and then expanded to 640x640 pixels by edge padding (such as zero padding or mirror padding). In this way, the input image is compatible with the model structure, the redundant calculation caused by too high resolution is reduced, and the geometric features of the key target area are reserved, laying a foundation for subsequent feature extraction.

[0035] Specifically, the step S32, the scaled first frame image information is normalized to obtain normalized first frame image information. Here, considering that factors such as light intensity and sensor noise in the image acquisition process will cause differences in pixel value distribution, which directly affects the stability of model inference. Therefore, in order to eliminate data distribution deviation and accelerate model convergence, the present application is based on the principle of data standardization, and the pixel values in the scaled first frame image information are mapped to a unified interval through linear transformation to realize numerical normalization. For example, the scaled image pixel value (usually an integer of 0-255) is converted to a floating point number between 0-1, and according to the standardization parameters (such as mean [0.485, 0.456, 0.406] and standard deviation [0.229, 0.224, 0.225]) preset in the model training stage, the RGB three channels are respectively executed. The operation of subtracting the mean and then dividing by the standard deviation. In this way, the distribution of the normalized image data is closer to the data distribution during model training, which helps to improve the consistency of feature extraction, while suppressing the interference introduced by light changes or sensor noise, and enhancing the adaptability of the model to complex environments.

[0036] Specifically, the step S33, the normalized first frame image information is format converted to obtain standardized first frame image information. It should be understood that the deep learning framework has specific requirements for the format of input data (such as channel order, data type), and the normalized image data is still stored in the traditional RGB format or array form in CPU memory, which cannot be directly processed by the model calculation unit efficiently. Therefore, in order to make the image data conform to the format specification of the model input and ensure that the computing device (such as GPU) can quickly read and operate, the present application is based on the principle of data format adaptation, and the normalized first frame image information is further converted to the format required by the model. Specifically, first, the RGB channel order of the normalized first frame image information is adjusted to the BGR order required by the model (such as some pre-trained models of the framework are trained based on the BGR format), then the image data is transferred from the CPU memory to the GPU video memory, and converted to a floating point tensor (Tensor) format (such as FloatTensor of PyTorch), while adding a batch dimension (Batch Dimension) to adapt to the batch inference interface of the model. In this way, the image data can be directly received by the computational graph (Computational Graph) of the model, avoiding inference failure caused by format incompatibility, while utilizing the parallel computing capability of GPU to accelerate the model forward propagation process and improve the detection real-time performance.

[0037] Specifically, the step S34 inputs the standardized first frame image information into the moving detection algorithm model to obtain the detection result, and the detection result is a list of potential target moving target information, and the potential target moving target information includes a bounding box coordinate of a potential target moving target, a target type, and a confidence score. Specifically, in order to accurately identify the potential moving target from the standardized first frame image information, the present application constructs a moving detection algorithm model based on the cooperative work of a convolutional neural network (CNN) and a region proposal network (RPN), and generates a detection result through forward propagation and post-processing fusion. Specifically, first, the standardized image tensor (i.e., the standardized first frame image information) is input into a pre-trained moving detection model (such as Faster R-CNN), multi-scale feature maps are extracted through multi-scale convolution processing of a backbone network, and then different levels of semantic and positional information are fused through a feature pyramid network (FPN). Subsequently, the region proposal network (RPN) generates candidate target regions based on an anchor box mechanism, and predicts the target class confidence score and the boundary box coordinate offset of each candidate target region through a classification branch and a regression branch, respectively. Finally, a structured list of detection results containing the bounding box coordinate of the potential target (such as [x_min, y_min, x_max, y_max]), the target type (such as pedestrian, vehicle, animal, etc.), and the confidence score (indicating the possibility of the potential target belonging to the corresponding target type, between 0 and 1) is output. For example, a certain detection result may contain the entry: {"bbox": [120, 80, 300, 400], "class": "car", "score": 0.92}. In this way, the powerful feature extraction and pattern recognition capabilities of the deep learning model can be used to separate the potential moving target from the complex background and provide accurate positioning information and category information, thereby providing accurate data support for subsequent effective target judgment based on physical constraints.

[0038] In the above deep learning-based moving object detection method, in step S4, in response to the detection result being a potential moving object, the bounding box coordinates of the potential moving object are extracted, and the area of the potential moving object is calculated based on the bounding box coordinates of the potential moving object. It should be understood that the detection result output by the moving object detection algorithm model only represents the position and type information of the potential object, and lacks quantitative description of the physical size of the object. In actual scenes, small detection regions (such as sensor noise, extremely small regions misjudged by the model) usually do not have actual detection significance and need to be further screened through physical measurement. Therefore, in order to screen out candidate objects with actual significance from potential objects, the present application determines the spatial range of the object in the image by extracting the bounding box coordinates of the potential moving object, and calculates the pixel area of the potential moving object using the bounding box coordinates as a preliminary basis for subsequent screening of effective objects. In the specific implementation process, when there is a potential moving object in the detection result, the bounding box coordinates [x_min, y_min, x_max, y_max] of each target are extracted one by one from the result list, and then the pixel area of the target in the image plane is calculated according to the positions of the four vertices of the bounding box coordinates, that is, (x_max - x_min) x (y_max - y_min). This area value reflects the occupied range of the potential moving object in the image, and the larger the area, the larger the object size in the actual scene, while the smaller the area, the smaller the object size in the actual scene. In this way, a preliminary screening condition based on the size of the object is provided for the subsequent determination of effective objects, which helps to avoid misjudgment caused by relying solely on model confidence.

[0039] In the above deep learning-based moving object detection method, in step S5, if the area of the potential moving object exceeds the designated threshold, the potential moving object is marked as an effective moving object; and if the area of the potential moving object does not exceed the designated threshold, the potential moving object is marked as invalid information. That is, the pixel area of the potential moving object is compared with the preset area threshold to further screen out objects with actual detection value. When the pixel area of the potential moving object is greater than or equal to the preset area threshold, it is considered that the occupied range of the object in the image is large enough to correspond to an object with a certain physical size in the actual scene, and therefore the object is marked as an effective moving object for subsequent processing or response. On the contrary, when the pixel area of the potential moving object is less than the preset area threshold, it is considered that the object may be a small region caused by environmental interference, sensor error or model misjudgment, and therefore the object is marked as invalid information and ignored or excluded. In this way, by combining the detection result of the deep learning model with the screening condition of the physical size, the accuracy and practicality of the moving object detection can be effectively improved, and the false positives and false negatives can be reduced.

[0040] In particular, considering that the setting of the calibration threshold needs to be based on the requirements of specific application scenarios and the characteristics of background noise, and the detection requirements in different scenarios (such as the attention of security monitoring to large targets and the identification of traffic detection to small and medium-sized vehicles) are different, it is difficult for a fixed calibration threshold to adapt to diversified application scenarios. Therefore, in order to improve the adaptability and flexibility of the calibration threshold, the present application further proposes a method of dynamically calibrating the threshold, which learns the noise level and target distribution characteristics in the standardized first frame image information by analyzing the target category distribution and confidence score of all potential moving targets in the standardized first frame image information, so as to automatically adjust the calibration threshold.

[0041] Figure 4 The flowchart of step S5 in the deep learning-based moving detection method according to the embodiment of the present application is shown in FIG. 5. As shown in FIG. 5, the step S5 includes: S51, extracting an initial calibration threshold; S52, extracting the target category and the confidence score from the list of potential target moving target information to obtain the global distribution of the target category and the set of confidence scores; S53, calculating a threshold fine-tuning coefficient based on the global distribution of the target category and the set of confidence scores; S54, fine-tuning the initial calibration threshold based on the threshold fine-tuning coefficient to obtain the calibration threshold. Figure 4

[0042] Specifically, the step S51 extracts an initial calibration threshold. In the technical solution of the present application, in order to establish a reference for judgment and realize flexible adjustment, based on historical data statistics, the dynamic fine-tuning process is started by loading a preset initial threshold. Specifically, the system reads the initial calibration threshold (such as the target area distribution statistical value or empirical value based on the training data set) from the configuration file. This threshold serves as the baseline for subsequent dynamic adjustment, ensuring that the basic filtering capability can be maintained when no significant category distribution deviation is detected, thereby retaining the scene universality and providing a starting point for subsequent threshold optimization based on real-time detection data, avoiding initial judgment deviation caused by complete reliance on dynamic calculation.

[0043] ​Specifically, the step S52 extracts the target categories and the confidence scores from the list of potential target moving target information to obtain a global distribution of target categories and a set of confidence scores. It should be understood that the distribution of different target categories and the confidence scores in the detection result implies the scene semantic information (for example, if there are a large number of pedestrians and high confidence, it may mean that the monitoring area is a public place; if the confidence score of the vehicle is concentrated, it may imply a traffic monitoring scene). Therefore, in order to quantify the scene characteristics and guide the threshold adjustment, based on the category distribution analysis and the confidence credibility principle, the present application captures the scene characteristics by statistically analyzing the global distribution of target categories and the set of confidence scores in the image, preliminarily evaluates the noise level and target distribution characteristics in the image, and mines the semantic tendency of the current scene to provide a basis for dynamically adjusting the calibrated threshold.

[0044] Specifically, the step S53 calculates a threshold fine-tuning coefficient based on the global distribution of target categories and the set of confidence scores. Wherein, Figure 5 The flow chart of step S53 in the deep learning-based moving detection method according to the embodiment of the present application is shown in FIG. 5. As shown in FIG. 5, the step S53 includes: S531, performing semantic embedding coding on each target category in the global distribution of target categories to obtain a global distribution of target category semantic embedding coding vectors; S532, performing parameterized feature fusion on the global distribution of target category semantic embedding coding vectors based on the set of confidence scores to obtain a global distribution of target category semantic embedding modulation coding vectors; S533, performing global collaborative coding fusion on the global distribution of target category semantic embedding modulation coding vectors to obtain a global target category semantic aggregation coding vector; S534, performing feature decoding on the global target category semantic aggregation coding vector to obtain the threshold fine-tuning coefficient. Figure 5

[0045] ​More specifically, the step S531, each target category in the global distribution of the target category is semantically embedded to obtain the global distribution of the target category semantic embedding encoding vector. It should be understood that, since the target category label (such as "pedestrian" "vehicle") is discrete symbol itself, it cannot directly represent the current scene implicit characteristics based on the semantic association between categories. Therefore, in order to map the category information to a high-dimensional semantic space to capture the implicit association, the present application is based on the word embedding technology, and the semantic embedding model (such as Word2Vec) is pre-trained to encode the semantics of each target category in the global distribution of the target category, to obtain the global distribution of the target category semantic embedding encoding vector. Among them, the classes with similar semantics (such as "sedan" and "SUV") are close in semantic embedding space, while the classes with large semantic difference (such as "pedestrian" and "bird") are far away. Based on this semantic similarity distribution characteristics, the present application can more accurately understand the target distribution under the current scene, further guide the dynamic adjustment of the calibration threshold.

[0046] More specifically, the step S532, based on the set of confidence scores, the global distribution of the target category semantic embedding encoding vector is parameterized feature fusion to obtain the global distribution of the target category semantic embedding modulation encoding vector. It should be understood that, since the high confidence target is more likely to be a real moving object, and the low confidence target may be noise or false detection. Therefore, in order to strengthen the influence weight of high confidence category on threshold adjustment, the present application is based on the attention mechanism and the confidence reliability principle, and the class confidence score of each potential moving target is used to parameterize the feature fusion of the corresponding target category semantic embedding encoding vector, thereby forming the global distribution of the target category semantic embedding modulation encoding vector. In this way, the class features of high confidence are enhanced, and the noise features of low confidence are suppressed, so that the threshold adjustment is more dependent on reliable detection results, avoiding the interference of threshold calculation caused by accidental misjudgment of model.

[0047] More specifically, the step S533, the global distribution of the target category semantic embedding modulation encoding vector is globally collaborative coding fusion to obtain the global target category semantic aggregation encoding vector. It should be understood that the semantic features of a single category only reflect local information, and further global aggregation of all category features is needed to capture the comprehensive influence of the overall semantic distribution of all potential moving targets in the scene on the area threshold (such as when there are vehicles and pedestrians, the size requirements of the two types of targets need to be balanced). Therefore, the present application further generates a global target category semantic aggregation encoding vector representing the overall semantic characteristics of the scene by performing feature aggregation processing on the global distribution of the target category semantic embedding modulation encoding vector, thereby realizing comprehensive modeling of the global semantics of the scene, and making the area threshold setting more in line with the detection requirements in the multi-category coexistence scene.

[0048] Figure 6 This is a flowchart of step S533 in the deep learning-based motion detection method according to an embodiment of this application. Figure 6 As shown, step S533 includes: S5331, performing basic representation fusion on the global distribution of the target category semantic embedding modulation coding vector to obtain a global target category basic representation fusion coding vector; S5332, extracting the semantic compensation stimulated features of each target category semantic embedding modulation coding vector relative to the global target category basic representation fusion coding vector in the global distribution of the target category semantic embedding modulation coding vector to obtain a global distribution of the global target category semantic compensation stimulated feature coding vector; S5333, based on the global distribution of the global target category semantic compensation stimulated feature coding vector, performing adaptive feature enhancement on the global target category basic representation fusion coding vector to obtain the global target category semantic aggregation coding vector.

[0049] More specifically, step S5331 includes: first, inputting the global distribution of the target category semantic embedding modulation coding vector into a K-Means clustering network to obtain K global target category semantic clustering center coding vectors, expressed by the formula:

[0050]

[0051]

[0052] in, This represents the global distribution of the semantic embedding modulation and coding vector of the target category. , , and These respectively represent the 1st, 2nd, and 3rd elements in the global distribution of the target category semantic embedding modulation and coding vector. The and the first Each target category semantic embedding modulation-coded vector, The number of vectors in the global distribution of the semantic embedding modulation-coding vector of the target category. This represents a K-Means clustering network. The set of encoding vectors representing the semantic clustering centers of the target category across the entire domain. , and These represent the 1st, 2nd, and 3rd nodes in the set of semantic clustering center encoding vectors for the global target category, respectively. Each target category semantic clustering center encoding vector.

[0053] It should be appreciated that considering that there can be multiple instances of the same category in the global distribution of the target category semantic embedding modulation coding vector. Therefore, in order to structure the inductive data and extract representative primitives, the present application is based on unsupervised clustering technology, and the high-dimensional vector space is divided into K dense regions by K-Means algorithm to realize preliminary clustering of different instances of the same target category. Specifically, by presetting the number of clusters K (such as K = 5), initializing the cluster center and then iteratively optimizing the sum of squared errors (SSE) to minimize the sum of squared errors (SSE), until the centroid is stable, to obtain K global target category semantic clustering center encoding vectors, each cluster global target category semantic clustering center encoding vector represents the common semantic primitive of the cluster. In this way, the original high-dimensional distribution is compressed into K cluster center encoding vectors, which not only preserves the key patterns of the global distribution, but also provides low-noise input for subsequent ground state aggregation, significantly reducing the computational complexity.

[0054] Next, a self-attention driven global representation fusion mechanism is performed on the K global target category semantic clustering center encoding vectors to obtain the global target category base representation fusion encoding vector, which is expressed by the formula as follows:

[0055]

[0056]

[0057]

[0058]

[0059]

[0060] wherein, , , respectively represent the query matrix, the key matrix and the value matrix, , and respectively represent the query embedding matrix, the key embedding matrix and the value embedding matrix, represents the attention interaction layer, represents the feature scale value of the key matrix, represents the normalized exponential function, represents the layer normalization operation, represents the global target category base representation fusion encoding vector.

[0061] It can be understood that, since the cluster centers generated by K-Means only reflect local density peaks, there is a lack of modeling of global semantic association. Therefore, in order to capture the context dependence across clusters and high-order semantic relationship, the present application is based on the self-attention mechanism and the principle of global context modeling, and the K global target category semantic cluster center encoding vectors are input into the self-attention layer as query (Q), key (K), and value (V), the attention weight matrix is calculated, and if a cluster center and another cluster center have complementarity in the semantic space (such as "pedestrian" and "bicycle" often co-occur in street scenes), the attention weight is higher, thereby explicitly mining the potential semantic relationship between different cluster centers, forming a deep understanding of the global target category association distribution. Finally, the global aggregation is performed to obtain the global target category basic representation fusion encoding vector. The global target category basic representation fusion encoding vector not only covers the local characteristics of each cluster, but also encodes the global association across clusters, forming a compact representation of the common semantic of the scene.

[0062] More specifically, the step S5332 can be represented by the formula:

[0063]

[0064]

[0065] wherein, represents feature concatenation, represents a Sigmoid activation function, and respectively represent a semantic compensation weight parameter matrix and a semantic compensation bias, represents the global distribution of the global target category semantic compensation excited feature encoding vector, , , and respectively represent , , and corresponding global target category semantic compensation excited feature encoding vectors.

[0066] Here, the global target class category-based representation fusion encoding vector only represents the common semantic features of the current scene global dominant target class category, but ignores the individualized semantic deviation of each target class category relative to the global base state (such as rare target class categories or abnormal features). Therefore, in order to explicitly separate individual-specific information, based on the residual learning principle, the present application separates the global common features from the individual-specific information by mining the individualized semantic compensation of each target class category semantic embedding modulation encoding vector relative to the global target class category-based representation fusion encoding vector, and obtains the global distribution of the global target class category semantic compensation excited feature encoding vector, so as to explicitly retain the individual difference information in the original data, and provide more fine semantic guidance for subsequent threshold adjustment and anomaly detection.

[0067] More specifically, the step S5333 comprises: first, calculating the excitation response weight coefficient of each global target class category semantic compensation excited feature encoding vector in the global distribution of the global target class category semantic compensation excited feature encoding vector to obtain the global distribution of the global target class category semantic compensation significance modulation weight factor.

[0068] In particular, considering that in the calculation of the semantic compensation features of each target class category semantic embedding modulation encoding vector relative to the global target class category-based representation fusion encoding vector, by defining the customized perturbation parameter of each target class category semantic embedding modulation encoding vector relative to the global target class category-based representation fusion encoding vector, ignoring the geometric constraint from high-dimensional space to low-dimensional base space, it is possible to cause the semantic compensation excited feature representation to be excessively inflated or the boundary to be blurred. Based on this, in one preferred example of the present application, calculating the excitation response weight coefficient of each global target class category semantic compensation excited feature encoding vector in the global distribution of the global target class category semantic compensation excited feature encoding vector to obtain the global distribution of the global target class category semantic compensation significance modulation weight factor comprises: first, based on the semantic difference boundary limitation between each target class category semantic embedding modulation encoding vector and the global target class category-based representation fusion encoding vector in the global distribution of the target class category semantic embedding modulation encoding vector, the feature modulation is performed on the each global target class category semantic compensation excited feature encoding vector to obtain the global distribution of the optimized global target class category semantic compensation excited feature encoding vector.

[0069] Specifically, first, the target class category semantic embedding modulation encoding vector relative to the global target class category-based representation fusion encoding vector Manhattan norm semantic compensation excitation coefficient and Euclidean norm semantic compensation excitation coefficient wherein, and respectively represent the Manhattan norm and the Euclidean norm of the computed vector; then, the sparse geodesic projection operator vector and the energy distribution geodesic projection operator vector , so as to hierarchically reduce the dimension mapping in the tangent space form, so that it is projected onto the orthogonal normal basis of the intrinsic space.

[0070] Then, the sparse geodesic projection operator vector and the energy distribution geodesic projection operator vector are taken as the hierarchical geometric measure eigenvalue to perform the boundary constrained space deformation as:

[0071]

[0072] wherein, indicates the corresponding reconstructed target category semantic embedding modulation coding vector, is a deformation gain parameter, used to compensate the influence of the hyperbolic space geometric coefficient .

[0073] That is, the measurement complexity of the global target category semantic compensation for the excited feature is converted from being determined by the high-dimensional set of target category semantic embedding modulation coding vectors to being determined by the geometric characteristics of the tangent space measure boundary.

[0074] In this way, the modified is , so that in the case of guiding the measurement framework from the original target category semantic embedding representation to the intrinsic space, the stable node of the tangent measure flow can be calibrated by making the measurement of the global target category semantic compensation for the excited feature dependent on the boundary guidance rather than the dimension transformation, to realize the effective quantization of the boundary and avoid the over-representation of each global target category semantic compensation for the excited feature coding vector.

[0075] Then, each optimized global target category semantic compensation for the excited feature coding vector in the global distribution of the optimized global target category semantic compensation for the excited feature coding vector is input into the feature refining compressor based on the ReLU function and then globally normalized to obtain the global distribution of the global target category semantic compensation significance modulation weight factor, which is expressed by the formula as:

[0076]

[0077]

[0078] wherein, indicates the attention score conversion vector, indicates a learnable weight parameter matrix, indicates a ReLU activation function, denotes an exponential function with base e, denotes a global distribution of global target category semantic compensation saliency modulation weight factors, , , and denote a global target category semantic compensation saliency modulation weight factor, respectively. , , and corresponding global target category semantic compensation saliency modulation weight factors.

[0079] It can be understood that, since the global distribution of the global target category semantic compensation excited feature encoding vector contains not only meaningful individual deviations (such as important abnormal targets), but also random noise or irrelevant details. Therefore, in order to distinguish between key features and interference signals, the present application performs semantic importance evaluation on each global target category semantic compensation excited feature encoding vector through a neural network based on feature importance evaluation and sparse coding principles to obtain an importance score, and normalizes it through Softmax to obtain a saliency weight factor, so as to quantify the importance of each global target category semantic compensation excited feature encoding vector to the current scene moving detection task, and provide accurate semantic weight guidance for subsequent personalized semantic compensation gain aggregation.

[0080] Then, based on the global distribution of the global target category semantic compensation saliency modulation weight factor, the global distribution of the global target category semantic compensation excited feature encoding vector is parameterized feature fusion to obtain the global distribution of the global target category semantic compensation excited feature saliency encoding vector, which can be expressed as:

[0081]

[0082]

[0083] wherein, denotes a global distribution of global target category semantic compensation excited feature saliency encoding vector, , , and denote a global target category semantic compensation excited feature saliency encoding vector, respectively. , , and corresponding global target category semantic compensation excited feature saliency encoding vector.

[0084] That is, in order to strengthen the key features and suppress noise, the present application is based on the attention-driven feature selection principle, and the corresponding global target class semantic compensation excited feature encoding vector is scaled element by element by the global target class semantic compensation saliency modulation weight factor, so that the semantic compensation features of high weight targets (such as “reverse vehicle”) are amplified, and the semantic compensation features of low weight targets (such as “flying bird”) are weakened, to focus on individual bias with high semantic value, filter redundant information, and improve the relevance and information density of subsequent aggregation.

[0085] Then, the global distribution of the global target class semantic compensation excited feature saliency encoding vector and the global target class basic representation fusion encoding vector are input into a deep feature fusion network based on a multilayer perceptron to obtain the global target class semantic aggregation encoding vector, which can be represented by a formula as follows:

[0086]

[0087] wherein, represents a multilayer perceptron, represents a global target class semantic aggregation encoding vector.

[0088] Here, in order to balance the global dominant target class semantic information and the individualized semantic bias of important individuals in the current scene, the present application performs position mean fusion on the global distribution of the global target class semantic compensation excited feature saliency encoding vector, and then performs deep fusion and gain superposition of feature layers based on a multilayer perceptron (MLP) model on the global target class semantic compensation excited feature saliency encoding vector and the global target class basic representation fusion encoding vector, thereby retaining the global semantic structure of the scene while accurately capturing the specificity information of key individuals, and providing a more rich and accurate semantic representation for the subsequent threshold calibration adaptive regulation task.

[0089] More specifically, the step S534, the global target class semantic aggregation encoding vector is decoded to obtain the threshold fine-tuning coefficient. Specifically, in order to convert the scene semantic information contained in the global target class semantic aggregation encoding vector into a specific threshold fine-tuning coefficient, the present application designs a feature decoder based on a deep neural network architecture to decode the global target class semantic aggregation encoding vector, so as to extract its deep semantic features and map them to the threshold adjustment space. Specifically, the feature decoder is based on a multi-layer perceptron (MLP) structure, which gradually analyzes the complex semantic correlation in the global target class semantic aggregation encoding vector through multiple layers of nonlinear transformation, understands the semantic tendency of the current scene, and restricts the result to the interval [-1, 1] through the Sigmoid function of the output layer, and outputs the threshold fine-tuning coefficient matched with the scene characteristics. The size and direction of the threshold fine-tuning coefficient determine the adjustment amplitude and direction of the initial calibration threshold. If the target class distribution in the scene is mainly large vehicles and has high confidence, the threshold fine-tuning coefficient is positive, prompting to increase the threshold to filter smaller interference; if there are more small and medium-sized targets, the threshold fine-tuning coefficient may be negative, allowing smaller targets to pass detection.

[0090] Specifically, the step S54, the initial calibration threshold is fine-tuned based on the threshold fine-tuning coefficient to obtain the calibration threshold. In the technical solution of the present application, in order to finally realize the accurate matching of the calibration threshold and the scene characteristics, the initial calibration threshold is dynamically scaled based on the threshold fine-tuning coefficient to realize threshold updating. Specifically, the initial calibration threshold Tinit and the threshold fine-tuning coefficient Δ are calculated according to the formula Tinit x (1+Δ) to obtain the updated calibration threshold Tadj. In this way, the calibration threshold can be adapted to the semantic characteristics of different scenes, effectively coping with complex and variable detection environments.

[0091] In summary, the deep learning-based moving detection method according to the embodiments of the present application is illustrated, which extracts the Y component data of the adjacent first frame image and the second frame image respectively to avoid color channel interference, and maps the high-dimensional image matrix to a low-dimensional matrix through data dimension reduction processing to reduce the calculation complexity. Subsequently, the frame difference information matrix is established by calculating the frame difference information of the Y component data of the adjacent two frames after dimension reduction, and the enable region in the frame difference information matrix is demarcated based on the enable mask, and the region frame difference value is compared with the dynamically adjusted preset threshold to determine whether there is a moving target. This method uses the sensitivity of the Y component to brightness changes to effectively improve the detection accuracy, while reducing the calculation complexity and hardware cost through data dimension reduction processing, which is suitable for monitoring scenes that are cost-sensitive and require accurate moving detection, and can optimize the control of device cost while ensuring detection performance.

[0092] Furthermore, this application also provides a deep learning-based motion detection system.

[0093] Figure 7 This is a block diagram of a deep learning-based motion detection system according to an embodiment of this application. Figure 7 As shown, the deep learning-based motion detection system 100 according to an embodiment of this application includes: an initialization module 110 for device startup and initialization; a first frame image acquisition module 120 for acquiring first frame image information; a motion detection module 130 for inputting the first frame image information into a motion detection algorithm model to obtain a detection result; an area calculation module 140 for extracting the bounding box coordinates of the potential moving target in response to the detection result indicating the existence of a potential moving target, and calculating the area of ​​the potential moving target based on the bounding box coordinates of the potential moving target; and a moving target validity determination module 150 for marking the potential moving target as a valid moving target if the area of ​​the potential moving target exceeds a calibration threshold; and marking the potential moving target as invalid information if the area of ​​the potential moving target does not exceed the calibration threshold.

[0094] Here, those skilled in the art will understand that the specific operations of each module in the aforementioned deep learning-based motion detection system have been described above. Figures 1 to 6 The description of deep learning-based motion detection methods is detailed here, and therefore, its repeated description will be omitted.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A motion detection method based on deep learning, characterized in that, include: S1: Device startup and initialization; S2: Obtain the image information of the first frame; S3: Input the image information of the first frame into the motion detection algorithm model to obtain the detection result; S4: In response to the detection result that a potential moving target exists, extract the bounding box coordinates of the potential moving target, and calculate the area of ​​the potential moving target based on the bounding box coordinates of the potential moving target; S5: If the area of ​​a potential moving target exceeds the calibration threshold, the potential moving target is marked as a valid moving target; and if the area of ​​a potential moving target does not exceed the calibration threshold, the potential moving target is marked as invalid information.

2. The deep learning-based motion detection method and system according to claim 1, characterized in that, Step S1 includes: Turn on the device; Initialize parameters; Configure deep learning hardware computing power; Configure the software parameters of the motion detection algorithm model.

3. The deep learning-based motion detection method and system according to claim 1, characterized in that, Step S3 includes: The first frame image information is scaled to obtain scaled first frame image information; The scaled first frame image information is normalized to obtain normalized first frame image information; The normalized first frame image information is converted to a standardized first frame image information; The standardized first frame image information is input into the motion detection algorithm model to obtain the detection result, which is a list of potential moving target information, including the bounding box coordinates, target type, and confidence score of the potential moving target.

4. The deep learning-based motion detection method and system according to claim 3, characterized in that, Step S5 includes: Extract the initial calibration threshold; Extract target categories and confidence scores from the list of potential target movement target information to obtain a global distribution of target categories and a set of confidence scores; Based on the global distribution of the target category and the set of confidence scores, the threshold fine-tuning coefficient is calculated; The initial calibration threshold is fine-tuned based on the threshold fine-tuning coefficient to obtain the calibration threshold.

5. The deep learning-based motion detection method according to claim 4, characterized in that, Based on the global distribution of the target category and the set of confidence scores, the threshold fine-tuning coefficient is calculated, including: Semantic embedding encoding is performed on each target category in the global distribution of the target category to obtain the global distribution of the target category semantic embedding encoding vector; Based on the set of confidence scores, parameterized feature fusion is performed on the global distribution of the target category semantic embedding coding vector to obtain the global distribution of the target category semantic embedding modulation coding vector; The global distribution of the target category semantic embedding modulation coding vector is subjected to global collaborative coding fusion to obtain a global target category semantic aggregate coding vector; The threshold fine-tuning coefficients are obtained by performing feature decoding on the semantic aggregation encoding vector of the global target category.

6. The deep learning-based motion detection method according to claim 5, characterized in that, The global distribution of the target category semantic embedding modulation coding vector is subjected to global collaborative coding fusion to obtain a global target category semantic aggregate coding vector, including: The global distribution of the target category semantic embedding modulation coding vector is fused with basic representations to obtain a global target category basic representation fusion coding vector; Extract the semantic compensation stimulated features of each target category semantic embedding modulation and coding vector relative to the global target category basic representation fusion coding vector in the global distribution of the target category semantic embedding modulation and coding vector to obtain the global distribution of the global target category semantic compensation stimulated feature coding vector; Based on the global distribution of the semantic compensation stimulated feature encoding vector of the global target category, adaptive feature enhancement is performed on the basic representation fusion encoding vector of the global target category to obtain the semantic aggregation encoding vector of the global target category.

7. The deep learning-based motion detection method according to claim 6, characterized in that, The global distribution of the target category semantic embedding modulation coding vector is fused with basic representations to obtain a global target category basic representation fusion coding vector, including: The global distribution of the target category semantic embedding modulation coding vector is input into the K-Means clustering network to obtain K global target category semantic clustering center coding vectors; A self-attention-driven global representation fusion mechanism is applied to the K global target category semantic clustering center encoding vectors to obtain the global target category basic representation fusion encoding vector.

8. The deep learning-based motion detection method according to claim 7, characterized in that, Based on the global distribution of the global target category semantic compensation stimulated feature encoding vector, adaptive feature enhancement is performed on the global target category basic representation fusion encoding vector to obtain the global target category semantic aggregate encoding vector, including: The activation response weight coefficients of each global target category semantic compensation activated feature encoding vector in the global distribution of the global target category semantic compensation activated feature encoding vector are calculated to obtain the global distribution of the global target category semantic compensation saliency modulation weight factor; Based on the global distribution of the semantic compensation saliency modulation weight factor of the whole domain target category, parameterized feature fusion is performed on the global distribution of the semantic compensation stimulated feature encoding vector of the whole domain target category to obtain the global distribution of the semantic compensation stimulated feature saliency encoding vector of the whole domain target category. The global distribution of the semantic compensation salient encoding vector of the global target category and the fusion encoding vector of the basic representation of the global target category are input into a deep feature fusion network based on a multilayer perceptron to obtain the semantic aggregate encoding vector of the global target category.

9. The deep learning-based motion detection method according to claim 8, characterized in that, Calculate the activation response weight coefficients of each global target category semantic compensation activated feature encoding vector in the global distribution of the global target category semantic compensation activated feature encoding vector to obtain the global distribution of the global target category semantic compensation saliency modulation weight factor, including: Based on the semantic difference boundary between each target category semantic embedding modulation coding vector and the global target category basic representation fusion coding vector in the global distribution of the target category semantic embedding modulation coding vector, feature modulation is performed on each global target category semantic compensation stimulated feature coding vector to obtain an optimized global distribution of the global target category semantic compensation stimulated feature coding vector; Each optimized global target category semantic compensation stimulated feature encoding vector in the global distribution of the optimized global target category semantic compensation stimulated feature encoding vector is input into a feature refining compressor based on the ReLU function and then subjected to global normalization to obtain the global distribution of the global target category semantic compensation saliency modulation weight factor.

10. A motion detection system based on deep learning, characterized in that, include: The initialization module is used for device startup and initialization. The first frame image acquisition module is used to acquire the first frame image information; The motion detection module is used to input the image information of the first frame into the motion detection algorithm model to obtain the detection result; An area calculation module is used to extract the bounding box coordinates of the potential moving target in response to the detection result that a potential moving target exists, and to calculate the area of ​​the potential moving target based on the bounding box coordinates of the potential moving target; The moving target validity determination module is used to mark a potential moving target as a valid moving target if the area of ​​the potential moving target exceeds a calibration threshold; and to mark the potential moving target as invalid information if the area of ​​the potential moving target does not exceed the calibration threshold.