Moving target detection method and system based on deep convolutional network and deep learning network, and storage medium

Through Yolov7 and FlowNet2.0 networks, detection and background boxes are generated, and optimization indicators are calculated to determine whether the target is a moving target, solving the problem of degradation of detection accuracy in complex scenarios, and achieving high-precision moving target recognition.

CN120495625APending Publication Date: 2025-08-15SHANGHAI CHENGTOU WATER (GRP) CO LTD WATER PROD BRANCH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510544232.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing motion object detection methods are affected by factors such as lighting, occlusion, and motion blur in complex scenarios, and the detection accuracy is prone to decrease.

Method used

The Yolov7 deep convolutional neural network is used to box the target in the image, generate position information of the detection box, and calculate the optical flow vector through the FlowNet2.0 deep learning network, expand the detection box to generate the background box, calculate optimization indicators and determine whether the target is a moving target.

Benefits of technology

Through local background analysis, eliminate uneven optical flow field errors, prevent false detection of edge targets, accurately identify small static targets, improve the detection accuracy of motion targets, and avoid false detection of multiple targets when occluded or crowded.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495625A_ABST
    Figure CN120495625A_ABST
Patent Text Reader

Abstract

The invention provides a moving target detection method based on a deep convolutional network and a deep learning network, and the method comprises the steps: S1, obtaining a first image, inputting the first image into a Yolov7 network, and generating the position information of a plurality of detection frames; s2, acquiring a second image of a frame adjacent to the first image, inputting the first image and the second image into a FlowNet2.0 network, and calculating an optical flow vector between the first image and the second image; s3, calculating a mean value of all pixel displacement of the target in each detection frame and a mean value of all pixel displacement angles of the target in each detection frame; s4, expanding each detection frame to generate a background frame; s5, calculating a mean value of all pixel displacements in each background frame and a mean value of all pixel displacement angles in each background frame; s6, calculating a first optimization index and a second optimization index; and S7, calculating a motion change value of the local background, setting a judgment function, and judging whether the target in the detection frame is a moving target or not. The precision of moving target detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of motion target detection technology, and in particular to a motion target detection method, system and storage medium based on a deep neural convolutional network and a deep learning network. Background Art

[0002] Motion target detection is an important research area in target detection. It can identify moving areas as foregrounds in a sequence of images or videos, separate them from the background, and annotate them, so as to extract regions of interest and perform subsequent processing. In visual scenes, moving targets often provide more information. By analyzing them, we can obtain the target's current position information or behavior status, and even predict upcoming events. With the development of hardware in recent years, motion target detection can be seen everywhere in life, from aerospace and military fields to production systems, smart transportation, and video surveillance systems in daily life.

[0003] Common traditional moving target detection methods include frame-difference-based methods, background modeling-based methods, and optical flow-based methods. Frame-difference-based methods are susceptible to interference from background changes and are prone to incomplete target detection and internal holes. Background modeling-based methods are highly sensitive to lighting and shadows, resulting in complex background modeling and limitations in practical applications. Optical flow-based methods perform calculations for each pixel, resulting in a large computational load and are sensitive to lighting changes, which can affect detection performance when lighting conditions are unstable.

[0004] Existing techniques for detecting moving targets based on optical flow and convolutional neural networks also exist. These methods first extract features from video frames, then use optical flow algorithms to estimate motion information between frames. This motion information is then used to guide feature aggregation, aggregating features from different frames to improve detection accuracy. These methods can accurately calculate motion information at the pixel level and capture minute details in images, but they require significant computing resources and data for training and inference. Overly complex models can reduce the algorithm's computational efficiency.

[0005] In dynamic scenes, the background and target move together, and the motion information is coupled, interfering with the accurate detection of moving targets. This is especially true in complex backgrounds with a large number of targets, which can easily lead to occlusion or crowding. Existing moving target detection methods can achieve good results in specific scenarios, but in complex scenes, detection accuracy is easily reduced due to factors such as lighting, occlusion, and motion blur. Summary of the Invention

[0006] The present invention proposes a motion target detection method, system and storage medium based on deep neural convolutional networks and deep learning networks to solve the technical problem that the existing motion target detection methods are easily affected by factors such as lighting, occlusion, motion blur, etc. in complex scenes, and the detection accuracy is easily reduced.

[0007] One aspect of the present invention is to provide a moving target detection method based on a deep neural convolutional network and a deep learning network, the moving target detection method comprising the following steps:

[0008] S1. Acquire a first image and input the acquired first image into a Yolov7 deep convolutional neural network;

[0009] The Yolov7 deep convolutional neural network selects each target in the first image through a detection frame and generates position information of multiple detection frames;

[0010] S2. Obtain a second image of a frame adjacent to the first image, and input the first image and the second image into a FlowNet2.0 deep learning network;

[0011] The optical flow vector between the first image and the second image is calculated using the FlowNet2.0 deep learning network to generate the horizontal velocity component u(x, y) and the vertical velocity component v(x, y) of the optical flow vector.

[0012] The horizontal velocity component u(x, y) of the optical flow vector contains the motion amplitude of each pixel (x, y) in the first image moving along the x-axis to the second image;

[0013] The vertical velocity component v(x, y) of the optical flow vector contains the motion amplitude of each pixel (x, y) in the first image moving along the y-axis to the second image;

[0014] S3, using the position information of each detection frame obtained in step S1 and the horizontal velocity component and the vertical velocity component of the optical flow vector obtained in step S2, calculate the mean Dm of all pixel displacements of the target in each detection frame and the mean Am of all pixel displacement angles of the target in each detection frame;

[0015] S4, expanding each detection frame to generate multiple background frames corresponding to each detection frame, and setting all pixels of the target in each detection frame to zero;

[0016] S5. Calculate the mean value De of the displacement of all pixels in each background frame and the mean value Ae of the displacement angle of all pixels in each background frame using the horizontal velocity component and the vertical velocity component of the optical flow vector obtained in step S2;

[0017] S6. Calculate a first optimization index and a second optimization index using the mean Dm of all pixel displacements of the target in each detection frame and the mean Am of all pixel displacement angles of the target in each detection frame obtained in step S3, and the mean De of all pixel displacements of each background frame and the mean Ae of all pixel displacement angles of each background frame obtained in step S5;

[0018] S7. Calculate the motion change value of the local background between each detection frame and each background frame using the first optimization index and the second optimization index, and set the judgment function:

[0019]

[0020] Among them, ID represents the motion change value of the local background between a detection frame and a background frame, IDa represents the set of motion change values of all pixels between a detection frame and a background frame, and IDm represents the mean of the set of motion change values of all pixels between a detection frame and a background frame;

[0021] When the value of the judgment function f(ID) is 1, the target in the detection frame corresponding to the background frame is a moving target;

[0022] When the value of the judgment function f(ID) is 0, the target in the detection frame corresponding to the background frame is a stationary target.

[0023] In a preferred embodiment, in step S1, the Yolov7 deep convolutional neural network selects each target in the first image through a detection frame, and the generated position information of each detection frame includes: the upper, lower, left and right boundary position information of the detection frame.

[0024] In a preferred embodiment, in step S3, the mean value Dm of all pixel displacements of the target within each detection frame and the mean value Am of all pixel displacement angles of the target within each detection frame are calculated by the following method:

[0025]

[0026] Where um represents the mean value of the motion amplitude of each pixel (x, y) of the target in the detection box along the x-axis, and u(x, y) represents the horizontal velocity component of the optical flow vector, which includes the motion amplitude of each pixel (x, y) in the first image moving along the x-axis to the second image.

[0027] vm represents the mean value vm of the motion amplitude of each pixel (x, y) of the target in the detection box along the y-axis direction, v(x, y) represents the vertical velocity component of the optical flow vector, which includes the motion amplitude of each pixel (x, y) in the first image moving along the y-axis direction to the second image;

[0028] t, b, l, and r represent the coordinate values of the upper, lower, left, and right boundaries of the detection box, respectively. Dm represents the mean of all pixel displacements of the target in each detection box, and Am represents the mean of all pixel displacement angles of the target in each detection box.

[0029] In a preferred embodiment, in step S4, each detection frame is enlarged by the following method:

[0030] With the center of the detection frame as the center, the width and height of the detection frame are expanded to 1.5 times the original size to generate a background frame corresponding to the detection frame.

[0031] In a preferred embodiment, in step S5, the mean value De of all pixel displacements in each background frame and the mean value Ae of all pixel displacement angles in each background frame are calculated by the following method:

[0032]

[0033] Wherein, ue represents the mean value of the motion amplitude of each pixel (x, y) in the background frame along the x-axis, and u(x, y) represents the horizontal velocity component of the optical flow vector, which includes the motion amplitude of each pixel (x, y) in the first image moving along the x-axis to the second image;

[0034] vm represents the mean value vm of the motion amplitude of each pixel (x, y) in the background frame along the y-axis direction, v(x, y) represents the vertical velocity component of the optical flow vector, which includes the motion amplitude of each pixel (x, y) in the first image moving along the y-axis direction to the second image;

[0035] w and h represent the width and height of the background frame respectively, De represents the mean value of the displacement of all pixels in each background frame, and Am represents the mean value of the displacement angle of all pixels in each background frame.

[0036] In a preferred embodiment, in step S6, the first optimization index and the second optimization index are calculated by the following method:

[0037] IDA=|Am-Ae|;

[0038] IDD=|Dm-De|;

[0039] Among them, IDA represents the first optimization index, IDD represents the second optimization index, Dm represents the mean of all pixel displacements of the target in each detection frame, Am represents the mean of all pixel displacement angles of the target in each detection frame, De represents the mean of all pixel displacements in each background frame, and Am represents the mean of all pixel displacement angles in each background frame.

[0040] In a preferred embodiment, in step S7, the motion change value of the local background between each detection frame and each background frame is calculated by the following method:

[0041] ID = IDA × IDD;

[0042] Among them, ID represents the motion change value of the local background between a detection frame and a background frame, IDA represents the first optimization index, and IDD represents the second optimization index.

[0043] Another aspect of the present invention is to provide a motion target detection system based on a deep neural convolutional network and a deep learning network, wherein the motion target detection system is used to execute a motion target detection method based on a deep neural convolutional network and a deep learning network provided by the present invention.

[0044] Another aspect of the present invention is to provide a computer storage medium, which is used to store computer execution instructions, and the computer execution instructions are used to execute a motion target detection method based on deep neural convolutional network and deep learning network provided by the present invention.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] The present invention proposes a motion target detection method, system and storage medium based on a deep neural network and a deep learning network. The method comprises the following steps: selecting each target in an image through a Yolov7 deep convolutional neural network to generate the position information of a detection frame; calculating the optical flow vector between two adjacent frames of images through a FlowNet2.0 deep learning network to generate the horizontal velocity component and the vertical velocity component of the optical flow vector; performing motion analysis on the target in the detection frame using the position information of the detection frame, as well as the horizontal velocity component and the vertical velocity component of the optical flow vector, and expanding the detection frame to generate a background frame corresponding to the detection frame; performing motion analysis on the background frame to calculate a first optimization index and a second optimization index, calculating the motion change of the local background between the detection frame and the background frame using the first optimization index and the second optimization index, and setting a judgment function to determine whether the target in the detection frame is a moving target.

[0047] The present invention proposes a motion target detection method, system and storage medium based on a deep neural convolutional network and a deep learning network. The method expands the detection frame to generate a background frame corresponding to the detection frame; performs motion analysis on the background frame, calculates a first optimization index and a second optimization index, uses the first optimization index and the second optimization index to calculate the motion change of the local background between the detection frame and the background frame, and sets a judgment function to judge whether the target in the detection frame is a moving target, thereby realizing motion analysis using the background around the detection frame (the local background between the detection frame and the background frame), eliminating the error caused by the uneven optical flow field of each part of the image background, and the influence on the subsequent judgment of whether the target is moving.

[0048] The present invention proposes a motion target detection method, system and storage medium based on a deep neural convolutional network and a deep learning network. The method uses the background around the detection frame (the local background between the detection frame and the background frame) for motion analysis, which can prevent the problem of false detection of edge targets caused by using the global background. In addition, when filtering the detection frame of a stationary target, the detection frame of the stationary target can be significantly removed, thereby retaining the detection frame of the moving target and realizing accurate detection of the moving target.

[0049] The present invention proposes a motion target detection method, system and storage medium based on a deep neural convolutional network and a deep learning network. The method uses the background around the detection frame (the local background between the detection frame and the background frame) for motion analysis, which can accurately identify small stationary targets and correctly eliminate them, effectively solving the problem of false detection of small targets caused by using the global background.

[0050] The present invention proposes a motion target detection method, system and storage medium based on a deep neural convolutional network and a deep learning network. When there is target occlusion or congestion due to a large number of targets in an image in a specific scenario, the background around the detection frame (the local background between the detection frame and the background frame) is used for motion analysis, which can effectively avoid the problem of detecting multiple crowded targets as the same target, thereby improving the accuracy of motion target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.

[0052] Figure 1 This is a flow chart of a motion target detection method based on deep neural convolutional network and deep learning network of the present invention.

[0053] Figure 2 This is a schematic diagram of selecting a target in a first image with a detection frame and expanding the detection frame to generate a background frame in an embodiment of the present invention.

[0054] Figure 3 FIG. 1 is a schematic diagram of setting all pixels of an object within a detection frame to zero in one embodiment of the present invention. DETAILED DESCRIPTION

[0055] In order to make the above and other features and advantages of the present invention more clear, the present invention is further described below with reference to the accompanying drawings. It should be understood that the specific embodiments given herein are for the purpose of explaining to those skilled in the art and are only exemplary and not restrictive.

[0056] Combine Figures 1 to 3 According to an embodiment of the present invention, a method for detecting moving objects based on a deep neural convolutional network and a deep learning network is provided, which includes the following steps:

[0057] Step S1: Acquire a first image and input the acquired first image into the Yolov7 deep convolutional neural network.

[0058] The Yolov7 deep convolutional neural network is a deep convolutional neural network consisting of three parts: a backbone network, a feature pyramid network, and an output network. The Yolov7 deep convolutional neural network can quickly detect the target in the first image through the detection box.

[0059] The Yolov7 deep convolutional neural network of the present invention selects each target in the first image through a detection frame and generates position information of multiple detection frames.

[0060] Furthermore, the Yolov7 deep convolutional neural network of the present invention selects each target in the first image through a detection frame, and generates position information of each detection frame, including: top, bottom, left, and right boundary position information (Top, Bottom, Left, Right) of the detection frame.

[0061] like Figure 2 As shown, in this embodiment, there are three targets in the first image, namely, the first target 100, the second target 200, and the third target 300. The Yolov7 deep convolutional neural network selects the three targets (the first target 100, the second target 200, and the third target 300) in the first image through detection boxes and generates position information of the three detection boxes.

[0062] Taking the first target 100 as an example, the Yolov7 deep convolutional neural network selects the first target 100 in the first image through the detection frame and generates the position information of the detection frame 101 of the first target 100, such as Figure 2 As shown, the green box is the detection box.

[0063] Similarly, the Yolov7 deep convolutional neural network selects the second target 200 and the third target 300 in the first image through the detection frame, and generates the position information of the detection frame of the second target 200 and the third target 300, such as Figure 2 As shown, the green box is the detection box.

[0064] Step S2: Obtain a second image of a frame adjacent to the first image, and input the first image and the second image into a FlowNet2.0 deep learning network.

[0065] The FlowNet2.0 deep learning network is a deep learning-based method for optical flow estimation. Optical flow is the surface motion of objects between consecutive frames of a video. The FlowNet2.0 deep learning network utilizes end-to-end training of convolutional neural networks to optimize the performance of optical flow estimation. It can handle input images of varying resolutions and sizes, while providing more accurate optical flow estimates and more robust performance.

[0066] According to an embodiment of the present invention, the optical flow vector between the first image and the second image (the image of the adjacent frame to the first image) is calculated through the FlowNet2.0 deep learning network, and the horizontal velocity component u(x, y) of the optical flow vector and the vertical velocity component v(x, y) of the optical flow vector are generated.

[0067] The output of the FlowNet2.0 deep learning network is a two-dimensional array, including the horizontal velocity component u(x, y) of the optical flow vector and the vertical velocity component v(x, y) of the optical flow vector.

[0068] The horizontal velocity component u(x, y) of the optical flow vector contains the motion amplitude of each pixel (x, y) in the first image moving along the x-axis to the second image;

[0069] The vertical velocity component v(x, y) of the optical flow vector contains the motion amplitude of each pixel (x, y) in the first image moving along the y-axis to the second image;

[0070] Where (x, y) represents the pixel coordinates of each pixel.

[0071] Step S3: Calculate the mean Dm of all pixel displacements of the target in each detection frame and the mean Am of all pixel displacement angles of the target in each detection frame using the position information of each detection frame obtained in step S1 and the horizontal velocity component and vertical velocity component of the optical flow vector obtained in step S2.

[0072] According to an embodiment of the present invention, the mean value Dm of all pixel displacements of the target within each detection frame and the mean value Am of all pixel displacement angles of the target within each detection frame are calculated by the following method:

[0073]

[0074] Where um represents the mean value of the motion amplitude of each pixel (x, y) of the target in the detection box along the x-axis, and u(x, y) represents the horizontal velocity component of the optical flow vector, which includes the motion amplitude of each pixel (x, y) in the first image moving along the x-axis to the second image.

[0075] vm represents the mean value vm of the motion amplitude of each pixel (x, y) of the target in the detection box along the y-axis direction, v(x, y) represents the vertical velocity component of the optical flow vector, which includes the motion amplitude of each pixel (x, y) in the first image moving along the y-axis direction to the second image;

[0076] t, b, l, and r represent the coordinate values of the upper, lower, left, and right boundaries of the detection box, respectively. Dm represents the mean of all pixel displacements of the target in each detection box, and Am represents the mean of all pixel displacement angles of the target in each detection box.

[0077] like Figure 2 As shown, taking the first target 100 as an example, the position information of the detection frame 101 of the first target 100 obtained in step S1 and the horizontal velocity component u(x, y) of the optical flow vector and the vertical velocity component v(x, y) of the optical flow vector obtained in step S2 are used to calculate the mean Dm of all pixel displacements of the first target 100 in the detection frame 101 of the first target 100 and the mean Am of all pixel displacement angles of the first target 100 in the detection frame 101.

[0078] In step S3 of the present invention, the position information of the detection frame, as well as the horizontal velocity component and the vertical velocity component of the optical flow vector are used to perform motion analysis on the target within the detection frame, thereby calculating the mean Dm of all pixel displacements of the target within each detection frame and the mean Am of all pixel displacement angles of the target within each detection frame.

[0079] Step S4: Enlarge each detection frame to generate multiple background frames corresponding to each detection frame, and set all pixels of the target in each detection frame to zero.

[0080] like Figure 2 As shown, the detection frame 101 of the first target 100 ( Figure 2 As an example, the detection frame 101 is enlarged to generate the background frame 102 ( Figure 2 ), which is the background frame 102 corresponding to the detection frame 101.

[0081] According to an embodiment of the present invention, each detection frame is enlarged by the following method: with the center of the detection frame as the center, the width and height of the detection frame are enlarged to 1.5 times of the original, and a background frame corresponding to the detection frame is generated.

[0082] It should be understood that the detection frame of each target in the first image (each detection frame) is enlarged according to the above method to generate multiple background frames corresponding to each detection frame (one detection frame corresponds to one background frame).

[0083] According to an embodiment of the present invention, all pixels of the target within each detection frame are set to zero. Figure 3 As shown, the detection frame 101 of the first target 100 ( Figure 2 As an example, all pixels of the first target within the detection frame 101 of the first target 100 are set to zero ( Figure 3 middle black area).

[0084] Step S5: Calculate the mean value De of all pixel displacements and the mean value Ae of all pixel displacement angles in each background frame using the horizontal velocity component and the vertical velocity component of the optical flow vector obtained in step S2.

[0085] According to an embodiment of the present invention, the mean value De of all pixel displacements in each background frame and the mean value Ae of all pixel displacement angles in each background frame are calculated by the following method:

[0086]

[0087] Wherein, ue represents the mean value of the motion amplitude of each pixel (x, y) in the background frame along the x-axis, and u(x, y) represents the horizontal velocity component of the optical flow vector, which includes the motion amplitude of each pixel (x, y) in the first image moving along the x-axis to the second image;

[0088] vm represents the mean value vm of the motion amplitude of each pixel (x, y) in the background frame along the y-axis direction, v(x, y) represents the vertical velocity component of the optical flow vector, which includes the motion amplitude of each pixel (x, y) in the first image moving along the y-axis direction to the second image;

[0089] w and h represent the width and height of the background frame respectively, De represents the mean value of the displacement of all pixels in each background frame, and Am represents the mean value of the displacement angle of all pixels in each background frame.

[0090] like Figure 3 As shown, taking the first target 100 as an example, the horizontal velocity component u(x, y) of the optical flow vector and the vertical velocity component v(x, y) of the optical flow vector obtained in step S2 are used to calculate the mean value De of the displacement of all pixels in the background frame 102 corresponding to the detection frame 101 of the first target 100 and the mean value Ae of the displacement angle of all pixels in the background frame 102 corresponding to the detection frame 101 of the first target 100.

[0091] In step S5 of the present invention, motion analysis is performed on the background frame using the background frame, the horizontal velocity component of the optical flow vector, and the vertical velocity component of the optical flow vector, thereby calculating the mean value De of the displacement of all pixels in each background frame and the mean value Ae of the displacement angle of all pixels in each background frame.

[0092] Step S6: Calculate the first optimization index and the second optimization index using the mean Dm of all pixel displacements of the target in each detection frame, the mean Am of all pixel displacement angles of the target in each detection frame, and the mean De of all pixel displacements of the target in each background frame, and the mean Ae of all pixel displacement angles of the target in each background frame obtained in step S5.

[0093] Specifically, for each detection frame, the first optimization index and the second optimization index are calculated by the following method:

[0094] IDA=|Am-Ae|;

[0095] IDD=|Dm-De|;

[0096] Among them, IDA represents the first optimization index, IDD represents the second optimization index, Dm represents the mean of all pixel displacements of the target in each detection frame, Am represents the mean of all pixel displacement angles of the target in each detection frame, De represents the mean of all pixel displacements in each background frame, and Am represents the mean of all pixel displacement angles in each background frame.

[0097] Step S7: Calculate the motion change value of the local background between each detection frame and each background frame using the first optimization index and the second optimization index, and set a determination function; use the determination function to determine whether the target in the detection frame is a moving target.

[0098] According to an embodiment of the present invention, the motion change value of the local background between each detection frame and each background frame is calculated by the following method:

[0099] ID = IDA × IDD;

[0100] Among them, ID represents the motion change value of the local background between a detection frame and a background frame, IDA represents the first optimization index, and IDD represents the second optimization index.

[0101] According to an embodiment of the present invention, the judgment function is set:

[0102]

[0103] Among them, ID represents the motion change value of the local background between a detection frame and a background frame, IDa represents the set of motion change values of all pixels between a detection frame and a background frame, and IDm represents the mean of the set of motion change values of all pixels between a detection frame and a background frame.

[0104] When the value of the judgment function f(ID) is 1, the target in the detection frame corresponding to the background frame is a moving target;

[0105] When the value of the judgment function f(ID) is 0, the target in the detection frame corresponding to the background frame is a stationary target.

[0106] Combine Figure 2 and Figure 3 Taking the first target 100 as an example, according to the method of step S6 and step S7, the motion change value of the local background between the detection frame 101 of the first target 100 and the background frame 102 is calculated, and the determination function f(ID) is set.

[0107] When the value of the determination function f(ID) is 1, the first object 100 in the detection frame 101 corresponding to the background frame 102 is a moving object.

[0108] When the value of the determination function f(ID) is 0, the first object 100 in the detection frame 101 corresponding to the background frame 102 is a stationary object.

[0109] If the target is a moving target, the detection frame of the target is retained; if the target is a stationary target, the detection frame of the target is removed and the target is also removed.

[0110] The present invention calculates the first optimization index and the second optimization index through steps S6 and S7, uses the first optimization index and the second optimization index to calculate the motion change of the local background between the detection frame and the background frame, and sets a judgment function to determine whether the target in the detection frame is a moving target, thereby realizing motion analysis using the background around the detection frame (the local background between the detection frame and the background frame), accurately identifying small stationary targets, correctly eliminating small stationary targets, effectively solving the problem of false detection of small targets, and improving the accuracy of moving target detection.

[0111] According to an embodiment of the present invention, a motion target detection system based on a deep neural convolutional network and a deep learning network is provided, which is used to execute a motion target detection method based on a deep neural convolutional network and a deep learning network of the present invention.

[0112] According to an embodiment of the present invention, a computer storage medium is provided for storing computer-executable instructions for executing a moving object detection method based on a deep neural convolutional network and a deep learning network according to the present invention.

[0113] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A moving target detection method based on deep neural convolutional network and deep learning network, characterized in that: The moving target detection method comprises the following steps: S1. Acquire a first image and input the acquired first image into a Yolov7 deep convolutional neural network; The Yolov7 deep convolutional neural network selects each target in the first image through a detection frame to generate position information of multiple detection frames; S2. Obtain a second image of a frame adjacent to the first image, and input the first image and the second image into a FlowNet2.0 deep learning network; The optical flow vector between the first image and the second image is calculated using the FlowNet2.0 deep learning network to generate the horizontal velocity component u(x, y) and the vertical velocity component v(x, y) of the optical flow vector. The horizontal velocity component u(x, y) of the optical flow vector contains the motion amplitude of each pixel (x, y) in the first image moving along the x-axis to the second image; The vertical velocity component v(x, y) of the optical flow vector contains the motion amplitude of each pixel (x, y) in the first image moving along the y-axis to the second image; S3, using the position information of each detection frame obtained in step S1 and the horizontal velocity component and the vertical velocity component of the optical flow vector obtained in step S2, calculate the mean Dm of all pixel displacements of the target in each detection frame and the mean Am of all pixel displacement angles of the target in each detection frame; S4, expanding each detection frame to generate multiple background frames corresponding to each detection frame, and setting all pixels of the target in each detection frame to zero; S5. Calculate the mean value De of the displacement of all pixels in each background frame and the mean value Ae of the displacement angle of all pixels in each background frame using the horizontal velocity component and the vertical velocity component of the optical flow vector obtained in step S2; S6. Calculate a first optimization index and a second optimization index using the mean Dm of all pixel displacements of the target in each detection frame and the mean Am of all pixel displacement angles of the target in each detection frame obtained in step S3, and the mean De of all pixel displacements of each background frame and the mean Ae of all pixel displacement angles of each background frame obtained in step S5; S7. Calculate the motion change value of the local background between each detection frame and each background frame using the first optimization index and the second optimization index, and set the judgment function: Among them, ID represents the motion change value of the local background between a detection frame and a background frame, IDa represents the set of motion change values of all pixels between a detection frame and a background frame, and IDm represents the mean of the set of motion change values of all pixels between a detection frame and a background frame; When the value of the judgment function f(ID) is 1, the target in the detection frame corresponding to the background frame is a moving target; When the value of the judgment function f(ID) is 0, the target in the detection frame corresponding to the background frame is a stationary target.

2. The moving target detection method according to claim 1, wherein: In step S1, the Yolov7 deep convolutional neural network selects each target in the first image through a detection frame, and generates position information of each detection frame, including: upper, lower, left, and right boundary position information of the detection frame.

3. The moving target detection method according to claim 1, wherein: In step S3, the mean value Dm of all pixel displacements of the target in each detection frame and the mean value Am of all pixel displacement angles of the target in each detection frame are calculated by the following method: Where um represents the mean value of the motion amplitude of each pixel (x, y) of the target in the detection box along the x-axis, and u(x, y) represents the horizontal velocity component of the optical flow vector, which includes the motion amplitude of each pixel (x, y) in the first image moving along the x-axis to the second image. vm represents the mean value vm of the motion amplitude of each pixel (x, y) of the target in the detection box along the y-axis direction, v(x, y) represents the vertical velocity component of the optical flow vector, which includes the motion amplitude of each pixel (x, y) in the first image moving along the y-axis direction to the second image; t, b, l, and r represent the coordinate values of the upper, lower, left, and right boundaries of the detection box, respectively. Dm represents the mean of all pixel displacements of the target in each detection box, and Am represents the mean of all pixel displacement angles of the target in each detection box.

4. The moving target detection method according to claim 1, wherein: In step S4, each detection box is enlarged by the following method: With the center of the detection frame as the center, the width and height of the detection frame are expanded to 1.5 times the original size to generate a background frame corresponding to the detection frame.

5. The moving target detection method according to claim 1, wherein: In step S5, the mean value De of all pixel displacements in each background frame and the mean value Ae of all pixel displacement angles in each background frame are calculated by the following method: Wherein, ue represents the mean value of the motion amplitude of each pixel (x, y) in the background frame along the x-axis, and u(x, y) represents the horizontal velocity component of the optical flow vector, which includes the motion amplitude of each pixel (x, y) in the first image moving along the x-axis to the second image; vm represents the mean value vm of the motion amplitude of each pixel (x, y) in the background frame along the y-axis direction, v(x, y) represents the vertical velocity component of the optical flow vector, which includes the motion amplitude of each pixel (x, y) in the first image moving along the y-axis direction to the second image; w and h represent the width and height of the background frame respectively, De represents the mean value of the displacement of all pixels in each background frame, and Am represents the mean value of the displacement angle of all pixels in each background frame.

6. The moving target detection method according to claim 1, wherein: In step S6, the first optimization index and the second optimization index are calculated by the following method: IDA=|Am-Ae|; IDD=|Dm-De|; Among them, IDA represents the first optimization index, IDD represents the second optimization index, Dm represents the mean of all pixel displacements of the target in each detection frame, Am represents the mean of all pixel displacement angles of the target in each detection frame, De represents the mean of all pixel displacements in each background frame, and Am represents the mean of all pixel displacement angles in each background frame.

7. The moving target detection method according to claim 1, wherein: In step S7, the motion change value of the local background between each detection frame and each background frame is calculated by the following method: ID = IDA × IDD; Among them, ID represents the motion change value of the local background between a detection frame and a background frame, IDA represents the first optimization index, and IDD represents the second optimization index.

8. A moving target detection system based on deep neural convolutional network and deep learning network, characterized in that: The moving target detection system is used to execute the moving target detection method according to any one of claims 1 to 7.

9. A computer storage medium, characterized in that The computer storage medium is used to store computer-executable instructions, and the computer-executable instructions are used to execute the moving target detection method according to any one of claims 1 to 7.