A CUDA-based parallel optical flow method and system for moving object detection

By integrating the Canny edge detection algorithm into the traditional optical flow method and using CUDA for parallel calculation, the problem of traditional optical flow method taking time and being susceptible to noise interference is solved, and high-precision and real-time motion object detection is achieved.

CN114913194BActive Publication Date: 2025-07-01QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210616829.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-01
Publication Date
2025-07-01
Estimated Expiration
2042-06-01

AI Technical Summary

Technical Problem

The traditional optical flow method motion object detection is unable to meet the demand for real-time high-precision detection in the field of intelligent monitoring because it takes a long time to calculate and is susceptible to noise interference from light changes.

Method used

The motion object detection method based on CUDA is adopted, and the Canny edge detection algorithm is integrated into the HS optical flow method. The Canny operator is used to distinguish between strong and weak edges, enhance the contour information of the detection results, and suppress noise through data fusion.

Benefits of technology

It improves the accuracy and anti-interference ability of motion target detection, meets the real-time detection needs, and reduces the detection time overhead through parallel computing optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114913194B_ABST
    Figure CN114913194B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of digital image processing, and provides a method and system for detecting moving targets based on CUDA parallel optical flow method, including acquiring video images and performing preprocessing; creating a first CPU thread to extract the target contour image from the preprocessed video images by using the optical flow method; creating a second CPU thread to perform edge detection on the preprocessed video images by using an edge detection algorithm; transmitting the edge detection result to the first CPU thread, and in the first CPU thread, fusing the edge detection result with the target contour image to obtain a fused image; detecting moving targets based on the fused image; by integrating the Canny edge detection algorithm into the HS optical flow method, the present invention can utilize the characteristic of the Canny operator to effectively distinguish strong and weak edges, strengthen the contour information of the moving target detection result of the HS optical flow method, retain the same information of both, effectively suppress environmental noise, enhance the anti-interference ability of the algorithm, and improve the detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital image processing, and particularly relates to a method and system for detecting moving targets based on CUDA parallel optical flow method. Background Art

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] CUDA (Compute Unified Device Architecture) is a development environment launched by NVIDIA for NVIDIA's GPUs, which can use the characteristics of GPU multi-threading for parallel computing in the CUDA development environment.

[0004] Moving target detection is to calculate the correlation between adjacent frames in the video image sequence collected, determine the moving targets in the image, and extract the moving targets from the complex background image.

[0005] The optical flow method is one of the moving target detection algorithms. This algorithm assigns motion vectors to each pixel point of the image to establish an image motion field, and judges whether there are moving targets by analyzing the change of motion vectors over time.

[0006] Traditional optical flow method for moving target detection calculates and detects moving targets through a single CPU thread. When constructing the optical flow motion field, due to a large number of image matrix iterative calculations involved, it takes a very long time, and the change of illumination will bring a large amount of noise data to the detection results. Facing the rapid development of the intelligent monitoring field and the high-efficiency and high-quality requirements in the fields of pedestrian detection, vehicle detection, etc., the traditional optical flow method cannot meet the real-time and high-precision detection requirements. Summary of the Invention

[0007] In order to solve the above problems, the present invention proposes a method and system for detecting moving targets based on CUDA parallel optical flow method. By integrating the Canny edge detection algorithm into the HS optical flow method, the present invention can use the characteristic of the Canny operator to effectively distinguish strong and weak edges, strengthen the contour information of the moving target detection result of the HS optical flow method, retain the same information of the two, effectively suppress environmental noise, enhance the anti-interference ability of the algorithm, and improve the detection accuracy.

[0008] According to some embodiments, the first solution of the present invention provides a method for detecting moving targets based on CUDA parallel optical flow method, and adopts the following technical solution:

[0009] A method for detecting moving targets based on CUDA parallel optical flow method includes:

[0010] Obtain a video image and perform preprocessing;

[0011] Create a first CPU thread, and use the optical flow method to extract the target contour image from the preprocessed video image;

[0012] Create a second CPU thread, and use the edge detection algorithm to perform edge detection on the preprocessed video image;

[0013] Transmit the edge detection result to the first CPU thread. In the first CPU thread, fuse the edge detection result with the target contour image to obtain a fused image;

[0014] Detect moving targets based on the fused image.

[0015] Further, the obtaining a video image and performing preprocessing includes:

[0016] Obtain a video image and perform grayscale conversion;

[0017] Process the grayscale video image using the Gaussian filtering algorithm to obtain a filtered video image;

[0018] Perform image augmentation on the filtered video image to obtain a preprocessed video image.

[0019] Further, the creating a first CPU thread and using the optical flow method to extract the target contour image from the preprocessed video image includes:

[0020] Based on the preprocessed video image, determine the image gradients of the preprocessed video image in the x direction and y direction;

[0021] Establish an HS motion model and perform iterative solution to extract the target contour image from the preprocessed video image.

[0022] Further, the establishing an HS motion model and performing iterative solution to extract the target contour image from the preprocessed video image includes:

[0023] a. Establish an HS motion model according to the optical flow constraint equation and constraint conditions;

[0024] b. Solve the HS motion model to obtain the result of the first iterative calculation of the optical flow components;

[0025] c. Substitute the result of the first iterative calculation of the optical flow components into the HS motion model to determine the mean value of the calculation result matrix;

[0026] d. Repeat step c;

[0027] e. Compare the result of the current iterative calculation with the result of the previous iterative calculation. If it is less than the first time, the iteration continues; otherwise, go to step f;

[0028] f. End the iteration, and extract the target contour image from the processed video image based on the iteration result.

[0029] Further, the creation of the second CPU thread and the use of the edge detection algorithm to perform edge detection on the preprocessed video image include:

[0030] Based on the preprocessed video image, calculate the image gradient using the Sobel convolution kernel to obtain the initial edge;

[0031] If the gradient intensity of the current point is the largest when compared with the gradient intensities of other points in the same direction, keep its value, otherwise set it to 0;

[0032] Compare the retained image gradient with two set size thresholds to screen out the edge nodes and complete the edge detection.

[0033] Further, transmit the edge detection result to the first CPU thread. In the first CPU thread, perform an "AND" operation fusion on the edge detection result and the target contour image to obtain a fused image.

[0034] Further, extracting the detection target based on the fused image includes:

[0035] Perform dilation processing on the fused image to eliminate the breakpoints and discontinuities of the edges in the fused image;

[0036] Obtain the final detection target image.

[0037] According to some embodiments, the second solution of the present invention provides a CUDA-based parallel optical flow method moving target detection system, adopting the following technical solution:

[0038] A CUDA-based parallel optical flow method moving target detection system includes:

[0039] An image acquisition module, configured to acquire a video image and perform preprocessing;

[0040] A target contour extraction module, configured to create a first CPU thread and use the optical flow method to extract a target contour image from the preprocessed video image;

[0041] An edge detection module, configured to create a second CPU thread and use the edge detection algorithm to perform edge detection on the preprocessed video image;

[0042] An image fusion module, configured to transmit the edge detection result to the first CPU thread. In the first CPU thread, fuse the edge detection result and the target contour image to obtain a fused image;

[0043] The moving target detection module is configured to detect moving targets based on the fused image.

[0044] According to some embodiments, the third solution of the present invention provides a computer-readable storage medium.

[0045] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in a method for detecting moving targets based on CUDA parallel optical flow method as described in the first aspect above.

[0046] According to some embodiments, the fourth solution of the present invention provides a computer device.

[0047] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a method for detecting moving targets based on CUDA parallel optical flow method as described in the first aspect above.

[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0049] By integrating the Canny edge detection algorithm into the HS optical flow method, the present invention can utilize the characteristic of the Canny operator to effectively distinguish strong and weak edges, strengthen the contour information of the moving target detection result of the HS optical flow method, retain the same information of both, effectively suppress environmental noise, enhance the anti-interference ability of the algorithm, and improve the detection accuracy.

[0050] In the parallel optimization stage of the present invention, the present proposal innovatively proposes a parallel HS-Canny algorithm. This algorithm ensures the real-time requirement of the actual detection process, can quickly and accurately detect moving targets, and enter subsequent target tracking, behavior analysis and other links more quickly. At the same time, using 2 GPUs effectively reduces the additional time overhead brought by integrating the Canny operator, further improves the real-time performance of the algorithm, and also improves the efficiency of processing a large number of pictures. Description of the Drawings

[0051] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0052] Figure 1 is a flowchart of a method for detecting moving targets based on CUDA parallel optical flow method according to an embodiment of the present invention;

[0053] Figure 2 is a hardware flowchart of the parallel HS-Canny algorithm according to an embodiment of the present invention;

[0054] Figure 3 It is a schematic diagram of using CUDA threads corresponding to the pixel coordinates of the picture to process corresponding elements in the embodiments of the present invention;

[0055] Figure 4(a) is the original video image in the embodiments of the present invention;

[0056] Figure 4(b) is the detection result of the traditional HS optical flow method in the embodiments of the present invention;

[0057] Figure 5(a) is the Canny edge detection result in the embodiments of the present invention;

[0058] Figure 5(b) is the detection result of the improved HS optical flow method in the embodiments of the present invention;

[0059] Figure 5(c) is the data fusion result in the embodiments of the present invention;

[0060] Figure 5(d) is the HS-Canny edge detection result in the embodiments of the present invention. Specific Embodiments

[0061] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0062] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0063] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0064] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0065] Embodiment 1

[0066] As Figure 1As shown in the figure, this embodiment provides a method for detecting moving objects based on the parallel optical flow method of CUDA. This embodiment takes the application of this method to a server as an example. It can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is realized through the interaction between the terminal and the server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, web servers, cloud communications, middleware services, domain name services, security services CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this. In this embodiment, the method includes the following steps:

[0067] Obtain a video image and perform preprocessing;

[0068] Create a first CPU thread and use the optical flow method to extract the target contour image from the preprocessed video image;

[0069] Create a second CPU thread and use the edge detection algorithm to perform edge detection on the preprocessed video image;

[0070] Transmit the edge detection result to the first CPU thread. In the first CPU thread, fuse the edge detection result with the target contour image to obtain a fused image;

[0071] Detect moving objects based on the fused image.

[0072] Extract the detected target based on the fused image, including:

[0073] Perform dilation processing on the fused image to eliminate the breakpoints and discontinuities of the edges in the fused image;

[0074] Obtain the final detected target image.

[0075] Among them, the dilation processing is specifically:

[0076] The dilation of set A by set B is defined as follows:

[0077]

[0078] Among them, the set is called the structuring element, represents the new set obtained by reflecting and translating B by z.

[0079] Interpreted in the image: A is a binary image, and B is a matrix similar to the size of the convolutional kernel, such as Traverse the image A with B. As long as there is a pixel point in A that coincides with B, then the pixel points in this area are all 255 pixel points of A and B.

[0080] Specifically, as Figure 1 shown, the method described in this embodiment specifically includes:

[0081] S1: Image preprocessing

[0082] Image preprocessing is to pre-denoise the image. The first steps of the HS optical flow method and the Canny edge detection algorithm both require image preprocessing. The image required for image preprocessing is a grayscale image, and the method is to read a single-channel image using the imread function in OpenCV. The image preprocessing stage is completed by the Gaussian filtering algorithm, and the parallel Gaussian filtering algorithm is completed by CUDA two-dimensional convolution. The two-dimensional Gaussian distribution formula is:

[0083]

[0084] Considering the problem of high global memory access latency, shared memory is used to optimize the memory access time. To ensure that the filtered image has the same size as the original image, pixel points need to be supplemented on the outside of the image, and the size of the image amplification is affected by the Gaussian convolutional kernel. Let the side length of the Gaussian convolutional kernel be L (L = 3, 5, 7, 9,...), then the size of the amplified image is (W + L - 1) × (H + L - 1), where W and H are the width and height of the original image. The values of the pixel points in the amplified area are 0, that is, 0 is filled at the edge. To reduce the additional storage and time overhead brought by image amplification, when copying image pixel points in shared memory, the amplified area determination condition is added and the indexing method of the thread coordinates under the global thread is changed.

[0085] S2: Parallel HS-Canny algorithm

[0086] This step is the core part of this proposal and the key technology to achieve the technical effects of this proposal. To improve the integrity of the moving target contour extraction by the HS optical flow method, the Canny edge detection algorithm is added to the HS optical flow method in this case. To reduce the extra time overhead brought by adding the Canny algorithm, two GPUs are used in this paper. The HS optical flow method is run on GPU 0, and the Canny edge detection is run on GPU 1. Finally, the results of these two algorithms are fused through an "AND" operation. To cater to the SIMT (Single Instruction Multiple Thread) mode of the GPU, the pthread function library is used to create two CPU threads, which serve as the Instructions for one GPU respectively. In each thread function, the cudaSetDevice function is used to select the corresponding GPU. The hardware process of the parallel HS-Canny algorithm is as Figure 2 shown.

[0087] (1) GPU 0: Acceleration and optimization of the HS optical flow method

[0088] The HS optical flow method is established based on the optical flow constraint equation, and the optical flow constraint equation is as follows:

[0089] I x u + I y v + I t = 0 (2)

[0090] Among them, I x 、I y 、I t represent the first-order gradients in the horizontal direction, vertical direction, and time of the image respectively, and u and v represent the optical flow components. Since there are two unknowns u and v in the above equation, one equation cannot solve them. The HS optical flow method solves u and v by adding additional constraint conditions. The HS optical flow method is mainly divided into 4 steps: image smoothing filtering, calculating the image gradient, establishing the HS motion model, and iterative solution. Among them, the image smoothing filtering has been completed in the image preprocessing. Calculating the image gradients in the x direction and y direction is to add the pixel coordinates at the corresponding positions of the gradient values in the x direction and y direction of two pictures (adjacent two frames) respectively, and the gradient value of the time direction matrix is the difference between the two pictures.

[0091] Calculating the image gradient is completed through convolution, specifically as follows:

[0092] The convolution kernel in the X direction convolves the grayscale images of the previous and next frames respectively, and then adds them together to obtain I x

[0093] The convolution kernel in the y direction Convolve the grayscale images of the front and back frames respectively, and then add them together, which is \(I\) in the motion model formula. y

[0094] There are two convolution kernels in the \(t\) direction, which are Convolve the grayscale images of the front and back frames respectively, and then add the results of the two (because one convolution kernel is all positive and the other is all negative, and the difference is obtained in this way).

[0095] Next, this section will focus on introducing the parallelization method for establishing the HS motion model and iterative solution.

[0096] The HS motion model is established based on the optical flow constraint equation. Since there are two unknowns \(u\) and \(v\) in the optical flow constraint equation, a single optical flow constraint equation cannot solve for them. The HS optical flow method adds two constraint conditions to this equation: ① The grayscale of the moving object remains unchanged within a very short time interval; ② The change of the velocity vector field within a given neighborhood is slow.

[0097] The HS motion model solves for \(u\) and \(v\) as follows:

[0098]

[0099] are the means of the neighborhoods of the optical flow components \(u\) and \(v\) (the variables in the above formula are all matrices, are the neighborhood means of each element in the matrices \(u\) and \(v\)), and \(\lambda\) is used to smooth the weight relationship between the error data term and the smooth constraint term. This formula is an iterative formula, \(n\) represents the number of iterations, and the \(u\) and \(v\) of the \(n\)th step are used to calculate the \(u\) and \(v\) of the \(n + 1\)th step. \(I\) x 、\(I\) y 、\(I\) t represent the first-order gradients in the horizontal, vertical, and time directions of the image respectively.

[0100] The above formula (3) is the HS motion model, and the specific solution process is as follows:

[0101] Before iteratively solving for \(u\) and \(v\), \(u\) and \(v\) are matrices with initial values all being 0. The steps for parallel solution are as follows:

[0102] ① Calculate \(u_{temp}\), \(v_{temp}\)

[0103] \(u_{temp}\) and \(v_{temp}\) are the values of \(u\) and \(v\) after one iteration of calculation, and the formula is as follows:

[0104]

[0105] where is the 4-neighborhood mean of the optical flow components u and v. When calculating the 4-neighborhood mean on the GPU, different treatments need to be made according to the position of the point to be calculated. The parallel method uses the idea of two-dimensional parallel convolution, and the convolution kernel is as follows:

[0106]

[0107] Then, it is necessary to judge the number of pixels in the neighborhood of the point to be calculated, and make additional judgments on the position of the pixel points in the kernel to indicate whether the point to be calculated corresponds to the edge area, non-edge area or the four corners of the actual image.

[0108] Obtain After that, substitute into formula (4) to calculate utemp and vtemp. The calculations involved in this step are addition, subtraction, multiplication and division of matrices. The parallel method belongs to the calculation of simple tasks with multiple data. Use the CUDA threads corresponding to the pixel coordinates of the image to process the corresponding elements, as follows Figure 3 shown.

[0109] ② Judge whether the iteration converges

[0110] After obtaining utemp and vtemp, substitute these values into the optical flow constraint equation and calculate the mean of the result matrix. The formula is expressed as follows:

[0111] result = I x utemp + I y vtemp + I t (5)

[0112]

[0113] Among them, result represents the result after one iteration of the HS optical flow method, mean represents the average value of the result matrix, total represents the number of pixel points in the result matrix, Sum() represents the summation function. The method is to use CUDA for reduction summation. To solve the bank conflict problem, let each thread add the pixel point of the current thread coordinate to the pixel point at a distance of (blockDim / 2). After each addition, the distance is halved until the distance is less than 1. After the in-block addition is completed, then use the global memory to add the results of each block to obtain the final result. This method will not have the problem of thread divergence in the warp, thus improving the calculation efficiency.

[0114] After each iteration, record the mean value of this iteration, and assign the utemp and vtemp calculated in ① above to u and v. That is, in the GPU, the device pointers utemp and vtemp point to the device pointers u and v respectively, and then the next iteration is performed. After the second iteration, compare with the mean value obtained in the first iteration. If it is less than the mean value of the first iteration, the iteration continues; otherwise, the iteration continues. The result in the formula is the result of the parallel HS optical flow method.

[0115] The iterative solution will solve for u and v, and substituting them into formula (2) can obtain the image with noise that only contains moving objects, realizing the extraction of the target contour.

[0116] (2) GPU 1: Acceleration and optimization of the Canny edge detection algorithm

[0117] The Canny edge detection algorithm is divided into four steps: image denoising, calculating the image gradient, non-maximum suppression, double-threshold screening, and edge tracking.

[0118] ① Image denoising

[0119] The image denoising method uses the method of image preprocessing in S1, which will not be repeated here.

[0120] ② Calculating the image gradient

[0121] Use the Sobel convolution kernel to calculate the image gradient to obtain possible edges. The Sobel convolution kernels in the horizontal and vertical directions are as follows:

[0122]

[0123] When convolving the image with the Sobel operator, the method of processing the image edge cannot continue to use the method of padding 0 at the edge described in S1. If the pixel value at the image edge is 0, and the value of the pixel point connected to the edge inside the image is not 0, then the gradient value obtained at the image edge after convolution will be very large, and subsequent calculations may lead to misjudgment of the edge. The solution is to make the amplified image pixel points equal to the outermost pixel points of the original image. The formulas for calculating the gradient magnitude and direction are as follows:

[0124]

[0125]

[0126] Among them, G x and G y represent the gradient magnitudes in the horizontal and vertical directions respectively, and θ is the angle magnitude.

[0127] ③ Non-maximum suppression

[0128] Non-maximum suppression is an edge thinning method. If the gradient intensity of the current point is the maximum compared with the gradient intensities of other points in the same direction, its value is retained; otherwise, it is set to 0. To reduce the computational load of each thread, instead of comparing the gradient intensities in 8 directions, the gradient intensities in 4 directions are compared. This way reduces the number of pixel comparisons, and all pixels exist during the comparison process, which reduces the computational complexity brought by supplementing non-existent pixels through interpolation algorithms, thereby improving the execution efficiency.

[0129] ④ Double-threshold screening and edge tracking

[0130] According to two thresholds, large and small, those higher than the large threshold are strong edges, which means that the pixel points must be edge points; those lower than the small threshold are not edges; those between the two are weak edges. For weak edges, if they are connected to strong edges, they are determined to be edges; otherwise, they are not edges.

[0131] S3: Data fusion

[0132] During the hardware operation, first, the image data of the previous and next 2 frames need to be read and passed into the CPU main thread. To further improve the acceleration effect, the grayscale images of these two pictures are respectively passed into GPU 0 and GPU 1 for image preprocessing. Then, the HS optical flow method is performed in GPU 0, and the Canny edge detection is performed in GPU 1. When the Canny algorithm finishes execution, the HS optical flow method is still continuing. At this time, the result of GPU 1 is passed into GPU 0. The method is to create a 2D array in advance as a global array to store the result of the Canny algorithm. When the Canny algorithm finishes execution, the result of Canny is passed into the pre-created empty array through cudaMemcpyDeviceToHost, that is, passed back to the CPU main thread. After the HS algorithm finishes execution, cudaSetDevice in the CPU main thread is set to 0. At this time, GPU 0 works and GPU 1 is idle. Then, using cudaMemcpyHostToDevice, the result of the Canny algorithm can be passed into GPU 0. In this way, the data in GPU 1 is passed into GPU 0 for final data fusion. During the "AND" operation process, a new double-threshold strategy is used to screen the pixel values of the two results respectively, filtering out the pixel values lower than the threshold, effectively improving the accuracy of the final detection result.

[0133] The HS optical flow method has obtained the moving target image, but there is a lot of environmental noise in this detection result; the Canny edge detection algorithm can obtain the edge information of the image. The purpose of data fusion is to solve the problem of environmental noise in the HS optical flow method.

[0134] HS-Canny algorithm detection result

[0135] The detection results of the traditional HS optical flow method for moving target detection and the Canny edge measurement algorithm are analyzed and improved. Figure 4(a) is the 1830th frame in the pedestrian video. This frame is the moment with the most pedestrians in the video. Figure 4(b) is the result of the traditional HS optical flow method for moving target detection. There are certain environmental noise points in the detection result, and the movement of the human figure also has a certain interference on the detection result. On the other hand, affected by the action of the moving target, the detected moving target is "thinner" than the original image. In short, the detection result is not ideal.

[0136] In order to solve the problem of human shadow in Figure 4(b), a large number of video frames were selected for analysis. The results show that the pixel value of human shadow detected by traditional HS optical flow method is low. The threshold can effectively suppress the appearance of human shadow in the detection result, but some stubborn (high pixel value) environmental noise is not removed. If the threshold is increased, it may lead to the loss of moving targets. Therefore, it is considered to eliminate environmental noise by performing "AND" operation with Canny edge detection algorithm. Figure 5(a) is the result of Canny edge detection. The human body shape in the detection result is different from the original image. Figure 1 The moving targets detected by the traditional HS optical flow method are relatively thin. At this time, the "AND" operation is performed. Since the pixels of the moving targets in the two pictures do not overlap, the fusion of the two pictures will hardly obtain the information of the moving targets.

[0137] Therefore, in view of the problem that the moving target in Figure 4(b) is too thin, the "expansion" operation in mathematical morphology is studied and adopted, which can make the moving target reach its original size and solve the problem of internal fracture of some moving targets. Figure 5(b) is the detection result of the HS optical flow method after threshold screening and mathematical morphology processing. It can be seen from the figure that the human shadow problem is effectively solved and the figure is fuller. Although some stubborn noise is also amplified with the morphological processing, the "AND" operation with the Canny edge detection result can effectively eliminate the amplified environmental noise. The "AND" operation result is shown in Figure 5(c). Due to the edge refinement of Canny edge detection, although the edge noise is eliminated, the fused edge is not obvious, and even a small part of the moving target has breakpoints and discontinuities on the edge. At this time, it only needs to be "expanded" to make the final detection result more obvious and clearer. Figure 5(d) is the final detection result of the HS-Canny algorithm in this paper.

[0138] Embodiment 2

[0139] This embodiment provides a CUDA-based parallel optical flow moving target detection system, including:

[0140] An image acquisition module is configured to acquire video images and perform preprocessing;

[0141] A target contour extraction module, configured to create a first CPU thread and extract a target contour image from the preprocessed video image using an optical flow method;

[0142] An edge detection module, configured to create a second CPU thread and perform edge detection on the preprocessed video image using an edge detection algorithm;

[0143] An image fusion module, configured to transmit the edge detection result to the first CPU thread, and in the first CPU thread, fuse the edge detection result with the target contour image to obtain a fused image;

[0144] A moving target detection module, configured to detect a moving target based on the fused image.

[0145] The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the first embodiment above. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0146] In the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0147] The proposed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the above modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0148] Embodiment Three

[0149] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in a method for detecting a moving target based on a CUDA-based parallel optical flow method as described in the first embodiment above are implemented.

[0150] Embodiment Four

[0151] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in a method for detecting a moving target based on a CUDA-based parallel optical flow method as described in the first embodiment above are implemented.

[0152] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) that contain computer-usable program code.

[0153] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0154] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0156] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above-described embodiment methods can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0157] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A motion target detection method based on CUDA parallel optical flow method, characterized in that Including: Obtain a video image and perform preprocessing; Create a first CPU thread, and use the optical flow method to extract the target contour image from the preprocessed video image; The creating of the first CPU thread and using the optical flow method to extract the target contour image from the preprocessed video image includes: Based on the preprocessed video image, determine the preprocessed video image in x direction, y direction image gradient; Establish an HS motion model and perform iterative solution to extract the target contour image from the preprocessed video image; Establish an HS motion model on the optical flow constraint equation. The establishing of the HS motion model and performing iterative solution to extract the target contour image from the preprocessed video image includes: a. Establish an HS motion model according to the optical flow constraint equation and the constraint conditions; b. Solve the HS motion model to obtain the result of the first iterative calculation of the optical flow components; c. Substitute the result of the first iterative calculation of the optical flow components into the HS motion model to determine the mean value of the calculation result matrix. The formula is expressed as follows: (5) (6) Among them, represents the result after one iteration of the HS optical flow method. There are two unknowns in the optical flow constraint equation , , 、 is 、 the values obtained after one iteration of calculation, 、 、 respectively represent the first-order gradients in the horizontal, vertical, and time directions of the image, represents the average value of the matrix, represents the number of pixel points in the matrix, represents the summation function. The method is to use CUDA for reduction summation, so that each thread adds the pixel point at the current thread coordinate to the pixel point at a distance of apart. After each addition, the distance is halved until the distance is less than 1. After the addition within the block is completed, the results of each block are then added using global memory to obtain the final result; d. Repeat step c; e. Compare the result of the current iterative calculation with the result of the previous iterative calculation. If it is less than the first time, the iteration continues; otherwise, go to step f; f. End the iteration, and extract the target contour image from the processed video image based on the iteration result; Create a second CPU thread, and use an edge detection algorithm to perform edge detection on the preprocessed video image; Transmit the edge detection result to the first CPU thread. In the first CPU thread, fuse the edge detection result with the target contour image to obtain a fused image; Detect moving targets based on the fused image.

2. The motion target detection method based on CUDA parallel optical flow method according to claim 1, characterized in that The obtaining of the video image and performing preprocessing includes: Obtain the video image and perform grayscale conversion; Process the grayscale video image using a Gaussian filtering algorithm to obtain the filtered video image; Perform image augmentation on the filtered video image to obtain the preprocessed video image.

3. The method for detecting moving objects by a parallel optical flow method based on CUDA according to claim 1, characterized in that, The creating of the second CPU thread and using the edge detection algorithm to perform edge detection on the preprocessed video image includes: Based on the preprocessed video image, calculate the image gradient using a Sobel convolution kernel to obtain the initial edge; If the gradient intensity of the current point is the largest compared with the gradient intensities of other points in the same direction, retain its value; otherwise, it is 0; Compare the retained image gradient with two set size thresholds to screen out edge nodes and complete edge detection.

4. The motion target detection method based on CUDA parallel optical flow method according to claim 1, wherein, Transmit the edge detection result to the first CPU thread. In the first CPU thread, perform an "AND" operation fusion on the edge detection result and the target contour image to obtain a fused image.

5. The motion target detection method based on CUDA parallel optical flow method according to claim 1, characterized in that Extracting the detection target based on the fused image includes: Perform dilation processing on the fused image to eliminate the breakpoints and discontinuities of the edges in the fused image; Obtain the final detection target image.

6. A motion target detection system based on CUDA parallel optical flow method, characterized in that, Including: An image acquisition module configured to obtain a video image and perform preprocessing; A target contour extraction module configured to create a first CPU thread and use the optical flow method to extract the target contour image from the preprocessed video image; The creating of the first CPU thread and using the optical flow method to extract the target contour image from the preprocessed video image includes: Based on the preprocessed video image, determine the image gradient of the preprocessed video image in the x direction, y direction; Establish an HS motion model and perform iterative solution to extract the target contour image from the preprocessed video image; Establish an HS motion model based on the optical flow constraint equation. The establishment of the HS motion model and iterative solution are carried out. Extract the target contour image from the preprocessed video image, including: a. Establish an HS motion model according to the optical flow constraint equation and the constraint conditions; b. Solve the HS motion model to obtain the result of the first iteration calculation of the optical flow components; c. Substitute the result of the first iteration calculation of the optical flow components into the HS motion model to determine the mean value of the calculation result matrix. The formula is expressed as follows: (5) (6) Among them, represents the result after one iteration of the HS optical flow method. There are two unknowns in the optical flow constraint equation , , 、 is 、 the values obtained through one iteration of calculation, 、 、 respectively represent the first-order gradients in the horizontal, vertical, and temporal directions of the image, represents the average value of the matrix, represents the number of pixel points within the matrix, represents the summation function. The method is to use CUDA for reduction summation, allowing each thread to add the pixel point at the current thread coordinate to the pixel point at a distance of apart. After each addition, the distance is halved until the distance is less than 1. After the in-block addition is completed, the results of each block are then added using global memory to obtain the final result;​​ d. Repeat step c; e. Compare the result of the current iteration calculation with the result of the previous iteration calculation. If it is less than the first time, the iteration continues; otherwise, go to step f; f. End the iteration and extract the target contour image from the processed video image based on the iteration result; An edge detection module, configured to create a second CPU thread and perform edge detection on the preprocessed video image using an edge detection algorithm; An image fusion module, configured to transmit the edge detection result to the first CPU thread. In the first CPU thread, fuse the edge detection result with the target contour image to obtain a fused image; A moving target detection module, configured to detect a moving target based on the fused image.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in a CUDA-based parallel optical flow method for moving target detection method according to any one of claims 1-5.

8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a CUDA-based parallel optical flow method for moving target detection method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Moving target detecting method based on differential fusion and image edge information

    CN102184552A

  • Movable target tracking method and system based on optical flow approach

    CN105023278A