Linear motion target feature point detection method based on random neighborhood cyclic background modeling

Through random neighborhood cyclic background modeling and connected domain analysis, the feature points of linear motion targets in staring scenes are detected, which solves the problem of balancing low computing power requirements and target detection effects, and realizes real-time target detection with low computing power.

CN120747155APending Publication Date: 2025-10-03中国人民解放军95859部队
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510665296.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The existing technology lacks a moving target feature point detection algorithm in gaze scenarios that takes into account both low computing power requirements and good target detection effects.

Method used

The random neighborhood cyclic background modeling method is adopted to detect the feature points of linear moving targets by constructing an initial background sample set and connected domain analysis, combined with two consecutive frame mask matrices.

Benefits of technology

The computational complexity is reduced, meeting the real-time processing requirements of ordinary computers and achieving effective target detection under low computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747155A_ABST
    Figure CN120747155A_ABST
Patent Text Reader

Abstract

The invention provides a linear motion target feature point detection method based on random neighborhood cyclic background modeling, and the method is divided into two stages of detection: the first stage is detection of a motion target pixel region, and the random neighborhood cyclic background modeling method is adopted to replace a traditional complete random background modeling method in the stage; the calculation amount is greatly reduced on the premise of ensuring the background modeling effect; in the second stage, target feature point detection is conducted, target linear motion constraint is introduced, and head feature point detection is achieved through the thought that a front frame and a rear frame are associated. The detection method has the advantages of being low in computer hardware requirement and small in calculation amount, and has very high engineering application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital image processing, and in particular to a method for detecting feature points of a linear motion target based on random neighborhood cyclic background modeling. Background Art

[0002] A typical optical measurement method involves using multiple cameras to gaze and measure an area, followed by the calculation of target motion trajectory information using digital image processing techniques and optical intersection methods. The key to this method is the effective extraction of pixel coordinates of the moving target's feature points in the image. For gaze scenarios, background cancellation and Gaussian background modeling are often used to separate the target from the background. While these methods have clear principles and low computational complexity, they suffer from poor interference resistance and are prone to high levels of noise in the detection results. Another class of methods, such as CodeBook, ViBE, and Yolo, are based on random background modeling and deep learning. These algorithms offer better target detection than traditional algorithms, but they require high computational complexity and higher device-side computing power, making them difficult to meet the demands of real-time dual-channel or even multi-channel image processing on standard desktop computers.

[0003] Therefore, there is currently a lack of moving target feature point detection algorithms in gaze scenarios that can simultaneously take into account low computing power requirements and good target detection effects. Summary of the Invention

[0004] In response to the problems existing in the prior art, the present invention provides a method for detecting feature points of linear motion targets based on random neighborhood cyclic background modeling, so as to solve the technical problem that the prior art lacks a feature point detection algorithm for motion targets in staring scenarios that takes into account both low computing power requirements and good target detection effects.

[0005] The present invention provides a method for detecting feature points of a linear motion target based on random neighborhood cyclic background modeling, comprising:

[0006] S1. Collect continuous image or video frame inputs and obtain a first frame input image. For each pixel point in the first frame input image, randomly select its neighboring point pixels to construct an initialization background sample set;

[0007] S2, obtaining each pixel point in each frame image after the first frame input image, calculating the grayscale difference between each pixel point and each sample value in the corresponding initial background sample set, and determining whether the number of pixels whose grayscale difference is less than the threshold grayscale is greater than a set value, if so, determining that the pixel is a background pixel, otherwise it is a target pixel;

[0008] S3, obtaining two consecutive mask matrices, and using connected domain analysis to obtain target regions of the two consecutive mask matrices respectively;

[0009] S4. Assuming a linear moving target, combine the target area of ​​two consecutive mask matrices to obtain the pixel coordinates of the target head feature points.

[0010] Optionally, for each pixel point in the first frame input image, randomly selecting its neighboring pixels to construct an initialized background sample set, including:

[0011] S101, collecting continuous image or video frame inputs and obtaining a first frame of input image;

[0012] S102. Define the pixel grayscale VL(x,y) at the pixel coordinate (x,y) in the first frame input image VL, define the neighborhood selection matrix cx = (-1, 0, 1, -1, 1, -1, 0, 1), cy = (1, 1, 1, 0, 0, -1, -1, -1), and generate a random integer k greater than or equal to 1 and less than or equal to 8;

[0013] S103. Use the grayscale value of VL(x+cx(k),y+cy(k)) as the background sample V1 of the pixel (x,y), and repeat the operation 15 times to construct the background sample set B(x,y)=(V1,V2,...V20) corresponding to the pixel (x,y). xy ;

[0014] S104 . Repeat steps S102 and S103 for each pixel in the first frame input image VL to construct a background sample set, thereby obtaining an initial background sample set of the first frame input image VL.

[0015] Optionally, the step of obtaining each pixel point in each frame of the image after the first frame of the input image, calculating the grayscale difference between the grayscale value of each pixel point and the grayscale difference between the grayscale value of each sample value in the corresponding initial background sample set, and determining whether the number of pixels whose grayscale difference is less than the threshold grayscale is greater than a set value, and if so, determining that the pixel is a background pixel, otherwise determining that the pixel is a target pixel, includes:

[0016] S201. For each image frame V after the first input image frame, generate a random integer n greater than or equal to 1 and less than or equal to 8, generate a random integer m greater than or equal to 1 and less than or equal to 15, and generate a mask matrix M with the same size as each image frame V and all elements being 0;

[0017] S202. For each pixel V(x,y) of each frame image V, construct a B(x,y)-V(x,y) data set, and determine whether the number of pixels in the B(x,y)-V(x,y) data set that are less than a threshold grayscale R=15 is greater than 3. If so, determine that V(x,y) is a background point and execute the operation; otherwise, determine that V(x,y) is a target point and execute S204.

[0018] S203, replace the pixel V(x,y) with m elements in its corresponding background sample set B(x,y), replace the pixel V(x+cx(n),y+cy(n)) with the m-th element in its corresponding background sample set B(x+cx(n),y+cy(n)), and add 1 to m in step S201. If m exceeds 15, m is reset to 1; add 1 to n in step S201. If n exceeds 8, n is reset to 1;

[0019] S204 . Update the mask matrix M and set M(x, y) to 255.

[0020] Optionally, the acquiring of two consecutive mask matrices and respectively acquiring target regions of the two consecutive mask matrices using connected domain analysis includes:

[0021] S301, defining two consecutive mask matrices as a first mask matrix M1 and a second mask matrix M2, respectively, and performing median filtering on the first mask matrix M1 and the second mask matrix M2;

[0022] S302, performing connected domain analysis on the first mask matrix M1 and the second mask matrix M2 respectively to obtain the number of pixels corresponding to each connected domain;

[0023] S303 , selecting a connected domain with the largest number of pixels, and determining whether the number of pixels in the connected domain is greater than 10. If so, determining that the connected domain is the target area; otherwise, no response is taken.

[0024] Optionally, the assumed linear motion target is combined with the target area of ​​two consecutive mask matrices to obtain the pixel coordinates of the target head feature points, including:

[0025] S401: Define the center pixel coordinates of the minimum bounding rectangle of the target area in the first mask matrix M1 as (x1, y1), and the first coordinate set T1 of the contour points of the connected domain of the target area; define the center pixel coordinates of the minimum bounding rectangle of the target area in the second mask matrix M2 as (x2, y2), and the second coordinate set T2 of the contour points of the connected domain of the target area:

[0026] S402. Calculate the velocity vector of the linear motion target, expressed as:

[0027] V=(x2,y2)-(x1,y1)

[0028] and determining whether the modulus |V| of the velocity vector is greater than 2 pixels. If so, searching for the coordinates of the target head feature point in the second coordinate set T2; otherwise, determining that no valid linear motion target is detected and no response is taken;

[0029] S403, define the coordinate vector set CT2 from the center point of the linear motion target to the contour point as:

[0030] CT2=T2-(x2,y2)

[0031] Calculate the projection and angle of each vector in the vector set CT2 on the velocity vector V;

[0032] S404 , searching the vector set for the coordinates of the contour point with an angle less than 90° and the longest projection length, and determining that the contour point coordinates are the target head feature point coordinates based on the assumed linear motion target.

[0033] Compared with the prior art, the present invention:

[0034] The detection process consists of two stages. The first stage is the detection of moving target pixel regions. This stage uses a random neighborhood cyclic background modeling method instead of the traditional completely random background modeling method, greatly reducing the computational effort while ensuring effective background modeling. The second stage is the detection of target feature points. This method introduces linear motion constraints and uses the idea of ​​correlating the previous and next frames to detect head feature points. This detection method has the advantages of low computer hardware requirements and small computational effort, and has strong engineering application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0037] Figure 1 Schematic diagram of the method of the present invention;

[0038] Figure 2 Schematic diagram of the method flow of the present invention;

[0039] Figure 3 Schematic diagram of detection results of moving pedestrians detected by surveillance camera video according to an embodiment of the present invention. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other implementation cases obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. The functional units with the same labels in the examples of the present invention have the same and similar structures and functions.

[0041] See also Figure 1 The present invention provides a method for detecting feature points of linear motion targets based on random neighborhood cyclic background modeling, comprising:

[0042] S1. Collect continuous image or video frame inputs and obtain a first frame input image. For each pixel point in the first frame input image, randomly select its neighboring point pixels to construct an initialization background sample set;

[0043] S2, obtaining each pixel point in each frame image after the first frame input image, calculating the grayscale difference between each pixel point and each sample value in the corresponding initial background sample set, and determining whether the number of pixels whose grayscale difference is less than the threshold grayscale is greater than a set value, if so, determining that the pixel is a background pixel, otherwise it is a target pixel;

[0044] S3, obtaining two consecutive mask matrices, and using connected domain analysis to obtain target regions of the two consecutive mask matrices respectively;

[0045] S4. Assuming a linear moving target, combine the target area of ​​two consecutive mask matrices to obtain the pixel coordinates of the target head feature points.

[0046] See also Figure 2 In this embodiment, S1 collects continuous image or video frame inputs and obtains a first frame input image. For each pixel point in the first frame input image, randomly selects its neighborhood point pixels to construct an initialized background sample set.

[0047] S101, collecting continuous image or video frame inputs and obtaining a first frame of input image;

[0048] S102. Define the pixel grayscale VL(x,y) at the pixel coordinate (x,y) in the first frame input image VL, define the neighborhood selection matrix cx = (-1, 0, 1, -1, 1, -1, 0, 1), cy = (1, 1, 1, 0, 0, -1, -1, -1), and generate a random integer k greater than or equal to 1 and less than or equal to 8;

[0049] S103. Use the grayscale value of VL(x+cx(k),y+cy(k)) as the background sample V1 of the pixel (x,y), and repeat the operation 15 times to construct the background sample set B(x,y)=(V1,V2,...V20) corresponding to the pixel (x,y). xy ;

[0050] S104 . Repeat steps S102 and S103 for each pixel in the first frame input image VL to construct a background sample set, thereby obtaining an initial background sample set of the first frame input image VL.

[0051] S2. Obtain each pixel point in each frame image after the first frame input image, calculate the grayscale difference between each pixel point grayscale value and each sample value in the corresponding initial background sample set, and determine whether the number of pixels whose grayscale difference is less than the threshold grayscale is greater than a set value. If so, the pixel is determined to be a background pixel, otherwise it is a target pixel.

[0052] S201. For each image frame V after the first input image frame, generate a random integer n greater than or equal to 1 and less than or equal to 8, generate a random integer m greater than or equal to 1 and less than or equal to 15, and generate a mask matrix M with the same size as each image frame V and all elements being 0;

[0053] S202. For each pixel V(x,y) of each frame image V, construct a B(x,y)-V(x,y) data set, and determine whether the number of pixels in the B(x,y)-V(x,y) data set that are less than a threshold grayscale R=15 is greater than 3. If so, determine that V(x,y) is a background point and execute the operation; otherwise, determine that V(x,y) is a target point and execute S204.

[0054] S203, replace the pixel V(x,y) with m elements in its corresponding background sample set B(x,y), replace the pixel V(x+cx(n),y+cy(n)) with the m-th element in its corresponding background sample set B(x+cx(n),y+cy(n)), and add 1 to m in step S201. If m exceeds 15, m is reset to 1; add 1 to n in step S201. If n exceeds 8, n is reset to 1;

[0055] S204 . Update the mask matrix M and set M(x, y) to 255.

[0056] S3. Obtain two consecutive mask matrices, and use connected component analysis to obtain target regions of the two consecutive mask matrices.

[0057] S301, defining two consecutive mask matrices as a first mask matrix M1 and a second mask matrix M2, respectively, and performing median filtering on the first mask matrix M1 and the second mask matrix M2;

[0058] S302, performing connected domain analysis on the first mask matrix M1 and the second mask matrix M2 respectively to obtain the number of pixels corresponding to each connected domain;

[0059] S303 , selecting a connected domain with the largest number of pixels, and determining whether the number of pixels in the connected domain is greater than 10. If so, determining that the connected domain is the target area; otherwise, no response is taken.

[0060] S4. Assuming a linear moving target, combine the target area of ​​two consecutive mask matrices to obtain the pixel coordinates of the target head feature points.

[0061] S401: Define the center pixel coordinates of the minimum bounding rectangle of the target area in the first mask matrix M1 as (x1, y1), and the first coordinate set T1 of the contour points of the connected domain of the target area; define the center pixel coordinates of the minimum bounding rectangle of the target area in the second mask matrix M2 as (x2, y2), and the second coordinate set T2 of the contour points of the connected domain of the target area:

[0062] S402. Calculate the velocity vector of the linear motion target, expressed as:

[0063] V=(x2,y2)-(x1,y1)

[0064] and determining whether the modulus |V| of the velocity vector is greater than 2 pixels. If so, searching for the coordinates of the target head feature point in the second coordinate set T2; otherwise, determining that no valid linear motion target is detected and no response is taken;

[0065] S403, define the coordinate vector set CT2 from the center point of the linear motion target to the contour point as:

[0066] CT2=T2-(x2,y2)

[0067] Calculate the projection and angle of each vector in the vector set CT2 on the velocity vector V;

[0068] S404. Find the coordinates of the contour point with an angle less than 90° and the longest projection length from the vector set, and combine it with the assumed linear motion target (based on the cylindrical target assumption of linear motion) to determine that the contour point coordinates are the target head feature point coordinates.

[0069] See also Figure 3 The video input resolution used is 640*512; the traditional random background modeling method takes 29ms to process in real time, while the patented algorithm takes 13ms to process.

[0070] In response to the need to simultaneously process two-way real-time video for target feature point detection in gaze scenarios, this paper designs a method for detecting linear motion target feature points based on random neighborhood cyclic background modeling. This method divides target feature point detection into two stages:

[0071] The first stage involves detecting the moving target pixel region. For each pixel in the first frame of the input image, an initial sample set is constructed, and foreground and background pixel determination is performed in subsequent images. After this determination is complete, the background model is updated. For each background pixel, a random starting point is used, using a cyclic extraction method to extract neighboring pixels to replace elements in the background model. To determine which background model element to update, the background model element is also randomly selected using a cyclic extraction method. This random starting point substitution method significantly reduces the computational effort required to generate random numbers, lowering the requirements for computer hardware. A target mask matrix is ​​then constructed, in which background pixels are set to 0 and foreground pixels are set to 255.

[0072] The second stage involves target feature point detection. Two consecutive mask matrices are obtained, and connected domain analysis is used to identify the target region in each frame. Based on the assumption of linear motion, inter-frame motion vectors are calculated. Feature points on the target's head are selected based on the projection and angle between the vectors formed by the target's connected domain outline pixels and the target's center pixels. This method balances target feature point detection performance with computational complexity, enabling real-time dual-channel image processing and possessing strong engineering application value.

[0073] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0074] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for detecting feature points of linear motion targets based on random neighborhood cyclic background modeling, characterized in that: include: S1. Collect continuous image or video frame inputs and obtain a first frame input image. For each pixel point in the first frame input image, randomly select its neighboring point pixels to construct an initialization background sample set; S2, obtaining each pixel point in each frame image after the first frame input image, calculating the grayscale difference between each pixel point and each sample value in the corresponding initial background sample set, and determining whether the number of pixels whose grayscale difference is less than the threshold grayscale is greater than a set value, if so, determining that the pixel is a background pixel, otherwise it is a target pixel; S3, obtaining two consecutive mask matrices, and using connected domain analysis to obtain target regions of the two consecutive mask matrices respectively; S4. Assuming a linear moving target, combine the target area of ​​two consecutive mask matrices to obtain the pixel coordinates of the target head feature points.

2. The method for detecting feature points of linear motion targets using random neighborhood cyclic background modeling as claimed in claim 1, wherein: For each pixel point in the first frame input image, randomly select its neighboring pixel points to construct an initialized background sample set, including: S101, collecting continuous image or video frame inputs and obtaining a first frame of input image; S102. Define the pixel grayscale VL(x,y) at the pixel coordinate (x,y) in the first frame input image VL, define the neighborhood selection matrix cx = (-1, 0, 1, -1, 1, -1, 0, 1), cy = (1, 1, 1, 0, 0, -1, -1, -1), and generate a random integer k greater than or equal to 1 and less than or equal to 8; S103. Use the grayscale value of VL(x+cx(k),y+cy(k)) as the background sample V1 of the pixel (x,y), and repeat the operation 15 times to construct the background sample set B(x,y)=(V1,V2,...V20) corresponding to the pixel (x,y). xy ; S104 . Repeat steps S102 and S103 for each pixel in the first frame input image VL to construct a background sample set, thereby obtaining an initial background sample set of the first frame input image VL.

3. The method for detecting feature points of linear motion targets using random neighborhood cyclic background modeling as claimed in claim 2, wherein: The method further comprises: obtaining each pixel point in each frame image after the first frame input image, calculating the grayscale difference between each pixel point grayscale value and each sample value in the corresponding initial background sample set, and determining whether the number of pixels whose grayscale difference is less than the threshold grayscale is greater than a set value; if so, determining that the pixel is a background pixel; otherwise, determining that the pixel is a target pixel, including: S201. For each image frame V after the first input image frame, generate a random integer n greater than or equal to 1 and less than or equal to 8, generate a random integer m greater than or equal to 1 and less than or equal to 15, and generate a mask matrix M with the same size as each image frame V and all elements being 0; S202. For each pixel V(x,y) of each frame image V, construct a B(x,y)-V(x,y) data set, and determine whether the number of pixels in the B(x,y)-V(x,y) data set that are less than a threshold grayscale R=15 is greater than 3. If so, determine that V(x,y) is a background point and execute the operation; otherwise, determine that V(x,y) is a target point and execute S204. S203, replace the pixel V(x,y) with m elements in its corresponding background sample set B(x,y), replace the pixel V(x+cx(n),y+cy(n)) with the m-th element in its corresponding background sample set B(x+cx(n),y+cy(n)), and add 1 to m in step S201. If m exceeds 15, m is reset to 1; add 1 to n in step S201. If n exceeds 8, n is reset to 1; S204 . Update the mask matrix M and set M(x, y) to 255.

4. The method for detecting feature points of linear motion targets using random neighborhood cyclic background modeling as claimed in claim 3, wherein: The step of obtaining two consecutive mask matrices and respectively obtaining target regions of the two consecutive mask matrices using connected domain analysis includes: S301, defining two consecutive mask matrices as a first mask matrix M1 and a second mask matrix M2, respectively, and performing median filtering on the first mask matrix M1 and the second mask matrix M2; S302, performing connected domain analysis on the first mask matrix M1 and the second mask matrix M2 respectively to obtain the number of pixels corresponding to each connected domain; S303 , selecting a connected domain with the largest number of pixels, and determining whether the number of pixels in the connected domain is greater than 10. If so, determining that the connected domain is the target area; otherwise, no response is taken.

5. The method for detecting feature points of linear motion targets using random neighborhood cyclic background modeling as claimed in claim 4, wherein: The assumed linear motion target is combined with the target area of ​​two consecutive mask matrices to obtain the pixel coordinates of the target head feature points, including: S401: Define the center pixel coordinates of the minimum bounding rectangle of the target area in the first mask matrix M1 as (x1, y1), and the first coordinate set T1 of the contour points of the connected domain of the target area; define the center pixel coordinates of the minimum bounding rectangle of the target area in the second mask matrix M2 as (x2, y2), and the second coordinate set T2 of the contour points of the connected domain of the target area: S402. Calculate the velocity vector of the linear motion target, expressed as: V=(x2,y2)-(x1,y1) and determining whether the modulus |V| of the velocity vector is greater than 2 pixels. If so, searching for the coordinates of the target head feature point in the second coordinate set T2; otherwise, determining that no valid linear motion target is detected and no response is taken; S403, define the coordinate vector set CT2 from the center point of the linear motion target to the contour point as: CT2=T2-(x2,y2) Calculate the projection and angle of each vector in the vector set CT2 on the velocity vector V; S404 , searching the vector set for the coordinates of the contour point with an angle less than 90° and the longest projection length, and determining that the contour point coordinates are the target head feature point coordinates based on the assumed linear motion target.