A video-based site occupancy recognition method

Through the monocular depth estimation algorithm and the connected area analysis, combined with the depth difference judgment, the problem of misidentification of traditional video detection algorithms when identifying site occupation is solved, and effective identification and accurate judgment of the occupation of non-fixed categories of objects is achieved.

CN115731496BActive Publication Date: 2025-06-06ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211488408.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-25
Publication Date
2025-06-06
Estimated Expiration
2042-11-25

AI Technical Summary

Technical Problem

Traditional video-based site occupation detection algorithms cannot distinguish between apparent changes in the channel and image changes caused by occupation, cannot distinguish between false occlusion and real occupation, and the object detection algorithm cannot recognize object occupation of non-fixed categories, and the output image area cannot be directly converted into site coordinates, which is prone to misjudgment.

Method used

The monocular depth estimation calculation method is used to estimate the depth image of the scene. Through the detection of the depth change area and the analysis of the connected area, the depth difference value is used to determine whether the three-dimensional space point is connected to the unoccupied area, and then the occupied state of the object is identified.

Benefits of technology

Effectively identifying the occupation of objects in the site, reducing the risk of false alarms, and being able to identify the occupation of objects of non-fixed categories, solving the problem of misidentification of traditional algorithms when scene appearance changes and occlusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115731496B_ABST
    Figure CN115731496B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of intelligent security technology, and discloses a method for identifying site occupancy based on video, which includes the following steps: Step 1, capture a frame of image in the case of no occupancy, and use a monocular depth estimation algorithm to estimate the depth image of the scene, so as to obtain the scene depth d0(x, y) corresponding to each pixel in the video image; Step 2, during the system monitoring process, for the video image captured at any moment, use a monocular depth estimation algorithm to estimate the depth image d. The method for identifying site occupancy proposed by the present invention can effectively identify the occupancy of objects in the scene, especially can effectively solve the problem of misidentification generated by traditional occupancy recognition algorithms when the appearance of the scene changes or is blocked, and effectively reduce the risk of false alarms; and there is no restriction on the category of the occupied object, and it can identify the occupancy of objects with unfixed categories.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent security technology, and in particular to a site occupancy recognition method based on video. Background Art

[0002] According to relevant national regulations, certain areas of public places, such as fire escapes, venue entrances and exits, traffic roads, etc., cannot be occupied by objects and people for a long time to avoid accidents or disasters; in order to automatically detect and identify the occupancy of these areas, the present invention provides a technical solution based on video analysis.

[0003] Traditional video-based site occupancy detection algorithms mainly use background modeling and change detection algorithms. This type of algorithm mainly estimates the video background image from the video data, compares the real-time acquired video image with the background image, calculates the changed area, and identifies the changed area as the occupied area.

[0004] Another video-based occupancy detection algorithm uses the currently more mature target detection technology to detect and locate some fixed categories of targets (such as vehicles and pedestrians) in video images, and then determines whether it is an occupancy behavior based on the length of time these targets appear in the specified area.

[0005] Traditional video-based site occupancy detection algorithms mainly use background modeling and change detection algorithms. These algorithms face two problems:

[0006] 1. It is impossible to distinguish the changes in channel appearance from the changes in the image caused by occupancy. For example, the image area of ​​the road surface caused by water accumulation has changed significantly, but this change is not caused by the road being occupied, and the traditional background subtraction method cannot distinguish it.

[0007] 2. Unable to distinguish between false occupancy and real occupation. Due to the influence of perspective projection, the image area of ​​the channel may be blocked by other objects, but it is not actually occupied. Traditional algorithms cannot distinguish this situation.

[0008] The target detection method has the following disadvantages:

[0009] 1. The current target detection algorithm uses supervised machine learning (mainly deep learning) technology and uses a large number of samples to train the target detection model. Once the model is trained, the types of targets that the model can detect remain fixed. Objects that do not appear in the training samples cannot be identified, so arbitrary object occupancy cannot be identified.

[0010] 2. The target detection algorithm outputs the target area on the image plane. Due to the perspective projection effect, this image area cannot be directly transformed into site coordinates, and it cannot accurately determine whether the object actually appears inside the unoccupied area, so it is easy to make misjudgments. Summary of the invention

[0011] In order to solve the technical problems raised in the background technology, the present invention provides a video-based site occupancy recognition method.

[0012] The present invention is implemented by the following technical solution: A method for identifying site occupancy based on video, comprising the following steps:

[0013] Step 1: Collect a frame of image without occupancy, use the monocular depth estimation algorithm to estimate the depth image of the scene, and obtain the scene depth d corresponding to each pixel in the video image 0 (x,y);

[0014] Step 2: During the system monitoring process, for the video images collected at any time, a monocular depth estimation algorithm is used to estimate the depth image d;

[0015] Step 3: Use d 0 ,d calculate the pixels whose depth has changed in the whole image, and obtain the area Ω where the depth has changed;

[0016] Step 4: Use the connected area detection algorithm to detect each connected area in Ω, and filter out the connected areas that overlap with the user-specified unoccupiable area R to obtain Ω i ,i=1...,k,Ω i represents the i-th connected region;

[0017] Step 5: For each connected region Ω i , calculate its boundary pixel set ω i ;

[0018] Step 6: For each connected region Ω i , according to ω i Each pixel (u, v) in R-Ω i The depth difference of the adjacent pixels (u', v') in the image is used to determine whether the 3D space point corresponding to the pixel (u, v) is connected to the physical space corresponding to the region R. When t < ξ, it is considered to be connected. At this time, Ω i The corresponding physical space occupies the physical space corresponding to R, and ξ is an empirical parameter.

[0019] Step 7: When Ω i The corresponding physical space occupies the physical space corresponding to R, then Ω i All pixels in are recorded as occupied;

[0020] Step 8: If a connected area in R that exceeds the area given by the user is continuously occupied for a period of time exceeding a threshold given by the user, the area is determined to be illegally occupied.

[0021] Optionally, the area Ω in step 3 is calculated using the following formula: Ω = {(u,v): |d 0 (u,v)-d(u,v)|>τ}, τ is an empirical parameter.

[0022] Optionally, the depth difference in step six is ​​t=|d(u,v)-d(u',v')|.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] The site occupancy identification method proposed in the present invention can effectively identify the occupancy of objects in the scene, and in particular can effectively solve the problem of misidentification caused by traditional occupancy identification algorithms when the scene appearance changes or is blocked, and effectively reduce the risk of false alarms; and there is no restriction on the category of occupied objects, and the occupancy of objects of non-fixed categories can be identified. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a frame of image collected at any time in the embodiment of the present invention;

[0026] Figure 2 A depth map obtained according to a monocular depth estimation algorithm in an embodiment of the present invention;

[0027] Figure 3 for Figure 2 The depth map of the image is segmented by difference with the depth map when it is not occupied, and the depth change area is obtained;

[0028] Figure 4 In the embodiment of the present invention, each pixel (u, v) based on the boundary pixel set in each connected area and its R-Ω i The depth difference of adjacent pixels (u', v') in is obtained;

[0029] Figure 5 for Figure 4 In the depth difference map in the connected area, the depth difference corresponding to each pixel (u, v) of the boundary pixel set is less than a certain threshold, indicating that this area is connected to the ground.

[0030] Figure 6 for Figure 4 In the depth difference map in the connected area, the depth difference corresponding to each pixel (u, v) of the boundary pixel set is greater than a certain threshold, indicating that this area is not connected to the ground. DETAILED DESCRIPTION

[0031] The present invention is further described below in conjunction with the accompanying drawings and specific implementation methods. It should be noted that, under the premise of no conflict, the various embodiments or technical features described below can be arbitrarily combined to form a new embodiment.

[0032] Example:

[0033] Please combine Figure 1 This embodiment is described by taking the case of parking a battery-powered vehicle next to an outdoor parking space as an example.

[0034] First, refer to Figure 1 , collect a frame of image at any time, and the area marked by the black line segment in the figure is the unoccupiable area R.

[0035] Using the monocular depth estimation algorithm, the depth map is calculated. According to the depth map and the reference depth map when not occupied, the depth change in the unoccupied area R can be obtained. Figure 2 Among them, the monocular depth estimation algorithm is the dense-prediction-transformer (DPT) algorithm. The DPT algorithm is based on the ViT backbone architecture and is a dense prediction structure based on the transformer encoder-decoder structure, which can be used for monocular depth estimation tasks.

[0036] The above depth map is differentially segmented with the reference depth map to obtain the depth change area, such as Figure 3 The white part indicates the area where the depth changes. i , in this case there is only one connected region;

[0037] According to each pixel (u, v) of the boundary pixel set in each connected region and its i The depth difference between adjacent pixels (u', v') in t = |d(u, v) - d(u', v')|, and the result is as follows Figure 4 As shown. Among them, the greater the brightness of the edge part, the greater t is;

[0038] According to the above difference map, it is determined whether the three-dimensional space point corresponding to the pixel edge (u, v) is connected to the physical space corresponding to the region R. When t < ξ, it is determined to be connected, which means Ω i The corresponding physical space occupies the physical space corresponding to R. ξ is an empirical parameter. Figure 5 The edge pixels satisfying t<ξ are shown. These edge points are the contact points between the object (electric car) and the ground.

[0039] Figure 6The edge pixels that satisfy t>ξ are displayed. These edge points are the contour points generated by the object (electric car) blocking the ground, not the contact points;

[0040] Depend on Figure 6 So, at this time Ω i The corresponding physical space is in the R area and has no contact with the ground, so it does not occupy the physical space corresponding to R, so Ω i All pixels in are recorded as unoccupied;

[0041] In the entire R region, there is only one connected region, which does not occupy R, so it is determined that R is not occupied.

[0042] The above-mentioned embodiments are only preferred embodiments of the present invention and cannot be used to limit the scope of protection of the present invention. Any non-substantial changes and substitutions made by technicians in this field on the basis of the present invention shall fall within the scope of protection required by the present invention.

Claims

1. A video-based site occupancy recognition method, It is characterized in that The steps include: Step 1: Collect a frame of image without occupancy, use the monocular depth estimation algorithm to estimate the depth image of the scene, and obtain the scene depth d corresponding to each pixel in the video image 0 (x,y); Step 2: During the system monitoring process, for the video images collected at any time, a monocular depth estimation algorithm is used to estimate the depth image d; Step 3: Use d 0 ,d calculate the pixels whose depth has changed in the whole image, and obtain the area Ω where the depth has changed; Step 4: Use the connected area detection algorithm to detect each connected area in Ω, and filter out the connected areas that overlap with the user-specified unoccupiable area R to obtain Ω i ,i=1...,k,Ω i represents the i-th connected region; Step 5: For each connected region Ω i , calculate its boundary pixel set ω i ; Step 6: For each connected region Ω i , according to ω i Each pixel (u, v) in R-Ω i The depth difference of the adjacent pixels (u', v') in the image is used to determine whether the 3D space point corresponding to the pixel (u, v) is connected to the physical space corresponding to the region R. When t < ξ, it is considered to be connected. At this time, Ω i The corresponding physical space occupies the physical space corresponding to R, and ξ is an empirical parameter; Step 7: When Ω i The corresponding physical space occupies the physical space corresponding to R, then Ω i All pixels in are recorded as occupied; Step 8: If a connected area in R that exceeds the area given by the user is continuously occupied for a period of time that exceeds a threshold given by the user, the area is determined to be illegally occupied.

2. A method for identifying site occupancy based on video according to claim 1, It is characterized in that The area Ω in step 3 is calculated using the following formula: Ω = {(u,v): |d 0 (u,v)-d(u,v)|>τ}, τ is an empirical parameter.

3. A method for identifying site occupancy based on video according to claim 1, It is characterized in that The depth difference in step six is ​​t=|d(u,v)-d(u',v')|.

Citation Information

Patent Citations

  • Video image depth estimation method based on image segmentation

    CN105069808A

  • Systems, methods, and media for detecting object-free space

    US20210233311A1