Illegal event detection method, device, equipment and storage medium

By using pixel point matching and Kalman filtering models in road violation event detection, combined with in-depth information, the problems of missed detection, missed detection and computing resource consumption in the existing technology are solved, and efficient detection of illegal incidents is achieved.

CN115170614BActive Publication Date: 2025-05-23CHONGQING UNISINSIGHT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210861210.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-20
Publication Date
2025-05-23
Estimated Expiration
2042-07-20

AI Technical Summary

Technical Problem

In the prior art, road violation event detection relies on two-dimensional images, resulting in mis-checking and missed inspection problems; while three-dimensional reconstruction technology requires a large amount of computing resources, has a slow processing speed, and is not suitable for illegal event detection.

Method used

By acquiring multiple road images at multiple historical moments, pixel points are matched to determine the view difference and image depth information, combining the Kalman filter model to predict the target box position, determine the position of the line segment to be detected, and judge the pinning status of the target to be detected to detect illegal events.

Benefits of technology

By introducing in-depth information, this method saves computing power, improves processing speed and efficiency, and can effectively detect illegal events, avoid mis-detection and missed inspections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115170614B_ABST
    Figure CN115170614B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, device and storage medium for detecting illegal events. The method matches pixels of multiple groups of first road images to be detected and second road images to be detected to obtain view differences, thereby determining image depth information, and performing target detection on one of the first road images to be detected and the second road images to be detected to obtain target detection frame positions, thereby determining predicted target frame position information, determining line segment position information to be detected based on motion vectors of the target to be detected and predicted target frame position information, determining the line-pressing state of the target to be detected through the line segment position information to be detected and the regular traffic line position information, so as to detect illegal events of the target to be detected, and predicting the position of the target to be detected by introducing depth information, thereby providing a solution that does not rely on processing of point cloud data sets, saves computing power, improves processing speed and efficiency, and can realize illegal event detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an illegal event detection method, device, equipment and storage medium. Background Art

[0002] Road violation detection, such as line crossing detection, often requires pre-configuration of scene information, which is usually generated by manual drawing or algorithm-generated two-dimensional images, resulting in the loss of spatial information, leading to false detections and missed detections due to the lack of spatial interaction in the violation detection process.

[0003] For the 3D reconstruction of traffic scenes, we usually obtain multi-view images of objects in the scene to obtain point clouds, detect feature information including point features and line features from the images, match the detected feature information, and complete pose estimation and depth estimation based on the matched feature information, and finally obtain the 3D scene model. However, the calculation of point cloud data sets often requires a lot of computing resources, with slow processing speed and low processing efficiency. Therefore, the above 3D reconstruction method is not suitable for the illegal event detection process. Summary of the invention

[0004] In view of the shortcomings of the prior art mentioned above, the purpose of the present invention is to provide a method, device, equipment and storage medium for detecting illegal events, which are used to solve the problems of false detection and missed detection in illegal event detection on two-dimensional images in related technologies, and the technical problems that traditional three-dimensional reconstruction relies on the processing of point cloud data sets, requires a lot of computing resources, has slow processing speed, low processing efficiency and is not suitable for illegal event detection.

[0005] In view of the above problems, the present invention provides a method for detecting illegal events, the method comprising:

[0006] Acquire a plurality of historical road images to be detected at a plurality of historical moments of a preset traffic area by a depth image acquisition device, wherein the historical road images to be detected include a first road image to be detected and a second road image to be detected, and the preset traffic area includes a regular traffic line;

[0007] performing pixel matching on a first pixel of the first road image to be detected and a second pixel of the second road image to be detected, using a position difference in a preset direction between the matched first pixel and the second pixel as a view difference, and determining pixel depth information of the first pixel based on the view difference, so as to determine image depth information of each of the first road images to be detected;

[0008] Performing target detection on each of the first road images to be detected to obtain detection target frame position information of the target to be detected in each of the first road images to be detected;

[0009] Determine the predicted target frame position information through a Kalman filter model based on the detection target frame position information and the image depth information;

[0010] Determining a motion vector of the target to be detected based on a target position of the target to be detected in each of the first road images to be detected;

[0011] Determine the position information of the line segment to be detected of the line segment to be detected of the target to be detected based on the predicted target frame position information and the motion vector;

[0012] The line-crossing state of the target to be detected is determined according to the line segment position information to be detected and the regular traffic line position information of the regular traffic line, so as to detect the illegal event of the target to be detected.

[0013] In one embodiment of the present invention, performing pixel matching on a first pixel of the first road image to be detected and a second pixel of the second road image to be detected includes:

[0014] Obtaining a first pixel value, first pixel position information, and a first pixel mean value of a pixel window where the first pixel is located, as well as a second pixel value, second pixel position information, and a second pixel mean value of a pixel window where the second pixel is located, of the first pixel point;

[0015] Determine an offset between the first pixel point and the second pixel point according to the first pixel position information and the second pixel position information;

[0016] If the offset is less than a preset offset threshold, determining a first difference value according to the first pixel value and the first pixel mean, and determining a second difference value according to the second pixel value and the second pixel mean;

[0017] Determine a correlation between the first pixel and the second pixel based on the first difference and the second difference;

[0018] If the correlation is greater than a preset correlation threshold, the first pixel point matches the second pixel point;

[0019] If the correlation is less than or equal to the preset correlation threshold, the first pixel point does not match the second pixel point.

[0020] In one embodiment of the present invention, the method for determining the relevance includes:

[0021]

[0022] Among them, ncc(I 1 I 2 ) is the correlation, I1 (x) is the first pixel value of the first pixel, I 2 (x) is the second pixel value of the second pixel, the offset between the first pixel and the second pixel is less than the preset offset threshold, μ 1 is the first pixel mean, μ 2 is the second pixel mean.

[0023] In an embodiment of the present invention, determining the pixel depth information of the first pixel point based on the view difference includes:

[0024] Obtaining a baseline and a focal length of the depth image acquisition device;

[0025] determining a view basis according to the view difference and the baseline;

[0026] The pixel depth information is determined based on the view basis and the focal length.

[0027] In one embodiment of the present invention, before determining the line-crossing state of the target to be detected according to the line segment position information to be detected and the regular traffic line position information of the regular traffic line, the illegal event detection method includes:

[0028] Acquiring a rotation state of the depth image acquisition device;

[0029] If the rotation state includes rotation, the regular traffic line position information of the regular traffic line in the first road image to be detected is detected by a preset image segmentation model.

[0030] In one embodiment of the present invention, the predicted target frame is a quadrilateral, and determining the position information of the line segment to be detected of the target to be detected based on the predicted target frame position information and the motion vector includes:

[0031] Connecting the opposite sides of the quadrilateral along the moving direction of the target object to obtain a relative connection line;

[0032] Determine two intersection points of the straight line formed by the motion vector passing through the line auxiliary point and the quadrilateral, and determine the line between the two intersection points as the line segment to be detected, and the line auxiliary point is a point on the relative line.

[0033] In one embodiment of the present invention, determining the line-crossing state of the target to be detected according to the line segment position information to be detected and the regular traffic line position information of the regular traffic line to detect the illegal event of the target to be detected includes:

[0034] Determine the intersection of the line segment to be detected and the regular traffic line according to the position information of the line segment to be detected and the regular traffic line position information of the regular traffic line;

[0035] If the line segment to be detected intersects with the regular traffic line, the line-crossing state is determined to be line-crossing, and the target to be detected has an illegal event;

[0036] If the line segment to be detected is separated from the regular traffic line, the line-crossing state is determined to be not crossing the line, and there is no illegal event for the target to be detected.

[0037] The embodiment of the present invention further provides a device for detecting illegal events, the device comprising:

[0038] An acquisition module, used to acquire a plurality of historical road images to be detected at a plurality of historical moments of a preset traffic area acquired by a depth image acquisition device, wherein the historical road images to be detected include a first road image to be detected and a second road image to be detected, and the preset traffic area includes a regular traffic line;

[0039] a matching module, configured to perform pixel matching on a first pixel of the first road image to be detected and a second pixel of the second road image to be detected, and determine pixel depth information of the first pixel based on a position difference in a preset direction between the matched first pixel and the second pixel as a view difference, so as to determine image depth information of each of the first road images to be detected;

[0040] A detection module, used to perform target detection on each of the first road images to be detected, and obtain detection target frame position information of the target to be detected in each of the first road images to be detected;

[0041] A predicted target frame determination module, used to determine the predicted target frame position information through a Kalman filter model based on the position information of each detected target frame and the image depth information;

[0042] A motion vector determination module, used to determine the motion vector of the target to be detected based on the target position of the target to be detected in each of the first road images to be detected;

[0043] A to-be-detected line segment determination module, used to determine the to-be-detected line segment position information of the to-be-detected line segment of the to-be-detected target based on the predicted target frame position information and the motion vector;

[0044] The illegal event detection module is used to determine the line-crossing state of the target to be detected according to the position information of the line segment to be detected and the regular traffic line position information of the regular traffic line, so as to detect the illegal event of the target to be detected.

[0045] An embodiment of the present invention further provides an electronic device, including a processor, a memory, and a communication bus;

[0046] The communication bus is used to connect the processor and the memory;

[0047] The processor is used to execute the computer program stored in the memory to implement the method as described in any one of the above embodiments.

[0048] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. The computer program is used to enable the computer to execute the method as described in any one of the above embodiments.

[0049] As described above, the illegal event detection method, device, equipment and storage medium provided by the present invention have the following beneficial effects:

[0050] By matching pixels of multiple groups of first road images to be detected and second road images to be detected, view difference is obtained, and then image depth information is determined. Target detection is performed on one of the first road images to be detected and the second road images to be detected, and the target detection frame position is obtained, and then the predicted target frame position information can be determined. The position information of the line segment to be detected is determined based on the motion vector of the target to be detected and the predicted target frame position information. The line-pressing state of the target to be detected is determined through the position information of the line segment to be detected and the position information of the regular traffic line, so as to detect illegal events of the target to be detected. By introducing depth information to predict the position of the target to be detected, a solution is provided that does not rely on the processing of point cloud data sets, saves computing power, improves processing speed and efficiency, and can realize illegal event detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a schematic diagram of an implementation environment of an illegal event detection system shown in an exemplary embodiment of the present application;

[0052] Figure 2 is a flow chart of a method for detecting illegal events shown in an exemplary embodiment of the present application;

[0053] Figure 3 is a schematic diagram of a line segment to be detected shown in an exemplary embodiment of the present application;

[0054] Figure 4 is a flowchart of a specific illegal event detection method shown in an exemplary embodiment of the present application;

[0055] Figure 5 is a block diagram of an illegal event detection device shown in an exemplary embodiment of the present application;

[0056] Figure 6 A schematic diagram of the structure of an electronic device provided by an embodiment. DETAILED DESCRIPTION

[0057] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0058] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0059] See also Figure 1 , Figure 1 FIG. 1 is a schematic diagram of an implementation environment of an illegal event detection system in a related art shown in an exemplary embodiment of the present application. Figure 1 As shown, the illegal event detection system includes a client 102 and a server 101, wherein the server 101 may include an independently operated server, a distributed server, or a server cluster composed of multiple servers. The server 101 may include a network communication unit, a processor, a memory, etc. The server 101 may be used to execute the roadside fitting prediction method provided in this embodiment. The client 102 may include physical devices of the type of smart phones, desktop computers, tablet computers, laptops, digital assistants, smart wearable devices, vehicle-mounted terminals, etc., and may also include software running in physical devices, such as web pages provided to users by some service providers, or applications provided to users by these service providers. The client 102 may also be used to execute the roadside fitting prediction method provided in this embodiment, or to execute the illegal event detection method provided in this embodiment through the interaction between the server and the client. This is not limited here.

[0060] See also Figure 2 , Figure 2 is a flowchart of an illegal event detection method shown in an exemplary embodiment of the present application. The method can be applied to Figure 1 It should be understood that the method can also be applied to other exemplary implementation environments and be specifically executed by devices in other implementation environments, and this embodiment does not limit the implementation environment to which the method is applicable.

[0061] Road violation detection, such as line crossing detection, often requires pre-configuration of scene information, which is usually generated by manual drawing or algorithm-generated two-dimensional images, resulting in the loss of spatial information, leading to false detections and missed detections due to the lack of spatial interaction in the violation detection process.

[0062] For the 3D reconstruction of traffic scenes, we usually obtain multi-view images of objects in the scene to obtain point clouds, detect feature information including point features and line features from the images, match the detected feature information, and complete pose estimation and depth estimation based on the matched feature information, and finally obtain the 3D scene model. However, the calculation of point cloud data sets often requires a lot of computing resources, with slow processing speed and low processing efficiency. Therefore, the above 3D reconstruction method is not suitable for the illegal event detection process.

[0063] To solve these problems, the embodiments of the present application respectively propose a method for detecting an illegal event, a device for detecting an illegal event, an electronic device, and a computer-readable storage medium, which will be described in detail below.

[0064] like Figure 2 As shown, in an exemplary embodiment, the illegal event detection method includes at least steps S201 to S207, which are described in detail as follows:

[0065] Step S201, acquiring a plurality of historical images of roads to be inspected at a plurality of historical moments in a preset traffic area through a depth image acquisition device.

[0066] Among them, the depth image acquisition device can be a depth camera, a binocular camera and other devices. For example, a binocular camera set in a preset traffic area such as an electronic police, a traffic checkpoint, a high-speed checkpoint scene, etc. Usually, the setting position of the depth image acquisition device will remain unchanged. In the application scenario of the above example, the depth image acquisition device often has a fixed point position and no complex posture transformation, resulting in insufficient multi-angle view resources, which is not suitable for complex point cloud data sets and linear reconstruction, and the judgment of traffic violations does not require information such as surrounding trees and houses. Therefore, on the one hand, it is not suitable to use the three-dimensional reconstruction scheme in the related technology for three-dimensional reconstruction.

[0067] The historical road images to be detected include a first road image to be detected and a second road image to be detected. When the depth image acquisition device is a binocular camera or other device, two images will be acquired at the same time, namely the first road image to be detected and the second road image to be detected.

[0068] The preset traffic area includes regular traffic lines, which may be guide lines, stop lines, crosswalk lines, and other traffic indication lines specified by those skilled in the art.

[0069] Step S202, pixel matching is performed on a first pixel point of a first road image to be detected and a second pixel point of a second road image to be detected, and a position difference in a preset direction between the matched first pixel point and the second pixel point is used as a view difference, and pixel depth information of the first pixel point is determined based on the view difference to determine image depth information of each first road image to be detected.

[0070] Each pixel in the first road image to be detected can be determined as a first pixel, and correspondingly, each pixel in the second road image to be detected can be determined as a second pixel. Since the first road image to be detected and the second road image to be detected are images of the same preset traffic area, the objects captured by the two have correspondence, and the first pixel and the second pixel can be matched.

[0071] In one embodiment, performing pixel matching on a first pixel of a first road image to be detected and a second pixel of a second road image to be detected includes:

[0072] Obtain a first pixel value, first pixel position information, and a first pixel mean value of a pixel window where the first pixel is located, and a second pixel value, second pixel position information, and a second pixel mean value of a pixel window where the second pixel is located, wherein the first pixel position is the coordinate information of the first pixel, the first pixel mean value may be the average value of the pixel values ​​of each pixel in the image area (pixel window) including the first pixel, and correspondingly, the second pixel position is the coordinate information of the second pixel, and the second pixel mean value may be the average value of the pixel values ​​of each pixel in the image area (pixel window) including the second pixel;

[0073] Determine an offset between a first pixel point and a second pixel point according to the first pixel position information and the second pixel position information, wherein when the image coordinate system is a two-dimensional coordinate system, the offset may be a coordinate difference in a direction of a coordinate axis parallel to the ground;

[0074] If the offset is less than the preset offset threshold, determining a first difference value according to the first pixel value and the first pixel mean, and determining a second difference value according to the second pixel value and the second pixel mean;

[0075] Determine a correlation between the first pixel and the second pixel based on the first difference and the second difference;

[0076] If the correlation is greater than a preset correlation threshold, the first pixel point matches the second pixel point;

[0077] If the correlation is less than or equal to the preset correlation threshold, the first pixel point does not match the second pixel point.

[0078] The preset relevant threshold value can be set by those skilled in the art as needed and is not limited here.

[0079] In one embodiment, the method of determining the relevance includes:

[0080]

[0081] Among them, ncc(I 1 I 2 ) is the correlation, I 1 (x) is the first pixel value of the first pixel, I 2 (x) is the second pixel value of the second pixel, the offset between the first pixel and the second pixel is less than the preset offset threshold, μ 1 is the first pixel mean, μ 2 is the second pixel mean.

[0082] For example, the correlation is between -1 and 1, where -1 indicates a very low correlation and 1 indicates a very high correlation. At this time, the preset correlation threshold may be a value greater than -1 and less than 1, such as 0.

[0083] In one embodiment, according to the feature points matched by the NCC, the difference in the horizontal direction (the direction of the coordinate axis parallel to the ground) is calculated as the view difference, and a set of view differences between each matched first pixel point and second pixel point is obtained:

[0084] D={d|d=|xl–xr|} Formula (2),

[0085] Wherein, D is a set, d is a view difference between a set of matched first pixel points and second pixel points, xl is the horizontal coordinate value of the first pixel point, and xr is the horizontal coordinate value of the second pixel point.

[0086] In one embodiment, performing pixel matching on a first pixel of a first road image to be detected and a second pixel of a second road image to be detected includes:

[0087] The target object is identified for each first road image to be detected to obtain a first target object area, the target object is identified for each second road image to be detected to obtain a second target object area, and pixel matching is performed on the first pixel point and the second pixel point in the first target object area and the second target object area.

[0088] In this way, subsequent pixel matching and subsequent processing can be performed only on the area of ​​interest, which can effectively save computing resources.

[0089] In one embodiment, determining pixel depth information of a first pixel point based on the view difference includes:

[0090] Get the baseline and focal length of the depth image acquisition device;

[0091] Determine the view basis according to the view difference and the baseline;

[0092] Determine pixel depth information based on view basis and focal length.

[0093] For example, continuing to take the binocular camera as an example, the baseline Tx of the binocular camera, the camera focal length f, and the view difference d obtained in the above embodiment are obtained, and the pixel depth information can be calculated:

[0094] z=f*Tx / d Formula (3),

[0095] Among them, z is the pixel depth information, f is the focal length, Tx is the baseline, and d is the view difference.

[0096] It should be noted that the above takes the first pixel as an example. Since the first pixel matches the second pixel and the view difference between the two is the same, those skilled in the art may also use the pixel depth information of the second pixel.

[0097] In one embodiment, the pixel depth information of the first pixel point is determined based on the view difference to determine the image depth information of each first road image to be detected. The image depth information can be obtained by taking the average value of the pixel depth information as the image depth information, or by performing weighted averaging on the pixel depth information according to the weight of the area in which it is located to obtain the image depth information.

[0098] Step S203 , performing target detection on each first road image to be detected, and obtaining detection target frame position information of the target to be detected in each first road image to be detected.

[0099] The execution order of step S202 and step S203 is not limited here. When it is necessary to reduce the amount of calculation for pixel matching, it is recommended to execute step S203 first to obtain the first target object area and the second target object area mentioned in the above embodiment.

[0100] Step S203 may be implemented in a manner known to those skilled in the art, such as by performing target detection on each first road image to be detected using a pre-trained target detection model.

[0101] Step S204: Determine the predicted target frame position information through a Kalman filter model based on the position information of each detected target frame and the depth information of each image.

[0102] Through steps S201 to S203, the detection target frame position information and image depth information at each historical moment can be obtained. The trajectory point prediction can be realized through the preset Kalman filter model to obtain the predicted target frame position information.

[0103] For example, taking the detection target frame as a rectangular frame as an example, the detection target frame position information is [W, H, X, Y], where W is the width of the detection target frame, Y is the height of the detection target frame, the above W and H can be preset values, and X and Y are the coordinate values ​​of the upper left corner of the detection target frame (or other corners determined by those skilled in the art). Based on the position information of each detection target frame and the corresponding image depth information D, Kalman filtering is performed. The image depth information increases the dimension of motion information. The predicted target frame position information will be more in line with the actual motion. The parameters of the filter [W, H, X, Y, D] passed into the Kalman filter are output as [Wout, Hout, Xout, Yout], and the output parameters are used as the predicted target frame position information.

[0104] Step S205 , determining a motion vector of the target to be detected based on the target position of the target to be detected in each of the first road images to be detected.

[0105] The motion vector may represent the motion trajectory of the target object, or the motion vector may be calculated from the motion trajectory. The motion vector may be determined by connecting target positions in a plurality of first road images to be detected and straightening the connecting line to obtain the motion vector. The motion vector may also be determined by other methods known to those skilled in the art.

[0106] Step S206, determining the position information of the line segment to be detected of the line segment to be detected of the target to be detected based on the predicted target frame position information and the motion vector.

[0107] In one embodiment, before executing step S206, the predicted target frame position information in the three-dimensional coordinate system of step S204 needs to be mapped to a value in the two-dimensional coordinate system. The position information of the line segment to be detected is determined by the predicted target frame position information in the two-dimensional coordinate system and the motion vector in the two-dimensional coordinate system.

[0108] In one embodiment, the predicted target frame is a quadrilateral, and determining the position information of the line segment to be detected of the target to be detected based on the predicted target frame position information and the motion vector includes:

[0109] Connect the opposite sides of the quadrilateral along the moving direction of the target object to obtain a relative connection line;

[0110] Determine the two intersection points of the straight line formed by the motion vector passing through the connecting line auxiliary point and the quadrilateral, and determine the connecting line between the two intersection points as the line segment to be detected, and the connecting line auxiliary point is a point on the relative connecting line.

[0111] The relative connection line may be a connection line of the midpoints of the relative sides, or a connection line of reference points at a distance of 1 / x from the left vertex of each side, where x is greater than 1. The connection line auxiliary point may be a point on the relative connection line at a preset distance from one side in the ground direction, and the preset distance may be 1 / 4 of the length of the relative connection line or other distances set by those skilled in the art.

[0112] For example, see Figure 3 , Figure 3 is a schematic diagram of a line segment to be detected shown in an exemplary embodiment of the present application, such as Figure 3 As shown in the figure, the predicted target box M is a quadrilateral. Due to the presence of depth information, the predicted target box in the two-dimensional coordinate system may not be a rectangle. The four sides of the quadrilateral are a, b, c, and d. The target object's forward direction can be the extension direction of the road. When the target object is a vehicle, it can also be the direction of the line connecting the front and rear of the vehicle. Figure 3 , the relative sides are a and c, the relative connecting line is the OP connecting line, the connecting line auxiliary point Q, the motion vector α passes through the straight line formed by the connecting line auxiliary point Q and the two intersection points R, S, and RS of the quadrilateral. The line segment obtained by connecting the line is the line segment to be detected.

[0113] The line segment to be detected obtained in the above manner can represent the "central axis" of the target object (target to be detected) or the projection of the road surface where the entity of the target object is located.

[0114] Step S207, determining the line-crossing state of the target to be detected according to the position information of the line segment to be detected and the regular traffic line position information of the regular traffic line, so as to detect the illegal event of the target to be detected.

[0115] The position information of the line segment to be detected and the position information of the regular traffic line can be obtained by methods known to those skilled in the art and are not limited here.

[0116] In one embodiment, before determining the line-crossing state of the target to be detected according to the line segment position information to be detected and the regular traffic line position information of the regular traffic line, the illegal event detection method includes:

[0117] Obtaining the rotation state of the depth image acquisition device;

[0118] If the rotation state includes rotation, regular traffic line position information of regular traffic lines in the first road image to be detected is detected by using a preset image segmentation model.

[0119] If the rotation state includes not rotating, the historically preset regular traffic line position information may be used as the current regular traffic line position information.

[0120] The preset image segmentation model may be a model obtained by a road perception algorithm or the like.

[0121] In one embodiment, determining the line-crossing state of the target to be detected according to the line segment position information to be detected and the regular traffic line position information of the regular traffic line to detect the illegal event of the target to be detected includes:

[0122] Determine the intersection of the line segment to be detected and the regular traffic line according to the position information of the line segment to be detected and the regular traffic line position information of the regular traffic line;

[0123] If the line segment to be detected intersects with the regular traffic line, the line-crossing state is determined to be line-crossing, and the target to be detected has a violation event;

[0124] If the line segment to be detected is separated from the regular traffic line, the line-crossing state is determined to be not crossing the line, and there is no illegal event in the target to be detected.

[0125] Intersection means that the line segment to be detected has an intersection with the regular traffic line, and separation means that the line segment to be detected has no intersection with the regular traffic line. Figure 3 , it can be seen that the line segment to be detected intersects T with the regular traffic line N, which means that there is an illegal event (crossing the line) in the target to be detected.

[0126] It should be noted that, since the line segment to be detected of the target to be detected is determined based on the predicted target frame position information and motion vector, it can be understood that illegal event detection is actually a prediction of illegal events, predicting the possibility of an illegal event occurring in the target to be detected at the next moment.

[0127] The illegal event detection method provided in the present embodiment performs pixel matching on multiple groups of first road images to be detected and second road images to be detected to obtain view difference, and then determine image depth information, and perform target detection on one of the first road images to be detected and the second road images to be detected to obtain the target detection frame position, and then determine the predicted target frame position information, determine the position information of the line segment to be detected based on the motion vector of the target to be detected and the predicted target frame position information, determine the line-pressing state of the target to be detected through the line segment position information to be detected and the regular traffic line position information, so as to detect illegal events of the target to be detected, and predict the position of the target to be detected by introducing depth information, thereby providing a solution that does not rely on the processing of point cloud data sets, saves computing power, improves processing speed and efficiency, and can realize illegal event detection.

[0128] The following is a further exemplary description of the illegal event detection method provided in the above embodiment through a specific embodiment. Figure 4 , Figure 4 FIG. 1 is a flowchart of a specific illegal event detection method shown in an exemplary embodiment of the present application. Figure 4 As shown, in an exemplary embodiment, the specific illegal event detection method is described in detail as follows:

[0129] First, obtain traffic video. For example, you can obtain the traffic video stream collected by the traffic checkpoint.

[0130] Second, decode the traffic video to obtain frame images. The decoding method can be implemented in a manner known to those skilled in the art. Obtain multiple historical road images to be detected.

[0131] Third, determine whether the depth image acquisition device is rotating.

[0132] 3.1: There are two judgment methods. One is to directly obtain the motion parameters of the camera rotation, and the other is to use VIBE background modeling to compare the background differences frame by frame to determine whether the camera is rotating.

[0133] 3.2: If the camera is rotating, enter the road perception algorithm and return to the second part to reacquire frame images until the camera stops rotating. Since the rotation will cause a large error in the subsequent depth information, we will no longer detect illegal events during the rotation process.

[0134] 3.3: If the camera does not rotate, enter the target detection algorithm and disparity map calculation.

[0135] Fourth, road perception.

[0136] For example, the image segmentation model detects the outline of the lane line, and then further calculates and fits the exact position of the lane line and the effective area of ​​the road (effective area road line segment, that is, regular traffic line). Each time the road is perceived, the road data is updated, and this step is only performed when the camera is rotating.

[0137] Fifth, disparity map calculation:

[0138] 5.1: Epipolar correction: Use a binocular camera to ensure that the epipolar lines of the left and right images are in the horizontal direction to facilitate NCC calculation.

[0139] 5:2: Feature matching.

[0140] The NCC algorithm is used to calculate the correlation of pixels in the horizontal direction, such as formula (1).

[0141] 5.3: Disparity map calculation.

[0142] According to the feature points matched by NCC, the difference in the horizontal direction is calculated as the view difference. The set of view differences is D = {d|d = |xl-xr|}. In order to reduce computing resources, only the view difference within the valid area is calculated here.

[0143] 5.4: Calculate depth.

[0144] According to the binocular camera baseline Tx, camera focal length f, and view difference d, the depth information can be calculated: z = f*Tx / d.

[0145] Sixth, target detection.

[0146] The target detection model is used to identify targets such as people, motor vehicles, and non-motor vehicles in the frame image, and one of them is used as the target to be detected.

[0147] Seventh, target tracking.

[0148] For example, the target can be tracked through position correlation and time correlation.

[0149] Eighth, illegal incident detection.

[0150] The process of judging illegal crossing is as follows:

[0151] a. Obtain the depth information of road segments and images.

[0152] b. Get the detection box of the vehicle.

[0153] c. Based on the vehicle's detection box (W, H, X, Y) and the corresponding depth information D. Perform Kalman filtering on the 3D trajectory to make the trajectory more realistic and predict the position P at the next moment (predicted target box position information). Depth information increases the dimension of motion information, and the predicted target box will be more consistent with the actual motion. The parameters input to the Kalman filter are (W, H, X, Y, D), and the output parameters are (Wout, Hout, Xout, Yout).

[0154] d. Map the predicted position P to 2D, record the vehicle's motion trajectory, and calculate the target's motion vector Dir on the plane;

[0155] e. Estimate the bottom line of the vehicle (the line connecting the left and right tires, i.e. the line segment to be detected) based on the motion vector Dir and the target frame;

[0156] f. Determine whether the bottom line intersects with the solid lane line (regular traffic line). If so, it is considered illegal.

[0157] The solution of the related art is to reconstruct the traffic scene using point cloud data and linear reconstruction, which requires a large amount of calculation. Taking up too much computing resources will affect the performance and indicators of the algorithm. The method provided in this embodiment uses a method of parallax image depth estimation based on the NCC algorithm to construct a spatial scene, and only calculates the scene in the effective area of ​​the road, which greatly reduces the amount of calculation.

[0158] The method provided in this embodiment maps a two-dimensional target to a three-dimensional space (predicts the prediction of a target box in three dimensions) for event judgment, and calculates the three-dimensional trajectory and spatial distribution of the vehicle based on the depth information, which is closer to the real scene and improves the detection accuracy.

[0159] The method provided in this embodiment takes into account the change of monitoring angle caused by camera rotation, and the two-dimensional line segments of the traffic scene need to be re-acquired to automatically acquire the two-dimensional line segments of the traffic scene, such as lane lines and zebra crossings.

[0160] It can be seen that the method provided in this embodiment greatly reduces the amount of calculation for 3D reconstruction, saves CPU computing resources, can perform depth estimation at the pixel level, is fast, and the image obtained using the binocular camera does not require horizontal correction of the view baseline.

[0161] The method provided in this embodiment also has the following advantages:

[0162] Compared with the traditional traffic events which are completed in two-dimensional images and have large errors in judging the motion trajectory of the target vehicle due to lack of spatial information, thus causing misjudgment of the event, the method provided in this embodiment maps the target to three-dimensional space and adds Kalman filtering, which can obtain the target trajectory more accurately.

[0163] The method provided in this embodiment can obtain the real spatial distribution of vehicles through three-dimensional space mapping, which is helpful for judging road congestion.

[0164] See also Figure 5 , Figure 5 is a block diagram of an illegal event detection device shown in an exemplary embodiment of the present application, such as Figure 5 As shown, this embodiment provides an illegal event detection device 500, which includes:

[0165] An acquisition module 501 is used to acquire a plurality of historical road images to be detected at a plurality of historical moments of a preset traffic area by a depth image acquisition device, wherein the historical road images to be detected include a first road image to be detected and a second road image to be detected, and the preset traffic area includes a regular traffic line;

[0166] A matching module 502 is used to perform pixel matching on a first pixel of a first road image to be detected and a second pixel of a second road image to be detected, and determine pixel depth information of the first pixel based on a position difference in a preset direction between the matched first pixel and the second pixel as a view difference, so as to determine image depth information of each first road image to be detected;

[0167] A detection module 503 is used to perform target detection on each first road image to be detected, and obtain detection target frame position information of the target to be detected in each first road image to be detected;

[0168] A predicted target frame determination module 504 is used to determine the predicted target frame position information through a Kalman filter model based on the position information of each detected target frame and the depth information of each image;

[0169] A motion vector determination module 505, configured to determine a motion vector of a target to be detected based on a target position of the target to be detected in each of the first road images to be detected;

[0170] A to-be-detected line segment determination module 506 is used to determine the to-be-detected line segment position information of the to-be-detected line segment of the to-be-detected target based on the predicted target frame position information and the motion vector;

[0171] The illegal event detection module 507 is used to determine the line-crossing state of the target to be detected according to the position information of the line segment to be detected and the regular traffic line position information of the regular traffic line, so as to detect the illegal event of the target to be detected.

[0172] In this embodiment, the device is substantially provided with a plurality of modules for executing the method in any of the above embodiments. The specific functions and technical effects may refer to the above embodiments and will not be described in detail here.

[0173] See also Figure 6 , an embodiment of the present invention further provides an electronic device 600, including a processor 601, a memory 602 and a communication bus 603;

[0174] The communication bus 603 is used to connect the processor 601 and the memory 602;

[0175] The processor 601 is configured to execute a computer program stored in the memory 602 to implement one or more of the methods described in the above embodiments.

[0176] The embodiment of the present invention further provides a computer-readable storage medium, characterized in that a computer program is stored thereon.

[0177] The computer program is used to enable a computer to execute any of the methods described in the first embodiment.

[0178] The embodiment of the present application also provides a non-volatile readable storage medium, which stores one or more modules (programs). When the one or more modules are applied to a device, the device can execute instructions (instructions) of the steps included in the embodiment 1 of the embodiment of the present application.

[0179] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0180] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0181] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0182] The flow chart and block diagram in the accompanying drawings illustrate the possible implementation architecture, function and operation of the method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0183] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.

Claims

1. A method for detecting illegal events, It is characterized in that The method comprises: Acquire a plurality of historical road images to be detected at a plurality of historical moments of a preset traffic area by a depth image acquisition device, wherein the historical road images to be detected include a first road image to be detected and a second road image to be detected, and the preset traffic area includes a regular traffic line; performing pixel matching on a first pixel of the first road image to be detected and a second pixel of the second road image to be detected, using a position difference in a preset direction between the matched first pixel and the second pixel as a view difference, and determining pixel depth information of the first pixel based on the view difference, so as to determine image depth information of each of the first road images to be detected; Performing target detection on each of the first road images to be detected to obtain detection target frame position information of the target to be detected in each of the first road images to be detected; Determine the predicted target frame position information through a Kalman filter model based on the detection target frame position information and the image depth information; Determining a motion vector of the target to be detected based on a target position of the target to be detected in each of the first road images to be detected; Determine the position information of the line segment to be detected of the line segment to be detected of the target to be detected based on the predicted target frame position information and the motion vector; The line-crossing state of the target to be detected is determined according to the line segment position information to be detected and the regular traffic line position information of the regular traffic line, so as to detect the illegal event of the target to be detected.

2. The illegal event detection method according to claim 1, It is characterized in that Performing pixel matching on a first pixel of the first road image to be detected and a second pixel of the second road image to be detected includes: Obtaining a first pixel value, first pixel position information, and a first pixel mean value of a pixel window where the first pixel is located, as well as a second pixel value, second pixel position information, and a second pixel mean value of a pixel window where the second pixel is located, of the first pixel point; Determine an offset between the first pixel point and the second pixel point according to the first pixel position information and the second pixel position information; If the offset is less than a preset offset threshold, determining a first difference value according to the first pixel value and the first pixel mean, and determining a second difference value according to the second pixel value and the second pixel mean; Determine a correlation between the first pixel and the second pixel based on the first difference and the second difference; If the correlation is greater than a preset correlation threshold, the first pixel point matches the second pixel point; If the correlation is less than or equal to the preset correlation threshold, the first pixel point does not match the second pixel point.

3. The illegal event detection method as claimed in claim 2, It is characterized in that The determination method of the relevance includes: Among them, ncc(I 1 I 2 ) is the correlation, I 1 (x) is the first pixel value of the first pixel, I 2 (x) is the second pixel value of the second pixel, the offset between the first pixel and the second pixel is less than the preset offset threshold, μ 1 is the first pixel mean, μ 2 is the second pixel mean.

4. The illegal event detection method according to claim 1, It is characterized in that Determining pixel depth information of the first pixel point based on the view difference includes: Obtaining a baseline and a focal length of the depth image acquisition device; determining a view basis according to the view difference and the baseline; The pixel depth information is determined based on the view basis and the focal length.

5. The illegal event detection method according to any one of claims 1 to 4, It is characterized in that Before determining the line-crossing state of the target to be detected according to the line segment position information to be detected and the regular traffic line position information of the regular traffic line, the illegal event detection method includes: Acquiring a rotation state of the depth image acquisition device; If the rotation state includes rotation, the regular traffic line position information of the regular traffic line in the first road image to be detected is detected by a preset image segmentation model.

6. The illegal event detection method according to any one of claims 1 to 4, It is characterized in that The predicted target frame is a quadrilateral, and determining the position information of the line segment to be detected of the target to be detected based on the predicted target frame position information and the motion vector includes: Connecting the opposite sides of the quadrilateral along the moving direction of the target object to obtain a relative connection line; Determine two intersection points of the straight line formed by the motion vector passing through the line auxiliary point and the quadrilateral, and determine the line between the two intersection points as the line segment to be detected, and the line auxiliary point is a point on the relative line.

7. The illegal event detection method according to any one of claims 1 to 4, It is characterized in that Determining the line-crossing state of the target to be detected according to the line segment position information to be detected and the regular traffic line position information of the regular traffic line to detect the illegal event of the target to be detected includes: Determine the intersection of the line segment to be detected and the regular traffic line according to the position information of the line segment to be detected and the regular traffic line position information of the regular traffic line; If the line segment to be detected intersects with the regular traffic line, the line-crossing state is determined to be line-crossing, and the target to be detected has an illegal event; If the line segment to be detected is separated from the regular traffic line, the line-crossing state is determined to be not crossing the line, and there is no illegal event for the target to be detected.

8. A device for detecting illegal events, It is characterized in that The device comprises: An acquisition module, used to acquire a plurality of historical road images to be detected at a plurality of historical moments of a preset traffic area acquired by a depth image acquisition device, wherein the historical road images to be detected include a first road image to be detected and a second road image to be detected, and the preset traffic area includes a regular traffic line; a matching module, configured to perform pixel matching on a first pixel of the first road image to be detected and a second pixel of the second road image to be detected, and determine pixel depth information of the first pixel based on a position difference in a preset direction between the matched first pixel and the second pixel as a view difference, so as to determine image depth information of each of the first road images to be detected; A detection module, used to perform target detection on each of the first road images to be detected, and obtain detection target frame position information of the target to be detected in each of the first road images to be detected; A predicted target frame determination module, used to determine the predicted target frame position information through a Kalman filter model based on the position information of each detected target frame and the image depth information; A motion vector determination module, used to determine the motion vector of the target to be detected based on the target position of the target to be detected in each of the first road images to be detected; A to-be-detected line segment determination module, used to determine the to-be-detected line segment position information of the to-be-detected line segment of the to-be-detected target based on the predicted target frame position information and the motion vector; The illegal event detection module is used to determine the line-crossing state of the target to be detected according to the position information of the line segment to be detected and the regular traffic line position information of the regular traffic line, so as to detect the illegal event of the target to be detected.

9. An electronic device, It is characterized in that including a processor, a memory and a communication bus; The communication bus is used to connect the processor and the memory; The processor is configured to execute the computer program stored in the memory to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, It is characterized in that A computer program is stored thereon, The computer program is used to make the computer execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Target vehicle line pressing detection method based on vehicle-mounted image

    CN110210363A

  • Road occupancy information determination method and apparatus

    WO2022142827A1