A monitoring method and terminal for road obstruction events
By detecting and converting blocked areas in the monitoring image into the world coordinate system, the problem of the inability to accurately assess the impact of road obstruction events in existing technologies is solved, precise positioning and impact assessment are achieved, and the practicality of the method is improved.
Patent Information
- Application Number
- CN202310696598.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-13
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2043-06-13
AI Technical Summary
Existing technologies make it difficult to accurately assess the impact of road obstruction events on traffic and require complex camera parameter settings.
By acquiring surveillance images, event target detection is performed, and blocked areas are extracted using image algorithms. These areas are converted into the world coordinate system, and the traffic impact is evaluated in combination with the road environment.
It achieves accurate positioning and impact assessment of road obstruction events, avoids the complex operation of camera parameter setting, and improves the coverage and practicality of the method.
Smart Images

Figure CN116912784B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent transportation technology, and in particular to a monitoring method and terminal for road obstruction events. Background Art
[0002] The smoothness of urban traffic is mainly affected by vehicle flow, pedestrian flow, and road obstruction events. Among them, vehicle flow and pedestrian flow have fixed changing patterns. Currently, various methods are available to monitor the real-time situation of pedestrian and vehicle flows. For monitoring road obstruction events, such as the accumulation of debris, garbage and slag, road flooding, fallen trees, street vendors, road occupation, construction roadblocks, and even drying grain on the road, the most common method is to use event targets to train a sample-based neural network model, and then perform target detection or image segmentation on the monitoring image. However, the results of these methods can only simply determine whether the event has occurred, and do not include information such as the severity of road obstruction and the precise location of the event. "CN115984772A" further analyzes road flooding events, uses internal and external parameters of the camera to perform an inverse perspective transformation on the image, and then estimates the area, and then analyzes the degree of road impact. This method still requires detailed settings of camera parameters and does not solve the problem of precise location positioning. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a monitoring method and terminal for road obstruction events, realize the evaluation of the impact of obstruction events on traffic obstacles, avoid complex operations such as camera parameter setting, and use algorithms to automatically complete identification and accurate positioning.
[0004] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0005] A method for monitoring road obstruction events, comprising the steps of:
[0006] S1. Obtain monitoring images, detect event targets and establish event areas;
[0007] S2. Using an image algorithm to extract a target object from the event area and generate a blocked area;
[0008] S3, converting the blocked area into a world coordinate system;
[0009] S4. Evaluate the impact of the blocked area on road traffic in combination with the road environment in the world coordinate system.
[0010] A monitoring terminal for road obstruction events includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, any step in a method for monitoring road obstruction events is implemented.
[0011] The beneficial effects of the present invention are: providing a monitoring method and terminal for road obstruction events, which are compatible with target detection and image segmentation calculations for various types of road obstruction events, avoiding complex operations such as camera parameter settings, and constructing a mapping relationship between the monitoring screen and the real-world plane based on the image algorithm, realizing the evaluation of the impact of obstruction events on traffic obstacles, and improving the coverage and practicality of this method in application. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a flow chart of a method for monitoring road obstruction events according to an embodiment of the present invention;
[0013] Figure 2 A schematic diagram of yaw angle calculation in a method for monitoring a road obstruction event in one embodiment of the present invention;
[0014] Figure 3 A schematic diagram of a blocked area in a method for monitoring road obstruction events in one embodiment of the present invention;
[0015] Figure 4 A schematic diagram of pedestrian target detection in a method for monitoring road obstruction events in one embodiment of the present invention;
[0016] Figure 5 A schematic diagram of vehicle target detection in a method for monitoring road obstruction events in one embodiment of the present invention;
[0017] Figure 6 Schematic diagram of a monitoring terminal for a road obstruction event in one embodiment of the present invention;
[0018] Description of labels:
[0019] 1. A monitoring terminal for road obstruction events; 2. A processor; 3. A storage device. DETAILED DESCRIPTION
[0020] To illustrate the technical content, achieved objectives and effects of the present invention in detail, the following description is given in conjunction with the embodiments and accompanying drawings.
[0021] Please refer to Figures 1 to 6 , a method for monitoring a road obstruction event, comprising the steps of:
[0022] S1. Obtain monitoring images, detect event targets and establish event areas;
[0023] S2. Using an image algorithm to extract a target object from the event area and generate a blocked area;
[0024] S3, converting the blocked area into a world coordinate system;
[0025] S4. Evaluate the impact of the blocked area on road traffic in combination with the road environment in the world coordinate system.
[0026] From the above description, it can be seen that the beneficial effects of the present invention are: it is compatible with target detection and image segmentation calculations of various types of road obstruction events, avoids complex operations such as camera parameter settings, and can construct a mapping relationship between the monitoring screen and the real-world plane based on the image algorithm, thereby realizing the evaluation of the impact of obstruction events on traffic obstacles, and improving the coverage and practicality of this method in application.
[0027] Furthermore, step S30 is further included between step S2 and step S3:
[0028] S301, using a neural network to establish corresponding rectangular frames for pedestrians and vehicles in the video stream and obtain key point coordinates;
[0029] S302, generating a corresponding cylindrical reference object according to the key point coordinates;
[0030] S303, performing camera self-calibration calculation using the columnar reference object;
[0031] S304, constructing an inverse perspective transformation relationship matrix from the video stream to the world coordinate system;
[0032] S305: Calculate the vehicle yaw angle, where the vehicle yaw angle is the angle between the camera direction and the north direction.
[0033] As can be seen from the above description, in this step, pedestrians and vehicles in the monitoring image are detected, and the camera is self-calibrated at the typical heights of people and vehicles to obtain the camera's intrinsic and extrinsic parameters. The camera's intrinsic and extrinsic parameters are then used to establish the inverse perspective transformation plane from the monitoring image to the world coordinate system. The camera's yaw angle relative to true north is then calculated using the known camera latitude and longitude information and road network information. This establishes the transformation relationship between each pixel of the road surface in the monitoring image and the coordinates in the real world.
[0034] Furthermore, the step S304 is specifically as follows:
[0035] S3041. Obtain a vehicle trajectory point sequence using a target tracking algorithm;
[0036] S3042: Obtain camera data of any frame in the video stream to establish an inverse perspective transformation relationship matrix, and obtain the actual length represented by a pixel in the inverse perspective transformation relationship matrix;
[0037] S3043, according to the inverse perspective transformation relationship matrix, calculate and obtain: the coordinates p0' of the midpoint in the original image on the inverse perspective transformation image, the coordinates v0' of the vertical vanishing point in the original image on the inverse perspective transformation image, and the coordinates {(x i ',y i ')}.
[0038] As can be seen from the above description, by tracking the vehicle in the picture, the vehicle trajectory obtained is also mapped to the inverse perspective transformation plane, thereby constructing lane information with the vehicle direction.
[0039] Furthermore, the step S305 is specifically as follows:
[0040] S3051. Establish a calculation model based on the latitude and longitude coordinates of the monitoring location where the video stream is generated and the nearby road network information, where the monitoring location is the origin in the calculation model;
[0041] S3052: performing an inverse perspective transformation on the road network information according to a first preset step size;
[0042] S3053, obtaining the lane direction of the current vehicle according to the coordinates of the vehicle trajectory point sequence on the inverse perspective transformed image and the inverse perspective transformed road network information;
[0043] S3054: Count the absolute values of the inner products of all vehicle trajectories and the corresponding lane directions and screen the maximum absolute value of the inner product. If multiple maximum inner product values are found, manually select the only one and proceed to the next step. Otherwise, proceed directly to the next step.
[0044] S3055. Calculate the yaw angle according to the maximum absolute value of the inner product;
[0045] S3056: Within the preset calibration angle range of the yaw angle, regard the second preset step length as the first preset step length, and return to step S3052 until the yaw angle meets the preset accuracy.
[0046] From the above description, we can see that the yaw angle cannot be obtained by camera self-calibration only from the image, so an algorithm is used to calculate the yaw angle. The specific example is as follows:
[0047] Please refer to Figure 2 Taking a three-way intersection as an example, the image is simplified based on the longitude and latitude coordinates of the monitoring location and the nearby road network information as shown in the figure. The monitoring location O is taken as the image origin, and the inverse perspective transformation matrix is used to obtain the length dL represented by a single pixel, and the image is facing due north;
[0048] Take 1° as the first preset step, take point O as the center, and rotate counterclockwise Figure 2 Calculate the position of the road represented by the three line segments AD, BD, and CD on the inverse perspective transformed image. That is, the four points A, B, C, and D are first rotated by a known angle, and then the v0 coordinate is offset to obtain the coordinates of the four points A', B', C', and D' of the road segment points on the inverse perspective transformed image.
[0049] Traverse the coordinates of the vehicle trajectory on the inverse perspective transformation in order {(x i ',y i ')}, backward search for two trajectory points whose actual distance is greater than the preset distance of 5 meters. The two points form a vector, and the absolute value of the inner product with the three line segments A'D', B'D', and C'D' is calculated. The maximum value is the direction of the corresponding lane, and the inner products of the other two lanes are discarded;
[0050] Count the absolute values of the inner products of all vehicle trajectories and the corresponding road directions. The corresponding rotation angle at the maximum value is the yaw angle ρ of the camera relative to the north direction.
[0051] Using ±1° as the preset calibration angle range and 1 / 10 of the first preset step size as the second preset step size, continue iterating the above steps within the ±1° range of the yaw angle ρ with a step size of 1 / 10 until the yaw angle reaches the preset accuracy;
[0052] Specifically, when the camera is in the middle of the road or the center of the intersection, this step will find more than two possible yaw angle values, which are eliminated by manually adding the approximate orientation of the preset position;
[0053] Calculate the actual position of the point p0' in the original image in the inverse perspective transformed image. Its actual distance in the ρ direction of the camera coordinate is d(p0',v0')×dL.
[0054] Furthermore, the step S2 is specifically as follows:
[0055] S21. Perform image semantic segmentation on the image using an image algorithm and obtain the category of each element;
[0056] S22, filtering out street scene categories;
[0057] S23: Eliminate elements belonging to the street scene category in the event area, and obtain a blocked area.
[0058] As can be seen from the above description, after establishing the event region, a frame of the event region is captured and semantically segmented to obtain the category of each pixel in the image. These categories are street scene categories, such as road, building, pedestrian, vehicle, sky, and tree. Specifically, using the Mapillary street scene dataset as an example, DeepLab V3 is trained and the output is a mask image, where pixel values 0 represent non-target events and scenes represent the categories of target events. Only street scene categories that are not easily confused with obstruction events and have distinct characteristics are selected, including pedestrians, road surface (combining various road surface categories, such as sidewalks and lanes), sky, and vegetation.
[0059] Furthermore, the step S23 is specifically as follows:
[0060] S231, connecting each element belonging to the street scene category into its corresponding regional pixels;
[0061] S232: If a certain region pixel is located within the event region, the region pixel in the event region is set to 0; if the intersection area of a certain region pixel and the event region is less than half of the event region, the region pixel in the event region is set to 0; if the region pixel is a pedestrian, the region pixel in the event region is set to 0;
[0062] S233: Define an area within the event area where the pixels are not 0 as the blocking area.
[0063] From the above description, we can know that this step removes non-target elements in the event area, that is, removes street scene elements, and generates a blocked area. Specifically, the semantic mask image of the image is connected by category to obtain multiple regions S scenes Each region contains only one street scene category. According to the judgment conditions, when the region pixel is completely located inside the event area, or the intersection area of the region pixel and the event area is less than half of the event area, or the region pixel is a pedestrian category in the street scene category, the region pixel in the event area is set to 0, and the area where the event area pixel is not 0 is extracted and defined as the blocking area.
[0064] Furthermore, the step S3 is specifically as follows:
[0065] S3. Perform an inverse perspective transformation on the blocked area to obtain the blocked area in a world coordinate system.
[0066] From the above description, it can be seen that the blocked area is transformed into the world coordinate system through the inverse perspective transformation.
[0067] Furthermore, the step S4 is specifically as follows:
[0068] S41. Generate a road area in a world coordinate system. The road area includes the blocked area and the street view area where the blocked area is located.
[0069] S42. Obtaining a lane direction of the road area according to the road network information rotated by the yaw angle;
[0070] S43, calculating the actual width and length of the blocked area and the actual width of the road area in the lane direction;
[0071] S44. Based on the traffic flow in the road area, a road congestion warning of a corresponding level is given.
[0072] As can be seen from the above description, in this step, the converted blocking event area is combined with the street view area to determine its impact on road traffic and complete the location calculation to facilitate the virtual effect display. An example is as follows:
[0073] Reference Figure 3 , on the inverse perspective transformation plane, overlap the blocked area S on the street view area S scenes Here, the people, vehicles, and road surfaces in the street scene category are combined to form the road area;
[0074] Using the road network information rotated by the yaw angle ρ, the lane direction of the blocking event and the surrounding road surface is obtained; here, the inverse perspective transformation image is transferred to the lane direction of the blocking event;
[0075] By finding the maximum and minimum values x and y of the blocked area, the actual width w×dL and actual length h×dL of the blocked area are calculated, and then the average width w of the blocked road area in the lane direction is calculated. road The above vehicle trajectories can be used to conveniently estimate the traffic flow at the location. By combining the type of congestion event, road section width, and congestion event width and length, different warning levels can be flexibly designed.
[0076] Furthermore, the step S4 further includes a step S5:
[0077] S5. Calculating the actual position of the lane center in the blocked area;
[0078] The calculation process is as follows:
[0079] The distance to the camera is calculated as: d(r0′, v0′)*dL;
[0080] The angle is calculated as: ρ′-dρ′;
[0081] in,
[0082] ρ′ is the value of the yaw angle in the original image on the inverse perspective transformed image;
[0083] r0′ is the coordinate of the center of the blocked lane in the original image on the inverse perspective transformed image.
[0084] d(r0',v0') is the pixel distance from the center of the blocked lane r0' to the vertical vanishing point v0' on the inverse perspective transformed image.
[0085] As can be seen from the above description, this step calculates the actual position of the center of the blocked lane. After the calculation is completed, the rendering can be pasted to the corresponding position of other display systems to generate a visual effect of the lane width and the blocking event situation.
[0086] A monitoring terminal for road obstruction events includes a memory 3, a processor 2, and a computer program stored in the memory 3 and executable on the processor 2. When the processor 2 executes the computer program, any step in the above-mentioned method for monitoring road obstruction events is implemented.
[0087] The present invention provides a method and terminal for monitoring road obstruction events, which are applied to obstacle detection programs in intelligent transportation. The following describes the method and terminal with reference to an embodiment:
[0088] Please refer to Figures 1 to 6 , Embodiment 1 of the present invention is: a method for monitoring a road obstruction event, comprising the steps of:
[0089] S1. Obtain monitoring images, detect event targets and establish event areas;
[0090] S2. Using an image algorithm to extract a target object from the event area and generate a blocked area;
[0091] S3, converting the blocked area into a world coordinate system;
[0092] S4. Evaluate the impact of the blocked area on road traffic in combination with the road environment in the world coordinate system.
[0093] That is, in this embodiment, the beneficial effects of the present invention are: it is compatible with target detection and image segmentation calculations of various road obstruction events, avoids complex operations such as camera parameter settings, and can construct a mapping relationship between the monitoring screen and the real-world plane based on the image algorithm, thereby realizing the evaluation of the impact of obstruction events on traffic obstacles, and improving the coverage and practicality of this method in application.
[0094] Specifically, in this embodiment, step S1 is as follows:
[0095] The images are monitored for road congestion events, and samples of common congestion events are collected for neural network training. Here, examples are taken of debris piles, gravel and slag piles, road waterlogging, fallen trees, street vendors, road occupation, construction roadblocks, and drying grain.
[0096] For non-fixed-shape, flat-surface blocking events on the road, such as stagnant water or drying grain, the DeepLab V3 image segmentation network is used to output a mask image of the target event, where pixel value 0 represents a non-target event, and c represents the target event category. Connected pixels of the same category form an event region S. These events will interfere with the passage of vehicles and pedestrians, but will not completely block traffic.
[0097] Other blocking events have morphological features and have a certain height. The yolov5 target detection network is used for identification, and the output is the coordinates of the center point of the target event rectangle, width and height information, and category label (x c ,y c ,w,h,c). Convert the rectangular box information into a mask image, where the pixel value inside the rectangular box is category c and the pixel value outside the rectangular box is 0;
[0098] Please refer to Figures 1 to 6 , the second embodiment of the present invention is: on the basis of the first embodiment,
[0099] The step S30 is further included between the step S2 and the step S3:
[0100] S301, using a neural network to establish corresponding rectangular frames for pedestrians and vehicles in the video stream and obtain key point coordinates;
[0101] S302, generating a corresponding cylindrical reference object according to the key point coordinates;
[0102] S303, performing camera self-calibration calculation using the columnar reference object;
[0103] S304, constructing an inverse perspective transformation relationship matrix from the video stream to the world coordinate system;
[0104] S305: Calculate the vehicle yaw angle, where the vehicle yaw angle is the angle between the camera direction and the north direction.
[0105] The step S304 is specifically as follows:
[0106] S3041. Obtain a vehicle trajectory point sequence using a target tracking algorithm;
[0107] S3042: Obtain camera data of any frame in the video stream to establish an inverse perspective transformation relationship matrix, and obtain the actual length represented by a pixel in the inverse perspective transformation relationship matrix;
[0108] S3043, according to the inverse perspective transformation relationship matrix, calculate and obtain: the coordinates p0' of the midpoint in the original image on the inverse perspective transformation image, the coordinates v0' of the vertical vanishing point in the original image on the inverse perspective transformation image, and the coordinates {(x i ',y i ')}.
[0109] The step S305 is specifically as follows:
[0110] S3051. Establish a calculation model based on the latitude and longitude coordinates of the monitoring location where the video stream is generated and the nearby road network information, where the monitoring location is the origin in the calculation model;
[0111] S3052: performing an inverse perspective transformation on the road network information according to a first preset step size;
[0112] S3053, obtaining the lane direction of the current vehicle according to the coordinates of the vehicle trajectory point sequence on the inverse perspective transformed image and the inverse perspective transformed road network information;
[0113] S3054: Count the absolute values of the inner products of all vehicle trajectories and the corresponding lane directions and screen the maximum absolute value of the inner product. If multiple maximum inner product values are found, manually select the only one and proceed to the next step. Otherwise, proceed directly to the next step.
[0114] S3055. Calculate the yaw angle according to the maximum absolute value of the inner product;
[0115] S3056: Within the preset calibration angle range of the yaw angle, regard the second preset step length as the first preset step length, and return to step S3052 until the yaw angle meets the preset accuracy.
[0116] That is, in this embodiment, to complete the construction of the traffic model, pedestrians and vehicles in the monitoring image are first detected, and the camera is self-calibrated at the typical heights of people and vehicles to obtain the camera's intrinsic and extrinsic parameters. The camera's intrinsic and extrinsic parameters are then used to establish the inverse perspective transformation plane from the monitoring image to the world coordinate system. The camera's yaw angle relative to true north is then calculated using the known camera latitude and longitude information and road network information. This allows the transformation relationship between each pixel on the road surface in the monitoring image and the real-world coordinates to be established. Specific examples are as follows:
[0117] Step S301:
[0118] Perform the reset operation and access the video stream;
[0119] like Figure 4As shown in the figure, the yolo-pose neural network is used to detect pedestrian key points in each frame of the video stream, and the coordinates of the center point, width and height, and category information of the pedestrian rectangle are obtained (x c ,y c ,w,h,c), key point coordinates (K x , K y ) i , where i represents the label of the pedestrian, and there are 17 key points, namely 1 nose, 2 left eye, 3 right eye, 4 left ear, 5 right ear, 6 left shoulder, 7 right shoulder, 8 left elbow, 9 right elbow, 10 left wrist, 11 right wrist, 12 left hip, 13 right hip, 14 left knee, 15 right knee, 16 left ankle, 17 right ankle;
[0120] like Figure 5 As shown, the monocular 3D target detection network Deep3Dbox is used to perform 3D bounding box detection (3D boundingbox) on the vehicle in the video stream, and the coordinates of the 8 corner points of the 3D rectangular box of the vehicle (K x , K y ) i , where i represents the vehicle label. By finding the minimum x, y value and the maximum x, y value of these 8 points, the coordinates of the center point, width, height and category information of the vehicle's 2D rectangular box (x c ,y c ,w,h,c).
[0121] Step S302:
[0122] Pedestrians and vehicles are considered as cylindrical reference objects perpendicular to the ground plane. Their position and length in the image are used for subsequent camera self-calibration calculations. Since it is necessary to establish a correlation with the actual height of people and the average height of vehicles, this step requires determining whether pedestrians are in an upright position. For vehicles, the longest of the four vertical lines in the 3D rectangle is used. The specific operation is as follows:
[0123] For pedestrians, we need to calculate the length of the torso. First, we find the midpoint of key points 5 and 6, and the midpoint of 11 and 12. These two points form a torso line segment with a length of l. body . Determine whether the angles between the line segments (12, 14), (14, 16), (11, 13), and (13, 15) and the trunk line segment are all less than 15 degrees, and whether the angle between the trunk and the vertical direction is less than 45 degrees. If the conditions are not met, it is considered as a non-upright state and no subsequent calculations are performed. For those that meet the upright conditions, the intersection of the line connecting the key points 15 and 16 and the straight line where the trunk line segment is located is regarded as the bottom coordinate of the columnar reference object (x bottom ,y bottom ). The average leg length is obtained by calculating the length of the two legs from the key points of the legs. leg , head length l headIt is 1.5 times the length from the nose to the torso. Finally, from the bottom along the torso line segment upwards, find the distance from the bottom l leg +l body +l head The point is the top coordinate of the cylindrical reference object (x top ,y top ), category c is pedestrian.
[0124] For the vehicle, calculate the lengths of the line segments (0, 4), (1, 5), (2, 6), and (3, 7) formed by the corner point pairs, find the longest one, and record it as the bottom coordinate (x bottom ,y bottom ) and the top coordinate (x top ,y top ), category c is vehicle.
[0125] Step 303 is specifically as follows:
[0126] First, we can use the cylindrical reference to find the vanishing point v0 on the image perpendicular to the ground plane. Assuming the focal length f is known, we can use this together with v0 to calculate the equation of the horizontal line.
[0127]
[0128] From the horizontal line and cross ratio invariance, we can get the ratio of the real height of the person and the camera height to satisfy the following formula, where the right formula is the distance between the midpoints of the image.
[0129]
[0130] Then, using the hypothesis test method, assume a focal length f and calculate the actual height of each pedestrian in the picture according to the above two formulas. The distribution of these heights must meet the average height and standard deviation in population information statistics. By traversing the possible values of the focal length f with a certain step size and calculating the height distribution with the largest likelihood function, the optimal value of the focal length f can be obtained, and thus the camera height h can be obtained. c .
[0131] Here, the mean and standard deviation of height need to be corrected. The mean height is set to 164 cm and the standard deviation is 6.7 cm. Furthermore, for vehicles, the mean height is set to 150 cm and the standard deviation is 10 cm. Finally, based on the relationship between the horizon and the camera angle, the pitch angle θ can be directly calculated. In video surveillance scenarios, the default roll angle ρ is 0 and the aspect ratio a is 1.
[0132]
[0133] Step S3041 is specifically as follows:
[0134] Use the deepsort method to track the vehicle. Specifically, select the rectangular box with the vehicle category in step a, intercept the image imgroi in the box, and use the feature extractor of the convolutional neural network structure to reduce the dimension of the imgroi image to obtain a 128-dimensional simplified feature for the similarity comparison between the two roi images. Associate the rectangular boxes between frames according to position, speed and similarity to obtain the vehicle trajectory point sequence {(x i ,y i )}, the sequence represents the trajectory of a vehicle, (x i ,y i ) is the midpoint of the rectangular frame, and i represents the sequence number of the trajectory point.
[0135] Step 3042 is specifically as follows:
[0136] First, establish the inverse perspective transformation matrix relationship:
[0137] Acquire a frame of a road monitoring screen video stream, and obtain the height, downward pitch angle, and equivalent focal length of the camera's monitoring installation; calculate the longitudinal half-field-of-view angle and the lateral half-field-of-view angle of the camera based on the equivalent focal length of the camera; calculate the relationship between the road monitoring screen image and the world coordinate system based on the width of the road monitoring screen image, the height of the road monitoring screen image, the longitudinal half-field-of-view angle, the lateral half-field-of-view angle, the height, and downward pitch angle of the camera's monitoring installation; select four points on the road monitoring screen image, and establish an inverse perspective transformation relationship matrix based on the selected four points; the selecting of the four points on the road monitoring screen image specifically transforms the road monitoring screen image through the relationship between the road monitoring screen image and the world coordinate system to obtain an inverse perspective transformation image;
[0138] Select two points on both sides of the bottom of the inverse perspective transformation image, calculate the width of the bottom of the inverse perspective transformation image, if the inverse perspective transformation image has a row with a width of 2-4 of the width of the bottom of the inverse perspective transformation image, then select the two points at both ends of the row, if not, then select the two points in the first row of the inverse perspective transformation image.
[0139] The relationship between the road monitoring screen image and the world coordinate system is calculated according to the following formula:
[0140]
[0141]
[0142] Where x and y are the image coordinates of the road monitoring screen, X P and Y P is the world coordinate system coordinate, h is the height of the camera monitoring installation, θ is the downward pitch angle, α is the longitudinal half field of view angle of the camera under the equivalent focal length f, and β is the horizontal half field of view angle of the camera under the equivalent focal length f.
[0143] Step S3043 is specifically as follows:
[0144] The actual length dL of a pixel in the world coordinate system is obtained in the inverse perspective transformation relationship;
[0145] Calculate the coordinates p0' of the original image's midpoint on the inverse perspective transformation, the coordinates v0' of the vertical vanishing point (that is, the position of the ground plane where the camera rod is located) on the inverse perspective transformation, and the coordinates {(x i ',y i ')}.
[0146] Step S305 is specifically as follows:
[0147] Since the yaw angle cannot be obtained by camera self-calibration from images alone, this step mainly relies on the calculated traffic direction and road information to determine the yaw angle ρ, which is the angle between the camera heading and the north direction.
[0148] like Figure 2 As shown in the figure, taking a three-way intersection as an example, the latitude and longitude coordinates of the monitoring location and the nearby road network information are simplified as shown in the figure. The monitoring location O is taken as the image origin, the length represented by a pixel is dL, and the image is facing due north.
[0149] Take 1° as the first preset step, take point O as the center, and rotate counterclockwise Figure 2 Calculate the position of the road represented by the three line segments AD, BD, and CD on the inverse perspective transformed image. That is, the four points A, B, C, and D are first rotated by a known angle, and then the v0 coordinate is offset to obtain the coordinates of the four points A', B', C', and D' of the road segment points on the inverse perspective transformed image.
[0150] Traverse the coordinates of the vehicle trajectory on the inverse perspective transformation in order {(x i ',y i ')}, backward search for two trajectory points whose actual distance is greater than the preset distance of 5 meters. The two points form a vector, and the absolute value of the inner product with the three line segments A'D', B'D', and C'D' is calculated. The maximum value is the direction of the corresponding lane, and the inner products of the other two lanes are discarded;
[0151] Count the absolute values of the inner products of all vehicle trajectories and the corresponding road directions. The corresponding rotation angle at the maximum value is the camera's yaw angle ρ relative to the north direction.
[0152] Using ±1° as the preset calibration angle range and 1 / 10 of the first preset step size as the second preset step size, continue iterating the above steps within the ±1° range of the yaw angle ρ with a step size of 1 / 10 until the yaw angle reaches the preset accuracy;
[0153] Specifically, when the camera is in the middle of the road or the center of the intersection, this step will find more than two possible yaw angle values, which are eliminated by manually adding the approximate orientation of the preset position;
[0154] Calculate the actual position of the point p0 in the original image in the inverse perspective transformed image. Its actual distance in the ρ direction of the camera coordinate is d(p0',v0')×dL.
[0155] Please refer to Figures 1 to 6 , the third embodiment of the present invention is: on the basis of the second embodiment,
[0156] The step S2 is specifically as follows:
[0157] S21. Perform image semantic segmentation on the image using an image algorithm and obtain the category of each element;
[0158] S22, filtering out street scene categories;
[0159] S23: Eliminate elements belonging to the street scene category in the event area, and obtain a blocked area.
[0160] In this embodiment, after establishing an event region, a frame of the event region is captured and semantically segmented to obtain the category of each pixel in the frame. These categories are street scene categories, such as road, building, pedestrian, vehicle, sky, and tree. Specifically, using the Mapillary street scene dataset as an example, DeepLab V3 is trained and the output is a mask image, where pixel values 0 represent non-target events and scenes represent the categories of target events. Only street scene categories that are not easily confused with obstruction events and have distinct characteristics are selected, including pedestrians, road surface (combining various road surface categories, such as sidewalks and lanes), sky, and vegetation.
[0161] The step S23 is specifically as follows:
[0162] S231, connecting each element belonging to the street scene category into its corresponding regional pixels;
[0163] S232: If a certain region pixel is located within the event region, the region pixel in the event region is set to 0; if the intersection area of a certain region pixel and the event region is less than half of the event region, the region pixel in the event region is set to 0; if the region pixel is a pedestrian, the region pixel in the event region is set to 0;
[0164] S233: Define an area within the event area where the pixels are not 0 as the blocking area.
[0165] Specifically, the following are some examples:
[0166] Use the mask image of image semantics to eliminate the pixels belonging to the street scene category inside the target detection rectangle.
[0167] The semantic mask image of the image is connected to the pixels in the region by category to obtain multiple regions Sscenes, each of which contains only one street scene category.
[0168] Traverse area S scenes If the category is not a pedestrian, calculate the intersection area between it and the rectangular frame. If the intersection area is less than half of the area Si, set the pixel value of the intersection part in the rectangular frame to zero; if the category is a pedestrian, directly set the pixel value of the intersection part to 0.
[0169] The area where the remaining pixels are not 0, that is, the area with the largest area in the rectangular box, is taken as the area S of the blocking event, and the blocking area with no fixed shape and the blocking area with a fixed shape are merged to obtain the mask image of the blocking event.
[0170] The step S3 is specifically as follows:
[0171] S3. Perform an inverse perspective transformation on the blocked area to obtain the blocked area in a world coordinate system.
[0172] From the above description, it can be seen that the blocked area is transformed into the world coordinate system through the inverse perspective transformation.
[0173] Although some blocking events have a certain height, directly converting them will occupy a larger area than the actual situation. However, in this scenario, the monitoring camera is at a high position, and the resulting error is within an acceptable range, which has little impact on the analysis and display of the recognition results. Similarly, the street view mask image is subjected to inverse perspective transformation to obtain the street view area S scenes .
[0174] The step S4 is specifically as follows:
[0175] S41. Generate a road area in a world coordinate system. The road area includes the blocked area and the street view area where the blocked area is located.
[0176] S42. Obtaining a lane direction of the road area according to the road network information rotated by the yaw angle;
[0177] S43, calculating the actual width and length of the blocked area and the actual width of the road area in the lane direction;
[0178] S44. Based on the traffic flow in the road area, a road congestion warning of a corresponding level is given.
[0179] In this step, the converted blocking event area is combined with the street view area to determine its impact on road traffic and calculate its location to facilitate the virtualization effect display. An example is shown below:
[0180] Reference Figure 3 , on the inverse perspective transformation plane, overlap the blocked area S on the street view area S scenes Here, the people, vehicles, and road surfaces in the street scene category are combined to form the road area;
[0181] Using the road network information rotated by the yaw angle ρ, the lane direction of the blocking event and the surrounding road surface is obtained; here, the inverse perspective transformation image is transferred to the lane direction of the blocking event;
[0182] By finding the maximum and minimum values x and y of the blocked area, the actual width w×dL and actual length h×dL of the blocked area are calculated, and then the average width w of the blocked road area in the lane direction is calculated. road The above vehicle trajectories can be used to conveniently estimate the traffic flow at the location. By combining the type of congestion event, road section width, and congestion event width and length, different warning levels can be flexibly designed.
[0183] After step S4, step S5 is also included:
[0184] S5. Calculating the actual position of the lane center in the blocked area;
[0185] The calculation process is as follows:
[0186] The distance to the camera is calculated as: d(r0′, v0′)*dL;
[0187] The angle is calculated as: ρ′-dρ′;
[0188] in,
[0189] ρ′ is the value of the yaw angle in the original image on the inverse perspective transformed image;
[0190] r0′ is the coordinate of the center of the blocked lane in the original image on the inverse perspective transformed image.
[0191] d(r0′, v0′) is the pixel distance from the center of the blocked lane r0′ to the vertical vanishing point v0′ on the inverse perspective transformed image.
[0192] As can be seen from the above description, this step calculates the actual position of the center of the blocked lane. After the calculation is completed, the rendering can be pasted to the corresponding position of other display systems to generate a visual effect of the lane width and the blocking event situation.
[0193] Please refer to Figure 6, embodiment 4 of the present invention is: a monitoring system for road obstruction events, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, any step of the above-mentioned method for monitoring road obstruction events is implemented.
[0194] In summary, the present invention provides a method and terminal for monitoring road obstruction events, which are compatible with target detection and image segmentation calculations for various types of road obstruction events, avoiding complex operations such as camera parameter settings. Based on image algorithms, a mapping relationship between the monitoring screen and the real-world plane can be constructed, thereby realizing the assessment of the impact of obstruction events on traffic obstacles and improving the coverage and practicality of this method in applications.
[0195] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent transformations made using the contents of the present invention's description and drawings, or directly or indirectly applied in related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for monitoring road obstruction events, characterized by: Including steps: S1. Obtain monitoring images, detect event targets and establish event areas; S2. Using an image algorithm to extract a target object from the event area and generate a blocked area; S3, converting the blocked area into a world coordinate system; S4. Evaluate the impact of the blocked area on road traffic in combination with the road environment in a world coordinate system; The step S30 is further included between the step S2 and the step S3: S301, using a neural network to establish corresponding rectangular frames for pedestrians and vehicles in the video stream and obtain key point coordinates; S302, generating a corresponding cylindrical reference object according to the key point coordinates; S303, performing camera self-calibration calculation using the columnar reference object; S304, constructing an inverse perspective transformation relationship matrix from the video stream to the world coordinate system; The step S304 is specifically as follows: S3041. Obtain a vehicle trajectory point sequence using a target tracking algorithm; S3042: Obtain camera data of any frame in the video stream to establish an inverse perspective transformation relationship matrix, and obtain the actual length represented by a pixel in the inverse perspective transformation relationship matrix; S3043, according to the inverse perspective transformation relationship matrix, calculate and obtain: the coordinates of the midpoint in the original image on the inverse perspective transformation image , the coordinates of the vertical vanishing point in the original image on the inverse perspective transformed image And the coordinate set of the vehicle trajectory point sequence on the inverse perspective transformed image ; S305: Calculate the vehicle yaw angle, where the vehicle yaw angle is the angle between the camera and the north direction; The step S305 is specifically as follows: S3051. Establish a calculation model based on the latitude and longitude coordinates of the monitoring location where the video stream is generated and the nearby road network information, where the monitoring location is the origin in the calculation model; S3052: performing an inverse perspective transformation on the road network information according to a first preset step size; S3053, obtaining the lane direction of the current vehicle according to the coordinates of the vehicle trajectory point sequence on the inverse perspective transformed image and the inverse perspective transformed road network information; S3054: Count the absolute values of the inner products of all vehicle trajectories and the corresponding lane directions and screen the maximum absolute value of the inner product. If multiple maximum inner product values are found, manually select the only maximum value and proceed to the next step. Otherwise, go directly to the next step; S3055. Calculate the yaw angle according to the maximum absolute value of the inner product; S3056: Within the preset calibration angle range of the yaw angle, regard the second preset step length as the first preset step length, and return to step S3052 until the yaw angle meets the preset accuracy.
2. The method for monitoring road obstruction events according to claim 1, characterized in that: The step S2 is specifically as follows: S21. Perform image semantic segmentation on the image using an image algorithm and obtain the category of each element; S22, filtering out street scene categories; S23: Eliminate elements belonging to the street scene category in the event area, and obtain a blocked area.
3. The method for monitoring road obstruction events according to claim 2, characterized in that: The step S23 is specifically as follows: S231, connecting each element belonging to the street scene category into its corresponding regional pixels; S232: If a certain region pixel is located within the event region, the region pixel in the event region is set to 0; if the intersection area of a certain region pixel and the event region is less than half of the event region, the region pixel in the event region is set to 0; if the region pixel is a pedestrian, the region pixel in the event region is set to 0; S233: Define an area within the event area where the pixels are not 0 as the blocking area.
4. The method for monitoring road obstruction events according to claim 3, characterized in that: The step S3 is specifically as follows: S3. Perform an inverse perspective transformation on the blocked area to obtain the blocked area in a world coordinate system.
5. The method for monitoring road obstruction events according to claim 4, characterized in that: The step S4 is specifically as follows: S41. Generate a road area in a world coordinate system; the road area includes the blocked area and the street view area where the blocked area is located; S42. Obtaining a lane direction of the road area according to the road network information rotated by the yaw angle; S43, calculating the actual width and length of the blocked area and the actual width of the road area in the lane direction; S44. Based on the traffic flow in the road area, a road congestion warning of a corresponding level is given.
6. The method for monitoring road obstruction events according to claim 5, characterized in that: After step S4, step S5 is also included: S5. Calculating the actual position of the lane center in the blocked area; The calculation process is as follows: The distance to the camera is calculated as: ; The angle is calculated as: ; in, ; is the value of the yaw angle in the original image on the inverse perspective transformed image; is the coordinate of the center of the blocked lane in the original image on the inverse perspective transformed image; The center of the blocked lane on the inverse perspective transformed image To the vertical vanishing point Pixel distance.
7. A monitoring terminal for road obstruction events, characterized by: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, any step of the method for monitoring a road obstruction event as described in claims 1 to 6 is implemented.
Citation Information
Patent Citations
Road waterlogging detection method based on video monitoring and terminal
CN115984772A
Road congestion determination method, terminal device and computer readable storage medium
CN108550259A
Traffic road condition analysis method and system and camera
CN111275960A