Unmanned vehicle road obstacle sensing system based on video monitoring
Data is obtained through the camera and lidar of the driverless vehicle, combined with object tracking algorithms and obstacle distance information, the obstacle identification and tracking problems in complex traffic environments are solved, the perception accuracy and matching success rate are improved, and the safe driving of driverless vehicles is ensured.
Patent Information
- Application Number
- CN202510518050.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-24
AI Technical Summary
It is difficult for autonomous vehicles to effectively identify and track road obstacles in complex traffic environments, especially in the case of interference and occlusion of multiple obstacles, resulting in a decrease in perceptual accuracy and matching success rate, affecting path planning and safe driving.
The road video is obtained through the camera of an unmanned vehicle, and combined with the lidar to obtain obstacle distance information, the object tracking algorithm is used for continuous tracking and detection, to determine the appearance matching degree and motion characteristics of the obstacle, calculate the occlusion degree, and thus improve tracking accuracy and position perception ability.
In the case of multiple obstacles associated interference and occlusion, the accuracy of obstacle detection and matching success rate are improved, ensuring effective tracking of obstacles, and providing reliable support for the path planning and safe driving of unmanned vehicles.
Smart Images

Figure CN120047537A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image communication, and particularly relates to an obstacle perception system for a driverless vehicle based on video surveillance. Background Art
[0002] A driverless vehicle needs to have the function of identifying and perceiving obstacles (including pedestrians, vehicles, roadblocks, etc.), and realizing detection, classification, positioning and dynamic behavior prediction. Usually, the driverless vehicle uses a camera to capture dynamic video data of the road scene in real time, providing an information source for the detection and recognition of obstacles, and then uses deep learning and computer vision algorithms to realize road obstacle perception. For example, an object tracking algorithm is used to track an obstacle for multiple frames, and the motion trajectory is predicted through the correlation matching of the obstacle in adjacent frame images.
[0003] However, the environment around a driverless vehicle is usually dynamic, and the type and position of obstacles will change over time. In complex traffic flows, the motion patterns of other traffic participants (such as vehicles, pedestrians, etc.) may cause the relative position of obstacles to change, easily resulting in multi-obstacle correlation interference and making multi-frame tracking difficult; or, an obstacle may be partially occluded by different objects at different time points, resulting in the failure of the correlation matching of the obstacle in adjacent frame images, thus affecting obstacle perception, unable to perform effective tracking, and unable to provide reliable support for the path planning and safe driving of the driverless vehicle. Summary of the Invention
[0004] In order to solve the above technical problems, the purpose of the present invention is to provide an obstacle perception system for a driverless vehicle based on video surveillance, and the specific technical solutions adopted are as follows: In a first aspect, an embodiment of the present application provides an obstacle perception system for a driverless vehicle based on video surveillance, including: An acquisition module, configured to acquire road videos through a camera of the driverless vehicle and acquire obstacle distance information around the driverless vehicle through a lidar; A first determination module, configured to continuously track and detect the road video through an object tracking algorithm to determine the appearance matching degree between obstacles in adjacent frame video images, and perform motion law analysis according to the obstacle distance information to determine the motion feature consistency of obstacles in adjacent frame video images; A second determination module, configured to determine the occlusion degree of an obstacle according to the appearance matching degree and the motion feature consistency when the obstacles in adjacent frame video images cannot be correspondingly matched during the continuous tracking and detection process; A perception module, configured to determine the tracking accuracy between obstacles in adjacent-frame video images according to the occlusion degree of the obstacles, and perform position perception of the obstacles according to the tracking accuracy.
[0005] In one implementation, the first determination module includes a first unit, a second unit, a third unit, and a fourth unit: The first unit is configured to continuously track and detect the road video through an object tracking algorithm, and determine the edges of each obstacle in each frame of video image; The second unit is configured to determine the consistency of the edge degree change between the obstacles in adjacent-frame video images according to the edges of each obstacle in each frame of video image, and determine the geometric feature consistency between the obstacles in adjacent-frame video images according to the consistency of the edge degree change; The third unit is configured to determine the color feature consistency between the obstacles in adjacent-frame video images according to the edges of each obstacle in each frame of video image; The fourth unit is configured to determine the appearance matching degree between the obstacles in adjacent-frame video images according to the first product of the geometric feature consistency and the color feature consistency.
[0006] In one implementation, the determining the color feature consistency between the obstacles in adjacent-frame video images according to the edges of each obstacle in each frame of video image includes: Determine a first frame of video image from several frames of video images, and determine a second frame of video image that is the previous frame of the first frame of video image; Respectively determine the first quantity of all pixel points and the second quantity of pixel points corresponding to different colors within the bounding boxes corresponding to the edges of each obstacle in the first frame of video image, and determine the third quantity of pixel points corresponding to different colors within the bounding boxes corresponding to the edges of each obstacle in the second frame of video image; Respectively determine the number of color types contained in the obstacles, and determine the color feature consistency between each obstacle in the first frame of video image and each obstacle in the second frame of video image according to the first quantity, the second quantity, the third quantity, and the number of color types, and return to the step of determining the first frame of video image from several frames of video images until the color feature consistency between the obstacles in adjacent-frame video images is determined.
[0007] In one implementation, the first determination module further includes a fifth unit and a sixth unit: The fifth unit is configured to determine the motion trend of each obstacle in each frame of video image according to the obstacle distance information corresponding to each frame of video image; The sixth unit is configured to determine the motion feature consistency of each obstacle in adjacent video images according to the motion trend of each obstacle in each frame of video image.
[0008] In one implementation, the determining the motion trend of each obstacle in each frame of video image according to the obstacle distance information corresponding to each frame of video image includes: Determine a third frame of video image from several frames of video images, determine several consecutive fourth frame of video images within a preset proximity range of the third frame of video image, and generate a position fitting line according to the obstacle distance information corresponding to the several consecutive fourth frame of video images; According to the obstacle distance information corresponding to each frame of video image, respectively determine the first position of each obstacle in any fourth frame of video image and the second position of the corresponding matching obstacle in the previous fourth frame of video image of the fourth frame of video image, and respectively determine the angle between the line connecting the first position and the second position and the corresponding position fitting line; Respectively determine the velocity variances of each obstacle in the third frame of video image and each obstacle in each fourth frame of video image; Obtain an average value of the angles by summing and averaging according to the angle and the number of the angles, and respectively determine the motion trend of each obstacle in the third frame of video image according to the second product of the average value of the angles and the velocity variances of the obstacles and the natural exponential function, and return to the step of determining the third frame of video image from several frames of video images until the motion trend of each obstacle in each frame of video image is determined.
[0009] In one implementation, the determining the motion feature consistency of each obstacle in adjacent video images according to the motion trend of each obstacle in each frame of video image includes: Determine the motion trend differences of each obstacle in adjacent video images according to the motion trend of each obstacle in each frame of video image; Respectively determine the motion feature consistency of each obstacle in adjacent video images according to the opposite number of the motion trend differences and the natural exponential function.
[0010] In one implementation, the second determination module includes a seventh unit, an eighth unit, a ninth unit, and a tenth unit: The seventh unit is configured to, when the obstacles in adjacent video images cannot be correspondingly matched during the continuous tracking and detection process, determine a fifth frame of video image from several frames of video images, and determine a sixth frame of video image that is the previous frame of the fifth frame of video image; The eighth unit is configured to predict the positions of each obstacle in the sixth-frame video image in the fifth-frame video image through a uniform acceleration motion model, and obtain a plurality of predicted positions; The ninth unit is configured to determine the average appearance matching degree corresponding to each obstacle according to the appearance matching degree between each obstacle in the adjacent-frame video images corresponding to the fifth-frame video image and the sixth-frame video image, and respectively determine the first difference value between the appearance matching degree between each obstacle in the adjacent-frame video images corresponding to the fifth-frame video image and the sixth-frame video image and the average appearance matching degree corresponding to each obstacle; The tenth unit is configured to respectively determine the occlusion degree of the obstacle at each predicted position in the fifth-frame video image according to the third product of the motion feature consistency of each obstacle in the adjacent-frame video images corresponding to the fifth-frame video image and the sixth-frame video image and the first difference value, and return the step of determining the fifth-frame video image from several frame video images when the obstacles in the adjacent-frame video images cannot be correspondingly matched during the continuous tracking and detection process until the occlusion degree of the obstacle at each predicted position in each frame video image is determined.
[0011] In one embodiment, the perception module includes a first processing unit, a second processing unit, and a third processing unit: The first processing unit is configured to determine the identity between each obstacle in the adjacent-frame video images according to the fourth product of the motion feature consistency and the appearance matching degree; The second processing unit is configured to determine the corrected matching degree between the obstacle at each predicted position in each frame video image and each obstacle in the previous-frame video image of this frame video image according to the identity between each obstacle in the adjacent-frame video images and the occlusion degree of the obstacle; The third processing unit is configured to determine the tracking accuracy between each obstacle in the adjacent-frame video images according to the corrected matching degree.
[0012] In one embodiment, the determining the corrected matching degree between the obstacle at each predicted position in each frame video image and each obstacle in the previous-frame video image of this frame video image according to the identity between each obstacle in the adjacent-frame video images and the occlusion degree of the obstacle includes: Respectively determine the sum value of the preset value and the occlusion degree of the obstacle at each predicted position in each frame video image; Respectively determine the corrected matching degree between the obstacle at each predicted position in each frame video image and each obstacle in the previous-frame video image of this frame video image according to the fifth product of the sum value and the identity between each obstacle in the adjacent-frame video images.
[0013] In one embodiment, determining the tracking accuracy between obstacles in adjacent frame video images according to the corrected matching degree includes: According to the corrected matching degrees between the obstacles at each predicted position in each frame of video image and the obstacles in the previous frame of video image of this frame, respectively determine the average value of the corrected matching degrees corresponding to each obstacle; Respectively determine the second difference values between the corrected matching degrees between the obstacles at each predicted position in each frame of video image and the obstacles in the previous frame of video image of this frame and the average value of the corrected matching degrees corresponding to each obstacle; Respectively determine the calculation results according to the opposite number of the second difference value and the natural exponential function, and respectively determine the tracking accuracy between obstacles in adjacent frame video images according to the calculation results and the sixth product of the average value of the corrected matching degrees corresponding to each obstacle.
[0014] The present invention has the following beneficial effects: Obtain road videos through the camera of the driverless vehicle and obtain the obstacle distance information around the driverless vehicle through lidar. Continuously track and detect the road video through the object tracking algorithm to determine the appearance matching degree between obstacles in adjacent frame video images, and analyze the motion law according to the obstacle distance information to determine the motion feature consistency of obstacles in adjacent frame video images. When the obstacles in adjacent frame video images cannot be correspondingly matched during the continuous tracking and detection process, determine the occlusion degree of the obstacles according to the appearance matching degree and the motion feature consistency. By analyzing the appearance matching degree and the motion feature consistency, it is beneficial to improve the detection accuracy and matching success rate of obstacles in the case of multi-obstacle associated interference. Determine the tracking accuracy between obstacles in adjacent frame video images according to the occlusion degree of the obstacles, and perform position perception of the obstacles according to the tracking accuracy. Performing position perception of the obstacles based on the tracking accuracy determined by the occlusion degree is beneficial to locate and perceive the obstacles when the obstacles are occluded, realize effective tracking of the obstacles, and provide reliable support for the path planning and safe driving of the driverless vehicle. Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for description in the embodiments or the prior art. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1Block diagram of an obstacle perception system for driverless vehicles on roads based on video surveillance provided by an embodiment of the present invention; Figure 2 Comparison diagram of video images and radar data provided by an embodiment of the present invention. Among them, (a) is a frame of video image obtained at a certain time, and (b) is a schematic diagram of radar data obtained at the same time; Figure 3 Schematic diagram of the bounding boxes of each obstacle provided by an embodiment of the present invention; Figure 4 Comparison diagram of the edges obtained by bounding box and contour detection provided by an embodiment of the present invention. Among them, (a) is a schematic diagram of the bounding box of a certain obstacle, and (b) is a schematic diagram of the edges obtained by contour detection of the obstacle within the bounding box of the obstacle; Figure 5 Schematic diagram of the corresponding position fitting line of a certain obstacle provided by an embodiment of the present invention. Detailed implementation manners
[0017] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following combines the accompanying drawings and preferred embodiments to detail the specific implementation manners, structures, features and effects of an obstacle perception system for driverless vehicles on roads based on video surveillance proposed according to the present invention. In the following description, different "an embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0019] It should be noted that the "exemplary" in the embodiments of the present application refers to examples listed for convenience of description, and other embodiments are not limited to the listed examples.
[0020] It should be noted that to ensure the significance of the calculation results, when performing fractional operations in the embodiments of the present invention, in the case of encountering a denominator of 0, a tuning factor greater than 0 needs to be added to the denominator to prevent the denominator from being 0. The value of the tuning factor is set by the implementer according to the actual situation, and no special limitation is made in this application.
[0021] The following specifically describes the specific solution of an obstacle perception system for driverless vehicles on roads based on video surveillance provided by the present invention with reference to the accompanying drawings.
[0022] Please refer to Figure 1, which shows a structural block diagram of a road obstacle perception system for driverless vehicles based on video surveillance provided by an embodiment of the present invention. The road obstacle perception system for driverless vehicles based on video surveillance may at least include: An acquisition module, configured to acquire a road video through a camera of the driverless vehicle and acquire obstacle distance information around the driverless vehicle through a lidar; A first determination module, configured to continuously track and detect the road video through an object tracking algorithm to determine the appearance matching degree between obstacles in adjacent frame video images, and perform motion law analysis based on the obstacle distance information to determine the motion feature consistency of obstacles in adjacent frame video images; A second determination module, configured to determine the occlusion degree of an obstacle according to the appearance matching degree and the motion feature consistency when the obstacles in adjacent frame video images cannot be correspondingly matched during the continuous tracking and detection process; A perception module, configured to determine the tracking accuracy between obstacles in adjacent frame video images according to the occlusion degree of the obstacles, and perform position perception of the obstacles according to the tracking accuracy.
[0023] The technical solution of the embodiment of the present application acquires a road video through a camera of the driverless vehicle and acquires obstacle distance information around the driverless vehicle through a lidar, continuously tracks and detects the road video through an object tracking algorithm to determine the appearance matching degree between obstacles in adjacent frame video images, and performs motion law analysis based on the obstacle distance information to determine the motion feature consistency of obstacles in adjacent frame video images. When the obstacles in adjacent frame video images cannot be correspondingly matched during the continuous tracking and detection process, the occlusion degree of the obstacles is determined according to the appearance matching degree and the motion feature consistency. By analyzing the appearance matching degree and the motion feature consistency, it is beneficial to improve the detection accuracy and matching success rate of obstacles when there are multi-obstacle correlation interferences. According to the occlusion degree of the obstacles, the tracking accuracy between obstacles in adjacent frame video images is determined, and the position perception of the obstacles is performed according to the tracking accuracy. Performing position perception of the obstacles based on the tracking accuracy determined by the occlusion degree is beneficial to perform positioning perception of the obstacles when the obstacles are occluded, realize effective tracking of the obstacles, and provide reliable support for the path planning and safe driving of the driverless vehicle.
[0024] In one embodiment, a camera can be installed on the top inside of a driverless vehicle (hereinafter referred to as the vehicle). The road video containing surrounding roads, obstacles, traffic signs, etc. is obtained in real time through the camera of the driverless vehicle, and the three-dimensional space information where the vehicle is located is mapped into a two-dimensional image (i.e., several frame video images contained in the video). In addition, a lidar is installed on the top outside of the vehicle at a corresponding position consistent with the camera. The lidar uses laser emission and reception of reflected signals to obtain radar data to measure the distance information of surrounding obstacles. Therefore, the distance information of obstacles around the vehicle can be obtained.
[0025] Optionally, after obtaining the road video and the obstacle distance information, preprocessing can be performed. The collected road video and obstacle distance information are preprocessed to reduce noise and enhance the quality of the data. For example, the road video is processed through a denoising filtering algorithm (such as Gaussian filtering) to remove noise caused by light changes or motion blur. The radar data is processed through Kalman filtering to remove invalid points and reduce noise generated by external interference. Then, data synchronization and fusion can be performed: synchronize the road video captured by the camera and the radar data through timestamps to ensure the temporal consistency of the two data, so that each frame of video image has corresponding radar data, that is, each frame of video image has corresponding obstacle distance information. As Figure 2 shown, Figure 2 is a comparison diagram of video images and radar data provided by an embodiment of the present invention. Among them, (a) is a frame of video image obtained at a certain time, and (b) is the radar data obtained at the same time.
[0026] In one embodiment, the first determination module includes a first unit, a second unit, a third unit, and a fourth unit, and is used to continuously track and detect the road video through an object tracking algorithm to determine the appearance matching degree between obstacles in adjacent frame video images. Specifically: The first unit is used to continuously track and detect the road video through an object tracking algorithm to determine the edges of each obstacle in each frame of video image.
[0027] In one embodiment, an object tracking algorithm such as (Tracktor algorithm) can be used to continuously track and detect the road video. The Tracktor algorithm is an object tracking algorithm based on object detection, which combines object detection and tracking and performs real-time object tracking through continuous detection results. When the Tracktor algorithm continuously tracks and detects the road video, it performs object detection on each frame of video image in the road video to generate the bounding boxes of each obstacle in each frame of video image. As Figure 3As shown in , since obstacles in the road environment (such as vehicles, pedestrians, traffic signs, etc.) have different shape characteristics, the edge detection algorithm can be used to extract the contour information of the obstacles in the bounding box, so as to obtain the edges of the obstacles, and finally obtain the edges of each obstacle in each frame of the video image (each obstacle can have multiple edges), as shown in Figure 4 As shown, Figure 4 A comparison diagram of the edge obtained by the bounding box and contour detection provided by an embodiment of the present invention, wherein (a) is the bounding box of an obstacle, and (b) is the edge obtained by contour detection of the obstacle within the bounding box of the obstacle. The Tracktor algorithm matches the obstacles in the adjacent video frames, such as calculating the position overlap of the bounding boxes for matching, and maintains and updates the position of the obstacles through data association and motion models to achieve continuous tracking of each obstacle. The Tracktor algorithm is an existing method, and its principle of continuous obstacle tracking will not be described in detail.
[0028] It should be noted that when calculating the position overlap of the bounding box for matching, if a match can be made, it means that the same obstacle is successfully found in the adjacent frames of video images. However, due to the dynamic and complex characteristics of traffic roads, there are many obstacles and dynamic movement, which may cause multiple obstacles to be associated with interference, occlusion, etc., resulting in matching failure. That is, in the adjacent frames of video images, the obstacle in the previous frame of video image cannot be found in the next frame of video image. Therefore, the existing method of only associating the bounding box is prone to matching errors and resulting in tracking failure. Therefore, the method of the present application is to make further improvements in the original object tracking algorithm, introduce dimensions such as appearance matching, motion feature consistency, and obstacle occlusion, to maximize the perception of obstacle position and improve the success rate and accuracy of tracking.
[0029] The second unit is used to determine the consistency of edge degree changes between obstacles in adjacent frames of video images based on the edges of obstacles in each frame of video image, and determine the consistency of geometric features between obstacles in adjacent frames of video images based on the consistency of edge degree changes.
[0030] In one embodiment, according to the edge of each obstacle in each frame of video image, the consistency of edge degree change between each obstacle in adjacent frames of video images is determined. , Stands for Dynamic TimeWarping, dynamic time warping, For the The obstacle is in Frame video image The chain code sequence of the edge of the strip (the chain code sequence is obtained by the existing method), For the The chain code sequence of the th edge of any obstacle in the frame video image and the adjacent previous frame video image, is the th obstacle's change consistency of edges from long to short (i.e., the change consistency of edge degree) in the th frame video image and its adjacent previous frame video image corresponding to any obstacle (when any obstacle corresponds to different obstacles, it can represent the change consistency of edge degree between each obstacle), The smaller it is, the more consistent the edge performance is.
[0031] Then, according to the change consistency of edge degree, determine the geometric feature consistency between each obstacle in adjacent frame video images : Among them, represents the geometric feature consistency of the th obstacle in the th frame video image and its adjacent previous frame video image corresponding to any obstacle (it can be understood that when any obstacle corresponds to different obstacles, at this time can represent the geometric feature consistency of the th obstacle in the th frame video image and its adjacent previous frame video image corresponding to each obstacle, that is, the geometric feature consistency between each obstacle in adjacent frame video images), represents the number of edges of the obstacle, represents the natural exponential function.
[0032] The third unit is used to determine the color feature consistency between each obstacle in adjacent frame video images according to the edges of each obstacle in each frame video image.
[0033] It should be noted that in the road environment, usually the colors of different obstacles are significantly different. The colors of vehicles, pedestrians, and road signs are often quite unique. For example, the body of a car usually has a fixed color (such as red, black, white), while pedestrians may wear different colored clothes. Therefore, analyzing the color feature consistency is beneficial to the matching and recognition of obstacles.
[0034] First, determine the first frame video image (such as the th frame video image) from several frame video images, and determine the second frame video image which is the previous frame of the first frame video image. For example, the previous frame of the th frame video image is the second frame video image.
[0035] Secondly, respectively determine the first quantity of all pixel points within the bounding box corresponding to the edges of each obstacle in the first frame video image (i.e., the number of pixel points within the bounding box corresponding to the edge of the th obstacle in the th frame of the video image) and the second quantity corresponding to pixel points of different colors (i.e., the number corresponding to pixel points of the th color within the bounding box corresponding to the edge of the th obstacle in the th frame of the video image), and determine the third quantity corresponding to pixel points of different colors within the bounding box corresponding to the edge of each obstacle in the second frame of the video image (i.e., the number corresponding to pixel points of the th color within the bounding box corresponding to the edge of any obstacle in the video image immediately preceding the th frame of the video image, and when any obstacle corresponds to different obstacles, it represents each obstacle).
[0036] Furthermore, respectively determine the number of color types contained in the obstacle , and based on the first quantity , the second quantity , the third quantity and the number of color types , determine the color feature consistency between each obstacle in the first frame of the video image and each obstacle in the second frame of the video image , and the formula is: where represents the color feature consistency between the th obstacle in the th frame of the video image and any obstacle in the immediately preceding frame of the video image adjacent to it (similarly, when any obstacle corresponds to different obstacles, it represents each obstacle, that is, the color feature consistency between each obstacle in the first frame of the video image and each obstacle in the second frame of the video image can be determined). Then, return to the step of determining the first frame of the video image from several frames of the video image until the color feature consistency between each obstacle of adjacent frames of the video image is determined , that is when taking different values, it represents the color feature consistency between each obstacle of different adjacent frames of the video image. Among them, represents the difference in the proportion of pixel points of the same color type within the bounding box between the th obstacle in the th frame of the video image and each obstacle in the immediately preceding frame of the video image adjacent to it. The smaller this formula is, the more consistent the color proportion is.
[0037] The fourth unit is used to determine the appearance matching degree between obstacles in adjacent frame video images according to the first product of geometric feature consistency and color feature consistency: Optionally, in the road environment of an autonomous vehicle, according to the geometric feature consistency and color feature consistency obtained from video monitoring, the movement trajectory of the same obstacle can be tracked, and the appearance matching degree of the obstacles corresponding to the bounding boxes in adjacent video frames can be comprehensively obtained. Specifically: according to the geometric feature consistency and the color feature consistency of the first product, determine the appearance matching degree between obstacles in adjacent frame video images : wherein, is the appearance matching degree of the th obstacle in the th frame video image and any obstacle in its adjacent previous frame video image.
[0038] It should be noted that introducing the appearance matching degree in the object tracking algorithm can perform preliminary obstacle matching. However, due to the complex and dynamic road environment, when there are different obstacles blocking or obstacles with similar appearances (such as cars of the same model and color), it is easy to have matching errors. Therefore, the embodiments of the present application further perform motion law analysis based on the obstacle distance information, predict the positions of obstacles in future frames, and determine whether occlusion occurs, identify and correct the obstacle matching errors caused by occlusion, and ensure continuous tracking of obstacles.
[0039] In one implementation manner, the first determination module further includes a fifth unit and a sixth unit, which are used to perform motion law analysis according to the obstacle distance information and determine the motion feature consistency of obstacles in adjacent frame video images. Specifically: The fifth unit is used to determine the motion trend of each obstacle in each frame video image according to the obstacle distance information corresponding to each frame video image.
[0040] First, determine the third frame video image from several frame video images. Similarly, taking the th frame video image as an example, assuming that the preset proximity range is the previous 5 frames, at this time, several consecutive fourth frame video images within the preset proximity range of the third frame video image can be determined, that is, determine the previous 5 consecutive fourth frame video images of the th frame video image, and generate a position fitting straight line according to the obstacle distance information corresponding to several consecutive fourth frame video images. Because based on the obstacle distance information corresponding to different frame video images, the positions of each obstacle in each frame video image can be determined respectively, and thus a corresponding position fitting straight line can be generated based on the change of positions, such as Figure 5As shown, it is the position fitting line corresponding to a certain obstacle.
[0041] Secondly, according to the obstacle distance information corresponding to each frame of video image, respectively determine the first position of each obstacle in any fourth frame of video image and the second position of the corresponding matched obstacle in the previous frame of the fourth frame of video image, and respectively determine the angle between the straight line connecting the first position and the second position and the corresponding position fitting line , that is, the th obstacle in the th frame of video image (the third frame of video image) adjacent to the th fourth frame of video image, the angle between the straight line connecting the first position where any obstacle is located and the second position of the matched obstacle in the previous frame (i.e., the -1th fourth frame of video image) and the position fitting line.
[0042] Then, respectively determine the velocity variance of each obstacle in the third frame of video image and each obstacle in each fourth frame of video image , which represents the velocity variance of the th obstacle in the th frame of video image (the third frame of video image) and each obstacle in each adjacent fourth frame of video image. Specifically, the velocity can be obtained by the ratio of the position distance to the time interval, and then the velocity variance is calculated, representing the continuity of the movement change of the obstacle. The smaller this value is, the more continuous the movement is and the more consistent the movement trend is.
[0043] Finally, according to the angle and the number of angles , for a preset adjacent range such as 5, perform summation and averaging to obtain the average angle , respectively according to the second product of the average angle and the velocity variance of each obstacle and the natural exponential function , determine the movement trend of each obstacle in the third frame of video image : Among them, represents the movement trend of the th obstacle in the th frame of video image (the third frame of video image), return the steps to determine the third frame of video image from several frames of video images until the movement trend of each obstacle in each frame of video image is determined , that is, when it is different values, it represents the movement trend of each obstacle in each frame of video image. Among them, Indicates the consistency of the position change direction of the th obstacle in the
[0044] current frame of video image and its adjacent frame of video image. The smaller this value is, the more consistent the movement direction of the obstacle in consecutive frames is.
[0045] It should be noted that in the road environment of driverless vehicles, the movement of long-time objects may produce large differences. However, due to traffic rules, the movement laws of objects are consistent in a short period of time. During the matching process of obstacles in the current frame of video image, it is necessary to be consistent with the movement trend of the possible same obstacle in the previous frame. According to the movement trend of obstacles in adjacent frames, the movement feature consistency of obstacles is obtained.
[0046] First, according to the movement trend of each obstacle in each frame of video image , determine the movement trend difference of each obstacle in adjacent frames of video image , is the movement trend of any obstacle in the adjacent previous frame of video image of the th frame of video image. Similarly, when any obstacle corresponds to different obstacles, it corresponds to the movement trends of each obstacle in the adjacent previous frame of video image of the th frame of video image.
[0047] Secondly, respectively determine the movement feature consistency of each obstacle in adjacent frames of video image according to the opposite number of the movement trend difference and the natural exponential function: where indicates the movement feature consistency of the th obstacle in the th frame of video image and any obstacle in its adjacent previous frame of video image. Similarly, when any obstacle corresponds to different obstacles, the movement feature consistency of each obstacle in adjacent frames of video image is obtained. reflects the difference in the movement trend of the th obstacle in the th frame of video image and any obstacle in its adjacent previous frame of video image. The smaller this formula is, the more consistent the movement trend is, and the more likely it is to be the same obstacle.
[0048] It should be noted that in a complex traffic road environment, there are multiple obstacles, and the vehicles block each other. If there is no obstacle with a high degree of identity in the current frame of video image compared to the previous frame of video image, it indicates that the obstacle may be blocked. In adjacent frames of video images, it appears different in appearance, but the movement trajectory of the position information may belong to the same object. Therefore, the occlusion degree of the obstacle can be obtained: if the obstacle matching fails in the previous video frame of the current frame, the uniform acceleration motion model is used to predict the position of the obstacle in the previous frame in the current frame, and the occlusion degree is obtained through the target information at this position.
[0049] In one implementation, the second determination module includes a seventh unit, an eighth unit, a ninth unit, and a tenth unit, which are used to determine the occlusion degree of the obstacle according to the appearance matching degree and the consistency of motion characteristics when the obstacles in adjacent frames of video images cannot be correspondingly matched during the continuous tracking and detection process. Specifically: The seventh unit is used to determine the fifth frame of video image from several frames of video images when the obstacles in adjacent frames of video images cannot be correspondingly matched during the continuous tracking and detection process, that is, when the Tracktor algorithm cannot correspondingly match the obstacles in adjacent frames of video images during the continuous tracking and detection of the obstacles. Similarly, taking the fifth frame of video image as an example, and determining the sixth frame of video image which is the previous frame of the fifth frame of video image, that is, the previous frame of the fifth frame of video image is the sixth frame of video image.
[0050] The eighth unit is used to predict the positions of the obstacles in the sixth frame of video image in the fifth frame of video image through the uniform acceleration motion model, and obtain several predicted positions.
[0051] The ninth unit is used to determine the average value of the appearance matching degree corresponding to each obstacle according to the appearance matching degree between the obstacles in the adjacent frame video images corresponding to the fifth frame of video image and the sixth frame of video image (that is, the average value of the appearance matching degree corresponding to each obstacle in the adjacent continuous video frames (the adjacent previous frame of video image) of the th obstacle in the th frame of video image), and respectively determine the appearance matching degree between the obstacles in the adjacent frame video images corresponding to the fifth frame of video image and the sixth frame of video image (that is, the appearance matching degree corresponding to each obstacle in the th obstacle in the th frame of video image and the adjacent previous frame of video image), and the first difference value from the average value of the appearance matching degree corresponding to the corresponding each obstacle .
[0052] The tenth unit is configured to determine the occlusion degree of obstacles at each predicted position in the fifth-frame video image respectively according to the third product of the motion feature consistency of each obstacle in the adjacent-frame video images corresponding to the fifth-frame video image and the sixth-frame video image and the first difference value: Wherein, represents the occlusion degree of the th obstacle at a certain predicted position in the th frame of video image (the fifth-frame video image); the larger is, the more likely it is that obstacles with the same motion trend are occluded in the current frame.
[0053] Then, return the step of determining the fifth-frame video image from several frames of video images when the obstacles in the adjacent-frame video images cannot be correspondingly matched during the continuous tracking and detection process, until the occlusion degree of obstacles at each predicted position in each frame of video image is determined , that is, different values can correspondingly determine the occlusion degree of obstacles at each predicted position in different frames of video images.
[0054] In one implementation, the perception module includes a first processing unit, a second processing unit, and a third processing unit, and is configured to determine the tracking accuracy between each obstacle in the adjacent-frame video images according to the occlusion degree of the obstacles. Specifically: The first processing unit is configured to determine the identity between each obstacle in the adjacent-frame video images according to the fourth product of the motion feature consistency and the appearance matching degree.
[0055] It should be noted that in a road environment, due to the complex situation of multiple obstacles, there may be tracking interference between obstacles. According to the appearance matching degree of obstacles in the adjacent-frame video images and the motion law of obstacles in consecutive video frames, the identity between each obstacle in the adjacent-frame video images can be obtained: Wherein, represents the identity of the th obstacle in the th frame of video image and any obstacle in its adjacent previous video frame. Similarly, when any obstacle corresponds to different obstacles, the identity between each obstacle in the adjacent-frame video images can be obtained.
[0056] The second processing unit is configured to determine the corrected matching degree between each obstacle at each predicted position in each frame of video image and each obstacle in the video image of the previous frame of this frame of video image according to the identity between each obstacle in the adjacent-frame video images and the occlusion degree of the obstacles.
[0057] It should be noted that when the obstacle is blocked, its matching degree is relatively low. According to the occlusion degree obtained from the movement law of the obstacle, the obstacle matching error caused by occlusion is corrected, and the corrected matching degree at the position of the obstacle is obtained.
[0058] First, assume that the preset value is 1, and determine the sum of the occlusion degrees of the preset value 1 and the obstacles at each predicted position in each frame of video image respectively value .
[0059] Secondly, according to the sum value and the identity between the obstacles in adjacent frames of video images , determine the corrected matching degree between the obstacles at each predicted position in each frame of video image and the obstacles in the previous frame of the video image of this frame: Among them, represents the corrected matching degree between the th obstacle at the predicted position in the th frame of video image and any obstacle in the adjacent previous frame of video image. When any corresponding value takes different values, the corrected matching degree between the obstacles at each predicted position in each frame of video image and the obstacles in the previous frame of the video image of this frame can be obtained .
[0060] It should be noted that the Tracktor algorithm is used to continuously track the obstacles, and the matching degree of the obstacles is continuously corrected according to their movement trajectories to ensure stable and accurate tracking of multiple obstacles in a complex and dynamic road environment. According to the corrected matching degree at the position of the obstacle obtained by target detection, the tracking accuracy of the obstacle is obtained.
[0061] The third processing unit is used to determine the tracking accuracy between the obstacles in adjacent frames of video images according to the corrected matching degree.
[0062] First, according to the corrected matching degree between the obstacles at each predicted position in each frame of video image and the obstacles in the previous frame of the video image of this frame , determine the average value of the corrected matching degrees corresponding to each obstacle respectively , that is obtained by taking the average.
[0063] Secondly, determine the second difference value between the corrected matching degree between the obstacles at each predicted position in each frame of video image and the obstacles in the previous frame of the video image of this frame and the average value of the corrected matching degrees corresponding to each obstacle respectively . The smaller this formula value is, the higher the probability that the obstacle and the obstacle in the previous consecutive frames are the same obstacle's motion trajectory (motion trend), and the greater the tracking accuracy.
[0064] Then, respectively according to the opposite number of the second difference value and the natural exponential function , determine the calculation result , and respectively according to the calculation result and the sixth product of the corrected matching degree mean corresponding to each obstacle, determine the tracking accuracy between each obstacle in adjacent frame video images. The specific formula is: Wherein, represents the tracking accuracy of the th obstacle in the th frame video image and any obstacle in its adjacent previous frame video image. Take different values corresponding to each obstacle, so as to determine the tracking accuracy between each obstacle in adjacent frame video images.
[0065] In one implementation, according to the tracking accuracy, perform obstacle position perception. Specifically: a tracking threshold can be set in advance, such as 0.7. If the tracking accuracy of a certain obstacle is greater than 0.7, it means that the probability that a certain obstacle in consecutive frame video images is the same obstacle's motion trajectory (motion trend) is high, then it is determined that the obstacle can be accurately detected in the current frame video image. And if the obstacle in the previous frame video image is successfully matched, at this time the tracking is successful, and the obstacle distance information of the obstacle is updated.
[0066] Optionally, when performing obstacle position perception, according to the obtained motion law of the obstacle, use the uniformly accelerated motion model to predict the future motion trajectory of the obstacle, judge the position change of the obstacle in advance, and provide a basis for path planning. And when the obstacle may be occluded, determine the occlusion degree and tracking accuracy of the obstacle through the above method for occlusion detection and repair to improve the tracking accuracy; and when the obstacle reappears in the field of view after being occluded, perform obstacle re-identification and matching. It can be judged whether it is the previously tracked obstacle by comparing the appearance matching degree or motion feature consistency of the obstacle, avoiding incorrect trajectory matching. Finally, on the basis of ensuring normal obstacle matching, the obstacle can be classified and dynamically analyzed for features, identifying the type of the obstacle (such as pedestrians, vehicles, static objects, etc.) and motion features (such as speed, acceleration, etc.). The unmanned vehicle system can then perform real-time path planning to adjust the driving route of the vehicle, ensure that the vehicle finds the optimal driving path in a dynamic environment, and safely bypass the obstacle.
[0067] The method of the embodiment of the present application collects the road video of the driverless vehicle, analyzes the movement law through the obstacle distance information of the obstacles in the continuous frame video images to distinguish different obstacles, prevents multi-obstacle interference in tracking, extracts the consistency of the movement characteristics of the obstacles in the continuous video frames, and combines the spatial position relationship determined by the obstacle distance information relative to the driverless vehicle to achieve continuous tracking of the obstacles, improving the perception ability of the driverless vehicle in complex and dynamic road environments. At the same time, during the process of tracking the obstacles, the obstacles in the frame video images are identified and located. By the associated matching of the obstacles in the continuous frame video images, the same obstacle in different frame video images is judged, and cross-frame continuous tracking is achieved, avoiding the limitations of the static obstacle detection method, and being able to stably track the obstacles (such as pedestrians, other vehicles, etc.) in the dynamic scene, ensuring the continuous recognition of the obstacles. In addition, to prevent the tracking failure of the obstacles caused by occlusion, analyzing the movement law of the obstacles to determine the consistency of the movement characteristics can improve the positioning accuracy of the obstacles and predict the behavior trend of the obstacles. Finally, by identifying and predicting the occlusion situation, the correction matching degree of the continuous obstacles is corrected. Even if the obstacle disappears briefly, its position can be restored through the trajectory, ensuring the stability and robustness of the obstacle perception system in complex environments, especially in high-density traffic or rapidly changing scenes, avoiding misjudgment caused by obstacle occlusion, enabling the road obstacle perception system of the driverless vehicle to adapt to environmental changes in real time, maintaining efficient and accurate obstacle tracking, and thus providing strong support for the safe driving, path planning and decision-making of the driverless vehicle.
[0068] It should be noted that the above sequence of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0069] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.
Claims
1. A road obstacle perception system for unmanned vehicles based on video surveillance, characterized in that: The system comprises: An acquisition module, used to acquire road video through a camera of the unmanned vehicle and acquire obstacle distance information around the unmanned vehicle through a laser radar; A first determination module is used to continuously track and detect the road video through an object tracking algorithm to determine the appearance matching degree between obstacles in adjacent frames of video images, and to perform motion law analysis based on the obstacle distance information to determine the consistency of motion features of obstacles in adjacent frames of video images; A second determination module is used to determine the degree of occlusion of the obstacle according to the appearance matching degree and the consistency of the motion features when obstacles in adjacent frame video images cannot be matched in the process of continuous tracking and detection; The perception module is used to determine the tracking accuracy between obstacles in adjacent frames of video images according to the degree of occlusion of the obstacles, and to perceive the position of the obstacles according to the tracking accuracy.
2. The road obstacle perception system for unmanned vehicles based on video surveillance according to claim 1 is characterized by: The first determining module includes a first unit, a second unit, a third unit and a fourth unit: The first unit is used to continuously track and detect the road video through an object tracking algorithm to determine the edge of each obstacle in each frame of the video image; The second unit is used to determine the consistency of edge degree changes between obstacles in adjacent frames of video images according to the edges of obstacles in each frame of video images, and determine the consistency of geometric features between obstacles in adjacent frames of video images according to the consistency of edge degree changes; The third unit is used to determine the consistency of color features between obstacles in adjacent frames of video images according to the edges of obstacles in each frame of video image; The fourth unit is used to determine the appearance matching degree between obstacles in adjacent frame video images according to the first product of the geometric feature consistency and the color feature consistency.
3. The road obstacle perception system for unmanned vehicles based on video surveillance according to claim 2 is characterized in that: Determining the consistency of color features between obstacles in adjacent frames of video images according to the edges of obstacles in each frame of video images includes: Determine a first frame of video image from a plurality of frames of video image, and determine a second frame of video image that is a frame before the first frame of video image; Respectively determine a first number of all pixels and a second number of pixels of different colors in a bounding box corresponding to an edge of each obstacle in the first frame of video image, and determine a third number of pixels of different colors in a bounding box corresponding to an edge of each obstacle in the second frame of video image; Determine the number of color types contained in the obstacles respectively, determine the consistency of color features of each obstacle in the first frame of video image and each obstacle in the second frame of video image according to the first number, the second number, the third number and the number of color types, and return to the step of determining the first frame of video image from the plurality of frames of video images until the consistency of color features between the obstacles in adjacent frames of video images is determined.
4. The unmanned vehicle road obstacle perception system based on video surveillance according to claim 2 or 3, characterized in that: The first determining module further includes a fifth unit and a sixth unit: The fifth unit is used to determine the movement trend of each obstacle in each frame of video image according to the obstacle distance information corresponding to each frame of video image; The sixth unit is used to determine the consistency of motion features of each obstacle in adjacent frames of video images according to the motion trend of each obstacle in each frame of video image.
5. The road obstacle perception system for unmanned vehicles based on video surveillance according to claim 4 is characterized in that: Determining the movement trend of each obstacle in each frame of video image according to the obstacle distance information corresponding to each frame of video image includes: Determine a third frame of video image from the plurality of frames of video image, determine a plurality of consecutive fourth frames of video image within a preset proximity range of the third frame of video image, and generate a position fitting straight line according to obstacle distance information corresponding to the plurality of consecutive fourth frames of video image; According to the obstacle distance information corresponding to each frame of video image, respectively determine the first position of each obstacle in any fourth frame of video image and the second position of the corresponding matched obstacle in the fourth frame of video image before the fourth frame of video image, and respectively determine the angle between the straight line connecting the first position and the second position and the corresponding position fitting straight line; respectively determining a speed variance of each obstacle in the third frame of video image and each obstacle in each of the fourth frame of video image; An average angle value is obtained by summing and averaging the angles and the number of the angles, and a movement trend of each obstacle in the third frame of video image is determined according to the second product of the average angle and the speed variance of each obstacle and a natural exponential function, and the process returns to the step of determining the third frame of video image from the plurality of frames of video images until the movement trend of each obstacle in each frame of video image is determined.
6. The road obstacle perception system for unmanned vehicles based on video surveillance according to claim 4 is characterized by: Determining the consistency of motion features of each obstacle in adjacent frames of video images according to the motion trend of each obstacle in each frame of video images includes: According to the movement trend of each obstacle in each frame of video image, the difference in the movement trend of each obstacle in adjacent frames of video image is determined; The consistency of motion features of each obstacle in adjacent frame video images is determined according to the opposite number of the motion trend difference and the natural exponential function.
7. The road obstacle perception system for unmanned vehicles based on video surveillance according to claim 4 is characterized by: The second determining module includes a seventh unit, an eighth unit, a ninth unit and a tenth unit: The seventh unit is used to determine a fifth frame of video image from a plurality of frames of video image when obstacles in adjacent frames of video image cannot be matched during continuous tracking and detection, and to determine a sixth frame of video image that is a frame previous to the fifth frame of video image; The eighth unit is used to predict the position of each obstacle in the sixth frame of video image in the fifth frame of video image by using a uniform acceleration motion model to obtain a plurality of predicted positions; The ninth unit is used to determine the appearance matching degree mean value corresponding to each obstacle according to the appearance matching degree between each obstacle in the adjacent frame video images corresponding to the fifth frame video image and the sixth frame video image, and respectively determine the appearance matching degree between each obstacle in the adjacent frame video images corresponding to the fifth frame video image and the sixth frame video image, and the first difference value of the appearance matching degree mean value corresponding to each obstacle; The tenth unit is used to determine the occlusion degree of the obstacles at the predicted positions in the fifth frame of video image according to the third product of the first difference value and the consistency of motion features of each obstacle in the adjacent frames of video images corresponding to the fifth frame of video image and the sixth frame of video image, and return to the step of determining the fifth frame of video image from a plurality of frames of video images when obstacles in adjacent frames of video images cannot be matched during continuous tracking and detection, until the occlusion degree of the obstacles at the predicted positions in each frame of video image is determined.
8. The road obstacle perception system for unmanned vehicles based on video surveillance according to claim 7 is characterized by: The perception module includes a first processing unit, a second processing unit and a third processing unit: The first processing unit is used to determine the identity between obstacles in adjacent frame video images according to a fourth product of the motion feature consistency and the appearance matching degree; The second processing unit is used to determine the corrected matching degree of obstacles at each predicted position of each frame of video image and each obstacle in the frame of video image of the previous frame of the frame of video image according to the identity between each obstacle in the adjacent frame of video image and the degree of occlusion of the obstacle; The third processing unit is used to determine the tracking accuracy between obstacles in adjacent frame video images according to the corrected matching degree.
9. The road obstacle perception system for unmanned vehicles based on video surveillance according to claim 8 is characterized by: Determining the corrected matching degree of obstacles at each predicted position of each frame of video image and each obstacle in a frame of video image of a previous frame of video image according to the identity between each obstacle in the adjacent frame of video image and the degree of occlusion of the obstacle comprises: Respectively determine the sum of the preset value and the degree of occlusion of the obstacle at each predicted position of each frame of video image; According to the fifth product of the sum and the identity between the obstacles in the adjacent frame video images, the corrected matching degree between the obstacles at the predicted positions of each frame video image and the obstacles in the frame video image of the previous frame of the frame video image is determined.
10. The road obstacle perception system for unmanned vehicles based on video surveillance according to claim 8, characterized in that: Determining the tracking accuracy between obstacles in adjacent frame video images according to the corrected matching degree includes: According to the corrected matching degree of each obstacle at each predicted position of each frame of video image and each obstacle in the frame of video image before the frame of video image, respectively determine the corrected matching degree mean value corresponding to each obstacle; Determine respectively a second difference value between the corrected matching degree of each obstacle at each predicted position of each frame of video image and each obstacle in the frame of video image of the previous frame of the frame of video image and the corrected matching degree mean value corresponding to each obstacle; The calculation results are determined according to the opposite number of the second difference value and the natural exponential function, and the tracking accuracy between the obstacles in the adjacent frame video images is determined according to the calculation results and the sixth product of the corrected matching degree means corresponding to each obstacle.
Citation Information
Patent Citations
Obstacle tracking method, obstacle tracking device and chip
CN112734811A
Road surface obstacle intelligent identification equipment based on vehicle track
CN113077494A
Path planning method and device for pilotless automobile, electronic equipment and medium
CN114894193A
Unmanned automatic tracking method and device and computer equipment
CN115690741A
Obstacle detection method and device
CN117452411A
Cited By
Vehicle control system and method and intelligent vehicle
CN120552897A