Vehicle early warning method and device, electronic equipment and storage medium
Through example segmentation and depth estimation technology, the target position in the vehicle rearview mirror image is accurately identified, which solves the problem of early warning and misjudgment when turning, and realizes accurate hazard warning when driving non-linear.
Patent Information
- Application Number
- CN202510359592.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
When the rearview mirror of a traditional vehicle is turning or driving in a non-linear direction, it is impossible to accurately identify whether the rear target is in a dangerous area, resulting in early warning and misjudgment.
By obtaining multi-frame driving images collected by the vehicle's electronic rearview mirror, performing instance segmentation and target tracking, calculating the historical average depth of the target mask, and predicting the target position in the next frame of image according to the depth change trend, triggering an early warning prompt.
When the vehicle turns or drives in a non-linear line, accurately identify the lane line area mask, reduce misjudgment of target recognition, and improve the effectiveness and reliability of hazard warnings.
Smart Images

Figure CN120299005A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicles, and in particular, to a vehicle warning method, device, electronic device, and storage medium. Background Art
[0002] Traditional vehicle rearview mirrors are limited by the vehicle body structure and angle, and the view that the driver can see is often limited. With the continuous development of vehicle intelligence, electronic rearview mirrors (Camera Monitor System, CMS) can provide a wider field of view, eliminate blind spots, and help the driver better understand the surrounding environment, thus gradually replacing traditional optical rearview mirrors. An electronic rearview mirror is an indirect vision device that obtains a specified field of view through a system composed of a camera and a monitor. It includes electronic devices such as a high-definition camera, a digital vision processing system, and a liquid crystal display. It captures images of the rear of the vehicle through an external camera and displays them on the in-vehicle display screen after data processing, so as to provide a wider and clearer field of view for the driver.
[0003] Currently, vehicle warning based on rearview mirrors is achieved by performing target detection within a preset area. Specifically, when the vehicle is driving straight, the preset area is usually defined based on the straight lane lines behind the vehicle. The system will detect whether there are vehicles or pedestrians entering the preset area to determine whether a warning is needed. However, in this warning judgment method, when the vehicle is turning or driving non-linearly, the position of the vehicle behind relative to the own vehicle will change. A vehicle that was originally outside the preset area during straight driving may enter the preset area when turning. The system cannot correctly identify whether the target is in a dangerous area, easily generates false judgments, and thus cannot perform accurate danger warnings. Summary of the Invention
[0004] In view of this, the present invention aims to propose a vehicle warning method, device, electronic device, and storage medium to solve the problem that current vehicle rearview mirror warnings are prone to false judgments during target detection within a preset area and cannot perform accurate danger warnings.
[0005] According to the first aspect of the present invention, a vehicle warning method is provided, and the method includes:
[0006] Obtain multiple frames of driving images collected by a vehicle electronic rearview mirror;
[0007] Perform instance segmentation on the driving images to obtain a target mask in each frame of the driving image; wherein, the target mask includes at least one of a target vehicle mask and a target pedestrian mask;
[0008] Perform target tracking on the target mask in the driving image to obtain a tracking identifier corresponding to the target mask;
[0009] Calculate the historical average depth of the target mask in the historical frame driving images according to the tracking identifier corresponding to the target mask, and predict the average depth of the target mask in the next frame driving image according to the change trend of the historical average depth;
[0010] When the average depth of the target mask in the next frame driving image is less than a preset safety threshold, trigger a warning prompt.
[0011] Optionally, perform instance segmentation on the driving images to obtain the target masks in each frame of driving images; wherein, the target masks include at least one of a target vehicle mask and a target pedestrian mask, and include:
[0012] Input the driving images into a pre-trained instance segmentation model for instance segmentation, and output a lane line mask;
[0013] Divide the lane line mask into a self-vehicle lane line area mask and an adjacent lane line area mask;
[0014] Identify the target vehicle mask and / or the target pedestrian mask in the self-vehicle lane line area mask and the adjacent lane line area mask to obtain the target masks in each frame of driving images.
[0015] Optionally, perform target tracking on the target masks in the driving images to obtain the tracking identifier corresponding to the target masks, including:
[0016] Perform image processing on the target masks in the driving images to obtain the circumscribed rectangles of the target masks;
[0017] Perform target tracking on the circumscribed rectangles of the target masks in multiple frames of driving images to obtain the target tracking trajectories in multiple frames of driving images;
[0018] Match the circumscribed rectangles of the target masks and the target tracking trajectories to obtain a matching result, and generate the tracking identifier of the target masks according to the matching result.
[0019] Optionally, match the circumscribed rectangles of the target masks and the target tracking trajectories to obtain a matching result, and generate the tracking identifier of the target masks according to the matching result, including:
[0020] Perform cascade matching on the circumscribed rectangles of the target masks and the target tracking trajectories to obtain a first matching result, unmatched target tracking trajectories and unmatched target masks;
[0021] Perform overlap matching on the unmatched target tracking trajectories and the unmatched target masks to obtain a second matching result;
[0022] Generate a tracking identifier for the target mask according to the first matching result and the second matching result.
[0023] Optionally, the cascaded matching of the circumscribed rectangle of the target mask and the target tracking trajectory to obtain a first matching result, unmatched target tracking trajectories, and unmatched target masks includes:
[0024] Calculate the cosine distance similarity between the circumscribed rectangle of the target mask and the target tracking trajectory to generate a first cost matrix;
[0025] Calculate the Mahalanobis distance between the circumscribed rectangle of the target mask and the target tracking trajectory to generate a second cost matrix;
[0026] Match the first cost matrix and the second cost matrix to obtain a first matching result, unmatched target tracking trajectories, and unmatched target masks.
[0027] Optionally, the calculating the historical average depth of the target mask in the historical frame driving images according to the tracking identifier corresponding to the target mask, and predicting the average depth of the target mask in the next frame driving image according to the change trend of the historical average depth includes:
[0028] Determine the historical frame driving images in which the tracking identifier appears according to the tracking identifier corresponding to the target mask;
[0029] Input the historical frame driving images into a pre-trained monocular depth estimation model, and output the historical average depth of the target mask in the historical frame driving images;
[0030] Adopt the Kalman filter algorithm to determine the change trend of the historical average depth, and predict the average depth of the target mask in the next frame driving image according to the change trend.
[0031] Optionally, before segmenting the driving image into instances to obtain the target mask in each frame of driving image, it further includes:
[0032] Obtain the original image collected by the vehicle electronic rearview mirror, the depth image corresponding to the original image, and the instance label of the original image;
[0033] Input the original image, the depth image, and the instance label into a pre-determined backbone network for alternating training to obtain an instance segmentation model and a monocular depth estimation model.
[0034] According to a second aspect of the present invention, there is provided a vehicle warning device, the device includes:
[0035] An image acquisition module, configured to acquire multiple frames of driving images collected by a vehicle electronic rearview mirror;
[0036] An instance segmentation module, configured to perform instance segmentation on the driving image to obtain a target mask in each frame of the driving image; wherein, the target mask includes at least one of a target vehicle mask and a target pedestrian mask.
[0037] A target tracking module, configured to perform target tracking on the target mask in the driving image to obtain a tracking identifier corresponding to the target mask;
[0038] A depth prediction module, configured to calculate a historical average depth of the target mask in historical frames of the driving image according to the tracking identifier corresponding to the target mask, and predict an average depth of the target mask in the next frame of the driving image according to a change trend of the historical average depth;
[0039] An early warning prompt module, configured to trigger an early warning prompt when the average depth of the target mask in the next frame of the driving image is less than a preset safety threshold.
[0040] According to another aspect of the present invention, there is also provided an electronic device, including:
[0041] A processor;
[0042] A memory for storing executable instructions of the processor;
[0043] Wherein, the processor is configured to execute the instructions to implement the vehicle early warning method as described above.
[0044] According to another aspect of the present invention, there is also provided a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the vehicle early warning method as described above are implemented.
[0045] The vehicle warning method provided by the embodiments of the present invention obtains multiple frames of driving images collected by the vehicle's electronic rearview mirror, performs instance segmentation on the driving images to obtain the target masks in each frame of the driving images, performs target tracking on the target masks in the driving images to obtain the tracking identifiers corresponding to the target masks, calculates the historical average depth of the target masks in the historical frame driving images according to the tracking identifiers corresponding to the target masks, and predicts the average depth of the target masks in the next frame of the driving images according to the change trend of the historical average depth. When the average depth of the target masks in the next frame of the driving images is less than the preset safety threshold, a warning prompt is triggered. By performing instance segmentation and depth estimation, the embodiments of the present invention perform pixel-level segmentation on the rearview mirror images, can accurately identify the lane line area masks when the vehicle turns or drives in a non-linear manner, reduce the misjudgment of target recognition, and use monocular depth estimation to dynamically predict the average depth of the target based on the change trend of the average depth of the target in the image, ensuring the accuracy of the depth information, greatly improving the accuracy of target detection and recognition, and further enhancing the effectiveness and reliability of the vehicle's electronic rearview mirror for hazard warning.
[0046] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically describes the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0048] Figure 1 is a flowchart of the steps of a vehicle warning method provided by the embodiments of the present invention;
[0049] Figure 2 is Figure 1 a flowchart of step 102 in the vehicle warning method provided by the embodiments of the present invention;
[0050] Figure 3 is Figure 1 a flowchart of step 103 in the vehicle warning method provided by the embodiments of the present invention;
[0051] Figure 4 is Figure 1 a flowchart of step 104 in the vehicle warning method provided by the embodiments of the present invention;
[0052] Figure 5It is a flowchart of steps of another vehicle warning method provided by an embodiment of the present invention;
[0053] Figure 6 It is a schematic structural diagram of a vehicle warning device provided by an embodiment of the present invention;
[0054] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0055] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the following will elaborate on various embodiments of the present invention in conjunction with the accompanying drawings. However, those of ordinary skill in the art can understand that in various embodiments of the present invention, many technical details are proposed to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented. The division of the following embodiments is for convenience of description and should not constitute any limitation on the specific implementation manners of the present invention. Various embodiments can be combined with and cited from each other on the premise of not being contradictory.
[0056] Referring to Figure 1 , a flowchart of steps of a vehicle warning method provided by an embodiment of the present invention is shown. The method may include:
[0057] Step 101, acquiring multiple frames of driving images collected by a vehicle electronic rearview mirror.
[0058] In an embodiment of the present invention, to solve the problem that the current vehicle rearview mirror warning performs target detection in a preset area and is prone to misjudgment and unable to perform accurate danger warnings, this embodiment introduces instance segmentation and monocular depth estimation. The instance segmentation technology based on deep learning can accurately identify and mark important elements such as pedestrians, vehicles, and lane lines around the vehicle. Monocular depth estimation can obtain the relative distance between the target and the vehicle in real time, and use Kalman filtering to accurately predict the distance between the target and the vehicle. Thus, when the average distance between vehicles or pedestrians in the own lane and adjacent lanes is less than the safety threshold, a warning prompt is made in a timely manner, and accurate and effective danger warnings are realized based on the vehicle electronic rearview mirror, significantly improving the accuracy and real-time performance of the driving assistance system.
[0059] It should be noted that the execution subject of the embodiment of the present invention is a driving assistance system (ADAS). The ADAS contains one or more processing units inside. The driving assistance system analyzes and judges the driving image data collected in real time by the vehicle electronic rearview mirror during driving, accurately identifies the dangerous targets in the image data, and controls the vehicle to perform danger warnings.
[0060] Specifically, during driving, the electronic rearview mirror captures multiple frames of driving images through a camera. The driving assistance system obtains the multiple frames of driving images captured by the vehicle's electronic rearview mirror. Among them, the driving images are images of the surrounding environment in the area that the electronic rearview mirror camera can capture. The driving images include visual information such as road targets, road types, road signs, and lane lines. The electronic rearview mirror camera can capture the lane lines behind the vehicle, including the lane lines of the vehicle's own lane and adjacent lanes, and can also capture other vehicles driving behind and pedestrians on the rear road. In some cases, the camera will also capture obstacles behind, such as roadblocks, construction signs, parked vehicles, etc. Therefore, the road target can be other vehicles, pedestrians, or any obstacle target, which is not specifically limited here.
[0061] Step 102: Perform instance segmentation on the driving images to obtain the target masks in each frame of driving image; among them, the target masks include at least one of the target vehicle mask and the target pedestrian mask.
[0062] In the embodiment of the present invention, the driving images are input into a pre-trained instance segmentation model for instance segmentation to extract the lane line masks in the driving images. The lane line masks are divided into the self-vehicle lane line area mask and the adjacent lane line area mask, and the target vehicle mask, target pedestrian mask, etc. in the self-vehicle lane line area mask and the adjacent lane line area mask are identified to obtain the target masks in each frame of driving image. The target masks include at least one of the target vehicle mask and the target pedestrian mask, and may also include obstacle targets in the self-vehicle lane line area mask and the adjacent lane line area mask, which is not specifically limited here.
[0063] Step 103: Perform target tracking on the target masks in the driving images to obtain the tracking identifiers corresponding to the target masks.
[0064] It should be noted that in this embodiment, after obtaining the target masks such as the target vehicle mask and the target pedestrian mask in the self-vehicle lane line area mask and the adjacent lane line area mask, according to the bounding rectangles of the target masks and in combination with the tracking algorithm, the tracking identifiers of the target masks can be obtained and paired and recorded in the vehicle and driving.
[0065] Specifically, perform image processing on the target masks in the driving images to obtain the bounding rectangles of the target masks. Perform target tracking on the bounding rectangles of the target masks in multiple frames of driving images to obtain the target tracking trajectories in multiple frames of driving images. Match the bounding rectangles of the target masks and the target tracking trajectories to obtain the matching results, and generate the tracking identifiers of the target masks according to the matching results.
[0066] Step 104: Calculate the historical average depth of the target mask in the historical frame driving images according to the tracking identifier corresponding to the target mask, and predict the average depth of the target mask in the next frame driving image according to the change trend of the historical average depth.
[0067] In the embodiment of the present invention, according to the tracking identifier corresponding to the target mask, the historical frame driving images in which the tracking identifier appears are determined, and the historical frame driving images are input into a pre-trained monocular depth estimation model to output the historical average depth of the target mask in the historical frame driving images. The Kalman filtering algorithm is used to determine the change trend of the historical average depth, and the average depth of the target mask in the next frame driving image is predicted according to the change trend.
[0068] It should be noted that through the tracking identifier of the target mask, that is, the unique tracking ID of the target, the historical frame driving images in which the tracking ID of the target appears are retrieved, and the historical frame driving images are input into a pre-trained monocular depth estimation model to output the historical average depth of the target mask in the historical frame driving images. The monocular depth estimation model is trained based on the backbone network of YOLOv8. The monocular depth estimation model can estimate the depth information of the target mask by reasoning on the historical frame images. The historical average depth is the average depth value of the target in the historical frame driving images. After obtaining the historical average depth of the target mask in the historical frame driving images, the Kalman filtering algorithm is used to determine the change trend of the historical average depth, and the average depth of the target mask in the next frame driving image is predicted according to the change trend.
[0069] Step 105: Trigger a warning prompt when the average depth of the target mask in the next frame driving image is less than a preset safety threshold.
[0070] In this embodiment, if in the self-lane and the adjacent lanes, when the average depth of the target mask in the next frame driving image is less than the preset safety threshold, a warning prompt is made according to the self-vehicle state. Among them, the warning prompt methods include light warning, sound prompt, and display prompt. Specifically, if the self-vehicle is driving normally and a vehicle or pedestrian is detected on the left or right side and meets the distance condition, a light warning can be made on the corresponding electronic rearview mirror, and the segmentation mask of the dangerous target is added to the display screen, which is more conducive to the user's judgment; if the self-vehicle turns on the left turn signal or the right turn signal, or the steering wheel has left or right operations, it can be considered that the vehicle has a tendency to change lanes to the left or right. If there are vehicles or pedestrians in the corresponding direction that meet the distance condition, in addition to the light warning and the segmented screen, a prompt sound message is additionally issued. If the user's vehicle speed is close to 0 or in the parking state, the distance conditions of the vehicles and pedestrians in both directions are continuously judged, and if the prompt situation is met, a light warning, a segmented screen, and a prompt sound message need to be issued to avoid collision accidents when opening the door.
[0071] The vehicle warning method provided in the embodiments of the present invention obtains multiple frames of driving images collected by the vehicle's electronic rearview mirror, performs instance segmentation on the driving images to obtain the target masks in each frame of the driving images, performs target tracking on the target masks in the driving images to obtain the tracking identifiers corresponding to the target masks, calculates the historical average depth of the target masks in the historical frame driving images according to the tracking identifiers corresponding to the target masks, and predicts the average depth of the target masks in the next frame of the driving images according to the change trend of the historical average depth. When the average depth of the target masks in the next frame of the driving images is less than the preset safety threshold, a warning prompt is triggered. The embodiments of the present invention perform pixel-level segmentation on the rearview mirror images through instance segmentation and depth estimation, can accurately identify the lane line area masks when the vehicle turns or drives in a non-linear manner, reduce the misjudgment of target recognition, and use monocular depth estimation to dynamically predict the average depth of the target based on the change trend of the average depth of the target in the image, ensuring the accuracy of the depth information, greatly improving the accuracy of target detection and recognition, and further enhancing the effectiveness and reliability of the vehicle's electronic rearview mirror for danger warning.
[0072] Further, referring to Figure 2 , shows Figure 1 The flowchart of step 102 in a vehicle warning method provided, this method is basically the same as the vehicle warning method provided in the first embodiment of the present invention, and step 102 may include:
[0073] Step 201, input the driving image into a pre-trained instance segmentation model for instance segmentation, and output the lane line mask.
[0074] Step 202, divide the lane line mask into the self-vehicle lane line area mask and the adjacent lane line area mask.
[0075] Step 203, identify the target vehicle mask and / or the target pedestrian mask in the self-vehicle lane line area mask and the adjacent lane line area mask, and obtain the target masks in each frame of the driving image.
[0076] It should be noted that in the embodiments of the present invention, the driving image is input into a pre-trained instance segmentation model for instance segmentation, and the lane line mask is output. The driving image collected by the electronic rearview mirror camera is usually an RGB image. The instance segmentation model can accurately identify the position and shape of the lane line by performing pixel-level segmentation on the image. The lane line mask output by the instance segmentation model includes the lane lines of the lane where the self-vehicle is located and the lane lines of the adjacent lanes. Among them, the instance segmentation model is pre-trained based on the backbone network of YOLOv8, which will not be elaborated here.
[0077] Specifically, according to the position and geometric relationship of the lane line mask, the lane line mask is divided into the ego-vehicle lane line area and the adjacent lane line area. Among them, the ego-vehicle lane line area mask represents the lane line area of the lane where the ego-vehicle is located, and the adjacent lane line area mask represents the lane line areas adjacent to the ego-vehicle (such as the left lane and the right lane). The divided mask areas can be used to calculate whether the target vehicle and pedestrians are in the ego-vehicle lane or adjacent lanes. By dividing the lane line mask, the ego-vehicle lane and adjacent lanes can be clearly distinguished, providing a basis for subsequent target recognition and warning. In the ego-vehicle lane line area mask and the adjacent lane line area mask, the target vehicle mask and the target pedestrian mask are recognized to obtain the target masks in each frame of driving image. The target masks include the pixel-level segmentation results of vehicles and pedestrians.
[0078] It should be noted that the instance segmentation model generates a segmentation mask by performing pixel-level classification on the input image. In this embodiment, the segmentation mask obtained by instance segmentation is a label matrix of each pixel in the image, which is used to represent the position and shape of each target instance in the image, has pixel-level accuracy, and can accurately describe the contour of the target. Specifically, the segmentation mask output by the model is usually a two-dimensional matrix, and each value in the matrix corresponds to a pixel in the image. The segmentation mask marks each target (such as a vehicle, a pedestrian) in the image with different label values. Through the label values of different targets in the segmentation mask, the target vehicle mask and the target pedestrian mask in the ego-vehicle lane line area mask and the adjacent lane line area mask can be accurately identified.
[0079] The embodiment of the present invention uses an instance segmentation model to perform pixel-level segmentation on the image, accurately identify lane lines, vehicles and pedestrians, accurately identify vehicles and pedestrians in the ego-vehicle lane and adjacent lanes, provide more accurate warning information for the driver, and can reduce the misjudgment of target recognition and improve the reliability of warning through the dynamic adjustment of the lane line area mask when the vehicle turns or drives non-linearly.
[0080] Further, referring to Figure 3 shows Figure 1 the flowchart of step 103 in a vehicle warning method provided, which is basically the same as the vehicle warning method provided in the first embodiment of the present invention. Step 103 may include:
[0081] Step 301, perform image processing on the target mask in the driving image to obtain the bounding rectangle of the target mask.
[0082] Step 302, perform target tracking on the bounding rectangles of the target masks in multiple frames of driving images to obtain the target tracking trajectories in multiple frames of driving images.
[0083] Step 303: Match the circumscribed rectangle of the target mask and the target tracking trajectory to obtain a matching result, and generate a tracking identifier for the target mask according to the matching result.
[0084] It should be noted that in the embodiments of the present invention, image processing is performed on the target mask in the driving image to extract the circumscribed rectangle of the target mask. Among them, the circumscribed rectangle of the target mask is the minimum circumscribed rectangle frame of the target mask. The circumscribed rectangle frame contains the position, size, and direction information of the target. The extraction of the circumscribed rectangle frame provides basic data for target tracking, facilitating subsequent trajectory matching and tracking identifier generation.
[0085] Specifically, perform target tracking on the circumscribed rectangles of the target masks in multiple frames of driving images to obtain the target tracking trajectories in multiple frames of driving images. In this embodiment, based on the circumscribed rectangles of the target masks in multiple frames of driving images, use a target tracking algorithm to track the target in multiple frames of images and generate the tracking trajectory of the target. Among them, the target tracking trajectory in multiple frames of driving images is the movement path of the target in consecutive frames. It should be noted that the target tracking algorithm can continuously track the target in multiple frames of images by combining the circumscribed rectangle frame of the target mask and feature information (such as RGB features). The tracking trajectory includes information such as the position and movement direction of the target.
[0086] In this embodiment, match the circumscribed rectangle of the target mask and the target tracking trajectory to obtain a matching result, and generate a tracking identifier for the target mask according to the matching result. In some embodiments, jump detection can also be performed on the target tracking trajectory, mainly performing effective jump detection on the target pedestrian mask. If a target jump occurs in the target tracking trajectory, the target tracking trajectory will be deleted, that is, if it does not meet the trajectory state, further target tracking matching will not be performed, and no specific limitation is made here.
[0087] Specifically, first perform cascade matching on the circumscribed rectangle of the target mask and the target tracking trajectory to obtain a first matching result, unmatched target tracking trajectories, and unmatched target masks, and then perform overlap matching (IOU matching) on the unmatched target tracking trajectories and target masks to obtain a second matching result. According to the first matching result and the second matching result, lock the same target appearing in different frames of driving images and generate a tracking identifier for the target mask. Among them, the tracking identifier of the target mask is the unique ID of the target to be detected. The tracking identifier is used to uniquely identify the target, facilitating subsequent trajectory management and early warning logic judgment.
[0088] The embodiments of the present invention complete target tracking and generate a unique tracking identifier under the requirement of real-time performance by accurately matching the target mask and the tracking trajectory, facilitating subsequent trajectory management and early warning logic judgment.
[0089] Specifically, in step 303, the circumscribed rectangle of the target mask and the target tracking trajectory are matched to obtain a matching result, and a tracking identifier of the target mask is generated according to the matching result. Specifically, it may include:
[0090] Perform cascade matching on the circumscribed rectangle of the target mask and the target tracking trajectory to obtain a first matching result, unmatched target tracking trajectories, and unmatched target masks;
[0091] Perform overlap matching on the unmatched target tracking trajectories and the unmatched target masks to obtain a second matching result;
[0092] Generate a tracking identifier of the target mask according to the first matching result and the second matching result.
[0093] In the embodiments of the present invention, in the above steps, first, cascade matching is performed on the circumscribed rectangle of the target mask and the target tracking trajectory to calculate the feature similarity between the circumscribed rectangle of the target mask and the target tracking trajectory, and the Hungarian algorithm is used to match the feature similarity to obtain a first matching result, unmatched target tracking trajectories, and unmatched target masks. Then, overlap matching is performed on the unmatched target tracking trajectories and the unmatched target masks to calculate the IOU distance between the unmatched target tracking trajectories and the unmatched target masks, and the Hungarian algorithm is used to match the IOU distance to obtain a second matching result. The IOU distance is used to measure the overlap degree between the circumscribed rectangle of the target mask and the target tracking trajectory.
[0094] In the embodiments of the present invention, through the combination of cascade matching and overlap matching, the overlap matching is used to process the unmatched targets in the cascade matching, further improving the coverage rate of the matching, and being able to efficiently match the target mask with the tracking trajectory to ensure the accuracy of target tracking.
[0095] Specifically, the step of performing cascade matching on the circumscribed rectangle of the target mask and the target tracking trajectory to obtain a first matching result, unmatched target tracking trajectories, and unmatched target masks may include:
[0096] First, calculate the cosine distance similarity between the circumscribed rectangle of the target mask and the target tracking trajectory to generate a first cost matrix;
[0097] Second, calculate the Mahalanobis distance between the circumscribed rectangle of the target mask and the target tracking trajectory to generate a second cost matrix;
[0098] Second, match the first cost matrix and the second cost matrix to obtain a first matching result, unmatched target tracking trajectories, and unmatched target masks.
[0099] In this embodiment, by calculating the cosine distance similarity between the circumscribed rectangle of the target mask and the target tracking trajectory, a first cost matrix is generated. The cosine distance similarity is used to measure the difference between two vectors, that is, to measure the feature similarity between the circumscribed rectangle of the target mask and the target tracking trajectory. The cosine distance reflects their similarity by calculating the cosine value of the angle between two vectors. Feature vectors are extracted from the circumscribed rectangle of the target mask and from the target tracking trajectory. The feature vectors include the label value and geometric features of the target. The cosine distance between the two feature vectors is calculated to obtain the first cost matrix composed of the cosine distances between the circumscribed rectangle of the target mask and the target tracking trajectory.
[0100] Specifically, the Mahalanobis distance between the circumscribed rectangle of the target mask and the target tracking trajectory is calculated to generate a second cost matrix. The Mahalanobis distance is used to measure the position similarity between the circumscribed rectangle of the target mask and the target tracking trajectory. Taking the first cost matrix and the second cost matrix as inputs, the Hungarian algorithm is used for matching to obtain the first matching result, the unmatched target tracking trajectories, and the unmatched target masks.
[0101] In the embodiment of the present invention, through the combination of cascade matching and overlap degree matching, the target mask can be efficiently matched with the tracking trajectory to ensure the accuracy of target tracking.
[0102] Further, referring to Figure 4 shows Figure 1 the flowchart of step 104 in a vehicle warning method provided by
[0103] Step 401: Determine the historical frame driving images in which the tracking identifier appears according to the tracking identifier corresponding to the target mask.
[0104] Step 402: Input the historical frame driving images into a pre-trained monocular depth estimation model, and output the historical average depth of the target mask in the historical frame driving images.
[0105] Step 403: Use the Kalman filter algorithm to determine the change trend of the historical average depth, and predict the average depth of the target mask in the next frame of driving image according to the change trend.
[0106] It should be noted that in the embodiment of the present invention, according to the tracking identifier corresponding to the target mask, the historical frame driving images in which the tracking identifier appears are determined. By the tracking identifier of the target mask, that is, the unique tracking ID of the target, the historical frame driving images in which the tracking ID of the target appears are retrieved. The historical frame driving images include the position, size, and depth information of the target at different time points. Through the tracking identifier, the driving images in which the target appears can be quickly located, providing a basis for depth estimation.
[0107] Specifically, the historical frame driving images are input into a pre-trained monocular depth estimation model to output the historical average depth of the target mask in the historical frame driving images. The monocular depth estimation model is trained based on the backbone network of YOLOv8. The backbone network is designed with a monocular depth estimation head. During training, the RGB images of the electronic rearview mirror and the aligned depth images are used. The monocular depth estimation model can estimate the depth information of the target mask by inferring the historical frame images. The historical average depth is the average depth value of the target in the historical frame driving images. After obtaining the historical average depth of the target mask in the historical frame driving images, the Kalman filter algorithm is used to determine the change trend of the historical average depth, and the average depth of the target mask in the next frame of driving images is predicted according to the change trend.
[0108] It should be noted that the Kalman filter algorithm is divided into a prediction stage and an update stage. In the prediction stage, the average depth of the target mask in the next frame is predicted based on the historical average depth. In the update stage, the prediction result is updated according to the actual observation value, that is, the average depth value of the current frame, to obtain the predicted average depth of the target mask in the next frame of driving images. By combining historical data and current observations, the Kalman filter algorithm can effectively filter noise and improve the accuracy of depth estimation.
[0109] In this embodiment, the specific process of the Kalman filter algorithm is divided into a prediction stage and an update stage. The prediction stage includes state prediction and error covariance prediction. The specific calculation formulas are as follows:
[0110] State prediction:
[0111]
[0112] Error covariance prediction:
[0113]
[0114] In the update stage, the Kalman gain is calculated:
[0115]
[0116] State update:
[0117]
[0118] Error covariance update:
[0119] P k|k-1 =(I - K k H k )P k|k-1 (Formula 5)
[0120] Among them, Equation 1 is for state prediction, Equation 1 is for error covariance prediction, Equation 3 is for Kalman gain, Equation 4 is for state update, and Equation 5 is for error covariance update. is the prior state estimate at time k. is the posterior state estimate at time k, P k|k-1 is the prior estimate error covariance, P k|k is the posterior estimate error covariance, K k is the Kalman gain, which is used to balance the credibility of the predicted value and the measured value, A k is the state transition matrix, B k is the control matrix, u k is the control input, Q k is the process noise covariance, H k is the observation matrix, R k is the observation noise covariance, z k is the observation value at time k.
[0121] In the embodiment of the present invention, through the monocular depth estimation model, the depth information of the target mask can be accurately estimated. By calculating the average depth in the historical frames, the distance change trend of the target is reflected. The Kalman filter algorithm combines historical data and current observation values, dynamically adjusts the prediction result, adapts to the movement change of the target, and quickly and accurately predicts the average depth of the target in the next frame, improving the accuracy of depth estimation.
[0122] Refer to Figure 5 , which shows the step flow chart of another vehicle warning method provided by the embodiment of the present invention. This method is basically the same as the vehicle warning method provided by the first embodiment of the present invention. The difference is that the method may further include:
[0123] Step 106, obtain the original image collected by the vehicle electronic rearview mirror, the depth image corresponding to the original image, and the instance label of the original image.
[0124] Step 107, input the original image, the depth image, and the instance label into the pre-determined backbone network for alternating training to obtain an instance segmentation model and a monocular depth estimation model.
[0125] In the embodiments of the present invention, the original image collected by the vehicle electronic rearview mirror, the depth image corresponding to the original image, and the instance labels of the original image are obtained in advance. The instance labels are respectively: the instance segmentation labels of lane lines, pedestrians, and vehicles. The original image is the RGB image in the electronic rearview mirror screen, and the depth image is the depth Depth image aligned with the original image. Training is carried out based on the backbone network of YOLOV8. The backbone network is designed with an instance segmentation head and a monocular depth estimation head. For the backbone network, each round of input image data participates in training. For the instance segmentation network head and the monocular depth estimation network head, they are alternately trained every other round. That is, in each round of training, the instance segmentation network head is trained using the original image and the instance labels, and another parameter is frozen. In the next round, the monocular depth estimation head is trained using the depth image corresponding to the original image. The original image, the depth image, and the instance labels are input into the pre-determined backbone network for alternating training to obtain an instance segmentation model and a monocular depth estimation model.
[0126] Specifically, for example, 130,000 segmentation training data and 40,000 monocular depth estimation data are collected, and a total of 500 loop iterations are designed for model training. Among them, the YOLOV8 backbone network is loaded. The instance segmentation head is trained in odd training cycles, and the monocular depth estimation head is trained in even training cycles. The backbone network participates in backpropagation in all training processes to obtain an instance segmentation model and a monocular depth estimation model.
[0127] Step 102, perform instance segmentation on the driving image to obtain the target mask in each frame of the driving image; wherein, the target mask includes at least one of a target vehicle mask and a target pedestrian mask.
[0128] Step 103, perform target tracking on the target mask in the driving image to obtain the tracking identifier corresponding to the target mask.
[0129] Step 104, calculate the historical average depth of the target mask in the historical frame driving image according to the tracking identifier corresponding to the target mask, and predict the average depth of the target mask in the next frame of the driving image according to the change trend of the historical average depth.
[0130] Step 105, trigger a warning prompt when the average depth of the target mask in the next frame of the driving image is less than the preset safety threshold.
[0131] The above steps 102 to 105 are as described in the previous section and will not be elaborated here.
[0132] Compared with the prior art, the embodiments of the present invention, on the basis of achieving the beneficial effects brought by the first embodiment, utilize the design idea and training method of deep learning instance segmentation and monocular estimation network. Under the same backbone network, it can greatly save the computing power consumed during vehicle operation. According to the idea of frozen training, two tasks can be effectively trained to further improve the data processing efficiency.
[0133] Referring to Figure 6 , a schematic structural diagram of a vehicle warning device provided by an embodiment of the present invention is shown. The device includes:
[0134] An image acquisition module 501, configured to acquire multiple frames of driving images collected by a vehicle electronic rearview mirror;
[0135] An instance segmentation module 502, configured to perform instance segmentation on the driving image to obtain a target mask in each frame of driving image; wherein, the target mask includes at least one of a target vehicle mask and a target pedestrian mask.
[0136] A target tracking module 503, configured to perform target tracking on the target mask in the driving image to obtain a tracking identifier corresponding to the target mask;
[0137] A depth prediction module 504, configured to calculate the historical average depth of the target mask in the historical frame driving image according to the tracking identifier corresponding to the target mask, and predict the average depth of the target mask in the next frame of driving image according to the change trend of the historical average depth;
[0138] A warning prompt module 505, configured to trigger a warning prompt when the average depth of the target mask in the next frame of driving image is less than a preset safety threshold.
[0139] Further, the instance segmentation module 502 includes:
[0140] A segmentation sub-module, configured to input the driving image into a pre-trained instance segmentation model for instance segmentation, and output a lane line mask;
[0141] A division sub-module, configured to divide the lane line mask into a self-vehicle lane line area mask and an adjacent lane line area mask;
[0142] An identification sub-module, configured to identify the target vehicle mask and / or the target pedestrian mask in the self-vehicle lane line area mask and the adjacent lane line area mask, and obtain the target mask in each frame of driving image.
[0143] Further, the target tracking module 503 includes:
[0144] A processing sub-module for performing image processing on the target mask in the driving image to obtain an external rectangle of the target mask;
[0145] A tracking sub-module for performing target tracking on the external rectangles of the target masks in multiple frames of driving images to obtain the target tracking trajectories in the multiple frames of driving images;
[0146] A matching sub-module for matching the external rectangle of the target mask and the target tracking trajectory to obtain a matching result, and generating a tracking identifier for the target mask according to the matching result.
[0147] Further, the matching sub-module includes:
[0148] A first matching unit for performing cascade matching on the external rectangle of the target mask and the target tracking trajectory to obtain a first matching result, unmatched target tracking trajectories, and unmatched target masks;
[0149] A second matching unit for performing overlap matching on the unmatched target tracking trajectories and the unmatched target masks to obtain a second matching result;
[0150] A generating unit for generating a tracking identifier for the target mask according to the first matching result and the second matching result.
[0151] Further, the first matching unit includes:
[0152] A first calculation sub-unit for calculating the cosine distance similarity between the external rectangle of the target mask and the target tracking trajectory to generate a first cost matrix;
[0153] A second calculation sub-unit for calculating the Mahalanobis distance between the external rectangle of the target mask and the target tracking trajectory to generate a second cost matrix;
[0154] A matching sub-unit for matching the first cost matrix and the second cost matrix to obtain a first matching result, unmatched target tracking trajectories, and unmatched target masks.
[0155] Further, the depth prediction module 504 includes:
[0156] A determination sub-module for determining the historical frame driving images in which the tracking identifier appears according to the tracking identifier corresponding to the target mask;
[0157] A depth estimation sub-module for inputting the historical frame driving images into a pre-trained monocular depth estimation model and outputting the historical average depth of the target mask in the historical frame driving images;
[0158] A predictor sub-module for determining the change trend of the historical average depth using the Kalman filtering algorithm and predicting the average depth of the target mask in the next frame of driving image according to the change trend.
[0159] Further, the device further includes:
[0160] A sample acquisition module for acquiring the original image collected by the vehicle electronic rearview mirror, the depth image corresponding to the original image, and the instance label of the original image;
[0161] A model training module for inputting the original image, the depth image, and the instance label into a pre-determined backbone network for alternating training to obtain an instance segmentation model and a monocular depth estimation model.
[0162] The vehicle warning device provided by the embodiment of the present invention obtains multiple frames of driving images collected by the vehicle electronic rearview mirror, performs instance segmentation on the driving images to obtain the target mask in each frame of driving image, performs target tracking on the target mask in the driving image to obtain the tracking identifier corresponding to the target mask, calculates the historical average depth of the target mask in the historical frame of driving image according to the tracking identifier corresponding to the target mask, and predicts the average depth of the target mask in the next frame of driving image according to the change trend of the historical average depth. When the average depth of the target mask in the next frame of driving image is less than the preset safety threshold, a warning prompt is triggered. The embodiment of the present invention performs pixel-level segmentation on the rearview mirror image through instance segmentation and depth estimation, can accurately identify the lane line area mask when the vehicle turns or drives non-linearly, reduces the misjudgment of target recognition, and uses monocular depth estimation to dynamically predict the average depth of the target based on the change trend of the average depth of the target in the image, ensuring the accuracy of the depth information, greatly improving the accuracy of target detection and recognition, and further enhancing the effectiveness and reliability of the vehicle electronic rearview mirror for hazard warning.
[0163] Refer to Figure 7 , the embodiment of the present invention also provides an electronic device, as Figure 7 shown, including a processor 601, a communication interface 602, a memory 603, and a communication bus 604. Among them, the processor 601, the communication interface 602, and the memory 603 complete mutual communication through the communication bus 604.
[0164] The processor 601, and the memory 603 for storing processor-executable instructions;
[0165] Among them, the processor 601 is configured to execute the instructions to implement the vehicle warning method described below:
[0166] Acquire multiple frames of driving images collected by the vehicle electronic rearview mirror;
[0167] Perform instance segmentation on the driving images to obtain target masks in each frame of driving images; wherein, the target masks include at least one of a target vehicle mask and a target pedestrian mask;
[0168] Perform target tracking on the target masks in the driving images to obtain tracking identifiers corresponding to the target masks;
[0169] Calculate the historical average depth of the target masks in historical frame driving images according to the tracking identifiers corresponding to the target masks, and predict the average depth of the target masks in the next frame of driving images according to the change trend of the historical average depth;
[0170] Trigger a warning prompt when the average depth of the target masks in the next frame of driving images is less than a preset safety threshold.
[0171] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0172] The communication interface is used for communication between the above terminal and other devices.
[0173] The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0174] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0175] In another embodiment provided by the present invention, a computer-readable storage medium is further provided. A computer program is stored on the readable storage medium. When the computer program is executed by a processor, the vehicle warning method described in any one of the above embodiments is implemented.
[0176] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0177] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0178] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the corresponding description in the method embodiment.
[0179] The above are only the preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.
Claims
1. A vehicle warning method, characterized in that, The method includes: Obtaining multiple frames of driving images collected by a vehicle's electronic rearview mirror; Performing instance segmentation on the driving images to obtain a target mask in each frame of the driving image; wherein, the target mask includes at least one of a target vehicle mask and a target pedestrian mask; Performing target tracking on the target mask in the driving image to obtain a tracking identifier corresponding to the target mask; Calculating the historical average depth of the target mask in the historical frame driving images according to the tracking identifier corresponding to the target mask, and predicting the average depth of the target mask in the next frame of driving image according to the change trend of the historical average depth; Triggering a warning prompt when the average depth of the target mask in the next frame of driving image is less than a preset safety threshold.
2. The method according to claim 1, characterized in that, The performing instance segmentation on the driving image to obtain a target mask in each frame of the driving image; wherein, the target mask includes at least one of a target vehicle mask and a target pedestrian mask, includes: Inputting the driving image into a pre-trained instance segmentation model for instance segmentation, and outputting a lane line mask; Dividing the lane line mask into a self-vehicle lane line area mask and an adjacent lane line area mask; Identifying the target vehicle mask and / or the target pedestrian mask in the self-vehicle lane line area mask and the adjacent lane line area mask to obtain the target mask in each frame of the driving image.
3. The method according to claim 1, wherein The performing target tracking on the target mask in the driving image to obtain a tracking identifier corresponding to the target mask, includes: Performing image processing on the target mask in the driving image to obtain an external rectangle of the target mask; Performing target tracking on the external rectangles of the target masks in multiple frames of driving images to obtain a target tracking trajectory in multiple frames of driving images; Matching the external rectangle of the target mask and the target tracking trajectory to obtain a matching result, and generating a tracking identifier of the target mask according to the matching result.
4. The method according to claim 3, wherein The matching the external rectangle of the target mask and the target tracking trajectory to obtain a matching result, and generating a tracking identifier of the target mask according to the matching result, includes: Performing cascade matching on the external rectangle of the target mask and the target tracking trajectory to obtain a first matching result, unmatched target tracking trajectories, and unmatched target masks; Performing overlap matching on the unmatched target tracking trajectories and the unmatched target masks to obtain a second matching result; Generating a tracking identifier of the target mask according to the first matching result and the second matching result.
5. The method according to claim 4, wherein The performing cascade matching on the external rectangle of the target mask and the target tracking trajectory to obtain a first matching result, unmatched target tracking trajectories, and unmatched target masks, includes: Calculating the cosine distance similarity between the external rectangle of the target mask and the target tracking trajectory to generate a first cost matrix; Calculating the Mahalanobis distance between the external rectangle of the target mask and the target tracking trajectory to generate a second cost matrix; Matching the first cost matrix and the second cost matrix to obtain a first matching result, unmatched target tracking trajectories, and unmatched target masks.
6. The method according to claim 1, characterized in that, Calculating the historical average depth of the target mask in the historical frame driving images according to the tracking identifier corresponding to the target mask, and predicting the average depth of the target mask in the next frame driving image according to the change trend of the historical average depth, includes: Determining the historical frame driving images in which the tracking identifier appears according to the tracking identifier corresponding to the target mask; Inputting the historical frame driving images into a pre-trained monocular depth estimation model, and outputting the historical average depth of the target mask in the historical frame driving images; Using the Kalman filtering algorithm to determine the change trend of the historical average depth, and predicting the average depth of the target mask in the next frame driving image according to the change trend.
7. The method according to claim 1, characterized in that, Before segmenting the driving images into instance masks to obtain the target masks in each frame of driving images, further includes: Obtaining the original images collected by the vehicle electronic rearview mirror, the depth images corresponding to the original images, and the instance labels of the original images; Inputting the original images, the depth images, and the instance labels into a pre-determined backbone network for alternating training to obtain an instance segmentation model and a monocular depth estimation model.
8. A vehicle warning device, characterized in that, The device includes: An image acquisition module, configured to acquire multiple frames of driving images collected by the vehicle electronic rearview mirror; An instance segmentation module, configured to perform instance segmentation on the driving images to obtain the target masks in each frame of driving images; wherein, the target masks include at least one of a target vehicle mask and a target pedestrian mask. A target tracking module, configured to perform target tracking on the target masks in the driving images to obtain the tracking identifiers corresponding to the target masks; A depth prediction module, configured to calculate the historical average depth of the target mask in the historical frame driving images according to the tracking identifier corresponding to the target mask, and predict the average depth of the target mask in the next frame driving image according to the change trend of the historical average depth; An early warning prompt module, configured to trigger an early warning prompt when the average depth of the target mask in the next frame driving image is less than a preset safety threshold.
9. An electronic device, characterized in that, Includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the vehicle early warning method according to any one of claims 1 to 7.
10. A readable storage medium, characterized in that, A computer program is stored on the readable storage medium, and when the computer program is executed by the processor, it implements the vehicle early warning method according to any one of claims 1 to 7.