Distance prediction method, model training method, planning and control system, and related devices thereof
The distance prediction method and model training approach address the high costs of updating high-precision maps by predicting the farthest reachable distance on each lane using a distance training sample set and landmark perception model, facilitating efficient and cost-effective route planning and intelligent driving.
Patent Information
- Application Number
- JP2024569655
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-23
- Filing Date
- 2023-05-19
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-05-19
AI Technical Summary
The high economic and time costs associated with frequent updates of high-precision maps for intelligent driving systems, which are necessary for accurate route planning and safe operation.
A distance prediction method and model training approach that allows for the prediction of the farthest reachable distance on each lane without relying on high-precision maps, using a distance training sample set and a landmark perception model to process global poses, extended navigation routes, and landmark information.
Enables efficient and cost-effective route planning and intelligent driving plan control by automatically predicting the farthest reachable distance on each lane, reducing the need for frequent high-precision map updates.
Smart Images

Figure 2025516993000001_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent driving technology, and specifically to a distance prediction method, a model training method, a planning control system, and related devices thereof.
Background Art
[0002] An intelligent driving system is a complex system that combines hardware and software, and the system includes multiple modules such as sensor integration, environmental perception, prediction, and planning control. A high-precision map, as an important support for the intelligent driving system, can provide information with higher accuracy and richer details compared to conventional navigation maps, thereby realizing the commercial application of intelligent driving technology. For example, in the process of planning control, the intelligent driving system performs route planning based on the high-precision map and determines whether the host vehicle continues to move forward based on the current lane or performs a lane change. However, the collection and creation of high-precision maps are very complex. In order to ensure the "freshness" of high-precision maps and meet the needs of safe use of intelligent driving, high-precision maps need to be updated frequently, and the workload of data collection is huge, requiring a large amount of investment in time, personnel, and materials. Therefore, the economic cost and time cost required for planning control based on high-precision maps are high.
Summary of the Invention
Problems to be Solved by the Invention
[0003] This application provides a distance prediction method, a model training method, a planning control system, and related devices thereof that can solve the problem of high economic cost and time cost required for planning control based on high-precision maps.
Means for Solving the Problems
[0004] The specific technical solutions are as follows.
[0005] In the first aspect, the embodiments of the present application provide a distance prediction method, and the method includes: obtaining a first global pose, a first extended navigation route, and first landmark information of a first vehicle; processing the first global pose, the first extended navigation route, and the first landmark information based on a distance prediction model, and obtaining the farthest reachable distance on each lane in the first landmark information; wherein the first global pose is the current global pose of the first vehicle, the first extended navigation route includes an extended route of a first in-vehicle navigation route determined based on the first global pose, and the first landmark information includes landmark information included in a first road environment image collected by the first vehicle at the first global pose; The training method of the distance prediction model includes: obtaining a distance training sample set; training using the distance training sample set to obtain the distance prediction model; each training sample in the distance training sample set includes a second extended navigation route, second landmark information, a second global pose, and a ground truth value of the farthest reachable distance, the second extended navigation route includes an extended route of a second in-vehicle navigation route, the second landmark information includes landmark information included in a second road environment image collected by a second vehicle on the second extended navigation route, the second global pose includes the global pose when the second vehicle collects the second road environment image, and the ground truth value of the farthest reachable distance includes the ground truth value of the farthest reachable distance for each lane in the second landmark information.
[0006] According to the above technical means, in the embodiment of the present application, first, a distance prediction model is trained and obtained based on a distance training sample set including a second extended navigation route, second landmark information, a second global pose, and a truth value of the farthest reachable distance for each lane in the second landmark information. Next, the first global pose, the first extended navigation route, and the first landmark information are input into the distance prediction model, and the farthest reachable distance on each lane in the first landmark information can be predicted. Since the second global pose, the second extended navigation route, the second landmark information required during model training, and the first global pose, the first extended navigation route, the first landmark information required during model application do not depend on a high-precision map, the embodiment of the present application can automatically train a distance prediction model capable of predicting the farthest reachable distance on each lane in front of the vehicle in a state of getting out of the frequent update of the high-precision map, and automatically predict the farthest reachable distance on each lane in front of the vehicle using the distance prediction model. As a result, it becomes easy to quickly perform route planning according to the farthest reachable distance on each lane in front of the vehicle. Furthermore, while saving economic costs and time costs, intelligent driving plan control can be realized.
[0007] In a first possible implementation form of the first aspect, when the target extended navigation route includes the first extended navigation route and / or the second extended navigation route, the method for obtaining the target extended navigation route is as follows. Obtaining a target in-vehicle navigation route and a target global pose when the target vehicle is traveling on the target in-vehicle navigation route; Extracting POI (Point of Interest) information corresponding to the target global pose from target data; Adding the POI information to the target in-vehicle navigation route to obtain the target extended navigation route. When the target extended navigation route is the first extended navigation route, the target in-vehicle navigation route is the first in-vehicle navigation route, the target vehicle is the first vehicle, and the target global pose is the first global pose; when the target extended navigation route is the second extended navigation route, the target in-vehicle navigation route is the second in-vehicle navigation route, and the target global pose is the second global pose. The target data includes navigation events and / or an in-vehicle navigation map. The POI information includes road attribute information related to the prediction of the farthest reachable distance. The POI information corresponding to the target global pose includes the POI information within a predetermined distance range in front of the target global pose on the target in-vehicle navigation route.
[0008] According to the above technical means, the embodiments of the present application can extract the POI information corresponding to the second global pose from navigation events and / or an in-vehicle navigation map, add the extracted POI information to the second in-vehicle navigation route, and obtain a second extended navigation route that contributes to the training of a distance prediction model with higher quality than the second in-vehicle navigation route. Also, the embodiments of the present application can extract the POI information corresponding to the first global pose from navigation events and / or an in-vehicle navigation map, add the extracted POI information to the first in-vehicle navigation route, and obtain a second extended navigation route that contributes to the prediction of the farthest reachable distance more than the first in-vehicle navigation route. Furthermore, the acquisition of the first extended navigation route and the second extended navigation route only needs to be performed according to navigation events and / or an in-vehicle navigation map without relying on a high-precision map.
[0009] In a second possible implementation of the first aspect, when the target landmark information includes the first landmark information and / or the second landmark information, the method for obtaining the target landmark information included in the target road environment image collected by the target vehicle at the target time point is as follows: obtaining input data of a landmark perception model; processing the input data based on the landmark perception model to obtain the target landmark information included in the target road environment image at the target time point, wherein the input data includes the target road environment image at the target time point and the target local pose at the target time point, the target local pose includes an offset amount with respect to the global pose at the starting point of the target of the global pose of the target vehicle when the target road environment image is collected, the target time point is any time point when the target vehicle is collecting the target road environment image on the target extended navigation route, when the target landmark information is the first landmark information, the target vehicle is the first vehicle, the target road environment image is the first road environment image, the target extended navigation route is the first extended navigation route, and the target local pose is the first local pose; when the target landmark information is the second landmark information, the target vehicle is the second vehicle, the target road environment image is the second road environment image, the target extended navigation route is the second extended navigation route, and the target local pose is the second local pose.
[0010] According to the above technical means, the embodiments of the present application can not only automatically sense the target landmark information included in the target road environment image at any time point based on the pre-trained landmark perception model, but also, since the input data of the landmark perception model is independent of the high-precision map and only includes the target road environment image and the target local pose, the target landmark information can also be obtained in a state of being extracted from the high-precision map.
[0011] In a third possible implementation of the first aspect, when the target time point is the time point when the target road environment image is collected for the Nth time on the target extended navigation route, the input data further includes the output result of the landmark perception model at the previous time point adjacent to the target time point, the output result at the previous time point includes the target landmark information included in the target road environment image collected at the previous time point, and N is a positive integer greater than or equal to 2.
[0012] According to the above technical means, when the target time point is not the time point when the target road environment image is collected for the first time on the target extended navigation route, the embodiments of the present application can use the output result of the landmark perception model at the previous time point as one of the input data of the landmark perception model corresponding to the target time point, thereby improving the accuracy of perception by the landmark perception model.
[0013] In a fourth possible implementation of the first aspect, the method for generating the landmark perception model is obtaining a landmark training sample set; using the landmark training sample set to perform training to obtain a landmark perception model, and includes Each training sample in the landmark training sample set includes a road environment sample image group, and a third local pose corresponding to each frame of the road environment sample image in the road environment sample image group and a landmark ground truth value corresponding to each frame of the road environment sample image, and the road environment image group includes road environment sample images of a plurality of consecutive frames. The landmark perception model is for perceiving and outputting landmark information in the road environment sample image.
[0014] According to the above technical means, the landmark perception model is trained based on a plurality of road environment sample image groups. In the machine learning process, when performing landmark information perception on each frame of the road environment sample image, when referring to the adjacent frames before and after (which may be one adjacent frame or a plurality of adjacent frames) in the road environment sample image group of the image of the current frame, when performing landmark perception on a single-frame road environment image based on the landmark perception model, for example, landmark information outside the detection range of the sensor, occluded landmark information, landmark information that cannot be clearly displayed due to poor image quality (such as at night or in a dazzling scene), etc., the invisible landmark information of the road environment image, including landmark information not included in the current image or that cannot be clearly displayed, can be perceived, achieving the perception effect that even invisible information can be acquired. In addition, since the acquisition of the third local pose and the landmark ground truth value in the landmark training sample set do not depend on the high-precision map, a distance prediction model capable of perceiving landmark information in the road environment sample image can be automatically trained in a state of getting out of the frequent update of the high-precision map.
[0015] In the fifth possible implementation form of the first aspect, the method for obtaining the landmark ground truth value is For each of the road environment sample image groups waiting for processing, based on the third local pose of each frame of the road environment sample image in the road environment sample image group, the step of obtaining the corresponding road measurement time series data of the road environment sample image group, The step of constructing a first local map based on the road measurement time series data, By respectively projecting the first local map onto each frame of the road environment sample image in the road environment sample image group, the step of obtaining the landmark ground truth value of each frame of the road environment sample image, including.
[0016] According to the above technical means, the embodiment of the present application constructs a first local map based on the corresponding road measurement time series data of the road environment sample image group, and projects the first local map onto the road environment sample image, so as to obtain the landmark truth value of the road environment sample image for each frame. As a result, on the premise of extracting from the high-precision map, it is possible to automatically mark the existence of landmark information invisible in the road environment sample image, providing a technical basis for bringing the technical effect that even invisible information can be obtained to the landmark perception model.
[0017] In the sixth possible implementation form of the first aspect, the method for obtaining the truth value of the farthest reachable distance is as follows: Generating a second local map including the second extended navigation route based on the continuous video stream, and driving according to the second extended navigation route in the second local map; Calculating the farthest reachable distance on each lane in front of each of the second global poses on the second extended navigation route based on the second local map; Taking the calculated farthest reachable distance as the truth value, performing truth value marking on the corresponding second landmark information, and obtaining the truth value of the farthest reachable distance for each lane in the second landmark information.
[0018] According to the above technical means, the embodiment of the present application can generate a second local map including the second extended navigation route and the second extended navigation route based on the continuous video stream without manual intervention, and realize the truth value marking of the farthest reachable distance in the second landmark information, thereby improving the efficiency of truth value marking of the farthest reachable distance.
[0019] In the second aspect, the embodiment of the present application provides a method for training a distance prediction model, and the method includes: The step of obtaining a distance training sample set, including the step of performing training using the distance training sample set to obtain a distance prediction model, Each training sample in the distance training sample set includes a second extended navigation route, second landmark information, a second global pose, and a ground truth value of the farthest reachable distance. The second extended navigation route includes an extended route of a second in-vehicle navigation route. The second landmark information includes landmark information included in a second road environment image collected by a second vehicle on the second extended navigation route. The second global pose includes the global pose when the second vehicle collects the second road environment image, The distance prediction model is for predicting the farthest reachable distance on each lane in front of an arbitrary vehicle.
[0020] According to the above technical means, the embodiment of the present application can train and obtain a distance prediction model based on a distance training sample set including a second extended navigation route, second landmark information, a second global pose, and a ground truth value of the farthest reachable distance for each lane in the second landmark information. Naturally, since the acquisition of these sample information does not depend on a high-precision map, it is possible to automatically train a distance prediction model that can predict the farthest reachable distance on each lane in front of a vehicle while getting out of the frequent update of the high-precision map, and use the distance prediction model to automatically predict the farthest reachable distance on each lane in front of the vehicle. As a result, it becomes easy to quickly perform route planning according to the farthest reachable distance on each lane in front of the vehicle. Furthermore, while saving economic costs and time costs, intelligent driving plan control can be realized.
[0021] In a first possible implementation form of the second aspect, the method for obtaining the second extended navigation route is, Obtaining a second in-vehicle navigation route and a second global pose when the second vehicle is traveling on the second in-vehicle navigation route; Extracting point of interest (POI) information corresponding to the second global pose from target data; Including adding the POI information to the second in-vehicle navigation route to obtain the second extended navigation route. The target data includes navigation events and / or in-vehicle navigation maps. The POI information includes road attribute information related to the prediction of the farthest reachable distance. The POI information corresponding to the second global pose includes the POI information within a predetermined distance range in front of the second global pose on the second in-vehicle navigation route.
[0022] According to the above technical means, the embodiments of the present application extract POI information corresponding to the second global pose from navigation events and / or in-vehicle navigation maps, and add the extracted POI information to the second in-vehicle navigation route to obtain a second extended navigation route that contributes to the training of a distance prediction model with higher quality than the second in-vehicle navigation route. Therefore, the acquisition of the second extended navigation route can be performed only according to navigation events and / or in-vehicle navigation maps without relying on high-precision maps.
[0023] In a second possible implementation form of the second aspect, the method for obtaining the second landmark information included in the second road environment image collected by the second vehicle at the target time is as follows: Obtaining input data of the landmark perception model; Processing the input data based on the landmark perception model to obtain the second landmark information included in the second road environment image at the target time. The input data includes the second road environment image at the target time point and the second local pose at the target time point. The second local pose includes an offset amount with respect to the global pose at the target start point of the global pose of the second vehicle when the second road environment image is collected. The target time point is any time point when the second vehicle is collecting the second road environment image on the second extended navigation route.
[0024] According to the above technical means, the embodiments of the present application can not only automatically sense the second landmark information included in the second road environment image at any time point based on the pre-trained landmark perception model, but also, since the input data of the landmark perception model is independent of the high-precision map and only includes the second road environment image and the second local pose, the second landmark information can be obtained in a state of being extracted from the high-precision map.
[0025] In a third possible implementation form of the second aspect, when the target time point is the time point when the second road environment image is collected for the Nth time on the second extended navigation route, the input data further includes the output result of the landmark perception model at the previous time point adjacent to the target time point. The output result at the previous time point includes the second landmark information included in the second road environment image collected at the previous time point, and N is a positive integer greater than or equal to 2.
[0026] According to the above technical means, when the target time point is not the time point when the second road environment image is collected for the first time on the second extended navigation route, the embodiments of the present application can use the output result of the landmark perception model at the previous time point as one of the input data of the landmark perception model corresponding to the target time point, thereby improving the accuracy of perception by the landmark perception model.
[0027] In a fourth possible implementation form of the second aspect, the method for generating the landmark perception model is The step of obtaining a landmark training sample set, The step of performing training using the landmark training sample set to obtain a landmark perception model, and Each training sample in the landmark training sample set includes a road environment sample image group, and a third local pose corresponding to each road environment sample image in each frame of the road environment sample image group and a landmark ground truth value corresponding to each road environment sample image in each frame of the road environment sample image group. The road environment image group includes road environment sample images of a plurality of consecutive frames, The landmark perception model is for perceiving and outputting landmark information in the road environment sample image.
[0028] According to the above technical means, the landmark perception model is trained based on a plurality of road environment sample image groups. In the machine learning process, when performing landmark information perception on each frame of road environment sample image, when referring to the adjacent frames (which may be one adjacent frame or a plurality of adjacent frames) before and after in the road environment sample image group of the image of the current frame, when performing landmark perception based on the landmark perception model on a single-frame road environment image, for example, landmark information outside the detection range of the sensor, occluded landmark information, landmark information that cannot be clearly displayed due to poor image quality (such as at night or in a dazzling scene), etc., invisible landmark information in the road environment image that is not included in the current image or cannot be clearly displayed can be perceived, achieving the perception effect that even invisible information can be obtained. In addition, since the acquisition of the third local pose and the landmark ground truth value in the landmark training sample set do not depend on the high-precision map, a distance prediction model capable of perceiving the landmark information of the road environment sample image can be automatically trained in a state of getting out of the frequent update of the high-precision map.
[0029] In a fifth possible implementation of the second aspect, the method for obtaining the landmark truth value is as follows: For each group of the road environment sample images waiting for processing, based on the third local pose of the road environment sample image for each frame in the road environment sample image group, obtaining corresponding road measurement time series data of the road environment sample image group; Constructing a first local map based on the road measurement time series data; Obtaining the landmark truth value of the road environment sample image for each frame by respectively projecting the first local map onto the road environment sample image for each frame in the road environment sample image group.
[0030] According to the above technical means, the embodiment of the present application constructs a first local map based on the corresponding road measurement time series data of the road environment sample image group, and projects the first local map onto the road environment sample image, so as to obtain the landmark truth value of the road environment sample image for each frame. Thereby, on the premise of getting out of the high-precision map, it is possible to automatically mark the existence of landmark information invisible in the road environment sample image, and provide a technical basis for bringing the technical effect that even invisible things can be obtained to the landmark perception model.
[0031] In a sixth possible implementation of the second aspect, the method for obtaining the truth value of the farthest reachable distance is as follows: Generating a second local map including the second extended navigation route based on a continuous video stream, and driving along the second extended navigation route in the second local map; Based on the second local map, respectively calculating the farthest reachable distance on each lane in front of each second global pose on the second extended navigation route; Using the calculated farthest reachable distance as a truth value, perform truth value marking on the corresponding second landmark information, and obtain the truth value of the farthest reachable distance for each lane in the second landmark information.
[0032] According to the above technical means, the embodiments of the present application can generate a second local map including a second extended navigation route and a second extended navigation route based on a continuous video stream without manual intervention, and can realize truth value marking of the farthest reachable distance in the second landmark information, thereby improving the efficiency of truth value marking of the farthest reachable distance.
[0033] In a third aspect, the embodiments of the present application provide a vehicle planning and control system, the system includes a positioning module, an extended route module, a sensing module, a distance prediction module, and a planning and control module. The positioning module is used to obtain the first global pose and the first local pose of the first vehicle, and the first local pose includes an offset amount with respect to the global pose at the target start point of the global pose of the first vehicle when the first road environment image is collected. The extended route module is used to determine a first extended navigation route based on the first global pose, and the first extended navigation route includes an extended route of the first in-vehicle navigation route. The sensing module is used to sense the first landmark information and the target object information around the first vehicle. The first landmark information includes the landmark information included in the first road environment image collected by the first vehicle at the first global pose. The target object information includes at least one of the traffic signal information in front of the first vehicle, the static object information around the first vehicle, and the dynamic object information around the first vehicle. The distance prediction module is used to obtain the farthest reachable distance on each lane in the first landmark information based on the method described in any embodiment of the first aspect. The motion planning and control module is used to determine the planned trajectory of the first vehicle based on the user's prediction information, the farthest reachable distance on each lane in the first landmark information, the first local pose, and the target object information, and to control the driving of the first vehicle based on the planned trajectory.
[0034] According to the above technical means, the motion planning and control system provided by the embodiments of the present application, which includes a positioning module, an extended route module, a perception module, a distance prediction module, and a motion planning and control module, automatically predicts the farthest reachable distance on each lane in the first landmark information using a distance prediction model that does not depend on a high-precision map and is automatically trained. Moreover, based on the user's prediction information, the farthest reachable distance on each lane in the first landmark information, the first local pose, and the target object information, it can determine the planned trajectory of the first vehicle and control the driving of the first vehicle based on the planned trajectory. As a result, while saving economic costs and time costs, it is possible to realize the motion planning and control of intelligent driving. In addition, global positioning and local positioning adopt a decoupling design, that is, when each module performs data processing, global positioning and local positioning do not affect each other, and only the positioning result of either one is required. Moreover, even if the accuracy of global positioning decreases, it will not affect local positioning, so it will not affect the real-time perception and motion planning and control results, nor will it cause incorrect emergency braking.
[0035] In a first possible implementation of the third aspect, the extension route module is used to: obtain the first global pose and the first in-vehicle navigation route determined based on the first global pose; extract point of interest (POI) information corresponding to the first global pose from target data; and add the POI information to the first in-vehicle navigation route to obtain the first extended navigation route. The target data includes navigation events and / or an in-vehicle navigation map. The POI information includes road attribute information related to prediction of the farthest reachable distance. The POI information corresponding to the first global pose includes the POI information within a predetermined distance range in front of the first global pose on the first in-vehicle navigation route.
[0036] In a second possible implementation of the third aspect, the sensing module is used to: obtain input data for a landmark sensing model; and process the input data based on the landmark sensing model to obtain the first landmark information included in the first road environment image at the target time point. The input data includes the first road environment image at the target time point and the first local pose at the target time point. The target time point is any time point when the first vehicle is collecting the first road environment image on the first extended navigation route.
[0037] When the target time point is the time point when the first road environment image is collected for the Nth time on the first extended navigation route, the input data obtained by the sensing module further includes the output result of the landmark sensing model at the previous time point adjacent to the target time point. The output result at the previous time point includes the first landmark information included in the first road environment image collected at the previous time point. N is a positive integer greater than or equal to 2.
[0038] In a fourth aspect, an embodiment of the present application provides a distance prediction device, which includes an acquisition unit and a distance prediction unit. The acquisition unit is used to acquire the first global pose, the first extended navigation route, and the first landmark information of the first vehicle. The first global pose is the current global pose of the first vehicle, the first extended navigation route includes an extended route of the first in-vehicle navigation route determined based on the first global pose, and the first landmark information includes landmark information included in the first road environment image collected by the first vehicle at the first global pose. The distance prediction unit processes the first global pose, the first extended navigation route, and the first landmark information based on a distance prediction model, and is used to obtain the farthest reachable distance on each lane in the first landmark information. Before processing the first global pose, the first extended navigation route, and the first landmark information based on the distance prediction model, the acquisition unit is further used to obtain a distance training sample set. Each training sample in the distance training sample set includes a second extended navigation route, second landmark information, a second global pose, and a ground truth value of the farthest reachable distance. The second extended navigation route includes an extended route of the second in-vehicle navigation route, the second landmark information includes landmark information included in the second road environment image collected by the second vehicle on the second extended navigation route, the second global pose includes the global pose when the second vehicle collects the second road environment image, and the ground truth value of the farthest reachable distance includes the ground truth value of the farthest reachable distance for each lane in the second landmark information. The device further includes a training unit that uses the distance training sample set for training to obtain the distance prediction model.
[0039] In the first possible implementation of the fourth aspect, the acquisition unit includes a first acquisition module, an extraction module, and an addition module. The first acquisition module is used to acquire a target vehicle navigation route and a target global pose when the target vehicle is traveling on the target vehicle navigation route when the target extended navigation route includes the first extended navigation route and / or the second extended navigation route. When the target extended navigation route is the first extended navigation route, the target vehicle navigation route is the first vehicle navigation route, the target vehicle is the first vehicle, and the target global pose is the first global pose. When the target extended navigation route is the second extended navigation route, the target vehicle navigation route is the second vehicle navigation route, and the target global pose is the second global pose. The extraction module is used to extract point of interest (POI) information corresponding to the target global pose from target data, where the target data includes navigation events and / or an in-vehicle navigation map. The POI information includes road attribute information related to the prediction of the farthest reachable distance. The POI information corresponding to the target global pose includes the POI information within a predetermined distance range in front of the target global pose on the target vehicle navigation route. The addition module is used to add the POI information to the target vehicle navigation route to obtain the target extended navigation route.
[0040] In the second possible implementation of the fourth aspect, the acquisition unit includes a second acquisition module and a processing module. The second acquisition module is used to acquire input data of the landmark perception model when the target landmark information includes the first landmark information and / or the second landmark information. The input data includes the target road environment image at the target time point and the target local pose at the target time point. The target local pose includes an offset amount with respect to the global pose at the target start point of the global pose of the target vehicle when the target road environment image is collected. The target time point is any time point when the target vehicle is collecting the target road environment image on the target extended navigation route. When the target landmark information is the first landmark information, the target vehicle is the first vehicle, the target road environment image is the first road environment image, the target extended navigation route is the first extended navigation route, and the target local pose is the first local pose. When the target landmark information is the second landmark information, the target vehicle is the second vehicle, the target road environment image is the second road environment image, the target extended navigation route is the second extended navigation route, and the target local pose is the second local pose. The processing module is used to process the input data based on the landmark perception model to obtain the target landmark information included in the target road environment image at the target time point.
[0041] In a third possible implementation manner of the fourth aspect, when the target time point is the time point when the target road environment image is collected for the Nth time on the target extended navigation route, the input data further includes the output result of the landmark perception model at the previous time point adjacent to the target time point. The output result at the previous time point includes the target landmark information included in the target road environment image collected at the previous time point, and N is a positive integer greater than or equal to 2.
[0042] In a fourth possible implementation manner of the fourth aspect, the acquisition unit Before the processing module processes the input data based on the landmark perception model and obtains the target landmark information included in the target road environment image at the target time point, the processing module further includes a first generation module for generating the landmark perception model. The first generation module includes an acquisition sub-module and a training sub-module. The acquisition sub-module is used to obtain a landmark training sample set. Each training sample in the landmark training sample set includes a road environment sample image group, a third local pose corresponding to each frame of the road environment sample image in the road environment sample image group, and a landmark ground truth value corresponding to each frame of the road environment sample image in the road environment sample image group. The road environment image group includes road environment sample images of a plurality of consecutive frames. The training sub-module is used to perform training using the landmark training sample set and obtain a landmark perception model for sensing and outputting landmark information in the road environment sample image.
[0043] In a fifth possible implementation form of the fourth aspect, the acquisition sub-module is used for each group of the road environment sample images waiting for processing to obtain corresponding road measurement time series data of the road environment sample image group based on the third local pose of each frame of the road environment sample image in the road environment sample image group, construct a first local map based on the road measurement time series data, and project the first local map onto each frame of the road environment sample image in the road environment sample image group respectively to obtain the landmark ground truth value of each frame of the road environment sample image.
[0044] In a sixth possible implementation form of the fourth aspect, the acquisition unit A second generation module for generating a second local map including the second extended navigation route based on a continuous video stream; A driving module for driving according to the second extended navigation route in the second local map; A calculation module for calculating, for each lane in front of each of the second global poses on the second extended navigation route, the farthest reachable distance respectively based on the second local map; A truth value marking module for performing truth value marking on the corresponding second landmark information using the calculated farthest reachable distance as a truth value, and obtaining the truth value of the farthest reachable distance for each lane in the second landmark information.
[0045] According to the above technical means, in the embodiments of the present application, first, a distance prediction model is trained and obtained based on a distance training sample set including a second extended navigation route, second landmark information, a second global pose, and a truth value of the farthest reachable distance for each lane in the second landmark information. Next, the first global pose, the first extended navigation route, and the first landmark information are input into the distance prediction model, and the farthest reachable distance on each lane in the first landmark information can be predicted. Since the second global pose, the second extended navigation route, the second landmark information required during model training, and the first global pose, the first extended navigation route, and the first landmark information required during model application do not depend on a high-precision map, the embodiments of the present application can automatically train a distance prediction model capable of predicting the farthest reachable distance on each lane in front of the vehicle in a state of getting out of the frequent update of the high-precision map, and can automatically predict the farthest reachable distance on each lane in front of the vehicle using the distance prediction model. As a result, it becomes easy to quickly perform route planning according to the farthest reachable distance on each lane in front of the vehicle. Furthermore, while saving economic costs and time costs, intelligent driving plan control can be realized.
[0046] In a fifth aspect, the embodiments of the present application provide a training device for a distance prediction model, the device including an acquisition unit and a training unit. The acquisition unit is used to acquire a distance training sample set, and each training sample in the distance training sample set includes a second extended navigation route, second landmark information, a second global pose, and a ground truth value of the farthest reachable distance. The second extended navigation route includes an extended route of a second in-vehicle navigation route. The second landmark information includes landmark information included in a second road environment image collected by a second vehicle on the second extended navigation route. The second global pose includes the global pose when the second vehicle collects the second road environment image. The training unit is used to perform training using the distance training sample set and obtain a distance prediction model for predicting the farthest reachable distance on each lane in front of an arbitrary vehicle.
[0047] In a first possible implementation manner of the fifth aspect, the acquisition unit includes a first acquisition module, an extraction module, and an addition module. The first acquisition module is used to acquire a second in-vehicle navigation route and a second global pose when the second vehicle is traveling on the second in-vehicle navigation route. The extraction module is used to extract point of interest (POI) information corresponding to the second global pose from target data. The target data includes navigation events and / or an in-vehicle navigation map. The POI information includes road attribute information related to the prediction of the farthest reachable distance. The POI information corresponding to the second global pose includes the POI information within a predetermined distance range in front of the second global pose on the second in-vehicle navigation route. The addition module is used to add the POI information to the second in-vehicle navigation route and obtain the second extended navigation route.
[0048] In a second possible implementation manner of the fifth aspect, the acquisition unit includes a second acquisition module and a processing module. The second acquisition module is used to acquire the input data of the landmark perception model. The input data includes the second road environment image at the target time point and the second local pose at the target time point. The second local pose includes an offset amount with respect to the global pose at the target start point of the global pose of the second vehicle when the second road environment image is collected. The target time point is any time point when the second vehicle is collecting the second road environment image on the second extended navigation route. The processing module is used to process the input data based on the landmark perception model to obtain the second landmark information included in the second road environment image at the target time point.
[0049] In a third possible implementation manner of the fifth aspect, when the target time point is the time point when the second vehicle is collecting the second road environment image for the Nth time on the second extended navigation route, the input data further includes the output result of the landmark perception model at the previous time point adjacent to the target time point. The output result at the previous time point includes the second landmark information included in the second road environment image collected at the previous time point. N is a positive integer greater than or equal to 2.
[0050] In a fourth possible implementation manner of the fifth aspect, the acquisition unit further includes a first generation module for generating the landmark perception model before processing the input data based on the landmark perception model to obtain the second landmark information included in the second road environment image at the target time point. The first generation module includes an acquisition sub-module and a training sub-module. The acquisition sub-module is used to acquire a landmark training sample set. Each training sample in the landmark training sample set includes a road environment sample image group, a third local pose corresponding to each road environment sample image in the road environment sample image group for each frame, and a landmark ground truth value corresponding to each road environment sample image in the road environment sample image group for each frame. The road environment image group includes road environment sample images of a plurality of consecutive frames. The training sub-module is used to perform training using the landmark training sample set and obtain a landmark perception model for perceiving and outputting landmark information in the road environment sample images.
[0051] In a fifth possible implementation manner of the fifth aspect, the acquisition sub-module is used for each road environment sample image group waiting for processing to: acquire corresponding road measurement time series data of the road environment sample image group based on the third local pose of each road environment sample image in the road environment sample image group for each frame; construct a first local map based on the road measurement time series data; and project the first local map onto each road environment sample image in the road environment sample image group for each frame, so as to obtain the landmark ground truth value of each road environment sample image for each frame.
[0052] In a sixth possible implementation manner of the fifth aspect, the acquisition unit a second generation module for generating a second local map including the second extended navigation route based on a continuous video stream; a driving module for driving according to the second extended navigation route in the second local map; a calculation module for calculating, based on the second local map, the farthest reachable distance on each lane in front of each second global pose on the second extended navigation route respectively. Using the calculated farthest reachable distance as a truth value, perform truth value marking on the corresponding second landmark information, and obtain the truth value of the farthest reachable distance for each lane in the second landmark information, including a truth value marking module.
[0053] According to the above technical means, the embodiments of the present application can train and obtain a distance prediction model based on a distance training sample set including a second extended navigation route, second landmark information, a second global pose, and the truth value of the farthest reachable distance for each lane in the second landmark information. Naturally, since the acquisition of these sample information does not depend on a high-precision map, it is possible to automatically train a distance prediction model that can predict the farthest reachable distance on each lane in front of the vehicle in a state of getting out of the frequent update of the high-precision map, and use the distance prediction model to automatically predict the farthest reachable distance on each lane in front of the vehicle. As a result, it becomes easy to quickly perform route planning according to the farthest reachable distance on each lane in front of the vehicle, and further, it is possible to realize intelligent driving plan control while saving economic costs and time costs.
[0054] In a sixth aspect, the embodiments of the present application provide a computer-readable storage medium storing a computer program, and when the program is executed by a processor, the method described in any one of the possible implementation forms of the first aspect or any one of the possible implementation forms of the second aspect is realized.
[0055] In a seventh aspect, the embodiments of the present application provide an electronic device, and the electronic device includes one or more processors, and a storage device coupled to the processor for storing one or more programs. When one or more programs are executed by one or more processors, the electronic device implements the method described in any one possible implementation form of the first aspect or any one possible implementation form of the second aspect.
[0056] In an eighth aspect, an embodiment of the present application provides a vehicle, and the vehicle includes the system described in any one possible implementation form of the third aspect, or the device described in any one possible implementation form of the third aspect or any one possible implementation form of the fourth aspect, or the electronic device described in the fourth aspect.
[0057] In a ninth aspect, an embodiment of the present application provides a computer program product including instructions, and when the instructions are executed by a computer or a processor, the computer or the processor executes the method described in any one possible implementation form of the first aspect or any one possible implementation form of the second aspect.
Brief Description of the Drawings
[0058] To more clearly illustrate the solution means of the embodiments of the present application or the prior art, the drawings necessary for use in the description of the embodiments or the prior art will be briefly described below. Of course, the drawings described below are some embodiments of the present application, and those skilled in the art can conceive of other drawings based on these drawings without creative effort.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
[0059] Hereinafter, with reference to the drawings according to the embodiments of the present application, the technical solution will be described clearly and completely. Of course, the described embodiments are only a part of the embodiments of the present application, not all of them. Those skilled in the art can obtain all other embodiments without creative labor based on the embodiments in the present application, and all of them belong to the protection scope of the present application.
[0060] In addition, the embodiments in the present application and the features in the embodiments can be combined with each other as long as they do not conflict. The terms "including" and "having" and their variants in the embodiments of the present application and the accompanying drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device including a series of steps or units is not limited to the listed steps or units, but further selectively includes steps or units not listed, or further selectively includes other steps or units specific to these processes, methods, products, or devices.
[0061] FIG. 1 is a flowchart of a distance prediction method, and the method can be applied to an electronic device or a computer device, specifically, a vehicle or a server. Specifically, the method can include S110 to S120.
[0062] In S110, the first global pose, the first extended navigation route, and the first landmark information of the first vehicle are acquired.
[0063] The first global pose is the current global pose of the first vehicle. The first extended navigation route includes the extended route of the first in-vehicle navigation route determined based on the first global pose. The first landmark information includes the landmark information included in the first road environment image collected by the first vehicle at the first global pose.
[0064] A pose includes the position and the attitude of the vehicle. A global pose includes the position and the attitude of the vehicle when based on global positioning. When acquiring a global pose, high-precision positioning such as RTK (Real-time kinematic) is not necessary, and it can be acquired only by road-level positioning such as navigation positioning. Global positioning can perform positioning with the earth as a reference system, that is, the coordinate system of global positioning can be a geodetic coordinate system.
[0065] After the user inputs a starting point (which may be the default current position) and an end point into the in-vehicle navigation software, the in-vehicle navigation software generates a first in-vehicle navigation route according to the starting point and the end point, and the navigation positioning system measures the first global pose of the first vehicle in real time. Here, the in-vehicle navigation software may be the software of the in-vehicle device or the software of the mobile terminal communicating with the in-vehicle device. Correspondingly, the in-vehicle navigation route may be the route generated by the in-vehicle navigation software of the in-vehicle device or the route generated by the in-vehicle navigation software of the mobile terminal communicating with the in-vehicle device. The embodiments of the present application do not limit the origin of the in-vehicle navigation route.
[0066] Among them, the specific acquisition methods of the first extended navigation route and the first landmark information can refer to the acquisition methods of the target extended navigation route and the target landmark information described later, and the description is omitted here.
[0067] In S120, the first global pose, the first extended navigation route, and the first landmark information are processed based on the distance prediction model to obtain the farthest reachable distance on each lane in the first landmark information.
[0068] After obtaining the first global pose, the first extended navigation route, and the first landmark information, the first global pose, the first extended navigation route, and the first landmark information are input into a pre-trained distance prediction model for calculation, and the farthest reachable distance on each lane in the first landmark information output from the distance prediction model is obtained. The farthest reachable distance obtained on each lane in the first landmark information may be indicated by an actual distance, such as 2000 meters, or, for example, 0 represents 0 meters, 1 represents (0, 200], 2 represents (200, 400] meters, 3 represents (400, 600] meters, 4 represents (600, 800] meters, 5 represents (800, 1000] meters, 6 represents (1000, 2000] meters, 7 represents 2000 meters or more, and may also be indicated by a symbol in a mapping relationship with the actual distance. As shown in FIG. 2, if there are 4 lanes in the first landmark information and the farthest reachable distances from left to right are 7, 7, 0, 7 respectively, it means that there are 3 lanes that can be traveled.
[0069] After obtaining the farthest reachable distance on each lane in the first landmark information, input the farthest reachable distance on each lane in the first landmark information, the user's prediction information, the first local pose, and the target object information around the first vehicle into the planning and control module, and it is possible to obtain planning and control results such as whether to change lanes, when to change lanes, whether to accelerate or decelerate, when to accelerate or decelerate, and the future planned route trajectory. The user's prediction information includes operations on the user's vehicle such as lever operations for lane changes, activation of lane change lights, and brake operations. The target object information includes at least one of traffic signal information in front of the first vehicle such as surrounding vehicles, pedestrians, obstacles, and traffic lights, static object information around the first vehicle, and dynamic object information around the first vehicle. When the farthest reachable distance of the lane in which the host vehicle is currently traveling becomes smaller, the planning and control module refers to the user's prediction information, the first local pose, and the target object information to determine whether to change lanes and when to change lanes.
[0070] The distance prediction method according to the embodiments of the present application first trains and obtains a distance prediction model based on a distance training sample set including a second extended navigation route, second landmark information, a second global pose, and a ground truth value of the farthest reachable distance for each lane in the second landmark information. Next, the first global pose, the first extended navigation route, and the first landmark information are input into the distance prediction model, and the farthest reachable distance on each lane in the first landmark information can be predicted. Since the second global pose, the second extended navigation route, the second landmark information required during model training, and the first global pose, the first extended navigation route, and the first landmark information required during model application do not depend on a high-precision map, the embodiments of the present application can automatically train a distance prediction model capable of predicting the farthest reachable distance on each lane in front of the vehicle in a state of getting out of the frequent update of the high-precision map, and can automatically predict the farthest reachable distance on each lane in front of the vehicle using the distance prediction model. As a result, it becomes easy to quickly perform route planning according to the farthest reachable distance on each lane in front of the vehicle. Furthermore, while saving economic costs and time costs, intelligent driving plan control can be realized.
[0071] FIG. 3 is a flowchart of a method for training a distance prediction model, and the method can be applied to an electronic device or a computer device, specifically a vehicle or a server. Specifically, the method can include S210 to S220.
[0072] In S210, a distance training sample set is obtained, and each training sample in the distance training sample set includes a second extended navigation route, second landmark information, a second global pose, and a ground truth value of the farthest reachable distance.
[0073] Among them, the second extended navigation route includes the extended route of the second in-vehicle navigation route, the second landmark information includes the landmark information included in the second road environment image collected by the second vehicle on the second extended navigation route, the second global pose includes the global pose when the second vehicle is collecting the second road environment image, and the truth value of the farthest reachable distance includes the truth value of the farthest reachable distance for each lane in the second landmark information. For the specific acquisition methods of the second extended navigation route and the second landmark information, reference can be made to the acquisition methods of the target extended navigation route and the target landmark information described later, and the description is omitted here. The navigation positioning system positions the second global pose of the second vehicle in real time.
[0074] Hereinafter, the acquisition method of the truth value of the farthest reachable distance will be described.
[0075] The truth value of the farthest reachable distance may be marked manually or automatically. The automatic marking method includes generating a second local map including the second extended navigation route based on a continuous video stream, driving according to the second extended navigation route in the second local map, calculating the farthest reachable distance on each lane in front of each second global pose on the second extended navigation route based on the second local map, performing truth value marking on the corresponding second landmark information with the calculated farthest reachable distance as the truth value, and obtaining the truth value of the farthest reachable distance for each lane in the second landmark information.
[0076] The continuous video stream can be generated by devices such as the image acquisition device of the second vehicle and the drive recorder as long as it can record the road environment information in front of or around the second vehicle during driving. The "continuous" here includes, in addition to the literal continuity, that the total number of frames of the video stream collected within a predetermined time exceeds a specific predetermined value.
[0077] As an additional explanation, the farthest reachable distance may be the farthest reachable distance within a certain distance range, and this distance range is usually larger than the distance range included in the road environment image. For example, if the road environment image has a range of 100 meters, the farthest reachable distance can be set as a distance range limited within a 2-kilometer range.
[0078] In S220, training is performed using the distance training sample set to obtain a distance prediction model.
[0079] The distance prediction model is for predicting the farthest reachable distance of vehicles on each lane ahead. In the embodiments of the present application, the distance prediction model can be obtained through multiple iterations of training. Therefore, each time an iteration of training is performed based on the distance training sample set, after obtaining the distance prediction model obtained in this iteration of training, at least one frame of the second road environment image is processed based on the distance prediction model obtained in this iteration of training, and the predicted value of the farthest reachable distance of the second road environment image for each frame of at least one frame of the second road environment image is obtained. At the same time, based on the difference between the predicted value of the farthest reachable distance and the true value of the farthest reachable distance corresponding thereto, a loss value is calculated. When the loss value is greater than the second loss threshold, the iteration of training is continued, and the training is stopped until the loss value is less than or equal to the second loss threshold, and the finally obtained distance prediction model after training is used as the finally required distance prediction model.
[0080] Note that the first vehicle may be the vehicle used when participating in the training of the distance prediction model, or it may be a vehicle that does not participate in the training of the distance prediction model. Therefore, the first vehicle may or may not be the second vehicle. The number of second vehicles participating in the distance prediction model may be one or more.
[0081] For the first road environment image and the second road environment image, after obtaining the distance prediction model by first training and converging it based on the second road environment image, when processing the information in the first road environment image using the distance prediction model, the second road environment image has an earlier collection time than the first road environment image. That is, the second road environment image belongs to the historical road environment images with respect to the first road environment image. After obtaining the distance prediction model by first training and converging it based on a group of second road environment images, in order to improve the quality of the distance prediction model, a group of second road environment images including road scenes waiting to be supplemented are added and collected for further training of the distance prediction model. The second road environment images for further training of the distance prediction model may have a later collection time than the first road environment image. That is, at this time, the second road environment image may not belong to the historical road environment images with respect to the first road environment image.
[0082] Accordingly, the second global pose may belong to the historical global poses with respect to the first global pose, and may also not belong to the historical global poses. Also, the second extended navigation route may belong to the historical extended navigation routes with respect to the first extended navigation route, and may also not belong to the historical extended navigation routes. The second landmark information may belong to the historical landmark information with respect to the first landmark information, and may also not belong to the historical landmark information.
[0083] The training method of the distance prediction model provided by the embodiments of the present application trains and obtains a distance prediction model based on a distance training sample set including a second extended navigation route, second landmark information, a second global pose, and the ground truth of the farthest reachable distance for each lane in the second landmark information. Naturally, since the acquisition of these sample information does not depend on a high-precision map, a distance prediction model capable of predicting the farthest reachable distance on each lane in front of the vehicle can be automatically trained in a state of escaping from the frequent update of the high-precision map, and the farthest reachable distance on each lane in front of the vehicle can be automatically predicted using the distance prediction model. Thereby, it becomes easy to quickly perform route planning according to the farthest reachable distance on each lane in front of the vehicle. Furthermore, while saving economic costs and time costs, intelligent driving plan control can be realized.
[0084] In one embodiment, when the target extended navigation route includes the first extended navigation route and / or the second extended navigation route, the method for obtaining the target extended navigation route will be described below.
[0085] After the user inputs a starting point (which may be the default current position) and an end point into the in-vehicle navigation software, the in-vehicle navigation software generates at least one target in-vehicle navigation route from the starting point to the end point, including road names, road types, and the geometric shape of the road (constituted by a set of global measurement points). In order to enable the distance prediction model to accurately predict the farthest reachable distance for each lane, several pieces of auxiliary information for the distance prediction model to make predictions can be added to the target in-vehicle navigation route.
[0086] An electronic device or a computer device first obtains a target in-vehicle navigation route and a target global pose when the vehicle is traveling on the target in-vehicle navigation route. Next, it extracts POI information corresponding to the target global pose from the target data. Finally, it can add the POI information to the target in-vehicle navigation route to obtain a target extended navigation route.
[0087] Here, the target data includes navigation events and / or an in-vehicle navigation map. A navigation event refers to an event played back in an in-vehicle navigation system. Examples of navigation events include "Enter the ramp 1 kilometer ahead", "There is a merging point 1 kilometer ahead", "There is a fork in the road ahead", "Please drive on the left fork", "You will enter the tunnel soon", etc. The POI information includes road attribute information related to the prediction of the farthest reachable distance, such as the number of lanes, merging and branching points, speed limit information, long solid lines, and other information. The POI information corresponding to the target global pose is the POI information within a predetermined distance range in front of the target global pose on the target in-vehicle navigation route. The predetermined distance range is greater than or equal to the length of the route included in the target road environment image collected by the vehicle at the target global pose.
[0088] As shown in FIG. 4, (a) is an in-vehicle navigation route including basic information such as road geometry, road type, and road name, and (b) is an extended navigation route after adding POI information to (a). To save storage space, the road geometry is removed, and a figure combining a simple line and an attribute icon is retained. Specifically, as shown in (c), the road can be indicated by a straight line, the merging and branching routes can be indicated by dots at the corresponding positions, and the corresponding attribute information can be given to the dots.
[0089] In addition, when the target extended navigation route is the first extended navigation route, the target in-vehicle navigation route is the first in-vehicle navigation route, the target vehicle is the first vehicle, and the target global pose is the first global pose; when the target extended navigation route is the second extended navigation route, the target in-vehicle navigation route is the second in-vehicle navigation route, and the target global pose is the second global pose.
[0090] Embodiments of the present application can extract POI information corresponding to the second global pose from navigation events and / or the in-vehicle navigation map, add the extracted POI information to the second in-vehicle navigation route, and obtain a second extended navigation route that contributes to the training of a distance prediction model of higher quality than the second in-vehicle navigation route. In addition, embodiments of the present application can extract POI information corresponding to the first global pose from navigation events and / or the in-vehicle navigation map, add the extracted POI information to the first in-vehicle navigation route, and obtain a second extended navigation route that contributes to the prediction of the farthest reachable distance from the first in-vehicle navigation route. Furthermore, the acquisition of the first extended navigation route and the second extended navigation route may be performed only according to navigation events and / or the in-vehicle navigation map without relying on a high-precision map.
[0091] In one embodiment, when the target landmark information includes the first landmark information and / or the second landmark information, the method for acquiring the target landmark information will be described below.
[0092] The target landmark information obtained from the target road environment image is not the landmark information marked on the target road environment image, but the target landmark information extracted from the target road environment image, and the obtained target landmark information is landmark information in which the geometric shape and the relative positional relationship between landmarks are retained, and its display effect is similar to that of a simple map.
[0093] A method for obtaining target landmark information included in a target road environment image collected by a vehicle at a target time point includes a step of obtaining input data of a landmark perception model, and a step of processing the input data based on the landmark perception model to obtain target landmark information included in the target road environment image at the target time point. The input data includes a target road environment image at the target time point and a target local pose at the target time point. The target local pose includes an offset amount with respect to the global pose at the target start point of the global pose of the target vehicle when the target road environment image is collected. The target time point is any time point when the target vehicle is collecting the target road environment image on the target extended navigation route.
[0094] Here, the coordinate system in which the local pose is located can be the boot coordinate system. In the boot coordinate system, the position point where the vehicle is powered is the origin of the coordinate system, so this coordinate system is also called the start coordinate system. Also, the local pose may be positioned based on a navigation positioning system without requiring high-precision positioning. The target start point can be any designated position, for example, the position when the target vehicle is powered on, or any designated position during the running of the target vehicle.
[0095] When there is a difference between the collection period of the target road environment image by the image collector and the positioning period of the navigation positioning system, for example, when the target road environment image is collected at the target time point but positioning is not performed at the target time point, the target local pose at the most recent positioning time point can be used as the target local pose at the target time point. However, since the digits of the collection period and the positioning period are relatively small, when there is a time difference, the error generated when collecting the target road environment image and the target local pose for the same target time point is negligible.
[0096] In addition, when the target landmark information is the first landmark information, the target vehicle is the first vehicle, the target road environment image is the first road environment image, the target extended navigation route is the first extended navigation route, and the target local pose is the first local pose; when the target landmark information is the second landmark information, the target vehicle is the second vehicle, the target road environment image is the second road environment image, the target extended navigation route is the second extended navigation route, and the target local pose is the second local pose.
[0097] The embodiments of the present application can not only automatically sense the target landmark information included in the target road environment image at any point in time based on the landmark perception model obtained through pre-training, but also, since the input data of the landmark perception model has nothing to do with the high-precision map and only includes the target road environment image and the target local pose, the target landmark information can be obtained in a state of being extracted from the high-precision map.
[0098] In one embodiment, the landmark perception model may be a generative model or other neural network models, and the embodiments of the present application do not limit the model algorithm specifically used for the landmark perception model. The generation method of the landmark perception model includes the steps of obtaining a landmark training sample set and using the landmark training sample set to perform training to obtain the landmark perception model. Each training sample in the landmark training sample set includes a road environment sample image group, the local pose corresponding to the road environment sample image for each frame in the road environment sample image group, and the landmark ground truth value corresponding to the road environment sample image for each frame, and is for the landmark perception model to sense and output the landmark information in the road environment sample image.
[0099] Here, the road environment image group includes road environment sample images of a plurality of consecutive frames. Here, "consecutive" includes not only literal continuity but also the case where the total number of frames of the road environment sample images collected within a predetermined time exceeds a specific predetermined value.
[0100] According to the above technical means, the landmark perception model is trained based on a plurality of road environment image groups. In the machine learning process, when performing landmark information perception on the road environment sample image for each frame, when referring to the adjacent frames (which may be one adjacent frame or a plurality of adjacent frames) before and after in the road environment image group of the image of the frame, when performing landmark perception based on the landmark perception model on a single-frame road environment image, for example, landmark information outside the detection range of the sensor, blocked landmark information, landmark information that cannot be clearly displayed due to poor image quality (such as at night or in a dazzling scene), etc., the invisible landmark information of the road environment image including the landmark information not included in the current image or that cannot be clearly displayed can be perceived, achieving the perception effect that even the invisible ones can be acquired. In addition, since the acquisition of the local pose and the landmark ground truth value in the landmark training sample set do not depend on the high-precision map, a distance prediction model capable of perceiving the landmark information of the road environment sample image can be automatically trained while getting out of the frequent update of the high-precision map.
[0101] In one embodiment, the landmark truth value may be marked manually or automatically. The automatic marking method includes, for each group of road environment sample images waiting for processing, the steps of obtaining the corresponding road measurement time series data of the road environment sample image group based on the third local pose of the road environment sample image for each frame in the road environment sample image group, constructing a first local map based on the road measurement time series data, and obtaining the landmark truth value of the road environment sample image for each frame by projecting the first local map onto the road environment sample image for each frame in the road environment sample image group.
[0102] The road measurement time series data includes the vehicle driving trajectory and the road environment images collected during the vehicle's driving process. If the road measurement time series data is actual time series data, it may be data recorded by a drive recorder, data collected by an in-vehicle sensor, data recorded when the server end interacts with the vehicle, or data recorded by other devices.
[0103] When obtaining the corresponding road measurement time series data of the road environment sample image group based on the third local pose of the road environment sample image for each frame in the road environment sample image group, first, obtain the third local pose of the road environment sample image for each frame in the road environment sample image group, and then search for the road measurement time series data including these third local poses so that the road measurement time series data including these third local poses can be used as the corresponding road measurement time series data of the road environment sample image group.
[0104] A method for projecting a first local map onto a road environment sample image and obtaining the ground truth value of landmarks in the road environment sample image includes the step of converting the first local map from a map coordinate system to an image coordinate system with respect to the first local map, and using the landmark information in the road environment sample image of the first local map converted to the image coordinate system as the ground truth value of the landmarks in the road environment sample image.
[0105] Embodiments of the present application construct a first local map based on corresponding road measurement time series data of a road environment sample image group, and project the first local map onto the road environment sample image, thereby obtaining the ground truth value of landmarks in the road environment sample image for each frame. As a result, on the premise of extracting from a high-precision map, it is possible to automatically mark the existence of landmark information that is invisible in the road environment sample image, and provide a technical basis for bringing about a technical effect that even invisible landmarks can be obtained to the landmark perception model.
[0106] Each time training iterations are performed based on a landmark training sample set, after obtaining the landmark perception model obtained in the current training iteration, at least one road environment sample image group is processed based on the landmark perception model obtained in the current training iteration, and the landmark prediction value of the road environment sample image for each frame in the at least one road environment sample image group is obtained. At the same time, a loss value is calculated based on the difference between the landmark prediction value and the corresponding ground truth value of the landmark. When the loss value is greater than a first loss threshold, the training iteration is continued, and the training is stopped until the loss value is less than or equal to the first loss threshold, and finally the finally obtained landmark perception model after training is used as the finally required landmark perception model.
[0107] After training the final landmark perception model, for each frame of road environment image collected by the second vehicle on the second extended navigation route, the road environment image of the frame and the third local pose corresponding to the road environment image of the frame are directly input into the landmark perception model for processing, and the landmark information included in the road environment image of the frame output from the landmark perception model can be obtained.
[0108] In order to further improve the accuracy of the target landmark information and further enhance the perception effect that even invisible things can be obtained, in the embodiment of the present application, after perceiving the target landmark information included in the target road environment image of one frame based on the landmark perception model, the target landmark information can also be used as the input information of the landmark perception model when perceiving the road environment image of the next frame. That is, when the target time point is the time point when the target road environment image is collected for the Nth time on the target extended navigation route, the input information of the landmark perception model can include, in addition to the target road environment image and the target local pose at the target time point, the output result of the landmark perception model at the previous time point adjacent to the target time point. The output result at the previous time point includes the target landmark information included in the target road environment image collected at the previous time point. Thereby, when the landmark perception model performs landmark perception on the road environment image of the current frame, it can also refer to the perception result of the previous frame as an intermediate value, and N is a positive integer greater than or equal to 2.
[0109] Based on the above method embodiments, another embodiment of the present application provides a vehicle planning and control system. As shown in FIGS. 5 and 6, the system includes a positioning module 310, an extended route module 320, a perception module 330, a distance prediction module 340, and a planning and control module 350. The positioning module 310 is used to obtain the first global pose and the first local pose of the first vehicle. The first local pose includes the offset amount with respect to the global pose at the target start point of the global pose of the first vehicle when the first road environment image is collected. The extended route module 320 is used to determine a first extended navigation route based on the first global pose. The first extended navigation route includes an extended route of the first in-vehicle navigation route. The sensing module 330 is used to sense the first landmark information and the target object information around the first vehicle. The first landmark information includes the landmark information included in the first road environment image collected by the first vehicle at the first global pose. The target object information includes at least one of traffic signal information in front of the first vehicle, static object information around the first vehicle, and dynamic object information around the first vehicle. The distance prediction module 340 is used to obtain the farthest reachable distance on each lane in the first landmark information based on the method described in any one of the embodiments of the above distance prediction method. The planning and control module 350 is used to determine the planned trajectory of the first vehicle based on the user's prediction information, the farthest reachable distance on each lane in the first landmark information, the first local pose, and the target object information, and to control the driving of the first vehicle based on the planned trajectory.
[0110] According to the above technical means, the planning and control system provided by the embodiments of the present application, which includes a positioning module, an extended route module, a perception module, a distance prediction module, and a planning and control module, does not rely on a high-precision map and automatically predicts the farthest reachable distance on each lane in the first landmark information using a trained distance prediction model. At the same time, based on the user's prediction information, the farthest reachable distance on each lane in the first landmark information, the first local pose, and the target object information, it determines the planned trajectory of the first vehicle, and controls the driving of the first vehicle based on the planned trajectory, thereby saving economic costs and time costs while realizing the planning and control of intelligent driving. Moreover, the global positioning and local positioning adopt a decoupling design, that is, when each module performs data processing, the global positioning and local positioning do not affect each other, and only one of the positioning results is required. Also, even if the accuracy of the global positioning decreases, it does not affect the local positioning, so it does not affect the real-time perception and planning and control results, nor does it apply an incorrect emergency brake.
[0111] In one possible implementation form, as shown in FIG. 6, the system further includes an in-vehicle navigation module 360. After the in-vehicle navigation module 360 obtains the start and end points (i.e., the starting point and the ending point) input by the user, it is used to generate a first in-vehicle navigation route and target data according to the start and end points. The starting point can be the default current position, that is, the first global pose at the current time. FIG. 6 takes the case of user 1 as an example.
[0112] In one possible implementation, the extension route module 320 is used to obtain the first global pose and the first in-vehicle navigation route determined based on the first global pose, extract point of interest (POI) information corresponding to the first global pose from target data, and add the POI information to the first in-vehicle navigation route to obtain the first extended navigation route. The target data includes navigation events and / or in-vehicle navigation maps. The POI information includes road attribute information related to the prediction of the farthest reachable distance. The POI information corresponding to the first global pose includes the POI information within a predetermined distance range in front of the first global pose on the first in-vehicle navigation route.
[0113] In one possible implementation, the sensing module 330 is used to obtain input data for a landmark sensing model, process the input data based on the landmark sensing model, and obtain the first landmark information included in the first road environment image at the target time. The input data includes the first road environment image at the target time and the first local pose at the target time. The target time is any time when the first vehicle is collecting the first road environment image on the first extended navigation route.
[0114] In one possible implementation, as shown in FIG. 6, the system further includes an image collector 370 for collecting a first road environment image.
[0115] In one possible implementation, when the target time point is the time point when the first road environment image is collected for the Nth time on the first extended navigation route, the input data acquired by the sensing module 330 further includes the output result of the landmark sensing model at the previous time point adjacent to the target time point, the output result at the previous time point includes the first landmark information included in the first road environment image collected at the previous time point, and N is a positive integer greater than or equal to 2.
[0116] According to an embodiment of the above method, another embodiment of the present application provides a distance prediction device. As shown in FIG. 7, the device includes an acquisition unit 410 and a distance prediction unit 420. The acquisition unit 410 is used to acquire the first global pose of the first vehicle, the first extended navigation route, and the first landmark information. The first global pose is the current global pose of the first vehicle. The first extended navigation route includes an extended route of the first in-vehicle navigation route determined based on the first global pose. The first landmark information includes the landmark information included in the first road environment image collected by the first vehicle at the first global pose. The distance prediction unit 420 processes the first global pose, the first extended navigation route, and the first landmark information based on a distance prediction model, and is used to acquire the farthest reachable distance on each lane in the first landmark information. The acquisition unit 410 is further used to obtain a distance training sample set before processing the first global pose, the first extended navigation route, and the first landmark information based on a distance prediction model. Each training sample in the distance training sample set includes a second extended navigation route, second landmark information, a second global pose, and a ground truth value of the farthest reachable distance. The second extended navigation route includes an extended route of a second in-vehicle navigation route. The second landmark information includes landmark information included in a second road environment image collected by a second vehicle on the second extended navigation route. The second global pose includes the global pose when the second vehicle collects the second road environment image. The ground truth value of the farthest reachable distance includes the ground truth value of the farthest reachable distance for each lane in the second landmark information. The apparatus further includes a training unit 430 for performing training using the distance training sample set to obtain the distance prediction model.
[0117] In one possible implementation form, the acquisition unit 410 includes a first acquisition module, an extraction module, and an addition module. The first acquisition module is used to obtain a target in-vehicle navigation route and a target global pose when the target vehicle is traveling on the target in-vehicle navigation route when the target extended navigation route includes the first extended navigation route and / or the second extended navigation route. When the target extended navigation route is the first extended navigation route, the target in-vehicle navigation route is the first in-vehicle navigation route, the target vehicle is the first vehicle, and the target global pose is the first global pose. When the target extended navigation route is the second extended navigation route, the target in-vehicle navigation route is the second in-vehicle navigation route, and the target global pose is the second global pose. The extraction module is used to extract interesting point (POI) information corresponding to the target global pose from the target data. The target data includes navigation events and / or in-vehicle navigation maps. The POI information includes road attribute information related to the prediction of the farthest reachable distance. The POI information corresponding to the target global pose includes the POI information within a predetermined distance range in front of the target global pose on the target in-vehicle navigation route. The addition module is used to add the POI information to the target in-vehicle navigation route to obtain the target extended navigation route.
[0118] In one possible implementation, the acquisition unit 410 includes a second acquisition module and a processing module. The second acquisition module is used to acquire the input data of the landmark perception model when the target landmark information includes the first landmark information and / or the second landmark information. The input data includes the target road environment image at the target time point and the target local pose at the target time point. The target local pose includes the offset amount with respect to the global pose at the target start point of the global pose of the target vehicle when the target road environment image is collected. The target time point is any time point when the target vehicle is collecting the target road environment image on the target extended navigation route. When the target landmark information is the first landmark information, the target vehicle is the first vehicle, the target road environment image is the first road environment image, the target extended navigation route is the first extended navigation route, and the target local pose is the first local pose. When the target landmark information is the second landmark information, the target vehicle is the second vehicle, the target road environment image is the second road environment image, the target extended navigation route is the second extended navigation route, and the target local pose is the second local pose. The processing module is used to process the input data based on the landmark perception model and obtain the target landmark information included in the target road environment image at the target time point.
[0119] In one possible implementation, when the target time point is the time point when the target road environment image is collected for the Nth time on the target extended navigation route, the input data further includes the output result of the landmark perception model at the previous time point adjacent to the target time point, the output result at the previous time point includes the target landmark information included in the target road environment image collected at the previous time point, and N is a positive integer greater than or equal to 2.
[0120] In one possible implementation, the acquisition unit 410 Before the processing module processes the input data based on the landmark perception model and obtains the target landmark information included in the target road environment image at the target time point, it further includes a first generation module for generating the landmark perception model. The first generation module includes an acquisition sub-module and a training sub-module. The acquisition sub-module is used to obtain a landmark training sample set. Each training sample in the landmark training sample set includes a road environment sample image group, a third local pose corresponding to each frame of the road environment sample image in the road environment sample image group, and a landmark ground truth value corresponding to each frame of the road environment sample image. The road environment image group includes road environment sample images of a plurality of consecutive frames. The training sub-module is used to perform training using the landmark training sample set and obtain a landmark perception model for perceiving and outputting landmark information in the road environment sample image.
[0121] In one possible implementation, the acquisition sub-module is used for each group of the road environment sample images waiting for processing to obtain corresponding road measurement time-series data of the road environment sample image group based on the third local pose of the road environment sample image for each frame in the road environment sample image group, construct a first local map based on the road measurement time-series data, and obtain the landmark ground truth value of the road environment sample image for each frame by respectively projecting the first local map onto the road environment sample image for each frame in the road environment sample image group.
[0122] In one possible implementation, the acquisition unit 410 a second generation module for generating a second local map including the second extended navigation route based on a continuous video stream, a driving module for driving according to the second extended navigation route in the second local map, a calculation module for calculating the farthest reachable distance on each lane in front of each second global pose on the second extended navigation route based on the second local map, a truth value marking module for performing truth value marking on the corresponding second landmark information with the calculated farthest reachable distance as the truth value and obtaining the truth value of the farthest reachable distance for each lane in the second landmark information.
[0123] The distance prediction device provided by the embodiments of the present application first trains and obtains a distance prediction model based on a distance training sample set including a second extended navigation route, second landmark information, a second global pose, and the ground truth of the farthest reachable distance for each lane in the second landmark information. Next, the first global pose, the first extended navigation route, and the first landmark information are input into the distance prediction model, and the farthest reachable distance on each lane in the first landmark information can be predicted. Since the second global pose, the second extended navigation route, the second landmark information required during model training, and the first global pose, the first extended navigation route, and the first landmark information required during model application all do not depend on a high-precision map, the embodiments of the present application can automatically train a distance prediction model capable of predicting the farthest reachable distance on each lane in front of the vehicle while getting out of the frequent update of the high-precision map, and can automatically predict the farthest reachable distance on each lane in front of the vehicle using the distance prediction model. As a result, it becomes easy to quickly perform route planning according to the farthest reachable distance on each lane in front of the vehicle, and furthermore, intelligent driving planning and control can be realized while saving economic costs and time costs.
[0124] According to an embodiment of the above method, another embodiment of the present application provides a training device for a distance prediction model. As shown in FIG. 8, the device includes an acquisition unit 510 and a training unit 520. The acquisition unit 510 is used to acquire a distance training sample set. Each training sample in the distance training sample set includes a second extended navigation route, second landmark information, a second global pose, and a ground truth value of the farthest reachable distance. The second extended navigation route includes an extended route of a second in-vehicle navigation route. The second landmark information includes landmark information included in a second road environment image collected by a second vehicle on the second extended navigation route. The second global pose includes the global pose when the second vehicle collects the second road environment image. The training unit 520 is used to perform training using the distance training sample set and obtain a distance prediction model for predicting the farthest reachable distance on each lane in front of an arbitrary vehicle.
[0125] In one possible implementation, the acquisition unit 510 includes a first acquisition module, an extraction module, and an addition module. The first acquisition module is used to acquire a second in-vehicle navigation route and a second global pose when the second vehicle is traveling on the second in-vehicle navigation route. The extraction module is used to extract point-of-interest (POI) information corresponding to the second global pose from target data. The target data includes navigation events and / or an in-vehicle navigation map. The POI information includes road attribute information related to the prediction of the farthest reachable distance. The POI information corresponding to the second global pose includes the POI information within a predetermined distance range in front of the second global pose on the second in-vehicle navigation route. The addition module is used to add the POI information to the second in-vehicle navigation route and obtain the second extended navigation route.
[0126] In one possible implementation, the acquisition unit 510 includes a second acquisition module and a processing module. The second acquisition module is used to acquire the input data of the landmark perception model. The input data includes the second road environment image at the target time point and the second local pose at the target time point. The second local pose includes an offset amount with respect to the global pose at the target start point of the global pose of the second vehicle when the second road environment image is collected. The target time point is any time point when the second vehicle is collecting the second road environment image on the second extended navigation route. The processing module is used to process the input data based on the landmark perception model to obtain the second landmark information included in the second road environment image at the target time point.
[0127] In one possible implementation, when the target time point is the time point when the second road environment image is collected for the Nth time on the second extended navigation route, the input data further includes the output result of the landmark perception model at the previous time point adjacent to the target time point. The output result at the previous time point includes the second landmark information included in the second road environment image collected at the previous time point, and N is a positive integer greater than or equal to 2.
[0128] In one possible implementation, the acquisition unit 510 further includes a first generation module for generating the landmark perception model before processing the input data based on the landmark perception model to obtain the second landmark information included in the second road environment image at the target time point. The first generation module includes an acquisition sub-module and a training sub-module. The acquisition sub-module is used to acquire a landmark training sample set, and each training sample in the landmark training sample set includes a road environment sample image group, a third local pose corresponding to each road environment sample image in the road environment sample image group for each frame, and a landmark ground truth value corresponding to each road environment sample image in the road environment sample image group for each frame. The road environment image group includes road environment sample images of a plurality of consecutive frames. The training sub-module is used to perform training using the landmark training sample set and obtain a landmark perception model for perceiving and outputting landmark information in the road environment sample images.
[0129] In one possible implementation, for each road environment sample image group waiting for processing, the acquisition sub-module is used to: obtain corresponding road measurement time series data of the road environment sample image group based on the third local pose of each road environment sample image in the road environment sample image group for each frame; construct a first local map based on the road measurement time series data; and obtain the landmark ground truth value of each road environment sample image for each frame by respectively projecting the first local map onto each road environment sample image in the road environment sample image group for each frame.
[0130] In one possible implementation, the acquisition unit 510 a second generation module for generating a second local map including the second extended navigation route based on a continuous video stream; a driving module for driving according to the second extended navigation route in the second local map; a calculation module for respectively calculating the farthest reachable distance on each lane in front of each second global pose on the second extended navigation route based on the second local map; Using the calculated farthest reachable distance as a truth value, perform truth value marking on the corresponding second landmark information, and include a truth value marking module for obtaining the truth value of the farthest reachable distance for each lane in the second landmark information.
[0131] The distance prediction model training device provided by the embodiments of the present application can train and obtain a distance prediction model based on a distance training sample set including a second extended navigation route, second landmark information, a second global pose, and the truth value of the farthest reachable distance for each lane in the second landmark information. Naturally, since the acquisition of these sample information does not depend on a high-precision map, a distance prediction model that can predict the farthest reachable distance on each lane in front of the vehicle can be automatically trained while getting out of the frequent update of the high-precision map, and the farthest reachable distance on each lane in front of the vehicle can be automatically predicted using the distance prediction model. As a result, it becomes easy to quickly perform route planning according to the farthest reachable distance on each lane in front of the vehicle. Furthermore, intelligent driving plan control can be realized while saving economic costs and time costs.
[0132] Based on the embodiments of the above method, another embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the program is executed by a processor, the method described in any one of the above embodiments is realized.
[0133] Based on the embodiments of the above method, another embodiment of the present application provides an electronic device or a computer device. As shown in FIG. 9, the computer device includes one or more processors 610, and a storage device 620 coupled to the processor 610 and storing one or more programs. When the one or more programs are executed by the one or more processors 610, the electronic device or computer device implements the method described in any one of the above embodiments.
[0134] Based on the above method embodiments, another embodiment of the present application provides a vehicle, which includes the system described in any one of the above embodiments, or the device described in any one of the above embodiments, or the above-described electronic device.
[0135] The vehicle includes a CPU (Central Processing Unit), a T-Box (Telematics Box), an image collector, and a navigation positioning device. Among them, the image collector is used to collect road environment images, and the navigation positioning device is used to position the vehicle to obtain the global pose and local pose of the vehicle. The CPU obtains road environment images, global poses, and local poses, and uses the above method embodiments of the distance prediction model training method to obtain a distance prediction module, and uses the above method embodiments of the distance prediction method to predict the farthest reachable distance on each lane in the current landmark information. The CPU also transmits road environment images, global poses, and local poses to the server through the T-Box, and the server can also use the above method embodiments of the distance prediction model training method to obtain a distance prediction module, and use the above method embodiments of the distance prediction method to predict the farthest reachable distance on each lane in the current landmark information.
[0136] Based on the above embodiments, another embodiment of the present application provides a computer program product including instructions. When the instructions are executed by a computer or a processor, the computer or the processor executes the method described in any one of the above embodiments.
[0137] The embodiments of the above-described apparatus correspond to the embodiments of the method and have the same technical effects as the embodiments of the method. For specific descriptions, refer to the embodiments of the method. The embodiments of the apparatus are obtained based on the embodiments of the method. For specific descriptions, the parts of the embodiments of the method can be referred to and will not be repeatedly described here. Those skilled in the art can understand that the drawings are only schematic diagrams of one embodiment, and the modules or flows related to the drawings are not necessarily required for the implementation of this application.
[0138] Those skilled in the art can understand that the modules in the apparatus according to the embodiments may be distributed in the apparatus according to the embodiments as described in the embodiments, or appropriate changes may be made and arranged in one or more apparatuses different from this embodiment. The modules according to the above embodiments may be combined as one module or further divided into a plurality of sub-modules.
[0139] Finally, it should be noted that the above embodiments are for explaining the technical solutions of this application and do not limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent substitutions for some of their technical features. It should be understood that these modifications or substitutions do not deviate from the essence of the corresponding technical solutions from the gist and scope of the technical solutions of the embodiments of this application.
Claims
1. A distance prediction method, comprising: obtaining a first global pose, a first extended navigation route, and first landmark information of a first vehicle; processing the first global pose, the first extended navigation route, and the first landmark information based on a distance prediction model, and obtaining a farthest reachable distance on each lane in the first landmark information; wherein the first global pose is the current global pose of the first vehicle, the first extended navigation route includes an extended route of a first in-vehicle navigation route determined based on the first global pose, and the first landmark information includes landmark information included in a first road environment image collected by the first vehicle at the first global pose; wherein a training method of the distance prediction model is: obtaining a distance training sample set; training using the distance training sample set to obtain the distance prediction model; each training sample in the distance training sample set includes a second extended navigation route, second landmark information, a second global pose, and a ground truth value of a farthest reachable distance, the second extended navigation route includes an extended route of a second in-vehicle navigation route, the second landmark information includes landmark information included in a second road environment image collected by a second vehicle on the second extended navigation route, the second global pose includes the global pose when the second vehicle collects the second road environment image, and the ground truth value of the farthest reachable distance includes the ground truth value of the farthest reachable distance for each lane in the second landmark information. The distance prediction method is characterized by the above.
2. When a target extended navigation route includes the first extended navigation route and / or the second extended navigation route, a method for obtaining the target extended navigation route is: obtaining a target in-vehicle navigation route and a target global pose when the target vehicle is traveling on the target in-vehicle navigation route; extracting point of interest (POI) information of interest corresponding to the target global pose from target data; adding the POI information to the target vehicle navigation route and obtaining the target extended navigation route; when the target extended navigation route is the first extended navigation route, the target vehicle navigation route is the first vehicle navigation route, the target vehicle is the first vehicle, and the target global pose is the first global pose; when the target extended navigation route is the second extended navigation route, the target vehicle navigation route is the second vehicle navigation route, and the target global pose is the second global pose; the target data includes navigation events and / or an in-vehicle navigation map, the POI information includes road attribute information related to prediction of the farthest reachable distance, and the POI information corresponding to the target global pose includes the POI information within a predetermined distance range in front of the target global pose on the target vehicle navigation route. The distance prediction method according to claim 1, characterized in that.
3. When the target landmark information includes the first landmark information and / or the second landmark information, the method for obtaining the target landmark information included in the target road environment image collected by the target vehicle at the target time point is as follows: obtaining input data of the landmark perception model; processing the input data based on the landmark perception model and obtaining the target landmark information included in the target road environment image at the target time point. The input data includes the target road environment image at the target time point and the target local pose at the target time point. The target local pose includes an offset amount with respect to the global pose at the target start point of the global pose of the target vehicle when the target road environment image is collected. The target time point is any time point when the target vehicle is collecting the target road environment image on the target extended navigation route. When the target landmark information is the first landmark information, the target vehicle is the first vehicle, the target road environment image is the first road environment image, the target extended navigation route is the first extended navigation route, and the target local pose is the first local pose. When the target landmark information is the second landmark information, the target vehicle is the second vehicle, the target road environment image is the second road environment image, the target extended navigation route is the second extended navigation route, and the target local pose is the second local pose. The distance prediction method according to claim 1, characterized in that.
4. When the target time point is the time point when the target road environment image is collected for the Nth time on the target extended navigation route, the input data further includes the output result of the landmark perception model at the previous time point adjacent to the target time point. The output result at the previous time point includes the target landmark information included in the target road environment image collected at the previous time point. N is a positive integer of 2 or more. The distance prediction method according to claim 3, characterized in that.
5. The generation method of the landmark perception model is Steps of obtaining a landmark training sample set; Steps of using the landmark training sample set for training to obtain a landmark perception model, including Each training sample in the landmark training sample set includes a road environment sample image group, a third local pose corresponding to each frame of the road environment sample image in the road environment sample image group, and a landmark ground truth value corresponding to each frame of the road environment sample image. The road environment image group includes road environment sample images of a plurality of consecutive frames. The landmark perception model is for perceiving and outputting landmark information in the road environment sample image, and the distance prediction method according to claim 3 is characterized by this.
6. The method for obtaining the landmark ground truth is as follows. For each group of the road environment sample images waiting for processing, based on the third local pose of each road environment sample image for each frame in the road environment sample image group, a step of obtaining corresponding road measurement time series data of the road environment sample image group; A step of constructing a first local map based on the road measurement time series data; A step of obtaining the landmark ground truth of each road environment sample image for each frame by respectively projecting the first local map onto each road environment sample image for each frame in the road environment sample image group, and the distance prediction method according to claim 5 is characterized by including this.
7. The method for obtaining the ground truth of the farthest reachable distance is as follows. Generating a second local map including the second extended navigation route based on a continuous video stream, and driving according to the second extended navigation route in the second local map; Based on the second local map, a step of respectively calculating the farthest reachable distance on each lane in front of each second global pose on the second extended navigation route; Taking the calculated farthest reachable distance as the ground truth, performing ground truth marking on the corresponding second landmark information, and obtaining the ground truth of the farthest reachable distance for each lane in the second landmark information, and the distance prediction method according to any one of claims 1 to 6 is characterized by including this.
8. A step of obtaining a distance training sample set; A step of using the distance training sample set for training to obtain a distance prediction model for predicting the farthest reachable distance on each lane in front of an arbitrary vehicle, and including this. Each training sample in the distance training sample set includes a second extended navigation route, second landmark information, a second global pose, and a ground truth value of the farthest reachable distance. The second extended navigation route includes an extended route of a second in-vehicle navigation route. The second landmark information includes landmark information included in a second road environment image collected by a second vehicle on the second extended navigation route. The second global pose includes the global pose when the second vehicle collected the second road environment image. A method for training a distance prediction model, characterized by the above.
9. A vehicle planning and control system, comprising a positioning module, an extended route module, a perception module, a distance prediction module, and a planning and control module. The positioning module is used to obtain a first global pose and a first local pose of a first vehicle. The first local pose includes an offset amount with respect to the global pose at the target start point of the global pose of the first vehicle when a first road environment image is collected. The extended route module is used to determine a first extended navigation route based on the first global pose. The first extended navigation route includes an extended route of a first in-vehicle navigation route. The perception module is used to perceive first landmark information and target object information around the first vehicle. The first landmark information includes landmark information included in a first road environment image collected by the first vehicle at the first global pose. The target object information includes at least one of traffic signal information in front of the first vehicle, static object information around the first vehicle, and dynamic object information around the first vehicle. The distance prediction module is used to obtain the farthest reachable distance on each lane in the first landmark information based on the distance prediction method according to any one of claims 1 to 7. The planned control module is used to determine the planned trajectory of the first vehicle based on the prediction information of the user, the farthest reachable distance on each lane in the first landmark information, the first local pose, and the target object information, and to control the driving of the first vehicle based on the planned trajectory. A vehicle planned control system characterized by this.
10. A distance prediction device, the distance prediction device includes an acquisition unit and a distance prediction unit, The acquisition unit is used to acquire the first global pose of the first vehicle, the first extended navigation route, and the first landmark information. The first global pose is the current global pose of the first vehicle. The first extended navigation route includes an extended route of the first in-vehicle navigation route determined based on the first global pose. The first landmark information includes landmark information included in the first road environment image collected by the first vehicle at the first global pose. The distance prediction unit processes the first global pose, the first extended navigation route, and the first landmark information based on a distance prediction model, and is used to obtain the farthest reachable distance on each lane in the first landmark information. Before the acquisition unit processes the first global pose, the first extended navigation route, and the first landmark information based on the distance prediction model, it is further used to acquire a distance training sample set. Each training sample in the distance training sample set includes a second extended navigation route, second landmark information, a second global pose, and a ground truth value of the farthest reachable distance. The second extended navigation route includes an extended route of the second in-vehicle navigation route. The second landmark information includes landmark information included in a second road environment image collected by a second vehicle on the second extended navigation route. The second global pose includes the global pose when the second vehicle collects the second road environment image. The ground truth value of the farthest reachable distance includes the ground truth value of the farthest reachable distance for each lane in the second landmark information. The distance prediction device further includes a training unit that uses the distance training sample set for training to obtain the distance prediction model. A distance prediction device characterized by this.
11. The acquisition unit includes a first acquisition module, an extraction module, and an addition module. The first acquisition module is used to acquire a target in-vehicle navigation route and a target global pose when the target vehicle is traveling on the target in-vehicle navigation route when the target extended navigation route includes the first extended navigation route and / or the second extended navigation route. When the target extended navigation route is the first extended navigation route, the target in-vehicle navigation route is the first in-vehicle navigation route, the target vehicle is the first vehicle, and the target global pose is the first global pose. When the target extended navigation route is the second extended navigation route, the target in-vehicle navigation route is the second in-vehicle navigation route, and the target global pose is the second global pose. The extraction module is used to extract interesting point (POI) information corresponding to the target global pose from the target data. The target data includes navigation events and / or in-vehicle navigation maps. The POI information includes road attribute information related to the prediction of the farthest reachable distance. The POI information corresponding to the target global pose includes the POI information within a predetermined distance range in front of the target global pose on the target in-vehicle navigation route. The additional module is used to add the POI information to the target in-vehicle navigation route to obtain the target extended navigation route. The distance prediction device according to claim 10, characterized in that.
12. The acquisition unit includes a second acquisition module and a processing module. The second acquisition module is used to acquire input data of the landmark perception model when the target landmark information includes the first landmark information and / or the second landmark information. The input data includes the target road environment image at the target time point and the target local pose at the target time point. The target local pose includes an offset amount relative to the global pose at the target start point of the global pose of the target vehicle when the target road environment image is collected. The target time point is any time point when the target vehicle is collecting the target road environment image on the target extended navigation route. When the target landmark information is the first landmark information, the target vehicle is the first vehicle, the target road environment image is the first road environment image, the target extended navigation route is the first extended navigation route, and the target local pose is the first local pose. When the target landmark information is the second landmark information, the target vehicle is the second vehicle, the target road environment image is the second road environment image, the target extended navigation route is the second extended navigation route, and the target local pose is the second local pose. The distance prediction device according to claim 10, wherein the processing module is used to process the input data based on the landmark perception model and obtain the target landmark information included in the target road environment image at the target time point.
13. When the target time point is the time point when the target road environment image is collected for the Nth time on the target extended navigation route, the input data further includes the output result of the landmark perception model at the previous time point adjacent to the target time point, the output result at the previous time point includes the target landmark information included in the target road environment image collected at the previous time point, and N is a positive integer of 2 or more. The distance prediction device according to claim 12, characterized in that.
14. The acquisition unit Before the processing module processes the input data based on the landmark perception model and obtains the target landmark information included in the target road environment image at the target time point, the distance prediction device according to claim 12 further includes a first generation module for generating the landmark perception model. The first generation module includes an acquisition sub-module and a training sub-module. The acquisition sub-module is used to obtain a landmark training sample set, and each training sample in the landmark training sample set includes a road environment sample image group, a third local pose corresponding to each frame of the road environment sample image in the road environment sample image group, and a landmark ground truth value corresponding to each frame of the road environment sample image. The road environment image group includes road environment sample images of a plurality of consecutive frames. The distance prediction device according to claim 12, characterized in that the training sub-module is used to perform training using the landmark training sample set and obtain a landmark perception model for sensing and outputting landmark information in the road environment sample image.
15. The acquisition sub-module is used for each of the road environment sample image groups waiting for processing to obtain corresponding road measurement time-series data of the road environment sample image group based on the third local pose of the road environment sample image for each frame in the road environment sample image group, construct a first local map based on the road measurement time-series data, and obtain the landmark ground truth value of the road environment sample image for each frame by projecting the first local map onto the road environment sample image for each frame in the road environment sample image group. The distance prediction device according to claim 14, characterized in that.
16. The acquisition unit A second generation module for generating a second local map including the second extended navigation route based on a continuous video stream, A driving module for driving along the second extended navigation route in the second local map, A calculation module for calculating, based on the second local map, the farthest reachable distance on each lane in front of each second global pose on the second extended navigation route, A ground truth marking module for performing ground truth marking on the corresponding second landmark information using the calculated farthest reachable distance as the ground truth value and obtaining the ground truth value of the farthest reachable distance for each lane in the second landmark information. The distance prediction device according to any one of claims 10 to 15, characterized in that it includes.
17. A training device for a distance prediction model, wherein the training device for the distance prediction model includes an acquisition unit and a training unit. The acquisition unit is used to acquire a distance training sample set, and each training sample in the distance training sample set includes a second extended navigation route, second landmark information, a second global pose, and a ground truth value of the farthest reachable distance. The second extended navigation route includes an extended route of a second in-vehicle navigation route. The second landmark information includes landmark information included in a second road environment image collected by a second vehicle on the second extended navigation route. The second global pose includes the global pose when the second vehicle collects the second road environment image. The training unit is used to perform training using the distance training sample set and obtain a distance prediction model for predicting the farthest reachable distance on each lane in front of an arbitrary vehicle. A training apparatus for a distance prediction model, characterized in that.
18. A computer-readable storage medium storing a computer program, wherein when the program is executed by a processor, the distance prediction method according to any one of Claims 1 to 7 or the training method of the distance prediction model according to Claim 8 is realized. A computer-readable storage medium, characterized in that.
19. An electronic device, one or more processors, and a storage device coupled to the processor and storing one or more programs, and includes when the one or more programs are executed by the one or more processors, the electronic device realizes the distance prediction method according to any one of Claims 1 to 7 or the training method of the distance prediction model according to Claim 8. An electronic device, characterized in that.
20. A vehicle comprising the system according to Claim 9, or the distance prediction device according to any one of Claims 10 to 16, or the training device for a distance prediction model according to Claim 17, or the electronic device according to Claim 19. A vehicle, characterized in that.
Citation Information
Patent Citations
Device for preventing collision
JP1998079100A
Recording device, recording program and recording method
JP2013127754A
Information processing system, information processing method, program and vehicle
JP2018118672A
Map information providing system for driving support and / or travel control of vehicle
JP2019064562A
Method of providing image by vehicle navigation device
US20220364874A1