Driving control method and device, network training method and device and vehicle
By generating implicit navigation field information through a navigation field network and optimizing trajectory planning by combining road environment and navigation information, the problem of insufficient vehicle flexibility in existing technologies is solved, enabling more flexible and accurate lane-level driving control and improving the safety and accuracy of autonomous driving.
Patent Information
- Application Number
- CN202511299596.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2026-01-23
AI Technical Summary
Existing path planning methods based on target point traction lack flexibility in scenarios such as vehicle game theory and obstacle avoidance, limiting the vehicle's decision-making space in the lateral and longitudinal directions and making it difficult to achieve accurate and flexible lane-level driving capabilities.
Implicit navigation field information is generated by the navigation field network and input into the trajectory generation network. By combining road environment, obstacles and navigation information, trajectory planning is optimized, the dependence on HD map is reduced and the flexibility and accuracy of vehicles in complex scenarios are improved.
It enables more flexible and real-time vehicle control in both map-less and map-based modes, improving the vehicle's ability to engage in strategic maneuvering, lane changing, and obstacle avoidance in complex scenarios, and ensuring a safer and more accurate autonomous driving experience.
Smart Images

Figure CN121386491A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of intelligent vehicles, and more particularly, to a driving control method, a network training method, an apparatus and a vehicle. BACKGROUND
[0002] Path planning of a vehicle is crucial in the driving process of the vehicle. At present, the path planning method based on goal point traction has some limitations, which are specifically manifested as: insufficient flexibility in vehicle game, obstacle avoidance and the like, which limits the decision space of the vehicle in the lateral and longitudinal directions.
[0003] How to enable the vehicle to obtain accurate and flexible lane-level driving capability has become a technical problem to be broken through in the field of intelligent driving. SUMMARY
[0004] The present application provides a driving control method, a network training method, an apparatus and a vehicle, which helps to improve the flexibility of the vehicle in game, lane changing and obstacle avoidance in complex scenes, and provides a guarantee for realizing a safer and more accurate automatic driving experience.
[0005] In a first aspect, a driving control method is provided, which can be executed by a vehicle, for example, can be executed by a computing platform of the vehicle, or can also be executed by a chip or circuit for the vehicle.
[0006] The method comprises: obtaining road environment information, obstacle information and navigation information, the navigation information comprising driving action information and lane selection information for guiding the vehicle to drive from a current position to a target position, the navigation information being information generated based on map information through a white-box algorithm; inputting the road environment information and the navigation information into a navigation field network to obtain first navigation field information, the navigation field network comprising an encoding network and a decoding network, the first navigation field information being feature information output by the encoding network; inputting the road environment information, the obstacle information and the first navigation field information into a trajectory generation network to obtain first trajectory information; and controlling the vehicle to drive based on the first trajectory information.
[0007] Based on the above technical solution, high-dimensional and implicit first navigation field information is obtained by using the navigation field network, and the first navigation field information is brought into the trajectory generation network to provide dense features with rich semantic information, which helps to improve the performance of trajectory planning and realize lane-level accuracy of the output trajectory, so as to control the vehicle to drive accurately according to the navigation intention.
[0008] In combination with the first aspect, in some implementation manners of the first aspect, inputting the road environment information and the navigation information into the navigation field network to obtain the first navigation field information comprises: obtaining first fusion features based on the road environment information and the navigation information; and obtaining the first navigation field information based on the first fusion features.
[0009] Based on the above technical solution, through the feature fusion step, the road environment information and the navigation information can establish the correlation between the static environment elements and the dynamic driving intention in the encoding stage. This implementation helps to improve the accuracy of the first navigation information generated by the navigation field network, provides high-quality prior knowledge for subsequent trajectory planning, and thus optimizes the generation result of the trajectory.
[0010] In combination with the first aspect, in some implementations of the first aspect, the method further includes: obtaining query point information, the query point information indicating point information on the road surface; inputting the road environment information, the obstacle information, and the first navigation field information into the trajectory generation network to obtain first trajectory information, including: obtaining second fusion features based on the road environment information, the obstacle information, the first navigation field information, and the query point information; obtaining the first trajectory information based on the second fusion features.
[0011] It should be understood that the distribution density and distribution regularity of the query points are not limited by the embodiments of the present application.
[0012] For example, the query points can be point information on the center line of the road where the vehicle is located, multiple query points on the center line of different lanes, point information distributed in a grid shape on the road, or point information distributed in a ring shape at an intersection, and the like.
[0013] For example, the query point information can also be point information at random positions within the current field of view of the vehicle.
[0014] Based on the above technical solution, by introducing the query points on the lane lines, more rich second fusion features are obtained, which helps to provide a more explicit focus of attention for the obtained first navigation information, and guide the navigation field network to focus on the features of specific query points. This helps to improve the generation accuracy of the trajectory generation network in trajectory generation.
[0015] In a second aspect, a network training method is provided, including: obtaining road environment information, obstacle information, navigation information, and second trajectory information, the navigation information including driving action information and lane selection information that guide the vehicle to travel from a current position to a target position, and the second trajectory information representing standard trajectory information that the vehicle can refer to; inputting the road environment information and the navigation information into a navigation field network to obtain first navigation field information, the navigation field network including an encoding network and a decoding network, and the first navigation field information being feature information output by the encoding network; inputting the road environment information, the obstacle information, and the first navigation field information into a trajectory generation network to obtain third trajectory information; and adjusting parameters of the navigation field network based on the second trajectory information and the third trajectory information.
[0016] Based on the technical scheme, the encoding network of the navigation field network extracts core features of input data from the input road environment information, obstacle information, navigation information and second trajectory information (trajectory true value), forms a compact and information-dense implicit feature (first navigation field information), calculates a predicted trajectory of the trajectory generation network based on the implicit feature, and adjusts parameters of the navigation field network according to the predicted trajectory and the trajectory true value, which helps to obtain a better navigation field model through training.
[0017] With reference to the second aspect, in some implementations of the second aspect, inputting the road environment information and the navigation information into the navigation field network to obtain the first navigation field information comprises: inputting the road environment information and the navigation information into the navigation field network to obtain the first navigation field information and second navigation field information, the second navigation field information being feature information output by the decoding network; and adjusting the parameters of the navigation field network based on the second trajectory information and the third trajectory information comprises: adjusting the parameters of the navigation field network based on the second trajectory information, the third trajectory information and the second navigation field information.
[0018] Based on the technical scheme, the second navigation field information (explicit navigation field) is used to assist in training the navigation field network model, which helps to provide richer and more stable training supervision signals, so that the navigation field network can learn more internal regularity feature representations that conform to the navigation intention.
[0019] With reference to the second aspect, in some implementations of the second aspect, the method further comprises: obtaining query point information, the query point information indicating point position information on a road surface; and the second navigation field information is obtained by the decoding network based on the road environment information, the navigation information and the query point information.
[0020] Based on the technical scheme, the introduction of the query point on the lane line helps to generate more accurate second navigation field information.
[0021] With reference to the second aspect, in some implementations of the second aspect, before the road environment information, the obstacle information, the navigation information and the second trajectory information are obtained, the method further comprises: obtaining map information; obtaining third navigation field information by a white-box algorithm based on the map information and the navigation information; and inputting the road environment information, the obstacle information and the first navigation field information into the trajectory generation network to obtain the third trajectory information comprises: inputting the road environment information, the obstacle information, the first navigation field information and the third navigation field information into the trajectory generation network to obtain the third trajectory information.
[0022] Based on the technical scheme, the first navigation field information output by the navigation field network and the third navigation field information determined by the white-box algorithm can be fused, the white-box algorithm can provide more reliable prior knowledge, thereby effectively constraining or calibrating the output of the navigation field network, and a more accurate predicted trajectory can be obtained.
[0023] With reference to the second aspect, in some implementations of the second aspect, before obtaining the third trajectory information, the method further includes: randomly setting a proportion of the third navigation field information as invalid information.
[0024] Based on the above technical solution, by randomly dropping out the third navigation information generated by the white-box algorithm during the training process, it can be avoided that the navigation field network excessively relies on the explicit navigation field output by the white-box algorithm, and it is helpful to still make accurate trajectory prediction in the graph-free mode.
[0025] In a third aspect, a driving control apparatus is provided, which includes: a memory configured to store a computer program; and a processor configured to execute the computer program stored in the memory to cause the apparatus to perform any possible method of the first aspect.
[0026] In a fourth aspect, a network training apparatus is provided, which includes: a memory configured to store a computer program; and a processor configured to execute the computer program stored in the memory to cause the apparatus to perform any possible method of the second aspect.
[0027] In a fifth aspect, a vehicle is provided, which includes any possible apparatus of the third aspect or the fourth aspect.
[0028] In a sixth aspect, a computer-readable storage medium is provided, which stores instructions, and the instructions are executed by a processor to cause the processor to implement any possible method of the first aspect or the second aspect.
[0029] In a seventh aspect, a computer program product is provided, which includes computer program code, and when the computer program code is run on a computer, the computer is caused to implement any possible method of the first aspect or the second aspect.
[0030] In an eighth aspect, a chip is provided, which includes a circuit configured to execute any possible method of the first aspect or the second aspect. BRIEF DESCRIPTION OF DRAWINGS
[0031] Figure 1 is a functional block diagram of a vehicle provided by an embodiment of the present application.
[0032] Figure 2 is a schematic diagram of an automatic driving system architecture provided by an embodiment of the present application.
[0033] Figure 3 is a schematic flowchart of a driving control method 300 provided by an embodiment of the present application.
[0034] Figure 4is a schematic diagram of a visual form of navigation information provided by an embodiment of the present application.
[0035] Figure 5 is a schematic diagram of a navigation field network provided by an embodiment of the present application.
[0036] Figure 6 is a schematic diagram of a navigation field network provided by an embodiment of the present application.
[0037] Figure 7 is a schematic diagram of a navigation field network provided by an embodiment of the present application.
[0038] Figure 8 is a schematic flow chart of a network training method 800 provided by an embodiment of the present application.
[0039] Figure 9 is a flow chart of a navigation field network training provided by an embodiment of the present application.
[0040] Figure 10 is a flow chart of a navigation field network application provided by an embodiment of the present application.
[0041] Figure 11 is a schematic block diagram of a driving control device 1100 provided by an embodiment of the present application.
[0042] Figure 12 is a schematic block diagram of a network training device 1200 provided by an embodiment of the present application. DETAILED DESCRIPTION
[0043] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B; in this document, "and / or" is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can mean that there are three cases of A alone, A and B together, and B alone. "At least one" means one or more. For example, "at least one of A and B" is similar to "A and / or B", which describes the association relationship of the associated objects, which means that there can be three relationships, for example, at least one of A and B, which means that there are three cases of A alone, A and B together, and B alone.
[0044] The prefix words such as "first", "second" are used in the embodiments of the present application only to distinguish different description objects, and have no limiting effect on the position, order, priority, quantity or content of the described objects. The use of ordinal words such as "first" in the embodiments of the present application does not constitute a limitation on the described objects, and the description of the described objects should be seen in the context of the claims or embodiments, and should not constitute an unnecessary limitation because of the use of such prefix words. In addition, in the description of the embodiments, unless otherwise stated, the meaning of "a plurality of" is two or more.
[0045] Figure 1 is a functional block diagram of a vehicle 100 provided by the embodiments of the present application. The vehicle 100 can include a perception system 110, a computing platform 120 and a display device 130, wherein the perception system 110 can include one or more sensors that sense information about the environment around the vehicle 100. For example, the perception system 110 can include a positioning system, which can be a global positioning system (GPS), or a Beidou system or other positioning system. For another example, the perception system 110 can include one or more of an inertial measurement unit (IMU), an acceleration sensor, a laser radar, a millimeter wave radar, an ultrasonic radar and a camera.
[0046] Some or all functions of the vehicle 100 can be controlled by the computing platform 120. The computing platform 120 can include one or more processors, such as processors 121 through 12n (n is a positive integer), which are circuits having a processing capability of signals. In one implementation, the processors can be circuits having an instruction reading and running capability, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a kind of microprocessor), a digital signal processor (DSP), or the like. In another implementation, the processors can be circuits having a certain function implemented by a logic relationship of hardware circuits, which is fixed or reconfigurable. For example, the processors can be hardware circuits implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as a field programmable gate array (FPGA). In the reconfigurable hardware circuit, the processor loads a configuration document to implement the hardware circuit configuration. It can be understood that the processor loads instructions to implement the functions of the above part or all units. In addition, the processor can also be a hardware circuit designed for artificial intelligence, which can be understood as a kind of ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), or the like. In addition, the computing platform 120 can further include a memory for storing instructions, and some or all of the processors 121 through 12n can call the instructions in the memory to implement corresponding functions.
[0047] The display device 130 in the cabin is mainly divided into two categories, the first category is a vehicle display screen, and the second category is a projection display screen, such as a head up display (HUD). The vehicle display screen is a physical display screen and is an important component of the in-vehicle infotainment system. Multiple display screens can be provided in the cabin, such as a digital instrument display screen, a center control screen, a display screen in front of a passenger (also referred to as a front passenger) at a co-driver position, a display screen in front of a left rear passenger, and a display screen in front of a right rear passenger, or even a vehicle window can be used as a display screen for display. The head up display, also known as a head-up display system, is mainly used for displaying driving information such as speed, navigation, etc. on a display device (such as a windshield) in front of the driver. This reduces the time for the driver to change his line of sight and avoids changes in the pupil caused by the driver changing his line of sight, thereby improving driving safety and comfort. The HUD includes, for example, a combiner-HUD (C-HUD) system, a windshield-HUD (W-HUD) system, and an augmented reality HUD (AR-HUD). It should be understood that other types of systems can also appear as the technology evolves, and the present application does not limit this.
[0048] The display device 130 described above is illustrated by taking the vehicle display screen and the projection display screen as examples, and embodiments of the present application are not limited thereto. For example, the display device 130 can also be a light display screen or a projection screen.
[0049] Optionally, the structure of the vehicle 100 described above is only schematic, and in actual applications, various components in the vehicle 100 described above can be added or deleted according to actual needs.
[0050] The vehicle 100 can include an intelligent driving system, which can include an advanced driving assistant system (ADAS) and an autonomous driving system (ADS). The intelligent driving system uses various sensors (including but not limited to laser radar, millimeter wave radar, camera, ultrasonic sensor, global positioning system, and inertial measurement unit) on the vehicle to obtain information from the surroundings of the vehicle, and analyzes and processes the obtained information to realize functions such as obstacle perception, target recognition, vehicle positioning, path planning, driver monitoring / reminding, etc., thereby improving the safety, automation level, and comfort of vehicle driving.
[0051] Figure 2A schematic block diagram of the intelligent driving system 200 provided by the embodiments of the present application is shown. The intelligent driving system mainly includes the following modules and networks: a map service module, a static perception module, a navigation field network, a dynamic perception module, a trajectory planning module, and a planning control module.
[0052] The static perception module and the dynamic perception module can perceive the environment around the vehicle body through sensors and output corresponding perception data. The map service module and the static perception module can send static road environment information and navigation information to the navigation field network, and the navigation field network outputs navigation field information. Meanwhile, the dynamic perception module can collect real-time dynamic obstacle information, which can be input into the trajectory planning network together with static obstacle information, static road environment information, and navigation field information. The trajectory planning network can generate a vehicle future trajectory conforming to the navigation logic, such as lane changing, straight driving, left turning, right turning, or U-turn, within a certain time based on the comprehensive information, and input the vehicle future trajectory into the planning control module. The planning control module analyzes the lateral action and longitudinal action of the vehicle future trajectory, and sends control signals to the vehicle actuators according to the analysis results to complete corresponding driving actions (such as turning, accelerating, and decelerating).
[0053] In order to accurately describe the technical content in the present application and accurately understand the content of the present application, before the specific embodiments are described, the following explanations or definitions are given for the terms used in the present specification.
[0054] Lane-level navigation: a high-precision navigation technology that can accurately restore the real road scene and accurately present the number of lanes, ground marking lines, entrances and exits, special lanes, and other information of the current road, and simultaneously combine high-precision satellite signals to display the user's lane in real time.
[0055] White-box algorithm: usually refers to a pre-set algorithm based on explicit rules and logic, whose internal state, decision process, and control logic are completely known to the developer, and has interpretability and traceability. In the embodiments of the present application, the white-box algorithm can be an algorithm module that takes map information and navigation information as input, performs calculation through a series of predetermined rules, and finally outputs a navigation field.
[0056] Navigation field: In physics, the distribution of a certain physical quantity in a region of space can be called a "field" (e.g. temperature field, electric field, magnetic field). In the embodiments of the present application, the "navigation field" can be used to describe the distribution field of continuous vectors or probabilities in the vehicle's surrounding environment that can guide the vehicle to move towards the destination. Among them, the navigation field can be mainly divided into "explicit navigation field" and "implicit navigation field". In the embodiments of the present application, the neural network outputting the navigation field can be referred to as a navigation field network.
[0057] Explicit navigation field: A data structure of the navigation field that can be directly visualized or interpreted, and has a clear physical meaning. For example, in a bird's eye view (BEV), each cell can include one or more direction vectors, and the size of the direction vectors can represent the degree of recommended vehicle travel. This BEV can be a vector form of navigation field; for another example, each cell in the BEV can correspond to a probability value, and the scalar value can represent the size of the probability of the future trajectory of the vehicle passing through the position. The higher the probability value, the more the position conforms to the intention of the vehicle navigation.
[0058] In the embodiments of the present application, the explicit navigation field can be generated by a white-box algorithm or a navigation field network. The explicit navigation field can be used as a supervised true value in training, and can also provide an explicit reference for trajectory planning.
[0059] Implicit navigation field: In the embodiments of the present application, the implicit navigation field can be high-dimensional and abstract feature information of the intermediate layer input of the navigation field network. The implicit navigation field itself does not have a direct visualized physical meaning, but in the training process of the navigation field network, it internalizes the navigation-related information (road topology, traffic rules, navigation intention) in its own data structure through encoding. For example, the implicit navigation field can be a high-dimensional tensor.
[0060] In the embodiments of the present application, the implicit navigation field can be used to provide information-intensive and real-time navigation guidance for trajectory prediction.
[0061] Map-based mode: Generally refers to a mode in which the automatic driving of the vehicle highly depends on the pre-made HD map. The vehicle can accurately locate itself in the HD map through positioning technology, so as to pre-judge the road environment information of the super field of view.
[0062] No-map mode: usually refers to a mode in which the automatic driving of a vehicle relies only on a standard definition map (SD map) and mainly relies on its own real-time perception ability to identify road structures and make decisions. The vehicle can perceive elements such as routes, road edges, and traffic signs in real time through sensors such as cameras and lidar, and can construct a local road environment awareness in real time, and make path planning accordingly.
[0063] Turn-by-turn (TBT) navigation information: information of a navigation service that can obtain a driving path according to a starting point and a target of a vehicle, and can be updated in real time. The TBT navigation information can provide specific turning instructions (such as left turn, right turn, or straight ahead) at each intersection, and can also indicate the angle of turning, for example. The TBT navigation information can be combined with voice instructions or visual instructions (such as arrow instructions on a center screen) to help the driver maintain attention during navigation.
[0064] Query point: in the embodiments of the present application, each query point can be used to indicate a point on the center line of a lane in a plurality of lanes, and the query point can be used to query the degree of passability or the degree of recommended driving of each position in a navigation field.
[0065] Lane remain: represents the remaining drivable distance of a vehicle in different lanes on a driving route that meets the navigation intention. Generally, the value of lane remain can be used to represent the degree of passability or the degree of navigation recommendation of the current lane, and the greater the value of lane remain, the higher the degree of passability. If lane remain is less than zero, it can indicate that the current vehicle driving route has deviated.
[0066] Route calculation result (sdlink): in the present application, it represents the information of a road from the current position of a vehicle to a target position.
[0067] Electric horizon (EHP) data: EHP is used to provide vehicles with information about road traffic beyond the range of sensors, such as road curvature, slope, lane shape, and traffic light information beyond the range of sensors. Generally, there are two cases of beyond the range of sensors, one is distance, and the other is field of view and occlusion and blind area in special scenarios such as intersections and curves.
[0068] In recent years, with the development of artificial intelligence, end-to-end autonomous driving technology has become a research hotspot due to its strong environmental perception and decision-making capabilities. This type of autonomous driving technology can use deep neural networks to improve the capabilities of intelligent driving systems in navigation, lane changing games, obstacle avoidance, and other scenarios through supervised learning or reinforcement learning.
[0069] However, end-to-end solutions still have some shortcomings in achieving lane-level navigation, a crucial aspect. For example, in one implementation, the vehicle relies on predefined goal points for navigation and guidance. This approach requires complex white-box algorithms, combined with HD maps and the vehicle's navigation information, to generate goal points in front of the vehicle, which then plans its trajectory based on these points. This implementation method has the following main drawbacks:
[0070] (1) High dependence on HD map: HD map is expensive to produce and maintain, and has a low update frequency, making it difficult to respond to real-time changes in the real world, which leads to a decrease in the reliability of the planned trajectory;
[0071] (2) The strategy based on goal point traction may limit the solution space of trajectory planning: regardless of whether the scheme is based on a single goal point traction, a single row of goal points traction, or a multi-row of goal points traction, the vehicle's lateral and longitudinal maneuverability is insufficient, making it difficult to make better decisions that conform to the navigation intent in complex and sudden scenarios (such as lane changing in dense traffic flow, emergency obstacle avoidance, etc.), thus restricting the vehicle's driving ability.
[0072] Among them, the single-row goal point traction scheme and the multi-row goal point traction scheme need to couple the navigation information generated by the white-box algorithm to obtain lateral lane selection capability to a certain extent. Although their lateral flexibility is better than the single goal point traction scheme, their overall performance is still not good, and the flexibility of the vehicle's longitudinal movement is still restricted.
[0073] Therefore, there is an urgent need for a driving control method that does not rely on HD maps and goal points, and can achieve more flexible and real-time vehicle control in both map-less and map-based modes.
[0074] The following describes the driving control method, network training method apparatus, and vehicle provided in the embodiments of this application. The driving control method provided in the embodiments of this application, by using the implicit navigation field output by the navigation field network as one of the inputs to the trajectory generation network, helps the vehicle to reduce its over-reliance on HD maps and goal points for trajectory planning. This reduces the vehicle's map maintenance costs and enables lane-level driving control. While adhering to the vehicle's navigation intent, it improves the vehicle's flexibility in complex scenarios, including lane changing and obstacle avoidance, thus providing a guarantee for a safer and more accurate autonomous driving experience.
[0075] Figure 3 A schematic flowchart of a driving control method 300 provided in an embodiment of this application is shown. Figure 3The method 300 includes:
[0076] S301: Obtain road environment information, obstacle information, and navigation information, the navigation information including driving action information and lane selection information for guiding the vehicle to travel from a current position to a target position;
[0077] The road environment information can be a set of static elements for describing the physical layout, geometric properties, topological connection relationships, and rule constraints of the road, for example, can include perceived road elements and structure information, including road points, roads, lanes, lane road boundaries, etc.
[0078] For example, the road environment information can include but is not limited to: the position of the lane, the width of the lane, the number of lanes, the curvature of the curve, the coordinates of the traffic sign, the type of traffic sign (such as speed limit, stop, no entry, etc.), road markings (such as arrows, text, lane lines, pedestrian crossings), the number of lanes, the curvature of the curve, the layout of the intersection, and other geometric features of the road.
[0079] The obstacle information can include static obstacle information and dynamic obstacle information. The static obstacle information can be an object whose position and shape remain fixed for a relatively long period of time. The dynamic obstacle information can be an object whose position, pose, or shape changes over time.
[0080] For example, the static obstacle information can include but is not limited to: road infrastructure such as curbs, medians, central barriers, lampposts, utility poles, traffic sign posts, toll booths, safety islands, walls, etc.; temporary road facilities such as traffic cones, construction fences; stationary vehicles such as vehicles parked on the roadside; road debris such as dropped cargo, stones, low branches, etc.
[0081] For example, dynamic obstacles can include but are not limited to: vehicles in motion (including motor vehicles and non-motor vehicles), pedestrians within the field of view of the vehicle, animals crossing the road (such as poultry, wildlife, pets, etc.), falling cargo, etc.
[0082] The navigation information is used to guide the driving action information and lane selection information for the vehicle to travel from the current position to the target position.
[0083] For example, the driving action information can include turning information (left turn, right turn, straight ahead, U-turn, etc.), lane selection information (left lane, right lane, left two lanes, straight lane, etc.), entry / exit instructions (entering a roundabout, exiting a roundabout from a certain exit, entering a ramp, exiting a ramp, etc.), destination indication information (such as the destination is on your left / right, etc.).
[0084] For example, navigation information can also inform drivers in advance of each intersection or node where a change of direction is required, such as "Please turn left 100 meters ahead", "Please go straight for one kilometer ahead", or "Please move to the left lane as soon as possible".
[0085] For example, navigation information can be obtained based on map data and real-time traffic conditions using rule-transparent white-box algorithms (such as path planning methods based on fixed rules or shallow machine learning models). The obtained navigation information can be "a sequence of driving actions from the current location to the target location + lane selection suggestions".
[0086] In one implementation, the navigation information may include: TBT navigation information.
[0087] In one embodiment, navigation information may be presented as vector information; in another embodiment, navigation information may also be presented as graphical information. In one embodiment, navigation information may be mapped to a bird's-eye view (BEV).
[0088] Figure 4 This diagram illustrates a visual representation of navigation information. For example... Figure 4 As shown, the diagonally filled arrows represent the roads a vehicle must travel to reach its target location. The arrows indicate the direction the vehicle will travel along the road. For example, when a vehicle travels to its target location, it needs to pass through road 1, road 2, road 3, and road 4 in sequence (here are some examples of at least one road). Furthermore, the vehicle may pass through at least one intersection during its journey to the target location, such as at least one of intersection 1, intersection 2, and intersection 3.
[0089] For example, the driving direction at each intersection can be indicated by one or more of the following information: ① Intersection guidance, such as left turn, straight, or right turn; ② Intersection topology and the target road recommended by navigation (e.g., black arrows indicating the road a vehicle needs to enter after passing the intersection), wherein the topology at each intersection can indicate the two or more roads connected to the intersection, and can also indicate the direction of each road; ③ Lane information, which indicates the number of lanes included in the road, the guidance supported by each lane, and the availability status of each lane guidance (e.g., arrows with filled horizontal lines), wherein the availability status of each lane guidance can indicate whether a vehicle can drive in that lane based on that guidance during its journey to the target.
[0090] like Figure 4As shown, a set of arrows associated with the lane information of intersection 1 indicates that road 1 includes three lanes: the leftmost lane supports vehicles to go straight and turn left at intersection 1, the middle lane supports vehicles to go straight at intersection 1, and the rightmost lane supports vehicles to turn right at intersection 1. Further, the arrows filled with horizontal lines indicate that vehicles can go straight in the leftmost lane, or vehicles can go straight in the middle lane, and the white arrows not filled indicate that vehicles cannot go straight in the corresponding lanes according to the guidance of the white arrows.
[0091] S302: input the road environment information and the navigation information into a navigation field network to obtain first navigation field information, the navigation field network includes an encoding network and a decoding network, and the first navigation field information is feature information output by the encoding network;
[0092] In some embodiments of the present application, the first navigation field information can also be referred to as an implicit navigation field. It can be a high-dimensional, dense, abstract feature tensor that encodes the complex spatial and semantic relationships of the environment scene (road environment information) and navigation information around the vehicle, and can implicitly guide the trajectory planning of the vehicle.
[0093] Optionally, inputting the road environment information and the navigation information into the navigation field network to obtain the first navigation field information includes: obtaining first fusion features based on the road environment information and the navigation information; and obtaining the first navigation field information based on the first fusion features.
[0094] As can be seen from the above, the navigation field network includes an encoding network and a decoding network. Exemplarily, the encoding network includes a first network, a second network, and a third network. Inputting the road environment information and the navigation information into the encoding network to obtain the first navigation field information can include: inputting the road environment information into the first network to obtain an output result of the first network; inputting the output result of the first network and the navigation information into the second network to obtain an output result of the second network; and inputting the output result of the second network and part of the road environment information (such as ground arrows, road boundaries, intersection surfaces, sdlinks, etc.) into the third network to obtain the first navigation field information.
[0095] It should be noted that inputting the road environment information into the first network can be: inputting the encoded road environment information into the first network; inputting the output result of the first network and the navigation information into the second network can be: inputting the output result of the first network and the encoded navigation information into the second network; and inputting the output result of the second network and part of the road environment information into the third network can be: inputting the output result of the second network and the encoded part of the road environment information into the third network.
[0096] Exemplarily, Figure 5 A schematic diagram of the navigation field network provided by an embodiment of the present application is shown. As shown in FIG. 2, the navigation field network includes an encoding network and a decoding network.Figure 5 As shown, the navigation field network is divided into an encoding network and a decoding network. The encoding network can include an encoding layer a, an encoding layer b, an encoding layer c, a self-attention network, a cross-attention network a, and a cross-attention network b. The self-attention network can be an example of the first network, the cross-attention network a can be an example of the second network, and the cross-attention network b can be an example of the third network. The cross-attention network b can input the output first navigation field information into a cross-attention model c of the decoding network to obtain an output result of the cross-attention network c, and input the output result of the cross-attention network c into a decoding layer to obtain a lane remain or a probability value. As shown in FIG. 2B, the decoding network can include a decoding layer a, a decoding layer b, a decoding layer c, a cross-attention model a, and a cross-attention model b. The cross-attention model a can be an example of the fourth network, and the cross-attention model b can be an example of the fifth network. Figure 5 As shown, in some embodiments, the probability value or the lane remain output by the decoding network can be concatenated with the query point to obtain second navigation field information.
[0097] In some embodiments, the navigation field network can also include multiple self-attention networks and cross-attention networks, as shown in FIG. 2C. Figure 6 As shown, the encoding network can include two self-attention networks and n network model groups, each model group including two self-attention networks, a cross-attention network a, and a cross-attention network b. The input of the last group of the n network model groups can be the output of the previous group, the input of the cross-attention model a in each model group can be the output of the previous model and navigation information, and the input of the cross-attention model b in each model group can be the output of the previous model and part of the road environment information. For example, n can be 4 or 6. In addition, the number of self-attention networks and cross-attention networks in the network model group is not limited. Figure 5 The structure shown is only an example of the navigation field network.
[0098] In actual implementation, Figure 5 The decoding layer in FIG. 2B can be a MLP. Figure 6 The encoding layer shown can be an artificial neural network (ANN), such as a multilayer perceptron (MLP) or a convolutional neural network (CNN). Figure 5 The decoding layer in FIG. 2B can be a MLP. Figure 6 The decoding layer in FIG. 2B can be a MLP.
[0099] Optionally, the method 300 further comprises: obtaining query point information, the query point information being indicative of point information on the road surface; inputting the road environment information, the obstacle information, and the first navigation field information into the trajectory generation network to obtain first trajectory information, including: obtaining second fusion features based on the road environment information, the obstacle information, the first navigation field information, and the query point information; and obtaining the first trajectory information based on the second fusion features.
[0100] In some embodiments, as shown in Figure 7 The encoding network of the navigation field network can further include an encoding layer d, which can be used to process the obtained query point information to extract query point features. Alternatively, in another embodiment, the encoding layer d can input the road environment information, and process the road environment information by the encoding layer d to autonomously extract a query point set to obtain the query point features. The cross-attention model c simultaneously receives the first navigation field information and the query point features, and obtains the second navigation field information through the decoding layer.
[0101] S303: inputting the road environment information, the obstacle information, and the first navigation field information into the trajectory generation network to obtain first trajectory information;
[0102] In some embodiments, the road structure information, the obstacle information, and the first navigation field information (implicit navigation field) generated by the encoding network in the above navigation field network are input into the trajectory generation network together. The trajectory generation network performs trajectory prediction based on these different inputs. The first navigation field information can be feature information that is easier for machines to understand, which helps to improve the accuracy of trajectory prediction.
[0103] S304: controlling the vehicle to travel based on the first trajectory information.
[0104] The vehicle can calculate specific throttle, pedal, brake, and other instructions based on the third trajectory information predicted by the trajectory generation network, so as to control the vehicle to travel accurately along the predicted trajectory.
[0105] Figure 8 A schematic flowchart of a network training method 800 provided by an embodiment of the present application is shown. The method 800 is mainly used to describe how to train the navigation field network. As Figure 8 shown, the method 800 includes:
[0106] S801: obtaining road environment information, obstacle information, navigation information, and second trajectory information, the navigation information including travel action information and lane selection information guiding the vehicle to travel from a current position to a target position, and the second trajectory information representing standard trajectory information that the vehicle can refer to;
[0107] S802: input the road environment information and the navigation information into the navigation field network to obtain first navigation field information, the navigation field network comprises an encoding network and a decoding network, and the first navigation field information is feature information output by the encoding network;
[0108] In some embodiments, before training the navigation field network, a training data set needs to be obtained, wherein the training data set can include scene data and supervised true value information, the scene data can include lane center line, road boundary, ground arrow, TBT navigation information, and the supervised true value can include two categories, one is a trajectory in a real driving scene, and the other is an explicit navigation field generated by a white-box algorithm combined with a map and navigation information.
[0109] In addition, in addition to the training data set, a pre-trained navigation field network is also needed. The pre-trained network can be an initial navigation field network trained based on the scene data, the navigation field network can learn the mapping of the explicit navigation field from the road environment information and the navigation information, and the encoding part of the pre-trained network is retained, the hidden layer feature output by the encoding part is the implicit navigation field, that is, the first navigation field information, which encodes rich semantic information for navigation.
[0110] S803: input the road environment information, the obstacle information, and the first navigation field information into the trajectory generation network to obtain third trajectory information;
[0111] S804: adjust the parameters of the navigation field network based on the second trajectory information and the third trajectory information.
[0112] In some embodiments, the road environment information, the obstacle information, and the first navigation field information are input into the trajectory generation network as inputs of the trajectory generation network, all data information is fused through a series of interaction modules, and a future predicted trajectory of the vehicle is generated by decoding. Then, a loss value can be calculated based on the future predicted trajectory of the vehicle and the supervised true value information in the training data set.
[0113] Optionally, inputting the road environment information and the navigation information into the navigation field network to obtain the first navigation field information comprises: inputting the road environment information and the navigation information into the navigation field network to obtain first navigation field information and second navigation field information, the second navigation field information being feature information output by the decoding network; and adjusting the parameters of the navigation field network based on the second trajectory information and the third trajectory information comprises: adjusting the parameters of the navigation field network based on the second trajectory information, the third trajectory information, and the second navigation field information.
[0114] In some embodiments, the training effect can be enhanced by introducing an explicit navigation field in the training of the navigation field network. In the training process, the first navigation field information (implicit navigation field) and the second navigation field information (explicit navigation field) output by the encoding network of the navigation field network can be obtained at the same time, and the two are homologous. In the process of calculating the loss function and adjusting the parameters of the navigation field network, in addition to comparing the predicted trajectory and the trajectory true value, the second navigation field information is additionally introduced to assist in adjusting the parameters of the navigation field network.
[0115] Optionally, the method 800 further comprises: obtaining query point information, the query point information indicating point position information on the road surface; and the second navigation field information being obtained by the decoding network based on the road environment information, the navigation information, and the query point information.
[0116] In some embodiments, in order to generate an explicit navigation field (second navigation field information), the decoding network of the navigation field network also needs to input additional query point information, which can be distributed on the centerlines of different lanes and can be uniformly distributed or non-uniformly distributed. For each query point, a value (for example, a lane remain or a probability value) at the point position can be output by the navigation field network in combination with the road environment information and the navigation information. The values of all query points are combined to form a complete second navigation field information.
[0117] In some embodiments, in the training process of the navigation field network, the navigation field network needs to be more sensitive to serious prediction errors. For example, a negative loss can be introduced to dynamically adjust the parameters of the navigation field network. For example, in the case of detecting a serious error, the value of the loss function can be dynamically increased. By such a method, the probability of predicting dangerous and erroneous trajectories can be effectively reduced.
[0118] In the embodiments of the present application, two vehicle route planning modes, i.e., a graph mode and a graph-free mode, can be adapted.
[0119] For example, in the graph mode, the vehicle can use the explicit navigation field obtained based on the white-box algorithm as a supervision signal.
[0120] For example, in the graph-free mode, the vehicle can use a "random drop" training strategy to make the trajectory generation network more dependent on the implicit navigation field for trajectory planning.
[0121] Optionally, before the road environment information, the obstacle information, the navigation information, and the second trajectory information are acquired, the method further includes: acquiring map information; obtaining third navigation field information based on the map information and the navigation information by using a white-box algorithm; and inputting the road environment information, the obstacle information, and the first navigation field information into the trajectory generation network to obtain the third trajectory information, including: inputting the road environment information, the obstacle information, the first navigation field information, and the third navigation field information into the trajectory generation network to obtain the third trajectory information.
[0122] For example, in some scenarios with map mode, third navigation information can be generated based on a white-box algorithm, using high-precision maps and navigation information. The third navigation information has strong interpretability and is relatively accurate, and therefore, the third navigation information can be used as one of the inputs of the trajectory generation network, so as to generate a safer and more reliable trajectory.
[0123] Optionally, before the third trajectory information is obtained, the method 800 further includes: randomly setting a certain proportion of the third navigation field information as invalid information.
[0124] The training method provided in the embodiments of the present application considers enhancing the trajectory generation capability of a vehicle in a no-map mode. During training, when the third navigation information generated by the aforementioned white-box algorithm is used for training, the navigation field network can randomly discard a certain proportion of information in the third navigation information, for example, the certain proportion of information can be set as invalid. Through such an implementation, the navigation field network can not excessively rely on the third navigation information provided by the white-box algorithm, but tend to learn how to mainly rely on the first navigation field information (implicit navigation field) to work, which can enable the navigation field network of the present application to perform well in both the no-map mode and the map mode.
[0125] Figure 9 A training flowchart of a navigation field network provided in the embodiments of the present application is shown. Figure 9 The training flowchart of the navigation field network based on supervised learning is mainly shown, and the core is to construct a navigation field based on multi-source features and optimize the model. The following is a step-by-step introduction to the training flowchart of the navigation field network based on supervised learning: Figure 9
[0126] Firstly, Figure 9 The input layer in the navigation field network based on supervised learning is a supervised training set, mainly including "map, dynamic and static features" and "map, static road features". The "map, dynamic and static features" can be used to extract dynamic (such as vehicles and pedestrians) and static (such as buildings and road profiles) environmental information in a scene. The "map, static road features" can be used to focus on the static structure (such as lane lines and intersection layout) of the road network.
[0127] Next, the dynamic and static feature coding module can process the "map, dynamic and static features" to obtain scene features (such as semantic segmentation results, target detection features, etc.), providing a basis for subsequent interaction.
[0128] At the same time, inputting the "map, static road features" into the navigation field network can generate two kinds of navigation fields: implicit navigation field and explicit navigation field. Among them, the implicit navigation field can be the implicit law of learning the road structure and navigation information through the neural network. The explicit navigation field can be the interpretable information output by the navigation field network.
[0129] After that, the scene features and the implicit navigation field are jointly input into the interaction module, which can fuse the environmental perception and road information, and finally output the trajectory of the vehicle (predicting the vehicle trajectory in the future for a period of time).
[0130] In some embodiments, the real trajectory of the vehicle and the predicted trajectory of the vehicle can be compared to calculate the loss function (such as mean square error, cross entropy, etc.), and the optimizer of the navigation field network can adjust the parameters of the navigation field network (such as the weights of the dynamic and static coding module, the navigation field network) after receiving the loss value through gradient backpropagation.
[0131] In addition, the explicit navigation field generated by the navigation field network can also be used to calculate the loss function, so as to train a more accurate navigation field model.
[0132] Figure 10 An application flowchart of a navigation field network provided by an embodiment of the present application is shown. Figure 10 It mainly shows the "perception-trajectory planning" process in the automatic driving process, and focuses on the fusion of multi-source information and the construction of implicit navigation field. The following is a hierarchical introduction to Figure 10 :
[0133] The sensors of the vehicle can be used to acquire environmental information, and the sensors can be, for example, cameras, radars, lidars, etc., for collecting raw environmental data. Then inputting the raw environmental data into the perception network can obtain static road structure and dynamic and static road structure, wherein the static road structure can include, for example, the fixed topological relationship of the road (such as the lane, intersection layout, etc.), and the dynamic and static obstacles can include, for example, the positions of the obstacles in the environment (such as other vehicles, pedestrians) or stationary obstacles (such as roadblocks, buildings, walls) etc.
[0134] In an embodiment of the present application, an explicit navigation field (third navigation field information) can be generated based on map data and navigation information using a white-box algorithm. The map data can be high-precision map data, and the explicit navigation field can be one of the optional inputs of the trajectory generation network.
[0135] In the embodiments of the present application, the navigation information and the road environment information can be input into an encoding network of a navigation field network, to output an implicit navigation field (first navigation field information), which needs to be input into a trajectory generation network. In addition, the output of the encoding network and the lane-level query point can also be input into a decoding network as input, to output an explicit navigation field (second navigation field information), which can be used to calculate a loss value together with the aforementioned explicit navigation field (third navigation field information). The loss value can be used to adjust the parameters of the network through gradient backpropagation.
[0136] The trajectory generation network can integrate several types of input information: static road structure, static and dynamic obstacle features, implicit navigation field (first navigation field information), and white-box generated explicit navigation field (third navigation field information, which is optional). Finally, a driving trajectory that meets the safety and navigation needs is output.
[0137] Figure 11 A schematic block diagram of a driving control device 1100 provided by the embodiments of the present application is shown. The device 1100 includes an acquisition unit 1110 configured to acquire road environment information, obstacle information, and navigation information, the navigation information including driving action information and lane selection information for guiding a vehicle to drive from a current position to a target position, and the navigation information being information generated based on map information through a white-box algorithm. The acquisition unit 1110 is further configured to input the road environment information and the navigation information into a navigation field network to obtain first navigation field information, the navigation field network including an encoding network and a decoding network, and the first navigation field information being feature information output by the encoding network. The acquisition unit 1110 is further configured to input the road environment information, the obstacle information, and the first navigation field information into a trajectory generation network to obtain first trajectory information. A control unit 1120 is configured to control the vehicle to drive based on the first trajectory information.
[0138] Optionally, the acquisition unit 1110 is further configured to obtain first fusion features based on the road environment information and the navigation information. The acquisition unit 1110 is further configured to obtain the first navigation field information based on the first fusion features.
[0139] Optionally, the acquisition unit 1110 is further configured to acquire query point information, the query point information indicating point information on a road surface. The acquisition unit 1110 is further configured to obtain second fusion features based on the road environment information, the obstacle information, the first navigation field information, and the query point information. The acquisition unit 1110 is further configured to obtain the first trajectory information based on the second fusion features.
[0140] Figure 12A schematic block diagram of the network training apparatus 1200 provided by the embodiments of the present application is shown. The apparatus 1200 comprises: an obtaining unit 1210, configured to: obtain road environment information, obstacle information, navigation information, and second trajectory information, the navigation information comprising driving action information and lane selection information for guiding a vehicle to travel from a current position to a target position, and the second trajectory information representing standard trajectory information that can be referenced by the vehicle; the obtaining unit 1210 is further configured to: input the road environment information and the navigation information into a navigation field network to obtain first navigation field information, the navigation field network comprising an encoding network and a decoding network, and the first navigation field information being feature information output by the encoding network; and the obtaining unit 1210 is further configured to: input the road environment information, the obstacle information, and the first navigation field information into a trajectory generation network to obtain third trajectory information; and an adjusting unit 1220, configured to: adjust parameters of the navigation field network based on the second trajectory information and the third trajectory information.
[0141] Optionally, the obtaining unit 1210 is further configured to: input the road environment information and the navigation information into the navigation field network to obtain the first navigation field information and second navigation field information, the second navigation field information being feature information output by the decoding network; and the adjusting unit 1220 is configured to: adjust the parameters of the navigation field network based on the second trajectory information, the third trajectory information, and the second navigation field information.
[0142] Optionally, the obtaining unit 1210 is further configured to: obtain query point information, the query point information indicating point position information on a road surface; and the second navigation field information is obtained by the decoding network based on the road environment information, the navigation information, and the query point information.
[0143] Optionally, the obtaining unit 1210 is further configured to: obtain map information; the obtaining unit 1210 is further configured to: obtain third navigation field information by a white-box algorithm based on the map information and the navigation information; and the obtaining unit 1210 is further configured to: input the road environment information, the obstacle information, the first navigation field information, and the third navigation field information into the trajectory generation network to obtain the third trajectory information.
[0144] Optionally, the adjusting unit 1220 is further configured to: randomly set a certain proportion of the third navigation field information as invalid information.
[0145] It should be understood that the division of each unit in the above device is only a logical functional division, and all or part of the units can be integrated into a physical entity, or can be physically separated. In addition, the units in the device can be implemented in the form of processor calling software; for example, the device includes a processor, the processor is connected with a memory, the memory stores instructions, and the processor calls the instructions stored in the memory to implement any of the above methods or to realize the functions of the units of the device, wherein the processor is, for example, a general processor such as a CPU or a microprocessor, and the memory is a memory in the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of hardware circuit, and the functions of part or all of the units can be realized by the design of the hardware circuit, which can be understood as one or more processors; for example, in one implementation, the hardware circuit is an ASIC, and the functions of part or all of the units are realized by the design of the logical relationship of elements in the circuit; for example, in another implementation, the hardware circuit is a PLD, and taking FPGA as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by a configuration file, so as to realize the functions of part or all of the units. All units of the above device can be implemented in the form of processor calling software, or all units can be implemented in the form of hardware circuit, or part of the units can be implemented in the form of processor calling software, and the remaining part can be implemented in the form of hardware circuit.
[0146] Each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0147] In addition, all or part of each unit in the above device can be integrated together or can be independently implemented. In one implementation, the units are integrated together to realize the form of SoC. The SoC can include at least one processor for implementing any of the above methods or realizing the functions of the units of the device, and the types of the at least one processor can be different, such as CPU and FPGA, CPU and artificial intelligence processor, CPU and GPU, etc.
[0148] Embodiments of the present application also provide a driving control device, which includes a memory for storing a computer program, and a processor for executing the computer program stored in the memory to enable the device to perform the method or steps performed by the above embodiments.
[0149] The embodiment of the present application further provides a network training device, which comprises a memory for storing a computer program; and a processor for executing the computer program stored in the memory, so that the device executes the method or the steps executed by the above-mentioned embodiment.
[0150] Optionally, if the driving control device and the network training device are located in a vehicle, the processor can be the processor 121-12n shown in the figure. Figure 1
[0151] The embodiment of the present application further provides a vehicle, which can comprise the driving control device 1100.
[0152] The embodiment of the present application further provides a computer readable storage medium, which stores instructions, and the instructions are executed by a processor to enable the processor to implement the method in the above-mentioned embodiment.
[0153] The embodiment of the present application further provides a computer program product, which comprises computer program codes, and when the computer program codes are executed on a computer, the computer executes the method in the above-mentioned embodiment.
[0154] The embodiment of the present application further provides a chip, which comprises a circuit for executing the method in the above-mentioned embodiment.
[0155] In the implementation process, each step of the above-mentioned method can be completed by integrated logic circuits of hardware in the processor or instructions in the form of software. The method disclosed in the embodiment of the present application can be directly embodied as hardware processor execution completion or combined execution completion by hardware and software modules in the processor. The software module can be located in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable read-only memory, register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps of the above-mentioned method. To avoid repetition, it will not be described in detail here.
[0156] It should be understood that the memory in the embodiment of the present application can comprise read-only memory and random access memory, and provide instructions and data for the processor.
[0157] It should be further understood that in various embodiments of the present application, the size of the serial number of each process does not mean the execution order, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.
[0158] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0159] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0160] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the above-described device embodiments are merely schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0161] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment.
[0162] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.
[0163] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0164] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A travel control method characterized by comprising: The method comprises: obtaining road environment information, obstacle information and navigation information, the navigation information comprising driving action information and lane selection information for guiding a vehicle to travel from a current position to a target position, the navigation information being information generated based on map information by a white-box algorithm; inputting the road environment information and the navigation information into a navigation field network to obtain first navigation field information, the navigation field network comprising an encoding network and a decoding network, the first navigation field information being feature information output by the encoding network; inputting the road environment information, the obstacle information and the first navigation field information into a trajectory generation network to obtain first trajectory information; controlling the vehicle to travel based on the first trajectory information.
2. The method of claim 1, wherein, The inputting of the road environment information and the navigation information into the navigation field network to obtain the first navigation field information comprises: obtaining first fusion features based on the road environment information and the navigation information; obtaining the first navigation field information based on the first fusion features.
3. The method according to claim 1 or 2, characterized in that, The method further comprises: obtaining query point information, the query point information indicating point information on a road surface; The inputting of the road environment information, the obstacle information and the first navigation field information into the trajectory generation network to obtain the first trajectory information comprises: obtaining second fusion features based on the road environment information, the obstacle information, the first navigation field information and the query point information; obtaining the first trajectory information based on the second fusion features.
4. A network training method, comprising: The method comprises: obtaining road environment information, obstacle information, navigation information and second trajectory information, the navigation information comprising driving action information and lane selection information for guiding a vehicle to travel from a current position to a target position, the second trajectory information representing standard trajectory information that can be referenced by the vehicle; inputting the road environment information and the navigation information into a navigation field network to obtain first navigation field information, the navigation field network comprising an encoding network and a decoding network, the first navigation field information being feature information output by the encoding network; inputting the road environment information, the obstacle information and the first navigation field information into a trajectory generation network to obtain third trajectory information; adjusting parameters of the navigation field network based on the second trajectory information and the third trajectory information.
5. The method of claim 4, wherein, The inputting of the road environment information and the navigation information into the navigation field network to obtain the first navigation field information comprises: inputting the road environment information and the navigation information into the navigation field network to obtain the first navigation field information and second navigation field information, the second navigation field information being feature information output by the decoding network; The adjusting of the parameters of the navigation field network based on the second trajectory information and the third trajectory information comprises: adjusting the parameters of the navigation field network based on the second trajectory information, the third trajectory information and the second navigation field information.
6. The method of claim 5, wherein, The method further comprises: obtaining query point information, the query point information indicating point information on a road surface; The second navigation field information is obtained by the decoding network based on the road environment information, the navigation information, and the query point information.
7. The method according to any one of claims 4-6, characterized in that, Before the road environment information, the obstacle information, the navigation information, and the second trajectory information are obtained, the method further includes: obtaining map information; obtaining third navigation field information based on the map information and the navigation information by a white-box algorithm; The third trajectory information is obtained by inputting the road environment information, the obstacle information, the first navigation field information, and the third navigation field information into the trajectory generation network. Before the third trajectory information is obtained, the method further includes:
8. The method of claim 7, wherein, randomly setting a certain proportion of the third navigation field information as invalid information. The device includes:
9. A travel control device characterized by comprising: an obtaining unit, configured to obtain road environment information, obstacle information, and navigation information, the navigation information including driving action information and lane selection information for guiding a vehicle to travel from a current position to a target position, the navigation information being information generated based on map information by a white-box algorithm; The obtaining unit is further configured to input the road environment information and the navigation information into a navigation field network to obtain first navigation field information, the navigation field network including an encoding network and a decoding network, and the first navigation field information being feature information output by the encoding network; The obtaining unit is further configured to input the road environment information, the obstacle information, and the first navigation field information into a trajectory generation network to obtain first trajectory information. a control unit, configured to control the vehicle to travel based on the first trajectory information.
10. The device of claim 9, wherein The obtaining unit is further configured to obtain first fusion feature based on the road environment information and the navigation information. The obtaining unit is further configured to obtain the first navigation field information based on the first fusion feature.
11. The device of claim 9 or 10, wherein The obtaining unit is further configured to obtain query point information, the query point information indicating point information on a road surface; The obtaining unit is further configured to obtain second fusion feature based on the road environment information, the obstacle information, the first navigation field information, and the query point information. The obtaining unit is further configured to obtain the first trajectory information based on the second fusion feature. The device includes:
12. A network training apparatus, comprising: an obtaining unit, configured to obtain road environment information, obstacle information, and navigation information, the navigation information including driving action information and lane selection information for guiding a vehicle to travel from a current position to a target position, the navigation information being information generated based on map information by a white-box algorithm; The obtaining unit is further configured to input the road environment information and the navigation information into a navigation field network to obtain first navigation field information, the navigation field network including an encoding network and a decoding network, and the first navigation field information being feature information output by the encoding network; The acquisition unit is further configured to input the road environment information, the obstacle information, and the first navigation field information into a trajectory generation network to obtain third trajectory information. The adjustment unit is configured to adjust parameters of the navigation field network based on the second trajectory information and the third trajectory information.
13. The apparatus of claim 12, wherein, The acquisition unit is further configured to input the road environment information and the navigation information into the navigation field network to obtain the first navigation field information and second navigation field information, the second navigation field information being feature information output by the decoding network. The adjustment unit is configured to adjust parameters of the navigation field network based on the second trajectory information, the third trajectory information, and the second navigation field information.
14. The apparatus of claim 13, wherein, The acquisition unit is further configured to acquire query point information, the query point information indicating point information on a road surface; and the second navigation field information is obtained by the decoding network based on the road environment information, the navigation information, and the query point information.
15. The apparatus of any one of claims 12-14, wherein, The acquisition unit is further configured to acquire map information. The acquisition unit is further configured to obtain third navigation field information by a white-box algorithm based on the map information and the navigation information. The acquisition unit is further configured to input the road environment information, the obstacle information, the first navigation field information, and the third navigation field information into a trajectory generation network to obtain third trajectory information.
16. The apparatus of claim 15, wherein, The adjustment unit is further configured to randomly set a certain proportion of the third navigation field information as invalid information.
17. A travel control device characterized by comprising: comprising: a memory configured to store a computer program; a processor configured to execute the computer program stored in the memory, so that the apparatus executes the method of any one of claims 1-3.
18. A network training apparatus, comprising: comprising: a memory configured to store a computer program; a processor configured to execute the computer program stored in the memory, so that the apparatus executes the method of any one of claims 4-8.
19. A vehicle characterized by comprising: comprising: the apparatus of any one of claims 9-11, or the apparatus of any one of claims 12-16, or the apparatus of claim 17, or the apparatus of claim 18.
20. A computer-readable storage medium, characterized in that, instructions stored thereon, which, when executed by a processor, cause the processor to implement the method of any one of claims 1-3, or the method of any one of claims 4-8.
21. A computer program product, characterised in that, The computer program product comprises computer program code which, when executed on a computer, causes the computer to implement the method of any one of claims 1-3, or the method of any one of claims 4-8.
22. A chip, characterized by The chip comprises a circuit configured to execute the method of any one of claims 1-3, or the method of any one of claims 4-8.