Port driving scene construction method and device, equipment and storage medium

By utilizing the port driving model and optimized reward function built on the experience of agents and experts based on the large language model, combined with the reinforcement learning algorithm, the problem that existing data sets are difficult to apply to the port environment is solved, and a port driving scenario with a high degree of authenticity is constructed, the performance of the autonomous driving algorithm is verified and the application prospects are improved.

CN120068613APending Publication Date: 2025-05-30QINGKE LINGJING (ANHUI) TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510126238.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing autonomous driving data sets are difficult to directly apply to port environments, resulting in the low degree of authenticity of the built autonomous driving simulation scenarios and the inability to effectively verify the performance of the autonomous driving algorithm.

Method used

By obtaining the driving status data of the selected confrontation vehicles, inputting them into the port driving model, training to obtain the human driver model, using the agent and expert experience based on the large language model to build an optimized reward function, and combining the reinforcement learning algorithm, a port driving scenario with a higher degree of reality is constructed.

Benefits of technology

It has realized the construction of a port driving scenario that is highly realistic against vehicles, which can more fully verify the performance of autonomous driving algorithms and improve the application prospects of autonomous driving technology in port environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068613A_ABST
    Figure CN120068613A_ABST
Patent Text Reader

Abstract

The invention discloses a port driving scene construction method and device, equipment and a storage medium. The method comprises the following steps: acquiring driving state data of a selected confrontation vehicle; inputting the obtained driving state data of the selected confrontation vehicle into a port driving model to obtain action data of the selected confrontation vehicle output by the port driving model; controlling the selected confrontation vehicle to move in the traffic simulation environment according to the obtained action data of the selected confrontation vehicle to obtain a driving scene of the confrontation vehicle; wherein the port driving model is obtained by training a human driver model through port vehicle behavior characteristic information and an optimized reward function, the optimized reward function is jointly constructed by an intelligent agent based on a large language model and expert experience, and the human driver model comprises a generator adopting a preset reinforcement learning algorithm model. According to the method, a port driving scene which is high in truth degree and is used for resisting vehicles can be constructed, so that an automatic driving algorithm can be fully verified subsequently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of autonomous driving simulation construction, and particularly to a method, device, equipment and storage medium for constructing a driving scenario in a port. Background Art

[0002] The port autonomous driving simulation test refers to testing the driving behavior of an autonomous driving algorithm in a virtual port autonomous driving test scenario, so as to determine whether the autonomous driving algorithm can achieve an intelligent port autonomous driving effect, such as automatically avoiding obstacles and intelligently coping with emergencies. In order to better verify the performance of the autonomous driving algorithm, the construction of the adversarial vehicle driving scenario is particularly important.

[0003] Compared with the autonomous driving scenarios of traditional cities or highways, since the port is a closed commercial environment, data collection is strictly restricted, and the road network structure, driving behavior and traffic rules are significantly different from those of conventional traffic scenarios, so that the existing autonomous driving data sets are difficult to be directly applied to the port, which brings more technical challenges to the construction of the adversarial vehicle port driving scenario, and further results in a lower authenticity of the constructed scenario. Summary of the Invention

[0004] Based on the above technical status quo, the present application proposes a method, device, equipment and storage medium for constructing a port driving scenario, which can construct a port autonomous driving simulation scenario with a relatively high authenticity for adversarial vehicles.

[0005] The first aspect of the present application provides a method for constructing a port driving scenario, including:

[0006] Obtaining the driving state data of a selected adversarial vehicle;

[0007] Inputting the obtained driving state data of the selected adversarial vehicle into a port driving model to obtain the action data of the selected adversarial vehicle output by the port driving model;

[0008] Controlling the movement of the selected adversarial vehicle in a traffic simulation environment according to the obtained action data of the selected adversarial vehicle to obtain a driving scenario of the adversarial vehicle;

[0009] Wherein, the port driving model is obtained by training a human driver model through port vehicle behavior feature information and an optimized reward function, the optimized reward function is jointly constructed by an intelligent agent based on a large language model and expert experience, and the human driver model includes: a generator adopting a preset reinforcement learning algorithm model.

[0010] The second aspect of the present application provides a port driving scenario generation device, including:

[0011] A data acquisition unit for acquiring the driving state data of a selected adversarial vehicle;

[0012] A data processing unit for inputting the acquired driving state data of the selected adversarial vehicle into a port driving model to obtain the action data of the selected adversarial vehicle output by the port driving model;

[0013] A scenario generation unit for controlling the movement of the selected adversarial vehicle in a traffic simulation environment according to the obtained action data of the selected adversarial vehicle to obtain a port driving scenario of the adversarial vehicle;

[0014] Wherein, the port driving model is obtained by training a human driver model through port vehicle behavior feature information and an optimized reward function, the optimized reward function is jointly constructed based on an agent of a large language model and expert experience, and the human driver model includes: a generator adopting a preset reinforcement learning algorithm model.

[0015] A third aspect of the present application proposes an electronic device, including:

[0016] A memory and a processor;

[0017] The memory is connected to the processor and is used for storing programs;

[0018] The processor is used for implementing the above-mentioned port driving scenario construction method by running the programs in the memory.

[0019] A fourth aspect of the present application proposes a storage medium, on which a computer program is stored, and when the computer program is run by a processor, the above-mentioned port driving scenario construction method is implemented.

[0020] In the port driving scenario construction method proposed by the present application, since in the port driving model for generating the driving scenario of the adversarial vehicle, an optimized reward function is constructed by using an agent of a large language model and expert experience, and port vehicle behavior feature information is used as training data, and a reinforcement learning algorithm is adopted in the training process, a port driving scenario of the adversarial vehicle with a relatively high degree of authenticity can be constructed, so as to fully verify the autonomous driving algorithm subsequently. Description of the Drawings

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0022] Figure 1Schematic flowchart of a method for constructing a port driving scenario according to an embodiment of the present application;

[0023] Figure 2 Schematic flowchart of a method for constructing a port driving model according to an embodiment of the present application;

[0024] Figure 3 Schematic diagram of the training process of a human driver model according to an embodiment of the present application;

[0025] Figure 4 Schematic diagram of the design process of a reward function according to an embodiment of the present application;

[0026] Figure 5 Schematic flowchart of another method for constructing a port driving scenario according to an embodiment of the present application;

[0027] Figure 6 Schematic diagram of the structure of a port driving scenario generation device according to an embodiment of the present application;

[0028] Figure 7 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0029] Compared with the automatic driving scenarios of traditional cities or highways, the special environment of ports brings more technical challenges to the development of automatic driving technology, for the following reasons:

[0030] First of all, since ports are closed commercial environments, data collection is strictly restricted, resulting in a serious lack of training data. In addition, there are significant differences between the road network structure, driving behaviors, and traffic rules in ports and conventional traffic scenarios, making it difficult to directly apply existing automatic driving data sets to ports, and it is also difficult for related automatic driving algorithms to reach the performance levels in traditional scenarios. These factors have led to the fact that port automatic driving technology currently mainly relies on line-tracking Automated Guided Vehicles (AGVs), so its flexibility and applicability are greatly reduced.

[0031] Secondly, different from warehouses or mines in closed environments, ports have limited space and huge modification costs, and it is not easy to demarcate dedicated automatic driving areas, resulting in the common phenomenon of mixed traffic of automatic driving vehicles and human-driven vehicles. In actual operations, in order to speed up the completion of transportation tasks, human drivers often adopt aggressive driving behaviors such as sudden lane changes, emergency braking, or not yielding at intersections, and even seize the driving rights of automatic driving vehicles by means of forced lane changes and overtaking, resulting in inefficient operation of automatic driving vehicles.

[0032] Due to the above factors, the actual performance of autonomous vehicles in ports is far lower than expected, and the future application prospects are unclear. This makes it difficult for port management to make investment decisions on broader upgrades of autonomous driving technologies, further delaying the improvement of port autonomous driving technologies and their commercialization process. In order to improve the performance of port autonomous driving algorithms, the first thing to consider should be how to generate relatively realistic driving scenarios for adversarial vehicles. Only by generating relatively realistic driving scenarios for adversarial vehicles can the autonomous driving algorithms of the vehicles under test be fully simulated and verified, significantly improving the performance of autonomous driving algorithms, thereby accelerating the application of autonomous driving technologies in ports.

[0033] The technical solution of the embodiment of the present application is applicable to generating driving scenarios for adversarial vehicles. By adopting the technical solution of the embodiment of the present application, a relatively realistic port driving scenario for adversarial vehicles can be generated, so as to subsequently fully verify the safety and reliability of autonomous driving algorithms in the port environment.

[0034] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0035] Exemplary method

[0036] An embodiment of the present application provides a method for constructing a port driving scenario, as Figure 1 shown. The method includes:

[0037] Step 100: Obtain the driving state data of the selected adversarial vehicle.

[0038] The driving state data may consist of observation data, and the observation data includes one or more of the following: lateral distance, longitudinal distance, lateral speed, longitudinal speed, and yaw angle of the vehicle.

[0039] Step 110: Input the obtained driving state data of the selected adversarial vehicle into the port driving model to obtain the action data of the selected adversarial vehicle output by the port driving model; wherein, the port driving model is trained by the port vehicle behavior characteristic information and the optimized reward function, and the optimized reward function is jointly constructed by an intelligent agent based on a large language model and expert experience. The human driver model includes a generator adopting a preset reinforcement learning algorithm model.

[0040] The action data may include information such as speed, acceleration, position, direction, and steering angle.

[0041] A large language model (LLM) refers to a deep learning model trained using a large amount of text data to enable it to generate natural language text or understand the meaning of language text. The problem discussed in reinforcement learning (RL) is how an agent can maximize the rewards it can obtain in a complex and uncertain environment. By perceiving the rewards of actions based on the state of the environment it is in, to guide better actions in order to obtain the maximum benefit, this is called learning through interaction, and such a learning method is called reinforcement learning.

[0042] Due to the high intensity and complexity of port operations and the linkage characteristics with maritime transportation, the behavior of port vehicles has unique characteristics. The following are several important aspects of the behavior characteristics of port vehicles: 1. Port vehicles need to perform high-density and high-frequency operations. Ports are usually a high-density and high-frequency operating environment. Especially in container ports, vehicles need to frequently perform tasks such as handling, loading and unloading, and transportation, and the intervals between tasks are short. Therefore, vehicles may start, stop, accelerate and decelerate frequently in a short period of time, and sometimes may need to respond quickly to sudden traffic or work situations. 2. Port vehicles require multi-task parallel and collaborative operations. Port vehicles usually need to switch efficiently between multiple operating tasks. Different types of vehicles (such as stackers, container trailers, forklifts, etc.) will cooperate with other vehicles during operation, so there may be a need for parallel work and collaborative scheduling when vehicles are operating. For example, a container trailer may cooperate with a stacker, and multiple forklifts may operate in the same area at the same time. 3. The traffic environment and route planning of port vehicles are complex. The road environment in the port is relatively complex, especially in the container yard area, where there may be narrow passages, frequent turns and temporary obstacles. Therefore, the driving of port vehicles usually needs to adapt to the complex traffic environment, and may frequently change lanes, slow down, avoid obstacles, etc. At the same time, flexible route planning is required to ensure operational efficiency. 4. The weight and load of port vehicles vary greatly. Port vehicles often need to bear heavier loads when carrying containers or other goods. Depending on the changes in load, the driving characteristics of the vehicle (such as acceleration, braking distance, turning radius, etc.) will also be different. Therefore, port vehicles may require longer stopping distances, slower starting speeds and smoother handling when heavily loaded, and may accelerate and turn faster when unloaded or lightly loaded. 5. The driving environment of port vehicles is complex. Ports are usually located on the coast and are greatly affected by weather and ground conditions. Bad weather (such as strong winds, rain, snow, and dense fog) may cause poor visibility and affect the driving stability and safety of the vehicle. Therefore, port vehicles will show slower driving speeds, increase braking distances, and adopt more conservative driving strategies when facing bad weather to cope with low visibility or slippery road conditions. In summary, port vehicle behavior feature information can come from several typical violations or aggressive driving behaviors such as deviation of vision from the driving route, operation of other on-board equipment, use of mobile phones, sudden braking / sudden acceleration, unwarranted slowing / parking / reversing, and frequent overtaking. From the above content, it can be seen that port vehicles have unique driving behaviors, so the acquisition of port vehicle behavior feature information is crucial to the training process of the human driver model. Therefore, in the port driving scene construction method provided in the embodiment of the present application, the port vehicle behavior feature information is obtained, so that the large language model learns the port vehicle behavior features and provides rich real data.

[0043] In the related art, the design of the reward function is mainly divided into two methods: one is the expert-based design, which relies on the prior experience accumulation of experts. Due to the complexity and variability of different port scenarios, this method often requires a large amount of manpower and time to continuously adjust the reward function. The other is based on inverse reinforcement learning, which relies on a large amount of data to solve the reward function that conforms to the rules of the dataset. Although this method reduces the dependence on experience, its strong dependence on data limits its direct application in port environments with scarce data.

[0044] Different from the above two methods, general large language models already have extensive basic experience knowledge, which can reduce the dependence on expert experience and datasets. Therefore, the method of designing the reinforcement learning reward function based on large language models has received more attention in recent years. However, this method also has some deficiencies. For example, although large language models have extensive knowledge, they are relatively weak in professional fields. Therefore, the design effect of the reward function is not good. In the port driving scenario construction method provided by the embodiments of the present application, the reward function is jointly constructed by an intelligent agent based on a large language model and expert experience. At the same time, the advantages of the large language model and expert experience are utilized, avoiding the problem of inaccurate or ineffective design caused by the lack of professional field knowledge of the large language model when designing the reward function alone, and also avoiding the problem of inaccurate design caused by low data utilization rate, expert personal bias or limited perspective when relying entirely on expert experience for design. Therefore, the design effect of the reward function is improved.

[0045] In the process of jointly constructing the reward function by an intelligent agent based on a large language model and expert experience, the important indicators in the process of constructing the reward function can be first clarified by experts in the port field, and then the reward function can be designed by the intelligent agent based on the large language model; or the intelligent agent based on the large language model can first conduct a preliminary design of the reward function, and then the expert can directly adjust the preliminarily designed reward function to obtain the final reward function; or the intelligent agent based on the large language model can first design the reward function, and then embed the preliminarily designed reward function in the human driver model. The expert can review the output results of the model according to his professional knowledge and practical experience, point out the deficiencies or errors in the answer, and provide specific feedback to guide the model to make adjustments. After 1-2 rounds of feedback, the expert completes the final link of code synthesis to obtain the final reward function.

[0046] The intelligent agent in the embodiments of the present application can be Port-Bot. Port-Bot is an intelligent agent specifically designed for port management and automation, mainly used to improve the efficiency and accuracy of port operations. As an important hub of global trade and logistics, ports involve very complex tasks, including cargo handling, transportation coordination, inventory management, etc., thus solving a series of challenges in port operations.

[0047] The reinforcement learning algorithm is a machine learning method that learns how to take actions to maximize cumulative rewards through the interaction between an agent and the environment. The core idea of reinforcement learning is to let the agent gradually optimize its behavior strategy during the trial-and-error process to achieve specific goals. In the port driving scenario construction method provided by the embodiments of this application, the reinforcement learning algorithm is introduced, and through a large amount of port vehicle behavior characteristic information, the reinforcement learning algorithm can capture the actual behavior patterns of adversarial drivers in different situations, so that the human driver model can more accurately grasp the adversarial driving habits in the real port environment during the training process.

[0048] The human driver model can be a generative adversarial network model. The generative adversarial network model can include a generator and a discriminator. The generator is responsible for generating driving behavior strategies according to the current environment and situation (such as traffic conditions, road conditions, behaviors of other vehicles, etc.). These behaviors usually manifest as decisions (such as accelerating, braking, changing lanes, turning, etc.). The generator usually learns through training how to simulate a human driver to execute optimal or reasonable driving decisions in a given environment. The task of the discriminator is to evaluate whether the driving behavior strategies generated by the generator are reasonable, realistic, or conform to the expected driving patterns. It "guides" the generator to improve by determining whether the generated behaviors can meet the expected standards of safety, efficiency, comfort, etc., so as to get closer and closer to the real driver behavior.

[0049] The inputs of the generator can include: the current state of the adversarial vehicle (position, speed, acceleration, etc.), surrounding traffic information (relative positions, speeds of other vehicles, etc.), and environmental information (road conditions, weather, traffic signals, etc.). The outputs of the generator can include: the generated actions of the adversarial vehicle (such as controlling acceleration, steering angle, etc.). These behaviors are to simulate the decisions of the adversarial driver in this environment. The inputs of the discriminator can include: the behavior strategies output by the generator (i.e., the generated driving behaviors) and some real adversarial driving behavior data (such as the behaviors of human adversarial drivers under the same conditions). The outputs of the discriminator can include: "true / false" judgments to evaluate whether the generated behaviors can be considered as reasonable simulations of adversarial driver behaviors.

[0050] Step 102: Control the selected adversarial vehicle to move in the traffic simulation environment according to the obtained action data of the selected adversarial vehicle to obtain the driving scenario of the adversarial vehicle.

[0051] Controlling the selected countermeasure vehicle to move in the traffic simulation environment based on the obtained action data of the selected countermeasure vehicle requires creating or importing a traffic network model that reflects an actual or hypothetical scenario through traffic simulation software, such as SUMO, AIMSUN, VISSIM, CarMaker, or MATLAB / Simulink. This should include road layouts, lanes, intersections, traffic lights, and other infrastructure elements. Then, the action data is converted into behavioral instructions that the simulation platform can understand and apply and input into the simulation platform to obtain the driving scenario of the countermeasure vehicle. Examples of the conversion of the behavioral data of the countermeasure vehicle into corresponding instructions are as follows: If the action data is "speed 30 km / h", then the action data can be converted into a behavioral instruction like "drive at a speed of 30 km / h". If the action data is "acceleration 5 m / s 2 ²", then the action data can be converted into a behavioral instruction like "drive with an acceleration of 5 m / s²". If the action data is "deflection angle of 30 degrees to the east", then the action data can be converted into a behavioral instruction like "drive with a deflection angle of 30 degrees to the east in the current driving direction".

[0052] And since the movement of the selected countermeasure vehicle in the traffic simulation environment is a continuous process with constantly changing motion states, the action data of the selected countermeasure vehicle is obtained from time to time. Whenever the action data of the countermeasure vehicle at a certain moment is obtained, the selected countermeasure vehicle is controlled to move in the traffic simulation environment based on the obtained action data of the selected countermeasure vehicle. Thus, a rich and variable motion state of the countermeasure vehicle close to that of a real vehicle is constructed. Specifically, for any two adjacent moments (the two moments are respectively referred to as the previous moment and the next moment), after obtaining the action data of the countermeasure vehicle at the previous moment, the selected countermeasure vehicle is controlled to move in the traffic simulation environment based on the action data of the countermeasure vehicle at the previous moment and maintain this motion state until the action data of the countermeasure vehicle at the next moment is obtained. Then, the selected countermeasure vehicle is controlled to move in the traffic simulation environment based on the action data of the countermeasure vehicle at the next moment. At this time, the motion state of the countermeasure vehicle has changed compared to the motion state at the previous moment.

[0053] In the port driving scenario construction method provided by the embodiments of this application, in the port driving model for generating the driving scenario of the countermeasure vehicle, an intelligent agent based on a large language model and expert experience is used to construct a reward function, and the port vehicle behavior characteristic information is used as training data. And a reinforcement learning algorithm is adopted during the training process. Therefore, a port driving scenario of the countermeasure vehicle with a relatively high degree of authenticity can be constructed to fully verify the autonomous driving algorithm subsequently.

[0054] In an exemplary embodiment, as Figure 2 shown, the method further includes:

[0055] Step 200: Obtain the port vehicle behavior characteristic information.

[0056] Step 210: Based on the obtained port vehicle behavior characteristic information, use the agent to design an initial reward function, and obtain an optimized reward function that integrates expert experience based on the designed initial reward function.

[0057] The optimized reward function can be obtained through the following process: The agent based on the large language model conducts a preliminary design of the reward function, and then the preliminarily designed reward function is built into the human driver model. Experts can review the output results of the model according to their professional knowledge and practical experience, point out the deficiencies or errors in the answers, and provide specific feedback to guide the model to make adjustments. After 1 - 2 rounds of feedback, the experts complete the final link of code synthesis to obtain the final reward function, that is, the optimized reward function.

[0058] Step 220: According to the simulation data from the natural scene traffic flow data set and the obtained optimized reward function, train the pre - constructed human driver model, and use the generator in the trained human driver model as the port driving model.

[0059] Input the simulation data into the generator of the human driver model to obtain the action data output by the generator of the human driving behavior strategy model. Calculate the reward value corresponding to the optimized reward function according to the simulation data and the action data, and adjust the parameters of the generator of the human driving behavior strategy model according to the reward value; then input the simulation data into the generator of the human driver model again to obtain the action data output by the generator of the human driving behavior strategy model with adjusted parameters. Calculate the reward value corresponding to the optimized reward function according to the simulation data and the action data, and adjust the parameters of the generator of the human driving behavior strategy model again according to the reward value. Repeat the above process until the obtained reward value meets the preset conditions to obtain the trained human driver model.

[0060] The simulation data of the natural scene traffic flow dataset is the core foundation of the reinforcement learning human driver model, enabling the model to make realistic action feedback in the constructed simulation environment based on various attributes of the data and dynamically simulate the interaction process of vehicles in the port. Due to the natural data shortage in the port environment and the diverse road network forms in the actual map, such as straight roads and intersections, complex traffic datasets in the real world, such as the NGSIM dataset and the INTERACTION dataset, can be used to achieve effective alignment and utilization of the existing datasets in the port road network environment; among them, the NGSIM dataset provides a large amount of high-quality and fine-grained traffic flow data. These data usually include vehicle trajectory data from different cities and road networks, information such as vehicle position, speed, and acceleration; the INTERACTION dataset is a high-quality dataset commonly used in the fields of traffic flow, autonomous driving, traffic behavior analysis, etc. It usually contains the interaction behaviors of vehicles and pedestrians recorded on actual roads or in specific scenarios. The core value of this dataset is to help researchers analyze issues such as individual behavior, traffic safety, and verification of autonomous driving algorithms through detailed traffic flow and behavior data. Using these complex traffic data, the simulation environment can be adjusted according to the specific needs of the port to make it closer to the complexity of actual port operations. In addition, by integrating the knowledge and feedback of experts, the training process of reinforcement learning will be optimized to improve the algorithm's ability to handle dynamic and complex interactions in the simulated port environment.

[0061] The training process of the human driver model can be as Figure 3 shown, covering the whole process from data collection, processing to model training and testing, specifically including:

[0062] The generator is responsible for generating updates to the policy network and the value network. The policy network decides the actions to be taken based on the current state, and the actions are sent to the simulation environment. The value network evaluates the value of the current state, that is, the possible reward obtained after taking a certain action. The output of the value network is used to update the policy network to optimize the policy.

[0063] The simulation environment uses complex traffic actual data to simulate the real traffic environment, improving the training effect of the policy network and the value network. It uses expert knowledge to optimize the simulation process to ensure the accuracy and effectiveness of the simulation environment. The simulation environment also receives the actions output by the policy network and updates the state according to these actions. After the state of the simulation environment is updated, it will be fed back to the policy network and the value network, and the policy network and the value network update their parameters according to the fed-back state and reward to optimize the policy.

[0064] In summary, this process is an iterative optimization process. Through the state update and reward feedback in the simulation environment, the policy network and the value network continuously learn and optimize to achieve better decision-making effects.

[0065] The simulation environment is a key component for the reinforcement learning network to obtain behavioral feedback. Usually, depending on different simulation objectives, its focus will also vary. For example, a simulation environment that emphasizes the perception function will focus more on comprehensive rendering effects, while one oriented towards the planning and control function will pay more attention to dynamic attributes. By customizing the input and output of the simulation environment, the model performance can be evaluated more accurately. The input and output contents concerned in the method for generating a port driving scenario provided by the embodiments of the present application may include:

[0066] Input: The input of the simulation environment usually only includes the behaviors generated by the reinforcement learning policy, and the execution effects of these behaviors are determined by the built-in reward function. To accurately simulate the characteristics of the port environment, complex traffic data and expert knowledge are integrated as input to participate in the configuration of the reward function and simulation parameters.

[0067] Output: The output of the simulation environment mainly includes two categories: state and reward. The state output depends on different simulation objectives, and different simulation platforms with different focuses can be selected to provide corresponding state information; the reward output is achieved through the designed reward function, which can effectively reflect the port characteristics and ensure an effective mapping from specific behaviors to reward values. At the same time, this process is not restricted by specific simulation platforms, such as simulation tools like SUMO (Simulation of Urban MObility) and CARLA. Among them, SUMO (Simulation of Urban MObility) is an open-source, portable, and modular traffic flow simulation tool that can simulate the movement of vehicles in a road network, including different types of transportation tools such as cars, trucks, buses, and bicycles. SUMO can be used for researching and optimizing urban traffic flow, planning public transportation systems, testing intelligent transportation systems, developing autonomous driving vehicle algorithms, etc. CARLA (CAR Learning to Act) is an open-source urban driving simulation platform specifically designed for autonomous driving research and development. It provides a high-fidelity 3D urban environment, supports the simulation of multiple sensors, and has a powerful API interface, enabling researchers and developers to easily test, verify, and train autonomous driving algorithms.

[0068] In an exemplary embodiment, the characteristic information of the port vehicle is obtained through a questionnaire survey. The recipients of the questionnaire survey are port fleet management personnel who meet the preset conditions. The questionnaire survey is designed using the cross-sectional survey method and stipulates the use of the Likert scale response format.

[0069] Given the limitations of large language models in terms of professional knowledge and the scarcity of port data, in order to build a realistic human driver model, it is crucial to understand the behavioral and psychological characteristics of drivers. Therefore, the characteristic information of port vehicles can be obtained through a questionnaire survey. The recipients of the questionnaire survey are port fleet managers who meet the preset conditions. The preset conditions can be the conditions set for port fleet managers in terms of management years and large fleet management experience. The management years can be more than 10 years of port fleet management experience, and the large fleet management experience can be managing a fleet of more than 200 container trucks.

[0070] Cross-sectional survey methods are widely used in medical and business research and can quickly capture the characteristics of the respondents through fewer questions at a single point in time. The questionnaire used to obtain the characteristic information of port vehicles can cover the following main contents:

[0071] Covered topics: including questions about risk behaviors (such as speeding, using mobile phones), interactions with autonomous vehicles (such as aggressive behaviors like overtaking and cutting in), and the overall impact on traffic efficiency, etc. The survey particularly focuses on the aggressive driving behaviors that need to be set in the reinforcement learning reward function.

[0072] Answer format: The questions mainly adopt the Likert scale, such as "never", "seldom", "sometimes", "often", enabling the respondents to indicate the frequency or degree of agreement with statements related to driving behaviors. This format is suitable for accurately measuring questions related to perception and attitude.

[0073] In an exemplary embodiment, the design of the initial reward function using the agent based on the obtained port vehicle behavior characteristic information may include the following steps:

[0074] Expand the obtained port vehicle behavior characteristic information using zero-shot prompting technology;

[0075] Obtain the prompt words designed using the preset technology; wherein, the preset technology includes at least one of the following: chain of thought prompting, example diffusion technology, detail focusing technology;

[0076] Control the agent to design the initial reward function according to the expanded port vehicle behavior characteristic information and the designed prompt words.

[0077] After obtaining the characteristic information of port vehicles through a questionnaire survey, the large language model has initially understood the driving behavior preferences of drivers in the port. However, since the cross-sectional survey method usually requires questions to be concise, the amount of data obtained is relatively small. To enrich and deeply understand these survey results, the traditional method is to invite domain experts to supplement and improve the information, but its flexibility is low and it is not conducive to wide application. Therefore, the port driving scenario construction method provided by the embodiments of this application can also support the large language model to complete information expansion and enrichment to generate more detailed content. The specific implementation method can be as follows:

[0078] 1. Use zero-shot prompting technology (also known as zero-shot prompting) for information expansion: By specifically designing zero-shot prompt words, complete and detailed information can be provided to the large language model at one time, thus effectively generating rich and accurate content. This method simplifies the need for multi-round logical interaction with the large language model, can significantly reduce the complexity of content generation, and at the same time can effectively avoid the error accumulation caused by multi-round interaction and reduce the risk of the model hallucinating.

[0079] 2. Obtain the prompt words designed using preset technologies: Some techniques of cognitive science are borrowed when designing prompt words. For example, when using the Chain-of-thought (CoT) technology, the logical reasoning ability of the model is strengthened through simple instructions such as "Let’s think step by step". In addition, the concept of "example diffusion" is also adopted to guide the large language model to use specific examples (such as adding "vivid examples or anecdotes" in the prompt words), prompting the model to perform creative expansion and comprehensive analysis based on the existing information, and improving the practicality of the generated content. Similarly, the concept of "detail focus" is also applied to guide the model to pay more attention to the details of the description (such as adding "Pay close attention to detail in your descriptions" in the instructions), which helps the model capture more details about the behavior patterns and interactions of human drivers, and enhances the comprehensiveness and accuracy of the reward function design.

[0080] Based on the above - enriched information, the large - language model already has the prerequisite conditions for generating a reward function. It is possible to guide the model to generate a specific reward function or corresponding source code by directly inputting prompt words into the agent. However, such an operation will lead to the following problems: 1) Incorrect output due to the mismatch between the reward function and the input information: For example, the design goal is to generate a model that can simulate the behavior of human drivers to complete adversarial driving against autonomous vehicles, which is essentially different from the careful driving behavior planning concerned in the training of conventional autonomous driving models. Direct input is likely to cause information confusion in the large - language model, wrongly encouraging careful driving behavior and punishing adversarial driving, thus violating the design goal. 2) The source code directly generated by the reward function does not match the code environment: The update speed of the code environment is usually faster than the training update speed of the large - language model. Direct input will result in lagging synthesized code. In addition, current code synthesis mainly targets common code libraries and has weak adaptability to specific fields such as autonomous driving adversarial training. Especially when calling functions that are uncommon but crucial in the calling platform, it is extremely error - prone.

[0081] To solve these problems, the following two key steps can be taken:

[0082] 1. Use zero - shot prompting for self - reflection and CoT prompting: Before the large - language model designs the output answer, add a description of the steps to solve the problem and the CoT logical process to enhance the richness of the output content. At the same time, explicitly add an internal verification process to actively monitor and evaluate the quality of the answer. This method allows the model to show its thinking and verification process before providing the answer. This process not only enhances the transparency of the answer content but also allows the model to self - correct in logical inferences, reducing the occurrence of errors.

[0083] 2. Introduce expert access in a timely manner to close the loop of code synthesis: After the large - language model designs a preliminary reward function scheme, experts can review and optimize the model output based on their professional knowledge and practical experience. Experts can point out the deficiencies or errors in the answer and provide specific feedback to guide the model to make more precise adjustments. After 1 - 2 rounds of feedback, the experts complete the final link of code synthesis to ensure the quality of the code and the practicality of the workflow.

[0084] Regarding when experts should intervene and how to complete the final code implementation of the reward function, it has been found through research that large - language models are most effective when presenting the reward function in the form of mathematical expressions and language explanations. This is because direct code implementation may lead to frequent coding errors due to differences in the platforms relied on. Experts can use the output of the large - language model as a guide, significantly reducing the difficulty of reward function design and improving accuracy while ensuring the practicality of the code - writing process.

[0085] In the case where the agent is Port - Bot, the design process of the reward function can be asFigure 4 As shown in, specifically including:

[0086] Port-Bot takes the questionnaire results and the expert's experiential knowledge as input, and uses zero-shot prompting and chain-of-thought techniques (such as detail focusing and example diffusion) to optimize the large language model to enrich information, improve the understanding and reasoning ability of port domain knowledge, and then uses zero-shot prompting to make the large language model perform self-reflection and CoT prompting, and conducts cyclic verification through expert feedback, continuously adjusting and optimizing the design of the reward function. This method not only improves the adaptability of the reward function design, but also enhances its accuracy, thus effectively supporting the construction of the adversarial vehicle driving scenario for complex port conditions.

[0087] In an exemplary embodiment, the initial reward function includes: the initial reward function in the straight road scenario and the initial reward function in the intersection scenario;

[0088] Correspondingly, the optimized reward function includes: the optimized reward function in the straight road scenario and the optimized reward function in the intersection scenario.

[0089] In order to accurately train the human driver model, the training method in the straight road and intersection respectively can more realistically reflect the interaction characteristics of vehicles in different traffic scenarios in the port environment, thus effectively simulating various challenges in the real driving environment.

[0090] The port driving scenario construction method provided by the embodiment of the present application obtains the port vehicle behavior characteristics by distributing questionnaires to senior port fleet managers, and combines zero-shot prompting and chain-of-thought techniques, uses the logical reasoning and semantic extension ability of the large language model, designs the Port-Bot workflow to sort out and summarize the questionnaire content, thereby enhancing the knowledge understanding ability of the general large language model in the port domain. Then, on the basis of the content summarized by Port-Bot, further explore the code and function synthesis ability of the general large language model, guide Port-Bot to complete the design of the reinforcement learning reward function, and provide a training basis for simulating the human driver model. Finally, train the human driver model based on the designed reward function and reinforcement learning algorithm, and complete the construction of the port driving environment in the simulator, so as to achieve the purpose of simulating the aggressive driving behavior and verifying the algorithm performance of the port autonomous driving algorithm in typical scenarios such as straight roads and intersections.

[0091] In an exemplary embodiment, the initial reward function in the straight road scenario and the optimized reward function in the straight road scenario both include: a mixed reward function obtained by combining an aggressive reward function, an interfering with traffic reward function, a risk participation reward function, and a speed incentive reward function;

[0092] The initial reward function and the optimized reward function in the intersection scenario both include: a mixed reward function obtained by combining a right-of-way preemption reward function, an aggressiveness reward function, a traffic interference reward function, and a risk participation reward function;

[0093] Among them, the aggressiveness reward function is used to reward the attack behavior of the adversarial vehicle, and the attack behavior includes: inserting in front of the vehicle under test or occupying the lane where the vehicle under test is located;

[0094] The traffic interference reward function is used to reward the interference behavior of the adversarial vehicle on the normal traffic flow;

[0095] The risk participation reward function is used to reward the risk behavior of the adversarial vehicle, and the risk behavior includes: overtaking or braking when the distance from the vehicle under test is less than a preset distance;

[0096] The speed reward function is used to reward the adversarial vehicle for competing with the vehicle under test by controlling the speed;

[0097] The right-of-way preemption reward function is used to reward the adversarial vehicle for reaching the designated point of the preset intersection prior to the vehicle under test;

[0098] Among them, the interference behavior on the normal traffic flow includes: the attack behavior of the adversarial vehicle, the risk behavior of the adversarial vehicle, and the adversarial vehicle reaching the designated point of the preset intersection prior to the vehicle under test.

[0099] The aggressiveness reward function is used to test the ability of the autonomous driving vehicle to handle sudden changes in the traffic environment. The aggressiveness reward function in the optimized reward function in the straight road scenario can be represented by R Aggressive The aggressiveness reward function in the optimized reward function in the intersection scenario can be represented by R'; Aggressive The traffic interference reward function is used to test the ability of the autonomous driving vehicle to adapt to the unpredictable and uncooperative behaviors of other road users. The traffic interference reward function in the optimized reward function in the straight road scenario can be represented by R TrafficDisruption The traffic interference reward function in the optimized reward function in the intersection scenario can be represented by R'; TrafficDisruption The risk participation reward function is used to test the reaction ability of the autonomous driving system to sudden operations. The risk participation reward function in the optimized reward function in the straight road scenario can be represented by R RiskEngagement The risk participation reward function in the optimized reward function in the intersection scenario can be represented by R'; RiskEngagement The speed incentive reward function is used to simulate the competitive driving behavior of some human drivers. The speed incentive reward function in the optimized reward function in the straight road scenario can be represented by R SpeedIt is indicated that the right-of-way preemption reward function is used to reward the adversarial vehicle for preemption of key points at intersections through acceleration and deceleration behaviors, which pose a safety threat to the vehicle under test during turning or passing through intersections. Optimizing the right-of-way preemption reward function in the intersection scenario can be represented by R Negotiation It is indicated.

[0100] In one exemplary embodiment, the steps of obtaining the hybrid reward function by combining the aggressiveness reward function, the traffic interference reward function, the risk participation reward function, and the speed incentive reward function include:

[0101] Obtaining the hybrid reward function according to the aggressiveness reward function, the traffic interference reward function, the risk participation reward function, the speed incentive reward function, and their respective preset weight coefficients;

[0102] The steps of obtaining the hybrid reward function by combining the right-of-way preemption reward function, the aggressiveness reward function, the traffic interference reward function, and the risk participation reward function include:

[0103] Obtaining the hybrid reward function according to the right-of-way preemption reward function, the aggressiveness reward function, the traffic interference reward function, the risk participation reward function, and their respective preset weight coefficients.

[0104] The hybrid reward function obtained by combining the aggressiveness reward function, the traffic interference reward function, the risk participation reward function, and the speed incentive reward function is the initial reward function and the optimized reward function in the straight road scenario, while the hybrid reward function obtained by combining the right-of-way preemption reward function, the aggressiveness reward function, the traffic interference reward function, and the risk participation reward function is the initial reward function and the optimized reward function in the intersection scenario.

[0105] The hybrid reward function R obtained by combining the aggressiveness reward function, the traffic interference reward function, the risk participation reward function, and the speed incentive reward function StrightLane can be expressed as:

[0106] R StrightLane = ω 1 R Aggressive + ω 2 R TrafficDisruption + ω 3 R RiskEngagement + ω 4 R Speed

[0107] where ω 1 , ω 2 , ω 3 and ω 4 respectively represent the weights corresponding to the aggressiveness reward function, the traffic interference reward function, the risk participation reward function, and the speed incentive reward function in the straight road scenario, ω1 , ω 2 , ω 3 and ω 4 are used to balance different types of adversarial behaviors, reflect the relative importance or expected frequency of each behavior in the straight - road simulation scenario, and aim to better simulate the complex driving environment in the real world.

[0108] The mixed - reward function R obtained by combining the right - of - way preemption reward function, aggression reward function, traffic - interference reward function, and risk - participation reward function Intersection can be expressed as:

[0109] R Intersection = ω 1 'R Negotiation + ω 2 'R' Aggressive + ω 3 'R' TrafficDisruption + ω 4 'R' RiskEngagement

[0110] ω 1 ', ω 2 ', ω 3 ' and ω 4 ' respectively represent the weights corresponding to the right - of - way preemption reward function, aggression reward function, traffic - interference reward function, and risk - participation reward function in the intersection scenario. ω 1 ', ω 2 ', ω 3 ' and ω 4 ' are used to balance different types of adversarial behaviors, reflect the relative importance or expected frequency of each behavior in the intersection simulation scenario, and aim to better simulate the complex driving environment in the real world.

[0111] In an exemplary embodiment, the design of the aggression reward function in the optimized reward function under the straight - road scenario is as follows:

[0112] Construct a reward function for rewarding an adversarial vehicle for suddenly inserting into the lane where the vehicle under test is located, based on the function term for ensuring that the adversarial vehicle and the vehicle under test are in the same lane, and the function term for encouraging at least a preset safety distance between the adversarial vehicle and the vehicle under test.

[0113] According to the description of the aggression reward, the behavior of "suddenly inserting" can be selected as the specific implementation goal. The aggression reward function R in the optimized reward function under the straight - road scenario Aggressive can actually be designed as:

[0114]

[0115] Wherein, xadv , xegorespectively represent the abscissas of the adversarial vehicle and the vehicle under test, yadv 、 ynpc respectively represent the ordinates of the adversarial vehicle and the vehicle under test, dsafe 、 wlane respectively represent the safety distance and the lane width. The key design points are as follows: 1) An adversarial vehicle (adv) and a cooperative adversarial vehicle (npc) are introduced into the scenario to simulate the natural "sudden insertion" behavior. When the cooperative adversarial vehicle brakes suddenly or stops temporarily, the adversarial vehicle suddenly inserts into the lane where the vehicle under test (ego) is located, by ensuring that adv and npc are as much as possible in the same lane to stimulate the natural "sudden insertion" behavior. 2) To further ensure the naturalness of the "sudden insertion" behavior, by terms to encourage the adversarial vehicle to maintain at least a minimum safety distance from the vehicle under test to prevent inevitable collisions and maintain the effectiveness of the test. 3) Use an exponential function to represent the reward function, limit the variable range to the interval (0, 1), focus on rewarding eligible behaviors, and give less or no reward to ineligible behaviors.

[0116] In an exemplary embodiment, the risk participation reward function in the optimized reward function under the straight road scenario is designed as follows:

[0117] Give a preset reward value when the adversarial vehicle and the vehicle under test are not in the same lane, and give a preset reward value and the reward value obtained from the collision reward function term when the adversarial vehicle and the vehicle under test are in the same lane, to construct a reward function that rewards the adversarial vehicle for sudden braking behavior; wherein, the collision reward function term is used to encourage the adversarial vehicle to collide with the vehicle under test.

[0118] According to the description of the risk participation reward, the behavior of "sudden braking" can be selected as the specific implementation goal, and the risk participation reward function R in the optimized reward function under the straight road scenario RiskEngagement can actually be designed as:

[0119]

[0120] Wherein, TTC represents the time required for the vehicle under test to collide with the adversarial vehicle while maintaining the current driving direction and speed, TTC max represents the maximum time required for a collision, which is set according to the road characteristics of the scenario, such as speed limit rules, etc., LaneID adv 、LaneID ego respectively represent the lane numbers of the adversarial vehicle and the vehicle under test, R SameLane represents the preset basic reward when the adversarial vehicle and the vehicle under test are in the same lane, which can be set to 0.1.

[0121] The key points of this design are: 1) Using R TTC Rewards are given for smaller TTCs to enhance the risky driving characteristics of the adversarial vehicle, that is, the reaction ability of the vehicle under test is tested by shortening its response time; 2) This function limits the test scenario to the same lane as much as possible, highlighting the natural occurrence of "sudden braking" behavior, while avoiding the adversarial vehicle from directly colliding with the vehicle under test across lanes, avoiding unavoidable collision scenarios.

[0122] In an exemplary embodiment, the interference traffic reward function in the optimization reward function in the straight road scenario is designed as follows:

[0123] By reusing the risk participation reward design in the optimization reward function in the straight road scenario and shortening the collision time, a traffic interference reward function R is constructed that rewards the vehicle for sudden braking at a higher frequency. TrafficDisruption .

[0124] According to the description of the traffic interference reward, the traffic interference behavior is realized by reusing the risk participation reward method. The key points of this design are: 1) By shortening the TTC time and increasing the frequency of the sudden behavior of the adversarial vehicle, the reaction ability and ability to adapt to traffic interference of the tested vehicle can be tested more frequently; 2) By reusing the risk participation reward mechanism, it can not only reduce the design complexity, but also effectively avoid reward deception and accelerate the convergence of the model to the target behavior.

[0125] In an exemplary embodiment, the speed incentive reward function in the optimization reward function in the straight road scenario is designed as follows:

[0126] Based on the function items for ensuring that the adversarial vehicle and the vehicle under test are in the same lane and there is a speed difference, and the function items for encouraging that there is at least a preset safety distance between the adversarial vehicle and the vehicle under test, a reward function is constructed to reward the adversarial vehicle for competing with the vehicle under test in the lane where the vehicle under test is located.

[0127] According to the description of the speed incentive reward function, using the aggressive reward R RiskEngagement Design idea, speed incentive reward function R Speed It can actually be designed as:

[0128]

[0129] in, xadv , xego Represent the horizontal coordinates of the adversarial vehicle and the tested vehicle, yadv , ynpc They represent the ordinates of the adversarial vehicle and the cooperative adversarial vehicle, respectively, and w lane Indicates lane width, Speed diffIndicates the speed difference between the adversarial vehicle and the vehicle under test. The key design point is to ensure that the two vehicles maintain a certain speed difference on the same lane, thereby continuously exerting pressure on the vehicle under test and testing its performance and safety in a variable-speed driving environment.

[0130] In an exemplary embodiment, the design of the right-of-way preemption reward function in the optimized reward function for the intersection scenario is as follows:

[0131] Construct a reward function for the adversarial vehicle to preempt the designated point of the preset intersection in terms of time according to the function item used to encourage the adversarial vehicle to spend less time than the vehicle under test to reach the designated point of the preset intersection.

[0132] According to the description of the right-of-way preemption reward, the behavior of "jumping the red light at the intersection" can be selected. The right-of-way preemption reward function R in the optimized reward function for the intersection scenario Negotiation Can actually be designed as:

[0133]

[0134] where ego t and adv t respectively represent the time for the vehicle under test and the adversarial vehicle to reach the designated point of the preset intersection. The key design point is to encourage the adversarial vehicle to preempt the key point first, so as to verify whether the vehicle under test can flexibly adjust the route and avoid risks.

[0135] In an exemplary embodiment, the implementation of the aggressiveness reward function in the optimized reward function for the intersection scenario is as follows:

[0136] Construct a reward function for the adversarial vehicle to preempt the designated point of the preset intersection in terms of distance according to the function item used to encourage the adversarial vehicle to reach the designated point of the preset intersection faster than the vehicle under test when there is a speed difference between the adversarial vehicle and the vehicle under test.

[0137] According to the description of the aggressiveness reward, the specific implementation goal of key point preemption can still be utilized. The aggressiveness reward function R' in the optimized reward function for the intersection scenario Aggressive Can actually be designed as:

[0138]

[0139] where d adv represents the distance between the adversarial vehicle and the key point of the preset intersection, and Speed ego 、Speed adv respectively represent the speeds of the vehicle under test and the adversarial vehicle.

[0140] The key design points are as follows: 1) Similar to the time constraint of right-of-way preemption, this time, from the spatial perspective, the adversarial vehicle is encouraged to preempt key points; 2) The adversarial vehicle is encouraged to maintain a large speed difference from the vehicle under test, and the adversarial vehicle travels relatively slowly, so as to create a fact of occupying key points and test whether the vehicle under test can handle the situation of occupying the intersection.

[0141] In an exemplary embodiment, the design of the interference traffic reward function in the optimized reward function under the intersection scenario is as follows:

[0142] By reusing the design of the aggressive reward function in the optimized reward function under the intersection scenario, a reward function for rewarding the vehicle under test for occupying the lane where the vehicle under test is located is constructed.

[0143] The interference traffic reward function R' in the optimized reward function under the intersection scenario TrafficDisruption By reusing the aggressive reward function R' Aggressive is designed and implemented. According to the description of the interference traffic reward, the content of the aggressive reward can be reused to reduce the complexity of the reward function, and the meaning of occupying the road at a low speed is already included in its content.

[0144] In an exemplary embodiment, the risk participation reward design in the optimized reward function under the intersection scenario is as follows:

[0145] Give a reward value when the adversarial vehicle collides with the vehicle under test, and give another reward value when the adversarial vehicle and the vehicle under test have other situations than collision, and construct a reward function for punishing collision behavior.

[0146] The risk participation reward R' in the optimized reward function under the intersection scenario RiskEngagement Can actually be designed as:

[0147]

[0148] The design idea of this function is: Considering that the vehicle distances are generally close in the intersection scenario, it is not appropriate to use the safety distance to avoid inevitable collisions anymore. Therefore, a direct punishment for the collision state can be directly adopted to enhance the naturalness of the adversarial scenario.

[0149] In an exemplary embodiment, training the pre-constructed human driver model with the optimized reward function obtained according to the simulation data from the natural scene traffic flow dataset includes:

[0150] Repeat the following steps until the reward value corresponding to the calculated optimized reward function meets the preset conditions to obtain a trained human driver model:

[0151] Input the simulation data from the natural scene traffic flow dataset into the generator of the pre-constructed human driver model to obtain the action data output by the generator of the human driver model; the simulation data includes: the driving state data of the training adversarial vehicle.

[0152] Calculate the reward value corresponding to the optimized reward function according to the simulation data and the action data, and adjust the parameters of the human driver model generator according to the obtained reward value.

[0153] In an exemplary embodiment, the preset reinforcement learning algorithm model includes: a Proximal Policy Optimization (PPO) network model. When the preset reinforcement learning algorithm model is a PPO network model, the objective function of the preset reinforcement learning algorithm model includes: a CLIP function.

[0154] By horizontally comparing multiple types of reinforcement learning methods, the PPO network can be selected as the specific implementation. The PPO algorithm is based on the policy gradient method and trains the agent by optimizing the policy to maximize the long-term return. Its core idea is to use the proximal policy optimization method to control the amplitude of policy update and avoid the problem of performance degradation caused by too drastic policy update. Specifically, the PPO algorithm uses a clipping function, that is, the CLIP function, to limit the update amplitude of the network, restricting the difference between the new policy and the old policy within a given range, thus effectively ensuring the balance of the network update speed and amplitude, and is suitable for constructing the adversarial scenario of the autonomous driving scenario.

[0155] In an exemplary embodiment, as Figure 5 shown, the method further includes:

[0156] Step 300: Obtain the driving state data of the background vehicle;

[0157] Step 310: Input the obtained driving state data of the background vehicle into the surrogate model to obtain the action data of the background vehicle output by the surrogate model;

[0158] Step 320: Control the background vehicle to move in the traffic simulation environment according to the action data of the background vehicle to obtain the port driving scenario of the background vehicle.

[0159] During the reinforcement learning training process, in addition to the vehicle under test or the adversarial vehicle, there are many other traffic participants that form a complete simulation environment. To accurately simulate the interactions between vehicles, it is crucial to describe the behaviors of these participants. These participants can be referred to as background vehicles. Since the Intelligent Driver Model (IDM) and the MOBIL model (lane-changing model) have the characteristics of a small number of parameters and clear meanings, the background vehicles can use the IDM model and the MOBIL model to control their straight-line driving and lane-changing behaviors, realizing vehicle driving simulation with a relatively low computational complexity. With the characteristics of a small number of parameters and clear meanings, it is relatively easy to analyze their extreme performance. When using the IDM model and the MOBIL model to control the driving behaviors of background vehicles, the key parameters of the IDM and MOBIL models can also be adjusted by analyzing the characteristics of the existing data and matching the characteristics of the port environment, including the target speed for straight-line driving, the safe distance between vehicles, the maximum acceleration, the comfortable deceleration, as well as the minimum acceleration gain and maximum braking at intersections.

[0160] Controlling the movement of the background vehicle in the traffic simulation environment based on the obtained action data of the background vehicle also requires a traffic simulation software to convert the action data into behavior instructions that the simulation platform can understand and apply, and then input them into the simulation platform to obtain the driving scenario of the background vehicle.

[0161] Similarly, since the movement of the background vehicle in the traffic simulation environment is a continuous process with constantly changing motion states, the action data of the background vehicle is obtained from time to time. Whenever the action data of the background vehicle at a certain moment is obtained, the movement of the background vehicle in the traffic simulation environment is controlled based on the obtained action data of the background vehicle. Therefore, a rich and variable motion state of the background vehicle similar to that of real vehicles is constructed. Specifically, for any two adjacent moments (the two moments are respectively referred to as the previous moment and the next moment), after obtaining the action data of the background vehicle at the previous moment, the movement of the background vehicle in the traffic simulation environment is controlled based on the action data of the background vehicle at the previous moment, and this motion state is maintained until the action data of the background vehicle at the next moment is obtained. Then, the movement of the background vehicle in the traffic simulation environment is controlled based on the action data of the background vehicle at the next moment.

[0162] Exemplary device

[0163] The embodiment of the present application also provides a device for generating a port driving scenario, as Figure 6 shown, including:

[0164] A data acquisition unit 400, configured to acquire the driving state data of a selected adversarial vehicle;

[0165] A data processing unit 410 for inputting the obtained driving state data of the selected adversarial vehicle into a port driving model to obtain the action data of the selected adversarial vehicle output by the port driving model;

[0166] A scenario generation unit 420 for controlling the movement of the selected adversarial vehicle in a traffic simulation environment according to the obtained action data of the selected adversarial vehicle to obtain a port driving scenario of the adversarial vehicle;

[0167] Wherein, the port driving model is obtained by training a human driver model through port vehicle behavior feature information and an optimized reward function, the optimized reward function is jointly constructed based on an agent of a large language model and expert experience, and the human driver model includes: a generator adopting a preset reinforcement learning algorithm model.

[0168] In an exemplary embodiment, the data processing unit 410 is further configured to:

[0169] Obtain port vehicle behavior feature information;

[0170] Design an initial reward function based on the obtained port vehicle behavior feature information by using the agent, and obtain an optimized reward function that integrates expert experience based on the designed initial reward function;

[0171] Train a pre-constructed human driver model according to simulation data from a natural scene traffic flow dataset and the obtained optimized reward function, and use the generator in the trained human driver model as the port driving model.

[0172] In an exemplary embodiment, the feature information of the port vehicle is obtained through a questionnaire survey. The recipients of the questionnaire survey are port fleet management personnel meeting preset conditions. The questionnaire survey is designed using a cross-sectional survey method and is stipulated to adopt a Likert scale response format.

[0173] In an exemplary embodiment, the data processing unit 410 is further configured to:

[0174] Expand the obtained port vehicle behavior feature information by using zero-shot prompting technology;

[0175] Obtain prompt words designed by using preset technologies; wherein, the preset technologies include at least one of the following: chain of thought prompting, example diffusion technology, detail focusing technology;

[0176] Control the agent to design an initial reward function according to the expanded port vehicle behavior feature information and the designed prompt words.

[0177] In an exemplary embodiment, the initial reward function includes: an initial reward function in a straight road scenario and an initial reward function in an intersection scenario;

[0178] Correspondingly, the optimized reward function includes: an optimized reward function in a straight road scenario and an optimized reward function in an intersection scenario.

[0179] In an exemplary embodiment, both the initial reward function in the straight road scenario and the optimized reward function in the straight road scenario include: a mixed reward function obtained by combining an aggressiveness reward function, a traffic interference reward function, a risk participation reward function, and a speed incentive reward function;

[0180] Both the initial reward function in the intersection scenario and the optimized reward function in the intersection scenario include: a mixed reward function obtained by combining a right-of-way preemption reward function, an aggressiveness reward function, a traffic interference reward function, and a risk participation reward function;

[0181] Among them, the aggressiveness reward function is used to reward the attack behavior of the adversarial vehicle, and the attack behavior includes: inserting in front of the vehicle under test or occupying the lane where the vehicle under test is located;

[0182] The traffic interference reward function is used to reward the interference behavior of the adversarial vehicle against the normal traffic flow;

[0183] The risk participation reward function is used to reward the risk behavior of the adversarial vehicle, and the risk behavior includes: overtaking or braking when the distance from the vehicle under test is less than a preset distance;

[0184] The speed reward function is used to reward the adversarial vehicle for competing driving with the vehicle under test by controlling the speed;

[0185] The right-of-way preemption reward function is used to reward the adversarial vehicle for reaching the designated point of the preset intersection prior to the vehicle under test;

[0186] Among them, the interference behavior against the normal traffic flow includes: the attack behavior of the adversarial vehicle, the risk behavior of the adversarial vehicle, and the adversarial vehicle reaching the designated point of the preset intersection prior to the vehicle under test.

[0187] In an exemplary embodiment, the data processing unit 410 is further configured to:

[0188] Obtain a mixed reward function according to the aggressiveness reward function, the traffic interference reward function, the risk participation reward function, the speed incentive reward function, and the respectively corresponding preset weight coefficients;

[0189] Obtain a mixed reward function according to the right-of-way preemption reward function, the aggressiveness reward function, the traffic interference reward function, the risk participation reward function, and the respectively corresponding preset weight coefficients.

[0190] In an exemplary embodiment, the design of the aggressive reward function in the optimized reward function in the straight road scenario is as follows:

[0191] Construct a reward function for rewarding the opposing vehicle for suddenly inserting into the lane where the vehicle under test is located, based on the function item for ensuring that the opposing vehicle and the vehicle under test are in the same lane, and the function item for encouraging at least a preset safety distance between the opposing vehicle and the vehicle under test.

[0192] In an exemplary embodiment, the design of the risk participation reward function in the optimized reward function in the straight road scenario is as follows:

[0193] Give a preset reward value when the opposing vehicle and the vehicle under test are not in the same lane, and give a preset reward value and the reward value obtained from the collision reward function item when the opposing vehicle and the vehicle under test are in the same lane, to construct a reward function for rewarding the opposing vehicle for performing a sudden braking behavior; wherein, the collision reward function item is used to encourage the opposing vehicle to collide with the vehicle under test.

[0194] In an exemplary embodiment, the design of the traffic interference reward function in the optimized reward function in the straight road scenario is as follows:

[0195] Adopt the design of reusing the risk participation reward in the optimized reward function in the straight road scenario and shorten the collision time, to construct a reward function for rewarding the opposing vehicle for performing a sudden braking behavior at a higher frequency.

[0196] In an exemplary embodiment, the design of the risk participation reward function in the optimized reward function in the straight road scenario is as follows:

[0197] Construct a reward function for rewarding the opposing vehicle for competing with the vehicle under test in the lane where the vehicle under test is located, based on the function item for ensuring that the opposing vehicle and the vehicle under test are in the same lane and there is a speed difference, and the function item for encouraging at least a preset safety distance between the opposing vehicle and the vehicle under test.

[0198] In an exemplary embodiment, the design of the right-of-way seizure reward function in the optimized reward function in the intersection scenario is as follows:

[0199] Construct a reward function for rewarding the opposing vehicle for seizing the preset intersection designated point in terms of time, based on the function item for encouraging the opposing vehicle to reach the preset intersection designated point in less time than the vehicle under test.

[0200] In an exemplary embodiment, the implementation of the aggressive reward function in the optimized reward function in the intersection scenario is as follows:

[0201] Construct a reward function for rewarding the adversarial vehicle for preempting the designated point of the preset intersection in terms of distance according to a function item for encouraging the adversarial vehicle to reach the designated point of the preset intersection faster than the vehicle under test when there is a speed difference between the adversarial vehicle and the vehicle under test.

[0202] In an exemplary embodiment, the design of the interference traffic reward function in the optimized reward function in the intersection scenario is as follows:

[0203] Adopt the design of reusing the aggressive reward function in the optimized reward function in the intersection scenario to construct a reward function for rewarding the vehicle under test for occupying the lane where the vehicle under test is located.

[0204] In an exemplary embodiment, the risk participation reward design in the optimized reward function in the intersection scenario is as follows:

[0205] Give one reward value when the adversarial vehicle collides with the vehicle under test, and give another reward value when the adversarial vehicle has a situation other than collision with the vehicle under test, and construct a reward function for punishing the collision behavior.

[0206] Training the pre-constructed human driver model with the optimized reward function obtained according to the simulation data from the natural scene traffic flow data set includes:

[0207] Repeat the following steps until the reward value corresponding to the calculated optimized reward function meets the preset conditions to obtain a trained human driver model:

[0208] Input the simulation data from the natural scene traffic flow data set into the generator of the pre-constructed human driver model to obtain the action data output by the generator of the human driver model; the simulation data includes: the driving state data of the training adversarial vehicle.

[0209] Calculate the reward value corresponding to the optimized reward function according to the simulation data and the action data, and adjust the parameters of the generator of the human driver model according to the obtained reward value.

[0210] In an exemplary embodiment, the preset reinforcement learning algorithm model includes: a PPO network model. When the preset reinforcement learning algorithm model is a PPO network model, the objective function of the preset reinforcement learning algorithm model includes: a CLIP function.

[0211] In an exemplary embodiment, the scenario generation unit 420 is further configured to:

[0212] Obtain the driving state data of the background vehicle;

[0213] Input the obtained driving state data of the background vehicle into the surrogate model to obtain the action data of the background vehicle output by the surrogate model;

[0214] Control the background vehicle to move in the traffic simulation environment respectively according to the action data of the background vehicle, so as to obtain the port driving scenario of the background vehicle.

[0215] The port driving scenario construction device provided in this embodiment belongs to the same inventive concept as the autonomous driving test scenario generation method provided in the above embodiments of the present application, and can execute the port driving scenario construction method provided in any of the above embodiments of the present application, and has the corresponding functional modules and beneficial effects of the execution method. For the technical details not described in detail in this embodiment, reference may be made to the specific processing content of the port driving scenario construction method provided in the above embodiments of the present application, which will not be elaborated here.

[0216] Exemplary electronic device

[0217] An embodiment of the present application also provides an electronic device, such as Figure 7 shown, including: a memory 500 and a processor 510;

[0218] The memory 500 is connected to the processor 510 and is used to store programs;

[0219] The processor 510 is used to implement the port driving scenario construction method described in any of the above embodiments by running the programs in the memory 500.

[0220] Specifically, the above electronic device may further include: a bus, a communication interface 520, an input device 530, and an output device 540.

[0221] The processor 510, the memory 500, the communication interface 520, the input device 530, and the output device 540 are interconnected through the bus. Among them:

[0222] The bus may include a path for transmitting information between various components of the computer system.

[0223] The processor 510 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0224] The processor 510 may include a main processor, and may also include a baseband chip, a modem, etc.

[0225] The program for implementing the technical solution of the present invention is stored in the memory 500, and the operating system and other key services may also be stored. Specifically, the program may include program codes, and the program codes include computer operation instructions. More specifically, the memory 500 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash memory, and so on.

[0226] The input device 530 may include devices for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor, etc.

[0227] The output device 540 may include devices for allowing information to be output to a user, such as a display screen, a printer, a speaker, etc.

[0228] The communication interface 520 may include devices of any transceiver type for communicating with other devices or communication networks, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0229] The processor 510 executes the program stored in the memory 500 and calls other devices, and can be used to implement each step of any one of the port driving scenario construction methods provided in the above embodiments of the present application.

[0230] Exemplary computer program product and storage medium

[0231] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the port driving scenario construction method according to various embodiments of the present application described in any of the above embodiments of this specification.

[0232] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages, such as Java, C++, etc., and also include conventional procedural programming languages, such as the "C" language or similar programming languages. The program code may be executed entirely on a user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0233] In addition, an embodiment of the present application further provides a storage medium, on which a computer program is stored. When the computer program is run by a processor, the method for constructing a port driving scenario described in any of the above embodiments is implemented.

[0234] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0235] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0236] The steps in the methods of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs. The technical features recorded in each embodiment can be replaced or combined.

[0237] The modules and sub-modules in the devices and terminals in the embodiments of the present application can be combined, divided, and deleted according to actual needs.

[0238] In several embodiments provided by the present application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are only illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical, or other forms.

[0239] The modules or sub-modules described as separate components may or may not be physically separated. The components as modules or sub-modules may or may not be physical modules or sub-modules, that is, they can be located in one place, or can be distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0240] In addition, in each embodiment of the present application, each functional module or sub-module can be integrated into one processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated into one module. The above-mentioned integrated module or sub-module can be implemented in the form of hardware, or in the form of a software functional module or sub-module.

[0241] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0242] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, a software unit executed by a processor, or a combination of the two. The software unit can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0243] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0244] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for constructing a port driving scene, characterized in that: include: Acquiring driving status data of the selected confrontation vehicle; Inputting the obtained driving state data of the selected opposing vehicle into the port driving model to obtain the action data of the selected opposing vehicle output by the port driving model; Controlling the selected antagonistic vehicle to move in a traffic simulation environment according to the obtained motion data of the selected antagonistic vehicle to obtain a driving scene of the antagonistic vehicle; Among them, the port driving model is obtained by training the human driver model through the port vehicle behavior characteristic information and the optimized reward function. The optimized reward function is jointly constructed by the intelligent agent based on the large language model and expert experience. The human driver model includes: a generator using a preset reinforcement learning algorithm model.

2. The method according to claim 1, characterized in that The method further comprises: Obtain information on the behavior characteristics of port vehicles; Based on the obtained port vehicle behavior characteristic information, the agent is used to design an initial reward function, and an optimized reward function is obtained by integrating expert experience on the basis of the designed initial reward function; A pre-built human driver model is trained based on the simulation data from a natural scene traffic flow dataset and the obtained optimized reward function, and the generator in the trained human driver model is used as the port driving model.

3. The method according to claim 2, characterized in that The initial reward function includes: an initial reward function in a straight road scenario and an initial reward function in an intersection scenario; Correspondingly, the optimization reward function includes: an optimization reward function in a straight road scenario and an optimization reward function in an intersection scenario.

4. The method according to claim 3, characterized in that The initial reward function in the straight-line scenario and the optimized reward function in the straight-line scenario both include: a hybrid reward function obtained by combining an aggressive reward function, a traffic interference reward function, a risk participation reward function, and a speed incentive reward function; The initial reward function in the intersection scenario and the optimized reward function in the intersection scenario both include: a mixed reward function obtained by combining a road right-of-way grabbing reward function, an aggressive reward function, a traffic interference reward function, and a risk participation reward function; The aggressive reward function is used to reward the aggressive behavior of the antagonistic vehicle, and the aggressive behavior includes: inserting in front of the tested vehicle or occupying the lane where the tested vehicle is located; The traffic interference reward function is used to reward the interference behavior of the adversarial vehicle on the normal traffic flow; The risk participation reward function is used to reward risky behaviors of the adversarial vehicle, the risky behaviors including: overtaking or braking when the distance to the tested vehicle is less than a preset distance; The speed reward function is used to reward the adversarial vehicle for driving competitively with the tested vehicle by controlling the speed; The road right grabbing reward function is used to reward the adversarial vehicle to arrive at a designated point at a preset intersection before the tested vehicle; Among them, the interference behaviors to normal traffic flow include: attacking behaviors of antagonistic vehicles, risky behaviors of antagonistic vehicles, and antagonistic vehicles giving priority to the tested vehicles to arrive at the designated point of the preset intersection.

5. The method according to claim 2, characterized in that: The method of training a pre-built human driver model based on the optimized reward function obtained based on the simulated data from the natural scene traffic flow data set includes: Repeat the following steps until the reward value corresponding to the calculated optimized reward function meets the preset conditions to obtain a trained human driver model: Inputting simulation data from a natural scene traffic flow data set into a pre-built generator of a human driver model to obtain motion data output by the generator of the human driver model; the simulation data includes: driving state data of an adversarial vehicle for training; A reward value corresponding to the optimized reward function is calculated according to the simulation data and the action data, and parameters of the human driver model generator are adjusted according to the obtained reward value.

6. The method according to claim 1, characterized in that The preset reinforcement learning algorithm model includes: a PPO network model. When the preset reinforcement learning algorithm model is a PPO network model, the objective function of the preset reinforcement learning algorithm model includes: a CLIP function.

7. The method according to claim 1, characterized in that The method further comprises: Obtain driving status data of background vehicles; Inputting the obtained driving state data of the background vehicle into a substitution model to obtain the action data of the background vehicle output by the substitution model; According to the motion data of the background vehicles, the background vehicles are respectively controlled to move in the traffic simulation environment to obtain the port driving scene of the background vehicles.

8. A port driving scene generation device, characterized in that: include: A data acquisition unit, used to acquire driving state data of the selected confrontation vehicle; A data processing unit, used for inputting the obtained driving state data of the selected opposing vehicle into the port driving model to obtain the action data of the selected opposing vehicle output by the port driving model; A scene generation unit, used for controlling the selected antagonistic vehicle to move in a traffic simulation environment according to the obtained motion data of the selected antagonistic vehicle, so as to obtain a port driving scene of the antagonistic vehicle; Among them, the port driving model is obtained by training the human driver model through the port vehicle behavior characteristic information and the optimized reward function, and the optimized reward function is jointly constructed based on the intelligent agent and expert experience of the large language model. The human driver model includes: a generator using a preset reinforcement learning algorithm model.

9. An electronic device, characterized in that: include: Memory and processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the port driving scene construction method as described in any one of claims 1 to 7 by running the program in the memory.

10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the method for constructing a port driving scene according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Automatic driving system multi-agent confrontation test method based on large language model

    CN121614411A

  • Multi-agent adversarial testing method for autonomous driving systems based on large language models

    CN121614411B