Driving simulation scene construction method and device, electronic equipment and program product
By training pedestrian and animal behavior simulation models using reinforcement learning algorithms and shadow proxy methods, a target driving simulation scenario is constructed, which solves the problem of insufficient interaction between simulated pedestrian and animal behavior and the environment, and improves the realism of the simulation scenario.
Patent Information
- Application Number
- CN202510983508.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-07-17
AI Technical Summary
In existing driving simulation scenarios, the behavior of simulated pedestrians and animals lacks interactivity with the traffic simulation environment, resulting in low realism of the simulation scenarios.
By using pedestrian behavior simulation models and animal behavior simulation models, target simulated pedestrians and target simulated animals are trained using reinforcement learning algorithms and shadow proxy methods, and their behavior is simulated to construct a target driving simulation scenario.
It improves the realism of driving simulation scenarios, making the behavior of simulated pedestrians and animals more consistent with changes in actual driving scenarios, thus enhancing the realism of the simulation scenarios.
Smart Images

Figure CN120509322B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of automatic driving, and particularly relates to a driving simulation scene construction method and device, an electronic device, and a program product. BACKGROUND
[0002] In the technical field of automatic driving, a driving simulation scene needs to be constructed, and then the automatic driving technology is verified and trained through the constructed driving simulation scene. When constructing the driving simulation scene, a traffic simulation environment, a simulation pedestrian, and a simulation animal need to be constructed.
[0003] Currently, when constructing the simulation pedestrian and the simulation animal in the driving simulation scene, the simulation pedestrian and the simulation animal are usually set as obstacles, or the behavior rules of the simulation pedestrian and the simulation animal are set to be too simple (such as setting the behavior rules of the simulation pedestrian and the simulation animal to move along a preset route). It can be seen that the behaviors of the simulation pedestrian and the simulation animal constructed at present will not change with the change of the traffic simulation environment, that is, the behaviors of the simulation pedestrian and the simulation animal constructed at present lack interaction with the traffic simulation environment, and in a real driving scene, the behaviors of pedestrians and animals will change with the change of the traffic environment. Therefore, compared with the real driving scene, the real driving simulation scene constructed at present is less realistic. SUMMARY
[0004] Therefore, the embodiments of the present application provide a driving simulation scene construction method, device, electronic device, and program product to solve the technical problem that the real driving simulation scene constructed at present is less realistic.
[0005] In a first aspect, the embodiments of the present application provide a driving simulation scene construction method, comprising:
[0006] An initial driving simulation scene constructed in advance is acquired, and the initial driving simulation scene at least includes a traffic simulation environment, an initial simulation pedestrian, and an initial simulation animal;
[0007] A behavior of the initial simulation pedestrian is simulated based on the traffic simulation environment through a pedestrian behavior simulation model to obtain a target simulation pedestrian, and the pedestrian behavior simulation model is trained through a reinforcement learning algorithm and a shadow agent mode;
[0008] A behavior of the initial simulation animal is simulated based on the traffic simulation environment through an animal behavior simulation model to obtain a target simulation animal, and the animal behavior simulation model is trained through only a reinforcement learning algorithm;
[0009] A target driving simulation scene is constructed according to the traffic simulation environment, the target simulation pedestrian, and the target simulation animal.
[0010] Optionally, the pedestrian behavior simulation model comprises a target pedestrian strategy model and a shadow agent model, the shadow agent model being trained by a preset pedestrian walking strategy; and the pedestrian behavior simulation model is specifically trained by the following manner:
[0011] obtaining a pre-trained initial pedestrian strategy model and the shadow agent model;
[0012] simulating behaviors of the same training pedestrian by the initial pedestrian strategy model and the shadow agent model respectively according to the same training environment, to obtain training pedestrian behaviors corresponding to the initial pedestrian strategy model and the shadow agent model;
[0013] determining punishment reward information by comparing the training pedestrian behaviors corresponding to the initial pedestrian strategy model and the shadow agent model;
[0014] feeding back the punishment reward information to the initial pedestrian strategy model, to instruct the initial pedestrian strategy model to optimize according to the punishment reward information, to obtain the target pedestrian strategy model.
[0015] Optionally, the pedestrian behavior simulation model further comprises an evaluation model trained; and the simulating behaviors of the initial simulation pedestrian based on the traffic simulation environment by the pedestrian behavior simulation model to obtain the target simulation pedestrian comprises:
[0016] simulating behaviors of the initial simulation pedestrian based on the traffic simulation environment by the target pedestrian strategy model and the shadow agent model respectively, to obtain a first simulation pedestrian output by the target pedestrian strategy model and a second simulation pedestrian output by the shadow agent model;
[0017] determining the target simulation pedestrian from the first simulation pedestrian and the second simulation pedestrian by the evaluation model.
[0018] Optionally, after the simulating behaviors of the initial simulation pedestrian based on the traffic simulation environment by the pedestrian behavior simulation model to obtain the target simulation pedestrian, the method further comprises:
[0019] if the first simulation pedestrian is determined as the target simulation pedestrian, outputting target reward information to the target pedestrian strategy model, to instruct the target pedestrian strategy model to optimize according to the target reward information;
[0020] If the second simulated pedestrian is determined as the target simulated pedestrian, target penalty information is output to the target pedestrian strategy model to instruct the target pedestrian strategy model to optimize according to the target penalty information.
[0021] Optionally, the animal behavior simulation model is obtained by training in the following manner:
[0022] An initial animal strategy model is obtained by pre-training;
[0023] The behavior of a training animal is simulated according to the training environment by using the initial animal strategy model, to obtain a training animal behavior;
[0024] Penalty reward information of the training animal behavior is determined according to a preset penalty reward function and the training animal behavior; the penalty reward function is used to reward behaviors of the training animal behavior that are close to a food source and far away from a vehicle, and is used to punish behaviors of the training animal behavior that cause collisions and / or excessive energy consumption;
[0025] The penalty reward information is fed back to the initial animal strategy model to instruct the initial animal strategy model to optimize according to the penalty reward information, to obtain the animal behavior simulation model.
[0026] Optionally, the initial driving simulation scene is constructed in the following manner:
[0027] Initial driving simulation scene data is obtained; the initial driving simulation scene data includes initial simulated pedestrian data and initial simulated animal data, and includes one or more of the following traffic simulation environment data: map data, road topology data, building data, obstacle data, traffic sign data, dynamic traffic signal data, dynamic weather data, dynamic vehicle data, and virtual dynamic sensor data;
[0028] The initial driving simulation scene is constructed according to the initial simulated pedestrian data, the initial simulated animal data, and the traffic simulation environment data.
[0029] Optionally, the method further comprises:
[0030] Training environment data corresponding to each training environment is obtained;
[0031] The strategy generalization model trained is used to perform strategy generalization processing on the pedestrian behavior simulation model and the animal behavior simulation model based on the training environment data, and the pedestrian behavior simulation model after strategy generalization processing and the animal behavior simulation model after strategy generalization processing are output; the pedestrian behavior simulation model after strategy generalization processing has a motion difference degree between each target simulation pedestrian obtained according to the changed traffic simulation environment lower than a first preset threshold, and the animal behavior simulation model after strategy generalization processing has a motion difference degree between each target simulation animal obtained according to the changed traffic simulation environment lower than a second preset threshold.
[0032] In a second aspect, an embodiment of the present application provides a driving simulation scene construction device, comprising:
[0033] An initial scene acquisition unit is configured to acquire a pre-constructed initial driving simulation scene, wherein the initial driving simulation scene at least includes a traffic simulation environment, initial simulation pedestrians, and initial simulation animals.
[0034] A simulation pedestrian simulation unit is configured to simulate behaviors of the initial simulation pedestrians based on the traffic simulation environment by using a pedestrian behavior simulation model to obtain target simulation pedestrians, wherein the pedestrian behavior simulation model is trained by using a reinforcement learning algorithm and a shadow agent method.
[0035] A simulation animal simulation unit is configured to simulate behaviors of the initial simulation animals based on the traffic simulation environment by using an animal behavior simulation model to obtain target simulation animals, wherein the animal behavior simulation model is trained by using only a reinforcement learning algorithm.
[0036] A target scene construction unit is configured to construct a target driving simulation scene according to the traffic simulation environment, the target simulation pedestrians, and the target simulation animals.
[0037] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements each step of the driving simulation scene construction method according to the first aspect when executing the computer program.
[0038] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executable on a processor to implement each step of the driving simulation scene construction method according to the first aspect.
[0039] In a fifth aspect, an embodiment of the present application provides a computer program which, when running on an electronic device, causes the electronic device to perform each step of the method for constructing a driving simulation scene according to the first aspect.
[0040] The method, device, electronic device and program product for constructing a driving simulation scene provided by the embodiments of the present application have the following beneficial effects:
[0041] In the method for constructing a driving simulation scene provided by the embodiments of the present application, first, an initial driving simulation scene is obtained, wherein the initial driving simulation scene at least includes a traffic simulation environment, initial simulation pedestrians and initial simulation animals; then, a pedestrian behavior simulation model is used to simulate the behavior of the initial simulation pedestrians based on the traffic simulation environment, to obtain target simulation pedestrians; the pedestrian behavior simulation model is trained by a reinforcement learning algorithm and a shadow agent method; then, an animal behavior simulation model is used to simulate the behavior of the initial simulation animals based on the traffic simulation environment, to obtain target simulation animals; the animal behavior simulation model is trained by only the reinforcement learning algorithm; finally, a target driving simulation scene is constructed according to the traffic simulation environment, the target simulation pedestrians and the target simulation animals. The target simulation pedestrians and the target simulation animals in the driving simulation scene constructed by the method are both simulated based on the traffic simulation environment, so the driving simulation scene constructed is more similar to a real driving scene; in addition, in the method, the pedestrian behavior simulation model used to simulate the simulation pedestrians is trained by the reinforcement learning algorithm and the shadow agent method, and the animal behavior simulation model used to simulate the simulation animals is trained by only the reinforcement learning algorithm, so that the behavior of the target simulation pedestrians is more normative, and the behavior of the target simulation animals is more random, which corresponds to the behavior of the pedestrians and the behavior of the animals in the real driving scene, so the real driving simulation scene constructed can be further improved. BRIEF DESCRIPTION OF DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0043] Figure 1 An implementation flowchart of the method for constructing a driving simulation scene provided by the embodiments of the present application is shown in the figure.
[0044] Figure 2 A structural schematic diagram of the driving simulation scene construction device provided by the embodiments of the present application is shown in the figure.
[0045] Figure 3 A structural schematic diagram of an electronic device is provided for an embodiment of the present application. DETAILED DESCRIPTION
[0046] It should be noted that the terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application. In the description of the embodiments of the present application, unless otherwise specified, "a plurality of" means two or more than two, "at least one", "one or more" means one, two or more than two. The terms "first", "second" are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the "first", "second" features can explicitly or implicitly include one or more features.
[0047] In this specification, the reference to "one embodiment" or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. Therefore, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in other some embodiments" and the like appearing in various places in the specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically noted. The terms "include", "contain", "have" and their variants mean "including but not limited to", unless otherwise specifically noted.
[0048] The execution subject of the driving simulation scene construction method provided by the embodiments of the present application can be an electronic device, which can execute each step of the driving simulation scene construction method provided by the embodiments of the present application. The electronic device can include, but is not limited to, a mobile phone, a tablet computer, a notebook computer, a desktop computer, and the like.
[0049] The driving simulation scene construction method provided by the embodiments of the present application can be applied to various scenarios that need to construct a driving simulation scene. For example, when a user needs to construct a driving simulation scene for optimizing or verifying an automatic driving technology, each step of the driving simulation scene construction method provided by the embodiments of the present application can be executed through an electronic device, so as to construct a driving simulation scene with high authenticity.
[0050] Please refer to Figure 1 , Figure 1 An implementation flowchart of a driving simulation scene construction method is provided for an embodiment of the present application. The driving simulation scene construction method can include S101-S104, which are described in detail as follows:
[0051] In S101, an initial driving simulation scene previously constructed is acquired, the initial driving simulation scene at least including a traffic simulation environment, initial simulation pedestrians, and initial simulation animals.
[0052] In the embodiments of the present application, when a target driving simulation scene needs to be constructed, the electronic device can first acquire an initial driving simulation scene including at least a traffic simulation environment, initial simulation pedestrians, and initial simulation animals.
[0053] The initial simulation pedestrians can be simulation pedestrians that have not been subjected to behavior simulation, and the initial simulation animals can be simulation animals that have not been subjected to behavior simulation.
[0054] The traffic simulation environment can at least include any one or more of the following: a road, a vehicle, a building, an obstacle, a plant, weather, and a sensor, etc.
[0055] Optionally, the electronic device can acquire initial driving simulation scene data, the initial driving simulation scene data including initial simulation pedestrian data and initial simulation animal data, and including one or more of the following traffic simulation environment data: map data, road topology data, building data, obstacle data, traffic sign data, dynamic traffic signal data, dynamic weather data, dynamic vehicle data, and virtual dynamic sensor data.
[0056] After the initial simulation pedestrian data, the initial simulation animal data, and the traffic simulation environment data are acquired, a dynamic initial driving simulation scene is constructed.
[0057] In actual applications, the electronic device can use three-dimensional map construction technology, rendering technology, physical simulation engines, sensor simulation technology, real-time path planning algorithms, and collision detection algorithms to construct a dynamic initial driving simulation scene according to the initial driving simulation scene data.
[0058] In actual applications, the initial driving simulation scene data can be input by a user to generate an initial driving simulation scene required by the user; in addition, the initial driving simulation scene data can also be randomly generated to generate a random initial driving simulation scene, so that a random target driving simulation scene can be constructed to meet various actual requirements.
[0059] In S102, a pedestrian behavior simulation model is used to simulate the behavior of the initial simulation pedestrians based on the traffic simulation environment to obtain target simulation pedestrians; the pedestrian behavior simulation model is trained by a reinforcement learning algorithm and a shadow agent method.
[0060] In the embodiments of the present application, after the initial driving simulation scene is acquired, the electronic device can first acquire a pedestrian behavior simulation model trained by a reinforcement learning algorithm and a shadow agent method.
[0061] The "shadow agent mode" in the present application is explained as follows:
[0062] In the field of automatic driving, the "shadow agent mode" refers to a technology in which the automatic driving system still runs but does not control the vehicle in the manned state, and only compares the actual operation of the driver with the simulated decision-making to optimize the algorithm itself. The "shadow agent mode" in the present application is similar to the "shadow agent mode" in the field of automatic driving. The "shadow agent mode" in the present application refers to: first, training a shadow agent model according to a preset pedestrian walking strategy, then comparing the training pedestrian behavior corresponding to the shadow agent model with the training pedestrian behavior corresponding to the initial pedestrian strategy model pre-trained by the reinforcement learning algorithm, and optimizing the initial pedestrian strategy model by the reinforcement learning algorithm, and finally obtaining a target pedestrian strategy model.
[0063] Based on this, in a possible implementation, the pedestrian behavior simulation model can include a target pedestrian strategy model and a shadow agent model, and the electronic device can train the pedestrian behavior simulation model by the reinforcement learning algorithm and the shadow agent mode described below. Details are as follows:
[0064] First, the electronic device can obtain a pre-trained initial pedestrian strategy model and a shadow agent model.
[0065] In the present implementation, the pre-trained initial pedestrian strategy model can be trained by a preset reinforcement learning algorithm.
[0066] The shadow agent model can be trained by a preset pedestrian walking strategy. The preset pedestrian walking strategy can be determined according to the real behavior of real pedestrians in various traffic environments. Based on this, the user can obtain the real behavior of real pedestrians in various traffic environments, and obtain the pedestrian walking strategy according to the real behavior of real pedestrians in various traffic environments, so as to construct the shadow agent model according to the pedestrian walking strategy.
[0067] After obtaining the initial pedestrian strategy model and the shadow agent model, the electronic device can optimize the initial pedestrian strategy model by the reinforcement learning algorithm and the shadow agent model to obtain a target pedestrian strategy model.
[0068] Specifically, the electronic device can simulate the behavior of the same training pedestrian according to the same training environment by the initial pedestrian strategy model and the shadow agent model respectively, to obtain the training pedestrian behavior corresponding to the initial pedestrian strategy model and the training pedestrian behavior corresponding to the shadow agent model.
[0069] Afterwards, the electronic device can determine the punishment reward information by comparing the training pedestrian behavior corresponding to the initial pedestrian strategy model and the training pedestrian behavior corresponding to the shadow agent model. Specifically, when the training pedestrian behavior corresponding to the initial pedestrian strategy model and the training pedestrian behavior corresponding to the shadow agent model are the same, the electronic device can determine the punishment reward information as the reward information; when the training pedestrian behavior corresponding to the initial pedestrian strategy model and the training pedestrian behavior corresponding to the shadow agent model are different, the electronic device can determine the punishment reward information as the punishment information.
[0070] Finally, the electronic device can feed back the punishment reward information to the initial pedestrian strategy model, so as to instruct the initial pedestrian strategy model to optimize according to the punishment reward information to obtain the target pedestrian strategy model, so that the pedestrian behavior simulation model can be obtained.
[0071] In the implementation, the shadow agent model is trained by the preset pedestrian walking strategy, and the preset pedestrian walking strategy can be determined according to the real behavior of real pedestrians in various traffic environments. Based on this, it can be considered that the target simulation pedestrian obtained by the shadow agent model has strong normativity, that is, the difference between the target simulation pedestrian obtained by the shadow agent model and the real pedestrian is low.
[0072] Since the target pedestrian strategy model is obtained by training and optimizing the shadow agent model, it can be considered that the target simulation pedestrian obtained by the target pedestrian strategy model also has strong normativity. In addition, since the initial pedestrian strategy model used to train the target pedestrian strategy model is trained by the preset reinforcement learning algorithm, it can be considered that the target simulation pedestrian obtained by the target pedestrian strategy model also has certain randomness, that is, the difference between the target simulation pedestrian obtained by the target pedestrian strategy model and the real pedestrian is high under a certain probability.
[0073] In summary, the same point of the target pedestrian strategy model and the shadow agent model is that both the target pedestrian strategy model and the shadow agent model can simulate the behavior of the simulation pedestrian according to the input traffic simulation environment to obtain the simulation pedestrian. The difference between the target pedestrian strategy model and the shadow agent model is that the normativity of the target simulation pedestrian output by the shadow agent model is higher than that of the target simulation pedestrian output by the target pedestrian strategy model, and the target simulation pedestrian output by the target pedestrian strategy model has certain randomness.
[0074] The target simulation pedestrian obtained by the target pedestrian strategy model and the shadow agent model together has the following advantages: if only the shadow agent model is used to obtain the target simulation pedestrian, the target simulation pedestrian obtained completely follows the preset pedestrian walking strategy, and since the preset pedestrian walking strategy is not an absolutely optimal pedestrian walking strategy, the target simulation pedestrian obtained only by using the shadow agent model is not an optimal simulation pedestrian in some cases. The target simulation pedestrian obtained by the target pedestrian strategy model and the shadow agent model together retains the normativity of the target simulation pedestrian output by the shadow agent model and the randomness of the target simulation pedestrian output by the target pedestrian strategy model, and the beneficial effect of retaining the randomness of the output target simulation pedestrian is that when the preset pedestrian walking strategy corresponds to a non-optimal simulation pedestrian in a certain scenario, retaining the randomness of the output target simulation pedestrian can enable the pedestrian behavior simulation model to still output an optimal simulation pedestrian in the scenario, and therefore, compared with the target simulation pedestrian obtained only by using the shadow agent model, the target simulation pedestrian obtained by the target pedestrian strategy model and the shadow agent model together is an optimal simulation pedestrian in various cases.
[0075] In a possible implementation, the pedestrian behavior simulation model further includes a trained evaluation model. The evaluation model can be used to evaluate the first simulation pedestrian output by the target pedestrian strategy model and the second simulation pedestrian output by the shadow agent model to determine the target simulation pedestrian from the first simulation pedestrian output by the target pedestrian strategy model and the second simulation pedestrian output by the shadow agent model.
[0076] Based on this, by the pedestrian behavior simulation model, the behavior of the initial simulation pedestrian is simulated based on the traffic simulation environment to obtain the target simulation pedestrian, which can include: first, by the target pedestrian strategy model and the shadow agent model, the behavior of the initial simulation pedestrian is simulated based on the traffic simulation environment to obtain the first simulation pedestrian output by the target pedestrian strategy model and the second simulation pedestrian output by the shadow agent model; and then, by the evaluation model, the target simulation pedestrian is determined from the first simulation pedestrian and the second simulation pedestrian.
[0077] The user can input the evaluation rule into the initial evaluation model to obtain the evaluation model. In actual applications, the evaluation rule can be set according to actual needs, and the user can determine the evaluation rule according to the characteristics of the target simulation pedestrian that the pedestrian behavior simulation model is expected to output. For example, when the user wants the behavior of the target simulation pedestrian output by the pedestrian behavior simulation model to be more risk-averse, the evaluation rule can be: the higher the possibility of collision with a vehicle, the lower the score of the behavior of the target simulation pedestrian; and the lower the possibility of collision with a vehicle, the higher the score of the behavior of the target simulation pedestrian.
[0078] In the implementation, after obtaining the target simulation pedestrian, if the first simulation pedestrian is determined as the target simulation pedestrian, the target reward information is output to the target pedestrian strategy model to instruct the target pedestrian strategy model to optimize according to the target reward information; if the second simulation pedestrian is determined as the target simulation pedestrian, the target penalty information is output to the target pedestrian strategy model to instruct the target pedestrian strategy model to optimize according to the target penalty information.
[0079] By continuously optimizing the target pedestrian strategy model through the reinforcement learning algorithm, the target simulation pedestrian obtained through the pedestrian behavior simulation model can be more satisfied with the evaluation rules input by the user, and finally a pedestrian behavior simulation model satisfying the user can be obtained, thereby improving the authenticity of the driving simulation scene constructed.
[0080] In the embodiment of the application, after obtaining the trained pedestrian behavior simulation model, the electronic device can input the traffic simulation environment and the initial simulation pedestrian into the trained pedestrian behavior simulation model to instruct the trained pedestrian behavior simulation model to output the target simulation pedestrian according to the traffic simulation environment and the initial simulation pedestrian.
[0081] It can be understood that different traffic simulation environments and initial simulation pedestrians can make the pedestrian behavior simulation model output different target simulation pedestrians. Compared with the prior art in which the behavior of the simulation pedestrian does not change with the change of the traffic simulation environment, the target simulation pedestrian output by the method has higher authenticity.
[0082] In S103, the behavior of the initial simulation animal is simulated based on the traffic simulation environment through the animal behavior simulation model to obtain a target simulation animal; the animal behavior simulation model is trained only through a reinforcement learning algorithm.
[0083] In the embodiment of the application, after obtaining the initial driving simulation scene, the electronic device can first obtain the animal behavior simulation model trained only through the reinforcement learning algorithm.
[0084] The "only through the reinforcement learning algorithm" in the application is explained as follows:
[0085] In contrast to the "through the reinforcement learning algorithm and the shadow agent method" in the application, the "only through the reinforcement learning algorithm" in the application refers to a method of training the animal behavior simulation model only using the reinforcement learning algorithm without using the shadow agent method.
[0086] In one possible implementation, the electronic device can use the "only through the reinforcement learning algorithm" described below to train the animal behavior simulation model. Details are as follows:
[0087] First, the electronic device can first obtain a pre-trained initial animal strategy model.
[0088] In the implementation, the pre-trained initial animal strategy model can be obtained by a preset reinforcement learning algorithm and based on a preset animal action strategy.
[0089] After obtaining the initial animal strategy model, the electronic device can optimize the initial animal strategy model by the reinforcement learning algorithm to obtain the animal behavior simulation model.
[0090] Specifically, the electronic device can simulate the behavior of the training animal according to the training environment by the initial animal strategy model to obtain the training animal behavior.
[0091] Then, the electronic device can determine the penalty reward information of the training animal behavior according to the preset penalty reward function and the training animal behavior, wherein the penalty reward function is used to reward the behavior of approaching the food source and the behavior of moving away from the vehicle in the training animal behavior, and is used to punish the behavior of causing collision and / or excessive energy consumption in the training animal behavior.
[0092] For example, when a certain training animal behavior is the behavior of approaching the food source, the penalty reward information of the training animal behavior can be determined as reward information; when a certain training animal behavior is the behavior of moving away from the vehicle, the penalty reward information of the training animal behavior can be determined as reward information; and when a certain training animal behavior is the behavior of causing collision and / or excessive energy consumption, the penalty reward information of the training animal behavior can be determined as penalty information.
[0093] After obtaining the penalty reward information, the electronic device can feed back the penalty reward information to the initial animal strategy model to instruct the initial animal strategy model to optimize according to the penalty reward information to obtain the animal behavior simulation model.
[0094] In the implementation, since the initial animal strategy model used to train the animal behavior simulation model is obtained by the preset reinforcement learning algorithm, it can be considered that the target simulation animal obtained by the animal behavior simulation model has high randomness.
[0095] The following compares the pedestrian behavior simulation model with the animal behavior simulation model.
[0096] Since the target pedestrian strategy model and the shadow agent model are included in the pedestrian behavior simulation model, the target simulation pedestrian obtained through the shadow agent model has strong normativity, and the target simulation pedestrian obtained through the target pedestrian strategy model has certain normativity and randomness. Since the animal behavior simulation model does not include the shadow agent model, it can be considered that the target simulation animal output by the animal behavior simulation model has higher randomness and lower normativity than the target simulation pedestrian output by the target pedestrian strategy model, and the target simulation pedestrian output by the target pedestrian strategy model has lower randomness and higher normativity.
[0097] In actual driving scenarios, the randomness of animal behavior is high and the normativity is low, and the randomness of pedestrian behavior is low and the normativity is high. Therefore, the target simulation animal and the target simulation pedestrian output by the method correspond to the behavior of pedestrians and animals in actual driving scenarios, and it can be seen that the driving simulation scene constructed by the application has high reality.
[0098] In S104, a target driving simulation scene is constructed according to the traffic simulation environment, the target simulation pedestrian, and the target simulation animal.
[0099] In the embodiment of the application, after obtaining the target simulation pedestrian and the target simulation animal, the electronic device can combine the traffic simulation environment, the target simulation pedestrian, and the target simulation animal into a target driving simulation scene.
[0100] It can be seen from the above that in the method for constructing a driving simulation scene provided in the embodiments of the present application, an initial driving simulation scene is first obtained, wherein the initial driving simulation scene at least includes a traffic simulation environment, initial simulation pedestrians and initial simulation animals; then the behavior of the initial simulation pedestrians is simulated based on the traffic simulation environment by using a pedestrian behavior simulation model to obtain target simulation pedestrians; the pedestrian behavior simulation model is trained by using a reinforcement learning algorithm and a shadow agent method; the behavior of the initial simulation animals is simulated based on the traffic simulation environment by using an animal behavior simulation model to obtain target simulation animals; the animal behavior simulation model is trained by using only a reinforcement learning algorithm; and finally, a target driving simulation scene is constructed according to the traffic simulation environment, the target simulation pedestrians and the target simulation animals. The target simulation pedestrians and the target simulation animals in the driving simulation scene constructed by the method are both simulated based on the traffic simulation environment, so that the driving simulation scene constructed by the method is more similar to a real driving scene. In addition, in the method, the pedestrian behavior simulation model used to simulate the simulation pedestrians is trained by using a reinforcement learning algorithm and a shadow agent method, and the animal behavior simulation model used to simulate the simulation animals is trained by using only a reinforcement learning algorithm, so that the behavior of the target simulation pedestrians is more normative, and the behavior of the target simulation animals is more random, which corresponds to the behavior of the pedestrians and the behavior of the animals in a real driving scene, and thus the reality of the driving simulation scene constructed by the method can be further improved.
[0101] In a possible implementation, in order to make the pedestrian behavior simulation model and the animal behavior simulation model applicable to various different traffic simulation environments, and to improve the robustness of the pedestrian behavior simulation model and the animal behavior simulation model, the electronic device can first obtain training environment data corresponding to each training environment, wherein the training environment can be used when training the pedestrian behavior simulation model and the animal behavior simulation model. Then, the electronic device can perform strategy generalization processing on the pedestrian behavior simulation model and the animal behavior simulation model based on the training environment data by using the trained strategy generalization model, and output the pedestrian behavior simulation model after strategy generalization processing and the animal behavior simulation model after strategy generalization processing, wherein the action difference degree between each target simulation pedestrian obtained by the pedestrian behavior simulation model after strategy generalization processing according to a changed traffic simulation environment is lower than a first preset threshold, and the action difference degree between each target simulation animal obtained by the animal behavior simulation model after strategy generalization processing according to a changed traffic simulation environment is lower than a second preset threshold. In actual applications, the first preset threshold and the second preset threshold can be set according to actual requirements.
[0102] The pedestrian behavior simulation model after strategy generalization processing and the animal behavior simulation model after strategy generalization processing can be obtained through the above method. Since the action difference degrees between the target simulation pedestrians and the action difference degrees between the target simulation animals output by the pedestrian behavior simulation model after strategy generalization processing and the animal behavior simulation model after strategy generalization processing are low, it can be considered that the pedestrian behavior simulation model after strategy generalization processing and the animal behavior simulation model after strategy generalization processing can be applied to various different traffic simulation environments.
[0103] Based on the method for constructing a driving simulation scene provided in the above embodiment, an embodiment of the present application further provides a device for constructing a driving simulation scene for implementing the above method embodiment. Please refer to Figure 2 , Figure 2 FIG. 1 is a structural schematic diagram of a device for constructing a driving simulation scene provided in an embodiment of the present application. As shown in FIG. 1, the device 20 for constructing a driving simulation scene can include an initial scene acquisition unit 21, a simulation pedestrian simulation unit 22, a simulation animal simulation unit 23, and a target scene construction unit 24. Wherein: Figure 2
[0104] The initial scene acquisition unit 21 is configured to acquire a pre-constructed initial driving simulation scene, and the initial driving simulation scene at least includes a traffic simulation environment, initial simulation pedestrians, and initial simulation animals.
[0105] The simulation pedestrian simulation unit 22 is configured to simulate the behavior of the initial simulation pedestrians based on the traffic simulation environment through a pedestrian behavior simulation model to obtain target simulation pedestrians, and the pedestrian behavior simulation model is trained through a reinforcement learning algorithm and a shadow agent method.
[0106] The simulation animal simulation unit 23 is configured to simulate the behavior of the initial simulation animals based on the traffic simulation environment through an animal behavior simulation model to obtain target simulation animals, and the animal behavior simulation model is trained through only a reinforcement learning algorithm.
[0107] The target scene construction unit 24 is configured to construct a target driving simulation scene according to the traffic simulation environment, the target simulation pedestrians, and the target simulation animals.
[0108] Optionally, the pedestrian behavior simulation model includes a target pedestrian strategy model and a shadow agent model, and the shadow agent model is trained through a preset pedestrian walking strategy. The device 20 for constructing a driving simulation scene can further include a first training unit, wherein:
[0109] The first training unit is specifically configured to:
[0110] acquire a pre-trained initial pedestrian strategy model and a shadow agent model;
[0111] The initial pedestrian strategy model and the shadow agent model are used to simulate behaviors of the same training pedestrian respectively according to the same training environment, and the training pedestrian behavior corresponding to the initial pedestrian strategy model and the training pedestrian behavior corresponding to the shadow agent model are obtained.
[0112] The training pedestrian behavior corresponding to the initial pedestrian strategy model and the training pedestrian behavior corresponding to the shadow agent model are compared to determine the punishment reward information.
[0113] The punishment reward information is fed back to the initial pedestrian strategy model to instruct the initial pedestrian strategy model to optimize according to the punishment reward information, and a target pedestrian strategy model is obtained.
[0114] Optionally, the simulation pedestrian simulation unit 22 is specifically configured to:
[0115] The behaviors of the initial simulation pedestrian are simulated based on the traffic simulation environment by the target pedestrian strategy model and the shadow agent model, and a first simulation pedestrian output by the target pedestrian strategy model and a second simulation pedestrian output by the shadow agent model are obtained.
[0116] The target simulation pedestrian is determined from the first simulation pedestrian and the second simulation pedestrian by the evaluation model.
[0117] Optionally, the driving simulation scene construction device 20 can further include an optimization unit. Wherein:
[0118] The optimization unit is specifically configured to:
[0119] If the first simulation pedestrian is determined as the target simulation pedestrian, target reward information is output to the target pedestrian strategy model to instruct the target pedestrian strategy model to optimize according to the target reward information.
[0120] If the second simulation pedestrian is determined as the target simulation pedestrian, target punishment information is output to the target pedestrian strategy model to instruct the target pedestrian strategy model to optimize according to the target punishment information.
[0121] Optionally, the driving simulation scene construction device 20 can further include a second training unit. Wherein:
[0122] The second training unit is specifically configured to:
[0123] The pre-trained initial animal strategy model is obtained;
[0124] The behavior of the training animal is simulated according to the training environment by the initial animal strategy model, and the training animal behavior is obtained.
[0125] The punishment reward information is fed back to the initial animal strategy model, so as to instruct the initial animal strategy model to be optimized according to the punishment reward information, and obtain the animal behavior simulation model.
[0126] The punishment reward information is fed back to the initial animal strategy model, so as to instruct the initial animal strategy model to be optimized according to the punishment reward information, and obtain the animal behavior simulation model.
[0127] Optionally, the initial scene acquisition unit 21 is specifically configured to:
[0128] acquire initial driving simulation scene data; the initial driving simulation scene data includes initial simulation pedestrian data and initial simulation animal data, and includes one or more of the following traffic simulation environment data: map data, road topology data, building data, obstacle data, traffic sign data, dynamic traffic signal data, dynamic weather data, dynamic vehicle data, and virtual dynamic sensor data;
[0129] According to the initial simulation pedestrian data, the initial simulation animal data and the traffic simulation environment data, a dynamic initial driving simulation scene is constructed.
[0130] Optionally, the driving simulation scene construction apparatus 20 can further include a strategy generalization unit, wherein:
[0131] The strategy generalization unit is specifically configured to:
[0132] acquire training environment data corresponding to each training environment;
[0133] The strategy generalization model trained is used to perform strategy generalization processing on the pedestrian behavior simulation model and the animal behavior simulation model based on the training environment data, and output the strategy generalization processed pedestrian behavior simulation model and the strategy generalization processed animal behavior simulation model; the action difference degree between each target simulation pedestrian obtained according to the changed traffic simulation environment is lower than a first preset threshold value for the strategy generalization processed pedestrian behavior simulation model, and the action difference degree between each target simulation animal obtained according to the changed traffic simulation environment is lower than a second preset threshold value for the strategy generalization processed animal behavior simulation model.
[0134] It should be noted that the information interaction, execution process and the like between the above units, since based on the same concept as the method embodiments of the present application, the specific functions and the technical effects brought by them can be referred to the method embodiment part, and will not be repeated here.
[0135] Please refer to Figure 3 , Figure 3 A structural schematic diagram of an electronic device provided by the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, the electronic device includes a processor 10, a memory 20 and a communication interface 30.Figure 3 As shown, the electronic device 3 provided by the embodiment can include a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30. For example, a program corresponding to the method for constructing a driving simulation scene. The processor 30 implements the steps in the above-mentioned embodiments of the method for constructing a driving simulation scene when executing the computer program 32, for example Figure 1 As shown in S101-S104. Alternatively, the processor 30 implements the functions of each module / unit in the above-mentioned embodiments of the driving simulation scene construction device when executing the computer program 32, for example Figure 2 The functions of the units 21-24 as shown.
[0136] For example, the computer program 32 can be divided into one or more modules / units, one or more modules / units are stored in the memory 31 and executed by the processor 30 to complete the present application. One or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which is used to describe the execution process of the computer program 32 in the electronic device 3. For example, the computer program 32 can be divided into an initial scene acquisition unit 21, a simulation pedestrian simulation unit 22, a simulation animal simulation unit 23, and a target scene construction unit 24. The specific functions of each unit are described in the related description in the corresponding embodiments, which will not be described here. Figure 2
[0137] Those skilled in the art can understand that Figure 3 is only an example of the electronic device 3 and does not constitute a limitation on the electronic device 3, which can include more or fewer components than shown, or combine certain components, or different components.
[0138] The processor 30 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0139] The memory 31 can be an internal storage unit of the electronic device 3, such as a hard disk or a memory of the electronic device 3. The memory 31 can also be an external storage device of the electronic device 3, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, or the like equipped on the electronic device 3. Further, the memory 31 can also include both the internal storage unit and the external storage device of the electronic device 3. The memory 31 is used to store computer programs and other programs and data required by the electronic device. The memory 31 can also be used to temporarily store data that has been output or will be output.
[0140] It can be clearly understood by those skilled in the art that, for the convenience and brevity of description, only the division of the above functional units is taken as an example, and in actual application, the above functions can be completed by different functional units according to needs, that is, the internal structure of the driving simulation scene construction device is divided into different functional units to complete all or part of the above described functions. Each functional unit in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or in the form of software functional unit. In addition, the specific name of each functional unit is only for convenient distinction, and does not limit the protection scope of the present application. The specific working process of the unit in the above system can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
[0141] The embodiment of the present application further provides a computer readable storage medium, the computer readable storage medium stores a computer program, and the computer program is executed by a processor to realize the steps in each of the method embodiments.
[0142] The embodiment of the present application provides a computer program product, when the computer program product is run on a terminal device, the terminal device realizes the steps in each of the method embodiments.
[0143] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.
[0144] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software manner depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0145] The above-described embodiments are only used to illustrate the technical solutions of the present application, but not limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalent replacements; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A method of constructing a driving simulation scenario, characterized by, The method comprises the following steps: acquiring a pre-constructed initial driving simulation scene, the initial driving simulation scene comprising at least a traffic simulation environment, initial simulation pedestrians and initial simulation animals; simulating behaviors of the initial simulation pedestrians based on the traffic simulation environment by using a pedestrian behavior simulation model to obtain target simulation pedestrians; the pedestrian behavior simulation model is trained by using a reinforcement learning algorithm and a shadow agent method; the pedestrian behavior simulation model comprises a target pedestrian strategy model and a shadow agent model, and the shadow agent model is trained by using a preset pedestrian walking strategy; simulating behaviors of the initial simulation animals based on the traffic simulation environment by using an animal behavior simulation model to obtain target simulation animals; the animal behavior simulation model is trained by using only a reinforcement learning algorithm; constructing a target driving simulation scene according to the traffic simulation environment, the target simulation pedestrians and the target simulation animals.
2. The method of claim 1, wherein, The pedestrian behavior simulation model is trained in the following manner: acquiring a pre-trained initial pedestrian strategy model and the shadow agent model; simulating behaviors of the same training pedestrians by using the initial pedestrian strategy model and the shadow agent model respectively according to the same training environment to obtain training pedestrian behaviors corresponding to the initial pedestrian strategy model and the shadow agent model; determining penalty reward information by comparing the training pedestrian behaviors corresponding to the initial pedestrian strategy model and the shadow agent model; feeding back the penalty reward information to the initial pedestrian strategy model to instruct the initial pedestrian strategy model to be optimized according to the penalty reward information to obtain the target pedestrian strategy model.
3. The method of claim 2, wherein, The pedestrian behavior simulation model further comprises an evaluation model trained; the pedestrian behavior simulation model simulates behaviors of the initial simulation pedestrians based on the traffic simulation environment to obtain target simulation pedestrians, and the method comprises the following steps: simulating behaviors of the initial simulation pedestrians based on the traffic simulation environment by using the target pedestrian strategy model and the shadow agent model respectively to obtain first simulation pedestrians output by the target pedestrian strategy model and second simulation pedestrians output by the shadow agent model; determining the target simulation pedestrians from the first simulation pedestrians and the second simulation pedestrians by using the evaluation model.
4. The method of claim 3, wherein, After the pedestrian behavior simulation model simulates behaviors of the initial simulation pedestrians based on the traffic simulation environment to obtain target simulation pedestrians, the method further comprises the following steps: if the first simulation pedestrians are determined as the target simulation pedestrians, outputting target reward information to the target pedestrian strategy model to instruct the target pedestrian strategy model to be optimized according to the target reward information; if the second simulation pedestrians are determined as the target simulation pedestrians, outputting target penalty information to the target pedestrian strategy model to instruct the target pedestrian strategy model to be optimized according to the target penalty information.
5. The method of claim 1, wherein, The animal behavior simulation model is trained in the following manner: acquiring a pre-trained initial animal strategy model; The initial animal strategy model is used to simulate behaviors of the training animals according to the training environment, to obtain training animal behaviors; According to a preset punishment reward function and the training animal behaviors, punishment reward information of the training animal behaviors is determined; the punishment reward function is used to reward behaviors of the training animal behaviors that are close to food sources and far away from vehicles, and is used to punish behaviors of the training animal behaviors that cause collisions and / or excessive energy consumption; The punishment reward information is fed back to the initial animal strategy model, so that the initial animal strategy model is optimized according to the punishment reward information, to obtain the animal behavior simulation model.
6. The method of claim 1, wherein, The initial driving simulation scene is constructed by the following method: Obtaining initial driving simulation scene data; the initial driving simulation scene data includes initial simulation pedestrian data and initial simulation animal data, and includes one or more of the following traffic simulation environment data: map data, road topology data, building data, obstacle data, traffic sign data, dynamic traffic signal data, dynamic weather data, dynamic vehicle data, and virtual dynamic sensor data; According to the initial simulation pedestrian data, the initial simulation animal data, and the traffic simulation environment data, a dynamic initial driving simulation scene is constructed.
7. The method according to any one of claims 1 to 6, characterized in that, Further comprising: Obtaining training environment data corresponding to each training environment; Through the trained strategy generalization model, the pedestrian behavior simulation model and the animal behavior simulation model are subjected to strategy generalization processing based on the training environment data, and the pedestrian behavior simulation model after strategy generalization processing and the animal behavior simulation model after strategy generalization processing are output; the action difference degree between each of the target simulation pedestrians obtained according to the changed traffic simulation environment is lower than a first preset threshold value, and the action difference degree between each of the target simulation animals obtained according to the changed traffic simulation environment is lower than a second preset threshold value.
8. A construction device of a driving simulation scenario, characterized by, Comprising: An initial scene acquisition unit is configured to acquire a pre-constructed initial driving simulation scene, wherein the initial driving simulation scene at least includes a traffic simulation environment, initial simulation pedestrians, and initial simulation animals; A simulation pedestrian simulation unit is configured to simulate behaviors of the initial simulation pedestrians based on the traffic simulation environment by using a pedestrian behavior simulation model, to obtain target simulation pedestrians; the pedestrian behavior simulation model is trained by using a reinforcement learning algorithm and a shadow agent method; The pedestrian behavior simulation model includes a target pedestrian strategy model and a shadow agent model, and the shadow agent model is trained by using a preset pedestrian walking strategy; A simulation animal simulation unit is configured to simulate behaviors of the initial simulation animals based on the traffic simulation environment by using an animal behavior simulation model, to obtain target simulation animals; the animal behavior simulation model is trained by using only a reinforcement learning algorithm; A target scene constructing unit is configured to construct a target driving simulation scene according to the traffic simulation environment, the target simulation pedestrian and the target simulation animal.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor implements each step in the method for constructing a driving simulation scene according to any one of claims 1 to 7 when executing the computer program.
10. A computer program product, characterised in that, The computer program product implements each step in the method for constructing a driving simulation scene according to any one of claims 1 to 7 when executed by the processor.
Citation Information
Patent Citations
Traffic scene generation method and device and medium
CN118708475A
Methods for training a behavioral model
DE102023200230A1