An automatic driving scene generation method, device and vehicle

By employing structured extraction, world model-based visual information generation, long-tail safety assessment, risk attribution, and multi-objective scoring, this approach addresses the issue of insufficient long-tail risk scenario generation in autonomous driving simulation testing. It achieves efficient and automated simulation scenario generation, thereby improving the safety and coverage of testing.

CN122113659APending Publication Date: 2026-05-29HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU DIANZI UNIV
Filing Date
2026-04-01
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing methods for generating scenarios in autonomous driving simulation testing lack the ability to cover low-frequency, high-risk long-tail safety scenarios, resulting in low test coverage, high labor costs, and an inability to meet the high reliability testing requirements of autonomous driving systems.

Method used

By extracting structured prompts from the input scene, generating visual information of the scene using a world model, and performing preliminary judgment of long-tail safety scenes, risk attribution analysis and counterfactual reasoning, candidate scenes are generated, and multi-objective comprehensive scoring and iterative re-judgment are performed to output simulation scene code.

Benefits of technology

It enables automated, high-quality, closed-loop iterative generation of non-safety simulation scenarios, improving the safety and comprehensiveness of autonomous driving testing, increasing the coverage and testing efficiency of long-tail high-risk scenarios, and reducing labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113659A_ABST
    Figure CN122113659A_ABST
Patent Text Reader

Abstract

The application discloses an automatic driving scene generation method and device and a vehicle, and relates to the technical field of automatic driving scene generation, and in particular to a method for generating a scene by using a world model, a long-tail safe scene, risk attribution analysis, counterfactual reasoning, and multi-objective comprehensive scoring and iterative re-determination screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of autonomous driving simulation testing technology, and relates to an autonomous driving scenario generation method, device and vehicle. Background Technology

[0002] With the rapid development of autonomous driving technology, simulation testing has become a key means of verifying the safety and reliability of autonomous driving systems. Current methods for constructing autonomous driving scenarios largely rely on manual design, rule enumeration, or simple data collection, making it difficult to efficiently cover low-frequency, high-risk, long-tail safety scenarios. Furthermore, existing scenario generation methods lack the ability to attribute risks to safety scenarios and iterate counterfactually, resulting in insufficient scenario diversity and realism, low test coverage, and high labor costs, failing to meet the high reliability testing requirements of autonomous driving systems.

[0003] Therefore, it is necessary to provide improved technical solutions to overcome the above-mentioned technical problems existing in the prior art. Summary of the Invention

[0004] The purpose of this application is to provide an autonomous driving scenario generation method, device, and vehicle to solve the problems of insufficient coverage of autonomous driving simulation scenarios, difficulty in generating long-tail risk scenarios, and low testing efficiency in the prior art, so as to realize the automated, high-quality, closed-loop iterative generation of non-safety simulation scenarios and improve the safety and comprehensiveness of autonomous driving testing.

[0005] To achieve the above objectives, the technical solution adopted in this application is as follows: Firstly, this application provides a method for generating autonomous driving scenarios, the method comprising: The input scene is structurally extracted to generate prompt words, and the scene visual information is generated based on the prompt words using a world model; The visual information of the scene is used to make a preliminary judgment on the long-tail safety scene to obtain an initial safety scene, and the initial safety scene is subjected to risk attribution analysis and counterfactual reasoning to generate scene rewriting conditions. Candidate scenarios are generated based on the scenario rewriting conditions, and the candidate scenarios are subjected to multi-objective comprehensive scoring and iterative re-judgment to generate unsafe simulation scenarios, and the simulation scenario code is output.

[0006] In one embodiment, the step of structurally extracting the input scene to generate prompt words includes: Obtain the original scene description text of the autonomous driving test scene to be generated from at least one of the user's natural language input and the historical simulation scene library; The original scene description is structurally extracted to obtain a set of structured scene elements including autonomous vehicles, background participants, behaviors, positional relationships, and road environment. By combining the structured scene element set with physical rule constraints and driving common sense constraints, prompt words that can be used to drive the world model are generated.

[0007] In one embodiment, generating scene visual information based on the cue words using a world model includes: The prompt words are input into the world model to generate continuous motion trajectories, interactive behaviors, and scene visual information of autonomous vehicles and background participants.

[0008] In one embodiment, the step of performing a preliminary long-tail security scene assessment on the scene visual information to obtain an initial security scene includes: The visual information of the scene is input into the visual language model, and the temporal behavior and interaction relationship are inferred through the visual language model to complete the initial judgment of the long-tail security scene. If the scenario is determined to be unsafe, simulation scenario code is generated using a large language model. If the scenario is determined to be a safe scenario, then the safe scenario is determined to be the initial safe scenario.

[0009] In one embodiment, the step of performing risk attribution analysis and counterfactual reasoning on the initial security scenario to generate scenario rewriting conditions includes: The visual language model is used to analyze the key risk factors, key time intervals, key spatial areas, key participants, and risk causal links of the initial security scenario, thereby completing the security attribution analysis. Based on the results of the security attribution analysis, a set of counterfactual suggestions, including behavior modification, state adjustment, and environmental parameter changes, is generated as the conditions for scenario rewriting.

[0010] In one embodiment, generating candidate scenes based on scene rewriting conditions includes: The set of structured scene elements and the set of counterfactual suggestions are input into a large language model, and a set of candidate scene examples is generated after multiple samplings. Based on the candidate scene example set, prompt words are generated, and candidate scenes are generated based on the prompt words using the world model.

[0011] In one embodiment, the step of performing multi-objective comprehensive scoring and iterative review and screening on candidate scenarios to generate unsafe simulation scenarios includes: The candidate scenes are evaluated and ranked using a visual language model, which includes physical authenticity scoring, safety criticality scoring, and long-tail scoring, to obtain the optimal candidate scene. The optimal candidate scene is input into the visual language model for long-tail safety scene re-judgment. If the scenario is determined to be unsafe, the scenario code is output; if the scenario is determined to be safe, the process returns to the safety attribution analysis step to iterate and generate candidate scenarios again.

[0012] In one embodiment, the weights of the multi-objective comprehensive score can be dynamically adjusted according to the type of autonomous driving test task, which includes extreme weather tests, urban congestion tests, and highway emergency scenario tests.

[0013] Secondly, this application provides an autonomous driving scene generation device, the device comprising: The scene visual information generation module is used to extract the input scene in a structured manner to generate prompt words, and generate scene visual information based on the prompt words through a world model; The scene rewriting condition generation module is used to perform a long-tail safety scene preliminary judgment on the scene visual information to obtain an initial safety scene, and to perform risk attribution analysis and counterfactual reasoning on the initial safety scene to generate scene rewriting conditions. The unsafe simulation scenario generation module is used to generate candidate scenarios based on the scenario rewriting conditions, perform multi-objective comprehensive scoring and iterative review and screening on the candidate scenarios to generate unsafe simulation scenarios, and output simulation scenario code.

[0014] Thirdly, this application provides a vehicle including the autonomous driving scene generation device described above.

[0015] This application discloses an autonomous driving scene generation method, apparatus, and vehicle. The method includes: performing structured extraction on an input scene to generate prompt words, and generating scene visual information based on the prompt words using a world model; performing a preliminary long-tail safety scene judgment on the scene visual information to obtain an initial safe scene, and performing risk attribution analysis and counterfactual reasoning on the initial safe scene to generate scene rewriting conditions; generating candidate scenes based on the scene rewriting conditions, and performing multi-objective comprehensive scoring and iterative re-judgment screening on the candidate scenes to generate unsafe simulation scenes, and outputting simulation scene code. This invention can efficiently discover low-frequency, high-risk long-tail scenes, improve the systematicness, scene coverage, and simulation credibility of autonomous driving safety testing, reduce the cost of manual scene construction, and ensure the safety and reliability of autonomous driving system testing. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating the autonomous driving scene generation method provided in an embodiment of the present invention.

[0018] Figure 2 This is a schematic diagram of the structure of the autonomous driving scene generation device provided in an embodiment of the present invention. Detailed Implementation

[0019] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0020] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.

[0021] It should be understood that although the terms first, second, third, etc., may be used herein to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this document, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if," as used herein, can be interpreted as "when," "when," or "in response to determination." Furthermore, as used herein, the singular forms "a," "an," and "the" are intended to also include the plural forms unless the context indicates otherwise. It should be further understood that the terms "comprising," "including," indicate the presence of the stated feature, step, operation, element, component, item, kind, and / or group, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, kinds, and / or groups. The terms "or" and "and / or" as used herein are to be interpreted as inclusive, or mean any one or any combination thereof. Therefore, "A, B, or C" or "A, B, and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A, B, and C". Exceptions to this definition will only occur if the combination of elements, functions, steps, or operations is inherently mutually exclusive in some way.

[0022] It should be understood that although the steps in the flowcharts of this application's embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0023] It should be noted that step designations such as S110 and S120 are used in this document for the purpose of more clearly and concisely describing the corresponding content, and do not constitute a substantial limitation on the order. In specific implementation, those skilled in the art may execute S120 first and then S110, etc., but these should all be within the protection scope of this application.

[0024] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0025] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.

[0026] First Embodiment See Figure 1 This application provides an autonomous driving scene generation method, which can be executed by an autonomous driving scene generation device provided in this application. The device can be implemented in software and / or hardware. In this embodiment, the autonomous driving scene generation device is taken as the executing entity of the method. The autonomous driving scene generation method provided in this embodiment includes the following steps: Step S110: Perform structured extraction of the input scene to generate prompt words, and generate scene visual information based on the prompt words using a world model.

[0027] In one implementation, the input scene is structurally extracted to generate prompt words, including: Obtain the original scene description text of the autonomous driving test scenario to be generated from at least one of the user's natural language input and the historical simulation scene library; perform structured extraction on the original scene description to obtain a set of structured scene elements including autonomous vehicles, background participants, behaviors, positional relationships, and road environment; combine the set of structured scene elements with physical rule constraints and driving common sense constraints to generate prompt words that can be used to drive the world model.

[0028] It is understandable that the original scene description is structurally extracted and parsed to separate key elements such as autonomous vehicles, background vehicles, pedestrians and other traffic participants, dynamic behaviors, positional relationships and road environment, forming a standardized set of structured scene elements. The above-mentioned set of structured scene elements is then processed using a large language model, which introduces physical motion rules, traffic regulations and driving common sense constraints to generate prompt words that meet semantic consistency and visual adaptability and can be used to drive the world model.

[0029] In one implementation, generating scene visual information based on cue words using a world model includes: Input the prompt words into the world model to generate continuous motion trajectories, interactive behaviors, and complete scene visual information, including autonomous vehicles and background participants.

[0030] Step S120: Perform a long-tail security scene preliminary judgment on the scene visual information to obtain an initial security scene, and perform risk attribution analysis and counterfactual reasoning on the initial security scene to generate scene rewriting conditions.

[0031] In one embodiment, a preliminary long-tail safety scene assessment is performed on the scene visual information to obtain an initial safety scene, including: The visual information of the scene is input into the visual language model, and the temporal behavior and interaction relationship are inferred through the visual language model to complete the initial judgment of the long-tail safe scene; if it is determined to be an unsafe scene, the simulation scene code is generated through the large language model; if it is determined to be a safe scene, the safe scene is determined as the initial safe scene.

[0032] It is understandable that the visual information of the scene is input into the visual language model, which then understands and infers the temporal behavior and interaction relationships of the scene to complete the initial judgment of the long-tail safety scene and output the judgment result. If the scene is judged to be unsafe, the reproducible simulation scene code is directly generated through the large language model. If the scene is judged to be safe, the scene is used as the initial safe scene and enters the subsequent safety attribution and counterfactual reasoning process.

[0033] In one implementation, risk attribution analysis and counterfactual reasoning are performed on the initial security scenario to generate scenario rewriting conditions, including: By analyzing the key risk factors, key time intervals, key spatial areas, key participants, and risk causal links in the initial security scenario using a visual language model, a security attribution analysis is completed. Based on the results of the security attribution analysis, a set of counterfactual suggestions, including behavior modification, state adjustment, and environmental parameter changes, is generated as conditions for scenario rewriting.

[0034] It is understandable that by inputting the initial security scenario into the visual language model, and analyzing key risk factors, key time intervals, key spatial areas, key participants, and corresponding risk causal links, a security attribution analysis is completed to clarify the core conditions for maintaining the security of the scenario. Based on the results of the security attribution analysis, a set of counterfactual suggestions is generated. The counterfactual suggestions include actionable instructions such as behavior modification, state adjustment, and environmental parameter change, which serve as conditions for scenario rewriting and are used to transform the security scenario into a long-tail risk scenario with non-security criticality.

[0035] Step S130: Generate candidate scenarios based on the scenario rewriting conditions, and perform multi-objective comprehensive scoring and iterative re-judgment screening on the candidate scenarios to generate unsafe simulation scenarios, and output simulation scenario code.

[0036] It is understandable that the set of structured scene elements and the set of counterfactual suggestions are input into a large language model, and multiple candidate scene example sets are generated through multiple sampling.

[0037] In one embodiment, candidate scenarios are subjected to multi-objective comprehensive scoring and iterative review and screening to generate unsafe simulation scenarios, including: The candidate scenes are evaluated and ranked using a visual language model, which includes physical authenticity scoring, safety criticality scoring, and long-tail scoring, to obtain the optimal candidate scene. The optimal candidate scene is then input into the visual language model for long-tail safety scene re-judgment. If it is determined to be an unsafe scene, the scene code is output. If it is determined to be a safe scene, the process returns to the safety attribution analysis step to iterate and generate candidate scenes again.

[0038] Understandably, for each candidate scene example, a large language model generates corresponding prompt words, which drive the world model to generate visual information for the candidate scene, resulting in a set of candidate scenes. The visual language model performs multi-objective comprehensive scoring on the candidate scenes, with scoring dimensions including physical realism score, safety criticality score, and long-tail score. The scenes are then sorted according to their comprehensive scores to select the optimal candidate scene. The optimal candidate scene is then input into the visual language model again for long-tail safety re-judgment. If the re-judgment determines it to be an unsafe scene, the large language model outputs reproducible simulation scene code. If the re-judgment still determines it to be a safe scene, the process returns to the safety attribution analysis step, and attribution and candidate scene generation are performed again until the iteration converges to obtain a qualified unsafe simulation scene.

[0039] In one implementation, the weights of the multi-objective comprehensive score can be dynamically adjusted according to the type of autonomous driving test task, which includes extreme weather tests, urban congestion tests, and highway emergency scenario tests.

[0040] In summary, the implementation method of this application achieves closed-loop automated generation of unsafe simulation scenarios through structured extraction, world model generation of visual information, long-tail safety judgment, risk attribution, counterfactual reasoning, multi-objective scoring and iterative re-judgment. This can effectively improve the coverage of long-tail high-risk scenarios, enhance the realism and efficiency of testing, reduce labor costs, and enhance the safety and reliability of autonomous driving system testing.

[0041] Second Embodiment Based on the first embodiment of this application, an autonomous driving scene generation device is provided in this embodiment, see below. Figure 2 The device includes: The scene visual information generation module 21 is used to extract structured information from the input scene to generate prompt words, and to generate scene visual information based on the prompt words through the world model.

[0042] The scene rewriting condition generation module 22 is used to perform a preliminary judgment of the long-tail security scene on the scene visual information to obtain the initial security scene, and to perform risk attribution analysis and counterfactual reasoning on the initial security scene to generate scene rewriting conditions.

[0043] The unsafe simulation scenario generation module 23 is used to generate candidate scenarios based on scenario rewriting conditions, and to perform multi-objective comprehensive scoring and iterative review and screening on the candidate scenarios to generate unsafe simulation scenarios, and output simulation scenario code.

[0044] In one embodiment, the scene visual information generation module 21 is further configured to: Obtain the original scene description text of the autonomous driving test scenario to be generated from at least one of the user's natural language input and the historical simulation scene library; perform structured extraction on the original scene description to obtain a set of structured scene elements including autonomous vehicles, background participants, behaviors, positional relationships, and road environment; combine the set of structured scene elements with physical rule constraints and driving common sense constraints to generate prompt words that can be used to drive the world model.

[0045] In one embodiment, the scene visual information generation module 21 is further configured to: Input the prompt words into the world model to generate continuous motion trajectories, interactive behaviors, and scene visual information, including autonomous vehicles and background participants.

[0046] In one embodiment, the scene rewriting condition generation module 22 is further configured to: The visual information of the scene is input into the visual language model, and the temporal behavior and interaction relationship are inferred through the visual language model to complete the initial judgment of the long-tail safe scene; if it is determined to be an unsafe scene, the simulation scene code is generated through the large language model; if it is determined to be a safe scene, the safe scene is determined as the initial safe scene.

[0047] In one embodiment, the scene rewriting condition generation module 22 is further configured to: By analyzing the key risk factors, key time intervals, key spatial areas, key participants, and risk causal links in the initial security scenario using a visual language model, a security attribution analysis is completed. Based on the results of the security attribution analysis, a set of counterfactual suggestions, including behavior modification, state adjustment, and environmental parameter changes, is generated as conditions for scenario rewriting.

[0048] In one embodiment, the non-safety simulation scene generation module 23 is further configured to: The candidate scenes are evaluated and ranked using a visual language model, which includes physical authenticity scoring, safety criticality scoring, and long-tail scoring, to obtain the optimal candidate scene. The optimal candidate scene is then input into the visual language model for long-tail safety scene re-judgment. If it is determined to be an unsafe scene, the scene code is output. If it is determined to be a safe scene, the process returns to the safety attribution analysis step to iterate and generate candidate scenes again.

[0049] In one embodiment, the non-safety simulation scene generation module 23 is further configured to: The weights of the multi-objective comprehensive score can be dynamically adjusted according to the type of autonomous driving test task, which includes extreme weather tests, urban congestion tests, and highway emergency scenario tests.

[0050] It should be noted that the description of the above-mentioned autonomous driving scene generation device is similar to the description of the above-mentioned autonomous driving scene generation method, and the beneficial effects of the same method will not be repeated. For technical details not disclosed in the embodiments of the autonomous driving scene generation device of this invention, please refer to the description of the embodiments of the autonomous driving scene generation method of this invention.

[0051] Third Embodiment This embodiment provides a vehicle equipped with an autonomous driving scenario generation device as described in the second embodiment, which can be used for simulation testing, scenario library construction, and safety verification of autonomous driving systems, thereby improving the reliability and safety of the vehicle's autonomous driving function.

[0052] In this application, the same or similar terms, concepts, technical solutions and / or application scenario descriptions are generally described in detail only when they appear for the first time. When they appear again, they are generally not repeated for the sake of brevity. When understanding the technical solutions and other contents of this application, the same or similar terms, concepts, technical solutions and / or application scenario descriptions that are not described in detail later can be referred to their previous relevant detailed descriptions.

[0053] In this application, the descriptions of the various embodiments have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0054] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0055] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. For those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for generating autonomous driving scenarios, characterized in that, The method includes: The input scene is structurally extracted to generate prompt words, and the scene visual information is generated based on the prompt words using a world model; The visual information of the scene is used to make a preliminary judgment on the long-tail safety scene to obtain an initial safety scene, and the initial safety scene is subjected to risk attribution analysis and counterfactual reasoning to generate scene rewriting conditions. Candidate scenarios are generated based on the scenario rewriting conditions, and the candidate scenarios are subjected to multi-objective comprehensive scoring and iterative re-judgment to generate unsafe simulation scenarios, and the simulation scenario code is output.

2. The method according to claim 1, characterized in that, The step of extracting structured input scenarios to generate prompt words includes: Obtain the original scene description text of the autonomous driving test scene to be generated from at least one of the user's natural language input and the historical simulation scene library; The original scene description is structurally extracted to obtain a set of structured scene elements including autonomous vehicles, background participants, behaviors, positional relationships, and road environment. By combining the structured scene element set with physical rule constraints and driving common sense constraints, prompt words that can be used to drive the world model are generated.

3. The method according to claim 2, characterized in that, The process of generating scene visual information based on the prompt words using a world model includes: The prompt words are input into the world model to generate continuous motion trajectories, interactive behaviors, and scene visual information of autonomous vehicles and background participants.

4. The method according to claim 3, characterized in that, The step of performing a long-tail safety scene preliminary judgment on the scene visual information to obtain an initial safety scene includes: The visual information of the scene is input into the visual language model, and the temporal behavior and interaction relationship are inferred through the visual language model to complete the initial judgment of the long-tail security scene. If the scenario is determined to be unsafe, simulation scenario code is generated using a large language model; if the scenario is determined to be safe, the safe scenario is identified as the initial safe scenario.

5. The method according to claim 4, characterized in that, The step of performing risk attribution analysis and counterfactual reasoning on the initial security scenario to generate scenario rewriting conditions includes: The visual language model is used to analyze the key risk factors, key time intervals, key spatial areas, key participants, and risk causal links of the initial security scenario, thereby completing the security attribution analysis. Based on the results of the security attribution analysis, a set of counterfactual suggestions, including behavior modification, state adjustment, and environmental parameter changes, is generated as the conditions for scenario rewriting.

6. The method according to claim 5, characterized in that, The generation of candidate scenarios based on the scenario rewriting conditions includes: The structured scene element set and the counterfactual suggestion set are input into a large language model, and a candidate scene example set is generated after multiple samplings. Based on the candidate scene example set, prompt words are generated, and candidate scenes are generated based on the prompt words using the world model.

7. The method according to claim 6, characterized in that, The process of generating unsafe simulation scenarios by performing multi-objective comprehensive scoring and iterative review and screening on the candidate scenarios includes: The candidate scenes are evaluated and ranked using a visual language model, which includes physical authenticity scoring, safety criticality scoring, and long-tail scoring, to obtain the optimal candidate scene. The optimal candidate scene is input into the visual language model for long-tail safety scene re-judgment. If the scenario is determined to be unsafe, the scenario code is output; if the scenario is determined to be safe, the process returns to the safety attribution analysis step to iterate and generate candidate scenarios again.

8. The method according to claim 7, characterized in that, The weights of the multi-objective comprehensive score can be dynamically adjusted according to the type of autonomous driving test task, which includes extreme weather test, urban congestion test, and highway emergency scenario test.

9. An autonomous driving scene generation device, characterized in that, The device includes: The scene visual information generation module is used to extract the input scene in a structured manner to generate prompt words, and generate scene visual information based on the prompt words through a world model; The scene rewriting condition generation module is used to perform a long-tail safety scene preliminary judgment on the scene visual information to obtain an initial safety scene, and to perform risk attribution analysis and counterfactual reasoning on the initial safety scene to generate scene rewriting conditions. The unsafe simulation scenario generation module is used to generate candidate scenarios based on the scenario rewriting conditions, perform multi-objective comprehensive scoring and iterative review and screening on the candidate scenarios to generate unsafe simulation scenarios, and output simulation scenario code.

10. A vehicle, characterized in that, Includes the autonomous driving scene generation device as described in claim 9.