Method, computer device, and computer program for traffic scenario generation
A multi-guided diffusion model with DPO algorithm and multitask learning addresses the challenge of balancing realism, diversity, and controllability in traffic scenario generation for autonomous vehicles, enhancing scenario quality and adherence to traffic rules.
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- NAVER LABS CORP
- Filing Date
- 2025-01-14
- Publication Date
- 2026-07-21
AI Technical Summary
Existing traffic simulation methods struggle to balance realism, diversity, and controllability in generating scenarios for autonomous vehicle safety testing, with rule-based approaches lacking realism and learning-based approaches lacking controllability, and guided sampling often deviating from data distribution.
A multi-guided diffusion model using a direct preference optimization (DPO) algorithm and multitask learning strategy for fine-tuning a traffic scenario generation model, combining a diffusion model with guided sampling to enhance realism, diversity, and controllability by optimizing guide scores and adhering to traffic priorities.
Generates traffic scenarios that simultaneously satisfy diversity, controllability, and realism, ensuring adherence to traffic rules and improving the precision of guide sampling through DPO algorithm fine-tuning.
Smart Images

Figure P1020250005318_ABST
Abstract
Description
Technology Field
[0001] The following description concerns a technology for generating traffic scenarios to verify the safety of autonomous vehicles. Background Technology
[0002] Recently, as the automotive industry and various sensing technologies have advanced and communication infrastructure has expanded, interest and research on vehicles that autonomously travel to their destinations without driver intervention have been actively underway.
[0003] An autonomous vehicle refers to a vehicle capable of driving itself to a set destination without driver intervention by utilizing various cameras, sensors, and high-precision maps to perceive, judge, and act upon surrounding vehicles, pedestrians, and objects, thereby controlling the vehicle.
[0004] Since the failure to ensure safety of such autonomous vehicles can lead to major casualties, evaluation and testing regarding how quickly and accurately they recognize and respond to various unexpected situations and events are essential.
[0005] The creation of realistic, diverse, and controllable traffic scenarios is crucial for testing autonomous vehicles and ensuring their safety.
[0006] While generating realistic and diverse traffic scenarios is crucial for verifying the safety of autonomous vehicles, collecting data through actual driving entails costs and risks.
[0007] As an example of technology for autonomous driving testing, Korean registered patent No. 10-2579590 (registered on September 13, 2023) discloses a technology for generating an evaluation scenario for autonomous vehicle driving ability based on the Road Traffic Act. The problem to be solved
[0008] A traffic scenario generation model combining a diffusion model and guided sampling can be fine-tuned using the direct preference optimization (DPO) algorithm and a multi-task learning strategy. means of solving the problem
[0009] A method for generating a traffic scenario for a computer device comprising at least one processor, wherein the method comprises the step of learning a traffic scenario generation model for generating a traffic scenario for an autonomous vehicle through a direct preference optimization (DPO) algorithm by the at least one processor.
[0010] According to one aspect, the traffic scenario generation model may be a diffusion model combined with guided sampling.
[0011] According to another aspect, the learning step may include a step of fine-tuning a target model pre-trained as a traffic scenario generation model through direct preference optimization (DPO) including guided sampling.
[0012] According to another aspect, the fine-tuning step may include the step of fine-tuning a segment of the transformer according to a guide selected through a guidance conditional layer.
[0013] According to another aspect, the fine-tuning step may include: compiling a DPO dataset containing instances designated as win samples and lose samples using a prompt encompassing traffic information and Gaussian noise; and fine-tuning the target model using a DPO loss based on the DPO dataset.
[0014] According to another aspect, the fine-tuning step may include: generating a plurality of traffic scenarios through guide sampling using a reference model together with the target model; distinguishing winning samples and losing samples based on guidance scores for the plurality of traffic scenarios; and fine-tuning the target model using the winning samples and losing samples according to the guidance scores.
[0015] According to another aspect, the step of fine-tuning the target model may include the step of training the target model so that the winning samples have a higher score than the reference model’s winning samples and the target model’s losing samples have a lower score than the reference model’s losing samples.
[0016] According to another aspect, the learning step may include a step of training a backbone model with scene-level traffic information using a diffusion transformer and CFS (classifier-free sampling).
[0017] According to another aspect, the backbone model may be structured into an encoder that tokenizes input data, a transformer block that processes the tokenized input data together with a scene context, and a transformer decoder that reconstructs the data processed by the transformer block.
[0018] According to another aspect, the step of training the backbone model can train the backbone model to predict a clean trajectory free of noise from a noisy trajectory through a noise removal process that functions inversely to a forward diffusion process.
[0019] A computer program stored on a computer-readable recording medium is provided to execute the above traffic scenario generation method on the computer device.
[0020] The computer device comprises at least one processor implemented to execute a readable instruction on a computer device, wherein the at least one processor processes the process of learning a traffic scenario generation model that generates traffic scenarios for an autonomous vehicle through a direct preference optimization (DPO) algorithm. Effects of the invention
[0021] According to embodiments of the present invention, by fine-tuning a traffic scenario generation model that combines a diffusion model and guided sampling using a DPO algorithm and a multitask learning strategy, it is possible to generate a traffic scenario that simultaneously satisfies diversity, controllability, and realism. Brief explanation of the drawing
[0022] FIG. 1 is a drawing illustrating an example of a network environment according to an embodiment of the present invention. FIG. 2 is a block diagram illustrating an example of a computer device according to an embodiment of the present invention. Figure 3 illustrates an example of a traffic simulation process. Figure 4 illustrates the learning process of a multi-guided diffusion model using DPO in an embodiment of the present invention. FIG. 5 illustrates the inference process of a traffic scenario generation model in one embodiment of the present invention. FIG. 6 illustrates the learning process of a traffic scenario generation model in one embodiment of the present invention. FIGS. 7 and 8 illustrate the fine-tuning process of a traffic scenario generation model in an embodiment of the present invention. FIG. 9 illustrates a model fine-tuning algorithm using DPO in an embodiment of the present invention. Specific details for implementing the invention
[0023] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings.
[0025] Embodiments of the present invention relate to a technology for generating traffic scenarios for verifying the safety of an autonomous vehicle.
[0026] Embodiments including those specifically disclosed in this specification can generate traffic scenarios that simultaneously satisfy diversity, controllability, and realism through a multi-guided diffusion model using direct preference optimization (DPO).
[0027] A traffic scenario generation device according to embodiments of the present invention may be implemented by at least one computer device, and a traffic scenario generation method according to embodiments of the present invention may be performed through at least one computer device included in the traffic scenario generation device. At this time, a computer program according to an embodiment of the present invention may be installed and run on the computer device, and the computer device may perform a traffic scenario generation method according to embodiments of the present invention under the control of the run computer program. The above-described computer program may be stored on a computer-readable recording medium to be combined with the computer device to execute the traffic scenario generation method on the computer.
[0028] FIG. 1 is a diagram illustrating an example of a network environment according to an embodiment of the present invention. The network environment of FIG. 1 illustrates an example including a plurality of electronic devices (110, 120, 130, 140), a plurality of servers (150, 160), and a network (170). FIG. 1 is an example for explaining the invention, and the number of electronic devices or servers is not limited to that shown in FIG. 1. Furthermore, the network environment of FIG. 1 is merely an example of one of the environments applicable to the present embodiments, and the environments applicable to the present embodiments are not limited to the network environment of FIG. 1.
[0029] Multiple electronic devices (110, 120, 130, 140) may be fixed terminals or mobile terminals implemented as computer devices. Examples of multiple electronic devices (110, 120, 130, 140) include smartphones, mobile phones, navigation systems, computers, laptops, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablet PCs, etc. For example, FIG. 1 shows the shape of a smartphone as an example of an electronic device (110), but in embodiments of the present invention, the electronic device (110) may substantially refer to one of various physical computer devices capable of communicating with other electronic devices (120, 130, 140) and / or servers (150, 160) via a network (170) using a wireless or wired communication method.
[0030] The communication method is not limited and may include not only communication methods utilizing communication networks (e.g., mobile communication networks, wired internet, wireless internet, broadcasting networks) that the network (170) may include, but also short-range wireless communication between devices. For example, the network (170) may include any one or more networks such as a PAN (personal area network), LAN (local area network), CAN (campus area network), MAN (metropolitan area network), WAN (wide area network), BBN (broadband network), and the Internet. Additionally, the network (170) may include any one or more network topologies such as a bus network, a star network, a ring network, a mesh network, a star-bus network, a tree or hierarchical network, but is not limited thereto.
[0031] Each of the servers (150, 160) may be implemented as a computer device or multiple computer devices that communicate with multiple electronic devices (110, 120, 130, 140) through a network (170) to provide commands, code, files, content, services, etc. For example, the server (150) may be a system that provides services (e.g., autonomous vehicle safety verification services, etc.) to multiple electronic devices (110, 120, 130, 140) connected through the network (170).
[0032] FIG. 2 is a block diagram illustrating an example of a computer device according to an embodiment of the present invention. Each of the plurality of electronic devices (110, 120, 130, 140) or servers (150, 160) described above can be implemented by the computer device (200) illustrated in FIG. 2.
[0033] As illustrated in FIG. 2, such a computer device (200) may include memory (210), a processor (220), a communication interface (230), and an input / output interface (240). The memory (210) is a computer-readable recording medium and may include a non-perishable mass storage device such as RAM (random access memory), ROM (read only memory), and a disk drive. Here, a non-perishable mass storage device such as a ROM and a disk drive may be included in the computer device (200) as a separate permanent storage device distinct from the memory (210). Additionally, an operating system and at least one program code may be stored in the memory (210). These software components may be loaded into the memory (210) from a computer-readable recording medium separate from the memory (210). This separate computer-readable recording medium may include a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, or a memory card. In another embodiment, software components may be loaded into memory (210) via a communication interface (230) rather than a computer-readable recording medium. For example, software components may be loaded into memory (210) of a computer device (200) based on a computer program installed by files received through a network (170).
[0034] The processor (220) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor (220) via memory (210) or a communication interface (230). For example, the processor (220) may be configured to execute instructions received according to program code stored in a recording device such as memory (210).
[0035] The communication interface (230) may provide a function for the computer device (200) to communicate with other devices (e.g., storage devices described above) through the network (170). For example, requests, commands, data, files, etc. generated by the processor (220) of the computer device (200) according to program code stored in a recording device such as memory (210) may be transmitted to other devices through the network (170) under the control of the communication interface (230). Conversely, signals, commands, data, files, etc. from other devices may be received by the computer device (200) through the communication interface (230) of the computer device (200) via the network (170). Signals, commands, data, etc. received through the communication interface (230) may be transmitted to the processor (220) or memory (210), and files, etc. may be stored in a storage medium (the permanent storage device described above) that the computer device (200) may further include.
[0036] The input / output interface (240) may be a means for interfacing with an input / output device (250). For example, the input device may include a device such as a microphone, keyboard, or mouse, and the output device may include a device such as a display or speaker. As another example, the input / output interface (240) may be a means for interfacing with a device in which the functions for input and output are integrated into one, such as a touchscreen. The input / output device (250) may be composed of a computer device (200) and a single device.
[0037] Additionally, in other embodiments, the computer device (200) may include fewer or more components than the components of FIG. 2. However, it is not necessary to clearly illustrate most of the prior art components. For example, the computer device (200) may be implemented to include at least some of the input / output devices (250) described above, or may include other components such as a transceiver, a database, etc.
[0038] Below, specific embodiments of the technology for generating traffic scenarios will be described.
[0039] The challenge of developing traffic models is to strike a balance between realism, diversity, and controllability in traffic simulations.
[0040] Figure 3 illustrates an example of a traffic simulation process.
[0041] There are two main approaches to traffic simulation methods: rule-based and learning-based approaches.
[0042] Rule-based approaches analyze vehicle movement and control it along a fixed path; while intuitively understandable to users, they lack the expressiveness necessary to accurately replicate actual driver behavior, often resulting in movements that deviate significantly from real-world driving patterns.
[0043] Learning-based approaches train deep generative models, such as variational auto-encoders (VAEs) and diffusion models, using real-world traffic data. In particular, diffusion models demonstrate great potential in traffic simulations as they generate highly realistic scenarios.
[0044] The problem is that rule-based approaches guarantee controllability but lack realism, while learning-based approaches guarantee realism but lack controllability.
[0045] Recently, traffic scenarios can be automatically generated by combining a learning diffusion-based traffic model based on actual driving data with a heuristic-driven approach known as guided sampling.
[0046] While this combination can improve the realism and controllability of traffic scenarios, there is a problem in that samples generated through guided sampling often deviate from the data distribution and exhibit unrealistic behavior.
[0047] In this embodiment, a rule-based approach and a learning-based approach can be combined to generate realistic and controllable traffic scenarios, and furthermore, focus can be placed on scenario diversity to address important aspects of traffic simulation.
[0048] The present invention is a fine-tuning method including guided sampling, which can generate traffic scenarios that balance realism, diversity, and controllability through a multi-guided diffusion model that utilizes a new training strategy to strictly adhere to traffic priorities even when using various guide combinations.
[0049] The multi-guide diffusion model according to the present invention adopts a multi-task learning framework to enable a single diffusion model to process various guide inputs.
[0050] In particular, the present invention can perform fine-tuning using the DPO algorithm to increase guide sampling precision. The DPO algorithm optimizes preferences based on guide scores, thereby effectively addressing the complexities and challenges associated with slope calculations, which are costly and often non-differentiable, during the guide sampling fine-tuning process.
[0051] Figure 4 illustrates the training process of a traffic scenario generation model in one embodiment of the present invention.
[0052] The traffic scenario generation model according to the present invention can undergo a backbone model training process (stage 1: training) and a fine-tuning process (stage 2: fine-tuning) in two stages of training.
[0053] In the initial training process (stage 1), a robust data-driven backbone model can be trained by utilizing a diffusion transformer and CFS (classifier-free sampling) to pre-learn traffic information at the scene level. At this stage, the diffusion transformer can be trained using actual driving data.
[0054] By introducing a guidance conditional layer similar to the task conditional layer in multitask learning to specify guides suitable for sampling, a single model can skillfully manage various guide combinations.
[0055] In the subsequent training process (stage 2), the pre-trained diffusion model can be fine-tuned by optimizing based on preferences for inference samples using guidance scores rather than relying on human preference feedback.
[0056] In the initial training process (stage 1) where the diffusion transformer is trained using actual driving data, both the realism and diversity of the model can be improved, and in the subsequent training process (stage 2), the realism and controllability of the model can be further enhanced.
[0057] Problem Formulation
[0058] The present invention aims to solve the major challenges of traffic simulation by utilizing DPO to generate more realistic, diverse, and controllable traffic scenarios.
[0059] The state at time step t, including 2D location, heading, speed, and yaw rate Let us assume that. Likewise, the action at time step t, called control, has acceleration and yaw acceleration. It is assumed that
[0060] In a traffic scenario at time step t, assuming a target tgt and N neighbor agents, a local semantic map M i map encoder Map features for each agent i extracted Introduces.
[0061] In addition, the past and future trajectory characteristics for each agent i are, respectively , It is defined as, and the time horizon H for j∈{P,F} j are trajectory encoders for j∈{P,F} respectively It is extracted as follows. At this time, the extracted features define the scene context as C=(M,P,F), where for all agents , , am.
[0062] Upcoming H F Future motion trajectory with respect to time As, the future state trajectory It is indicated as.
[0063] The objective of the present invention is to provide new and diverse trajectories T that exhibit both realistic traffic behavior and rule-compliant traffic behavior. s It is to generate.
[0064] To this end, based on a unicycle dynamics model for simple vehicle dynamics, state s t and operation a t When an action is applied at s t+1 =f(s t ,a tApply a transition function f that calculates ).
[0065] Using f, T through rollout from the initial state s0 a Predict and T s By obtaining this, the physical feasibility of the state trajectory occurring in the denoising process can be guaranteed.
[0066] Diffusion model training for traffic scenario generation
[0067] The training process of the diffusion model includes a forward diffusion process and a noise removal process that functions inversely. The forward process gradually introduces noise into a clean (noise-free) trajectory over K diffusion steps until it becomes completely noisy.
[0068] Assume that represents the operational trajectory of the k-th diffusion step. The original clean trajectory Starting from, the forward diffusion process can be defined as Equation 1.
[0069] [Mathematical Formula 1]
[0070]
[0071] Each β for k=1, 2,...,K k specifies a predetermined dispersion schedule that controls the level of noise added at each diffusion step k. If a sufficiently large K is used Approximates .
[0072] To reverse the diffusion process, the Improved Denoising Diffusion Probabilistic Models methodology can be utilized. This approach is particularly effective for guided sampling because it directly generates a clean trajectory.
[0073] The traffic model is designed to learn an inverse process that transforms sampled noise back into a plausible trajectory. Each iteration of the inverse process is explicitly conditioned on the scene context C to ensure that the output matches predefined constraints. Given the scene context C, the inverse process can be formulated as shown in Equation 2.
[0074] [Mathematical Formula 2]
[0075]
[0076] Here, and and are the mean and variance of the inverse process at diffusion step k, respectively, and θ represents the model parameters.
[0077] Model output It generates a noisy trajectory using known dynamics. Note that it is generated in. In the noise removal step, the traffic model learns how to parameterize the mean of the Gaussian distribution at each diffusion step k.
[0078] In the case of guided sampling, the clean trajectory at each diffusion step k Predicts. To further enhance diversity, a strategy combining CFS and clean trajectory-guided sampling can be implemented during the test period.
[0079] In the case of CFS, future condition models are utilized. and a model without future conditions It trains simultaneously. Through this approach, the model can generate trajectories that reflect future conditions to varying degrees during test time.
[0080] The trajectories predicted by the two models can be merged using CFS weight w as in Equation 3.
[0081] [Mathematical Formula 3]
[0082]
[0083] Here, is a model modulated by CFS and controlled by w It represents the clean trajectory. When set to w=1.0, future information is fully integrated, whereas w=0.0 generates a trajectory without considering future information. This mechanism ensures that the generated trajectory adheres to the guided sampling principle.
[0084] Previous studies directly perturb the average of noise predicted by the network to match the guide. However, approaches relying on a learned loss function require training on the noise level spectrum and face numerical problems regarding the analysis loss function.
[0085] To avoid these problems, the guide formula can be extended to work with arbitrary guide functions using a reconstruction guidance strategy.
[0086] The Clean trajectory can be perturbed as shown in Equation 4.
[0087] [Mathematical Formula 4]
[0088]
[0089] Here, α represents the guidance strength, and represents the selected specific guide. Then, the noise average Calculate as described in Equation 2 Process as network output.
[0090] The process of generating a clean trajectory can be executed at each diffusion step as illustrated in Fig. 5, ensuring strong controllability. In the inference step, CFS controls the amount of future feature information provided to the samples, whereas guided sampling can guide sample generation to the desired result during the noise removal process.
[0091] Inspired by the Vision Transformer (ViT), the present invention can apply a Diffusion Transformer (DiT) customized to generate traffic scenarios.
[0092] The forward path of DiT is explained in detail as follows.
[0093] Figure 6 illustrates the DiT architecture. Referring to Figure 6, the DiT (600) can be structured around three core components: an encoder (610) that tokenizes input data, a transformer block (620) that processes the tokenized noisy input data along with a scene context C, and a transformer decoder (630) that reconstructs the data into its original format.
[0094] DiT (600) is trained to predict a clean trajectory from a noisy trajectory. The input trajectory and traffic information can be first encoded through an encoder (610) and then processed through a transformer block (620).
[0095] Noisy trajectory and traffic information is tokenized using an MLP encoder (610) to form trajectory tokens Tτ∈R 6ХDH and scene context token C∈R DH It generates, where D H represents the hidden dimension. These tokens are processed by the transformer block (620). Based on previous research highlighting the benefits of initializing residual blocks with an identity function to improve training efficiency in large-scale training, this approach can be applied to the Diffusion U-Net model.
[0096] The model of the present invention further utilizes context information to adjust parameters γ and β, and also introduces a dimension-wise scaling parameter α that is integrated immediately before the residual connection of the model architecture.
[0097] Following the transformer block (620), the transformer decoder (630) can convert the context token into a clean action. The clean action is then refined through known dynamics to generate a clean trajectory.
[0098] Diffusion model fine-tuning
[0099] The fine-tuning process of the diffusion model can be summarized into two stages.
[0100] First, an initial model (Multi-Guided Diffusion Model) (referred to as the 'MuDi model') is implemented by applying a multi-task learning approach and fine-tuning the model using a guided condition layer. Then, a final model (Multi-Guided Diffusion Model with Direct Preference Optimization) (referred to as the 'MuDi-Pro model') is implemented by further fine-tuning the model with DPO.
[0101] Figures 7 and 8 illustrate a fine-tuning framework, Figure 7 illustrates the process of fine-tuning a transformer segment according to a guide selected through a guide condition layer, and Figure 8 illustrates the process of further fine-tuning a target model using DPO loss.
[0102] Referring to Fig. 7, a new network architecture featuring a guide condition layer is applied based on recent advancements in task condition approaches and multitask diffusion fine-tuning.
[0103] Figure 7 illustrates a framework for selecting a specific guide. The architecture of the guide condition layer (701) is separated into two main components: a guide embedding layer and a guide encoding network.
[0104] The ruler is a unique vector in the provided guide While it converts to, the latter is a guide vector unique guide potential space Converting to, D H / 4 is the hidden dimension size D H It represents 1 / 4 of the. Then, the guide latent layer is merged with the scene context to be processed in the transformer block (620).
[0105] To optimize fine-tuning efficiency, modifications are limited to the Transformer block and the concluding layer. Through this strategy, the Transformer block can adjust its functionality based on the provided guide, and a single shared layer can act as a multi-decoder, a common feature of multi-task learning.
[0106] Referring to FIG. 8, the present invention can improve the efficiency of guide sampling by utilizing DPO to improve the model, unlike a method of capturing human preferences.
[0107] In the sampling phase of the diffusion model, heuristic-based preferences for intuition can be provided through a preference dataset based on guide scores directly derived from guide loss. Within the DPO framework, prompt c is used, which encompasses all traffic information and Gaussian noise. Then, the DPO dataset Compile.
[0108] The DPO dataset contains win samples derived from the same prompt c and lost sample Instances designated as are included. In a pair of samples, the sample with the higher preference score is identified as the winning sample, while the sample with the lower preference score is labeled as the losing sample.
[0109] After preparing the dataset, the model can be fine-tuned using DPO loss as follows.
[0110] [Mathematical Formula 5]
[0111]
[0112] Here, σ represents the sigmoid function and β acts as a hyperparameter controlling normalization.
[0113] Model p θ is the target model designated for fine-tuning, and model p ref represents the reference model that remains unchanged during the DPO process. The model represents the preferred distribution. It is reparameterized to enable direct optimization for.
[0114] The DPO process is It can improve the model's ability to generate trajectories that are realistic and adhere to established rules by running over epochs.
[0115] The details of the fine-tuning process described above are the same as the algorithm in Fig. 9.
[0116] Referring to Fig. 9, the target model p is created by copying the already trained model. θ and reference model p ref After preparing the target model p θ and reference model p ref Multiple scenarios can be generated through guide sampling.
[0117] Multiple scenarios generated through guided sampling are scored by reflecting the user's intent, and winning and losing samples can be distinguished based on these guide scores.
[0118] Target model p using preference (win or lose) based on guide scores θ It can be fine-tuned. In this case, to increase the precision of guide sampling, the DPO algorithm is used to target model p θ It can be fine-tuned (Equation 5).
[0119] Target model p θ Among the outputs, the winning sample is the reference model p ref To have a higher score than the winning sample, target model p θ Among the outputs, the losing sample is reference model p ref It learns to have a lower score than the losing sample.
[0120] In other words, the present invention, through preference data based on guide scores directly derived from guide loss, enables a target model p without a reward model. θ It is easy to train the model because it can be fine-tuned.
[0121] Accordingly, the present embodiment can provide a scene-level diffusion transformer to ensure the generation of realistic traffic scenarios, and a single model can reinforce multiple guides through a learning method similar to multitask learning, and can ensure a balance between realism and controllability by fine-tuning a data-driven model through a DPO including guide sampling.
[0122] As such, according to the embodiments of the present invention, by fine-tuning a traffic scenario generation model combining a diffusion model and guide sampling using a DPO algorithm, it is possible to generate a traffic scenario that simultaneously satisfies diversity, controllability, and realism.
[0123] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.
[0124] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or collectively. Software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.
[0125] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may continuously store a program executable by a computer, or temporarily store it for execution or download. Additionally, the medium may be various recording or storage means in the form of a single or several hardware combined, and may not be limited to a medium directly connected to a computer system but may exist distributed over a network. Examples of media may include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to store program instructions, including ROM, RAM, and flash memory. Additionally, other examples of media may include recording or storage media managed by app stores that distribute applications or sites and servers that supply or distribute various other software.
[0126] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results can be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
[0127] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.
Claims
Claim 1 A method for generating a traffic scenario of a computer device comprising at least one processor, wherein the method comprises the step of learning a traffic scenario generation model that generates a traffic scenario of an autonomous vehicle through a direct preference optimization (DPO) algorithm by the at least one processor. Claim 2 A method for generating traffic scenarios according to claim 1, characterized in that the traffic scenario generating model is a diffusion model combined with guided sampling. Claim 3 A traffic scenario generation method according to claim 1, wherein the learning step comprises fine-tuning a target model pre-trained as a traffic scenario generation model through direct preference optimization (DPO) including guided sampling. Claim 4 A method for generating traffic scenarios according to claim 3, wherein the fine-tuning step comprises the step of fine-tuning a segment of a transformer according to a guide selected through a guidance conditional layer. Claim 5 A method for generating traffic scenarios according to claim 3, wherein the fine-tuning step comprises: compiling a DPO dataset including instances designated as win samples and lose samples using a prompt encompassing traffic information and Gaussian noise; and fine-tuning the target model using a DPO loss based on the DPO dataset. Claim 6 A method for generating traffic scenarios according to claim 3, wherein the fine-tuning step comprises: generating a plurality of traffic scenarios through guide sampling using a reference model together with the target model; distinguishing winning samples and losing samples based on a guidance score for the plurality of traffic scenarios; and fine-tuning the target model using the winning samples and losing samples according to the guidance score. Claim 7 A method for generating traffic scenarios according to claim 6, wherein the step of fine-tuning the target model comprises the step of learning such that the winning sample of the target model has a higher score than the winning sample of the reference model and the losing sample of the target model has a lower score than the losing sample of the reference model. Claim 8 A method for generating traffic scenarios according to claim 1, wherein the learning step comprises the step of learning a backbone model with scene-level traffic information using a diffusion transformer and CFS (classifier-free sampling). Claim 9 A method for generating traffic scenarios according to claim 8, wherein the backbone model is structured with an encoder that tokenizes input data, a transformer block that processes the tokenized input data together with a scene context, and a transformer decoder that reconstructs the data processed in the transformer block. Claim 10 A traffic scenario generation method according to claim 8, wherein the step of training the backbone model is characterized by training the backbone model to predict a noise-free clean trajectory from a noisy trajectory through a noise removal process that functions inversely to a forward diffusion process. Claim 11 A computer program stored on a computer-readable recording medium to execute the traffic scenario generation method of any one of claims 1 to 10 on the computer device. Claim 12 A computer device comprising at least one processor implemented to execute readable instructions on a computer device, wherein the at least one processor processes the process of learning a traffic scenario generation model that generates traffic scenarios for an autonomous vehicle through a direct preference optimization (DPO) algorithm. Claim 13 A computer device characterized in that, in paragraph 12, the traffic scenario generation model is a diffusion model combined with guided sampling. Claim 14 A computer device according to claim 12, wherein at least one processor fine-tunes a target model pre-trained as a traffic scenario generation model through direct preference optimization (DPO) including guided sampling. Claim 15 A computer device according to claim 14, wherein at least one processor fine-tunes a segment of a transformer according to a guide selected through a guidance conditional layer. Claim 16 A computer device according to claim 14, wherein at least one processor compiles a DPO dataset including instances designated as win samples and lose samples using a prompt encompassing traffic information and Gaussian noise, and fine-tunes the target model using a DPO loss based on the DPO dataset. Claim 17 A computer device according to claim 14, wherein at least one processor generates a plurality of traffic scenarios through guide sampling using a reference model together with the target model, distinguishes winning samples and losing samples based on a guidance score for the plurality of traffic scenarios, and fine-tunes the target model using the winning samples and losing samples according to the guidance score. Claim 18 A computer device according to claim 17, wherein at least one processor learns that the winning sample of the target model has a higher score than the winning sample of the reference model and the losing sample of the target model has a lower score than the losing sample of the reference model. Claim 19 A computer device according to claim 12, wherein at least one processor learns a backbone model using scene-level traffic information using a diffusion transformer and CFS (classifier-free sampling). Claim 20 A computer device according to claim 19, wherein at least one processor learns the backbone model to predict a noise-free clean trajectory from a noisy trajectory through a noise removal process that functions inversely to a forward diffusion process.