Device and method for generating first agent, in particular for interaction between the first agent and second agent, and device and method for training at least one model for generating the first agent
A multi-layered neural network approach for agent generation and training addresses the inefficiencies of existing methods by using pre-trained models to map descriptions onto representations, enhancing data efficiency and domain knowledge transfer for realistic agent behavior.
Patent Information
- Application Number
- JP2025019261
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-09
- Filing Date
- 2025-02-07
- Publication Date
- 2025-08-22
AI Technical Summary
Existing methods struggle to efficiently generate and train agents for interactions, particularly in dialogue scenarios, requiring extensive training data and lacking efficient knowledge transfer across domains.
A method involving pre-trained neural networks to map descriptions of agent behavior onto representations, using a multi-layered model architecture to generate output variables that influence agent behavior, allowing for less data-intensive training and domain knowledge transfer.
Enables efficient generation and training of agents with realistic behavior, reducing the need for extensive data and facilitating domain knowledge transfer, thereby improving interaction scenarios.
Smart Images

Figure 2025123202000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention is particularly directed to an apparatus and method for generating a first agent for a dialogue between a first agent and a second agent, as well as an apparatus and method for training at least one model for generating the first agent. Summary of the Invention [Means for solving the problem]
[0002] Disclosure of the Invention It is envisaged that a method for generating a first agent, in particular for a dialogue between a first agent and a second agent, includes mapping a description of the behavior of the first agent, in particular in the dialogue between the first agent and the second agent, onto a first representation using a first model configured to map the description onto the first representation, the first representation being mapped onto a second representation using a second model configured to map the first representation onto the second representation, and the second representation being mapped onto output variables using a third model configured to map the second representation onto output variables for influencing the behavior of the first agent, the description being pre-defined in a natural language, in particular in textual or audio format, or in a formal language or digital graphic format, and in particular the behavior of the first agent, in the dialogue between the first agent and the second agent, being pre-defined depending on the output variables. Each model may be a stochastic, probabilistic or deterministic model.
[0003] For example, an anomaly in the behavior of a first agent in a dialogue between the first agent and the second agent is identified depending on the dialogue.
[0004] It may be assumed that the output variables include a trajectory of the first agent and / or that the output variables include controller parameters for a controller of the first agent, and that the behavior of the first agent is determined depending on the behavior of the controller for the first agent.
[0005] It may be assumed that the first model comprises a pre-trained artificial neural network, and / or the second model comprises a pre-trained artificial neural network, and / or the third model comprises a pre-trained artificial neural network.
[0006] A method for training at least one model for generating a first agent, in particular for an interaction between a first agent and a second agent, is envisaged, in which a description of the behavior of the first agent, in particular in an interaction between the first agent and a second agent, is mapped onto a first representation using a first model configured to map the description onto the first representation, the first representation is mapped onto a second representation using a second model configured to map the first representation onto the second representation, and the second representation is mapped onto output variables using a third model configured to map the second representation onto output variables, the description being predefined in a natural language or a formal language or a digital graphic format, references for the description and the output variables being predefined, the references characterizing realistic behavior in the real world that is appropriate for the description of the first agent, and the second model being trained depending on the differences between the output variables and the reference, whereby the second model is trained for output of output variables that define behavior appropriate for the description of the first agent that is physically and as realistic as possible in the real world.
[0007] It may be assumed that the first model comprises a pre-trained artificial neural network and / or that the third model comprises a pre-trained artificial neural network, meaning that the training is based on an already available model.
[0008] The first model and / or the third model remain unchanged during training, allowing the second model to be trained exclusively, thereby requiring less training data than if the second model were trained together with the first model and / or the third model.
[0009] The reference may be envisaged to include the trajectory of the first agent and / or control parameters for the controller of the first agent.
[0010] An apparatus for generating an interaction, or for training at least one model, or for training a first agent for an interaction between the first agent and a second agent includes at least one processor and at least one memory, wherein the at least one processor is configured to execute instructions that, when executed by the at least one processor, cause the apparatus to perform the method, and the at least one memory stores the instructions.
[0011] The data structure includes at least one data field for describing the behavior of a first agent, in particular in a dialogue between the first agent and a second agent in a natural or formal language, the data structure includes at least one data field for a first representation of the description, the data structure includes at least one data field for a second representation of the description, and the data structure includes at least one data field for an output variable for influencing the behavior of the first agent.
[0012] It may be envisaged that the data structure includes at least one data field for a first model configured to map the description onto the first representation, the data structure includes at least one data field for a second model configured to map the first representation onto the second representation, and / or the data structure includes at least one data field for a third model configured to map the second representation onto output variables.
[0013] A computer program may be envisaged which comprises computer executable instructions, which when executed by the computer on a computer cause the method to proceed.
[0014] Further advantageous embodiments can be gleaned from the following description and drawings. [Brief explanation of the drawings]
[0015] [Figure 1] FIG. 1 is a schematic diagram of an apparatus for machine learning or for generating a behavior of a first agent. [Figure 2] FIG. 1 is a schematic diagram of a model for machine learning or for generating the behavior of a first agent. [Figure 3] FIG. 2 is a schematic diagram of an exemplary interaction. [Figure 4] 1 is a flowchart with steps of a method for generating a behavior of a first agent. [Figure 5] 1 is a flowchart with steps of a method for machine learning. DETAILED DESCRIPTION OF THE INVENTION
[0016] In FIG. 1 an apparatus 100 for machine learning, or in particular for generating a behavior of a first agent in an interaction between said first agent and a second agent, is shown schematically.
[0017] The device 100 includes at least one processor 102 and at least one memory 104 .
[0018] The at least one processor 102 is configured to execute instructions that, when executed by the at least one processor 102, cause the device 100 to perform methods for machine learning or for generating behaviors, as described below.
[0019] The at least one memory 104 stores instructions. The at least one memory 104 includes, for example, a non-volatile memory. The at least one memory 104 includes, for example, a volatile memory.
[0020] The device 100 includes, in this example, an interface 106, which is configured to receive, for example, a description of a behavior.
[0021] The interface 106 may be configured to capture the description in textual, audio or digital graphic form.
[0022] The interface 106 may be configured to request the description by output in textual or audio form. The device 100 may, for example, be configured to request and capture the description in a dialogue with a user.
[0023] For example, device 100 may be configured to automatically query and populate the description with information needed for the description, and may be configured to generate the description if, for example, device 100 identifies that it has captured the information needed for the description.
[0024] FIG. 2 shows a schematic representation of a model for machine learning or for generating behavior.
[0025] The first model 202 is configured in this example to map input variables 204 of the first model 202 onto output variables 206 of the first model 202. The first model 202 may be, for example, an encoder, or a sequential and / or autoregressive encoder, or a transformer architecture.
[0026] The second model 208 is configured in this example to map the output variables 206 of the first model 202 onto the output variables 210 of the second model 208. The second model 208 is, for example, a translator that translates the output variables 206 of the first model 202 into the output variables 210 of the second model 208, i.e., into the input variables of the third model 212.
[0027] The third model 212 is configured in this example to map the output variables 210 of the second model 208 onto the output variables 214 of the third model 212. This third model 212 is, for example, a decoder.
[0028] The input variables 204 of the first model 202, in this example, include a description in natural language, a formal language, or a digital graphic format. The output variables 206 of the first model 202, in this example, include a first representation, which may be, for example, a first embedding or a first sequence of tokens. The output variables 210 of the second model 208, in this example, include a second representation, which may be, for example, a second embedding or a second sequence of tokens. The output variables 214 of the third model 212 define the behavior.
[0029] The first model 202 may, for example, comprise an artificial neural network. The second model 208 may, for example, comprise an artificial neural network. The third model 212 may, for example, comprise an artificial neural network.
[0030] The apparatus 100 includes, for example, a model. The apparatus 100 is configured to receive, for example, a description via an interface 106 and to determine output variables for the description using the model.
[0031] It may be assumed that the device 100 is configured to perform a test, for example, the device 100 being configured to examine the behavior of a second agent in a dialogue with a first agent depending on the dialogue, and the device 106 being configured to identify an anomaly in the dialogue and output as a result of the test via the interface 106 the presence of an anomaly or the absence of any anomaly identified in the test.
[0032] The dialogue is not limited to the agent to be examined in the test or the agent that can be synthetically moved in the dialogue. Multiple agents that need to be examined using the dialogue in the test may be assumed. Also, multiple agents that can be synthetically moved in the dialogue in the test may be assumed. This description, for example, describes the behavior of each agent that should be synthetically moved. The behavior of each synthetically movable agent is defined, for example, by the output variable 214.
[0033] 3 illustrates an exemplary scenario 300 having an example interaction with a first vehicle 302 for a first agent and a second vehicle 304 for a second agent. The scenario 300 includes a trajectory 306 of the first vehicle 302 and a trajectory 308 of the second vehicle 304.
[0034] Scenario 300 is illustrated for the following example description in natural language:
[0035] Here, in a highway approach scenario, a first vehicle is traveling on a highway approach road and needs to cut in behind a second vehicle traveling on the highway.
[0036] In this example, the first vehicle 302 includes at least one sensor for capturing sensor data characterizing the surroundings of the vehicle 302 .
[0037] In this example, the first vehicle 302 includes a controller for controlling the behavior of the first vehicle 302 in dependence on the sensor data, the controller being parameterized, for example, by controller parameters.
[0038] In testing, for example, the second vehicle 304 is examined in interaction with the first vehicle 302 .
[0039] The pictorial representation of the scenario 300 shows a trajectory 306 in digital graphic form as an exemplary description of the behavior of the first vehicle 302. The pictorial representation is exemplary. The apparatus 100 may be configured to use the scenario 300 in a formal language that allows automated processing in testing, for example, by a test bench or by a simulation environment in which the testing is performed.
[0040] It may be assumed that the descriptions provided as input variables 204 of the first model 202 are supplemented with information in natural language for framework conditions, which may be defined in natural language, for example, as follows:
[0041] These include the geometry of the highway, weather conditions such as dryness, wind, rain, snow, and ice, visibility conditions such as fog, darkness, and brightness, and traffic rules such as speed limits and no overtaking.
[0042] The scenario 300 includes, for example, a map description including framework conditions, the map description being written in a formal language that allows automated processing by, for example, a test bench or a simulation environment.
[0043] In FIG. 4 a flow chart with the steps of the method is shown.
[0044] These agents are road users in one example. A first agent is, for example, a first vehicle 302. A second agent is, for example, a second vehicle 304.
[0045] In one example from robotics, the first agent is a robot and the second agent is a human model.
[0046] The method is based on a predefined description.
[0047] The description may be predefined, for example, in a natural language, in particular in textual or audio form, or in a formal language or in digital graphic form. The user may enter the description, for example, in textual form, or may speak the description in audio form. The user may represent the description in digital graphic form, for example in the form of a diagram.
[0048] The description may be requested, for example, by output in text or audio format, or may be requested and captured in a dialogue with the user, for example.
[0049] The information required for the description can be queried, for example, automatically, and added to the description.
[0050] A description is generated, for example, only when it is identified that the information required for the description has been captured.
[0051] For example, descriptions of rare interactions in the real world are captured using language, among other things. An example of a rare interaction is an interaction where the second agent is an autonomous vehicle, where the first agent unexpectedly drives out in front of the autonomous vehicle from a parking lot that is not visible to the autonomous vehicle. An example of a rare interaction is an interaction with an aggressive behavior pattern of the agent, for example, when the agent merges with other agents in traffic on a highway. An example of a rare interaction is an interaction under extreme weather conditions in which the first agent and / or the second agent are present. An example of a rare interaction is an interaction under very rare environmental conditions, for example, an interaction on a day when the first agent represents a human dressed in, for example, Halloween, Fasching, or Carnival costume. In this context, "rare" means that the interaction occurs extremely rarely in training data captured in the real world.
[0052] The method is based on a first model 202, a second model 208, and a third model 212. The first model 202, the second model 208, and the third model 212 are pre-configured in the method in this example, e.g., the respective artificial neural networks are pre-trained.
[0053] The first model 202 may be, for example, a pre-trained model for text and / or audio input, configured to map descriptions onto a first representation in text or audio form, such that the use of a pre-trained first model 202 allows for greater data efficiency compared to individual models trained to map descriptions directly onto output variables 214.
[0054] The behavior of the first agent is generated, for example, in a specific domain, such as transportation or robotics. The first model 202 and / or the third model 212 are pre-trained, for example, in another domain or without content specialization in the domain in which the agent's behavior is generated. The domain in which the first model 202 and / or the third model 212 are pre-trained is a different domain from the domain in which the behavior of the first agent is generated, for example, a general domain.
[0055] The pre-trained first model 202 and / or the pre-trained third model 212 enable knowledge transfer, particularly from the domain in which the first model 202 or the third model 212 is pre-trained to the domain in which the first agent is generated. For example, each pre-trained model may include general knowledge about the behavior of the first agent. For example, the behavior of a pedestrian represented by the first agent can be better generated based on general knowledge if each pre-trained model is pre-trained with a large amount of data including human behavior. Thus, subsequent training of each pre-trained model to generate the first agent requires only a small amount of data including pedestrian behavior to generate the behavior of the first agent in a relatively good manner. This would otherwise be possible only with a larger amount of data including pedestrian behavior.
[0056] The method includes step 402 .
[0057] In step 402 , the predefined description is mapped onto the first representation using the first model 202 .
[0058] The method includes step 404 .
[0059] In step 404 , the first representation is mapped onto the second representation using the second model 208 .
[0060] The method includes step 406 .
[0061] In step 406 , the second representation is mapped onto the output variables 214 using the third model 212 .
[0062] The method may include further steps for training and / or testing the behavior of the second agent.
[0063] Steps 402-406 are repeated to generate training data including, for example, output variables 214 for training and / or testing the behavior of the second agent. For example, a number of output variables 214 are determined from a catalog of descriptors and added to the training data.
[0064] The catalog includes, for example, a description of interactions within the domain. An example of a domain description is Koopman, P., Osyk, B., Weast, J., et al. (2019) "Autonomous Vehicles Meet the Physical World: RSS, Variability, Uncertainty, and Proving Safety." Romanovsky, A., Troubitsyna, E., Bitsch, F. (eds.) "Computer Safety, Reliability, and Security," SAFECOMP 2019. Lecture Notes in Computer Science (), vol. 11698. Springer, Cham. https: / / doi.org / 10.1007 / 978-3-030-26601-1_17 describes the operational design domain.
[0065] For example, a method for training includes step 408 .
[0066] In step 408, the behavior of the second agent is trained in an interaction between the first agent and the second agent.
[0067] The behavior of the first agent is determined, for example, by a trajectory preset in a scenario for the first agent.
[0068] The first agent may include a controller configured to determine a behavior of the first agent depending on the scenario, the behavior of the first controller being determined, for example, by a controller within the first agent.
[0069] For example, the second agent is trained using reinforced learning, i.e., reinforcement learning, such that the second agent is paid a reward if the second agent does not collide with the first agent moving on a preset trajectory, and no reward is paid otherwise.
[0070] The second agent may include a sensor and a controller configured to determine a behavior of the second agent depending on information about the first agent measured by the sensor. The behavior of the second controller may be determined, for example, by a controller within the second agent. For example, the controller of the second agent may be trained using reinforcement learning.
[0071] For example, a method for testing includes step 410 .
[0072] The behavior of the second agent is monitored in step 410. It may be assumed that the behavior of the first agent is monitored or that the behavior of the first agent generated using the model is used as a basis for testing.
[0073] The behavior of the first agent is determined, for example, by a trajectory or a controller within the first agent.
[0074] For example, a method for testing includes step 412 .
[0075] In step 412, an anomaly in the behavior of the second agent is identified depending on the interaction between the second agent and the first agent. For example, an anomaly in the behavior of the second agent is identified depending on the interaction.
[0076] For example, an anomaly in the behavior of a second agent may be identified upon collision between the first agent and the second agent.
[0077] Using this method, multiple agents may be generated, including their interaction behavior, relative to each other and potentially to third parties. The generated agents may be expected to interact with one or more external agents, either in a real-world test environment or in a simulation, to ultimately test how the external agents interact with the generated agents.
[0078] The tests themselves may be performed, for example, in a real test environment or in a simulation where agents interact, with simulation being a priority, especially for tests that may lead to collisions, in order to avoid endangering human life or preventing harm to real agents.
[0079] For example, if a merging scenario is described in which two agents are traveling with almost no space between them in an approach road area on a highway, the behavior of the two agents traveling on the highway will be generated in accordance with the described merging scenario.
[0080] For example, for an external agent whose behavior is not generated using a description but is controlled by an external pre-set control device, it is tested whether the external agent can nevertheless merge without incident.
[0081] Additionally, it may be envisaged that tests are carried out on the domain in which the first agent is generated and on further descriptions preset by the user.
[0082] For example, further tests may be performed using dialogue captured in the real world, or the dialogue captured in the real world may be randomly selected from a pre-defined set of dialogues captured in the real world.
[0083] FIG. 5 shows a flow chart with steps of a method for training at least one model of a plurality of models.
[0084] The method for training is based on a first model 202, a second model 208, and a third model 212. In this example, the first model 202 and the third model 212 are pre-configured in the method for training at least one model of the plurality of models. For example, the neural networks of each of the first model 202 and the third model 212 are pre-trained.
[0085] Pre-trained models, due to the information stored in the pre-training and due to the density of information contained in the language, offer knowledge transfer and / or greater data efficiency compared to models that are trained only on the domain in which the first agent is generated.
[0086] The method for training at least one model of the plurality of models is based on training data, which in this example include training data spots each including a predetermined description and a reference assigned to the predetermined description. Alternatively, the method can be based on a pre-trained model.
[0087] The method for training at least one model of a plurality of models includes step 502 .
[0088] In step 502, the description and references from the training data spots are pre-established.
[0089] In step 504, the predefined descriptions of the training data spots are mapped onto the first representation using the first model 202.
[0090] The method for training at least one model of a plurality of models includes step 506 .
[0091] In step 506, for the training data spots, the first representation is mapped onto the second representation using the second model 208.
[0092] The method for training at least one model of the plurality of models includes step 508 .
[0093] In step 508, for the training data spots, the second representation is mapped onto the output variables 214 using the third model 212.
[0094] The references include, for example, output variables for realistic and descriptive behavior in the real world, such as trajectories defining the behavior of the first agent, controller parameters for the controller of the first agent, and / or maps defining framework conditions for the behavior of the first agent.
[0095] Steps 502 through 508 are performed on training data spots from the training data in this example.
[0096] The method for training at least one model of a plurality of models includes step 510 .
[0097] In step 510, the second model 208 is trained depending on the difference between the output variables 214 and the reference.
[0098] In this example, the second model 208 is trained based on an objective function that includes the respective differences determined for the training data spots, for example, the second model 208 is determined to minimize, in particular, the objective function that depends on the sum of the differences.
[0099] For example, the parameters of the neural network comprising the second model 208 are determined differentially using gradient descent.
[0100] In this example, the first model 202 and the third model 212 remain unchanged during the training of the second model 208. The first model 202 and / or the third model 212 may be expected to train similarly during training.
[0101] The output variables of the models are, for example, aggregated into vectors. For example, a trajectory or several controller parameters are described in a scenario by a vector output by the third model 212. A map is, for example, described by a vector output by the third model 212, which vector contains parameters of the formal language in which the map is described.
[0102] The first model 202, the second model 208, and / or the third model 212 may be stochastic, probabilistic, or deterministic models.
[0103] The first model 202, the second model 208, and the third model 212 are configured for reverse mapping in one example. This reverse mapping enables behavioral interpretability because the models are correlated with the language. This means that the third model 212 is configured to map its output variables 214 onto the output variables 210 of the second model 208. The second model 208 is configured to map its output variables 210 onto the output variables 206 of the first model 202. The first model 202 is configured to map its output variables 206 onto its input variables 204.
[0104] For the backward mapping, it can be assumed that for the output variable 214 of the third model 212, successive images of a video are used, in which case this video is mapped onto a text description of the video content, meaning that from one behavioral video, an appropriate text description is generated.
[0105] In a method for describing the behavior of a second agent, particularly in interaction with a first agent, it is assumed that the output variables 214 of the third model 212 comprise the behavior to be described.
[0106] For purposes of illustration, it is assumed that the output variables 214 of the third model 212 are mapped to the output variables 210 of the second model 208 .
[0107] For purposes of illustration, it is assumed that the output variables 210 of the second model 208 are mapped to the output variables 206 of the first model 202 .
[0108] In this method for explanation, it is assumed that the output variables 206 of the first model 202 are mapped to the input variables 204 of the first model 202. The input variables 204 of the first model 202 contain a description of the behavior to be explained, in particular the interaction to be explained between the first agent and the second agent.
Claims
1. In particular, a method for generating a first agent (302) for a dialogue (300) between the first agent (302) and a second agent (304), comprising: In particular, a description of the behavior of the first agent (302) in the interaction (300) between the first agent (302) and the second agent (304) is mapped (402) onto the first representation using a first model (202) configured to map the description onto the first representation; The first representation is mapped (404) onto a second representation using a second model (208) configured to map the first representation onto the second representation; the second representation is mapped (406) onto output variables (214) using a third model (212) configured to map the second representation onto output variables (214) for influencing the behavior of the first agent (302); said description being preset in a natural language, in particular in textual or audio form, or in a formal language or digital graphic form, In particular, the behavior of the first agent (302) in the interaction (300) between the first agent (302) and the second agent (304) is preset (408) depending on the output variables (214); A method characterized by:
2. 2. The method of claim 1, wherein an anomaly in the behavior of the second agent in the interaction between the first agent and the second agent is identified depending on the interaction.
3. the output variables (214) include a trajectory of the first agent (302); and / or 3. The method of claim 1, wherein the output variables (214) include controller parameters for a controller of the first agent (302), and the behavior of the first agent (302) is determined (410) depending on the behavior of the controller for the first agent (302).
4. 4. The method of claim 1, wherein the first model (202) comprises a pre-trained artificial neural network, and / or the second model (208) comprises a pre-trained artificial neural network, and / or the third model (212) comprises a pre-trained artificial neural network.
5. 1. A method for training at least one model for generating a first agent (302) particularly for a dialogue (300) between a first agent (302) and a second agent (304), comprising: In particular, a description of the behavior of the first agent (302) in the interaction (300) between the first agent (302) and the second agent (304) is mapped (504) onto the first representation using a first model (202) configured to map the description onto the first representation; The first representation is mapped (506) onto the second representation using a second model (208) configured to map the first representation onto the second representation; the second representation is mapped (508) onto output variables (214) using a third model (212) configured to map the second representation onto output variables (214) for influencing the behavior of the first agent (302); the description is preset in a natural language or in a formal language or digital graphic format, References for the description and the output variables (214) are preset (502); the reference characterizes behavior that is realistic in the real world and appropriate to the description of the first agent (302); The second model is trained (510) depending on the difference between the output variables (214) and the reference. A method characterized by:
6. 6. The method of claim 5, wherein the first model (202) comprises a pre-trained artificial neural network and / or the third model (212) comprises a pre-trained artificial neural network.
7. The method of claim 6 , wherein the first model (202) and / or the third model (212) are kept unchanged (510) during training.
8. The method of any one of claims 5 to 7, wherein the reference comprises a trajectory of the first agent (302) and / or controller parameters for a controller of the first agent (302).
9. 1. An apparatus (100) for generating a dialogue, or for training at least one model, or for training a first agent (302) for a dialogue (300) between a first agent (302) and a second agent (304), comprising: The device (100) comprises: at least one processor (102); At least one memory (104); Including, The at least one processor (102) is configured to execute instructions that, when executed by the at least one processor (102), cause the apparatus (100) to perform the method of any one of claims 1 to 8; the at least one memory (104) storing the instructions; 1. An apparatus (100) comprising:
10. In data structures, the data structure includes at least one data field for describing the behavior of the first agent (302) in a dialogue between the first agent (302) and the second agent (304), in particular in a natural or formal language; the data structure includes at least one data field for a first representation of the description; the data structure includes at least one data field for a second representation of the description; the data structure includes at least one data field for an output variable (214) for influencing the behavior of the first agent (302); 1. A data structure comprising:
11. 11. The data structure of claim 10, wherein the data structure includes at least one data field for a first model (202) configured to map the description onto the first representation, the data structure includes at least one data field for a second model (208) configured to map the first representation onto the second representation, and / or the data structure includes at least one data field for a third model (212) configured to map the second representation onto the output variables (214).
12. 9. A computer program comprising computer executable instructions, which when executed by the computer on the computer cause the method of any one of claims 1 to 8 to proceed.
Citation Information
Cited By
Drilling position determination system, drilling control system, and work machine
US12497751B2