Method for generating textual description of decision made automatically when controlling robotic device

By using a multi-module control processing chain and encoder technology, textual descriptions of autonomous robot decisions are generated, solving the problem of difficult-to-interpret decisions in autonomous control and improving safety and development efficiency.

CN121005000APending Publication Date: 2025-11-25ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510663350.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-24
Filing Date
2025-05-22
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

In the autonomous control of robotic devices, existing technologies struggle to provide easily understandable explanations for automated decision-making, impacting safety and development efficiency.

Method used

By using a multi-module control processing chain, combined with rule-based intermediate steps and machine learning models, a textual description of the autonomous robot device's decision-making is generated. An encoder is used to encode the decision-making process into a textual description, and the annotation text that best matches the decision-making process is selected for interpretation.

Benefits of technology

It improves the interpretability of autonomous robot equipment decisions, shortens the development cycle, and enhances safety and user understanding of the decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121005000A_ABST
    Figure CN121005000A_ABST
Patent Text Reader

Abstract

According to various embodiments, a method for generating a textual description of a decision that is automatically made when controlling a robotic device is described, comprising: processing data containing information about an environment of the robotic device by means of a control processing chain having a plurality of modules, wherein at least some of the modules output protocols with respect to rule-based intermediate steps performed by the respective modules when controlled, encoding inputs of the at least some of the modules, outputs of the at least some of the modules and the protocols output by the at least some of the modules into decision process codes, and selecting, from the set of textual descriptions, a textual description of at least one decision made when the data is processed by the control processing chain according to the decision process encoding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to methods for generating decisions that are made automatically when controlling robotic devices. Background Technology

[0002] In the autonomous control of robotic devices, especially autonomous vehicles, complex processing chains employing machine learning models are typically used. A key challenge here is interpretability—understanding why such processing chains (especially machine learning models) make specific decisions. For example, it is interesting to consider whether an autonomous robotic device should be reconfigured if it is not operating as expected (at least superficially) or even misbehaving. Such understanding can also improve safety, for instance, by explaining control decisions to the user in advance and allowing the user to overturn those decisions if necessary. Therefore, it is desirable to provide easily understandable explanations of the decisions automatically made when controlling robotic devices.

[0003] Alec Radford et al.’s open paper “Learning transferable visual models from natural language surveillance” (8748-8763), presented at the International Conference on Machine Learning (PMLR, 2021), hereinafter referred to as “Reference 1”, describes the CLIP method for jointly training an image encoder and a text encoder to find the correct pairing of an image and the matching text in an input image and text. Summary of the Invention

[0004] According to different implementations, a method is provided for generating text descriptions of decisions made automatically when controlling a robotic device. The method comprises: processing data containing information about the environment of the robotic device through a control processing chain having multiple modules, wherein at least some modules output protocols regarding rule-based intermediate steps executed by corresponding modules during control; encoding the inputs of at least some modules, the outputs of at least some modules, and the protocols output by at least some modules into a decision process code; and selecting, based on the decision process code, text descriptions from a set of text descriptions for at least one decision made when processing data through the control processing chain.

[0005] The above method achieves the following: providing informative annotations about decisions related to the control actions that autonomous robotic devices (e.g., autonomous vehicles (AVs) or autonomous robots) have performed or intend to perform.

[0006] This allows developers or users to understand the reasons behind the autonomous robotic device's decisions regarding corresponding control actions. Therefore, for example, during test runs of an autonomous vehicle, developers do not need to guess why the vehicle performs specific actions. This can help shorten development cycles because specific test scenarios can be repeated using the reasons why the autonomous vehicle (i.e., the corresponding control device that selects the control action) makes a particular decision. Such annotations are also of interest to users, as they can consider specific situations, such as when the autonomous vehicle makes a particular decision, during operation. If the user disagrees with the decision, they can reconfigure the autonomous vehicle so that it does not make the decision (and avoids undesirable actions arising from it). Informing the user of driving decisions in advance also improves safety, as the user can identify erroneous situations that could lead to potentially safety-critical actions and can veto autonomous control to perform safe actions.

[0007] Different implementation methods are described below.

[0008] Example 1 is a method for generating decisions that are made automatically when controlling a robotic device, as described above.

[0009] Example 2 is based on the method of Example 1, which involves selecting text descriptions from the set of text descriptions by evaluating the consistency between the (text) encoding of the text description and the encoding of the decision process for each text description from the set of text descriptions, and selecting the text description with the best evaluation (e.g., Euclidean distance in the encoding space).

[0010] This enables the simple generation of annotations without requiring a generative model, as pre-defined annotation text can also be used. The encoder used to encode the text descriptions of the set of text descriptions can also be trained along with the processing chain, or after its training.

[0011] Example 3 is a text description of a set of text descriptions generated by means of a generative model (e.g., a large language model (LLM)) according to the method of Example 1 or 2, wherein the generative model receives inputs containing internal states, variable values ​​and / or intermediate results of a control processing chain.

[0012] The use (and training) of such models requires corresponding costs, but with proper training, they can improve the quality and diversity of text descriptions. The input to generative models can also at least partially include the inputs and / or outputs of the processing chain.

[0013] Example 4 is based on any one of Examples 1 to 3, and includes displaying the generated text description on the display of a robot device.

[0014] This explains the decision made to the user.

[0015] Example 5 is based on any one of Examples 1 to 4, wherein the robotic device is an autonomous vehicle controlled in a traffic scenario.

[0016] In this context, annotations of control actions are particularly important, for example, for the user (i.e., the driver, who is often a layman in this case) in order to understand the vehicle's behavior.

[0017] Example 6 is a control device (control device for robotic equipment) designed to perform the method according to any one of Examples 1 to 5.

[0018] The control device is particularly capable of implementing processing chains.

[0019] Example 7 is a computer program having instructions that, when executed by a processor, cause the processor to perform a method according to any one of Examples 1 to 5.

[0020] Example 8 is a computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform a method according to any one of Examples 1 to 5. Attached Figure Description

[0021] In the accompanying drawings, similar reference numerals generally refer to the same parts in completely different views. The drawings are not necessarily to scale, and are instead intended to illustrate the principles of the invention. In the following description, different aspects are described with reference to the accompanying drawings.

[0022] Figure 1 The vehicle is shown.

[0023] Figure 2 The diagram shows a processing chain consisting of modules used to control the vehicle.

[0024] Figure 3 A flowchart is shown, illustrating a method according to one embodiment for generating textual descriptions of decisions made automatically when controlling a robotic device. Detailed Implementation

[0025] The following detailed description refers to the accompanying drawings, which, for illustrative purposes, illustrate specific details and aspects of the invention that may be practiced. Other aspects may be used and structural, logical, and electrical changes may be made without departing from the scope of the invention. Different aspects of this disclosure are not necessarily mutually exclusive, as some aspects of this disclosure may be combined with one or more other aspects of this disclosure to form new aspects.

[0026] Below, we describe the different examples in more detail.

[0027] Figure 1 Vehicle 101 is shown.

[0028] exist Figure 1 In the example, vehicle 101 (e.g., a motor vehicle, such as a private vehicle or a truck) is equipped with a vehicle control device (e.g., an electronic control unit (ECU)) 102.

[0029] The vehicle control unit 102 has data processing components, such as a processor (e.g., a CPU (Central Processing Unit)) 103 and a memory 104. The memory stores control software 107 and data. The vehicle control unit 102 operates according to the control software, and the processor 103 processes the data. The processor 103 executes the control software 107.

[0030] For example, the stored control software (computer program) has instructions that, when executed by the processor, cause the processor 103 to perform driving assistance functions or even autonomously control the vehicle.

[0031] For example, control software 107 is transmitted from computer system 105 to vehicle 101, for example, via network 106 (or by means of a storage medium, such as a memory card). This can also occur during operation (or at least when vehicle 101 is with a user), as control software 107 is updated to a new version, for example, over time.

[0032] For example, the control software 107 can be trained using machine learning (ML), that is, the control software 107 implements one or more ML models 108 (or "machine learning models") trained on training data, which in this example is trained by the computer system 105. Therefore, the computer system 105 implements an ML training algorithm for training one or more ML models 108.

[0033] Control software 107 derives control actions (e.g., steering, braking, etc.) for the vehicle from input data 109, said input data being available to the control software and containing information about the environment, or said control software deriving information about the environment from said input data (such as by detecting other traffic participants, such as other vehicles). For example, said input data includes sensor data, such as information obtained from the vehicle's cameras or via communication with other vehicles or devices on the roadside.

[0034] Sensor data (and potentially additional information, such as digital maps or information sent to the vehicle by other traffic participants or infrastructure units) is processed by a (control) processing chain to control the vehicle, such as a modular processing chain, like... Figure 2 As shown in the image. In the following text, the vehicle to be controlled is also referred to as a self-propelled vehicle.

[0035] Figure 2 The processing chain consists of modules 201, 202, and 203 for controlling the vehicle.

[0036] In the example, the processing chain includes a perception module 201, a prediction module 202, and a planning module 203 (i.e., a module chain). These modules can be implemented at least in part by an ML (machine learning) model, such as a neural network, where it is assumed below that the driving strategy (or general decision-making strategy) derived from the processing chain has been trained.

[0037] The perception module 201 obtains control input data 204 containing information about the traffic scene. For example, this is sensor data (e.g., camera data, LiDAR data, radar data), map information, and / or information received (e.g., via V2X (vehicle-to-everything) communication).

[0038] The perception module 201 detects the vehicle's environment, for example, by locating the vehicle (or other objects), object detection (e.g., detecting other traffic participants), and object tracking (e.g., of other traffic participants). It provides perception results 205, such as a list of objects, occupancy gates of the vehicle's environment, etc.

[0039] The prediction module 202 makes predictions about the future state of the vehicle environment based on prediction results, such as the future trajectories of other traffic participants; however, it can also derive (“predict”) the possible behavior of the vehicle itself. The prediction module 202 provides prediction results (e.g., predictions of the trajectories (or trajectory ranges) of other traffic participants).

[0040] Based on priority, planning module 203 searches for a safer, more comfortable, and / or faster trajectory for the vehicle based on the prediction results. The output of the planning module is a planning result 207, such as a specification of the vehicle's trajectory or behavior in the form of waypoints (e.g., a specification of boundary conditions that must be followed). This output (possibly by another module) is then converted into control actions (braking, steering, etc.). Alternatively, planning module 203 itself can provide the (low-level) control actions.

[0041] Therefore, the processing chain performs localization, perception, prediction, and planning, for example:

[0042] • Localization and perception aim to accurately pinpoint the vehicle's location within the environment or provide a reliable model of the 3D environment.

[0043] • Predict other vehicles or their intentional actions.

[0044] • Plan the route that the bicycle should take.

[0045] Typically, the autonomous vehicle (i.e., the vehicle control unit 102 implementing the processing chain) makes control decisions based on a specific strategy that sets optimal actions based on detections of its environment. According to different implementations, the interpretability and clarity of the actions performed by the vehicle (i.e., its control decisions) are improved by generating and outputting textual annotations of the autonomous vehicle's behavior, for example, from the control actions (i.e., the reasons for selecting the control actions).

[0046] Ensuring passenger comfort, trust, and safety is a crucial pillar of autonomous driving. Systems that annotate the driving behavior of autonomous vehicles in real time can help achieve this goal. Such systems provide insights into the vehicle's decision-making processes and facilitate a deeper understanding of its operational logic.

[0047] Here, based on different implementations, not only are the control actions selected by the end-to-end architecture (for machine learning (ML), such as neural networks) described (i.e., outputting annotated text for driving trajectories), but also the decisions performed by the "classic" (non-ML-based) rule-based components (which perform rule-based intermediate processing steps) are annotated, such as filtering or pruning the graph by means of a threshold (the nodes of the graph represent different behaviors, and therefore the nodes are deleted).

[0048] Furthermore, depending on the implementation method, the following annotations are generated, which can be used by the corresponding developer (e.g., control software 107) to understand error conditions and improve control. This achieves a shorter development cycle and a shorter time to market (TTM).

[0049] Depending on the implementation, the vehicle control unit 102 thus generates explanatory notes on the control decisions made by means of its modular processing chain. The vehicle control unit generates the explanatory notes based on input 204, intermediate results (perception results 205 and prediction results 206; the intermediate results are inputs to the following modules), and output (planning results 207), as well as a protocol (“log”) 215 describing the decisions made in the rule-based components of modules 201, 202, and 203.

[0050] Intermediate results may include interpretable representations (e.g., occupancy grids or object recognition) as well as latent features (which are forwarded between modules or sub-modules of the aforementioned modules) (e.g., feature maps extracted from camera images by the image processing module of perception module 201, which are then processed by a neural object recognition network).

[0051] Input 204, intermediate results 205, 206, and output 207 are encoded by a first encoder 208 to produce a wide range of descriptive features. Similarly, the protocol is encoded by a second encoder (text encoder) 209. The first encoder 208 and the second encoder 209 are implemented, for example, by an ML (machine learning) model (e.g., a neural network), which is trained, for example, along with or after the processing chain to produce annotations.

[0052] All the codes 210 thus generated are linked together to produce a common “decision process code” 211 for the decision process in the corresponding traffic scenario (the information is also contained in the decision process code 211 via the decision process code), the decision process generating the corresponding driving decision (i.e. the corresponding behavior or one or more control actions).

[0053] In addition, a series of possible annotation texts 212 (e.g., explanations of why a decision was made) are encoded into corresponding annotation codes 214 by means of a third encoder (text encoder) 213 (this can also be done in advance, for example in computer system 105, and the annotation codes 214 can be loaded into vehicle 100).

[0054] The control device 102 compares the common code 211 with the annotation code 214 and selects the annotation text whose annotation code 214 has the best consistency with the common code 211.

[0055] In order to generate corresponding and compatible codes, i.e., to train encoders 208, 209, 213 such that the encoders generate common codes 211 that are well consistent with the annotation codes 214 of the annotations adapted thereto for a specific control decision (i.e., the processing performed through the processing chain), a scheme similar to that used for generating codes for images, such as the scheme described in Reference 1, wherein the encoding of the images is consistent with the encoding of the text descriptions adapted thereto.

[0056] By comparing the common code 211 with each annotation code 214 (e.g., calculating the corresponding Euclidean distance between the codes respectively), a score (or evaluation, i.e., a "score") is associated with each annotation code 214, the score indicating how well the corresponding annotation (text description) fits the vehicle's decision in the traffic scenario. The control device 102 selects the most suitable annotation (i.e., the annotation with the highest score) and outputs it, for example, on a screen 110 at the dashboard of the vehicle 101. Thus, the annotation is selected based on the common code (i.e., the decision process code) 211.

[0057] Annotated text 212 can be pre-generated, or it can be generated based on the internal states and variables of the processing chain and / or input 204, intermediate results (perception results 205 and prediction results 206) and output (planning results 207), and optionally a protocol (“log”) 215, for example, through a (possibly pre-trained) LLM (Large Language Model). For this purpose, the annotated text is, for example, input into the inner layers of one or more deep neural networks (DNNs) implementing the processing chain and trained to learn the relevance of some parts of the decision policy to the embedding space of the LLM, providing a textual basis (or for selecting multiple actions) for the action chosen by the decision policy. The aim here is to provide rich textual terms (text “snippets”) and to provide the user with additional information that would not otherwise be available.

[0058] In general, the processing chain (which includes one or more trained neural networks) provides textual descriptions in addition to the planning result 206 (i.e., the control action), which annotate the decisions made when generating the planning result 206 (i.e., when selecting a control action or, for example, a trajectory). As mentioned above, generating the annotations can involve a large language model (LLM) that has been pre-trained to incorporate knowledge about the world. The LLM can be trained together with the control policy (e.g., through training on a traffic scenario labeled with control actions and annotations). In this way, the actions selected by the control policy can be interpreted.

[0059] Examples of control decisions and accompanying notes are as follows:

[0060] • Control decision: Turn five degrees to the left and accelerate by 1 m / s 2

[0061] Note: The reason for turning five degrees to the left is that the predicted road curvature is eight degrees, and additional mitigation measures should be used to avoid over-steering. Furthermore, since a sufficient distance from the vehicle ahead and a speed below the applicable speed limit should allow for acceleration.

[0062] • Control Decision: Braking

[0063] Note: The reason for braking is that the speed of the truck in front is lower than its own speed. Because the other lane is blocked by an oncoming black truck, overtaking is impossible, and the "exit lane" maneuver cannot be performed.

[0064] • Control decision: Driving at 10km / h

[0065] Note: The vehicle on the right is stopped and can overtake. The estimated clearance from the oncoming lane is wide enough for passage. Therefore, proceeding straight at a low speed is the preferred course of action.

[0066] While the above embodiments relate to autonomous driving, the processing methods described herein can also be applied to other fields, such as robotics, manufacturing, and so on. These processing methods can be applied to any application in which a robotic device is trained to perform actions requiring safety and interpretability in an environment. Therefore, the term "robotic device" can be understood to refer to any engineered system (with mechanically controlled components), such as computer-controlled machines, vehicles, household appliances, power tools, manufacturing machines, personal assistants, or access control systems.

[0067] In summary, a method is provided according to different implementations, such as... Figure 3 The method shown.

[0068] Figure 3 A flowchart 300 is shown, illustrating a method according to one embodiment for generating textual descriptions of decisions made automatically when controlling a robotic device (i.e., an autonomous robotic device, such as an autonomous vehicle).

[0069] In 301, data (particularly sensor data) containing information about the robot's environment (or "surrounding environment") is processed through a control processing chain having multiple (linked) modules (each with corresponding inputs and outputs). Modules are, for example, at least one perception module, a prediction module, and a planning module (or sub-modules thereof). In other words, the data is processed through a control pipeline that is at least partially modularized, where modules undertake different control tasks (i.e., decision-making processes for selecting control actions).

[0070] At least some modules output protocols (i.e., “logs”) regarding rule-based intermediate steps (intermediate decisions) performed by the respective modules during control. This includes, for example, outputting anomalies that occur here (i.e., in rule-based intermediate steps). Rule-based intermediate steps are non-ML (machine learning) based intermediate steps, i.e., “classical” intermediate steps, such as if-then-else operations, safety checks, graph (e.g., tree) operations, such as pruning or expanding nodes (the protocol, for example, provides insight into when a node (representing a specific action) is removed). The textual description generated for such graph operations would ultimately include, for example, “forced action was taken because the node used for light braking had been removed.”

[0071] In 302, the inputs (including intermediate results) of at least some modules in the module, the outputs of at least some modules in the module, and the protocol output by at least some modules in the module are encoded into a (common) decision process code. For example, codes are generated separately (for the inputs, protocol, each intermediate result, and output) and then the (separate) codes are combined into a decision process code as described above (e.g., simply concatenated together).

[0072] In 303, based on the decision process encoding, a text description is selected from a (preset) set of text descriptions (from a preset data set or generated online, for example by means of an LLM) for at least one decision made when processing data through a control processing chain.

[0073] When necessary, the robotic device is controlled based on the processing results (e.g., the planning results described above) obtained by processing data through a (control) processing chain. However, as described in the text, the user can disable this (automatic) control (e.g., control actions).

[0074] According to one implementation, the encoding is generated by one or more encoders (e.g., implemented via an ML model), which are trained together with a processing chain (which may also be implemented at least in part via an ML model) or after its training (e.g., by means of training examples containing annotations as ground truths, or also by human feedback on the generated annotations).

[0075] Based on this text description, users (or developers) can change the configuration of the processing chain when necessary, such as deciding whether a decision through the processing chain is reasonable.

[0076] Figure 3 The method can be executed by one or more computers having one or more data processing units. The term "data processing unit" can be understood as any type of entity that implements data or signal processing. For example, data or signals can be processed according to at least one (i.e., one or more) specific functions performed by the data processing unit. The data processing unit may include integrated circuits such as analog circuits, digital circuits, logic circuits, microprocessors, microcontrollers, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable gate arrays (FPGAs), or any combination thereof, or constituted by them. Any other manner in which the corresponding functions described in more detail herein are implemented can also be understood as a data processing unit or logic circuit device. One or more of the method steps described in detail herein can be executed (e.g., implemented) by a data processing unit through one or more specific functions performed by the data processing unit.

[0077] Therefore, depending on the implementation scheme, this method is particularly computer-implemented.

[0078] The detection of the corresponding control status (or control scenario, such as the environment of the robot equipment) can be based on sensor data from different sensors, such as video, radar, lidar, ultrasonic, motion, and thermal imaging sensors.

Claims

1. A method for generating a textual description of a decision made automatically when controlling a robotic device (101), the method having: processing (301) data containing information about an environment of the robotic device (101) by a control processing chain having a plurality of modules (201, 202, 203), wherein at least some of the modules (201, 202, 203) output protocols (215) about rule-based intermediate steps performed by the respective module when controlling; encoding inputs of at least some of the modules (201, 202, 203), outputs of at least some of the modules (201, 202, 203) and protocols output by at least some of the modules (201, 202, 203) into a decision process encoding (211) by an encoder (208, 209) in order to generate descriptive features of a decision process related to the automatically made decision; and selecting (303) from a set of textual descriptions (212) a textual description of at least one decision made when processing the data by the control processing chain according to the decision process encoding (211).

2. The method according to claim 1, having: selecting a textual description from a set of textual descriptions (212) by evaluating for each textual description of the set of textual descriptions (212) a consistency of the encoding (214) of the textual description (212) with the decision process encoding (211) and selecting the textual description of the textual descriptions (212) having the best evaluation.

3. The method according to claim 1 or 2, having: generating the set of textual descriptions (212) of textual descriptions by means of a generative model obtaining inputs containing internal states, variable values and / or intermediate results of the control processing chain.

4. The method according to any one of claims 1 to 3, wherein the encoder (208, 209) is implemented by a machine learning model.

5. The method according to any one of claims 1 to 4, having: displaying the generated textual description on a display (110) of the robotic device (101).

6. The method according to any one of claims 1 to 5, wherein the robotic device is an autonomous vehicle (101) controlled in a traffic scenario.

7. A control device (102) designed to perform the method according to any one of claims 1 to 6.

8. A computer program having instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 6.

9. A computer readable medium storing instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 6.