Model generation method, data acquisition method, and control program
By prioritizing datasets with appropriate reaction speeds and training a neural network control model to fit ground truth data, the method ensures that machine learning models for controlling moving objects achieve optimal reaction speeds, addressing the variability in driver response speeds.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-06-09
- Publication Date
- 2026-03-17
AI Technical Summary
Existing methods for training machine learning models for controlling moving objects, such as vehicles or drones, do not ensure that the models acquire the ability to perform control with appropriate reaction speeds due to variations in driver or operator reaction speeds in the collected learning data.
A method involving prioritizing datasets with appropriate reaction speeds for machine learning, ensuring that the control model is trained to match ground truth data, using a neural network to derive control commands that fit predetermined response conditions.
This approach increases the likelihood of obtaining a trained machine learning model capable of controlling moving objects with appropriate reaction speeds, enhancing the control model's ability to respond effectively to environmental events.
Smart Images

Figure 0007831410000003 
Figure 0007831410000004 
Figure 0007831410000005
Abstract
Description
Technical Field
[0001] The present disclosure relates to a model generation method, a data collection method, and a control program.
Background Art
[0002] Patent Document 1 proposes a driving support device configured to acquire information indicating a driving operation and information indicating a driving situation at the time of the driving operation, determine whether the driving situation is appropriate for learning based on the acquired information, and determine that the driving operation in the driving situation determined to be inappropriate is excluded from the learning target.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] An object of the present disclosure is to provide a technique for increasing the probability of obtaining a trained machine learning model that has acquired the ability to perform control of a moving body at an appropriate reaction speed, or a control technique of a moving body using the trained machine learning model obtained thereby.
Means for Solving the Problems
[0005] A model generation method according to a first aspect of this disclosure is performed by a computer. The model generation method includes acquiring a plurality of datasets, each consisting of a combination of training data showing the environment in which a mobile object moves in a time series and ground truth data showing control commands for the mobile object in the environment in a time series; and performing machine learning of a control model using the acquired plurality of datasets. Performing machine learning includes training the control model for each dataset so that the result of deriving control commands for the mobile object from the training data using the control model fits the ground truth data. Using the plurality of datasets includes prioritizing the use of datasets that are evaluated as appropriate because the response speed of the control commands to events shown by the ground truth data fits predetermined conditions. The control model may be composed of a neural network.
[0006] A data collection method relating to a second aspect of this disclosure is performed by a computer. The data collection method includes collecting a plurality of datasets, each consisting of a combination of training data showing the environment in which a mobile object moves in a time series and ground truth data showing control commands to the mobile object in the environment in a time series, and outputting the collected plurality of datasets for use in machine learning. The collection of the plurality of datasets includes prioritizing the collection of datasets that are evaluated as appropriate because the response rate of the control commands to events shown by the ground truth data meets predetermined conditions.
[0007] A control program according to a third aspect of this disclosure is a program that causes a computer to perform the following actions: acquire target data indicating the environment in which a target mobile object moves; derive control commands from the acquired target data using a trained control model; and control the movement of the target mobile object according to the result of deriving the control commands. The trained control model is a combination of training data showing the environment in which a training mobile object moves in chronological order and ground truth data showing control commands for the training mobile object in the environment in chronological order. The results are generated by performing machine learning using multiple datasets, each composed of a combination of these datasets. Performing the machine learning includes training the control model for each dataset so that the result of deriving the control command for the mobile body from the training data using the control model fits the ground truth data. Using the multiple datasets for the machine learning includes prioritizing the use of datasets that are evaluated as appropriate because the response rate of the control command to events indicated by the ground truth data fits predetermined conditions. [Effects of the Invention]
[0008] According to this disclosure, it is possible to provide a technique for increasing the probability of obtaining a trained machine learning model that has acquired the ability to perform control of a moving object with an appropriate reaction speed, or a technique for controlling a moving object using such a trained machine learning model. [Brief explanation of the drawing]
[0009] [Figure 1] Figure 1 schematically illustrates an example of a scenario in which this disclosure applies. [Figure 2] Figure 2 schematically shows an example of the hardware configuration of a model generation device according to an embodiment. [Figure 3] Figure 3 schematically shows an example of the hardware configuration of the control device according to the embodiment. [Figure 4]Figure 4 schematically shows an example of the software configuration of the model generation device according to the embodiment. [Figure 5] Figure 5 schematically shows an example of the software configuration of the control device according to the embodiment. [Figure 6] Figure 6 is a flowchart showing an example of a processing procedure related to machine learning of a control model by the model generation device according to the embodiment. [Figure 7A] Figure 7A schematically shows an example of an event according to the embodiment. [Figure 7B] Figure 7B schematically shows an example of a method for evaluating the reaction rate in the event shown in Figure 7A. [Figure 8A] Figure 8A schematically shows an example of an event according to the embodiment. [Figure 8B] Figure 8B schematically shows an example of a method for evaluating the reaction rate in the event shown in Figure 8A. [Figure 9] Figure 9 schematically shows an example of an event according to the embodiment. [Figure 10A] Figure 10A schematically shows an example of an event according to the embodiment. [Figure 10B] Figure 10B schematically shows an example of a method for evaluating the reaction rate in the event shown in Figure 10B. [Figure 11A] Figure 11A schematically shows an example in the embodiment where datasets whose reaction rates are evaluated as appropriate under predetermined conditions are used preferentially. [Figure 11B] Figure 11B schematically shows an example in the embodiment where datasets whose reaction rates are evaluated as appropriate under predetermined conditions are used preferentially. [Figure 12] Figure 12 is a flowchart showing an example of a processing procedure related to the control of the movement of a mobile body by a control device according to the embodiment. [Figure 13] Figure 13 schematically illustrates an example of another scenario to which this disclosure applies. [Figure 14] Figure 14 schematically shows an example of the hardware configuration of a data acquisition device of another form. [Figure 15] FIG. 15 schematically shows an example of the software configuration of a data collection device according to another form. [Figure 16] FIG. 16 is a flowchart showing an example of a processing procedure regarding data collection by a data collection device according to another form.
Mode for Carrying Out the Invention
[0010] According to the method proposed by Patent Document 1, it can be expected that an automatic driving model that has acquired the ability to execute only appropriate driving operations is generated by excluding driving operations in an inappropriate driving situation from the learning target. However, the inventors of the present invention have found that the conventional method has the following problems.
[0011] That is, assume a scenario where the ability to perform automatic driving of a vehicle is acquired by a model through machine learning. In this case, the ability acquired by the trained machine learning model depends on the learning data used for machine learning. The learning data can be collected from various drivers. At this time, the abilities of the drivers are not always constant. For example, there are drivers with a fast reaction (execution of driving operations) speed to events such as deceleration of the preceding vehicle, while there are also drivers with a slow reaction speed to such events. Therefore, the reaction speeds of the drivers represented in the collected learning data can be dispersed.
[0012] Regarding this point, in the conventional method, it only excludes driving operations determined to be inappropriate from the learning target. Differences in reaction speeds can also occur in appropriate driving operations. Therefore, with the conventional method, it was not always possible to obtain a trained machine learning model that had acquired the ability to perform automatic driving at an appropriate reaction speed.
[0013] Furthermore, this problem can occur regardless of the type of vehicle (e.g., number of wheels (two-wheeled, four-wheeled, etc.), size (large, regular, small, etc.), power source (electric, fuel, etc.)). Moreover, this problem is not limited to vehicle control. The same issues apply to other moving objects besides vehicles. Therefore, similar problems can arise when controlling any other moving object (e.g., flying objects (drones, etc.), ships, etc.).
[0014] In contrast, the model generation method according to the first aspect of this disclosure is performed by a computer. The model generation method includes acquiring a plurality of datasets, each consisting of a combination of training data showing the environment in which a mobile body moves in a time series and ground truth data showing control commands for the mobile body in the environment in a time series, and performing machine learning of a control model using the acquired plurality of datasets. Performing machine learning includes training the control model for each dataset so that the result of deriving control commands for the mobile body from the training data using the control model fits the ground truth data. Using the plurality of datasets includes prioritizing the use of datasets that are evaluated as appropriate because the response rate of the control commands to events shown by the ground truth data fits predetermined conditions.
[0015] The capabilities of a trained model generated by machine learning depend on the dataset used for the machine learning. In the first aspect of this disclosure, datasets with appropriate reaction speeds are given priority for machine learning. This increases the likelihood of obtaining a trained machine learning model that has acquired the ability to control a moving object with appropriate reaction speeds.
[0016] Furthermore, the data collection method according to the second aspect of this disclosure is performed by a computer. The data collection method includes collecting a plurality of datasets, each consisting of a combination of training data showing the environment in which a mobile object moves in a time series and ground truth data showing control commands to the mobile object in the environment in a time series, and outputting the collected plurality of datasets for use in machine learning. The collection of the plurality of datasets includes prioritizing the collection of datasets that are evaluated as appropriate because the response rate of the control commands to events shown by the ground truth data meets predetermined conditions.
[0017] In a second aspect of this disclosure, datasets for which the response rate is deemed appropriate are collected preferentially. By doing so, similar to the first embodiment described above, datasets in which the reaction rate is evaluated as appropriate can be given priority for use in machine learning. Therefore, the probability of obtaining a trained machine learning model that has acquired the ability to perform control of a moving object with an appropriate reaction rate can be increased.
[0018] Furthermore, a control program according to a third aspect of this disclosure is a program that causes a computer to perform the following actions: acquire target data indicating the environment in which a target mobile object moves; derive control commands from the acquired target data using a trained control model; and control the operation of the target mobile object according to the result of deriving the control commands. The trained control model is generated by performing machine learning using a plurality of datasets, each consisting of a combination of training data showing the environment in which a training mobile object moves in time series and ground truth data showing control commands for the training mobile object in time series in the environment. Performing the machine learning includes training the control model for each dataset so that the result of deriving control commands for the mobile object from the training data using the control model fits the ground truth data. Using the plurality of datasets for the machine learning includes prioritizing the use of datasets that are evaluated as appropriate because the response speed of the control commands to events indicated by the ground truth data fits predetermined conditions.
[0019] As described in each of the above embodiments, by prioritizing the use of datasets with appropriately evaluated reaction speeds in machine learning, it is possible to obtain a trained machine learning model that has acquired the ability to perform control of a moving object with an appropriate reaction speed. According to the third embodiment of this disclosure, it can be expected that by using such a trained control model (machine learning model), it will be possible to perform control of a moving object with an appropriate reaction speed.
[0020] Hereinafter, embodiments relating to one aspect of this disclosure (hereinafter also referred to as "this embodiment") will be described based on the drawings. However, this embodiment described below is merely illustrative in all respects of this disclosure. Various improvements or modifications may be made without departing from the scope of this disclosure. In implementing this disclosure, specific configurations may be adopted as appropriate depending on the embodiment. Although the data appearing in this embodiment is described in natural language, more specifically, it is specified in computer-recognizable pseudo-language, commands, parameters, machine code, etc.
[0021] [1. Application Examples] Figure 1 schematically shows an example of a scenario in which this disclosure is applied. The system according to this embodiment comprises a model generation device 1 and a control device 2.
[0022] The model generation device 1 according to this embodiment is one or more computers configured to generate a trained control model 5 by performing machine learning. In this embodiment, the model generation device 1 acquires a plurality of datasets 4, each composed of a combination of training data 41 and ground truth data 45. The training data 41 is configured to show the environment in which a mobile object (training mobile object) moves in a time series. The ground truth data 45 is configured to show the true values of control commands to the mobile object in the environment shown by the corresponding training data 41 in a time series.
[0023] The model generation device 1 performs machine learning on the control model 5 using the acquired multiple datasets 4. Performing this machine learning involves training the control model 5 so that, for each dataset 4, the result of deriving the control command for the mobile object from the training data 41 using the control model 5 fits the corresponding ground truth data 45. In this machine learning, the use of multiple datasets 4 is indicated by the ground truth data 45. This includes prioritizing the use of datasets in which the response rate to events of the control commands is evaluated as appropriate when the response rate meets predetermined conditions. This machine learning makes it possible to generate a trained control model 5 that has acquired the ability to derive control commands according to the environment in which the moving object is moving.
[0024] On the other hand, the control device 2 in this embodiment is one or more computers configured to control the movement of the target mobile object M using a trained control model 5. In this embodiment, the control device 2 acquires target data 221 that represents the environment in which the target mobile object M moves. The control device 2 uses the trained control model 5 to derive control commands from the acquired target data 221. Then, the control device 2 controls the movement of the target mobile object M according to the result of deriving the control commands.
[0025] As described above, in this embodiment, the model generation device 1 prioritizes the use of datasets with appropriate reaction speeds for machine learning, thereby generating a trained control model 5. Since the capabilities of the trained model generated by machine learning depend on the dataset used for the machine learning, according to this embodiment, it is possible to obtain a trained control model 5 that has acquired the ability to perform control of a moving object with an appropriate reaction speed. Furthermore, in the control device 2 according to this embodiment, it is possible to perform control of the target moving object M with an appropriate reaction speed by using such a trained control model 5.
[0026] (Mobile) The type of mobile body (mobile body M) is not particularly limited, as long as it can be moved automatically by mechanical control, and may be appropriately selected depending on the embodiment. The mobile body (mobile body M) may be any mobile device such as a vehicle, an aircraft, a ship, or a robotic device. The aircraft may be at least one of an unmanned aircraft such as a drone and a manned aircraft.
[0027] In one example, as shown in Figure 1, the moving object (moving object M) may be a vehicle. In this case, the model generation device 1 can be expected to acquire a trained control model 5 that has acquired the ability to perform vehicle control with an appropriate reaction speed. The control device 2 can also be expected to be able to perform control of the target vehicle with an appropriate reaction speed by using such a trained control model 5.
[0028] If the mobile entity is a vehicle, the type of vehicle may be selected arbitrarily. Vehicles may be selected from, for example, two-wheeled vehicles, three-wheeled vehicles, four-wheeled vehicles, etc. The power source of the vehicle may be selected from, for example, electricity, fuel, etc. If the vehicle is an automobile, the size of the vehicle may be selected from large, medium, semi-medium, regular, large special, small special, etc. If the vehicle is a two-wheeled vehicle, the size of the vehicle may be selected from large, regular, etc. As a typical example, the mobile entity (mobile entity M) may be an automobile with Level 2 or higher autonomous driving capability.
[0029] (Environment / Sensors) The environment is an event observed in at least one of the mobile object itself and its surroundings. In one example, at least a portion of the environment may be observed by one or more sensors S placed inside or outside the mobile object (mobile object M). Accordingly, the training data 41 and the target data 221 may each include sensor data SD obtained by one or more sensors S.
[0030] The type of sensor S is not particularly limited as long as it can observe any environment in which a moving object is moving, and may be appropriately selected depending on the embodiment. For example, one or more sensors S may include a camera (image sensor), radar, LiDAR (Light Detection and Ranging), sonar (ultrasonic sensor), infrared sensor, GNSS (Global Navigation Satellite System) / GPS (Global Positioning Satellite) module, etc.
[0031] (Control command) The control commands relate to the operation of the moving object. The configuration of the control commands may be determined as appropriate depending on the embodiment. For example, the control commands may consist of acceleration, deceleration, steering, or a combination thereof. Acceleration and deceleration may include gear changes. In this case, the model generation device 1 can be expected to acquire a trained control model 5 that has acquired the ability to perform control of acceleration, deceleration, steering, or a combination thereof with an appropriate reaction speed. Furthermore, the control device 2 can be expected to be able to perform control of acceleration, deceleration, steering, or a combination thereof with an appropriate reaction speed by using such a trained control model 5.
[0032] In one example, if the moving object (moving object M) is a vehicle, the control command may include at least one of acceleration, deceleration, and steering of the vehicle. If it includes at least one of acceleration, deceleration, and steering, the control command may be expressed as a path. Accordingly, the control model 5 may be expressed as a path planner.
[0033] Furthermore, the control commands may include commands relating to the operation of the mobile object. For example, if the mobile object (mobile object M) is a vehicle, the control commands may include vehicle operations such as turning on the turn signals, hazard lights, horn, and communication processing (e.g., sending data to a center, making an emergency call, etc.).
[0034] (Dataset) Each dataset 4 may be generated as appropriate. Each dataset 4 may be generated automatically by computer operation, or it may be generated manually, at least partially by operator operation. Typically, operation data and environmental data of a mobile body may be collected while a subject controls the mobile body entirely manually. Environmental data (observation data) may be obtained by sensors S mounted on the mobile body. Operation data may be obtained by recording manual operations by the subject. The training data 41 for each dataset 4 may be generated from the environmental data. The ground truth data 45 may be generated from the operation data. That is, typically, dataset 4 may be generated from the results of manual operations by a subject (e.g., manual driving of a vehicle). When obtaining dataset 4 using a mobile body, the mobile body from which dataset 4 is obtained may include the mobile body M that uses the generated trained control model 5, or it may not include the mobile body M. That is, the training mobile body may include the target mobile body M, or it may not include the target mobile body M. The training mobile body may include a mobile body used by an end user. In this case, the end user may be a subject. Furthermore, the training vehicles may include vehicles used experimentally.
[0035] However, the method for generating dataset 4 is not limited to this example and may be appropriately selected depending on the embodiment. In another example, dataset 4 may be generated from the results of manual operations by a subject, similar to the example above, but manual operations may include operations by the subject during the operation of partial automatic control, such as override operations on arbitrary automatic control. In another example, at least a portion of dataset 4 may be obtained by a virtual method such as simulation. In another example, at least a portion of dataset 4 may be obtained by a reinforcement learning framework. Furthermore, in another example, at least a portion of dataset 4 may be obtained by data augmentation of an arbitrary dataset. Data augmentation is performed by generating new training data by changing the attribute values of the training data. For example, if the training data includes images, the parameter changes may consist of image processing such as translation, scaling, rotation, and noise application to the images. When at least a portion of the dataset 4 is obtained by data augmentation, the reaction speed and reaction speed-dependent parameters are applied to the training data of any dataset (the original dataset). One or more new datasets may be generated by changing the values of attributes other than at least one of the attributes and assigning corresponding ground truth data. Multiple datasets 4 may include the newly generated datasets.
[0036] (Control model) The control model 5 is comprised of a machine learning model having one or more computational parameters that can be adjusted by machine learning. One or more computational parameters are used for the calculation of the desired inference (in this case, the derivation of control commands). Machine learning involves adjusting (optimizing) the values of the computational parameters using training data (in this case, multiple datasets 4). The configuration and type of the machine learning model are not particularly limited and may be appropriately selected depending on the embodiment. The machine learning model may consist of, for example, a neural network, a support vector machine, a regression model, a decision tree model, etc.
[0037] As an example, control model 5 may be composed of a neural network. The structure of the neural network may be determined as appropriate depending on the embodiment. The structure of the neural network may be specified, for example, by the number of layers from the input layer to the output layer, the type of each layer, the number of nodes (neurons) contained in each layer, the connection relationships between the nodes in each layer, etc. In one example, the neural network may have a recursive structure. Furthermore, the neural network may include any layers such as fully connected layers, convolutional layers, pooling layers, deconvolutional layers, unpooling layers, normalization layers, dropout layers, LSTM (Long short-term memory), etc. The neural network may also have any mechanisms such as an attention mechanism. Model 5 (neural network) may include any model such as a GNN (Graph neural network), a diffusion model, or a generative model (e.g., Generative Adversarial Network, Transformer, etc.). When a neural network is used as the control model 5, the weights of the connections between each node included in the control model 5 (neural network) and the thresholds of each node are examples of computational parameters.
[0038] The input / output configuration of the control model 5 is not particularly limited and can be appropriately selected depending on the embodiment, as long as control commands can be derived from the environment of the moving object. For example, the control model 5 may be configured to derive control commands for one or more time points from environmental data for one or more time points. Also, the control model 5 may be configured to accept time-series data depending on its structure. As one example, the control model 5 may be configured to accept time-series data by being configured in a recursive manner. As another example, the control model 5 may be configured to take environmental data for multiple time points as input at once. Alternatively, the control model 5 may be configured structurally to be unable to accept time-series data. For example, the control model 5 may be configured to derive a control command for one time point from environmental data for one time point. In this case, the control model 5 may be used to obtain calculation results for time-series data by sequentially receiving data for each time point in the time-series data and sequentially outputting the calculation results. Furthermore, the control model 5 may be configured to derive control commands immediately. Alternatively, the control model 5 may be configured to derive control commands for multiple future time points at once. In this case, at least a portion of the control commands derived collectively may be used to control the mobile body (mobile body M).
[0039] The processing performed by control model 5 is not particularly limited, and can be appropriately selected depending on the embodiment, as long as it is involved in at least a part of the inference process that derives control commands from the environment of the moving object. For example, control model 5 may be configured to perform surrounding recognition and path planning (route / trajectory planning). Control model 5 may also be configured to perform motion planning (action / control planning). In other words, control model 5 may be an end-to-end model.
[0040] Furthermore, if the movement of the mobile object can be controlled by the output of control model 5, then the output of control model 5 The force type may be appropriately selected depending on the embodiment. The control model 5 may be configured to directly output control commands. Alternatively, control commands may be obtained by performing arbitrary information processing (interpretation processing) on the output of the control model 5. The control commands may be configured to directly indicate control quantities (control instruction values, control output quantities) of a moving body, such as accelerator control quantity, brake control quantity, and steering angle. Alternatively, the control commands may be configured to indirectly indicate control quantities of a moving body, such as path and post-control state. In this case, control quantities of a moving body may be obtained from the control commands by performing arbitrary information processing. For example, if the moving body is a vehicle, control quantities of the vehicle may be obtained by applying the inference results obtained from the control model 5 to a vehicle model. The vehicle model may have various parameters such as accelerator, brake, and steering, and may be appropriately configured to derive control quantities from indirect information (path, post-control state, etc.).
[0041] Furthermore, the ground truth data 45 for each dataset 4 may be configured as appropriate to directly or indirectly indicate the control commands. If the control commands are derived by performing arbitrary arithmetic operations from the output of the control model 5, the ground truth data 45 may be provided for the control commands derived from the output of the control model 5 (i.e., the ground truth data 45 may be configured to directly indicate the control commands). Alternatively, the ground truth data 45 may be provided for the output of the control model 5 (i.e., the ground truth data 45 may be configured to indirectly indicate the control commands).
[0042] (event) An event may include any occurrence that may be involved in the operation of a moving object. Furthermore, an event may include any occurrence that can be detected by sensor S. Detection by sensor S means determination based on sensor values. That is, the start time of the event may be determined by sensor data. The detection (determination, identification) method may be determined appropriately depending on the event. Sensor data may be analyzed in any way, thereby determining the start time of the event (i.e., detecting the occurrence of the event).
[0043] Accordingly, the mobile body (mobile body M) may be equipped with a sensor S. The training data 41 may include sensor data SD obtained by the sensor S. The start time of events in the training data 41 may be identified by the sensor data SD. This makes it possible to mechanically evaluate the reaction speed in each dataset 4, thereby improving the efficiency of identifying the dataset 4 to be used preferentially in machine learning. In other words, the process of selecting the dataset 4 to be used preferentially can be automated at least partially, thereby reducing the amount of work required. The start time of the operation in response to an event is shown in the ground truth data 45. Therefore, in this configuration, the reaction speed to an event (i.e., the time from the time of event occurrence to the start time of operation) can be identified from the training data 41 and the ground truth data 45.
[0044] (Example of an event) For example, if the moving object (moving object M) is a vehicle, the events may include at least one of the following: deceleration of a preceding vehicle relative to the vehicle, a vehicle cutting in alongside, the appearance of a parked or stopped vehicle, the appearance of an obstacle, and a change in traffic signals. Obstacles may include any object that could obstruct the vehicle's movement. Obstacles may be, for example, pedestrians, bicycles, etc. In this case, the model generation device 1 can be expected to acquire a trained control model 5 that has acquired the ability to perform vehicle control with an appropriate reaction speed in response to at least one of these events. Furthermore, the control device 2 can be expected to use such a trained control model 5 to perform control of the target vehicle with an appropriate reaction speed in response to at least one of these events.
[0045] Furthermore, if the events targeted for autonomous driving (automatic control) include the deceleration of the preceding vehicle, The control command may include a deceleration command in response to the preceding vehicle. If the target event includes a cut-in by a parallel vehicle, the control command may include a deceleration command in response to the parallel vehicle. If the target event includes the occurrence of a parked or stopped vehicle, the control command may include at least one of deceleration and steering in response to the parked or stopped vehicle. If the target event includes the occurrence of an obstacle, the control command may include at least one of deceleration and steering in response to the obstacle. If the target event includes a change in traffic signals, the control command may include an acceleration or deceleration command in response to the traffic signals. These events may also be adapted for other types of moving objects (e.g., aircraft, ships, etc.).
[0046] (Meets the specified conditions) The specified conditions may be defined as appropriate to allow for the evaluation of an appropriate response rate depending on the event. For example, the specified conditions may be defined so that a faster response rate is considered more appropriate. In this case, datasets with faster response rates to events (evaluated as having an appropriate response rate) may be given priority for machine learning. However, the specified conditions are not limited to this example. In another example, the specified conditions may define a range (upper and lower limits) of appropriate response rates. In this case, datasets whose response rates fall within the range defined by the specified conditions may be given priority for machine learning, while other datasets (i.e., datasets with response rates faster or slower than the range defined by the specified conditions) may not be given priority for machine learning.
[0047] (Priority use) Prioritizing the use of datasets with appropriately evaluated reaction rates may be configured in any way that the preferred datasets are more readily reflected in the training of the control model 5 than the non-preferred datasets.
[0048] For example, one can simply decide whether or not to use a dataset based on whether or not it should be prioritized. That is, prioritizing the use of datasets in which the reaction rate is evaluated as appropriate can be achieved by using datasets in which the reaction rate is evaluated as appropriate for training control model 5, and not using datasets in which the reaction rate is not evaluated as appropriate for training control model 5. This method makes it extremely easy to incorporate the evaluation of reaction rate into machine learning.
[0049] As another example, a dataset from among multiple datasets 4 that is evaluated as having an appropriate reaction rate is designated as the first dataset, and a dataset that is not evaluated as having an appropriate reaction rate (i.e., does not meet the specified conditions) is designated as the second dataset. Prioritizing the use of datasets that are evaluated as having an appropriate reaction rate can be achieved by setting a higher sampling probability for the first dataset in machine learning and a lower sampling probability for the second dataset in machine learning than for the first dataset.
[0050] In this case, prioritizing the use of datasets with appropriately assessed reaction rates may be further constructed by excluding at least a portion of the second dataset from machine learning (i.e., setting the sampling probability to 0). Alternatively, prioritizing the use of datasets with appropriately assessed reaction rates may be further constructed by not excluding the second dataset from machine learning (i.e., not setting the sampling probability to 0). In real-world situations, due to external factors, slow-response operations may be performed as appropriate operations. For example, sudden operations such as sudden braking or sudden steering are considered to be operations with slow response rates. Also, a long period between the detection of an event and the need for an operation may make it appear as if the reaction rate is slow. In this respect, by not excluding the second dataset from machine learning, these operations can also be taught to the control model 5. Therefore, an improvement in the robustness of operations to events can be expected.
[0051] As yet another example, prioritizing the use of datasets that are evaluated as having an appropriate reaction speed may be achieved by increasing the training weights of the first dataset among multiple datasets 4 and decreasing the training weights of the second dataset. For example, increasing the training weights may be achieved by increasing the learning rate, and decreasing the training weights may be achieved by decreasing the learning rate. In other words, prioritizing the use of datasets that are evaluated as having an appropriate reaction speed may be achieved by making the amount of parameter updates of the control model 5 in a single training session larger for the preferred dataset. As yet another example, prioritizing use may include both the sampling probability and weight methods described above.
[0052] Priority levels for use may be set as appropriate. For example, there may be two priority levels (i.e., prioritize / do not prioritize). In another example, there may be three or more priority levels. In this case, the priority levels may differ among the datasets that are prioritized (i.e., there may be superiority or inferiority). Similarly, the priority levels may differ among the datasets that are not prioritized.
[0053] (Controlling actions) In one example, controlling the movement of the target mobile object M may be done by directly controlling the target mobile object M. In another example, the mobile object (mobile object M) may be equipped with a dedicated control device, such as a controller. In this case, controlling the movement of the target mobile object M by the control device 2 may be done by indirectly controlling the target mobile object M by providing the dedicated control device with a derived result.
[0054] (System Configuration) In one example, as shown in Figure 1, the model generation device 1 and the control device 2 may be configured to communicate (connect) with each other via a network. The type of network is not particularly limited and may be appropriately selected from, for example, the Internet, wireless communication network, mobile communication network, telephone network, dedicated network, etc. However, the method of exchanging data between the model generation device 1 and the control device 2 is not limited to this example and may be appropriately selected depending on the embodiment. In another example, data may be exchanged using a storage medium.
[0055] Furthermore, in the example shown in Figure 1, the model generation device 1 and the control device 2 are separate computers. However, the system configuration is not limited to this example. In another example, the model generation device 1 and the control device 2 may be configured as a single computer. Also, at least one of the model generation device 1 and the control device 2 may be configured as multiple computers.
[0056] Furthermore, in the example shown in Figure 1, the control device 2 is mounted inside the mobile body M. However, the placement of the control device 2 is not limited to this example. The control device 2 may be placed outside the mobile body M, as long as it can directly or indirectly control the operation of the mobile body M.
[0057] [2 Example Configurations] [Example Hardware Configuration] <Model Generator> Figure 2 schematically shows an example of the hardware configuration of the model generation device 1 according to this embodiment. As shown in Figure 2, the model generation device 1 according to this embodiment is a computer in which a control unit 11, a storage unit 12, a communication interface 13, an input device 14, an output device 15, and a drive 16 are electrically connected.
[0058] The control unit 11 is a hardware processor, a CPU (Central Processing Unit), It includes RAM (Random Access Memory), ROM (Read Only Memory), etc., and is configured to perform information processing based on programs and various data. The control unit 11 (CPU) is an example of processor resources.
[0059] The storage unit 12 may be composed of, for example, a hard disk drive, a solid-state drive, etc. The storage unit 12 (and RAM, ROM) is an example of memory resources. In this embodiment, the storage unit 12 stores various information such as a model generation program 81, multiple datasets 4, and training result data 125.
[0060] The model generation program 81 is a program that causes the model generation device 1 to perform information processing related to machine learning of the control model 5 (Figure 6, described later). The model generation program 81 includes a series of instructions for said information processing. The learning result data 125 is configured to show information about the generated trained control model 5. In this embodiment, the learning result data 125 is generated as a result of executing the model generation program 81.
[0061] The communication interface 13 is an interface for wired or wireless communication over a network. The communication interface 13 may consist of, for example, a wired LAN (Local Area Network) module, a wireless LAN module, etc. The model generation device 1 may perform data communication with another computer (for example, the control device 2) via the communication interface 13.
[0062] The input device 14 is, for example, a device for inputting data such as a mouse or keyboard. The output device 15 is, for example, a device for outputting data such as a display or speaker. The operator can operate the model generation device 1 by using the input device 14 and the output device 15. The input device 14 and the output device 15 may be integrated together, for example, by a touch panel display.
[0063] Drive 16 is a device for reading various information, such as programs, stored in the storage medium 91. At least one of the model generation program 81, the multiple datasets 4, and the learning result data 125 may be stored in the storage medium 91 instead of or together with the storage unit 12. The storage medium 91 is configured to store various information (stored programs, etc.) by electrical, magnetic, optical, mechanical, or chemical means so that a machine such as a computer can read the information. The model generation device 1 may obtain at least one of the model generation program 81 and the multiple datasets 4 from the storage medium 91.
[0064] In Figure 2, a disk-type storage medium such as a CD or DVD is shown as an example of a storage medium 91. However, the type of storage medium 91 is not limited to disk type. Other storage media include, for example, semiconductor memory such as flash memory. The type of drive 16 may be appropriately selected according to the type of storage medium 91.
[0065] Furthermore, regarding the specific hardware configuration of the model generation device 1, components can be omitted, replaced, and added as appropriate depending on the embodiment. For example, the control unit 11 may include multiple hardware processors. Hardware processors include microprocessors, FPGAs (field-programmable gate arrays), DSPs (digital signal processors), It is composed of ECU (Electronic Control Unit), GPU (Graphics Processing Unit), etc. This may be done. At least one of the communication interface 13, input device 14, output device 15, and drive 16 may be omitted. The model generation device 1 may consist of multiple computers. In this case, the hardware configuration of each computer is identical. It is acceptable, or it is not necessary, for them to match. Model generation device 1 may be a computer designed specifically for the service provided, as well as a general-purpose server device, a general-purpose PC (Personal Computer). ), industrial PCs, terminal devices (e.g., tablet PCs, etc.) may also be used.
[0066] <Control device> Figure 3 schematically shows an example of the hardware configuration of the control device 2 according to this embodiment. As shown in Figure 3, the control device 2 according to this embodiment is a computer in which a control unit 21, a storage unit 22, a communication interface 23, an input device 24, an output device 25, a drive 26, and an external interface 27 are electrically connected.
[0067] The control units 21 to 26 and the storage medium 92 of the control device 2 may be configured in the same way as the control units 11 to 16 and the storage medium 91 of the model generation device 1. The control unit 21 (CPU) is an example of the processor resources of the control device 2, and the storage unit 22 (and RAM, ROM) is an example of the memory resources of the control device 2. In this embodiment, the storage unit 22 stores various information such as the control program 82 and the learning result data 125.
[0068] The control program 82 is a program that causes the control device 2 to execute information processing (Figure 12, described later) related to the automatic control of the target mobile object M by the trained control model 5. The control program 82 includes a series of instructions for said information processing. At least one of the control program 82 and the learning result data 125 may be stored in the storage medium 92 instead of or together with the storage unit 22. The control device 2 may retrieve at least one of the control program 82 and the learning result data 125 from the storage medium 92.
[0069] The control device 2 may communicate data with another computer (for example, the model generation device 1) via the communication interface 23. The operator can operate the control device 2 using the input device 24 and the output device 25. The input device 24 and the output device 25 may be integrated into a single unit, such as a touch panel display.
[0070] The external interface 27 is an interface for connecting to an external device. The external interface 27 may be, for example, a USB (Universal Serial Bus) port, a dedicated port, etc. The type and number of external interfaces 27 may be appropriately determined according to the type and number of external devices to be connected. In this embodiment, the control device 2 may be connected to the sensor S via the external interface 27. At least a portion of the target data 221 may consist of sensor data obtained by the sensor S. Note that the method of connecting the sensor S is not limited to this example. In another example, the sensor S may be connected via the communication interface 23.
[0071] Regarding the specific hardware configuration of the control device 2, components can be omitted, replaced, and added as appropriate depending on the embodiment. For example, the control unit 21 may include multiple hardware processors. Hardware processors may consist of microprocessors, FPGAs, DSPs, ECUs, GPUs, etc. At least one of the communication interface 23, input device 24, output device 25, drive 26, and external interface 27 may be omitted. The control device 2 may consist of multiple computers. In this case, the hardware configurations of each computer may or may not be the same. The control device 2 may be a computer designed specifically for the service provided, a general-purpose computer, a mobile phone including a smartphone, a tablet PC (Personal Computer), etc. If the mobile body M is a vehicle, the control device 2 may be an in-vehicle device.
[0072] [Example Software Configuration] <Model Generator> Figure 4 schematically shows an example of the software configuration of the model generation device 1 according to this embodiment. The control unit 11 of the model generation device 1 loads the model generation program 81 stored in the memory unit 12 into RAM, and the CPU executes the instructions contained in the model generation program 81. As a result, the model generation device 1 operates as a computer equipped with a learning data acquisition unit 111, a learning processing unit 112, and a storage processing unit 113 as software modules. In other words, in this embodiment, each software module of the model generation device 1 is realized by the control unit 11 (CPU).
[0073] The learning data acquisition unit 111 is configured to acquire multiple datasets 4, each composed of a combination of training data 41 and ground truth data 45. The learning processing unit 112 is configured to perform machine learning of the control model 5 using the acquired datasets 4. In this embodiment, performing machine learning includes training the control model 5 for each dataset 4 so that the result of deriving control commands for a mobile object from the training data 41 using the control model 5 fits the corresponding ground truth data 45. Furthermore, using multiple datasets 4 in this machine learning includes prioritizing the use of datasets in which the response speed to events of the control commands indicated by the ground truth data 45 is evaluated as appropriate because it meets predetermined conditions. By performing this machine learning, a trained control model 5 is generated.
[0074] The storage processing unit 113 is configured to store the trained control model 5 generated by machine learning. In one example, the storage processing unit 113 may be configured to generate learning result data 125 that shows the trained control model 5 generated as a result of machine learning. The structure of the learning result data 125 is not particularly limited and may be determined as appropriate depending on the embodiment, as long as it can hold information for performing calculations on the trained control model 5. For example, the learning result data 125 may be configured to include information showing the values of calculation parameters adjusted by machine learning. In some cases, the learning result data 125 may be configured to include information showing the structure of the trained control model 5 (e.g., the structure of the neural network). The storage processing unit 113 may be configured to store the generated learning result data 125 in a predetermined memory area. The learning result data 125 may be provided to the control device 2 at any time.
[0075] <Control device> Figure 5 schematically shows an example of the software configuration of the control device 2 according to this embodiment. The control unit 21 of the control device 2 loads the control program 82 stored in the storage unit 22 into RAM, and the CPU executes the instructions contained in the control program 82. As a result, as shown in Figure 5, the control device 2 according to this embodiment operates as a computer equipped with an acquisition unit 211, an output unit 212, and an operation control unit 213 as software modules. That is, in this embodiment, similar to the model generation device 1, each software module of the control device 2 is also realized by the control unit 21 (CPU).
[0076] The acquisition unit 211 is configured to acquire target data 221 that indicates the environment in which the target mobile object M moves. The derivation unit 212 holds the learning result data 125 and includes a trained control model 5 generated by the model generation device 1. The derivation unit 212 is configured to derive control commands from the acquired target data 221 using the trained control model 5. The operation control unit 213 is configured to control the operation of the target mobile object M according to the result of deriving the control commands (i.e., the control commands derived by the trained control model 5).
[0077] <Other> In this embodiment, an example is described in which each software module of the model generation device 1 and the control device 2 is implemented by a general-purpose CPU. However, some or all of the above software modules may be implemented by one or more dedicated processors. Each of the above modules may also be implemented as a hardware module. Regarding the software configuration of the model generation device 1 and the control device 2, modules may be omitted, replaced, and added as appropriate, depending on the embodiment.
[0078] [3 Examples of operation] [Model Generator] Figure 6 is a flowchart showing an example of a processing procedure related to machine learning of the control model 5 by the model generation device 1 according to this embodiment. The following processing procedure is an example of a model generation method executed by a computer. However, the following processing procedure of the model generation device 1 is merely an example, and each step may be modified as much as possible. Furthermore, depending on the embodiment, steps in the following processing procedure can be omitted, replaced, and added as appropriate.
[0079] <Step S101> In step S101, the control unit 11 operates as a training data acquisition unit 111. That is, the control unit 11 acquires multiple datasets 4, each composed of a combination of training data 41 and ground truth data 45.
[0080] The training data 41 is configured to show the environment in which the mobile object moves in a time series. In one example, the training data 41 may include sensor data SD obtained by sensor S. In addition, the training data 41 may include any information that can be involved in control, such as set speed, speed limit, map information, and navigation information. The ground truth data 45 is configured to show, directly or indirectly, control commands to the mobile object in the environment shown by the corresponding training data 41 in a time series.
[0081] As described above, each dataset 4 may be generated (collected) as appropriate. Each generated dataset 4 may be stored in the model generation device 1 (at least one of the memory unit 12 and the storage medium 91). Alternatively, each dataset 4 may be stored on another computer such as a network server (e.g., NAS: Network Attached Storage). In this case, when performing machine learning, the control unit 21 may acquire each dataset 4 via the network, external storage device, storage medium 91, etc. Each dataset 4 may be stored in database format.
[0082] The generation of at least a portion of the multiple datasets 4 may be performed by the model generation device 1. The generation of at least a portion of the multiple datasets 4 may be performed by a computer other than the model generation device 1. If the datasets 4 are generated by another computer, the control unit 11 may acquire the datasets 4 generated by the other computer via, for example, a network, an external storage device, a storage medium 91, etc.
[0083] The number of datasets 4 to be acquired may be determined as appropriate depending on the embodiment. Once multiple datasets 4 have been acquired, the control unit 11 proceeds to the next step S102.
[0084] <Step S102> In step S102, the control unit 11 operates as a learning processing unit 112. That is, the control unit 11 performs machine learning on the control model 5 using the acquired datasets 4. In machine learning, the control unit 11 trains the control model 5 for each dataset 4 so that the result of deriving control commands for the mobile object from the training data 41 using the control model 5 fits the corresponding ground truth data 45.
[0085] The control model 5 (machine learning model) comprises one or more computational parameters for performing computational processing to solve the inference task. Training the control model 5 involves optimizing (adjusting) the values of the computational parameters of the control model 5 (machine learning model) according to the given training data (multiple datasets 4). The machine learning method may be appropriately determined according to the embodiment of the machine learning model used for the control model 5, such as the type and structure of the machine learning model. Any method may be employed for adjusting the computational parameters, such as backpropagation or solving an optimization problem.
[0086] (An example of a machine learning method) As a typical example, if the control model 5 is composed of a neural network, the control unit 11 first prepares the control model 5 to be processed by machine learning. The structure of the control model 5 to be prepared, the initial values of the weights of the connections between each neuron, and the initial values of the thresholds of each neuron may be given by a template or by operator input. When retraining is performed, the control unit 11 may prepare the control model 5 based on the learning result data obtained from past machine learning. Next, the control unit 11 uses the training data 41 of each dataset 4 as input data and the ground truth data 45 as teacher signals (labels) to perform the learning process (supervised learning) of the control model 5.
[0087] As shown in Figure 4, as an example of the learning process, in the first step, the control unit 11 receives training data 41 from each dataset 4 and performs forward propagation calculations for the control model 5. As a result of this calculation, the control unit 11 obtains output values from the control model 5 that correspond to the inference results (direct or indirect derivation results of control commands) for the training data 41. In the second step, the control unit 11 calculates the error between the obtained output values and the corresponding ground truth data 45. In the third step, the control unit 11 calculates the gradient of the calculated error. Then, the control unit 11 calculates the error in the values of the calculation parameters of the control model 5 (weights of connections between each node, thresholds for each node, etc.) by backpropagating the calculated error gradient using the backpropagation method. In the fourth step, the control unit 11 updates the values of the calculation parameters based on the calculated error. The extent to which the values of the calculation parameters are updated may be adjusted by the learning rate.
[0088] The control unit 11 repeats the first to fourth steps described above to adjust the values of the calculation parameters of the control model 5 for each dataset 4 so that the sum of errors between the output values output from the control model 5 and the ground truth data 45 becomes smaller. This adjustment of the calculation parameter values may be repeated until predetermined conditions are met, such as adjusting the set number of iterations or the calculated sum of errors falling below a threshold. The threshold may be set appropriately depending on the embodiment. In addition, the machine learning conditions such as the objective function (cost function, loss function, error function), learning rate, and optimization algorithm for calculating the errors may be set appropriately depending on the embodiment.
[0089] The adjustment of the calculation parameters of this control model 5 may be performed on a mini-batch. For example, before executing the processing of the first to fourth steps described above, the control unit 11 may generate a mini-batch by extracting arbitrary samples (datasets) from multiple datasets 4. The size of the mini-batch may be set appropriately depending on the embodiment. The control unit 11 may then perform the processing of the first to fourth steps described above on the datasets 4 included in the generated mini-batch. If the first to fourth steps are repeated, the control unit 11 may generate a mini-batch again and perform the processing of the first to fourth steps described above on the newly generated mini-batch.
[0090] Furthermore, machine learning methods are not limited to supervised learning examples like this; other methods are also applicable. The method may be adopted at least partially. As another example, deep reinforcement learning may be adopted. In this case, some of the multiple datasets 4 may be obtained as the results of episodes in reinforcement learning. Alternatively, the control unit 11 may generate multiple new datasets from the results of episodes in reinforcement learning, and use the generated multiple new datasets and the multiple datasets 4 to adjust (optimize) the computation parameters of the control model 5. Any method may be adopted for deep reinforcement learning, such as R2D3: Recurrent Replay Distributed DQN from Demonstrations.
[0091] (Specific examples of events) In this embodiment, in the machine learning described above, the control unit 11 prioritizes using datasets in which the response rate to the control command events indicated by the ground truth data 45 is evaluated as appropriate because the response rate meets predetermined conditions.
[0092] As described above, an event may include any occurrence that may be involved in the operation of a moving object. Furthermore, an event may include any occurrence that can be detected by sensor S. For example, the moving object (moving object M) may be a vehicle, and the event may include at least one of the following: deceleration of a preceding vehicle relative to the vehicle, a vehicle cutting in alongside, the appearance of a parked or stopped vehicle, the appearance of an obstacle, and a change in traffic signals.
[0093] The specified conditions may be defined as appropriate to evaluate an appropriate reaction rate depending on the event. The start time of the event may be determined by sensor data obtained by sensor S. The reaction rate may be defined by the time between the start time of the event and the start time of the operation control (operation) for that event. A specific example is shown below.
[0094] (A) Slowing down the preceding vehicle Figure 7A schematically shows an example of an event (deceleration of the preceding vehicle). In the example in Figure 7A, vehicle MA is the vehicle that encounters the event. That is, during the training phase, it can be assumed that dataset 4 is obtained from vehicle MA. On the other hand, vehicle MB is the vehicle preceding vehicle MA.
[0095] If vehicle MA encounters deceleration of the preceding vehicle MB, it may perform any operation to address the deceleration of the preceding vehicle MB. One example of a control operation that vehicle MA may perform is a deceleration operation corresponding to the preceding vehicle MB. Therefore, in one example, the reaction time may be defined as the time from the time when the preceding vehicle MB began to decelerate (the time when deceleration was detected) to the time when vehicle MA began to perform a deceleration operation.
[0096] The deceleration of the preceding vehicle MB is determined by, for example, the speed, position, or distance between vehicles, as measured by cameras, radar, LiDAR, etc. This can be detected by a sensor (sensor S). The indicators for detecting the deceleration of the preceding vehicle MB may be defined as appropriate. For example, collision risk indicators such as time-to-collision (TTC) and margin-to-collision (MTC) may be used as indicators for detecting the deceleration of the preceding vehicle MB.
[0097] On the other hand, deceleration operations in the vehicle MA may be detected by at least one of the vehicle MA's speed, acceleration, and braking amount. For example, the timing of a deceleration operation in the vehicle MA may be detected by threshold evaluation of at least one of the vehicle MA's speed, acceleration, and braking amount.
[0098] Therefore, in one example of this embodiment, the reaction speed of the deceleration operation to the deceleration of the preceding vehicle MB may be defined by the time from the time the deceleration of the preceding vehicle MB is detected (start time of the event) to the time the deceleration operation of vehicle MA is detected (start time of the operation). The shorter the time, the faster the reaction. A fast response rate is considered good, while a longer response time is considered slower.
[0099] Figure 7B schematically shows an example of a method for evaluating the reaction speed of a response operation to the deceleration of a preceding vehicle. In the example in Figure 7B, a collision risk index is used as an indicator for detecting the deceleration of the preceding vehicle MB, and the acceleration (negative acceleration) of vehicle MA is used as an indicator for detecting the deceleration operation. A threshold TA is set for the collision risk index to detect the start time of the event (deceleration of the preceding vehicle MB), and a threshold TB is set for the acceleration of vehicle MA to detect the start time of the deceleration operation. The threshold TB may be set considering external factors such as the influence of the road surface. Each threshold (TA, TB) may be set manually by the operator, or it may be set at least partially automatically by statistical quantities, etc. In the example in Figure 7B, the reaction speed can be calculated from the time from the start time of the event detected by threshold TA (based on the collision risk index) to the start time of the deceleration operation detected by threshold TB (based on the acceleration of vehicle MA).
[0100] Furthermore, the operation (operation control) of vehicle MA in relation to the preceding vehicle MB is not limited to deceleration as in the example above, and may be appropriately selected depending on the embodiment. In another example, when vehicle MA encounters deceleration of the preceding vehicle MB, it may perform a lane change. In this case, the start time of the lane change operation may be detected by at least one of the vehicle MA's speed, acceleration, brake amount, accelerator amount, and steering amount, and the reaction speed may be calculated accordingly. The steering amount may be measured, for example, by steering torque. In yet another example, the start time of the lane change operation may be detected by vehicle operation such as turn signal operation. In addition, signals from the preceding vehicle MB, such as whether or not the brake lights of the preceding vehicle MB are illuminated, may be used as indicators for detecting the deceleration of the preceding vehicle MB.
[0101] (B) Cut-in of parallel running vehicle Figure 8A schematically shows an example of an event (cut-in by a parallel vehicle). In the example in Figure 8A, vehicle MC is the vehicle that encounters the event. That is, in the training phase, it can be assumed that dataset 4 is obtained from vehicle MC. On the other hand, vehicle MD is a vehicle running parallel to vehicle MC. In the example in Figure 8A, it is assumed that parallel vehicle MD cuts in front of vehicle MC.
[0102] If a vehicle encounters a cut-in by a parallel vehicle MD, the vehicle control controller (MC) may perform any operation to address the cut-in by the parallel vehicle MD. One example of a control operation that the vehicle MC may perform is a deceleration operation in response to the cut-in by the parallel vehicle MD. For example, the reaction speed may be defined as the time from the time the parallel vehicle MD started its cut-in (the time the cut-in was detected) to the time the vehicle MC started its deceleration operation.
[0103] The cut-in of the parallel vehicle MD is, for example, based on the speed, position, or vehicle of cameras, radar, LiDAR, etc. The distance between vehicles can be detected by a sensor (sensor S). An indicator for detecting the cut-in of a parallel vehicle MD may be defined as appropriate. For example, distance indicators such as the lap amount and the distance between the parallel vehicle MD and the white line may be used as indicators for detecting the cut-in of a parallel vehicle MD. For example, the cut-in of a parallel vehicle MD may be detected when the lap amount falls below a certain value, or when the parallel vehicle MD reaches the white line. The lap amount is the distance in the vehicle width direction (left-right direction with respect to the direction of travel of the vehicle, up-down direction in the figure) between the other vehicle (parallel vehicle MD) and the predicted path MCA of the own vehicle (vehicle MC). The predicted path MCA may be, for example, the predicted path range of vehicle MC shown by the dotted line in the figure. On the other hand, similar to (A) above, deceleration operations in vehicle MC may be detected by at least one of the vehicle MC's speed, acceleration, and brake amount. Therefore, in this example, the reaction speed of the deceleration operation to the cut-in of the parallel vehicle MD is determined by the time from the time the cut-in of the parallel vehicle MD is detected (event start time) to the time the deceleration operation of the vehicle MC is detected (operation start time). It may be defined.
[0104] Figure 8B schematically shows an example of a method for evaluating the reaction speed of response operations to a cut-in by a parallel vehicle. In the example in Figure 8B, the lap amount is used as an indicator for detecting the cut-in of the parallel vehicle MD, and the acceleration (negative acceleration) of the vehicle MC is used as an indicator for detecting the deceleration operation. The threshold TC is set for the lap amount to detect the start time of the event (cut-in of the parallel vehicle MD), and the threshold TD is set for the acceleration of the vehicle MC to detect the start time of the deceleration operation. The threshold TD may be set considering external factors such as the influence of the road surface. Each threshold (TC, TD) may be set manually by the operator, or it may be set at least partially automatically by statistical quantities, etc. In the example in Figure 8B, the reaction speed can be calculated from the time from the event start time detected by threshold TC (based on the lap amount) to the start time of the deceleration operation detected by threshold TD (based on the acceleration of the vehicle MC).
[0105] Furthermore, the operation (operation control) of the vehicle MC in relation to the parallel vehicle MD is not limited to deceleration as in the example above, and may be appropriately selected depending on the embodiment. In another example, if the vehicle MC encounters a cut-in from the parallel vehicle MD, it may perform a lane change. In this case, the start time of the lane change operation may be detected by at least one of the vehicle MC's speed, acceleration, brake amount, accelerator amount, and steering amount, and the reaction speed may be calculated accordingly. In yet another example, the start time of the lane change operation may be detected by vehicle operation such as turn signal operation. In addition, signals from the parallel vehicle MD, such as whether or not the stop lamps or turn signal lamps are illuminated, may be used as indicators for detecting the cut-in from the parallel vehicle MD.
[0106] Furthermore, the manner in which a parallel vehicle's MD cuts in can vary depending on the situation. For example, in cut-ins at impassable locations such as merging lanes or construction zones, the vehicle MC may need to react earlier to ensure that the parallel vehicle's MD can reliably merge, compared to other cut-ins (e.g., normal lane changes). Therefore, cut-in events may be categorized according to the situation.
[0107] (C) Occurrence of parked or stopped vehicles Figure 9 schematically illustrates an example of an event (the occurrence of a parked vehicle). In the example in Figure 9, vehicle ME is the vehicle that encountered the event. That is, during the training phase, it can be assumed that dataset 4 is obtained from vehicle ME. On the other hand, vehicle MF is the parked vehicle. In the example in Figure 9, it is assumed that vehicle MF parked in front of vehicle ME.
[0108] If a parked vehicle MF is encountered, the vehicle ME may perform any operation to deal with the parked vehicle MF. For example, the occurrence of a parked vehicle MF may be treated similarly to the deceleration of the preceding vehicle MB in (A) above, with the preceding vehicle MB being replaced by the parked vehicle MF. That is, for example, one example of a control operation that the vehicle ME may take is a deceleration or avoidance operation (such as changing lanes) in response to the occurrence of a parked vehicle MF. The reaction speed may be defined as the time from the time the parked vehicle MF occurred (the time the parked vehicle MF was detected) to the time when the vehicle ME started the deceleration or avoidance operation.
[0109] The occurrence of parked vehicle MF is, for example, due to speed, position, or distance between vehicles detected by cameras, radar, LiDAR, etc. It can be detected by a sensor (sensor S) related to distance. For example, the above collision risk index may be used as an indicator for detecting a parked vehicle MF. In addition to the above collision margin time and collision margin, indicators such as the distance to the parked vehicle MF and the distance between vehicles / driving speed (THW: Time Head Way) may be used as collision risk indicators for detecting a parked vehicle MF. In one example, the reaction speed to the occurrence of a parked vehicle MF is relative to the value of the collision risk index. The start time of an event detected by threshold determination may be calculated from the time from the event start time detected by threshold determination on the value of the vehicle ME's operating amount (speed, acceleration, brake amount, accelerator amount, and steering amount) to the start time of the operation detected by threshold determination. In another example, the start time of an evasive operation may be detected by vehicle operation such as turn signal operation. In addition, signals from a parked vehicle MF, such as whether or not the hazard lights of the parked vehicle MF are illuminated, may be used as an indicator to detect the occurrence of a parked vehicle MF.
[0110] (D) Obstacles The occurrence of an obstacle is the same as the occurrence of a parked vehicle MF described above. In the above example, by replacing the parked vehicle MF with an obstacle, the reaction speed of an operation to the occurrence of an obstacle can be evaluated in the same way as the reaction speed of an operation to the occurrence of a parked vehicle MF. In one example of this embodiment, the reaction speed to the occurrence of an obstacle may be calculated by the time from the event start time detected by threshold determination of the value of the collision risk index to the start time of the operation detected by threshold determination of the value of the vehicle's movement amount. The collision risk index may be measured by replacing the parked vehicle MF with an obstacle. In another example, the start time of the avoidance operation may be detected by vehicle operation such as turn signal operation. As described above, the obstacle may be, for example, a pedestrian, a bicycle, etc.
[0111] (E) Changes in traffic lights Figure 10A schematically shows an example of an event (a change in traffic signals). In the example in Figure 10A, vehicle MG is the vehicle that encounters the event. That is, during the learning phase, it can be assumed that dataset 4 is obtained from vehicle MG.
[0112] When encountering a change in traffic signal MT, the vehicle's motor controller (MG) may perform any operation to address the change in traffic signal MT. For example, if the traffic signal MT changes from a proceed signal to a deceleration signal or a caution signal (from a green light to a yellow light), one possible control operation of the vehicle's MG is to decelerate in order to stop before the traffic signal MT or to accelerate in order to pass the road where the traffic signal MT is installed. In this example, the reaction time may be defined as the time from the time the traffic signal MT began to change (the time the change was detected) to the time when the vehicle's MG began to decelerate or accelerate.
[0113] Changes in the traffic signal MT can be detected, for example, by a sensor (sensor S) such as a camera. The indicator used to detect changes in the traffic signal MT may be the result of identifying the color of the illuminated signal in the traffic signal MT. The color of the illuminated signal in the traffic signal MT may be identified by any method. On the other hand, deceleration or acceleration operations in the vehicle MG may be detected by at least one of the vehicle MG's speed, acceleration, accelerator amount, and brake amount. Therefore, in this example, the reaction speed of deceleration or acceleration operations to changes in the traffic signal MT may be defined by the time from the time the change in the traffic signal MT is detected (event start time) to the time the deceleration or acceleration operation of the vehicle MG is detected (operation start time). Thresholds may be set separately for deceleration and acceleration operations.
[0114] Figure 10B schematically shows an example of a method for evaluating the response speed of response operations to changes in traffic signals. In the example in Figure 10B, it is assumed that the color change of the traffic signal MT occurs instantaneously. In this example in Figure 10B, the result of identifying the color of the illuminated signal on the traffic signal MT is used as an indicator for detecting changes in the traffic signal MT, and the acceleration of the vehicle MG is used as an indicator for detecting deceleration or acceleration operations. The threshold TE is set for the acceleration of the vehicle MG to detect the start time of deceleration operations, and the threshold TF is set for the acceleration of the vehicle MG to detect the start time of acceleration operations. Each threshold (TE, TF) may be set considering external factors such as the influence of the road surface. Each threshold (TE, TF) may be set manually by the operator, or at least partially automatically by statistical quantities, etc. This is also possible. In the example in Figure 10B, the reaction time for deceleration can be calculated from the time when the signal MT changes (event start time) to the start time of the deceleration operation detected by the threshold TE (based on the acceleration of the vehicle MG). Similarly, the reaction time for acceleration can be calculated from the time when the signal MT changes (event start time) to the start time of the acceleration operation detected by the threshold TF (based on the acceleration of the vehicle MG).
[0115] Note that the traffic light MT change event is not limited to the above example (switching from a green light to a yellow light), and may be set as appropriate depending on the embodiment. In another example, instead of the timing of switching to a yellow light, the timing of switching to a red light may be detected as the start time of the traffic light MT change event.
[0116] Furthermore, in the above example, the timing of the completion of the change in the traffic light MT (for example, the timing when the green light turns off and the yellow light turns on) is defined as the change time of the traffic light MT (event start time). However, the definition of the event start time is not limited to this example and may be set as appropriate depending on the embodiment. In another example, the traffic light MT may include a pedestrian signal in addition to a vehicle signal (Figure 10A is an image of a vehicle signal). In this case, any timing from the start of flashing of the pedestrian signal to the completion of the change in the vehicle signal may be defined as the change time of the traffic light MT (event start time).
[0117] (Specific examples of preferential use) The control unit 11 prioritizes the use of datasets from among the multiple datasets 4 that are evaluated as having appropriate reaction rates for machine learning. In other words, the control unit 11 incorporates datasets evaluated as having appropriate reaction rates into the training of the control model 5, rather than datasets evaluated as having inappropriate reaction rates. In this embodiment, the priority use may consist of at least one of the following three methods.
[0118] (1) Method 1 Figure 11A schematically shows an example of a first method in which datasets with appropriately evaluated reaction rates are used preferentially. As a first method, the control unit 11 may use datasets with appropriately evaluated reaction rates for training the control model 5, and not use datasets with inappropriately evaluated reaction rates for training the control model 5.
[0119] The graph in Figure 11A shows a histogram of the number of events (samples) against the reaction rate. The number of events may be the number of data points, or it may be aggregated in a way that differs at least partially from the data point collection. Whether the reaction rate is appropriate or not may be evaluated using any statistic. For example, the statistic may be the mean, mode, nth percentile, etc. n can be any number. The nth percentile may be, for example, the 50th percentile (median), the 25th percentile, etc.
[0120] For example, whether the reaction rate is appropriate or not (whether it meets the specified conditions) may be determined by this statistic (corresponding to the case where the set value is 0 in the example in Figure 11A). If a faster reaction rate is evaluated as more appropriate, a dataset with a reaction rate faster than the reference statistic may be used for machine learning as a dataset with an appropriate reaction rate. In other words, the control unit 11 may use the reference statistic as the upper limit of the reaction rate. If an appropriate range of reaction rates is defined, the upper and lower limits may be set from the corresponding statistics, respectively. As a specific example, assuming that the data is aggregated from the fastest reaction rates, the lower limit of the reaction rate may be defined by the 25th percentile value, and the upper limit of the reaction rate may be defined by the 75th percentile value.
[0121] In another example, instead of using the statistic directly as the standard, the control unit 11 may derive a standard value by performing an arbitrary operation on the statistic. Simply put, the control unit 11 may derive a standard value by adding or subtracting a set value from the statistic. Then, whether the reaction rate is appropriate or not may be distinguished by the derived standard value instead of the above statistic. That is, if a faster reaction rate is evaluated as more appropriate, a dataset with a reaction rate faster than the standard value may be used for machine learning as a dataset with an appropriate reaction rate. When a range of appropriate reaction rates is defined, the upper and lower limits may be derived from the same statistic, respectively. Alternatively, the upper and lower limits may be derived from different statistics, respectively. In each case, the magnitude of the set value used to derive the upper limit may be the same as or different from the magnitude of the set value used to derive the lower limit.
[0122] In the example in Figure 11A, the mode is used as the evaluation metric. The baseline value is calculated by subtracting a set value from the mode's reaction rate (the time between the event start time and the operation start time). Datasets with a reaction rate faster than the baseline value are evaluated as datasets with an appropriate reaction rate (the hatched area in Figure 11A). In this first method, the control unit 11 may exclude datasets that are not evaluated as having an appropriate reaction rate and perform the machine learning process using only datasets that are evaluated as having an appropriate reaction rate.
[0123] (2) Second method Figure 11B schematically shows an example of a second method in which datasets with appropriately evaluated reaction rates are given priority. As a second method, the control unit 11 may set a higher sampling probability for the first dataset among the multiple datasets 4 in which the reaction rates are evaluated as appropriate, and a lower sampling probability for the second dataset in which the reaction rates are not evaluated as appropriate. In other words, the control unit 11 may increase the number of samples taken in machine learning for datasets in which the reaction rates are evaluated as appropriate.
[0124] For example, the sampling probability (number of times) corresponds to the probability (number of times) of being extracted into a mini-batch in the machine learning described above. Therefore, a higher sampling probability (more sampling times) means that more times the values of the computational parameters of control model 5 are used for adjustment. This makes it possible to train control model 5 to reflect a dataset where the reaction rate is evaluated as appropriate, compared to a dataset where the reaction rate is not evaluated as appropriate.
[0125] The method for setting the sampling probability (number of samples) according to the reaction rate may be determined as appropriate depending on the embodiment. In one example, the control unit 11 may determine the sampling probability (number of samples) for each dataset 4 using a function formula that is appropriately designed so that the sampling probability (number of samples) increases as the reaction rate conforms to predetermined conditions.
[0126] In the second method, the control unit 11 may exclude at least a portion of the second dataset whose reaction rate is not evaluated as appropriate from the machine learning target (i.e., set the sampling probability to 0). Alternatively, the control unit 11 may not exclude the second dataset whose reaction rate is not evaluated as appropriate from the machine learning target (i.e., not set the sampling probability to 0).
[0127] In the example in Figure 11B, the given conditions are defined to be evaluated as more appropriate the faster the reaction rate. That is, the faster the reaction rate of the dataset, the higher the sampling probability is set. Also, the second dataset is not excluded from the machine learning study. As a specific example, the sampling probability P(i) of the i-th dataset among the multiple datasets 4 is This may be defined by the following formula 1.
[0128]
number
[0129] Additionally, as an optional measure, the training weights for each dataset 4 may be adjusted according to the sampling probability. For example, datasets with low sampling probabilities may not be reflected at all in the adjustment of the computational parameters of the control model 5. To avoid this, the training weights for datasets with lower sampling probabilities may be increased to the extent that the priority of sampling probabilities is not invalidated. As mentioned above, increasing the training weights may be achieved by increasing the learning rate. As a specific example, the reaction speed of the i-th dataset... The weight of training lol i This can be calculated from the sampling probability P(i) using the following equation 2.
[0130]
number
[0131] (3) Third method As a third method, the control unit 11 may set a large training weight for the first dataset among the multiple datasets 4 in which the reaction rate is evaluated as appropriate, and a small training weight for the second dataset in which the reaction rate is not evaluated as appropriate, and then execute the machine learning process described above.
[0132] For example, the control unit 11 may set a large training weight by increasing the learning rate in the above machine learning, or a small training weight by decreasing the learning rate. If the learning rate is large, the amount of update when adjusting the values of the computation parameters of the control model 5 in the fourth step of the above machine learning will be large. This makes it possible to train the control model 5 to reflect the dataset in which the reaction speed is evaluated as appropriate, compared to the dataset in which the reaction speed is not evaluated as appropriate.
[0133] The method for setting training weights according to reaction speed may be determined as appropriate depending on the embodiment. In one example, the control unit 11 may determine the training weights for each dataset 4 using a function formula that is appropriately designed so that the training weights increase as the reaction speed meets predetermined conditions.
[0134] In the third method, the control unit 11 may exclude at least a portion of the second dataset whose reaction speed is not evaluated as appropriate from machine learning (i.e., set the training weight to 0). The dataset with a training weight of 0 may not be used for machine learning. Alternatively, the control unit 11 may not exclude the second dataset whose reaction speed is not evaluated as appropriate from machine learning (i.e., not set the training weight to 0).
[0135] In the first to third methods described above, the extraction process (calculation process) for selecting datasets to be used preferentially may be performed by any component. For example, the extraction process may be performed by the control unit 11. That is, the control unit 11 may refer to the reaction rate of each dataset 4 and assign a priority to each dataset 4 for use in machine learning according to the first to third methods described above. In another example, even if the control unit 11 does not perform any processing, the extraction process may be achieved by the mechanism of the memory area that stores the datasets 4. As a specific example, if the datasets 4 are stored in a database, the extraction process may be achieved as a database operation. That is, the control unit 11 may extract each dataset 4 from the database in a state where a priority to use in machine learning has been assigned based on the reaction rate (for example, sorted in order of reaction rate in the database). Not using the datasets for machine learning may be achieved by not extracting them from the database. The control unit 11 may use the extracted datasets from the database as they are for machine learning, thereby performing machine learning that prioritizes the use of datasets whose reaction rates are evaluated as appropriate. In the third method described above, the control unit 11 may set the training weights according to the order in which the data are extracted (for example, if the data are extracted in order of fastest response time, the control unit 11 may assign heavier training weights to the datasets extracted earlier).
[0136] Furthermore, in the first to third methods described above, whether or not to use a dataset where the operation start time is earlier than the event start time (a dataset located to the left of the event start time in Figures 11A and 11B) for machine learning may be appropriately selected depending on the embodiment.
[0137] The control unit 11 may, by employing at least one of the first to third methods described above, prioritize the use of datasets from among the multiple datasets 4 that are evaluated as having appropriate response rates for machine learning. The first to third methods described above may be used in combination. In this embodiment, the control unit 11 can generate a trained control model 5 by performing the machine learning process described above. Once the machine learning of the control model 5 is complete, the control unit 11 proceeds to the next step S103.
[0138] The operations required may differ for each event. Therefore, the control unit 11 may generate a trained control model 5 for each event. In addition, the multiple datasets 4 obtained may include datasets relating to events other than the target event to which priority is assigned based on reaction speed. In this case, priority may be assigned to the datasets relating to the other events in any way. Alternatively, the datasets relating to the other events may be used for machine learning without priority being assigned.
[0139] <Step S103> Returning to Figure 6, in step S103, the control unit 11 operates as a storage processing unit 113. That is, the control unit 11 generates information about the trained control model 5 generated by machine learning as learning result data 125. The learning result data 125 may be configured as appropriate to include information for reproducing the trained control model 5. The control unit 11 stores the generated learning result data 125 in a predetermined memory area.
[0140] The predetermined memory area may be, for example, RAM in the control unit 11, memory unit 12, external memory device, memory media, or a combination thereof. The memory media may be, for example, a CD or DVD. The external storage device may be a semiconductor memory or the like, and the control unit 11 may store the learning result data 125 in the storage medium via the drive 16. The external storage device may be, for example, a data server such as a NAS. In this case, the control unit 11 may use the communication interface 13 to store the learning result data 125 in the data server via the network. The external storage device may also be, for example, an external storage device. The external storage device may be appropriately connected to the model generation device 1. For example, the model generation device 1 may further be equipped with an external interface, and may be connected to the external storage device via this external interface.
[0141] Once the machine learning results have been saved, the control unit 11 terminates the processing procedure of the model generation device 1 in this example of operation.
[0142] The generated learning result data 125 may be provided to the control device 2 at any timing and in any manner. For example, the control unit 11 may transfer the learning result data 125 to the control device 2 as part of the processing in step S103 or separately from the processing in step S103. The control device 2 may acquire the learning result data 125 by receiving this transfer. Alternatively, for example, the control device 2 may acquire the learning result data 125 by accessing the model generation device 1 or data server via a network using the communication interface 23. Alternatively, for example, the control device 2 may acquire the learning result data 125 via the storage medium 92. Alternatively, for example, the learning result data 125 may be pre-loaded into the control device 2.
[0143] Furthermore, the control unit 11 may update or generate new learning result data 125 by repeatedly executing the processes in steps S101 to S103 periodically or irregularly. During this repetition, at least a portion of the dataset 4 used for machine learning may be modified, corrected, added, deleted, etc. as appropriate. The control unit 11 may then update the learning result data 125 held by the control device 2 by providing the updated or newly generated learning result data 125 to the control device 2 in any way.
[0144] [Control device] Figure 12 is a flowchart showing an example of a processing procedure for the automatic control of a target mobile object M using a trained control model 5 by the control device 2 according to this embodiment. The following processing procedure is an example of a control method executed by a computer. However, the following processing procedure of the control device 2 is merely an example, and each step may be modified as much as possible. Furthermore, depending on the embodiment, steps in the following processing procedure can be omitted, replaced, and added as appropriate.
[0145] <Step S201> In step S201, the control unit 21 operates as an acquisition unit 211. That is, the control unit 21 acquires target data 221 that indicates the environment in which the target moving object M is moving. In one example, the target data 221 may include sensor data obtained by the sensor S. In addition, the target data 221 may include any information that can be involved in control, such as set speed, speed limit, map information, and navigation information. The control unit 21 may acquire various types of information by any method. Once the target data 221 is acquired, the control unit 21 proceeds to the next step S202.
[0146] <Step S202> In step S202, the control unit 21 operates as a derivation unit 212. That is, the control unit 21 uses the trained control model 5 to derive control commands from the acquired target data 221.
[0147] Furthermore, the control unit 21, at any time before executing step S202, will retrieve the learning results The control unit 21 may refer to data 125 to set up the trained control model 5 to a usable state (i.e., a state in which computational processing can be performed). If a trained control model 5 is generated for each event, the control unit 21 may identify an event that the target mobile object M is encountering or is likely to encounter. The encounter of an event or its probability can be detected by the sensor S. The control unit 21 may select a trained control model 5 from among the multiple trained control models 5 it holds that corresponds to the identified event. The control unit 21 may then use the selected trained control model 5 to derive a control command from the acquired target data 221.
[0148] The computational processing of the trained control model 5 may be appropriately determined according to the type, configuration, structure, etc., of the control model 5. For example, if the control model 5 is composed of a neural network, the control unit 21 inputs the target data 221 to the trained control model 5 and performs forward propagation computation of the trained control model 5. As a result of performing this computation, the control unit 21 can obtain an output value from the trained control model 5 that corresponds to the result of deriving the control command. If the output of the control model 5 is configured to indirectly indicate the control command, the control unit 21 may derive the control command by performing a predetermined computation on the output of the control model 5. Once the derivation of the control command is complete, the control unit 21 proceeds to the next step S203.
[0149] <Step S203> In step S203, the control unit 21 operates as an operation control unit 213. That is, the control unit 21 controls the movement of the target mobile object M according to the result of deriving the control command. The control unit 21 may directly or indirectly control the movement of the target mobile object M.
[0150] Once control of the target mobile body M is complete, the control unit 21 terminates the processing procedure of the control device 2 according to this example of operation. The control unit 21 may repeatedly execute the series of information processing steps S201 to S203. The timing of the repetition may be determined as appropriate depending on the embodiment. In one example, the control unit 21 may repeatedly execute the series of information processing steps S201 to S203 for a predetermined period (for example, from when the power source of the mobile body M is started until it is stopped). This allows the control device 2 to continuously perform automatic control of the mobile body M.
[0151] [Features] In this embodiment, in the process of step S102, datasets with a reaction rate evaluated as appropriate are given priority for machine learning to generate a trained control model 5. Since the capability of the trained model generated by machine learning depends on the dataset used for the machine learning, according to this embodiment, it can be expected that a trained control model 5 that has acquired the ability to perform control of a moving object with an appropriate reaction rate will be obtained. Furthermore, in the process of step S202, it can be expected that by using such a trained control model 5, it will be possible to perform control of the target moving object M with an appropriate reaction rate.
[0152] [4. Variant] While embodiments of this disclosure have been described in detail above, the above description is merely illustrative in all respects. It goes without saying that various improvements or modifications can be made without departing from the scope of this disclosure. For example, the following modifications are possible.
[0153] Figure 13 schematically illustrates an example of another scenario in which the present disclosure is applied. The modified system includes, in addition to the model generation device 1 and control device 2 described above, a data acquisition device 3. The data acquisition device 3 is one or more computers configured to collect a dataset 4. In this modified system, the data acquisition device 3 collects a combination of training data 41 and ground truth data 45. Multiple datasets 4, each composed of the above, are collected. Collecting multiple datasets 4 may include prioritizing the collection of datasets that are evaluated as appropriate because the response rate to events of the control commands indicated by the ground truth data 45 meets predetermined conditions. The data acquisition device 3 then outputs the collected multiple datasets 4 for use in machine learning.
[0154] In the example shown in Figure 13, the data acquisition device 3 is mounted on the mobile body MZ and is assumed to collect a dataset 4 in response to the subject's operation of the mobile body MZ. In this example, the data acquisition device 3 may generate training data 41 from sensor data obtained from the sensor S and other information (e.g., set speed, speed limit, map information, navigation information, etc.). The data acquisition device 3 may also generate ground truth data 45 from the results of the subject's operations during the period in which the information constituting the training data 41 is acquired. The data acquisition device 3 may then acquire the dataset 4 by associating the generated training data 41 and ground truth data 45. The data acquisition device 3 may be configured to execute the process of collecting the dataset 4 in response to commands from an external computer, such as the model generation device 1. Data exchange between each of the devices 1 to 3 may be carried out in any manner.
[0155] The configuration of the data acquisition device 3 is not limited to the example shown in Figure 13. In another example, the data acquisition device 3 may not be mounted on the mobile body MZ, but may be positioned separately from the mobile body MZ. In yet another example, the data acquisition device 3 may be integrated with the control device 2 or the model generation device 1. That is, the control device 2 may also function as the data acquisition device 3 (in this case, the mobile body MZ is the mobile body M). Alternatively, the model generation device 1 may also function as the data acquisition device 3.
[0156] [Hardware configuration] Figure 14 schematically shows an example of the hardware configuration of the data acquisition device 3 according to this modified example. As shown in Figure 14, the data acquisition device 3 according to this modified example is a computer in which a control unit 31, a storage unit 32, a communication interface 33, an input device 34, an output device 35, a drive 36, and an external interface 37 are electrically connected.
[0157] The control unit 31 to the external interface 37 and the storage medium 93 of the data acquisition device 3 may be configured in the same way as the control unit 21 to the external interface 27 and the storage medium 92 of the control device 2. The control unit 31 (CPU) is an example of the processor resources of the data acquisition device 3, and the storage unit 32 (and RAM, ROM) is an example of the memory resources of the data acquisition device 3. In this modified example, the storage unit 32 stores various information such as the data acquisition program 83 and the data set 4.
[0158] The data acquisition program 83 is a program that causes the data acquisition device 3 to perform information processing related to the acquisition of the dataset 4 (Figure 16, described later). The data acquisition program 83 includes a series of instructions for said information processing. The dataset 4 may be stored as a result of the execution of the data acquisition program 83. At least one of the data acquisition program 83 and the dataset 4 may be stored in the storage medium 93 in place of or together with the storage unit 32. The data acquisition device 3 may retrieve the data acquisition program 83 from the storage medium 93.
[0159] The data acquisition device 3 may perform data communication with another computer (e.g., model generation device 1) via the communication interface 33. The operator (e.g., subject) can operate the data acquisition device 3 using the input device 34 and the output device 35. The input device 34 and the output device 35 may be integrated together, for example, by a touch panel display. The data acquisition device 3 can communicate via the external interface 37. It may be connected to sensor S. However, the method of connecting sensor S is not limited to this example. In another example, data acquisition device 3 may be connected to sensor S via communication interface 23.
[0160] Regarding the specific hardware configuration of the data acquisition device 3, components can be omitted, replaced, and added as appropriate depending on the embodiment. For example, the control unit 31 may include multiple hardware processors. Hardware processors may consist of microprocessors, FPGAs, DSPs, ECUs, GPUs, etc. At least one of the communication interface 33, input device 34, output device 35, drive 36, and external interface 37 may be omitted. The data acquisition device 3 may consist of multiple computers. In this case, the hardware configurations of each computer may or may not be the same. The data acquisition device 3 may be a computer designed specifically for the service provided, a general-purpose server device, a general-purpose computer, a mobile phone including a smartphone, a tablet PC (Personal Computer), etc. When the mobile unit MZ is a vehicle, The data acquisition device 3 may be an in-vehicle device.
[0161] [Software Configuration] Figure 15 schematically shows an example of the software configuration of the data acquisition device 3 according to this embodiment. The control unit 31 of the data acquisition device 3 loads the data acquisition program 83 stored in the storage unit 32 into RAM, and the CPU executes the instructions contained in the data acquisition program 83. As a result, as shown in Figure 15, the data acquisition device 3 according to this embodiment operates as a computer equipped with an acquisition unit 311 and an output unit 312 as software modules.
[0162] The collection unit 311 is configured to collect multiple datasets 4, each composed of a combination of training data 41 and ground truth data 45. The collection of multiple datasets 4 by the collection unit 311 may include prioritizing the collection of datasets that are evaluated as appropriate because the response rate to events of the control commands indicated by the ground truth data 45 meets predetermined conditions. The output unit 312 is configured to output the collected datasets 4 for use in machine learning.
[0163] In this modified example, similar to the model generation device 1 and control device 2 described above, each software module of the data acquisition device 3 is also implemented by the control unit 31 (CPU). In other words, this describes an example in which each software module of the data acquisition device 3 is implemented by a general-purpose CPU. However, some or all of the software modules of the data acquisition device 3 may be implemented by one or more dedicated processors. Each of the above modules may also be implemented as a hardware module. Regarding the software configuration of the data acquisition device 3, modules may be omitted, replaced, and added as appropriate, depending on the embodiment.
[0164] [Example of operation] Figure 16 is a flowchart showing an example of the processing procedure for collecting a dataset 4 by the data acquisition device 3 according to this modified example. The following processing procedure is an example of a data acquisition method performed by a computer. However, the following processing procedure is merely an example, and each step may be modified as much as possible. Furthermore, depending on the embodiment, steps in the following processing procedure can be omitted, replaced, and added as appropriate.
[0165] (Step S301) In step S301, the control unit 31 operates as an acquisition unit 311 and collects multiple datasets 4, each composed of a combination of training data 41 and ground truth data 45. In this case, in one example, the control unit 31 may collect the dataset 4 regardless of whether the reaction rate is appropriate or not. In another example, the control unit 31 may prioritize collecting datasets that are evaluated as appropriate because the reaction rate to the control command event indicated by the ground truth data 45 meets predetermined conditions.
[0166] Collecting dataset 4 may involve either acquiring a new dataset or selecting from an already acquired dataset. Acquiring a new dataset may involve generating a new dataset or obtaining a dataset from another computer. Prioritizing collection may be performed at least in either the stage of acquiring a new dataset or the stage of selecting from an already acquired dataset. The stage of acquiring a new dataset may correspond to the initial entry stage for saving (memorizing) the data.
[0167] Prioritizing data collection may involve increasing the amount of datasets where reaction rates are deemed appropriate and decreasing the amount of datasets where reaction rates are deemed inappropriate. For example, reducing the amount of datasets where reaction rates are deemed inappropriate may include not collecting datasets for at least some of the inappropriate reaction rates.
[0168] If a device for acquiring datasets (first device) and a device for storing acquired datasets (second device) are provided separately, priority collection may be carried out in one of the following three stages. (1) Whether or not to continue to retain the target dataset in the storage of the first device. (2) Whether the first device transmits (transfers) the target dataset to the second device. (3) Whether to continue to retain the target dataset in the storage of the second device or not An example of the first device is a terminal device (user terminal, in-vehicle device, etc.), and an example of the second device is a server device. The data acquisition device 3 may be either the first or second device. In the example in Figure 13, the data acquisition device 3 is assumed to be the first device.
[0169] In one example, the data acquisition device 3 may be the first or second device, and may prioritize data acquisition at the stage of (1) or (3) above. In this case, prioritizing the acquisition of datasets that are evaluated as having an appropriate response rate may be achieved by maintaining datasets that are evaluated as having an appropriate response rate under predetermined conditions, and deleting datasets that are evaluated as having an inappropriate response rate under predetermined conditions, from among the datasets temporarily stored in the memory area (RAM, memory unit 32, storage medium 93, etc.) of the data acquisition device 3. By not maintaining datasets that are evaluated as having an inappropriate response rate, the memory area of the data acquisition device 3 can be efficiently used to acquire appropriate datasets.
[0170] In another example, the data acquisition device 3 may be the first device, and may be prioritized for data acquisition in step (2) above. In this case, the second device is, for example, an external storage device such as an external server. In one example, the external server may be the model generation device 1 or a network server (such as a NAS). The transmission process is performed in step S302, which will be described later. Therefore, if this configuration is adopted, in step S301, the control unit 31 may evaluate the response rate of the obtained dataset in order to determine whether or not to transmit it. That is, prioritizing data acquisition may include evaluating the response rate.
[0171] Each dataset 4 may be acquired by any method. As described above, dataset 4 may be acquired through the operation of the mobile device MZ by the subject. Operation of the mobile device MZ may include not only fully manual operation but also override operations on any automatic control. In addition, dataset 4 may be acquired by methods such as simulation and data augmentation. Multiple Once dataset 4 is acquired, the control unit 31 proceeds to the next step S302.
[0172] (Step S302) In step S302, the control unit 31 operates as an output unit 312. That is, the control unit 31 outputs multiple collected datasets 4 for use in machine learning.
[0173] Outputting for use in machine learning may involve maintaining the collected datasets 4 in a distinguishable (identifiable) state so that they can be used for machine learning. Therefore, outputting the collected datasets 4 may consist of saving the collected datasets 4 to any memory area. Any memory area may be RAM, memory unit 32, external storage device, etc. The external storage device may include the model generation device 1, a network server (NAS, etc.), or other external servers.
[0174] In one example, the control unit 31 may transmit multiple datasets 4 to the model generation device 1 via a network. The model generation device 1 may then generate a trained control model 5 by executing steps S102 and S103. If the data acquisition device 3 is configured integrally with the model generation device 1, the collection of datasets 4 in step S301 may be reflected in obtaining a dataset of a specified batch size in step S101 or in machine learning. In this case, the output processing in step S302 may be included in step S102.
[0175] In another example, as described above, the data acquisition device 3 may reflect prioritizing data collection at stage (2) above. In this case, outputting multiple datasets 4 may be configured by sending datasets evaluated as having an appropriate response rate under predetermined conditions to the second device, and omitting the transmission of datasets evaluated as having an inappropriate response rate under predetermined conditions to the second device. In one example, the second device may be an external server such as the model generation device 1 or a network server. This reduces the communication cost associated with outputting datasets 4 by omitting the communication processing of datasets evaluated as having an inappropriate response rate.
[0176] Once the output of multiple datasets 4 is complete, the control unit 31 terminates the processing procedure of the data acquisition device 3 according to this example of operation. The control unit 31 may repeatedly execute the series of information processing steps S301 to S302. The timing of the repetition may be determined as appropriate depending on the embodiment. In a typical example, the control unit 31 may start executing the processing of step S301 in response to a data acquisition command from an external computer such as a model generation device 1. The control unit 31 may then repeatedly execute the series of information processing steps S301 to S302 until it receives a command to stop data acquisition from the external computer. In this way, the data acquisition device 3 may be configured to continuously acquire datasets 4 while a data acquisition instruction is given.
[0177] (Features) In this modified example, the data acquisition device 3 prioritizes the collection of datasets whose reaction rates are evaluated as appropriate during the processing of step S301. This ensures that, from the processing of step S302 onward, datasets with appropriately evaluated reaction rates are given priority for use in machine learning. Therefore, even with this modified example, it is possible to obtain a trained machine learning model (control model 5) that has acquired the ability to perform control of a moving object with an appropriate reaction rate.
[0178] [5 Supplement] The processes and means described in this disclosure are free, insofar as they do not result in technical inconsistencies. It can be implemented in combination with other methods.
[0179] Furthermore, a process described as being performed by a single device may be divided and executed by multiple devices. Conversely, a process described as being performed by different devices may be executed by a single device. In a computer system, the hardware configuration used to implement each function can be flexibly changed.
[0180] The present disclosure can also be realized by supplying a computer program implementing the functions described in the embodiments above to a computer, and having one or more processors in the computer read and execute the program. Such a computer program may be provided to the computer by a non-temporary computer-readable storage medium that can be connected to the computer's system bus, or it may be provided to the computer via a network. Non-temporary computer-readable storage mediums include, for example, any type of disk such as magnetic disks (floppy disks, hard disk drives (HDDs), etc.), optical disks (CD-ROMs, DVDs, Blu-ray discs, etc.), read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic cards, flash memory, optical cards, semiconductor drives (solid-state drives, etc.), and any type of medium suitable for storing electronic instructions. [Explanation of symbols]
[0181] 1...Model generation device, 11...Control unit, 12...Storage unit, 13...Communication interface, 14...Input device, 15...Output device, 16...Drive, 81...Model generation program, 91...Storage medium, 111...Learning data acquisition unit, 112...Learning processing unit, 113...Storage processing unit, 125...Learning result data, 2... Control device, 21...Control unit, 22...Storage unit, 23...Communication interface, 24...Input device, 25...Output device, 26...Drive, 27…External interface, 82... control program, 92... storage medium, 211...Acquisition unit, 212...Derivation unit, 213...Operation control unit, 221...Target data, 3…Data acquisition device, 31...Control unit, 32...Storage unit, 33...Communication interface, 34...Input device, 35...Output device, 36...Drive, 37…External interface, 83...Data acquisition program, 93...Storage medium, 311...Collection unit, 312...Output unit, 4…Dataset, 41...Training data, 45...Correct answer data, 5…Control model, M...Mobile object, S...Sensor
Claims
1. A model generation method performed by a computer, The aforementioned model generation method is: This involves acquiring multiple datasets, each composed of a combination of training data showing the environment in which a mobile object moves in a time series and ground truth data showing control commands for the mobile object in the said environment in a time series. Using the acquired multiple datasets, machine learning is performed on the control model. Includes, Performing the machine learning described above includes training the control model for each dataset such that the result of deriving the control command for the mobile body from the training data using the control model fits the ground truth data, and Using the aforementioned multiple datasets includes prioritizing the use of datasets that are evaluated as appropriate because the response rate of the control command to the event, as indicated by the ground truth data, fits predetermined conditions. Model generation method.
2. Prioritizing the use of datasets in which the aforementioned reaction rates are evaluated as appropriate means that The sampling probability in machine learning for the first dataset, among the multiple datasets, whose reaction rate is evaluated as appropriate, is increased, and The second dataset, whose reaction rate is not evaluated as appropriate, is not excluded from the machine learning program, and the sampling probability of the second dataset in the machine learning program is set lower than that of the first dataset. It is composed of, The model generation method according to claim 1.
3. The mobile body is equipped with a sensor, The training data includes sensor data obtained by the sensor, The start time of the event in the training data is determined by the sensor data. The model generation method according to claim 1 or 2.
4. The aforementioned moving object is a vehicle. The model generation method according to claim 1 or 2.
5. The aforementioned event includes at least one of the following: deceleration of a preceding vehicle, cutting in of a parallel vehicle, the appearance of a parked or stopped vehicle, the appearance of an obstacle, and a change in traffic signals. The model generation method according to claim 4.
6. The control command includes at least one of the following: acceleration, deceleration, and steering of the vehicle. The model generation method according to claim 4.
7. A data collection method performed by a computer, The aforementioned data collection method is: Collecting multiple datasets, each composed of a combination of training data showing the environment in which a mobile object moves in a time series and ground truth data showing control commands to the mobile object in the said environment in a time series, To output multiple collected datasets for use in machine learning, Includes, The collection of the aforementioned multiple datasets includes prioritizing the collection of datasets that are evaluated as appropriate because the response rate of the control command to the event, as indicated by the ground truth data, meets predetermined conditions. Data collection methods.
8. Prioritizing the collection of datasets whose reaction rates are evaluated as appropriate comprises maintaining datasets whose reaction rates are evaluated as appropriate under predetermined conditions and deleting datasets whose reaction rates are evaluated as inappropriate under predetermined conditions from among the datasets temporarily stored in the computer's memory. The data collection method according to claim 7.
9. Outputting the aforementioned multiple datasets means A dataset in which the reaction rate is evaluated to be appropriate under the predetermined conditions is sent to an external server, and To omit sending to the external server datasets that are evaluated as having an inappropriate reaction rate under the aforementioned predetermined conditions, Composed of, The data collection method according to claim 7.
10. On the computer, To acquire target data that shows the environment in which the target moving object is moving, Using a trained control model, control commands are derived from acquired target data, The operation of the target moving object is controlled according to the result of deriving the control command, A control program for executing, The aforementioned trained control model is generated by performing machine learning using multiple datasets, each consisting of a combination of training data showing the environment in which the training mobile body moves in a time series and ground truth data showing the control commands for the training mobile body in the environment in a time series. Performing the machine learning described above includes training the control model for each dataset such that the result of deriving the control command for the mobile body from the training data using the control model fits the ground truth data, and Using the aforementioned multiple datasets for machine learning includes prioritizing the use of datasets that are deemed appropriate because the response rate of the control commands to events, as indicated by the ground truth data, meets predetermined conditions. Control program.
11. The moving object in question is a vehicle. The control program according to claim 10.
12. The aforementioned event includes at least one of the following: deceleration of a preceding vehicle, cutting in of a parallel vehicle, the appearance of a parked or stopped vehicle, the appearance of an obstacle, and a change in traffic signals. The control program according to claim 11.
13. The control command includes at least one of the following: acceleration, deceleration, and steering of the vehicle. The control program according to claim 11.
Citation Information
Patent Citations
Parking assisting device, and vehicle provided with parking assisting device
JP2003276540A
Preceding vehicle follow-up control method and preceding vehicle follow-up control device
JP2011051498A
Information distribution device
JP2014081947A
Control system and semiconductor device
JP2016031586A
Driving support device and method of learning driving characteristics
JP2019127207A