Data collection method and model generation method
A centralized data collection method addresses inefficiencies in collecting in-vehicle data by identifying underperforming scenes and generating scene-specific collection conditions, enhancing model training and accuracy.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- DENSO CORP
- Filing Date
- 2025-11-21
- Publication Date
- 2026-06-04
AI Technical Summary
Existing systems face challenges in efficiently collecting diverse data from multiple in-vehicle devices due to the heterogeneity of data types and the need for targeted data collection based on specific scenes and evaluation criteria.
A centralized data collection method that identifies scenes where machine learning models underperform, generates and distributes collection conditions to in-vehicle devices, and stores data linked to scene identification information, enabling efficient data collection and model training.
This approach allows for targeted data collection, improving the accuracy of machine learning models by addressing data collection inefficiencies and enhancing model evaluation through scene-specific data gathering.
Smart Images

Figure JP2025040874_04062026_PF_FP_ABST
Abstract
Description
Data collection method and model generation method Cross-reference to related applications
[0001] This international application claims the benefit of Japanese Patent Application No. 2024-208264, filed with the Japan Patent Office on November 29, 2024, the entire disclosure of which is incorporated herein by reference.
[0002] This disclosure relates to a technique for collecting data from a plurality of vehicles.
[0003] Patent Document 1 describes a system for collecting vehicle data from a plurality of in-vehicle devices mounted on each of a plurality of vehicles.
[0004] Japanese Unexamined Patent Application Publication No. 2023-84379
[0005] As a result of the inventors' detailed examination, it has been found that since the data handled by vehicles is diverse, there is a need to efficiently collect data from a plurality of in-vehicle devices.
[0006] This disclosure efficiently collects data from a plurality of in-vehicle devices.
[0007] One aspect of this disclosure is a data collection method executed by a center configured to receive vehicle data including at least information regarding the mounted vehicle from a plurality of in-vehicle devices mounted on each of a plurality of vehicles.
[0008] The center stores the vehicle data received from the plurality of in-vehicle devices.
[0009] The center uses the stored vehicle data to identify a scene in which the machine learning model does not meet a preset evaluation criterion based on the result of model evaluation obtained by performing machine learning including annotation, model training, and model evaluation on the machine learning model, and for the identified scene, generates collection condition data indicating a collection condition including a collection data item indicating the type of vehicle data to be collected and a collection start condition for starting data collection, and stores the generated collection condition data in association with scene identification information for identifying the scene.
[0010] The center distributes the collection condition data to the plurality of in-vehicle devices.
[0011] The center stores vehicle data uploaded from multiple in-vehicle devices based on collection condition data, linking it with scene identification information.
[0012] The center accepts registrations from users of at least one of the following: a target scene, which is the scene to be evaluated for the model; and target evaluation criteria, which are the evaluation criteria for the model evaluation of the target scene. The center then performs the model evaluation of the target scene based on the target evaluation criteria, and if the evaluation result of the model evaluation of the target scene is unsatisfactory, it generates collection condition data for collecting vehicle data for the target scene.
[0013] The data collection method described herein, configured in this manner, can collect vehicle data by narrowing it down by scene, thus enabling efficient data collection from multiple in-vehicle devices. Furthermore, the data collection method described herein can efficiently collect data from multiple in-vehicle devices by identifying scenes in which the evaluation results of the model evaluation are poor and generating collection condition data.
[0014] Another aspect of this disclosure is a model generation method that generates a machine learning model by performing model training using vehicle data collected by the data collection method of this disclosure.
[0015] The model generation method described herein is a method that uses the data collection method described herein, and by performing this method, the same effects as the data collection method described herein can be obtained.
[0016] This is a block diagram showing the configuration of the data collection system. This is a block diagram showing the configuration of the data collection device. This is a block diagram showing the configuration of the center. This is a functional block diagram showing the functional configuration of the center. This is a diagram explaining the management of collection condition data and collected data. This is a diagram explaining the traceability of the AI model. This is a flowchart showing the data collection management process. This is a flowchart showing the first half of the model learning process. This is a flowchart showing the second half of the model learning process. This is a flowchart showing the collection condition generation process. This is a flowchart showing the distribution process. This is a flowchart showing the data collection process. This is a diagram explaining the registration of test cases. This is a diagram explaining the calculation of statistical features. This is a diagram explaining the generation of collection conditions. This is a flowchart showing the test case registration process. This is a flowchart showing the automatic generation of collection conditions process. This is a flowchart showing the automatic distribution process. This is a diagram showing the processing cycle of the data collection system.
[0017] Embodiments of the present disclosure will be described below with reference to the drawings.
[0018] As shown in Figure 1, the data acquisition system 1 of this embodiment comprises a plurality of data acquisition devices 2 and a center 3.
[0019] The data acquisition device 2 is mounted on the vehicle and has the function of communicating data with the center 3 via a wide-area wireless communication network NW.
[0020] Center 3 is a device that manages the data acquisition system 1. Center 3 has the function of communicating data with multiple data acquisition devices 2 via a wide-area wireless communication network NW.
[0021] As shown in Figure 2, the data acquisition device 2 comprises a control unit 11, a CAN communication unit 12, a storage unit 13, and a communication unit 14. CAN stands for Controller Area Network.
[0022] The control unit 11 is an electronic control device centered around a microcomputer equipped with a CPU 21, ROM 22, RAM 23, etc. The various functions of the microcomputer are realized by the CPU 21 executing a program stored in a non-transitional physical recording medium. In this example, the ROM 22 corresponds to the non-transitional physical recording medium that stores the program. Furthermore, the execution of this program executes a method corresponding to the program. Note that some or all of the functions executed by the CPU 21 may be configured hardware-wise by one or more ICs, etc. Also, the number of microcomputers constituting the control unit 11 may be one or more.
[0023] The CAN communication unit 12 is connected to multiple ECUs via communication lines to enable data communication and transmits and receives data according to the CAN communication protocol. Specifically, the multiple ECUs connected to the CAN communication unit 12 include an engine ECU for engine control, a brake ECU for brake control, a steering ECU for steering control, a suspension ECU for suspension control, and an ECU for controlling the on / off state of the lights. In Figure 2, only ECUs 16, 17, and 18 are shown as ECUs connected to the CAN communication unit 12. ECU stands for Electronic Control Unit.
[0024] The memory unit 13 is a memory device for storing various types of data.
[0025] The communications unit 14 performs data communication with the center 3 via the wide-area wireless communication network NW.
[0026] As shown in Figure 3, Center 3 comprises a control unit 31, a communication unit 32, and a storage unit 33.
[0027] The control unit 31 is an electronic control device centered around a microcomputer equipped with a CPU 41, ROM 42, RAM 43, etc. The various functions of the microcomputer are realized by the CPU 41 executing a program stored in a non-transitional physical recording medium. In this example, the ROM 42 corresponds to the non-transitional physical recording medium that stores the program. Furthermore, the execution of this program executes a method corresponding to the program. Note that some or all of the functions executed by the CPU 41 may be configured hardware-wise by one or more ICs, etc. Also, the number of microcomputers constituting the control unit 31 may be one or more.
[0028] The communication unit 32 performs data communication with multiple data acquisition devices 2 via a wide-area wireless communication network NW.
[0029] The memory unit 33 is a storage device for storing various types of data. The memory unit 33 is equipped with a collection condition database (hereinafter referred to as the collection condition DB) 33a, a collection data database (hereinafter referred to as the collection data DB) 33b, a test case database (hereinafter referred to as the test case DB) 33c, a scene feature extraction data database (hereinafter referred to as the scene feature extraction data DB) 33d, a training dataset database (hereinafter referred to as the training dataset DB) 33e, an evaluation dataset database (hereinafter referred to as the evaluation dataset DB) 33f, and an evaluation result database (hereinafter referred to as the evaluation result DB) 33g.
[0030] Center 3, as a functional block realized by the CPU 41 executing a program stored in ROM 42, includes, as shown in Figure 4, a data collection unit 101, a collected data / product storage unit 102, a collected data management / analysis unit 103, a machine learning unit 104, a collection condition management unit 105, a CI / CD unit 106, a UI unit 107, authentication / authorization units 108, 109, 110, 111, 112, a collection condition file transmission unit 113, and an AI model file transmission unit 114. CI stands for Continuous Integration. CD stands for Continuous Delivery. UI stands for User Interface. AI stands for Artificial Intelligence.
[0031] The data acquisition unit 101 collects data from multiple data acquisition devices 2 installed in each of the multiple vehicles.
[0032] The collected data and product storage unit 102 stores the data collected by the data collection unit 101 and the products generated by the collected data management and analysis unit 103, the machine learning unit 104, and the collection condition management unit 105.
[0033] The collected data management and analysis unit 103 includes a preprocessing unit 121 and a data collection status confirmation unit 122.
[0034] The preprocessing unit 121 performs preprocessing on the collected data to make it easier to use in machine learning (for example, changing the resolution, changing to grayscale, etc.).
[0035] The data collection status confirmation unit 122 checks the amount of data collected.
[0036] The machine learning unit 104 includes an annotation unit 123, a model training unit 124, and a model evaluation unit 125.
[0037] The annotation unit 123 performs annotation on the data that has been preprocessed by the preprocessing unit 121, adding information (for example, labels) for machine learning.
[0038] The model training unit 124 trains a machine learning model using the data generated by the annotation unit 123. The machine learning model is, for example, a model that takes image data captured by an in-vehicle camera as input data, determines an object in the image, and outputs a determination result.
[0039] The model evaluation unit 125 evaluates the accuracy of the machine learning model based on the training results of the model training unit 124.
[0040] The collection condition management unit 105 manages the collection conditions used to collect data for training the machine learning model. The collection condition management unit 105 includes a collection condition generation unit 126.
[0041] The collection condition generation unit 126 generates collection condition data indicating the collection conditions for collecting data for improving the accuracy of the machine learning model based on the evaluation results of the model evaluation unit 125.
[0042] The CI / CD unit 106 includes a collection condition build and distribution unit 127, an AI logic build and distribution unit 128, and a distribution status management unit 129.
[0043] The collection condition build and distribution unit 127 converts the collection condition data generated by the collection condition generation unit 126 into a format recognizable by the data collection device 2 and distributes it to the data collection device 2.
[0044] The AI logic build and distribution unit 128 converts the machine learning model (i.e., AI logic) generated by the machine learning unit 104 into a format executable by the data collection device 2 and distributes it to the data collection device 2.
[0045] The distribution status management unit 129 checks the distribution status of the collection condition data and the AI logic.
[0046] The UI unit 107 is a device for exchanging information between data engineers, AI engineers, software engineers, test engineers, and operators and the center 3.
[0047] A data engineer is an engineer related to data used for machine learning. An AI engineer is an engineer related to the generation of machine learning models. A software engineer is an engineer related to applications generated by combining multiple machine learning models. A test engineer is an engineer who evaluates machine learning models and applications. An operator manages the collection condition data and the distribution status of AI logic.
[0048] The authentication and authorization unit 108 performs authentication and authorization for the access of the data engineer to the center 3. The authentication and authorization unit 108 performs authentication and authorization so that, for example, the data engineer can only access the collection data management and analysis unit 103.
[0049] The authentication and authorization unit 109 performs authentication and authorization for the access of the AI engineer to the center 3. The authentication and authorization unit 109 performs authentication and authorization so that, for example, the AI engineer can only access the machine learning unit 104.
[0050] The authentication and authorization unit 110 performs authentication and authorization for the access of the software engineer to the center 3. The authentication and authorization unit 110 performs authentication and authorization so that, for example, the software engineer can only access the machine learning unit 104.
[0051] The authentication and authorization unit 111 performs authentication and authorization for the access of the test engineer to the center 3. The authentication and authorization unit 111 performs authentication and authorization so that, for example, the test engineer can only access the collection condition management unit 105.
[0052] The authentication and authorization unit 112 performs authentication and authorization for the access of the operator to the center 3. The authentication and authorization unit 112 performs authentication and authorization so that, for example, the operator can only access the CI / CD unit 106.
[0053] The collection condition file transmission unit 113 transmits the collection condition file generated by the collection condition build and distribution unit 127 to a plurality of data collection devices 2.
[0054] The AI model file transmission unit 114 transmits the AI model file generated by the AI logic build and distribution unit 128 to multiple data acquisition devices 2.
[0055] Furthermore, the authentication and authorization units 108-112 are configured to allow different access permissions to be set for each of the multiple applications (for example, ADAS applications, multimedia applications) that utilize the machine learning model handled by Center 3. ADAS stands for Advanced Driver Assistance System.
[0056] The data collection conditions are either manually generated by engineers or automatically generated by an AI model, as shown by arrows L1 and L2 in Figure 5.
[0057] The data collection conditions file shown in Figure 5 includes conditions C1, C2, and C3, etc. Conditions C1 and C2 are the conditions for executing data collection. Data collection is executed when both conditions C1 and C2 are met.
[0058] Condition C3 indicates the timing for uploading the collected data to Center 3.
[0059] A collection condition profile is attached to the collection condition file. The collection condition profile includes a collection condition ID and a scene tag. ID stands for identification. "CJD-001" shown in Figure 5 is the collection condition ID. "night rain" shown in Figure 5 is the scene tag. A scene includes ambient brightness, vehicle behavior, driver actions, road conditions, application execution status, etc. Therefore, the collection condition ID corresponds to a scene. A scene tag is a string set by engineers to identify a scene.
[0060] The collection condition file to which the collection condition profile has been added is stored in the collection condition DB33a, as indicated by arrow L3.
[0061] The data collection conditions file is distributed to the data collection device 2 mounted on the vehicle, as indicated by arrow L4.
[0062] The data acquisition device 2 acquires image data and sensor data, at least one of them, based on the acquisition condition file obtained from the center 3. Data D1 shown in Figure 5 is image data. Data D2 shown in Figure 5 is sensor data showing the detection results of sensors mounted on the vehicle.
[0063] The data collection device 2 adds the collection condition ID included in the collection condition profile corresponding to the collection condition file to the collected data. As shown by arrow L5, the data collection device 2 uploads the data with the added collection condition ID to the center 3.
[0064] Center 3 stores the data uploaded from data acquisition device 2 in the collected data DB 33b. As shown by arrow L6, the data stored in the collected data DB 33b is linked to scenes by the collection condition ID. In this way, the data stored in the collected data DB 33b is linked to the scene tag corresponding to the collection condition ID (for example, "night rain" in Figure 5).
[0065] The data stored in the collected data DB33b is used by engineers, as indicated by arrow L7.
[0066] As shown in Figure 6, the collected data metadata D11 stores data classified according to multiple collection conditions. The collected data metadata D11 shown in Figure 6 includes a first collection condition storage area R1 that stores data with collection condition ID "T-1" and a second collection condition storage area R2 that stores data with collection condition ID "T-2".
[0067] The collected data metadata D11 stores data for each of the multiple vehicles within the collection condition storage area, where data is stored for each collection condition.
[0068] The first collection condition storage area R1 includes a first vehicle storage area R11 that stores data acquired from a vehicle identified as "VIN-1", and a second vehicle storage area R12 that stores data acquired from a vehicle identified as "VIN-2".
[0069] The second collection condition storage area R2 includes a third vehicle storage area R13 for storing data acquired from the vehicle identified as "VIN-3".
[0070] The vehicle storage area, which stores data for each vehicle, stores the date and time of data acquisition, and for each data item, it stores the data ID and data type.
[0071] The first vehicle storage area R11 indicates, for the vehicle identified as "VIN-1", the date and time of data acquisition, that the data type of the data with data ID "Data-1" is sensor data, and that the data type of the data with data ID "Data-2" is image data.
[0072] The second vehicle storage area R12 indicates, for the vehicle identified as "VIN-2", the date and time of data acquisition, that the data type of the data with data ID "Data-3" is sensor data, and that the data type of the data with data ID "Data-4" is image data.
[0073] The third vehicle storage area R13 indicates, for the vehicle identified as "VIN-3", the date and time of data acquisition, that the data type of the data with data ID "Data-5" is sensor data, and that the data type of the data with data ID "Data-6" is image data.
[0074] Annotation job data D12 stores data categorized by multiple annotation jobs.
[0075] The annotation job data D12 shown in Figure 6 includes a first annotation job storage area R21 that stores data with annotation job ID "Anot-1", and a second annotation job storage area R22 that stores data with annotation job ID "Anot-2".
[0076] The annotation job storage area, where data is stored for each annotation job, stores the data by associating the collection condition ID with the data ID for each collection condition ID.
[0077] The first annotation job storage area R21 stores the data by associating the collection condition ID "T-1" with the data ID "Data-2", and the collection condition ID "T-2" with the data ID "Data-6".
[0078] The second annotation job storage area R22 stores the collection condition ID, which is "T-1", and the data ID, which is "Data-4", in association with each other.
[0079] As indicated by arrows L11 and L12, in an annotation job with annotation job ID "Annot-1", image data collected under the collection condition specified by the collection condition ID "T-1" and specified by the data ID "Data-2" is used, and image data collected under the collection condition specified by the collection condition ID "T-2" and specified by the data ID "Data-6" is used.
[0080] As indicated by arrow L13, in annotation jobs with annotation job ID "Annot-2", image data collected under the collection condition specified by the collection condition ID "T-1" and specified by the data ID "Data-4" is used.
[0081] Model training job data D13 stores the training job ID and one or more annotation job IDs in association.
[0082] The model training job data D13 shown in Figure 6 stores the model training job ID, "Train-1", the annotation job ID, "Anot-1", and the annotation job ID, "Anot-2", in association with each other.
[0083] As indicated by arrows L14 and L15, in a model training job with model training job ID "Train-1", the image data generated by the annotation job identified by annotation job ID "Anot-1" and the image data generated by the annotation job identified by annotation job ID "Anot-2" are used.
[0084] Next, the procedure for the data collection management process performed by Center 3 will be explained. The data collection management process is a process that is repeatedly executed while the control unit 31 is operating.
[0085] When the collected data management process is executed, the CPU 41 of the control unit 31 determines in S10 whether or not new collected data has been saved to the collected data DB 33b, as shown in Figure 7. The collected data includes, for example, in-vehicle camera videos, in-vehicle camera still images, application operation results, sensor system recognition results, and CAN data.
[0086] If no new data is saved to the data collection database DB 33b, the CPU 41 terminates the data collection management process. On the other hand, if new data is saved to the data collection database DB 33b, the CPU 41 creates data collection metadata for the newly saved data in S20 and saves the created data collection metadata to the data collection database DB 33b. The data collection metadata includes, for example, a data ID, a date and time of collection, a data type (e.g., CAN, application operation, video, still image, sensor), a collection condition ID, a collected vehicle ID, and a raw data storage path.
[0087] In S30, the CPU 41 performs preprocessing of the collected data. Preprocessing of the collected data may include, for example, changing the resolution or converting to grayscale. Privacy-conscious processing (for example, mosaic processing) may also be performed as part of the preprocessing of the collected data.
[0088] In S40, the CPU 41 stores the pre-processed collected data (hereinafter referred to as pre-processed collected data) in the collected data DB 33b. The collected data DB 33b comprises a first data storage area for storing collected data and a second data storage area for storing pre-processed collected data.
[0089] In S50, CPU 41 updates the collected data metadata. Specifically, CPU 41 adds the save path of the preprocessed collected data saved in S40 to the collected data metadata.
[0090] In S60, the CPU 41 aggregates the number of collected data for each scene. Specifically, the CPU 41 identifies the number of collected data stored in the collected data DB 33b for each scene based on the collection condition ID associated with the collected data.
[0091] In S70, the CPU 41 updates the collected data metrics based on the aggregation results in S60. The collected data metrics include, for example, the collection condition ID, date and time, and the number of collected data. Specifically, the CPU 41 updates the date and time and the number of collected data in the collected data metrics for each scene.
[0092] CPU 41, in S80, checks the time elapsed from the start of collection to the present (hereinafter referred to as the collection period) and the number of collected data for each collection condition.
[0093] Note that the processing described later in S90 to S140 is performed for each of the multiple scenes.
[0094] In S90, the CPU 41 determines whether the collection period is equal to or greater than the generation determination time predetermined for each collection condition. If the collection period is less than the generation determination time, the CPU 41 terminates the collected data management process. On the other hand, if the collection period is equal to or greater than the generation determination time, the CPU 41 determines in S100 whether the number of collected data is less than the generation determination number predetermined for each collection condition.
[0095] If the number of collected data points is greater than or equal to the number of data points to be generated, the CPU 41 terminates the collected data management process. On the other hand, if the number of collected data points is less than the number of data points to be generated, the CPU 41 determines in S110 whether or not it is necessary to relax the collection conditions. For example, if there is other data that can be included in the collection conditions, it is determined that it is necessary to relax the collection conditions. If it is determined that it is necessary to relax the collection conditions, the CPU 41 starts the collection condition generation process, which will be described later, in S120 and terminates the collected data management process.
[0096] On the other hand, if there is no need to relax the collection conditions, the CPU 41 determines in S130 whether or not it is necessary to change the distribution settings. For example, if it is possible to change the vehicle types to be distributed or the regions to be distributed, it is determined that a change in the distribution settings is necessary. If a change in the distribution settings is not necessary, the CPU 41 terminates the data collection management process. On the other hand, if a change in the distribution settings is necessary, the CPU 41 starts the distribution setting generation process in S140 and terminates the data collection management process. In the distribution setting generation process, the CPU 41 expands the distribution region or increases the number of vehicle types to be distributed.
[0097] Next, we will explain the procedure for the model training process performed by Center 3. The model training process starts when the collected data used to train the machine learning model is updated, or when a predetermined execution cycle has elapsed.
[0098] When the model training process is executed, the CPU 41 of the control unit 31 reads the annotation type included in the collection profile metadata of the collected data used to train the machine learning model in S210, as shown in Figure 8. The collection profile metadata includes the collection condition ID, creation date and time, collection start date and time, scene tag, annotation type, and collection condition status.
[0099] Scene tags include "rainy" and "night," and can be set in combination. The annotation type indicates the type of annotation. The collection condition status indicates whether or not the data has been delivered to the vehicle.
[0100] In S220, the CPU 41 executes an annotation job based on the annotation type read in S210. Executing the annotation job generates annotation results and annotation job data. The annotation job data includes the annotation job ID, start date and time, data ID and collection condition ID pairs, status (e.g., job in progress, successful, failed), and result storage path.
[0101] CPU 41 checks the status included in the annotation job data (hereinafter referred to as annotation status) in S230.
[0102] In S240, CPU 41 determines whether the annotation job is complete based on the confirmation results in S230. If the annotation job is not complete, CPU 41 proceeds to S230. If the annotation job is complete, in S250, the annotation status included in the annotation job data is updated.
[0103] In S260, the CPU 41 saves the annotation results to the storage unit 33.
[0104] In S270, the CPU 41 saves the annotation job data to the storage unit 33.
[0105] In S280, the CPU 41 determines whether the annotation results have been updated. If the annotation results have been updated, the CPU 41 proceeds to S300. If the annotation results have not been updated, the CPU 41 determines in S290 whether the pre-set training execution cycle has elapsed.
[0106] If the training execution cycle has not yet elapsed, the CPU 41 proceeds to S280. On the other hand, if the training execution cycle has elapsed, the CPU 41 proceeds to S300.
[0107] When transitioning to S300, CPU 41 reads the model training settings. The model training settings include the training setting ID, pairs of required data types and quantities, and the path to the server configuration file. The pairs of required data types and quantities indicate the number of new data items needed to start model training. The path to the server configuration file indicates the computing environment used for model training.
[0108] As shown in Figure 9, in S310, the CPU 41 determines whether the start conditions indicated by the loaded model training settings are met. The start conditions indicated by the model training settings are set by pairs of required data types and data quantities.
[0109] If the start condition is not met, the CPU 41 waits until the start condition is met by repeating the process in S310. When the start condition is met, the CPU 41 starts model training in S320. By executing the model training, model training results and model training job data are generated. The model training job data includes the training job ID, start date and time, a list of annotation job IDs, status (e.g., job running, successful, failed), and the training result storage path.
[0110] In S330, the CPU 41 determines whether or not model training is complete. If model training is not complete, the CPU 41 repeats the process in S330 until model training is complete. Once model training is complete, the CPU 41 updates the status of the model training job data in S340.
[0111] In S350, the CPU 41 saves the model training results to the memory unit 33.
[0112] In S360, the CPU 41 saves the annotation job data to the storage unit 33.
[0113] CPU 41 evaluates the model training results in S370. The evaluation of the model training results includes checking whether the machine learning model is working properly by using scenes in which the machine learning model did not work properly as evaluation data, and detecting weaknesses in the machine learning model by checking whether the machine learning model is working properly by adding noise to the annotation results or cropping parts of the image.
[0114] In S380, the CPU 41 determines whether or not it is necessary to update the data collection conditions based on the evaluation results in S370. For example, the CPU 41 determines that it is necessary to update the data collection conditions if it identifies a scene in which the machine learning model does not meet the pre-set evaluation criteria.
[0115] If it is necessary to update the collection conditions, the CPU 41 starts the collection condition generation process described later in S390 and terminates the model learning process. On the other hand, if it is not necessary to update the collection conditions, the CPU 41 sets the machine learning model obtained by model training as a vehicle-distributable model in S400 and terminates the model learning process.
[0116] Next, the procedure for generating collection conditions performed by Center 3 will be explained. The collection condition generation process is started by the process in S120 or S390.
[0117] When the data collection condition generation process is executed, the CPU 41 of the control unit 31 generates new data collection conditions in S510, as shown in Figure 10. Specifically, if the CPU 41 identifies a scene in which the machine learning model does not meet the pre-set evaluation criteria, it generates collection conditions to collect the identified scene. Also, if the CPU 41 detects a weakness in the machine learning model, it generates collection conditions to collect data to overcome that weakness. Furthermore, if it is necessary to relax the collection conditions, the CPU 41 generates data collection conditions by changing thresholds such as the collection start period and the number of data points to be collected to loosen the restrictions.
[0118] In S520, the CPU 41 saves the collection condition file, which indicates the collection conditions generated in S510, to the collection condition DB 33a. As described above, a collection condition profile is attached to the collection condition file.
[0119] In S530, the CPU 41 saves the collection profile metadata of the collection conditions generated in S510 to the collection conditions DB 33a.
[0120] In S540, the CPU 41 determines whether the collection conditions generated in S510 are suitable for distribution to the vehicle. If the collection conditions are not suitable for distribution to the vehicle, the CPU 41 terminates the collection condition generation process. On the other hand, if the collection conditions are suitable for distribution to the vehicle, the CPU 41 sets the collection conditions generated in S510 as collection conditions that can be distributed to the vehicle in S550, and terminates the collection condition generation process.
[0121] Next, the procedure for the distribution process performed by Center 3 will be explained. The distribution process is started when the operator performs an input operation to start the distribution process via the UI unit 107.
[0122] When the distribution process is executed, the CPU 41 of the control unit 31, as shown in Figure 11, sets the machine learning model and collection conditions to be distributed from among the machine learning models and collection conditions that can be distributed to vehicles, based on the input operations performed by the operator via the UI unit 107 in S610.
[0123] The CPU 41, using S620, sets the vehicles to be targeted for distribution (i.e., targets) using individual vehicle data such as the target vehicle type, target region, and frequency of function usage.
[0124] CPU 41 builds the machine learning model and collection conditions set in S610 for distribution at S630.
[0125] CPU 41 starts the process of distributing the machine learning model and data collection conditions built in S630 to the target set in S620 at S640.
[0126] Note that the processing steps S650 to S700, which will be described later, are performed for each of the multiple targets.
[0127] CPU 41 queries the vehicle via S650 to determine whether installation is possible. Conditions where installation is not possible include insufficient storage due to other high-priority applications running while the vehicle is in motion, or the inability to communicate via Wi-Fi. Wi-Fi is a registered trademark.
[0128] The CPU 41 determines in S660 whether the vehicle to be distributed is in a state where it can be installed. If the vehicle is not in a state where it can be installed, the CPU 41 repeats the process of S660 and waits until the vehicle becomes in a state where it can be installed.
[0129] Then, when the vehicle is ready for installation, the CPU 41, in S670, distributes the machine learning model and collection conditions set in S610 to the vehicle to be distributed.
[0130] CPU 41 checks the distribution processing status on S680.
[0131] CPU 41 determines at S690 whether the installation is complete based on the verification results at S680. If the installation is not complete, CPU 41 moves to S680. On the other hand, if the installation is complete, CPU 41 updates the distribution processing status at S700 and terminates the distribution process.
[0132] Next, the procedure for the data acquisition process performed by the control unit 11 of the data acquisition device 2 will be explained. The data acquisition process is a process that is repeatedly performed while the control unit 11 is operating.
[0133] When data acquisition processing is performed, the CPU 21 of the control unit 11 performs rule-based judgment in S810, as shown in Figure 12. In rule-based judgment, the CPU 21 makes a determination based on a numerical value, for example, "whether or not the value detected by the sensor is above a threshold."
[0134] CPU 21 performs ambiguous condition determination in S820. In ambiguous condition determination, CPU 21 uses an AI application to determine conditions that are difficult to express numerically, such as "heavy rain" or "light rain," using images captured by the in-vehicle camera.
[0135] In S830, the CPU 21 determines whether the collection start condition included in the collection conditions has been met, based on the judgment results of S810 and S820. The collection conditions include a collection data item indicating the type of data to be collected, a collection start condition for starting data collection, data sampling conditions, video data trimming conditions, and an upload start condition indicating the timing for uploading the data.
[0136] The types of data collected include, for example, in-vehicle data such as camera data, sensor data, CAN data, and AI logic operation logs.
[0137] The conditions for initiating data collection can be set by, for example, a driver overwrite operation, a rule-based system based on CAN data values, the results of application operations, a specific scene, or a combination of these conditions.
[0138] Driver overrides include, for example, the driver pressing the brake pedal while autonomous driving is in progress, the driver taking control of the steering wheel while autonomous driving is in progress, switching the wipers to manual mode while they are on, canceling music shuffle playback, canceling voice recognition while it is in use, or canceling suggestions from the in-vehicle agent. These are all actions taken by the driver to manually intervene in functions that the vehicle system is automatically controlling.
[0139] The rule-based criteria, based on CAN data values, include, for example, illuminance falling below a predetermined value, or acceleration exceeding a predetermined value. Specific scenes include, for example, daytime, nighttime, rain, fog, and snow.
[0140] If the collection conditions are not met, the CPU 21 proceeds to S870. On the other hand, if the collection conditions are met, the CPU 21 collects in-vehicle data in S840 based on the collection conditions.
[0141] CPU 21 converts the collected data into an upload format using S850.
[0142] CPU 21 sends the data, which has been converted to upload format by S860, to the uploader and then moves on to S870.
[0143] When the process moves to S870, it is determined whether the upload start condition included in the collection conditions has been met. If the upload start condition has not been met, the CPU 21 terminates the data collection process. On the other hand, if the upload start condition has been met, the CPU 21, at S880, uploads the collected data corresponding to the collection conditions for which the upload start condition has been met, along with the collection condition ID corresponding to the collection conditions, to Center 3 via the uploader, and terminates the data collection process. The collected data uploaded to Center 3 is the data that has been converted to the upload format at S850.
[0144] The user (for example, a test engineer) creates test cases for each machine learning model by performing input operations via the UI unit 107, as shown by arrow L21 in Figure 13.
[0145] A test case includes at least a target scene, a target evaluation method, and a target evaluation criterion. In this embodiment, the target scene is the scene targeted in the test case, for example, "bright night" or "dark night." The target evaluation method is an evaluation method for the target scene, for example, the F1 score. The target evaluation criterion is an evaluation criterion for the target scene, for example, that the F1 score is 0.8 or higher. The target evaluation method may also be, for example, IoU. IoU is an abbreviation for Intersection over Union.
[0146] When a test case is input via the UI unit 107, the center 3, as indicated by arrow L22, refers to the collection condition file stored in the collection condition DB 33a to determine whether or not a scene tag matching the target scene included in the input test case is included in the collection condition file.
[0147] If a scene tag matching the target scene is included in the collection conditions file, Center 3 saves the entered test case to Test Case DB33c, as indicated by arrow L23. This registers the test case in Center 3.
[0148] As shown by arrow L31 in Figure 14, the data acquisition device 2 uploads data that matches the collection conditions included in the collection conditions file to the center 3.
[0149] Center 3, as indicated by arrow L32, searches for scenes corresponding to the collection condition ID attached to the uploaded data, classifies the uploaded data by scene, and saves it to the collected data DB33b. Figure 14 shows the state in which the data stored in the collected data DB33b is classified into "bright night" scenes and "dark night" scenes.
[0150] Center 3 extracts feature quantities from each of the multiple classified data for each scene, as indicated by arrow L33 and process P1, and calculates statistical features using the extracted feature quantities. In this embodiment, the feature quantities are brightness and time, and the statistical features are brightness range and time period.
[0151] Center 3 stores the uploaded data, the extracted features, and the calculated statistical features for each scene in the Scene Feature Extracted Data DB33d, as indicated by arrows L34 and L35. In this embodiment, Center 3 classifies the data and features for each scene by storing each of the data in association with the extracted features and scene tags.
[0152] Center 3, as shown by arrow L41 in Figure 15, selects multiple data sets from the multiple data stored in the scene feature extraction data DB33d to be used for training the machine learning model, and saves the selected multiple data sets as a training dataset DB33e.
[0153] Center 3, as indicated by arrow L42, selects multiple data sets from the multiple data stored in the scene feature extraction data DB33d to be used for evaluating the machine learning model, and saves the selected multiple data sets as an evaluation dataset DB33f.
[0154] Center 3, as indicated by arrow L43, links the evaluation dataset stored in evaluation dataset DB33f with the test cases corresponding to the evaluation dataset.
[0155] Furthermore, the data sets that make up the training dataset and the data sets that make up the evaluation dataset are different from each other. For example, at least one of the following differs between the training dataset and the evaluation dataset: the vehicle used to collect the data, the roads traveled during data collection, and the date and time the data was collected.
[0156] Center 3 performs model training using the training dataset, as indicated by process P2 and arrow L44. As a result, Center 3 generates a machine learning model 200, as indicated by arrow L45.
[0157] Center 3, as indicated by process P3 and arrows L46 and L47, instructs the generated machine learning model 200 to identify scenes for each of the multiple data sets that make up the evaluation dataset. Center 3 then evaluates the identification results of the machine learning model 200 for each scene based on the target evaluation method and target evaluation criteria of the test cases associated with the evaluation dataset.
[0158] Center 3 saves the evaluation results from process P3 to evaluation result DB33g, as indicated by arrow L48.
[0159] Figure 15 shows that all three data points classified as "bright night" scenes received a "correct" scene determination result and an overall evaluation of "good."
[0160] Figure 15 shows that of the three data points classified as "dark night" scenes, one data point had a "correct" scene determination result, while the other two data points had "incorrect" scene determination results, resulting in an overall evaluation of "poor."
[0161] In this way, Center 3 evaluates whether each data point is correct or not, and also makes an overall evaluation of whether each scene is good or bad.
[0162] Center 3 determines that the overall evaluation of a target scene is "good" if, for example, the percentage of correct data in the target scene (e.g., "bright night," "dark night") is equal to or greater than a predetermined evaluation threshold (e.g., 90%).
[0163] Center 3 extracts feature quantities from data where the scene judgment result is "incorrect," calculates statistical features using one or more extracted feature quantities, and inputs the calculated statistical features into the collection condition generation AI 150.
[0164] However, if the distribution of features for data where the scene judgment result is "incorrect" is uniform within the scene, Center 3, as indicated by arrow L50, inputs the statistical features stored in the scene feature extraction data DB33d for scenes where the overall evaluation is "poor" into the collection condition generation AI150.
[0165] As indicated by arrow L51, the data collection condition generation AI 150 generates data collection conditions for collecting scene data corresponding to the input statistical features. In Figure 15, because the overall evaluation of "bright night" in the model evaluation is "good" and the overall evaluation of "dark night" in the model evaluation is "poor", the data collection condition generation AI 150 generates data collection conditions for collecting "dark night" data.
[0166] When new collection conditions are generated, Center 3 distributes the generated collection conditions to multiple data acquisition devices 2.
[0167] Having received the data collection conditions for collecting data for "dark nights," the data collection device 2 uploads new data for "dark nights." This generates new training and evaluation datasets for "dark nights."
[0168] Center 3 generates machine learning model 200 by training the model using the "dark night" training dataset.
[0169] Center 3 will evaluate the generated machine learning model 200 using both the evaluation dataset for "bright nights" and the evaluation dataset for "dark nights."
[0170] Next, we will explain the procedure for the test case registration process performed by Center 3. The test case registration process is a process that is repeatedly executed while the control unit 31 is operating.
[0171] When the test case registration process is executed, the CPU 41 of the control unit 31 determines in S910, as shown in Figure 16, whether or not a test case has been input via the UI unit 107. If no test case has been input, the CPU 41 terminates the test case registration process.
[0172] On the other hand, when a test case is input, the CPU 41, in S920, obtains a scene tag from the collection condition DB 33a that matches the target scene included in the input test case.
[0173] CPU 41 determines at S930 whether or not a scene tag matching the target scene has been obtained. If no scene tag matching the target scene has been obtained, CPU 41 terminates the test case registration process.
[0174] On the other hand, if a scene tag matching the target scene is obtained, the CPU 41 registers the test case by saving the input test case to the test case DB 33c in S940, and then terminates the test case registration process.
[0175] Next, the procedure for the automatic collection condition generation process performed by Center 3 will be explained. The automatic collection condition generation process is a process that is repeatedly executed while the control unit 31 is operating.
[0176] When the automatic collection condition generation process is executed, the CPU 41 of the control unit 31 determines in S1010 whether or not it has received collected data from the data collection device 2, as shown in Figure 17.
[0177] If no data has been received, the CPU 41 terminates the automatic generation of collection conditions process. On the other hand, if data has been received, in S1020, the CPU 41 searches the collection condition DB 33a for a scene corresponding to the collection condition ID, based on the collection condition ID attached to the received data.
[0178] In S1030, the CPU 41 classifies the collected data received in S1010 based on the scenes obtained through the search in S1020 and stores it in the collected data DB 33b. As a result, the collected data is stored in the collected data DB 33b in a state classified by scene.
[0179] In S1040, CPU 41 extracts features from each of the multiple collected data that have been classified for each scene, and calculates statistical features using the extracted features.
[0180] In step S1050, CPU 41 saves multiple collected data, multiple extracted features, and calculated statistical features for each scene into the scene feature extraction data DB 33d.
[0181] In S1060, CPU 41 performs annotation on multiple collected data stored in DB 33d, which contains scene feature extraction data.
[0182] In S1070, CPU 41 selects multiple data sets from the multiple collected data stored in the scene feature extraction data DB 33d to be used for training the machine learning model, and saves the selected multiple collected data sets as a training dataset in the training dataset DB 33e.
[0183] In S1080, CPU 41 selects multiple data sets from the multiple collected data stored in the scene feature extraction data DB33d to be used for evaluating the machine learning model, and saves the selected multiple collected data sets as an evaluation dataset in the evaluation dataset DB33f.
[0184] In S1090, CPU 41 performs training using the training dataset. This generates a machine learning model 200.
[0185] In S1100, the CPU 41 performs model evaluation. Specifically, the CPU 41 has the generated machine learning model 200 determine the scene for each of the multiple collected data that make up the evaluation dataset, and evaluates the scene determination result of the machine learning model 200 for each scene based on the target evaluation method and target evaluation criteria of the test cases associated with the evaluation dataset.
[0186] In S1110, the CPU 41 determines whether or not there are any scenes with an overall evaluation of "poor". If there are no scenes with an overall evaluation of "poor", the CPU 41 terminates the automatic generation of collection conditions. On the other hand, if there are scenes with an overall evaluation of "poor", the CPU 41 determines in S1120 whether or not the distribution of features of the collected data for which the scene judgment result is "incorrect" is uniform within the scene.
[0187] Here, if the distribution of features in the collected data for which the scene judgment result is "incorrect" is uniform within the scene, the CPU 41 obtains the calculated statistical features in S1130 and proceeds to S1150. That is, for scenes whose overall evaluation is "poor", the CPU 41 obtains the statistical features stored in the scene feature extraction data DB 33d.
[0188] On the other hand, if the distribution of features in data where the scene determination result is "incorrect" is not uniform within the scene, the CPU 41 obtains the features of the collected data where the scene determination result is "incorrect" in S1140, calculates statistical features using one or more of the obtained features, and proceeds to S1150.
[0189] When the process moves to S1150, the CPU 41 inputs the statistical features acquired in S1130 or calculated in S1140 into the collection condition generation AI 150, generating collection conditions for collecting scene data corresponding to the input statistical features, and then terminates the automatic collection condition generation process.
[0190] Next, the procedure for the automatic distribution process performed by Center 3 will be explained. The automatic distribution process is a process that is repeatedly executed while the control unit 31 is operating.
[0191] When the automatic distribution process is executed, the CPU 41 of the control unit 31 determines in S1210, as shown in Figure 18, whether or not new collection conditions have been generated by the collection condition generation AI 150 in the automatic collection condition generation process. If no new collection conditions have been generated by the collection condition generation AI 150, the CPU 41 terminates the automatic distribution process.
[0192] On the other hand, if new collection conditions are generated by the collection condition generation AI 150, the CPU 41 distributes the newly generated collection conditions to the data collection device 2 of the vehicle to which the newly generated collection conditions are to be distributed in S1220, and terminates the automatic distribution process.
[0193] When Center 3 receives registration of test cases from the user, including the target scene, target evaluation method, and target evaluation criteria, it extracts the feature quantities of each of the multiple collected data stored for each scene, as shown in processing P11 of Figure 19.
[0194] Next, Center 3 performs annotation on the multiple collected data stored for each scene, as shown in process P12.
[0195] Next, Center 3, as shown in process P13, trains a machine learning model using multiple collected data stored for each scene.
[0196] Next, as shown in process P14, Center 3 performs a model evaluation on the target scene based on the target evaluation criteria for the machine learning model generated by training.
[0197] Next, as shown in processing P15, Center 3 automatically generates collection conditions based on the features of the collected data where the scene judgment result is "incorrect" in target scenes with an overall evaluation of "poor".
[0198] Next, Center 3 distributes the automatically generated collection conditions to multiple data collection devices 2, as shown in process P16.
[0199] Next, as shown in process P17, the data acquisition device 2 collects vehicle data based on the newly distributed collection conditions.
[0200] Next, as shown in processing P18, the data acquisition device 2 uploads the collected vehicle data to the center 3.
[0201] Next, Center 3 extracts the features of the newly uploaded collected data, as shown in processing P11.
[0202] In this way, when the data collection system 1 receives registration of target scenes and target evaluation criteria from the user, it repeats a cycle in which it sequentially executes processes P11 to P18 until the overall evaluation is satisfactory, without requiring input of target scenes and target evaluation criteria from the user in each cycle.
[0203] The data acquisition system 1 configured in this way comprises a plurality of data acquisition devices 2 and a center 3.
[0204] Multiple data acquisition devices 2 are installed in each of the multiple vehicles and are configured to transmit vehicle data that includes at least information about the vehicle on which they are installed. Center 3 is configured to receive vehicle data from the multiple data acquisition devices 2.
[0205] Center 3 stores vehicle data received from multiple data acquisition devices 2.
[0206] Center 3 uses the saved vehicle data to perform machine learning on a machine learning model, including annotation, model training, and model evaluation. Based on the results of the model evaluation obtained, Center 3 identifies scenes in which the machine learning model does not meet the pre-set evaluation criteria. For the identified scenes, Center 3 generates collection condition data that includes collection data items indicating the type of vehicle data to be collected and collection start conditions for initiating data collection. Center 3 then stores the generated collection condition data in association with scene tags that identify the scenes.
[0207] Center 3 distributes the collection condition data to multiple data collection devices 2.
[0208] Center 3 stores vehicle data uploaded from multiple data collection devices 2 based on collection condition data, linking it to scene tags.
[0209] Center 3 accepts registrations from users for target scenes, which are the scenes to be evaluated for model evaluation, and target evaluation criteria, which are the evaluation criteria for the model evaluation of the target scenes. It then performs model evaluation of the target scenes based on the target evaluation criteria, and if the overall evaluation of the model evaluation of the target scenes is poor, it generates collection condition data for collecting vehicle data for the target scenes.
[0210] Because Center 3 can collect vehicle data by narrowing it down by scene, it can efficiently collect data from multiple data acquisition devices 2. Furthermore, Center 3 can efficiently collect data from multiple data acquisition devices 2 by identifying scenes where the overall evaluation of the model is poor and generating collection condition data.
[0211] Furthermore, Center 3 generates collection condition data based on the feature quantities of one or more vehicle data in the target scene where the overall model evaluation was judged to be poor. In this embodiment, the feature quantities are brightness and time. This allows Center 3 to identify the vehicle data to be collected by feature quantities and efficiently acquire the desired vehicle data.
[0212] Furthermore, Center 3 generates collection condition data based on the statistical features of one or more vehicle data points that do not meet the target evaluation criteria among the one or more vehicle data points in the target scene that were evaluated as having a poor overall model evaluation. In this embodiment, the statistical features are the brightness range and the time period. This allows Center 3 to represent the features of the vehicle data to be collected using statistical features, thereby simplifying the collection condition data.
[0213] Furthermore, when Center 3 receives registration of target scenes and target evaluation criteria from a user, it sequentially repeats a cycle of performing model evaluations on the target scenes based on the target evaluation criteria, generating collection condition data based on the evaluation results of the model evaluation, distributing the generated collection condition data to multiple data collection devices 2, saving vehicle data received from the multiple data collection devices 2, and performing annotation and model training using the saved vehicle data, without requiring the user to input the target scenes and target evaluation criteria in each of the multiple cycles. As a result, Center 3 can minimize the amount of input operations required from the user for target scenes and target evaluation criteria, thereby improving user convenience.
[0214] Furthermore, when Center 3 receives registrations for "bright night" and "dark night" as target scenes, it performs model evaluations for both "bright night" and "dark night." If the overall evaluation for "bright night" is good and the overall evaluation for "dark night" is poor, and collection condition data for "dark night" is generated, Center 3 performs a model evaluation of the machine learning model 200, which has been trained using the uploaded vehicle data for "dark night" based on the generated collection condition data, for both "bright night" and "dark night." This allows Center 3 to detect if the model evaluation of the machine learning model 200 for "bright night" deteriorates due to the model training of the machine learning model 200 for "dark night" being performed again.
[0215] Center 3 also extracts feature vectors for each of the multiple vehicle data sets in the target scene and stores the vehicle data linked to the feature vectors and scene tags. This makes it easy for Center 3 to extract vehicle data for each target scene.
[0216] Center 3 also accepts registrations from users for target scenes, which are the scenes to be evaluated by the model, and target evaluation criteria, which are the evaluation criteria for the model evaluation of the target scenes. Center 3 then performs a model evaluation of the target scenes based on the target evaluation criteria, and if the overall evaluation of the model evaluation of the target scenes is poor, it generates a machine learning model 200 by performing model training using vehicle data collected by a data collection method that generates collection condition data for collecting vehicle data for the target scenes.
[0217] Since such a center 3 can efficiently collect data from multiple data collection devices 2, it can efficiently generate machine learning models 200.
[0218] In the embodiments described above, the data acquisition device 2 corresponds to an in-vehicle device, the scene tag corresponds to scene identification information, the overall evaluation corresponds to the evaluation result of the model evaluation, "bright night" corresponds to the first scene, and "dark night" corresponds to the second scene.
[0219] Although one embodiment of the present disclosure has been described above, the present disclosure is not limited to the above embodiment and can be implemented in various modified forms.
[0220] [Modification 1] In the above embodiment, a form was shown in which the user registers both the target scene and the target evaluation criteria. However, there are also cases in which the user registers only one of the target scene or the target evaluation criteria. For example, if there is only one scene to be judged in the machine learning model, registration of the target scene is unnecessary. Also, if the target evaluation criteria have been set in advance for the target scene, registration of the target evaluation criteria is unnecessary.
[0221] [Modification 2] In the above embodiment, the statistical features were shown to be in the range from the lower limit to the upper limit of the features of vehicle data that do not meet the target evaluation criteria. However, the statistical features are not limited to the range from the lower limit to the upper limit, and may be, for example, quartiles.
[0222] The control units 11, 31 and their methods described in this disclosure may be implemented by a dedicated computer provided by configuring a processor and memory programmed to execute one or more functions embodied by a computer program. Alternatively, the control units 11, 31 and their methods described in this disclosure may be implemented by a dedicated computer provided by configuring a processor by one or more dedicated hardware logic circuits. Alternatively, the control units 11, 31 and their methods described in this disclosure may be implemented by one or more dedicated computers configured by a combination of a processor and memory programmed to execute one or more functions and a processor configured by one or more hardware logic circuits. Furthermore, the computer program may be stored as instructions executed by the computer on a computer-readable non-transitional tangible recording medium. The methods for realizing the functions of each part included in the control units 11, 31 do not necessarily need to include software, and all of its functions may be realized using one or more hardware components.
[0223] Multiple functions of one component in the above embodiment may be realized by multiple components, or one function of one component may be realized by multiple components. Furthermore, multiple functions of multiple components may be realized by one component, or one function realized by multiple components may be realized by one component. Also, some parts of the configuration of the above embodiment may be omitted. Furthermore, at least some parts of the configuration of the above embodiment may be added to or replaced with the configuration of other above embodiments.
[0224] In addition to the data acquisition device 2 and center 3 described above, this disclosure can also be realized in various forms, such as a system comprising the data acquisition device 2 and center 3, a program for causing the computer to function as the data acquisition device 2 and center 3, a non-transitional physical recording medium such as semiconductor memory on which this program is recorded, and a data acquisition method. [Technical Concept Disclosed in This Specification] [Item 1] A data collection method performed in a center (3) configured to receive vehicle data, which includes at least information about the vehicle, from a plurality of in-vehicle devices (2) installed in each of a plurality of vehicles, comprising: storing the vehicle data received from the plurality of in-vehicle devices; using the stored vehicle data, identifying scenes in which the machine learning model does not meet pre-set evaluation criteria based on the results of a model evaluation obtained by performing machine learning, including annotation, model training, and model evaluation, on a machine learning model; generating collection condition data for the identified scenes, which includes a collection data item indicating the type of vehicle data to be collected and a collection start condition for starting data collection; storing the generated collection condition data in association with scene identification information that identifies the scene; distributing the collection condition data to the plurality of in-vehicle devices; and storing the vehicle data uploaded from the plurality of in-vehicle devices based on the collection condition data, in association with the scene identification information. A data collection method that accepts registration from a user of at least one of a target scene, which is a scene to be evaluated for the model evaluation, and a target evaluation criterion, which is an evaluation criterion for the model evaluation of the target scene; performs the model evaluation of the target scene based on the target evaluation criterion; and generates collection condition data for collecting the vehicle data of the target scene if the evaluation result of the model evaluation of the target scene is unsatisfactory.
[0225] [Item 2] A data collection method described in Item 1, wherein the data collection method accepts registration of both the target scene and the target evaluation criteria.
[0226] [Item 3] A data collection method according to Item 1 or Item 2, wherein the data collection method generates the collection condition data based on the feature quantities of one or more vehicle data in the target scene in which the evaluation result was evaluated as poor.
[0227] [Item 4] A data collection method according to Item 3, wherein the data collection method generates the collection condition data based on the statistics of the feature quantities of one or more of the vehicle data in the target scene that were evaluated as having a poor evaluation result and which do not meet the target evaluation criteria.
[0228] [Item 5] A data collection method according to any one of Items 1 to 4, wherein, upon receiving registration of the target scene and the target evaluation criteria from the user, the method repeats a cycle in which the following are performed sequentially until the evaluation result is satisfactory: performing the model evaluation on the target scene based on the target evaluation criteria; generating the collection condition data based on the evaluation result of the model evaluation; distributing the generated collection condition data to a plurality of the in-vehicle devices; saving the vehicle data received from the plurality of the in-vehicle devices; and performing the annotation and model training using the saved vehicle data, without requiring input of the target scene and the target evaluation criteria from the user in each of the plurality of cycles.
[0229] [Item 6] A data collection method according to Item 5, wherein when registration of a first scene and a second scene is accepted as the target scenes, the model evaluation is performed for each of the first scene and the second scene, and when the evaluation result for the first scene in the model evaluation is good and the evaluation result for the second scene in the model evaluation is poor, the model evaluation of the machine learning model, which has been trained using the vehicle data of the second scene uploaded based on the generated collection condition data, is performed for both the first scene and the second scene.
[0230] [Item 7] A data collection method according to Item 3 or Item 4, wherein the feature quantities are extracted for each of the multiple vehicle data in the target scene, and the vehicle data is stored linked to the feature quantities and the scene identification information.
[0231] [Item 8] A center (3) configured to receive vehicle data, which includes at least information about the vehicle on which it is installed, from a plurality of in-vehicle devices (2) installed in each of a plurality of vehicles, stores the vehicle data received from the plurality of in-vehicle devices, the center uses the stored vehicle data to perform machine learning on a machine learning model, including annotation, model training, and model evaluation, and based on the results of the model evaluation obtained, identifies scenes in which the machine learning model does not meet pre-set evaluation criteria, generates collection condition data for the identified scenes, which includes a collection data item indicating the type of vehicle data to be collected and a collection start condition for starting data collection, stores the generated collection condition data in association with scene identification information that identifies the scene, distributes the collection condition data to the plurality of in-vehicle devices, and stores the vehicle data uploaded from the plurality of in-vehicle devices based on the collection condition data, in association with the scene identification information. A model generation method comprising: the center receiving registration from a user of at least one of a target scene which is a scene to be evaluated by the model and target evaluation criteria which are evaluation criteria for the model evaluation of the target scene; performing the model evaluation of the target scene based on the target evaluation criteria; and, if the evaluation result of the model evaluation of the target scene is poor, performing the model training using the vehicle data collected by the data collection method which generates collection condition data for collecting the vehicle data of the target scene, thereby generating a machine learning model (200).
Claims
1. A data collection method performed in a center (3) configured to receive vehicle data, which includes at least information about the vehicle, from a plurality of in-vehicle devices (2) installed in each of a plurality of vehicles, comprising: storing the vehicle data received from the plurality of in-vehicle devices; using the stored vehicle data, performing machine learning, including annotation, model training, and model evaluation, on a machine learning model to identify scenes in which the machine learning model does not meet pre-set evaluation criteria based on the results of the model evaluation obtained; generating collection condition data for the identified scenes, which includes collection data items indicating the type of vehicle data to be collected and collection start conditions for starting data collection; storing the generated collection condition data in association with scene identification information that identifies the scene; distributing the collection condition data to the plurality of in-vehicle devices; and storing the vehicle data uploaded from the plurality of in-vehicle devices based on the collection condition data, in association with the scene identification information. A data collection method that accepts registration from a user of at least one of a target scene, which is a scene to be evaluated for the model evaluation, and a target evaluation criterion, which is an evaluation criterion for the model evaluation of the target scene; performs the model evaluation of the target scene based on the target evaluation criterion; and generates collection condition data for collecting the vehicle data of the target scene if the evaluation result of the model evaluation of the target scene is unsatisfactory.
2. A data collection method according to claim 1, wherein the data collection method accepts registration of both the target scene and the target evaluation criteria.
3. A data collection method according to claim 1, wherein the data collection method generates the collection condition data based on the feature quantities of one or more vehicle data in the target scene in which the evaluation result was evaluated as poor.
4. A data collection method according to claim 3, wherein the data collection method generates the collection condition data based on the statistics of the feature quantities of one or more of the vehicle data in the target scene that were evaluated as having a poor evaluation result and which do not satisfy the target evaluation criteria.
5. A data collection method according to any one of claims 1 to 4, wherein, upon receiving registration of the target scene and the target evaluation criteria from the user, the method repeats a cycle in which, until the evaluation result is satisfactory, the method sequentially repeats the following: performing the model evaluation on the target scene based on the target evaluation criteria; generating collection condition data based on the evaluation result of the model evaluation; distributing the generated collection condition data to a plurality of in-vehicle devices; saving the vehicle data received from the plurality of in-vehicle devices; and performing the annotation and model training using the saved vehicle data, without requiring input of the target scene and the target evaluation criteria from the user in each of the plurality of cycles.
6. A data collection method according to claim 5, wherein when registration of a first scene and a second scene is accepted as the target scenes, the model evaluation is performed for each of the first scene and the second scene, and when the evaluation result for the first scene in the model evaluation is good and the evaluation result for the second scene in the model evaluation is poor, the data collection method is performed for both the first scene and the second scene, and the model evaluation of the machine learning model, which has been trained using the vehicle data of the second scene uploaded based on the generated collection condition data, is performed for both the first scene and the second scene.
7. A data collection method according to claim 3 or claim 4, comprising extracting the feature quantities for each of the multiple vehicle data in the target scene, and storing the vehicle data linked to the feature quantities and the scene identification information.
8. A center (3) configured to receive vehicle data, which includes at least information about the vehicle on which it is installed, from a plurality of in-vehicle devices (2) installed in each of a plurality of vehicles, stores the vehicle data received from the plurality of in-vehicle devices, the center uses the stored vehicle data to perform machine learning on a machine learning model, including annotation, model training, and model evaluation, and based on the results of the model evaluation obtained, identifies scenes in which the machine learning model does not meet pre-set evaluation criteria, generates collection condition data for the identified scenes, which includes a collection data item indicating the type of vehicle data to be collected and a collection start condition for starting data collection, stores the generated collection condition data in association with scene identification information that identifies the scene, distributes the collection condition data to the plurality of in-vehicle devices, and stores the vehicle data uploaded from the plurality of in-vehicle devices based on the collection condition data, in association with the scene identification information. A model generation method comprising: the center receiving registration from a user of at least one of a target scene which is a scene to be evaluated by the model and target evaluation criteria which are evaluation criteria for the model evaluation of the target scene; performing the model evaluation of the target scene based on the target evaluation criteria; and, if the evaluation result of the model evaluation of the target scene is poor, performing the model training using the vehicle data collected by the data collection method which generates collection condition data for collecting the vehicle data of the target scene, thereby generating a machine learning model (200).