Information processing apparatus and method

The system efficiently collects images for retraining machine learning models by using local image features from sensors to automate the selection process, enhancing accuracy and reducing costs.

JP7861737B2Active Publication Date: 2026-05-19TOYOTA JIDOSHA KK
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
TOYOTA JIDOSHA KK
Filing Date
2023-08-30
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing methods for collecting learning images for machine learning models used in image recognition are inefficient and lack consistency in selecting images for retraining, leading to high human costs and varying selection criteria.

Method used

A system that utilizes local image features from a first sensor to identify and transmit similar images to a server for training data, using a control unit to acquire and store these images efficiently, reducing human intervention and ensuring consistent selection criteria.

Benefits of technology

This approach allows for the efficient collection of images suitable for retraining machine learning models, improving estimation accuracy and reducing human costs by automating the selection process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007861737000001
    Figure 0007861737000001
  • Figure 0007861737000002
    Figure 0007861737000002
  • Figure 0007861737000003
    Figure 0007861737000003
Patent Text Reader

Abstract

To efficiently collect images for learning of a machine learning model for use in image recognition.SOLUTION: A terminal receives a local image feature amount of one or more first objects included in a first image from a server, acquires a local image feature amount of one or more objects included in a captured image for each of a plurality of captured images of a camera on the basis of a detection result of objects by a first sensor that detects the objects by emitting a predetermined signal for a range including an imaging range of the camera, and transmits, among the plurality of captured images, a second captured image, in which the local image feature amount of the one or more objects included in the captured image is similar to the local image feature amount of one or more first objects, to the server as one of learning data for a machine learning model for use in image recognition. The server transmits the local image feature amount of the one or more first objects to the terminal, receives the second captured image from the terminal, and stores the second captured image in a storage section as one of the learning data for the machine learning model.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a machine learning model used for image recognition.

Background Art

[0002] There is disclosed a database construction system that automatically adjusts and collects supervised learning data for machine learning to perform object recognition from the output of another sensor using the detection result of a certain sensor as supervised data, and constructs a database of supervised learning data (for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present disclosure is to provide an information processing apparatus and a method capable of efficiently collecting learning images for a machine learning model used for image recognition.

Means for Solving the Problems

[0005] One aspect of the present disclosure is receiving local image feature amounts of one or more first objects included in a first image from a server, acquiring a detection result of an object by a first sensor that transmits a predetermined signal to detect an object for a range including an imaging range of a camera, acquiring local image feature amounts of one or more objects included in each of a plurality of captured images of the camera based on the detection result of the object by the first sensor, Among the plurality of captured images, a second captured image in which the local image features of one or more objects included in the captured image are similar to the local image features of the one or more first objects is transmitted to the server as one of the training data for a machine learning model used for image recognition. A control unit that executes This is an information processing device equipped with [a specific feature / feature].

[0006] Another aspect of this disclosure is, Transmitting local image features of one or more first objects contained in the first image to the first terminal, The first terminal receives a second image captured by the camera, wherein the local image features of one or more objects included in the image are similar to the local image features of the one or more first objects. The second captured image is stored in the memory unit as one of the training data for a machine learning model used for image recognition, A control unit that executes Equipped with, The local image features of one or more objects included in the aforementioned image are acquired based on the object detection result by a first sensor that detects objects by transmitting a predetermined signal over a range including the imaging range of the camera. It is an information processing device.

[0007] Another aspect of this disclosure is, The device, The server receives local image features of one or more first objects contained in the first image, The process involves obtaining the object detection result from a first sensor that detects objects by transmitting a predetermined signal within the range including the camera's imaging range, Based on the object detection result by the first sensor, local image features of one or more objects included in each of the multiple images captured by the camera are obtained. Among the plurality of captured images, a second captured image in which the local image features of one or more objects included in the captured image are similar to the local image features of the one or more first objects is transmitted to the server as one of the training data for a machine learning model used for image recognition. Execute, The aforementioned server, Transmitting local image feature quantities of the one or more first objects to the terminal, The terminal receives the second captured image, The second captured image is stored in the memory unit as one of the training data for the machine learning model, Execute It is a method. [Effects of the Invention]

[0008] According to this disclosure, it is possible to efficiently collect images for training machine learning models used in image recognition. [Brief explanation of the drawing]

[0009] [Figure 1] Figure 1 shows an example of the system configuration of a data collection system for training a machine learning model according to the first embodiment, and an example of the hardware configuration of the center server and vehicle. [Figure 2] Figure 2 shows an example of the data collection process for training according to the first embodiment. [Figure 3] Figure 3 shows an example of the functional configuration of the central server and the vehicle. [Figure 4] Figure 4 shows an example of the information included in the collection plan. [Figure 5] Figure 5 shows an example of a flowchart for the data collection process for retraining on the central server. [Figure 6] Figure 6 shows an example of a flowchart for the process of collecting data for vehicle retraining. [Modes for carrying out the invention]

[0010] A machine learning model can improve the accuracy of estimation by learning. For example, when a machine learning model performs image recognition to determine the type of an object captured in a captured image and its position in the captured image, the recognition result is evaluated based on whether the type and position of the object are correct or not. For re-learning to improve the estimation accuracy of the machine learning model used for image recognition, it is required to select images similar to those with poor recognition results and create teacher data corresponding to the selected images.

[0011] The selection of images for re-learning is often done by humans. However, for example, when selecting re-learning images from the video of an in-vehicle camera, a video of a predetermined time length must be viewed, which results in high human costs. Also, since there are variations in the criteria for determining images similar to those with poor recognition results among humans, it is difficult to ensure the consistency of the selection criteria for re-learning images. That is, the collection of learning images for a machine learning model that performs image recognition processing may be inefficient.

[0012] One aspect of the present disclosure, in view of the above problems, notifies a terminal of local image feature amounts of one or more first objects included in an image to be learned by a machine learning model, and causes the server to transmit a captured image of a camera in which the local image feature amounts of the objects included in the image are similar to the local image feature amounts of the first objects. As a result, learning images that are similar to the images to be learned and are beneficial for improving the accuracy of the machine learning model can be efficiently collected.

[0013] More specifically, one aspect of the present disclosure is an information processing device comprising a control unit. The control unit receives local image features of one or more first objects contained in a first image from a server. The control unit obtains the detection result of an object by a first sensor that detects objects by transmitting a predetermined signal over a range including the imaging range of a camera. Based on the detection result of the object by the first sensor, the control unit obtains local image features of one or more objects contained in each of a plurality of images captured by the camera. The control unit transmits to the server a second image from among the plurality of images in which the local image features of one or more objects contained in the image are similar to the local image features of one or more first objects, as one of the training data for a machine learning model used for image recognition.

[0014] The information processing device in one of these embodiments is, for example, a computer connected to an in-vehicle camera such as an ECU or data communication device (DCM) mounted on a vehicle, or a computer connected to a surveillance camera, etc. Vehicles include, for example, automobiles, motorcycles, bicycles, and railway vehicles. Furthermore, the information processing device in one of these embodiments may be a computer connected to a camera mounted on, for example, an aircraft or a ship. Also, the information processing device in one of these embodiments may be a computer such as a smartphone or a tablet terminal. The control unit includes, for example, a processor such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate). It may also be a circuit such as an array.

[0015] The camera may be provided with the information processing device, or it may be a separate device connected to the information processing device. Alternatively, the camera may be neither provided with nor connected to the information processing device. The first sensor is, for example, a millimeter-wave radar, sonar, or LiDAR (Light Detection and Ranging). The specified signal used is a signal transmitted via radio waves or sound waves.

[0016] Machine learning models include, for example, convolutional neural networks. This is a Network (CNN) model. However, in one aspect of this disclosure, the machine learning model is not limited to a specific algorithm.

[0017] Local image features are features corresponding to a single object within an image. Examples of local image features include values ​​related to the object's color and shape. More specifically, local image features include AKAZE, SIFT (Scaled Invariance Feature Transform), and HOG. Examples include Histograms of Oriented Gradients (Histograms of Oriented Gradients), SURF (Speeded-Up Robust Features), and LBP (Local Binary Pattern). However, local image features are not limited to these. Furthermore, local image features are not limited to one type of feature; multiple types of features may be used.

[0018] According to one aspect of this disclosure, the information processing device uses a plurality of images captured by a camera to process the images of the camera. A second image, similar to the first image, is sent to the server as one of the data points for training the machine learning model. For example, if the first image is an image with low estimation accuracy by the machine learning model, the server can collect similar images from the information processing device that are useful for improving the estimation accuracy of the machine learning model. This allows for efficient collection of data for training the machine learning model. In addition, by using the object detection results from the first sensor, it becomes easier to identify the position of objects in the camera's captured images, thereby reducing the processing load.

[0019] Another aspect of the present disclosure is an information processing device comprising a control unit that performs the following: transmitting local image features of one or more first objects included in a first image to a first terminal; receiving a second image captured by a camera from the first terminal, wherein the local image features of one or more objects included in the image are similar to the local image features of one or more first objects; and storing the second image captured image in a storage unit as one of the training data for a machine learning model used for image recognition. The information processing device is, for example, a server. The server is, for example, a dedicated computer and a PC or the like. The first terminal may acquire the local image features of one or more objects included in the image based on the object detection results by the first sensor.

[0020] In another aspect, this disclosure can be identified as a method by which a computer performs the processing performed by each of the above-mentioned information processing devices. In yet another aspect, this disclosure can also be identified as a program for causing a computer to perform the above method, and a non-temporary, computer-readable recording medium on which the program is recorded.

[0021] Embodiments of this disclosure will be described below with reference to the drawings. The configurations of the following embodiments are illustrative, and this disclosure is not limited to the configurations of these embodiments.

[0022] <First Embodiment> Figure 1 shows an example of the system configuration of the machine learning model training data collection system 100 according to the first embodiment, and an example of the hardware configuration of the center server 1 and vehicle 2. The training data collection system 100 is a system that collects images to be used as training data for a machine learning model that performs image recognition. The training data collection system 100 includes the center server 1 and vehicle 2. The training data collection system 100 includes multiple vehicles 2, but for convenience, only one vehicle 2 is shown in Figure 1.

[0023] The center server 1 and vehicle 2 are connected to network N1 and can communicate with each other through network N1. Network N1 is, for example, a public network such as the Internet. The vehicle information DB 3, weather information DB 4, and map information DB 5 are also connected to network N1.

[0024] Vehicle 2 is a so-called connected car, a vehicle equipped with communication capabilities. Information regarding the operation of Vehicle 2 is collected periodically and stored in Vehicle Information DB 3. The vehicle information stored in Vehicle Information DB 3 includes, for example, the location information, timestamp, and sensor information such as speed of Vehicle 2.

[0025] In the first embodiment, for example, a machine learning model used for image recognition for safe driving of vehicles such as advanced driver-assistance systems (ADAS) is assumed. Therefore, the center server 1 collects images captured by the vehicle's onboard camera as training data for the machine learning model used for image recognition. In the first embodiment, the center server 1 sets similar images from the vehicle's onboard camera images of images with low accuracy in the estimation results by the machine learning model as targets for collection as data for retraining the machine learning model. System 1 notifies vehicle 2 of the local image features of one or more objects contained in images where the accuracy of the machine learning model's estimation results is low. Vehicle 2 identifies images from multiple images captured by the onboard camera in which the local image features of one or more objects contained in the image are similar to those of one or more objects notified by center server 1, and transmits them to center server 1.

[0026] This allows the center server 1 to collect similar images from vehicle 2 to images with low accuracy in the estimation results by the machine learning model, which are suitable as data for retraining the machine learning model. Since images suitable for retraining the machine learning model are received precisely from vehicle 2, the human cost of manually selecting suitable images for retraining from video data of a predetermined length can be reduced. Therefore, the training data collection system 100 can improve the efficiency of collecting retraining data.

[0027] Next, the hardware configuration of each device will be described. Center Server 1 can be configured using a dedicated information processing device (computer), such as a server machine. Center Server 1 may also be a collection of one or more computers (cloud). Center Server 1 is an example of an "information processing device".

[0028] The center server 1 comprises a processor 101, memory 102, auxiliary storage device 103, and communication unit 104 as its hardware configuration. The memory 102 and auxiliary storage device 103 are recording media that can be read by a computer.

[0029] The auxiliary storage device 103 stores various programs and data used by the processor 101 when executing each program. The auxiliary storage device 103 is, for example, an EPROM (Erasable Programmable Read-Only Memory), a hard disk drive, or an SSD (Solid State Drive). The programs held in the auxiliary storage device 103 include, for example, This includes an operating system (OS), application programs, and a training data collection program. The training data collection program is a program that collects images captured by an in-vehicle camera from vehicle 2 that are suitable as training data for a machine learning model.

[0030] Memory 102 is a storage device that provides the processor 101 with a storage area and a work area for loading programs stored in the auxiliary storage device 103, and is also used as a buffer. Memory 102 includes, for example, semiconductor memory such as ROM (Read Only Memory) and RAM (Random Access Memory).

[0031] The processor 101 performs various processes by loading a program held in the auxiliary storage device 103 into the memory 102 and executing it. The processor 101 is, for example, a CPU, GPU, or DSP (Digital Signal Processor). There is not limited to one processor 101, but multiple processors may be provided. The processor 101 is an example of a "control unit".

[0032] The communication unit 104 is, for example, a NIC (Network Interface Card) or an optical line interface. The communication unit 104 is, for example, a wireless communication device that connects to a wireless network such as a wireless LAN and a mobile wireless communication network such as 5G, LTE (Long Term Evolution), or 6G. It may also be a circuit. Note that the hardware configuration of the center server 1 is not limited to that shown in Figure 1. The center server 1 may be a device equipped with an FPGA (Field-Programmable Gate Array) or ASIC (Application Specific Integrated Circuit) or other electrical circuit dedicated to performing the relevant processing.

[0033] Next, vehicle 2 is equipped with a DCM 210, an ECU 212, a camera 230, a millimeter-wave radar 240, and a position acquisition sensor 250. The DCM 210 and the ECU 220 are connected, for example, by an in-vehicle LAN. The ECU 220, camera 230, millimeter-wave radar 240, and position acquisition sensor 250 are connected, for example, by a CAN (Controller Area Network). However, the DCM 210, ECU 220, and camera 23 The method of connecting the 0, millimeter-wave radar 240, and position acquisition sensor 250 is not limited to these. In Figure 1, for convenience, only the hardware components related to the processing according to the first embodiment are extracted and shown, and the vehicle 2 also has other hardware components for controlling driving, etc.

[0034] The position acquisition sensor 250 is, for example, a GPS (Global Positioning System) receiver. The detected value from the position acquisition sensor is the current location information of vehicle 2. The location information is, for example, latitude and longitude. The position acquisition sensor 250 acquires the location information at a predetermined period and transmits it to the ECU. The data is output to the DCM 210 via 220. The location information is sent from the DCM 210 to the center server 1 as part of the vehicle information. Camera 230 may be a camera shared with, for example, a drive recorder.

[0035] The millimeter-wave radar 240 performs object detection processing at predetermined intervals. In object detection processing, the millimeter-wave radar 240 irradiates radio waves in a frequency band with wavelengths in millimeters and measures the reflected waves to obtain the presence of an object in the direction of irradiation, as well as positional information such as the distance and angle of the object. In the first embodiment, it is assumed that the imaging direction of the camera 230 and the irradiation direction of the millimeter-wave radar 240 are facing the same direction, and that the imaging range of the camera 230 is included in the irradiation range of the millimeter-wave radar 240.

[0036] The ECU 220 is, for example, an ECU that processes ADAS systems, or a multimedia ECU. However, the ECU 220 is not limited to these. In the first embodiment, the ECU 220 uses the object detection results from the millimeter-wave radar 240 to process images captured by the camera 230 to identify images having local image features similar to the local image features notified by the center server 1.

[0037] The ECU 220 has a hardware configuration that includes a processor 221, memory 222, auxiliary storage device 223, an interface 224 to the in-vehicle network, and an interface 225 to CAN. The processor 221, memory 222, and auxiliary storage device 223 are the same as those of the processor 101, memory 102, and auxiliary storage device 103, respectively. The auxiliary storage device 223 stores an application program for the client of the learning data acquisition system 100. The application program for the client of the learning data acquisition system 100 is installed by the center server 1, for example, via OTA (Over The Air). The application program for the client of the learning data acquisition system 100 is a program that identifies images from images captured by the camera 230 that have local image features similar to local image features notified by the center server 1. The ECU 220 is an example of an "information processing device". The processor 221 is an example of a "control unit".

[0038] The DCM 210 is responsible for the communication functions of vehicle 2. The DCM 210's hardware configuration includes a processor 211, memory 212, auxiliary storage device 213, wireless communication unit 214, and an interface 215 to the in-vehicle network. The processor 211, memory 212, and auxiliary storage device 213 are the same as those of processor 101, memory 102, and auxiliary storage device 103, respectively. The wireless communication unit 214 is a wireless communication circuit that connects to wireless networks such as wireless LAN and mobile wireless communication networks such as 5G, LTE, and 6G. Note that the hardware configurations of the center server 1 and vehicle 2 shown in Figure 1 are examples. However, this is not limited to the configuration shown in Figure 1.

[0039] Figure 2 shows an example of the data collection process for training according to the first embodiment. First, vehicle 2 periodically transmits vehicle information at predetermined intervals, and this vehicle information is stored in vehicle information DB 3.

[0040] In S11, the center server 1 extracts images with low estimation accuracy from the training data of test images. The training data of images is, for example, image data to which each object contained in the image is tagged with the type of object and its position information in the image. A predetermined evaluation metric is used to determine the estimation accuracy. The estimation accuracy of the machine learning model can be improved by training it with images similar to those with low estimation accuracy. Hereinafter, images with low estimation accuracy will be referred to as question images. The images of the retraining data to be collected are images similar to the question images. Hereinafter, the images of the retraining data to be collected may also be referred to as similar images to the question images. The question image is an example of the "first image".

[0041] In S12, the center server 1 performs scene analysis on the problem image. Scene analysis is performed, for example, by estimating from location information such as the time of shooting and the shooting location, or by using a predetermined image scene recognition technology. Scene analysis may also be performed, for example, by visual judgment by a person. Through scene analysis, information such as weather, time of day, and driving scene is obtained. Driving scene refers to information about the location where vehicle 2 is driving. Information about the driving scene includes, for example, whether the vehicle is driving on a highway, on a main road, in a tunnel, or on a mountain road. The information obtained as a result of the scene analysis is stored in the collection plan DB 17.

[0042] In S13, the center server 1 extracts local image features for one or more objects contained in the task image from the task image corresponding to the training data extracted in S11. The local image features are obtained using a feature extractor with algorithms such as AKAZE, SIFT, HOG, SURF, and LBP. The set of local image features for one or more objects extracted from the task image is stored in the collection plan DB 17.

[0043] The administrator of the training data collection system 100 creates a data collection plan for retraining the machine learning model. The collection plan includes information on when to start and end the collection of data for retraining, from which vehicle 2 data to be collected, how much data to collect for retraining, and what kind of images the collected data for retraining will be. This collection plan is created, for example, by referring to the results of the scene analysis in S12 and using the local image feature set of the task image acquired in S13.

[0044] In S21, the center server 1 refers to the vehicle information DB 3, weather information DB 4, and map information DB 5 to identify the vehicle to be instructed to collect data for retraining according to the collection plan. This makes it possible to identify vehicle 2 that is in a situation similar to the scene analysis results of the problem image as the target, and prevents instructing vehicle 2, which is unlikely to capture similar images to the problem image, to collect data for retraining. In S22, the center server 1 transmits, for example, an instruction to start data collection and the local image features of the problem image to vehicle 2 identified in S21, via OTA.

[0045] In S31, vehicle 2 detects an object using millimeter-wave radar 240. While the millimeter-wave radar 240 detects the presence of an object, it does not detect the type of object. The millimeter-wave radar 240 also acquires the size and position of the detected object. The millimeter-wave radar 240 is an example of a "first sensor."

[0046] In S32, Vehicle 2 converts the object detection result from S31 from the coordinates of the millimeter-wave radar 240 to the coordinates of the camera 230 using a predetermined coordinate transformation method. In S33, based on the coordinate transformation result, Vehicle 2 identifies the area in the captured image where the object is located as the position of the object to be recognized. The captured image used in S33 is the one taken at the time closest to the time when object detection by the millimeter-wave radar 240 occurred in S31.

[0047] In S34, vehicle 2 extracts local image features from the image captured by camera 230, specifically the location of the object identified in S33. The type of features used is the same as that used for the local image features of the problem image. If multiple objects are detected by millimeter-wave radar 240, the location of each object in the captured image is identified, and local image features are extracted for each object.

[0048] In S35, vehicle 2 compares the similarity between the local image features of the task image and the local image features of the images captured by multiple cameras 230. For the similarity comparison between the task image and the images captured by cameras 230, for example, BoVW (Bag-of-Visual Words) is used. These methods are used.

[0049] In S36, vehicle 2 transmits data from multiple captured images to the center server 1, for example, a predetermined number of captured images with the highest similarity, or captured images with a similarity equal to or greater than a predetermined threshold. Along with the captured image data, the position information of objects in the captured images, obtained by coordinate transformation in S32, is also transmitted. The position information of objects is, for example, information indicating the region within the captured image in which the object exists. For example, when the captured image is displayed, a frame indicating the approximate shape of the object is displayed at the position of the object based on the position information of the object, making it easier to see where the object is located. The captured image transmitted from vehicle 2 to the center server 1 is an example of a "second captured image".

[0050] In S41, the center server 1 stores the captured image data received from the vehicle 2 in the retraining image information DB 18 as one of the input data for training data. The center server 1 also stores the positional information of objects contained in the captured image received from the vehicle 2 in the retraining image information DB 18 as auxiliary information for creating training data.

[0051] Subsequently, training data is created using the captured image data received from vehicle 2 and the location information of the objects contained in the captured image. For example, training data is created when a person determines the type of object contained in the captured image, labels it, and assigns the location information of each object. The location information of the objects contained in the captured image can be used to display, for example, a frame indicating the area where the object is located on the captured image displayed on the screen. This makes it easier for the worker to identify the object to be labeled from the captured image, improving work efficiency.

[0052] By retraining the machine learning model using the image data and training data collected in this manner as input data for retraining, the estimation accuracy of the machine learning model can be improved. The machine learning model may be installed on the central server 1 or on a device other than the central server 1. Furthermore, the training of the machine learning model may be performed by the central server 1 or by a device other than the central server 1.

[0053] Figure 3 shows an example of the functional configuration of the center server 1 and the vehicle 2. The center server 1 includes, as a functional configuration, an evaluation unit 11, a control unit 12, an image feature extraction unit 13, a scene analysis unit 14, a test image information DB 16, a collection plan DB 17, and a retraining image information DB 18. These are, for example, generated by the processor 101 of the center server 1 under a predetermined program. The functionality may be achieved by executing the above. Furthermore, each or part of the functional components may correspond to a different program or a different hardware component (such as an FPGA).

[0054] The evaluation unit 11 evaluates the estimation results of the machine learning model for each image of the trained training data according to predetermined evaluation criteria and extracts the problem image. The control unit 12 controls the collection of retraining data to be used as the problem image, which corresponds to, for example, the processes S21, S22, and S41 in Figure 2. Details of the processing of the control unit 12 will be described later. The image feature extraction unit 13 extracts local image features from the image. The scene analysis unit 14 performs scene analysis of the image using a predetermined method.

[0055] The test image information DB 16, the data collection plan DB 17, and the retraining image information DB 18 are created, for example, in the storage area of ​​the auxiliary storage device 103 of the center server 1. The test image information DB 16 is training data that includes test images as input data and training data for those test images, and it holds information about the training data that the machine learning model has already trained on. The data collection plan DB 17 holds the data collection plan, the scene analysis results of the task images, and the local image feature sets. Details of the information included in the data collection plan will be described later. Retraining image information DB Item 18 stores similar images of the problem image collected from vehicle 2 as input data, auxiliary information for creating training data, and the training data itself.

[0056] Next, the vehicle 2 comprises, as a functional configuration, a control unit 21, an object detection unit 22, and an image feature extraction unit 23. These may be functions achieved, for example, by the processor 221 of the ECU 220 executing a predetermined program. Furthermore, each or part of the functional components may correspond to a different program or a different hardware component (such as an FPGA).

[0057] The control unit 21 controls the process of identifying similar images to the problem image from the images captured by the camera 230 and transmitting them to the center server 1, which corresponds to the processes from S31 to S36 in Figure 2. Details of the processing of the control unit 21 will be described later. The object detection unit 22 controls object detection by the millimeter-wave radar 240. The image feature extraction unit 23 extracts local image features from the image. The local image features acquired by the image feature extraction unit 23 are of the same type as the local image features acquired by the image feature extraction unit 13 of the center server 1.

[0058] Vehicle Information DB 3 holds vehicle information collected periodically from vehicle 2 at predetermined intervals. This vehicle information includes, for example, the location and speed of vehicle 2. Weather Information DB 4 holds weather information. This weather information includes information on weather, sunshine, precipitation, wind speed, and the occurrence of lightning for each area. Map Information DB 5 stores map information. Vehicle Information DB 3, Weather Information DB 4, and Map Information DB 5 may be databases managed by the learning data collection system 100, or they may be databases managed by external organizations. The functional configuration of the center server 1 and vehicle 2 is not limited to the example shown in Figure 3.

[0059] Figure 4 shows an example of the information included in the data collection plan. The data collection plan includes fields for plan ID, issue image ID, start conditions, target vehicle conditions, end conditions, and local image features. The plan ID field stores identification information for the data collection plan. The issue image ID field stores identification information for the issue image.

[0060] The "Start Conditions" field stores information defining the start conditions for the data collection process for retraining for the given image. The "Start Conditions" field includes a "Time Zone" field. The "Time Zone" field specifies the conditions for the data collection process for retraining for the given image. Information indicating the time period to be performed is stored. The information indicating the time period may be determined, for example, based on the time period in which the task image was captured, obtained by scene analysis of the task image, or it may be a time period arbitrarily specified by the administrator of the training data collection system 100. The start condition field may be empty, in which case the data collection process for retraining data for the task image will start upon input of a start instruction from the administrator of the training data collection system 100.

[0061] The Target Vehicle Conditions field stores information defining the conditions of Vehicle 2, which is the vehicle to be instructed to collect retraining data corresponding to the problem image. The Target Vehicle Conditions field includes Area, Weather, and Driving Scene fields. The Area field stores information indicating the geographical range from which retraining data corresponding to the problem image will be collected. The Weather field stores information indicating the weather conditions for which retraining data corresponding to the problem image will be collected. The Driving Scene field stores information indicating the driving scene for which retraining data corresponding to the problem image will be collected. The information stored in the Area, Weather, and Driving Scene fields may be determined, for example, based on the scene analysis results of the problem image, or may be arbitrarily specified by the administrator of the training data collection system 100. Note that the Area, Weather, and Driving Scene fields may be left blank. If no Target Vehicle Conditions are set, for example, Vehicle 2 to be instructed to collect retraining data corresponding to the problem image may be randomly identified.

[0062] The termination condition field stores information defining the termination conditions for the data collection process for retraining for the given problem image. The termination condition field includes fields for the number of images collected and the number of vehicles collected. The number of images collected field stores the lower limit of the number of images to be collected as data for retraining for the given problem image. When the number of images collected from vehicle 2 reaches the value in the number of images collected field, the data collection process for retraining for the given problem image is deemed complete. The number of vehicles collected field stores the number of vehicles 2 to be collected as data for retraining for the given problem image. When images have been collected from the number of vehicles 2 equal to the value in the number of vehicles collected field, the data collection process for retraining for the given problem image is deemed complete. The termination condition may be defined by either the number of images collected or the number of vehicles collected, or by both. If the termination condition is defined by both the number of images collected and the number of vehicles collected, the data collection process for retraining for the given problem image is deemed complete when both conditions are met. The termination condition field may be left empty; in that case, for example, a predetermined number of collected items may be used as the termination condition.

[0063] The local image features field stores a set of local image features extracted from the problem image. The set of local image features stored in the local image features field is then notified to vehicle 2.

[0064] Note that the data collection plan shown in Figure 4 is just one example, and each condition included in the collection plan can be flexibly set depending on what kind of images the training data collection system 100 wants to collect as training data. For example, the number of images to be collected from one vehicle 2 may be defined as the termination condition.

[0065] Furthermore, the image data collected from vehicle 2 may be further refined using, for example, vehicle information to obtain data for retraining. Vehicle information includes, for example, information such as the amount the brake pedal is pressed and the amount of the lower ring on the steering wheel, making it possible to detect when sudden braking or sudden steering has occurred. For example, if you want to collect images of sudden braking and sudden steering as training data, you may attach vehicle information to the images captured from vehicle 2 and obtain the images from vehicle 2 that have attached vehicle information indicating sudden braking or sudden steering as retraining data. Alternatively, along with the instruction to start data collection, sudden braking Instructions may be sent to transmit images of raking and sudden steering, causing vehicle 2 to transmit similar images to the problem images of sudden braking and sudden steering.

[0066] Figure 5 is an example of a flowchart for the data collection process for retraining on the center server 1. The process shown in Figure 5 is executed for each collection plan stored in the collection plan DB 17. The process shown in Figure 5 is executed repeatedly at predetermined intervals. The entity executing the process shown in Figure 5 is, for example, the processor 101 of the center server 1, but for convenience, the functional components will be described as the main components.

[0067] In OP101, the control unit 12 determines whether the start conditions included in the collection plan are met. If the start conditions are met (OP101:YES), the process proceeds to OP102. If the start conditions are not met (OP101:NO), the process shown in Figure 5 ends.

[0068] In OP102, the control unit 12 extracts multiple vehicles 2 that meet the target vehicle conditions included in the collection plan. In OP103, the control unit 12 sends a data collection start instruction and a set of local image features of the task image to each vehicle 2 extracted in OP102. In OP104, the control unit 12 receives image data from the vehicles 2 as retraining data and performs a reception process to store it in the retraining image information DB 18.

[0069] In OP105, the control unit 12 determines whether the termination conditions included in the collection plan have been met. If the termination conditions are met (OP105: YES), the process proceeds to OP106. If the termination conditions are not met (OP105: NO), the process proceeds to OP104.

[0070] OP106 sends a data collection completion instruction to the target vehicle. Subsequently, the process shown in Figure 5 is completed.

[0071] Figure 6 is an example of a flowchart for the data collection process for vehicle 2's relearning. The process shown in Figure 6 is executed repeatedly at predetermined intervals. The entity executing the process shown in Figure 6 is, for example, the processor 221 of the ECU 220, but for convenience, the functional components will be described as the main components.

[0072] In OP201, the control unit 21 determines whether or not it has received a data acquisition start instruction from the center server 1. If a data acquisition start instruction is received from the center server 1 (OP201:YES), the process proceeds to OP202. Along with the data acquisition start instruction, the local image feature set of the task image is also received. If a data acquisition start instruction is not received from the center server 1 (OP201:NO), the process shown in Figure 6 ends.

[0073] In OP202, the control unit 21 determines whether a predetermined time has elapsed. The predetermined time can be arbitrarily set within a range of, for example, 1 second to 10 seconds. If the predetermined time has elapsed (OP202: YES), the process proceeds to OP203. The control unit 21 remains in a standby state until the predetermined time has elapsed.

[0074] In OP203, the control unit 21 acquires multiple object detection results from the millimeter-wave radar 240 performed over the most recent predetermined time period from the object detection unit 22. In OP204, the control unit 21 converts each object detection result from the millimeter-wave radar 240 acquired in OP203 into camera coordinates. In OP205, the control unit 21 identifies the position of an object in each of the multiple images captured by the camera 230 over the most recent predetermined time period, based on the corresponding object detection result.

[0075] In OP206, the control unit 21 requests the image feature extraction unit 23 for each of the multiple images captured by the camera 230 over the most recent predetermined time period, and obtains local image features extracted from the location of the object identified in OP205. In OP207, the control unit 21 compares the similarity between the multiple images captured by the camera 230 over the most recent predetermined time period and the group of local image features of the problem image, and determines the image captured that is similar to the problem image. There may be one image or multiple images similar to the problem image. In OP208, the control unit 21 transmits the image captured that is similar to the problem image, along with the location information of the object in the image obtained in OP204, to the center server 1.

[0076] In OP209, the control unit 21 determines whether or not it has received a data collection termination instruction from the center server 1. If a data collection termination instruction is received from the center server 1 (OP209: YES), the process shown in Figure 6 is terminated. If a data collection termination instruction is not received from the center server 1 (OP209: NO), the process proceeds to OP202. As a result, the process from OP203 is performed after a predetermined time has elapsed. By making the execution frequency of the processes from OP203 to OP209 equal to a predetermined time interval, it is possible to suppress the impact on other processes of the ECU 220 due to the load from the data collection process for retraining.

[0077] Note that the data acquisition process for the center server 1 shown in Figure 5, and the data acquisition process for the vehicle 2 shown in Figure 6, are examples and can be modified as appropriate depending on the implementation.

[0078] <Effects of the First Embodiment> In the first embodiment, the center server 1 can pinpoint and collect images captured by an in-vehicle camera that are similar to the problem image from the vehicle 2, as training data for a machine learning model. This allows for the collection of images suitable for training a machine learning model and with consistent selection criteria, without human intervention, thereby improving the efficiency of the data collection process for training the machine learning model. Furthermore, training using training data with consistent selection criteria can further improve the estimation accuracy of the machine learning model.

[0079] Furthermore, in the first embodiment, the vehicle 2 uses the object detection results from the millimeter-wave radar 240 to identify the position of an object in the image captured by the corresponding camera 230. This makes it easier to identify the target region from which to extract local image features in the image captured by the camera 230, thereby reducing the load on the vehicle 2 for collecting data for retraining and shortening the processing time. In addition, by transmitting the position information of objects in the captured image along with the captured image determined to be a similar image to the problem image, it becomes easier for a person to visually identify the position of objects in the captured image, thereby improving the efficiency of the human work involved in creating training data.

[0080] Furthermore, the data collection process for retraining according to the first embodiment can be executed by the vehicle 2 simply by installing the client program for the training data collection system 100 via OTA (Over-the-Air). Therefore, in the first embodiment, data collection for retraining can be achieved without changing the hardware configuration of the vehicle 2.

[0081] Furthermore, in the first embodiment, a data collection plan is created based on the scene analysis results of the problem image, and the center server 1, based on the data collection plan, instructs vehicles under similar conditions to those in which the problem image was captured to begin collecting data for retraining at a timing that takes into account the circumstances under which the problem image was captured. This improves the probability that the images collected from vehicle 2 are more suitable as training data. In addition, even if the processing power of vehicle 2 is not high, it is possible to collect images that are more suitable as training data.

[0082] <Other Embodiments> The embodiments described above are merely examples, and this disclosure may be modified as appropriate without departing from its essence.

[0083] In the first embodiment, the center server 1 collects captured images from the vehicle 2 as training data for a machine learning model. However, the terminals from which training data is collected are not limited to the vehicle 2 and can be appropriately changed depending on the purpose of use of the machine learning model. For example, the technology described in the first embodiment may be applied to a system that collects captured images from cameras mounted on motorcycles, bicycles, and railway vehicles as training data. Furthermore, it is not limited to vehicles moving on land, but may be applied to a system that collects captured images from cameras mounted on aircraft, ships, etc., as training data. Alternatively, it may be applied to a system that collects captured images from surveillance cameras, fixed-point cameras, etc., as training data.

[0084] In the first embodiment, the vehicle 2 uses the object detection results of the millimeter-wave radar 240 to determine the position of objects in the captured image, but instead of the millimeter-wave radar 240, other sensors used in ADAS systems, such as sonar or LiDAR, may be used.

[0085] In the first embodiment, the ECU 220 performs the data collection process for relearning in vehicle 2, but the hardware component that performs the data collection process for relearning in vehicle 2 is not limited to the ECU 220. For example, the DCM 210, drive recorder system, and car navigation system may perform the data collection process for relearning in vehicle 2.

[0086] In the first embodiment, the center server 1 instructs the vehicle 2 to start and end the data collection process for retraining, but is not limited to this. For example, along with the instruction to start data collection, start and end conditions for each vehicle 2 may be transmitted, and the vehicle 2 may determine when to start and end the data collection process for retraining. The start conditions may be defined, for example, by the time period, area, weather, and driving scene included in the collection plan. The end conditions may be defined, for example, by the number of similar images collected for each vehicle 2.

[0087] The processes and methods described in this disclosure can be freely combined and implemented, provided that no technical inconsistencies arise.

[0088] Furthermore, a process described as being performed by a single device may be divided and executed by multiple devices. Conversely, a process described as being performed by different devices may be executed by a single device. In a computer system, the hardware configuration (server configuration) by which each function is implemented can be flexibly changed.

[0089] The present disclosure can also be realized by supplying a computer program implementing the functions described in the embodiments above to a computer, and having one or more processors in the computer read and execute the program. Such a computer program may be provided to the computer by a non-temporary computer-readable storage medium that can be connected to the computer's system bus, or it may be provided to the computer via a network. Non-temporary computer-readable storage mediums include, for example, any type of disk such as magnetic disks (floppy disks, hard disk drives (HDDs), etc.), optical disks (CD-ROMs, DVDs, Blu-ray discs, etc.), read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic cards, flash memory, optical cards, and any type of medium suitable for storing electronic instructions. [Explanation of symbols]

[0090] 1. Center Server 2. Vehicles 11. Evaluation Department 12, 21... Control Unit 13, 23...Image Feature Extraction Unit 14. Scene Analysis Department 16. Test Image Information Database 17. Collection Plan Database 18. Image information database for retraining 22. Object detection unit 100 · Learning Data Collection System 101, 221... Processors 102,222...memory 103, 223...Auxiliary storage device 104 Communications Department 210··DCM 220··ECU 230 Camera 240 mm wave radar 250 · Position acquisition sensor

Claims

1. The server receives local image features of one or more first objects contained in the first image, The first sensor detects an object by emitting a predetermined signal within the range including the camera's imaging range, and the first sensor detects an object. Based on the object detection result by the first sensor, local image features of one or more objects included in each of the multiple images captured by the camera are obtained. Of the plurality of captured images, a second captured image in which the local image features of one or more objects included in the captured image are similar to the local image features of the one or more first objects is transmitted to the server as one of the training data for a machine learning model used for image recognition. A control unit that executes An information processing device equipped with the following features.

2. The control unit, Between the plurality of captured images, one or more positions of one or more objects included in the captured images are identified based on the detection results of the first sensor, local image features of one or more objects included in the captured images are obtained from the identified one or more positions in the captured images, and the similarity between the obtained local image features of one or more objects included in the captured images and the local image features of the one or more first objects is compared. Based on the results of the comparison, the image among the plurality of captured images in which the local image features of one or more objects included in the captured image are similar to the local image features of the one or more first objects is determined to be the second captured image. Further execution, The server is transmitted, along with the second captured image, the position information in the second captured image of one or more objects included in the second captured image, which has been identified based on the detection result of the first sensor. The information processing apparatus according to claim 1.

3. Transmitting local image features of one or more first objects contained in the first image to the first terminal, The first terminal receives a second image captured by a camera, wherein the local image features of one or more objects included in the image are similar to the local image features of the one or more first objects. The second captured image is stored in the memory unit as one of the training data for a machine learning model used for image recognition, A control unit that executes Equipped with, The local image features of one or more objects included in the aforementioned image are acquired based on the object detection result by a first sensor that detects objects by transmitting a predetermined signal over a range including the imaging range of the camera. Information processing device.

4. The control unit, The first terminal is further identified from multiple terminals based on at least one of the terminal's location, weather, and time of day. Along with the second image captured from the first terminal, the object as detected by the first sensor The system receives the position in the second image of one or more objects included in the second image, which has been identified based on the detection results. The training data includes training data corresponding to the second image, created based on the positions of one or more objects in the second image in the second image. The information processing apparatus according to claim 3.

5. The device, The server receives local image features of one or more first objects contained in the first image, The first sensor detects an object by emitting a predetermined signal within the range including the camera's imaging range, and the first sensor detects an object. Based on the object detection result by the first sensor, local image features of one or more objects included in each of the multiple images captured by the camera are obtained. Of the plurality of captured images, a second captured image in which the local image features of one or more objects included in the captured image are similar to the local image features of the one or more first objects is transmitted to the server as one of the training data for a machine learning model used for image recognition. Execute, The aforementioned server, Transmitting local image feature quantities of the one or more first objects to the terminal, The terminal receives the second captured image, The second captured image is stored in the memory unit as one of the training data for the machine learning model, Execute method.