Information processing apparatus and information processing method

The device and method enhance the accuracy of road environment maps by assessing operator corrections to determine re-training needs, focusing on areas with high correction costs to improve machine learning model precision.

JP2026017893APending Publication Date: 2026-02-05TOYOTA JIDOSHA KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024118946
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing machine learning models for generating road environment maps in autonomous driving systems lack a reliable method to determine when re-training is necessary to improve accuracy, leading to potential errors in road object positioning and detection.

Method used

An information processing device and method that acquires operator corrections to map data, calculates the cost of these corrections, and determines whether re-learning of the machine learning model is necessary based on the correction cost, allowing targeted re-training on areas with low accuracy.

Benefits of technology

Improves the accuracy of road environment recognition by selectively re-training the machine learning model on areas with high correction costs, reducing overall estimation errors and maintaining map precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017893000001_ABST
    Figure 2026017893000001_ABST
Patent Text Reader

Abstract

To improve accuracy of a machine learning model for recognizing a road environment.SOLUTION: Position information of one or more road objects is acquired by inputting sensor data obtained by sensing a road environment to a machine learning model, a result of an operator correcting first map data in which the position information of the road objects is reflected is acquired, and it is determined whether or not relearning of the machine learning model is necessary based on a cost of a correction operation performed by the operator.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a technology for generating a map showing a road environment. [Background technology]

[0002] Systems for recognizing real-world environments using machine learning models are known. For example, Patent Literature 1 discloses a system that can evaluate the accuracy of a machine learning model by determining a similarity score for objects whose positions and features are known. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2023-131069 [Patent Document 2] Japanese Patent Publication No. 2020-086786 Summary of the Invention [Problem to be solved by the invention]

[0004] The present disclosure aims to improve the accuracy of machine learning models that recognize road environments. [Means for solving the problem]

[0005] One aspect of the present disclosure is The information processing device has a control unit that performs the following operations: acquiring position information of one or more road objects by inputting sensor data obtained by sensing the road environment into a machine learning model; acquiring the results of corrections made by an operator to first map data that reflects the position information of the road objects; and determining whether or not re-learning of the machine learning model is necessary based on the cost of the correction operation made by the operator.

[0006] One aspect of the present disclosure is This is an information processing method in which a computer executes the following steps: obtain position information of one or more road objects by inputting sensor data obtained by sensing the road environment into a machine learning model; obtain the results of corrections made by an operator to first map data that reflects the position information of the road objects; and determine whether or not re-learning of the machine learning model is necessary based on the cost of the correction operation made by the operator.

[0007] Another aspect is a program for causing a computer to execute the above-described information processing method, or a computer-readable storage medium that non-temporarily stores the program. [Effects of the Invention]

[0008] According to the present disclosure, it is possible to improve the accuracy of machine learning models that recognize road environments. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram for explaining an overview of a system according to an embodiment. [Figure 2] 2 is a diagram illustrating the configuration of a server device 1 and an in-vehicle device 2. [Figure 3] FIG. 2 is a diagram illustrating a learning data set stored in the storage unit 12. [Figure 4] FIG. 3 is a diagram for explaining the flow of processing in the control unit 11. [Figure 5] 4 is a flowchart of a process executed by the server device 1. [Figure 6] 4 is a flowchart of a process executed by the server device 1. DETAILED DESCRIPTION OF THE INVENTION

[0010] In recent years, there has been much research into autonomous driving systems in which a vehicle autonomously drives along a pre-set route. In autonomous driving, a vehicle determines its own position and attitude by comparing a pre-stored road map with the results of sensing the road environment.

[0011] In such systems, road maps must be kept up to date. To address this, attempts have been made to automatically generate road maps based on sensor data collected by the vehicle. For example, a machine learning model can generate a map for autonomous driving by identifying the locations of objects such as lane lines, road boundaries, traffic lights, and crosswalks based on camera images acquired by the vehicle.

[0012] The quality of road maps generated by machine learning models can vary depending on the model's learning history. However, it is not easy to objectively determine whether a machine learning model has been trained sufficiently. The information processing device according to this embodiment solves such a problem.

[0013] An information processing device according to one embodiment has a control unit that performs the following operations: acquiring location information of one or more road objects by inputting sensor data obtained by sensing the road environment into a machine learning model; acquiring the results of corrections made by an operator to first map data that reflects the location information of the road objects; and determining whether or not re-learning of the machine learning model is necessary based on the cost of the correction operation made by the operator.

[0014] The machine learning model is a model that takes data obtained by sensing the road environment as input and outputs position information of road objects. Road objects are typically objects related to vehicle driving control, such as road boundaries, lane boundaries, traffic lights, or crosswalks. Note that road objects do not necessarily have to be objects with three-dimensional shapes and may be road markings or the like. The sensor data input to the machine learning model may be, for example, data collected by multiple probe cars, or may be image data acquired by an image sensor mounted on the probe cars. The first map data is road map data that includes position information of road objects output by the machine learning model.

[0015] Depending on the accuracy of the machine learning model, the first map data generated may contain errors, such as incorrect positions of road boundaries or lane boundaries, or recording of non-existent objects. The control unit acquires the results of the corrections made by the operator to the first map data, and determines the cost of the corrections.

[0016] The cost can be calculated based on, for example, the amount of operation required for the correction or the number of operations. For example, when an operator corrects the first map data using predetermined software, the cost may be determined based on the amount of operation required for the software (for example, the amount of interaction with the user interface). The control unit then determines whether or not re-learning of the machine learning model is necessary based on the cost of the correction operation performed by the operator.

[0017] For example, the more the machine learning model outputs results that deviate from the actual road environment, the greater the cost required to correct the first map data. Based on the cost of the correction operation performed on the data, it is possible to evaluate whether the machine learning model has been properly trained (in other words, whether retraining is necessary).

[0018] In addition, the control unit may determine the cost of the correction operation performed by the operator for each of a plurality of unit areas included in the first map data, and if there is a first area, which is a unit area whose cost exceeds a predetermined threshold, determine that re-learning of the machine learning model is necessary.

[0019] For example, the first map may be divided into a plurality of unit areas (tiles), and the correction cost may be determined for each tile. For example, if there is a unit area whose cost exceeds a predetermined threshold, it may be determined that re-learning is necessary to improve the estimation accuracy for that unit area.

[0020] The control unit may also determine a real-world road environment corresponding to the first area, and re-train the machine learning model using a training dataset including the determined road environment.

[0021] If there is a unit area whose correction cost exceeds a predetermined threshold, it is estimated that the estimation accuracy of the machine learning model for the road environment included in that unit area is poor. Therefore, the control unit may, for example, classify the road environment included in the target first area and re-train the machine learning model using a training dataset corresponding to the obtained class. For example, if the class of the road environment corresponding to the first area is "road with multiple lanes," it is estimated that the estimation accuracy of the machine learning model for roads with multiple lanes is poor. Therefore, the machine learning model is re-trained using a training dataset corresponding to a similar road environment (e.g., an image of a road with multiple lanes) as training data. This can improve the estimation accuracy of the machine learning model.

[0022] The control unit may also determine the number of pieces of learning data to be used in the re-learning based on the cost of the correction operation performed by the operator. The control unit may determine, for example, that the more costly the correction operation, the more training data is needed.

[0023] The control unit may add the sensor data corresponding to the first region as input data and the corrected position information of the road object as training data to the learning data set. The result of the correction by the operator can also be used as training data.

[0024] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. The configurations of the following embodiments are examples, and the present disclosure is not limited to the configurations of the embodiments.

[0025] (First embodiment) An overview of a system according to a first embodiment will be described. The system according to this embodiment includes a vehicle (probe car) equipped with an on-board device 2, and a server device 1 that generates a road map based on sensor data collected from the vehicle.

[0026] An overview of the processing performed by each device will be described with reference to FIG. First, the vehicle (on-board device 2) collects data via sensors mounted on the vehicle while traveling and transmits the data as probe data to the server device 1. The probe data may be, for example, an image of the outside of the vehicle captured by an on-board camera. The image may be captured by a single camera, or may be a bird's-eye view image generated (synthesized) based on images captured by multiple cameras.

[0027] The server device 1 has a machine learning model and generates a road map using probe data received from multiple vehicles. In this embodiment, the machine learning model is a model that detects predetermined traffic-related objects from the probe data and outputs their attributes and location information. The predetermined objects are typically objects related to vehicle driving control (hereinafter referred to as road objects), such as road boundaries, lane boundaries, traffic lights, or crosswalks. The server device 1 inputs the probe data into the machine learning model to acquire the attributes and location information of road objects and maps them onto a road map. The data obtained by mapping road objects onto a road map is called road map data. The road map data is used, for example, for driving control of an autonomous vehicle.

[0028] The road map data generated by the server device 1 may contain errors. Therefore, in this embodiment, the system operator is allowed to perform operations to correct the road map data. This allows operations such as correcting the placement positions of road objects, deleting erroneously detected road objects, and adding undetected road objects.

[0029] Furthermore, the server device 1 determines whether or not re-learning of the machine learning model is necessary based on the cost of the correction operation performed by the operator, and performs partial re-learning of the machine learning model as necessary. This makes it possible to perform re-learning only on parts with low accuracy, thereby improving the overall estimation accuracy at low cost.

[0030] [Device configuration] FIG. 2 is a diagram showing an example of the configuration of the server device 1. As shown in FIG. The server device 1 is, for example, a computer such as a server device, a personal computer, a smartphone, a mobile phone, a tablet computer, a personal digital assistant, etc. The server device 1 includes a control unit 11, a storage unit 12, a communication unit 13, and an input / output unit 14.

[0031] The server device 1 can be configured as a computer having a processor (CPU, GPU, etc.), a main memory device (RAM, ROM, etc.), and an auxiliary memory device (EPROM, hard disk drive, removable media, etc.). The auxiliary memory device stores an operating system (OS), various programs, various tables, etc., and by executing the programs stored therein, various functions (software modules) that match predetermined purposes, as described below, can be realized. However, some or all of the functions may be realized as hardware modules using hardware circuits such as ASICs, FPGAs, etc.

[0032] The control unit 11 is a computing unit that executes predetermined programs to realize various functions of the server device 1. The control unit 11 can be realized by, for example, a hardware processor such as a CPU. The control unit 11 may also be configured to include RAM, ROM (Read Only Memory), cache memory, etc.

[0033] The control unit 11 is configured to have three software modules: a generation unit 111, a correction unit 112, and a relearning unit 113. Each software module may be realized by the control unit 11 (CPU) executing a program stored in the storage unit 12, which will be described later.

[0034] The generation unit 111 receives probe data from the vehicle (on-board device 2) and generates the received probe data. Road map data is generated based on the data. The generation unit 111 converts the probe data into feature quantities and inputs the feature quantities to a machine learning model (referred to as an estimation model) described later. Furthermore, the generation unit 111 obtains labels indicating the attributes of road objects and location information from the estimation model, and maps these onto a road map. The mapping result is stored in the storage unit 12 as road map data.

[0035] The correction unit 112 provides the road map data stored in the memory unit 12 to the operator of the device, allowing the operator to correct the contents. As described above, the road map data is generated using the results of estimation by the machine learning model, and therefore may contain errors. Examples of errors include incorrect recognition of road objects, failure to detect road objects themselves, and incorrect positions of road objects. Therefore, the correction unit 112 outputs the road map data and corresponding probe data (e.g., images acquired by an on-board camera) along with a user interface for correction, allowing the operator to correct the errors. The operator compares the two, corrects any erroneous parts, and saves them. The results of the correction are reflected in the road map data, and the contents of the correction are saved as a correction log, which is used for re-learning the estimation model, as described below.

[0036] The correction unit 112 divides the road map data into a plurality of unit areas (hereinafter referred to as tiles) and accepts a correction operation for each tile. In this embodiment, the road map data is divided into a plurality of tiles in a mesh shape, and when a correction operation is performed, the target tile can be identified. The correction unit 112 saves information for identifying the tile on which the correction operation was performed and information about the correction content as a correction log.

[0037] The re-learning unit 113 determines whether or not re-learning of the estimation model is necessary based on the correction log stored by the correction unit 112, and performs re-learning of the estimation model based on the result of the determination. When road map data is modified, the re-learning unit 113 first identifies the tile on which the modification operation was performed based on the modification log. Then, the re-learning unit 113 executes a process for calculating the cost of the modification (hereinafter, "modification cost") for each tile. In this embodiment, the modification cost is a dimensionless number that increases as the amount of modification increases.

[0038] The correction cost can be calculated based on, for example, the amount of operation required for the correction or the number of correction operations. For example, the greater the number of objects to be corrected, the greater the calculated correction cost. Also, the greater the amount of correction (for example, the distance traveled by an object), the greater the calculated correction cost. Note that if the correction unit 112 is capable of correcting road map data using curation software or the like, the re-learning unit 113 may acquire information regarding the amount of operation required for the correction and the number of operations from the software.

[0039] If there is a tile ("first region" in this disclosure) whose correction cost exceeds a predetermined threshold, the re-learning unit 113 determines that re-learning of the estimation model is necessary to improve the estimation accuracy for that tile. A specific method of the re-learning process will be described later.

[0040] The storage unit 12 is a means for storing information, and is configured with storage media such as RAM, a magnetic disk, a flash memory, etc. The storage unit 12 stores programs executed by the control unit 11, data used by the programs, etc.

[0041] The storage unit 12 also stores the estimation model, road map data, correction log, and learning data set described above.

[0042] The estimated model is used to annotate the camera images included in the probe data. This is a machine learning model for this purpose. The estimation model is trained to output labels indicating the attributes of road objects and location information when feature quantities corresponding to camera images are input. The estimation model is a model that has been trained in advance using training images and training data.

[0043] The road map data is data obtained by mapping the labels and position information of road objects output by the estimation model onto a predetermined coordinate system. The road map data is updated as needed by the generation unit 111 based on probe data received from the vehicle. The road map data can also be modified by the operator.

[0044] The correction log is a database that stores the history of corrections made to road map data. The correction log records the identifiers of tiles corrected by the operator, the details of the correction operation, and so on.

[0045] A training dataset is a collection of data (training data) used for retraining an estimation model. FIG. 3 shows the structure of a training dataset. The training data consists of a set of input images for the estimation model (e.g., images captured by an in-vehicle camera) and training data (e.g., attributes and location information of one or more traffic objects contained in the images). Furthermore, the training data is assigned a label indicating the road environment. The label is a label (tag) that describes the road environment contained in the input image and is selected from multiple pre-defined classes, such as "four or more lanes," "intersection," "motorway only," and "traffic lights present." While the attributes contained in the training data are attributes of each traffic object, the road environment label represents the attributes of the road environment itself. One or more road environment labels are assigned to each piece of training data. The training data set is used when the retraining unit 113 retrains the estimation model.

[0046] Returning to Figure 2, we continue the explanation. The communication unit 13 is a communication interface for connecting the server device 1 to a network. The communication unit 13 is configured to be able to communicate with the network via, for example, Ethernet (registered trademark), a wireless LAN, a cellular communication network, or the like.

[0047] The input / output unit 14 is a means for receiving input operations performed by an operator and presenting information to the operator. Specifically, the input / output unit 14 includes devices for input such as a mouse and a keyboard, and devices for output such as a display and a speaker. The input / output devices may be integrally configured with, for example, a touch panel display.

[0048] Next, the configuration of the vehicle-mounted device 2 mounted on the vehicle will be described.

[0049] The in-vehicle device 2 can be configured as a computer having a processor (CPU, GPU, etc.), a main memory device (RAM, ROM, etc.), and an auxiliary memory device (EPROM, hard disk drive, removable media, etc.). The auxiliary memory device stores an operating system (OS), various programs, various tables, etc., and by executing the programs stored therein, various functions (software modules) that match predetermined purposes, as described below, can be realized. However, some or all of the functions may be realized as hardware modules using hardware circuits such as ASICs, FPGAs, etc.

[0050] The in-vehicle device 2 includes a control unit 21, a storage unit 22, a communication unit 23, and a group of sensors 24.

[0051] Here, the sensor group 24 will be explained first. The sensor group 24 is a collection of multiple sensors mounted on the vehicle. In this embodiment, the sensor group 24 is configured to include an on-board camera and a GPS unit. In this embodiment, these are collectively referred to as "sensors."

[0052] An on-board camera is an optical unit that includes an image sensor for capturing images, and is mounted, for example, facing forward of the vehicle. The GPS unit is a unit for acquiring vehicle position information based on signals received from positioning satellites (also called GNSS satellites). The GPS unit includes, for example, a GPS antenna and a positioning module for determining the position information. The GPS antenna is an antenna that receives positioning signals transmitted from the positioning satellites. The positioning module is a module that calculates the position information based on the signals received by the GPS antenna.

[0053] The control unit 21 is a computing unit that executes predetermined programs to realize various functions of the in-vehicle device 2. The control unit 21 can be realized by, for example, a hardware processor such as a CPU. The control unit 21 may also be configured to include RAM, ROM (Read Only Memory), cache memory, etc.

[0054] In this embodiment, the control unit 21 of the in-vehicle device 2 is configured to have a data transmission unit 211 as a software module. The software module may be realized by the control unit 21 (CPU, etc.) executing a program stored in the storage unit 22. Note that the information processing executed by the software module is synonymous with the information processing executed by the control unit 21 (CPU, etc.).

[0055] The data transmission unit 211 acquires data from a plurality of sensors included in the sensor group 24 while the vehicle is traveling, generates probe data based on the data, and transmits the probe data to the server device 1. In this embodiment, the data transmission unit 211 acquires images (hereinafter referred to as camera images) from an on-board camera included in the sensor group 24, and simultaneously acquires position information from a GPS module. The data transmission unit 211 also determines the traveling direction of the vehicle based on changes in the position information over time. The data transmission unit 211 then generates probe data including the camera images, the position information of the vehicle, and data indicating the traveling direction of the vehicle, and periodically transmits the probe data to the server device 1.

[0056] The storage unit 22 is a means for storing information, and is configured with storage media such as RAM, a magnetic disk, a flash memory, etc. The storage unit 22 stores programs executed by the control unit 21, data used by the programs, etc.

[0057] The communication unit 23 is a wireless communication interface for connecting the in-vehicle device 2 to an external network. The communication unit 23 is configured to be able to communicate with the server device 1 via, for example, a wireless LAN or a cellular communication network such as 3G, 4G, or 5G.

[0058] The specific hardware configurations of the server device 1 and the in-vehicle device 2 may include omissions, substitutions, and additions of components as appropriate depending on the embodiment. For example, the control units 11 and 21 may include multiple hardware processors. The hardware processors may be configured with a microprocessor, FPGA, GPU, etc. Furthermore, input / output devices other than those illustrated (for example, an optical drive, etc.) may be added. Furthermore, the server device 1 and the in-vehicle device 2 may be configured with multiple computers. In this case, the hardware configurations of the computers may or may not be the same.

[0059] Next, a detailed flow of processing performed by the control unit 11 of the server device 1 will be described with reference to FIG.

[0060] First, the generation unit 111 receives probe data periodically transmitted from multiple vehicles under its management. The generation unit 111 converts camera images included in the probe data into feature quantities and inputs the feature quantities into an estimation model. As a result, labels indicating attributes of road objects included in the images and location information are obtained from the estimation model. The location information may be expressed in a coordinate system with the camera as the origin (camera coordinate system). In this case, the generation unit 111 may convert the coordinate system of the road objects into a geographic coordinate system using the vehicle's location information and direction.

[0061] The generation unit 111 reflects this information in the road map data. In this embodiment, the road map data is divided into multiple tiles. The generation unit 111 may determine a tile to be updated based on the geographic coordinates of the road object, and update the data for that tile. If the road map data does not contain data corresponding to the tile to be updated, the generation unit 111 may newly generate the target tile. If the road map data already contains data corresponding to the tile to be updated, the generation unit 111 may update the target tile based on the new and old data. If an updated tile is found, the generation unit 111 may notify the operator of the device.

[0062] The correction unit 112 generates a user interface for correcting road map data and provides it to the operator of the device. For example, the correction unit 112 may output a road map corresponding to the most recently updated tile on the screen, or may accept a tile designation from the operator and output a road map corresponding to the designated tile on the screen. The user interface output by the correction unit 112 may include an interface for adding, deleting, or correcting the position of a road object. When the operator completes the correction operation, the correction unit 112 reflects the content of the correction in the road map data and saves the content of the correction in a correction log. The correction log includes information for identifying the tile on which the correction operation was performed and information about the correction content. The information about the correction content includes, for example, an identifier of the road object to be corrected and information indicating what correction was made to the road object.

[0063] When the operator completes the correction operation, the re-learning unit 113 makes a determination regarding re-learning of the estimation model. The relearning unit 113 identifies tiles that affect the most recent modification by referring to the modification log. Furthermore, it calculates the cost of the modification (modification cost) for each identified tile. The modification cost can be calculated based on the following criteria:

[0064] (1) Calculation criteria based on the number of modified road objects This method increases the modification cost as the number of modified road objects increases. For example, if a larger number of road objects are modified, the modification cost is increased compared to when a smaller number of road objects are modified. For example, the re-learning unit 113 may determine the total number of road objects that have been corrected, calculate the correction cost for each road object, and add them up.

[0065] (2) Calculation criteria based on the number of correction operations This method increases the cost of modification as the number of modification operations increases. For example, if an operator performs Nu updates, Ni insertions, and Nd deletions, the modification cost D is CU × Nu + CI × Ni + CD × Nd, where CU is the cost of an update operation, CI is the cost of an insert operation, and CD is the cost of a delete operation.

[0066] (3) Calculation criteria based on the amount of movement of road objects When the position of a road object is corrected, the amount of movement of the road object (distance in the real world) is determined, and the correction cost is calculated according to the amount of movement. For example, if the road object can be represented by a graph, such as a lane boundary line, the amount of movement may be determined by the Gromov-Wasserstein distance.

[0067] (4) Calculation criteria based on the attributes of road objects When a label indicating an attribute of a road object is corrected, the correction cost is calculated based on the similarity between the labels before and after the correction. For example, when the labels before and after the correction are different from each other, the correction cost may be lower than when the labels before and after the correction are similar to each other.

[0068] The re-learning unit 113 calculates the overall modification cost for each tile based on these criteria. If there is a tile whose calculated modification cost exceeds a predetermined threshold, the re-learning unit 113 determines that the estimation model needs to be re-learned. When the re-learning unit 113 determines that re-learning of the estimation model is necessary, the re-learning unit 113 analyzes the target tile and, based on the results, acquires a learning data set to be used for re-learning.

[0069] For example, the re-learning unit 113 analyzes a camera image used when generating information about a target tile to determine the road environment of the image. The road environment is selected from a plurality of predefined classes. For example, suppose that the analysis of a camera image used when generating a certain tile derives the classes "roads with four or more lanes" and "motorway." In this case, it is estimated that the estimation model has poor estimation accuracy for the classes "roads with four or more lanes" and / or "motorway."

[0070] Therefore, the re-learning unit 113 extracts training data that matches the determined road environment from the training dataset and uses this to re-train the estimation model. As described with reference to FIG. 3, each piece of training data is associated with a label that indicates the class of the road environment. The re-learning unit 113 may use this as a key to acquire data to be used for re-learning. For example, in the above example, training data including images taken on roads with four or more lanes and / or expressways is extracted. By repeating this operation, the accuracy of the estimation model can be improved.

[0071] The road map data after the operator has made corrections can also be used as data for relearning. In this case, the probe data associated with the target tile is the input data, and the result of the correction (for example, a set of road object attributes and location information) is the training data. The relearning unit 113 may generate such data at the timing of relearning and add it to the learning dataset. Note that the road environment label corresponding to the target tile is inherited.

[0072] [flowchart] Next, the process executed by the server device 1 according to this embodiment will be described. 5 is a flowchart of the process executed by the server device 1. The process shown in the figure is executed periodically.

[0073] First, in step S11, the generation unit 111 collects probe data from a plurality of vehicles (on-board devices 2) under its management. The collection of the probe data may be triggered by the generation unit 111 or may be triggered by the on-board device 2. The collected probe data is stored in the storage unit 12. In this step, the collection process may be continued until a certain amount of probe data is accumulated. In this embodiment, the road map data is divided into multiple tiles, and the road map data is updated on a tile-by-tile basis. Therefore, for example, the collection of probe data may be continued until a sufficient amount of probe data is accumulated to recognize all road objects present in a tile.

[0074] In step S12, the generation unit 111 updates the road map data based on the collected probe data. In this step, the generation unit 111 converts the camera images included in the probe data into features and inputs them into the estimation model. The generation unit 111 also converts the location information output from the estimation model into a geographic coordinate system and maps it to the road map data together with labels indicating attributes. As a result, a road object is placed in the target tile of the road map data along with the attribute label and location information. Note that if a road object has already been placed in the target tile, the attribute label and location information of the existing road object may be updated using information obtained from the estimation model. For example, a weighted average may be performed based on the recency of the information and the likelihood (accuracy) of the estimation, and the result may be overwritten. The position information output from the estimation model may be expressed in a camera image system or in a geographic coordinate system. By inputting information about the position and direction of a vehicle into the estimation model, it is also possible to obtain position information of road objects expressed in a geographic coordinate system from the estimation model.

[0075] Next, in step S13, the generation unit 111 determines whether or not the tiles updated in step S12 require visual confirmation by an operator. For example, if the tile updated in step S12 is a new tile, i.e., a tile on which a road object has been placed for the first time, it may be determined that it is necessary to have the operator confirm whether or not correction is necessary. In this case, if the tile updated in step S12 is a new tile, the device starts processing from step S14 onwards. If the tile updated in step S12 is not a new tile, i.e., a tile on which a road object has already been placed, the processing returns to step S11.

[0076] In this example, the process from step S14 onward is started on the condition that "the tile updated in step S12 is a new tile," but other conditions may be used as long as it is possible to determine whether visual confirmation by an operator is necessary. For example, if the likelihood (accuracy) of the estimation made in step S12 is lower than a predetermined value, a positive determination may be made in this step.

[0077] In step S14, the correction unit 112 presents information about the target tile to the operator of the device, and allows the operator to determine whether or not correction is required. In this step, the modifying unit 112 generates and outputs a user interface for modifying the road map data. The user interface output by the modifying unit 112 includes, for example, an interface for adding, deleting, or modifying the position of a road object. Through this interface, the operator can delete erroneously determined road objects, add unrecognized road objects, correct the positions of road objects whose positions are incorrect, etc. When the correction operation is completed, the process proceeds to step S15. Note that if no correction is required, the process may end here.

[0078] In step S15, the correction unit 112 reflects the content of the correction in the road map data and saves the content of the correction in a correction log. As described above, the correction log records information for calculating the correction cost. The correction log may record a log of the correction operation, or may record the difference between the road map data before and after the correction.

[0079] Next, in step S16, the re-learning unit 113 calculates a cost (correction cost) corresponding to the correction made by the operator. The correction cost can be calculated based on, for example, the number of correction operations, the amount of operation, the number of road objects to be corrected, the amount of correction (amount of movement of position), etc. If there is a tile whose calculated correction cost exceeds a predetermined threshold, the re-learning unit 113 determines that re-learning of the estimation model is necessary to improve the estimation accuracy for that tile (step S16—Yes). In this case, the process proceeds to step S17. If there is no tile whose calculated correction cost exceeds the predetermined threshold, it determines that re-learning of the estimation model is not necessary (step S16—No), and the process ends.

[0080] FIG. 6 is a flowchart showing the process executed in step S17. First, in step S171, the re-learning unit 113 determines the road environment of a target tile, i.e., a tile whose correction cost exceeds a threshold. The road environment can be determined, for example, using camera images acquired in a geographical area corresponding to the tile. For example, if the probe data collected in step S11 is stored in the storage unit 12, the re-learning unit 113 may extract the probe data acquired in the target tile and acquire the camera image included in the extracted probe data. The re-learning unit 113 can determine the road environment by analyzing the camera image using a machine learning model. The machine learning model used here can be a model trained to input a camera image and output a classification result of the road environment. Note that the classification corresponds to the class of the road environment label assigned to each training data in the training dataset.

[0081] Next, in step S172, the re-learning unit 113 acquires a data set for re-learning that corresponds to the road environment identified in step S171 from the storage unit 12. For example, if the road environment class of "motorway" is identified in step S171, multiple pieces of training data to which the road environment label "motorway" is assigned are extracted from the training data set in the storage unit 12.

[0082] The number of pieces of learning data used for re-learning may be set according to the correction cost determined in step S16. For example, the larger the correction cost, the larger the number of pieces of learning data used for re-learning may be.

[0083] Next, in step S173, the re-learning unit 113 re-learns the estimation model using the extracted training data.

[0084] As described above, the server device according to this embodiment is configured to generate road map data using a machine learning model based on camera images acquired by a vehicle, and is also configured to allow an operator (curator) to modify the road map data. Furthermore, a dataset for retraining the machine learning model is determined based on the cost of the modification and the corresponding road environment. This configuration makes it possible to improve the estimation accuracy of the machine learning model at a lower cost.

[0085] (Variation) The above-described embodiment is merely an example, and the present disclosure can be modified and implemented as appropriate within the scope that does not deviate from the gist of the disclosure. For example, the processes and means described in this disclosure can be freely combined and implemented as long as no technical contradiction occurs.

[0086] In the embodiment, a camera image is exemplified as sensor data collected from a vehicle. ,Other sensor data can also be used, provided that the attributes and positions of road objects can be determined.

[0087] Furthermore, a process described as being performed by one device may be shared and executed by multiple devices. Alternatively, a process described as being performed by different devices may be executed by a single device. In a computer system, the hardware configuration (server configuration) by which each function is realized can be flexibly changed.

[0088] The present disclosure can also be realized by providing a computer program implementing the functions described in the above embodiments to a computer, and having one or more processors in the computer read and execute the program. Such a computer program may be provided to the computer via a non-transitory computer-readable storage medium connectable to the computer's system bus or via a network. Examples of non-transitory computer-readable storage media include any type of disk, such as a magnetic disk (e.g., a floppy disk, a hard disk drive (HDD), etc.), an optical disk (e.g., a CD-ROM, a DVD disk, a Blu-ray disk), a read-only memory (ROM), a random access memory (RAM), an EPROM, an EEPROM, a magnetic card, a flash memory, an optical card, or any type of medium suitable for storing electronic instructions. [Explanation of symbols]

[0089] 1. Server device 2...In-vehicle device 11,21 Control unit 12,22...Storage section 13,23···Communications Department 14...Input / output section 24 Sensor group

Claims

1. acquiring position information of one or more road objects by inputting sensor data obtained by sensing a road environment into a machine learning model; acquiring a result of an operator making corrections to the first map data in which the position information of the road object is reflected; determining whether or not re-learning of the machine learning model is necessary based on the cost of the correction operation performed by the operator; An information processing device having a control unit that executes the above.

2. the control unit determines a cost of the correction operation performed by the operator for each of a plurality of unit areas included in the first map data, and determines that re-learning of the machine learning model is necessary when there is a first area that is a unit area whose cost exceeds a predetermined threshold. The information processing device according to claim 1 .

3. the control unit determines real-world road environment attributes corresponding to the first area; retraining the machine learning model using a training dataset including the determined road environment attributes; The information processing device according to claim 2 .

4. the control unit determines the number of learning data to be used in the re-learning based on the cost of the correction operation performed by the operator. The information processing device according to claim 3 .

5. the control unit adds the sensor data corresponding to the first area as input data and the corrected position information of the road object as teacher data to the learning data set. The information processing device according to claim 3 .

6. acquiring position information of one or more road objects by inputting sensor data obtained by sensing a road environment into a machine learning model; acquiring a result of an operator making corrections to the first map data in which the position information of the road object is reflected; determining whether or not re-learning of the machine learning model is necessary based on the cost of the correction operation performed by the operator; An information processing method executed by a computer.

Citation Information

Patent Citations

  • Detection device and machine learning method

    JP2020086786A

  • Object data curation of map information using neural networks for autonomous systems and applications

    JP2023131069A