Ai model learning data generation apparatus, ai model learning data generation method, ai model learning data generation system, program, and ai model learning data structure
The system addresses the risk of contaminated learning data by comparing candidate images from in-vehicle devices with similar images to filter out tampered content, ensuring high accuracy in AI model training.
Patent Information
- Application Number
- JP2023213275
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-30
AI Technical Summary
Conventional methods for collecting learning data for AI models, particularly for in-vehicle devices, risk contamination with illegally altered images, leading to decreased learning accuracy due to unauthorized access and tampering.
A system and method that acquire candidate images from in-vehicle devices, compare them with images having similar shooting positions and directions from a plurality of images, and select images with similarity equal to or greater than a threshold as learning data, thereby filtering out potentially tampered images.
Effectively prevents the inclusion of illegally altered images in the learning data, thereby maintaining high learning accuracy by accurately detecting and excluding fraudulent images.
Smart Images

Figure 2025097152000001_ABST
Abstract
Description
Technical Field
[0001] The disclosed embodiments relate to an AI model learning data generation device, an AI model learning data generation method, an AI model learning data generation system, a program, and a data structure of AI model learning data.
Background Art
[0002] Conventionally, a technique of collecting captured images taken by a plurality of cameras and using them as learning data for an AI model such as image recognition AI (Artificial Intelligence) is known. Further, in such a technique, in order to suppress the collection of inappropriate images unsuitable for learning, there is one that performs an appropriateness determination of an image in accordance with a predetermined determination criterion in advance before being registered as learning data (see, for example, Patent Document 1).
[0003] In this appropriateness determination, for example, an image in which an object other than an object to be learned, such as a human hand, is captured, or an image with insufficient illuminance is determined as an inappropriate image.
[0004] The collection of such learning data is also required, for example, when learning an AI model for in-vehicle devices. However, in recent years, learning data for in-vehicle devices often uses camera images of an actually running vehicle. In this case, the camera image of the vehicle is uploaded to a server device that collects learning data via a network from the vehicle.
[0005] In addition, when collecting data from a vehicle via a network in this way, there is a concern that unauthorized access from a malicious attacker may occur. For such unauthorized access to in-vehicle devices, for example, a technique of detecting an attack frame in an in-vehicle network such as CAN (Controller Area Network) has been proposed (see, for example, Patent Document 2).
Prior Art Documents
Patent Documents
[0006] [Patent Document 1] Japanese Patent Application Laid-Open No. 2020-008904 [Patent Document 2] Japanese Patent Application Laid-Open No. 2017-111796 [Summary of the Invention] [Problems to be Solved by the Invention]
[0007] However, when the above-described conventional technology is used, there is a risk that the attacker may target the system, and the uploaded image may be replaced with an illegally altered image, etc., so that unintended data may be mixed into the learning data, resulting in a decrease in learning accuracy.
[0008] One aspect of the embodiment is made in view of the above, and an object is to provide an apparatus for generating learning data for an AI model, a method for generating learning data for an AI model, a system for generating learning data for an AI model, a program, and a data structure of learning data for an AI model that can effectively prevent a decrease in learning accuracy due to the inclusion of illegal images. [Means for Solving the Problems]
[0009] The apparatus for generating learning data for an AI model according to one aspect of the embodiment includes a controller. The controller acquires a candidate image, position information where the candidate image was taken, and a shooting direction, acquires an image whose shooting position and shooting direction are approximate to the candidate image as a comparison image from a plurality of images, calculates a similarity between the candidate image and the comparison image, selects the candidate image whose similarity is equal to or greater than a threshold value as learning data, and non-selects the candidate image whose similarity is less than the threshold value as the learning data. [Effects of the Invention]
[0010] According to one aspect of the embodiment, by comparing with a normal image considered to be normal, an illegally replaced illegal image can be detected with high accuracy, and a decrease in learning accuracy due to the inclusion of an illegal image can be effectively prevented.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Embodiments for Carrying Out the Invention
[0012] Hereinafter, with reference to the accompanying drawings, embodiments of an AI model learning data generation device, an AI model learning data generation method, an AI model learning data generation system, a program, and a data structure of AI model learning data disclosed in the present application will be described in detail. Note that the present invention is not limited by the embodiments shown below.
[0013] Also, hereinafter, it is assumed that the learning data generation device according to the embodiment is the server device 100 (both are referred to hereinafter from FIG. 1) included in the learning data generation system 1.
[0014] The learning data generation system 1 is a system that collects captured images taken by in-vehicle devices mounted on vehicles as "candidate images" of learning data used for training an AI model. Further, the learning data generation system 1 is a system that detects whether the collected captured images are forged images replaced by an attacker. Further, the learning data generation system 1 is a system that selects captured images determined not to be forged images as learning data.
[0015] The method for generating learning data for an AI model according to the embodiment is assumed to be a method for generating learning data executed by the controller 103 (see FIG. 4) of the server device 100 included in this learning data generation system 1. Further, hereinafter, it is assumed that the in-vehicle device according to the embodiment is a communication-type drive recorder 10 (see FIGS. 1 and subsequent) provided communicably with the server device 100.
[0016] In the following, when it is necessary to distinguish the same components, they may be numbered in the form of "-m" or "-n" (m and n are arbitrary natural numbers) after the reference numeral. For example, the drive recorder 10 is represented as drive recorders 10-1, 10-2,... 10-m. When there is no particular need to distinguish, it is simply represented as the drive recorder 10 without numbering.
[0017] Further, the expression "predetermined" in the following description may be read as "pre-determined".
[0018] First, the outline of the learning data generation method according to the embodiment will be described with reference to FIG. 1. FIG. 1 is a schematic explanatory diagram of the learning data generation method according to the embodiment.
[0019] In the learning data generation method according to the embodiment, as shown in FIG. 1, the server device 100 executes a learning data selection process of selecting learning data from the captured data of the target vehicle (the vehicle from which the images of the learning data are collected) that has been acquired. Note that this process will be sequentially performed for each captured data of the acquired target vehicle, but for the sake of easy understanding, the process of one captured data will be described.
[0020] The controller 103 of the server device 100 acquires imaging data (including image data, imaging position information, and imaging angle of view (including at least imaging direction information)) from the target vehicle (step S1). Note that the acquired imaging data is preferably imaging images at a preset time interval, imaging data at a preset specific position (for example, an intersection or a point where a specific type of facility exists (in the vicinity)), etc. Further, the controller 103 acquires a group of comparison image data (including image data, imaging position information, and imaging angle of view (including at least imaging direction information)) to be compared with the imaging data (step S2). Note that the comparison image data may be, for example, imaging image data in the vicinity of the imaging position of the collected image data (for example, within 50 m around the imaging position).
[0021] The controller 103 selects comparison image data whose imaging data, imaging position information, and imaging direction are approximate from the group of comparison image data (step S3). Then, the controller 103 calculates the similarity between the image of the candidate image data and the image of the comparison image data (step S4). Then, the controller 103 selects candidate images with a similarity equal to or higher than a threshold as learning data (step S5). Note that a plurality of images of the comparison image data to be compared with the candidate images may be selected and compared, and the plurality of comparison results may be statistically processed, and selection as learning data may be performed based on the processing result (for example, selection as learning data when the ratio of the determination that the similarity is equal to or higher than the threshold is equal to or higher than the threshold). Then, the above processing is sequentially performed for each imaging data.
[0022] Specifically, in step S1, the controller 103 acquires, from the drive recorder 10-1 of the target vehicle, candidate images IM1 of learning data captured by the camera 13a of the drive recorder 10, imaging position information, angle of view, and imaging data including the imaging direction at the time of imaging. The imaging direction may be estimated from the traveling direction of the vehicle and the mounting state of the camera 13a.
[0023] In step S2, based on the acquired shooting position information, the controller 103 acquires a comparison image data group IMG including comparison images IMC1 to IMC n captured within a range DR at a predetermined distance from the shooting point of the candidate image IM1. At this time, the controller 103 acquires the comparison image data group IMG captured by the drive recorder 10-2 or the camera 30-1 within the range DR. The camera 30 is, for example, a fixed camera installed on the street or road. The controller 103 acquires the comparison image data group IMG captured in the past within the range DR. Hereinafter, a specific comparison image included in the comparison image data group IMG may be represented as "comparison image IMC i " (i is an arbitrary natural number).
[0024] In step S3, the controller 103 calculates the shooting range IR-1 of the candidate image IM1 based on the shooting position information, the field of view angle, and the shooting direction of the candidate image IM1. At the same time, the controller 103 calculates the shooting ranges IRC-1 to IRC-n of the comparison images IMC1 to IMC n . The controller 103 selects the comparison image IMC i with the highest similarity of the shooting range with respect to the shooting range IR-1 of the candidate image IM1.
[0025] In step S4, the controller 103 calculates the similarity between the candidate image IM1 and the selected comparison image IMC i . At this time, the controller 103 calculates the similarity using an algorithm such as AKAZE (Accelerated KAZE). Note that the overlapping shooting range area between the candidate image IM1 and the comparison image IMC i may be cut out in advance, and the similarity between the cut-out images may be calculated. Also, scaling processing for matching the image sizes of both images may be performed as necessary.
[0026] In step S5, the controller 103 selects the candidate image IM1 with a similarity equal to or higher than the threshold as learning data. Further, the controller 103 detects the candidate image IM1 with a similarity lower than the threshold as an illegal image and does not select it as learning data. Then, the same process is performed for each candidate image IM i (where i is an arbitrary natural number), and a large amount of image data for learning data is selectively collected.
[0027] According to the learning data generation method according to the embodiment, the similarity is calculated by comparing the candidate image with the highest similarity in the shooting range with the comparison target image, and the illegal image is detected based on the similarity. That is, if the shooting location and the shooting direction are approximate, structures such as landmarks and natural objects such as mountains included in the captured images will be similar. However, if the structures and natural objects included in the images with approximate shooting locations and shooting directions are different, there is a high possibility that they are illegal images, and it is possible to exclude images with a high possibility of such illegal images from the learning data. That is, according to the learning data generation method according to the embodiment, it is possible to prevent a decrease in learning accuracy due to the inclusion of illegal images.
[0028] Hereinafter, a configuration example of a learning data generation system 1 including a server device 100 to which the learning data generation method according to the above-described embodiment is applied will be described more specifically.
[0029] FIG. 2 is a diagram showing a configuration example of a learning data generation system 1 according to the embodiment. As shown in FIG. 2, the learning data generation system 1 includes drive recorders 10-1, 10-2,... 10-m, cameras 30-1,... 30-n, a server device 100, and a learning device 200.
[0030] Each drive recorder 10, each camera 30, the server device 100, and the learning device 200 are communicably connected to each other via a network N such as the Internet, a mobile phone line network, or a C-V2X (Cellular Vehicle to Everything) communication network.
[0031] The drive recorder 10 records, in a ring buffer memory in an overwritable manner for a certain period, the captured images taken by the camera 13a mounted on the own device, and the captured data including the captured position information, the angle of view, the shooting direction, the time information, etc. of the images.
[0032] Also, when the drive recorder 10 detects an event that satisfies a predetermined collection condition requested from the server device 100 during the startup of the vehicle, it sets the captured data for a certain time before and after the detection time to be write-protected. Also, the drive recorder 10 transmits the captured data to be set to write-protected to the server device 100.
[0033] Also, each drive recorder 10 has a unique identification number and is managed by the server device 100 so as to be uniquely identifiable. Note that when transmitting the captured data, the drive recorder 10 transmits this identification number, and the server device 100 stores the received captured data in association with the identification number. Each drive recorder 10 transmits metadata including, for example, the position information of the vehicle to the server device 100 constantly during the startup of the vehicle.
[0034] As described above, the camera 30 is a fixed camera installed, for example, on the street or road. Each camera 30 has a unique identification number and is managed by the server device 100 so as to be uniquely identifiable. The position information of each camera 30 is also managed by the server device 100. Note that when transmitting the captured data, each camera 30 transmits this identification number, and the server device 100 stores the received captured data in association with the identification number.
[0035] The server device 100 is realized, for example, as a cloud server. The server device 100 is managed by, for example, an operator who operates a data center for collecting vehicle data, an operator who sells various learning data sets based on the collected data, etc. The server device 100 acquires the captured data transmitted from each drive recorder 10 and each camera 30.
[0036] The server device 100 acquires, as a target vehicle, any one of the vehicles equipped with the drive recorder 10, and acquires shooting data including a shooting image that is a candidate image for learning data from this target vehicle (see step S1). Further, the server device 100 executes steps S2 to S5 described with reference to FIG. 1 based on the acquired shooting data.
[0037] In addition, the server device 100 generates an arbitrary learning data set using an image group selected as learning data for artificial intelligence (AI). Specifically, tag data corresponding to the function of artificial intelligence (AI) (in the case of AI for estimating the traveling direction of a vehicle, data on the direction in which the vehicle has traveled in the captured image is used as tag data and is set manually by an operator, for example) is added to each image of the captured image group to generate a learning data set. The learning data set generated by the server device 100 is sold and used by, for example, an AI developer who develops an AI model such as an image recognition AI. The server device 100 distributes the generated learning data set to the learning device 200 managed by, for example, an AI developer who is a purchaser via the network N.
[0038] The learning device 200 executes machine learning such as deep learning using the learning data set distributed from the server device 100, and learns an AI model such as an image recognition AI. The learned AI model is appropriately distributed to in-vehicle devices such as the drive recorder 10. Then, the drive recorder 10 and the like incorporate and operate the distributed image recognition AI, and use the image recognition result of the captured image of the camera to give advice such as safe driving to the driver and the like.
[0039] Next, a configuration example of the drive recorder 10 will be described. FIG. 3 is a diagram showing a configuration example of the drive recorder 10 according to the embodiment. As shown in FIG. 3, the drive recorder 10 includes a communication unit 11, an HMI (Human Machine Interface) unit 12, a sensor unit 13, a storage unit 14, and a controller 15.
[0040] The communication unit 11 is realized by a network adapter or the like. The communication unit 11 is wirelessly connected to the network N, and transmits and receives information to and from the server device 100 and the learning device 200 via the network N.
[0041] The HMI unit 12 is a component that provides interface components for input and output to users or the like who use the drive recorder 10. The HMI unit 12 includes an input interface that receives input operations from a driver or the like. The HMI unit 12 also includes an output interface that presents visual information and audio information to users or the like. For example, the HMI unit 12 includes a touch panel display, a microphone, a speaker, and the like.
[0042] The sensor unit 13 is a group of various sensors mounted on the drive recorder 10. The sensor unit 13 includes, for example, a camera 13a and a GPS (Global Positioning System) sensor 13b.
[0043] The camera 13a is provided so as to be able to capture at least an image outside the vehicle. The camera 13a is attached near the front glass, near the dashboard, or the like, and captures a predetermined imaging range in front of the vehicle. The camera 13a may also be attached near the rear glass or the like and be able to capture a predetermined imaging range behind the vehicle.
[0044] The GPS sensor 13b measures the position of the vehicle based on signals from GPS satellites. Note that the sensor unit 13 may include various sensors other than the camera 13a and the GPS sensor 13b. For example, the sensor unit 13 may include a G sensor that measures the acceleration applied to the vehicle to detect a vehicle collision accident or the like.
[0045] In addition to the sensor unit 13, the drive recorder 10 can be connected to in-vehicle sensors 5, which are various types of sensor groups mounted on the vehicle. The in-vehicle sensors 5 include, for example, a vehicle speed sensor, an accelerator sensor, a brake sensor, and the like. The in-vehicle sensors 5 are connected to the drive recorder 10 via an in-vehicle network such as CAN. The drive recorder 10 stores this data in the storage unit 14 as vehicle driving data.
[0046] The storage unit 14 is realized by a storage device such as a ROM (Read Only Memory), a RAM (Random Access Memory), or a flash memory. In the example of FIG. 3, the storage unit 14 stores the photographed data 14a and the collection condition information 14b.
[0047] The collection condition information 14b is information including the aforementioned collection conditions requested from the server device 100. When the controller 15 described later receives the collection conditions from the server device 100 via the communication unit 11, it rewrites the collection condition information 14b in the storage unit 14. Also, when an event corresponding to this collection condition is detected, the photographed data captured by the camera 13a according to the collection condition is transmitted to the server device 100.
[0048] The controller 15 corresponds to a so-called processor. The controller 15 is realized by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphical Processing Unit), or the like. The controller 15 executes a program according to an illustrated embodiment stored in the storage unit 14, using the RAM as a work area. Also, the controller 15 can be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0049] While the vehicle is in operation, the controller 15 constantly transmits metadata including the vehicle's position information and the like to the server device 100 via the communication unit 11.
[0050] In addition, when the controller 15 detects an event corresponding to the collection condition included in the collection condition information 14b, it transmits the shooting data for a predetermined time before and after the event detection time to the server device 100 via the communication unit 11.
[0051] Next, a configuration example of the server device 100 will be described. FIG. 4 is a diagram showing a configuration example of the server device 100 according to the embodiment. As shown in FIG. 4, the server device 100 includes a communication unit 101, a storage unit 102, and a controller 103. In addition, the HMI unit 50 is appropriately connected to the server device 100.
[0052] The HMI unit 50 is an input / output interface used by the operator of the server device 100. For example, the HMI unit 50 includes a display, a microphone, a speaker, a keyboard, a mouse, and the like.
[0053] The communication unit 101 is realized by a network adapter or the like. The communication unit 101 is connected to the network N by wire or wirelessly, and transmits and receives information to and from the drive recorder 10, the camera 30, and the learning device 200 via the network N.
[0054] The storage unit 102 is realized by a storage device such as a ROM, a RAM, a flash memory, or an HDD (Hard Disk Drive). In the example of FIG. 4, the storage unit 102 stores a drive recorder DB (Database) 102a, a camera DB 102b, a collection condition information DB 102c, a collection data DB 102d, a comparison image DB 102e, a representative image DB 102f, a processing result DB 102g, and a learning data set DB 102h.
[0055] The drive recorder DB102a is a database of management information for each drive recorder 10 managed by the server device 100. This management information includes an identification number for identifying the drive recorder 10 and position information of the vehicle (drive recorder 10) that was most recently transmitted to the server device 100, etc.
[0056] The camera DB102b is a database of management information for each camera 30 managed by the server device 100. This management information includes an identification number for identifying the camera 30 and position information where the camera 30 is installed, etc.
[0057] The collection condition information DB102c is a database that stores the collection conditions of the captured data set via the HMI unit 50 from the operator etc. of the server device 100. The collection conditions stored in the collection condition information DB102c are appropriately transmitted to the drive recorder 10 of the target vehicle via the communication unit 101.
[0058] The collected data DB102d is a database that stores the captured data collected from each drive recorder 10 and each camera 30. The collected data DB102d also stores the captured data collected in the past.
[0059] The comparison image DB102e is a database that stores the acquired comparison image data group IMG. The controller 103 acquires the comparison image data group IMG captured by the drive recorder 10 and the camera 30 other than the target vehicle within the range DR of a predetermined distance from the shooting position of the candidate image based on the shooting position information of the candidate image, and stores it in the comparison image DB102e.
[0060] The controller 103 acquires the comparison image data group IMG captured by the drive recorder 10 and the camera 30 other than the target vehicle in the past within the range DR of a predetermined distance from the shooting position of the candidate image based on the shooting position information of the candidate image from the collected data DB102d, and stores it in the comparison image DB102e.
[0061] The representative image DB102f is a database that stores a pre-provided group of representative images. The representative images are images of representative scenes that can be assumed in advance, such as scenes of general intersections, roads, parking lots, etc. The representative images may be images actually captured from a vehicle, or may be images generated by an image generation AI or the like. The representative image is a comparison image IMC whose shooting range overlaps with the candidate image i When it does not exist in the comparison image data group IMG, an appropriate image (an image of the same scene as the candidate image) is selected and acquired by the controller 103 as the comparison image IMC i Note that the scene of the candidate image can be estimated by, for example, collating the shooting position of the candidate image with map data (using data on facilities (such as buildings) near each location), or the AI recognition result of the candidate image by an AI that estimates the scene from the image. Thereby, even when the comparison image IMC i shot under the same shooting conditions as the candidate image does not exist, a simple similarity determination using the representative image becomes possible.
[0062] The processing result DB102g is a database that stores the processing results related to the similarity determination processing between the candidate image and the comparison image performed by the controller 103. A specific example of the data structure of the processing result DB102g will be described later with reference to FIG. 7.
[0063] The learning dataset DB102h is a database that stores the learning dataset generated by the controller 103 based on the similarity determination result. A specific example of the data structure of the learning dataset DB102h will be described later with reference to FIG. 8.
[0064] The controller 103 corresponds to a so-called processor. The controller 103 is realized by a CPU, MPU, GPU, etc. The controller 103 executes the program according to the illustrated embodiment stored in the storage unit 102, using the RAM as a work area. Also, the controller 103 can be realized by an integrated circuit such as an ASIC or FPGA.
[0065] The controller 103 executes information processing according to the processing procedure shown as a flowchart in FIG. 6. The description using FIG. 6 will be described later.
[0066] Next, a configuration example of the learning device 200 will be described. FIG. 5 is a diagram showing a configuration example of the learning device 200 according to the embodiment. As shown in FIG. 5, the learning device 200 includes a communication unit 201, a storage unit 202, a controller 203, and an HMI unit 150.
[0067] The HMI unit 150 is an input / output interface used by the operator of the learning device 200. For example, the HMI unit 150 includes a display, a microphone, a speaker, a keyboard, a mouse, and the like.
[0068] The communication unit 201 is realized by a network adapter or the like. The communication unit 201 is connected to the network N by wire or wirelessly, and transmits and receives information to and from the drive recorder 10 and the server device 100 via the network N.
[0069] The storage unit 202 is realized by a storage device such as a ROM, a RAM, a flash memory, or an HDD. In the example of FIG. 5, the storage unit 202 stores a learning data set 202a and an AI model 202b.
[0070] The learning data set 202a is a learning data set generated by the server device 100 and transmitted from the server device 100 via the network N.
[0071] The AI model 202b is learned by the controller 203 using the learning data set 202a, and the generated AI model is stored.
[0072] The controller 203 corresponds to a so-called processor. The controller 203 is implemented by a CPU, MPU, GPU, etc. The controller 203 executes a program according to an embodiment (not shown) stored in the storage unit 202, using the RAM as a work area. Also, the controller 203 can be implemented by an integrated circuit such as an ASIC or FPGA.
[0073] The controller 203 acquires a desired learning dataset from the server device 100 via the network N based on a request operation such as that of an operator of the learning device 200, etc., and stores it as the learning dataset 202a. Also, the controller 203 executes machine learning such as deep learning using the learning dataset 202a.
[0074] Also, the controller 203 stores the learned AI model generated by executing machine learning as the AI model 202b. Also, the controller 203 appropriately transmits the learned AI model 202b to the drive recorder 10 via the network N.
[0075] Next, the processing procedure of the information processing executed by the server device 100 will be described with reference to FIG. 6. FIG. 6 is a flowchart showing the processing procedure of the learning data selection process executed by the server device 100 according to the embodiment. This process starts with a start instruction by an operator of the server device 100 or the like. Note that FIG. 6 shows the processing procedure for one candidate image. Therefore, every time a candidate image is acquired, the processing procedure shown in FIG. 6 is repeated.
[0076] First, the controller 103 of the server device 100 acquires the imaging data of the target vehicle (step S101). This imaging data includes the captured image that is a candidate image for the learning data, the imaging position information where this image was captured, the viewing angle, and the imaging direction.
[0077] Based on the shooting position information included in the shooting data, the controller 103 acquires a group of comparison image data IMG shot within a range DR of a predetermined distance from the shooting position of the shooting data (step S102). Specifically, the controller 103 extracts, from the group of images shot by the drive recorder 10 and the camera 30 other than the target vehicle, an image within a range DR of a predetermined distance from the shooting position of the shooting image based on the shooting position information included in the shooting data.
[0078] The controller 103 acquires, as a group of comparison image data IMG, an image shot within a range DR of a predetermined distance from the shooting position of the shooting image by the target vehicle from the past shooting data stored in the collection data DB102d.
[0079] Also, a method of newly causing another drive recorder or the like to shoot a comparison image and acquiring it as a group of comparison image data IMG may be applied, or this method may be added. In this case, the drive recorder 10 of the vehicle within a range DR of a predetermined distance from the shooting position of the shooting image is extracted. Then, for the extracted drive recorder 10, collection condition information (for example, start shooting immediately after reception) capable of collecting the shooting data at the shooting position of the shooting image by the target vehicle is transmitted, and a group of comparison image data IMG of the images shot based on the collection condition information is acquired from the extracted drive recorder 10. In addition, when the extracted drive recorder 10 stores a shooting image that matches the collection condition information (already shot), it is also effective to cause the drive recorder 10 to transmit the image data of the image and acquire it as a group of comparison image data IMG. Also, for the extracted camera 30, instruction data for transmitting the shooting data is transmitted, and a group of comparison image data IMG is acquired from the image shot by the camera 30 and the position information of the camera DB102b.
[0080] For each of the candidate image and the images in the group of comparison image data IMG, the controller 103 calculates a shooting range IR based on the shooting position information, the viewing angle, and the shooting direction (step S103). Then, the controller 103 determines that the candidate image and the comparison image IMC whose shooting ranges overlap by a determination threshold value or more iDetermine whether there is one (step S104).
[0081] A comparison image IMC in which the candidate image and the shooting range overlap i If there is one (step S104, Yes), the controller 103 selects this comparison image IMC i (step S105). Then, the controller 103 calculates the similarity of the shooting target between the candidate image and the selected comparison image IMC i (step S106).
[0082] Here, the existence of the comparison image IMC is determined by whether the candidate image and the shooting range overlap by a determination threshold value or more i For example, an image in which the position where the candidate image was taken is close and the shooting direction (for example, the direction in the horizontal plane (north, south, east, west)) is approximate is selected as the comparison image IMC i The camera 13a of the drive recorder 10 often shoots the front (or rear) at an angle close to parallel to the ground. However, between images with different shooting directions in the horizontal plane, for example, an image taken in the north-south direction and an image taken in the east-west direction, the images are completely different. Therefore, for example, a comparison image IMC in which the difference in the shooting position from the candidate image is equal to or less than a threshold distance (for example, about 5 m) and the difference in the shooting direction is equal to or less than a threshold angle (for example, about 30 degrees) is selected, and the others are excluded i
[0083] For the similarity calculation process in step S106, a known similarity determination algorithm can be used. The controller 103 performs, for example, a similarity calculation process using feature amount detection by AKAZE of OpenCV (Open Source Computer Vision Library).
[0084] When a plurality of comparison images IMC are selected in step S105 i the similarity in step S106 is calculated for all the comparison images IMC i . And in this case, the controller 103 uses a value obtained by performing statistical processing on the calculated similarities, for example, the maximum value of each similarity, for each comparison image IMC i As a representative value of the similarity to [object], this representative value is treated as the similarity in the processing after step S108.
[0085] On the other hand, when there is no comparison image IMC where the candidate image and the shooting range overlap in step S104 i (step S104, No), the controller 103 calculates the similarity between the candidate image and the representative image (step S107). The similarity calculation method here is the same as that in step S106.
[0086] In addition, when there are multiple representative images, the similarity in step S107 is calculated for all representative images. And in this case, the controller 103 uses, for example, the maximum value obtained by performing statistical processing on the calculated similarities as the representative value of the similarity to the representative image, and treats this representative value as the similarity in the processing after step S108.
[0087] The controller 103 determines whether the similarity calculated in step S106 or step S107 is equal to or greater than a predetermined threshold (step S108). Here, when the similarity is equal to or greater than the threshold (step S108, Yes), the controller 103 determines that the candidate image is a legitimate image that is not an illegal image and selects it as learning data (step S109). On the other hand, when the similarity is less than the threshold (step S108, No), the controller 103 determines that the candidate image is an illegal image and does not select it as learning data (step S110).
[0088] In addition, when there are multiple selected comparison images IMC (meeting the selection condition that the candidate image and the shooting range overlap by a judgment threshold or more) i other statistical methods can also be applied, for example, a method of selecting as learning data when the average of the similarities in each comparison image IMC i is equal to or greater than the threshold, or a method of selecting as learning data when the ratio of the comparison images IMC i with a similarity equal to or greater than the judgment threshold is a certain ratio.
[0089] Note that the similarity calculated in step S106 and step S107 is stored in the processing result DB102g. FIG. 7 is a diagram showing an example of the data structure of the processing result DB102g.
[0090] As shown in FIG. 7, the processing result DB102g includes records for each candidate image identified by a "data ID". "Image data" stores the image data of each candidate image. In the example of FIG. 7, each candidate image is represented as candidate images A to G. "Presence / absence of comparison image" is a flag value (here, "present" or "absent") indicating the presence / absence of the comparison image IMC i determined in step S104.
[0091] "First similarity" stores the similarity calculated in step S106 when there is a comparison image IMC i and, when there are multiple comparison images IMC i , a representative value of the similarity is stored. "Second similarity" stores the similarity with the representative image calculated in step S107 when there is no comparison image IMC i and, when there are multiple representative images, a representative value of the similarity is stored. i
[0092] Based on each record of this processing result DB102g, the controller 103 performs determination processing in steps S108 to S110. Here, it is assumed that a predetermined threshold value to be compared in step S108 is "70".
[0093] In this case, in the example of FIG. 7, the controller 103 determines candidate images A, B, D to G whose similarity with the comparison image IMC i or the representative image is equal to or greater than the threshold value "70" as legitimate images and selects them as learning data. Also, the controller 103 iCandidate images C with a similarity less than the threshold "70" to [the reference] are determined as illegal images and not selected as learning data. Here, an example where both the first similarity and the second similarity are determined using the same threshold "70" has been given, but the threshold may be set to different values for the first similarity and the second similarity. For example, since the representative image corresponds to an alternative comparison image after all, the second similarity, which is the similarity to this representative image, may have a higher threshold than the first similarity.
[0094] The controller 103 extracts each record of candidate images that are not illegal images selected as learning data from this processing result DB102g, generates a learning data set, and stores it in the learning data set DB102h. FIG. 8 is a diagram showing an example of the data structure of the learning data set DB102h.
[0095] As shown in FIG. 8, the controller 103 extracts records having a similarity of, for example, the aforementioned predetermined threshold "70" or more from the processing result DB102g, and generates a learning data set identified by the "data set ID".
[0096] At this time, the controller 103 may increase the learning data by performing data augmentation using candidate images with higher reliability. Since the high reliability can be regarded as the magnitude of the similarity, the controller 103 performs data augmentation using candidate images with relatively high similarity. For the data augmentation process, known data augmentation methods can be used. For example, the image is rotated, flipped horizontally, enlarged, image quality (resolution, brightness, contrast, etc.), modified, etc. to generate a transformed image and perform the data augmentation process.
[0097] FIG. 8 shows an example in which the controller 103 performs data augmentation using candidate image A with a threshold greater than the aforementioned threshold "70", i.e., "90" or more, to increase the learning data. In the example of FIG. 8, candidate images A-1 and A-2 with branch numbers assigned in the form of "-1" and "-2" represent data obtained by augmenting candidate image A as the data augmentation source. This can increase the contribution of the learning data with higher reliability to learning.
[0098] Further, for candidate images whose similarity to the representative image is equal to or greater than the threshold value, even if the similarity is equal to or greater than the threshold value, the controller 103 may select only a part of them as learning data by thinning them out. FIG. 8 shows an example in which, among the candidate images D to G that were equal to or greater than a predetermined threshold value “70” in the example of FIG. 7, the controller 103 thinned them out and extracted only the candidate images D and F. In other words, candidate images in which a representative image with a similarity equal to or greater than the threshold value exists are selected as learning data, and candidate images in which a representative image with a similarity equal to or greater than the threshold value does not exist are not selected as learning data.
[0099] As a result, as described above, it is possible to reduce the contribution of the candidate images targeted for learning, which correspond to the alternative comparison images, to the comparison target representative image.
[0100] Note that the above-mentioned “dataset ID” is assigned in units for generating one learning dataset. One learning dataset is generated from, for example, a group of images collected through a series of data collection operations for a certain learning. A certain learning is appropriately planned and executed, for example, by the operator of the server device 100. Alternatively, this learning is appropriately planned and executed based on a request from an AI development operator or the like.
[0101] Further, as shown in FIG. 8, each piece of learning data in the learning dataset generated by the controller 103 is associated with each of the calculated similarities described above. That is, the data structure of the learning data in the learning dataset DB102h includes a record having at least the captured image of the learning data selected by the controller 103 and the similarity associated with this learning data by the controller 103.
[0102] As a result, for example, in the learning device 200, it becomes easy to grasp learning data with higher reliability when performing learning. Further, based on this, it becomes possible to ensure the performance of the AI model to be learned.
[0103] Also, although the processing procedure in FIG. 6 is not shown for simplicity, in the processing of each consecutive candidate image, if it is determined that there is continuity in the illegal image determination, such as when the determination to proceed to step S110 in step S108 continues for a certain number of times or more, the controller 103 may stop collecting candidate images. Thereby, it is possible to effectively prevent illegal images from being mixed in due to continuous attacks. Further, when acquiring candidate images for a plurality of drive recorders 10, the unique identification number of the drive recorder 10 is acquired together with the shooting data, and when candidate images acquired from the same vehicle continue for a certain number of times or more, it is also possible to stop only the candidate images from a specific vehicle that has been attacked.
[0104] As described above, the server device 100 according to the embodiment (corresponding to an example of the "AI model learning data generation device") includes a controller 103. The controller 103 acquires a candidate image, position information where the candidate image was taken, and the shooting direction, and compares the candidate image with an image whose shooting position and shooting direction are approximate to the candidate image from a plurality of images to obtain a comparison image IMC i as, calculates the similarity between the candidate image and the comparison image IMC i selects the candidate image whose similarity is equal to or greater than the threshold as learning data, and does not select the candidate image whose similarity is less than the threshold as the learning data.
[0105] Therefore, according to the server device 100 of the embodiment, since the candidate image of the learning data is compared with other images taken at the same shooting location to verify normality, it is possible to prevent a fraudulent image that has been fraudulently tampered with by an attacker from being mixed in with the learning data. Therefore, it is possible to prevent learning of an AI model using a learning data set including inappropriate learning data. In other words, if the AI model is, for example, an AI model of an image recognition AI, it is possible to prevent the image recognition AI from being unable to perform normal image recognition. In other words, according to the server device 100 of the embodiment, it is possible to detect a fraudulently replaced fraudulent image with high accuracy by comparing with a genuine image that is considered to be normal, and it is possible to effectively prevent a decrease in learning accuracy due to the inclusion of a fraudulent image.
[0106] In the above embodiment, the server device 100 and the learning device 200 are configured as separate entities, but they may be integrated into one such that the server device 100 has a learning function. In this case, the server device 100 may stop learning in response to the above-mentioned continuous attack.
[0107] In the above-described embodiment, an example was given in which the server device 100 automatically detects fraudulent images and does not select them as learning data. However, in cases in which manual confirmation, such as annotation work, is involved, the images may be selected as learning data.
[0108] In the above-described embodiment, the controller 103 selects a comparison image whose shooting range based on the shooting position information, angle of view, and shooting direction is similar to that of the candidate image, but the comparison images to be selected may be further narrowed down by taking into account other information.
[0109] For example, the controller 103 may further acquire the time information when the candidate image was taken, and acquire a comparison image taken in the time zone corresponding to this time information. Further, the controller 103 may acquire date information, weather information, etc. in addition to the time information. Thereby, a more appropriate selection of the comparison image can be made in consideration of the surrounding environment that changes over time, etc., and the accuracy of the validity determination of the candidate image can be improved.
[0110] Also, in the above-described embodiment, specific examples of the predetermined threshold value of the similarity were given as numerical values such as "70", "80", "90", etc., but this is merely an example, and does not limit the actual value of the threshold.
[0111] Further effects and modifications can be easily derived by those skilled in the art. Therefore, a broader aspect of the present invention is not limited to the specific details and representative embodiments described and represented as above. Accordingly, various changes can be made without departing from the spirit or scope of the general inventive concept defined by the appended claims and their equivalents.
Explanation of Reference Numerals
[0112] 1 Learning data generation system 5 In-vehicle sensor 10 Drive recorder 11 Communication unit 12 HMI unit 13 Sensor unit 14 Storage unit 15 Controller 30 Camera 100 Server device 101 Communication unit 102 Storage unit 103 Controller 200 Learning device 201 Communication unit 202 Storage unit 203 Controller
Claims
1. Obtain a candidate image, position information where the candidate image was taken, and a shooting direction, Obtain, as a comparison image, an image from a plurality of images that is similar to the candidate image in terms of the shooting position and shooting direction, Calculate the similarity between the candidate image and the comparison image, Select, as learning data, the candidate image for which the similarity is equal to or greater than a threshold value, A controller that does not select, as the learning data, the candidate image for which the similarity is less than the threshold value, A learning data generation device for an AI model comprising the above.
2. The controller, Further obtains time information when the candidate image was taken, Obtains, as the comparison image, an image that is similar to the candidate image in terms of the time information, The learning data generation device for an AI model according to Claim 1.
3. The controller, Performs data augmentation of the learning data using the learning data for which the similarity is equal to or greater than the threshold value among the selected learning data, The learning data generation device for an AI model according to Claim 1.
4. The controller, When the comparison image cannot be obtained, obtains a plurality of representative images that are images of a pre-provided representative scene as the comparison image, The learning data generation device for an AI model according to Claim 1.
5. The controller, When there exists a representative image for which the similarity is equal to or greater than the threshold value, selects the candidate image as the learning data, The learning data generation device for an AI model according to Claim 4.
6. A method for generating learning data for an AI model executed by a controller, comprising: Obtaining a candidate image, position information where the candidate image was taken, and a shooting direction; Obtaining, as a comparison image, an image from a plurality of images that is similar to the candidate image in terms of the shooting position and shooting direction; Calculating the similarity between the candidate image and the comparison image; Selecting, as learning data, the candidate image for which the similarity is equal to or greater than a threshold value; Not selecting, as the learning data, the candidate image for which the similarity is less than the threshold value; A method for generating learning data for an AI model including the above.
7. Comprising a learning data generation device for an AI model and an in-vehicle device communicably provided with the learning data generation device, The in-vehicle device, Takes a candidate image, Transmits the candidate image, position information where the candidate image was taken, and a shooting direction to the learning data generation device, The learning data generation device, Receiving the candidate image, the position information, and the shooting direction transmitted from the in-vehicle device, Obtaining, as a comparison image, an image from a plurality of images that approximates the candidate image in terms of the shooting position and the shooting direction, Calculating the similarity between the candidate image and the comparison image, Selecting, as learning data, the candidate image whose similarity is equal to or greater than a threshold value, Not selecting, as the learning data, the candidate image whose similarity is less than the threshold value, A learning data generation system for an AI model.
8. Obtaining a candidate image, position information where the candidate image was taken, and a shooting direction, Obtaining, as a comparison image, an image from a plurality of images that approximates the candidate image in terms of the shooting position and the shooting direction, Calculating the similarity between the candidate image and the comparison image, Selecting, as learning data, the candidate image whose similarity is equal to or greater than a threshold value, Not selecting, as the learning data, the candidate image whose similarity is less than the threshold value, A program for causing a computer to execute the above.
9. A data structure of learning data generated by a learning data generation device for an AI model provided with a controller, wherein the controller obtains a candidate image, position information where the candidate image was taken, and a shooting direction, obtains, as a comparison image, an image from a plurality of images that approximates the candidate image in terms of the shooting position and the shooting direction, calculates the similarity between the candidate image and the comparison image, selects, as the learning data, the candidate image whose similarity is equal to or greater than a threshold value, does not select, as the learning data, the candidate image whose similarity is less than the threshold value, and the data structure of the learning data comprises a record having at least the learning data selected by the controller and the similarity associated with the learning data by the controller, A data structure of learning data for an AI model.
Citation Information
Patent Citations
Security processing method and server
JP2017111796A
Learning data collection apparatus, learning data collection system and learning data collection method
JP2020008904A