Data search device

The data search device addresses the inefficiencies in searching driving data by using feature amounts and customizable search conditions to efficiently retrieve relevant scenes from driving data, enhancing the accuracy and speed of data collection for AI testing.

WO2025109961A1PCT designated stage expired Publication Date: 2025-05-30ASTEMO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/038493
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-22
Filing Date
2024-10-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing methods for searching driving data, such as those used for AI testing in autonomous driving, are inefficient and prone to data oversights due to the need for manual visual checking and sorting of large datasets.

Method used

A data search device that accumulates past driving data with calculated feature amounts, allowing for user input-based search queries with customizable search conditions, including similarity, dissimilarity, and diversification, to extract relevant data.

Benefits of technology

Enables efficient and accurate retrieval of desired scenes from driving data, reducing the time and effort required for data collection and evaluation, while minimizing data oversights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024038493_30052025_PF_FP_ABST
    Figure JP2024038493_30052025_PF_FP_ABST
Patent Text Reader

Abstract

A data search device according to the present invention comprises: a past data management unit that accumulates past data together with a feature amount calculated on the basis of feature information included in the past data; and a search execution unit that executes a reception process for receiving a search query including user input data and a search condition, a feature amount calculation process for calculating a feature amount on the basis of the feature information included in the user input data, a similarity degree calculation process for comparing the feature amount related to the user input data with a feature amount related to the past data and calculating the degree of similarity, and a search condition application process for extracting data satisfying the search condition from the past data on the basis of the calculated degree of similarity. The search condition can be selected from among a plurality of conditions in which at least three types of relationship with the user input data have been set.
Need to check novelty before this filing date? Find Prior Art

Description

Data search device

[0001] The present invention relates to a data search device that extracts data that satisfies a search condition from past data.

[0002] Analysis of driving data collected from vehicles is essential for the development of new artificial intelligence (AI) for autonomous driving and for bug fixes. For example, if the operating conditions of an AI product suggest that the lane detection failure rate is high, it is necessary to search a database for past driving data including scenes measured under various conditions, such as driving in the rain, when lane marks are blurred, or driving at high speeds, and then comprehensively test the lane detection AI using the data contained in those scenes.

[0003] However, driving data is a combination of various data such as footage from a drive recorder, time-series numerical data from acceleration sensors, and operation logs obtained from an ECU (Electronic Control Unit), and as such, it is not possible to search for scenes that meet specific conditions, such as "when the lane markings are blurred." Therefore, in the past, in order to collect scenes that meet the conditions necessary for AI testing, it was necessary to visually check and sort through the huge amount of driving data accumulated in the past, which was time-consuming and often resulted in data being overlooked.

[0004] There are roughly two types of solutions to this problem. The first solution is to make the driving data searchable using newly added tags. For example, Non-Patent Document 1 discloses a method in which scenes such as overtaking maneuvers and lane changes are detected using AI, and then searched using the added tags.

[0005] The second solution is similarity search. For example, Patent Document 1 discloses a method for searching for a desired scene by searching general video based on the similarity of partial images, without relying on tags or keywords.

[0006] Japanese Patent Application Laid-Open No. 2007-323319

[0007] IVS (Intertempora Validation Suite) - dSPACE https: / / www.dspace.com / ja / jpn / home / products / sw / datenmanaagement / ivs.cfm (accessed August 30, 2023)

[0008] However, the method disclosed in Non-Patent Document 1 uses AI to assign tags, so searches can only be performed using tags defined during AI development. For example, even if an object detection AI is tagged with a "truck" tag to enable searching for scenes in which a truck is present ahead, this does not support applications where users only want to search for "green trucks," and the amount of data that must be visually confirmed may not be sufficiently reduced. Furthermore, if test data is created using tags assigned by another object detection AI to test defects in one object detection AI, the evaluation results of the former object detection AI may be affected by the technical limitations of the latter object detection AI, making it impossible to accurately evaluate the former object detection AI.

[0009] On the other hand, the method disclosed in Patent Document 1 can acquire images similar to a "green truck" by inputting the image of the "green truck" as a search query, thereby solving the problem of Non-Patent Document 1. However, there are problems in that, depending on the application, it may be desired to acquire images that are not similar (dissimilar) to the image inputted as the search query, or it may be desired to acquire a comprehensive variety of image data regardless of whether they are similar or dissimilar.

[0010] In view of the above, an object of the present invention is to provide a data search device that can search for and acquire a desired scene from driving data that includes a variety of information such as video and sensor data.

[0011] In order to solve the above problem, the data search device of the present invention includes a past data management unit that accumulates past data together with features calculated based on characteristic information contained in the past data; a reception process that receives a search query including user input data and search conditions; a feature calculation process that calculates the feature based on the characteristic information contained in the user input data; a similarity calculation process that compares the feature related to the user input data with the feature related to the past data to calculate the similarity; and a search condition application process that extracts data that satisfies the search conditions from the past data based on the calculated similarity, and the search condition can be selected from a plurality of conditions that have at least three types of relationship with the user input data.

[0012] According to the present invention, a data search device can be provided that can search for and acquire a desired scene from driving data that includes various information such as video and sensor data. Further features related to the present invention will become apparent from the description of this specification and the accompanying drawings. In addition, the problems, configurations, and effects described above will become apparent from the following description of the embodiments.

[0013] FIG. 1 is a diagram showing an example of a network configuration including a data search server in an embodiment of the present invention. FIG. 1 is a diagram showing an example of a hardware configuration of a data search server in an embodiment of the present invention. FIG. 2 is a diagram showing an example of a configuration of running data information in an embodiment of the present invention. FIG. 3 is a diagram showing an example of a configuration of sensor data information in an embodiment of the present invention. FIG. 4 is a diagram showing an example of a configuration of feature calculation information in an embodiment of the present invention. FIG. 5 is a diagram showing an example of a configuration of running data feature amount information in an embodiment of the present invention. FIG. 6 is a diagram showing an example of a configuration of search conditions in an embodiment of the present invention. FIG. 7 is a diagram showing an example of a configuration of input data feature amount information in an embodiment of the present invention. FIG. 8 is a diagram showing an example of a configuration of similarity calculation results in an embodiment of the present invention. FIG. 9 is a diagram showing an example of a configuration of search result information (by viewpoint) in an embodiment of the present invention. FIG. 10 is a diagram showing an example of a configuration of search result information (overall) in an embodiment of the present invention. FIG. 11 is a flowchart of a feature amount calculation process in an embodiment of the present invention. FIG. 12 is a flowchart of processes executed by a search execution unit and a search result presentation unit in an embodiment of the present invention. FIG. 13 is a diagram showing an example of a configuration of a search condition input screen in an embodiment of the present invention. FIG. 14 is a diagram showing an example of a configuration of a search result display screen in an embodiment of the present invention. FIG. 15 is a diagram showing an example of a configuration of a data selection screen in an embodiment of the present invention.

[0014] Specific embodiments of the present invention will be described below with reference to the drawings. (System Configuration) Figure 1 is a diagram showing an example of the configuration of a data search system including a data search server (data search device) 100 in one embodiment of the present invention. In this embodiment, the data search system has the data search server 100 and a user terminal 200. The data search server 100 is connected so as to be able to communicate with the user terminal 200 used by a user of the data search system.

[0015] The data search server 100 includes, as its functional units, a past data management unit 110, a search execution unit 120, and a search result presentation unit 130.

[0016] The past data management unit 110 of the data search server 100 accumulates driving data in a past data information storage unit 111, which includes video collected in advance from the vehicle's drive recorder, sensor data such as speed, latitude, and longitude collected from the ECU, and system logs such as accelerator operation status collected from the ECU. The past data management unit 110 reads data from the past data information storage unit 111, performs a predetermined feature amount calculation process based on characteristic information included in the past data, and records the calculated feature amounts in the past data feature amount information storage unit 113. This feature amount calculation process is executed by a feature amount calculation processing unit 112. The feature amount calculation process executes predefined processing depending on the viewpoint to be expressed as a feature amount and the calculation target, such as video or sensor data, included in the past data information storage unit 111.

[0017] Here, the term "perspective" in this embodiment refers to a concept expressed by a part of the driving data. For example, for video data, a high-level perspective of a subject, and lower-level perspectives such as a vehicle, pedestrian, building, train, motorcycle, animal, lost item, road surface condition, wetness, crosswalk, road marking, lane, unevenness, parking lot, tunnel, traffic light, road sign, utility pole, street light, guardrail, sidewalk, bridge, intersection, railroad track, expressway, fork, toll booth, service area, overpass, time of day, weather, light source, dirt on the camera lens, etc., which are types of subjects, can be defined.

[0018] In addition, for sensor data, it is possible to define perspectives based on sensor types such as speed, acceleration, latitude, and longitude, as well as perspectives based on operation records such as deceleration, acceleration, left turn, right turn, turn signal status, warning sound, and collision damage mitigation braking device operating status.

[0019] Furthermore, it is possible to define a higher-level perspective, that is, a situation defined by a combination of sensor data and video data, and lower-level perspectives, such as congestion, cutting in, merging, traffic accidents, lane changes, and overtaking, which are types of situations.

[0020] According to such a viewpoint, for example, in the case of a vehicle among the subjects included in the video data, an object detection AI that performs object detection on the video data can be applied, and a partial image determined to be a vehicle can be extracted. The partial image can be used as an input and the embedded vector output by a neural network can be treated as a feature. Alternatively, frame images can be extracted from the video data, and feature points in the image can be calculated using an OpenCV program or the like and treated as a feature. Furthermore, in the case of time-series numerical data such as speed included in the sensor data, in addition to the embedded vector obtained by processing partial time-series data extracted with a predetermined window size using a neural network, a vector obtained by normalizing the partial time-series data or a vector composed of statistical quantities such as maximum and minimum values ​​can also be treated as a feature.

[0021] In addition to data stored in advance in the past data information storage unit 111, new data may be added during operation of the data search server 100 of this embodiment. For newly added driving data, the feature amounts obtained by executing processing by the feature amount calculation processing unit 112 at the time the data is added to the past data information storage unit 111 may be recorded in the past data feature amount information storage unit 113.

[0022] Furthermore, when a new calculation method for a feature for the above viewpoint is added to the feature calculation processing unit 112, the added calculation method may be executed by the feature calculation processing unit 112 for all of the driving data already stored in the past data information storage unit 111, and the feature may be recorded in the past data feature information storage unit 113.

[0023] The search execution unit 120 in this embodiment executes a reception process for receiving a search query including user input data and search conditions, a feature calculation process for calculating feature values ​​based on characteristic information included in the user input data, a similarity calculation process for calculating similarity by comparing feature values ​​related to the user input data with feature values ​​related to past data, and a search condition application process for extracting data that satisfies the search conditions from the past data based on the calculated similarity.

[0024] The function of the search execution unit 120 will be described in more detail. The search execution unit 120 of the data search server 100 executes a search condition reception process by the search condition reception processing unit 121 in response to a request from the user terminal 200. The search condition reception processing unit 121 receives a set of user-input data and search conditions as a search query and stores it in the search query information storage unit 122. Here, the user-input data refers to data of the same type as the driving data stored in the past data information storage unit 111, and may be a portion of the driving data, such as including only video data. The search condition can be selected from a plurality of conditions with at least three types of relationship to the data. In this embodiment, the search condition is a value selected from three types: "similar," "diversified," and "dissimilar," and is set for each perspective. The search condition is used by the search condition application processing unit 127, which will be described later.

[0025] The feature calculation processing unit 123 of the search execution unit 120 reads out the user input data included in the search query information, and similarly to the feature calculation processing unit 112 of the past data management unit 110, calculates the feature using a predetermined method defined for each perspective based on the characteristic information included in the input data, and stores it in the input data feature information storage unit 124.

[0026] The similarity calculation processing unit 125 of the search execution unit 120 reads and compares the features stored in the input data feature information storage unit 124 and the past data feature information storage unit 113, calculates the similarity between the features belonging to the same viewpoint, and stores the results in the similarity calculation result storage unit 126. In this case, the similarity may be calculated using a known calculation method such as cosine similarity, Euclidean distance, or dynamic time warping.

[0027] The search condition application processing unit 127 of the search execution unit 120 reads the similarities stored in the similarity calculation result storage unit 126 and extracts and ranks data that satisfy the search conditions described in the search query information storage unit 122 based on the read similarities. Specifically, for a perspective where "similar" is specified as a search condition, data with high similarity is prioritized, and for a perspective where "dissimilar" is specified as a search condition, data with low similarity is prioritized. Furthermore, for a perspective where "diversification" is specified as a search condition, representative data from each region obtained by dividing the feature space by clustering is extracted and ranked. Specifically, ranking is performed to maximize the distance between data displayed as search results, and the ranking results are stored in the search result information storage unit 131 of the search result presentation unit 130. Any known calculation method, such as Euclidean distance, Mahalanobis distance, or Manhattan distance, may be used to calculate the distance between data. Furthermore, as a method for maximizing the distance between data, for example, a method of applying K-means clustering to the feature amounts related to the viewpoint and selecting the center point of each cluster may be used, but the method is not limited to this.

[0028] Finally, the search result display processing unit 132 of the search result presentation unit 130 reads out the search results stored in the search result information storage unit 131 and displays them on the user terminal 200 .

[0029] At this time, the data search server 100 may accept an additional search query in response to a request from the user terminal 200. This additional search query includes one or more additional data selected from the past data and additional search conditions. After the search condition acceptance processing unit 121 of the search execution unit 120 stores the accepted additional search query in the search query information storage unit 122, the feature calculation processing unit 123 calculates feature amounts based on characteristic information of the additional data included in the additional search query. Then, the similarity calculation processing unit 125 calculates the similarity between the feature amounts related to the additional data and the feature amounts related to the search results, the search condition application processing unit 127 updates the search result information storage unit 131, and the information displayed on the user terminal 200 is updated via the search result display processing unit 132 of the search result presentation unit 130.

[0030] The additional data included in the additional search query may be selected by a user interactively from a diagram in which user-input data and past data are plotted as points in a feature space, as shown on the search result display screen in Figure 15. Alternatively, when the past data is data to be classified by a machine learning model, the additional data may be selected from the past data based on a predetermined threshold value for the classification probability. Alternatively, the search execution unit 120 may further include an inquiry information unit containing the number of inquiries for the past data, and the additional data may be selected from the past data with a higher number of inquiry information items as priority.

[0031] (Hardware Configuration) Fig. 2 is a configuration diagram showing an example of the hardware configuration of the data search server 100. As shown in Fig. 2, the data search server 100 has a storage device 101, an arithmetic unit 103, a memory 104, and a communication unit 105, and each device is connected to each other via a bus 106.

[0032] The storage device 101 is composed of non-volatile storage elements such as a solid state drive (SSD) and a hard disk drive. The storage device 101 stores a program 102 that defines the operation of the arithmetic device 103 and various information used or generated by the arithmetic device 103. The memory 104 is composed of a volatile storage element such as a random access memory (RAM).

[0033] The arithmetic device 103 is configured with a processor such as a CPU (Central Processing Unit). The arithmetic device 103 has the functions of each functional unit described in FIG. 1 and reads the program 102 stored in the storage device 101 into the memory 104 and executes it. Specifically, the arithmetic device 103 reads feature quantity calculation processing programs 112′ and 123′, a search condition reception processing program 121′, a similarity calculation processing program 125′, a search condition application processing program 127′, and a search result display processing program 132′, and realizes processing by each functional unit shown in FIG. 1. The communication device 105 communicates with an external device such as the user terminal 200 shown in FIG. 1 via the network 107.

[0034] (Data Structure Example) From here, an example of the data structure of various data handled in this embodiment will be described. Fig. 3 is a diagram showing an example of the configuration of data stored in the past data information storage unit 111. The past data information storage unit 111 shown in Fig. 3 has fields 111a to 111c. Field 111a stores a data ID, which is identification information for identifying past data. Field 111b stores the video file name of the data. Field 111c stores the sensor data file name of the data.

[0035] FIG. 4 is a diagram showing an example of the configuration of the sensor data information 111-2, and shows an example of the contents of sensor data defined by the sensor data file name included in field 111c of the past data information storage unit 111 shown in FIG. 3. The sensor data information 111-2 shown in FIG. 4 has fields 111-2a to 111-2f. Field 111-2a stores a timestamp indicating the time the data was recorded. Fields 111-2b to 111-2f store the values ​​of various data measured at the time the data was recorded. Specifically, field 111-2b stores the measured value of the vehicle speed. Field 111-2c stores the measured value of the accelerator operation state. Field 111-2d stores the measured value of the brake operation state. Field 111-2e stores the measured value of latitude. Field 111-2f stores the measured value of longitude.

[0036] 5 is a diagram showing an example of the configuration of the feature calculation method 112-1, which is information stored as setting values ​​in the feature calculation processing unit 112 and the feature calculation processing unit 123. The feature calculation method 112-1 has fields 112-1a to 112-1c. Field 112-1a stores a value indicating the type of viewpoint. Field 112-1b stores a value indicating the type of target data. Field 112-1c is a field that stores a feature calculation method according to the viewpoint specified in field 112-1a and the target data specified in field 112-1b, and the example configuration of FIG. 5 shows a program execution method.

[0037] Fig. 6 is a diagram showing an example of the configuration of data stored in the past data feature amount information storage unit 113. The past data feature amount information storage unit 113 shown in Fig. 6 has fields 113a to 113c. Field 113a is a data ID for uniquely identifying data, and stores the same value as field 111a in Fig. 3. Field 113b is a field indicating a viewpoint, and stores one of the viewpoints defined in field 112-1a in Fig. 5. Field 113c is a field for storing a feature amount, and stores a vector obtained as a result of executing the calculation method defined in field 112-1c in Fig. 5.

[0038] Fig. 7 is a diagram showing an example of the configuration of data stored in the search query information storage unit 122. The search query information storage unit 122 shown in Fig. 7 has fields 122a to 122d. Field 122a is a field indicating the name of the input data, and stores any character string input via the user terminal 200. Field 122b is a field indicating the input file name, and stores the file name uploaded via the user terminal 200. Field 122c is a field indicating the viewpoint, and stores a value indicating the type of viewpoint specified via the user terminal 200. Field 122d is a search condition, and stores a value indicating the condition specified via the user terminal 200 from among "similarity," "diversification," and "dissimilarity."

[0039] Fig. 8 is a diagram showing an example of the configuration of data stored in the input data feature amount information storage unit 124. The input data feature amount information storage unit 124 in Fig. 8 has fields 124a to 124b. Field 124a is a field indicating a viewpoint, and stores a value indicating one of the viewpoints defined in field 112-1a in Fig. 5. Field 124b is a field for storing a feature amount, and stores a vector obtained as a result of executing the calculation method defined in field 112-1c in Fig. 5.

[0040] Fig. 9 is a diagram showing an example of the configuration of data stored in the similarity calculation result storage unit 126. The similarity calculation result storage unit 126 in Fig. 9 has fields 126a to 126c. Field 126a is a data ID that uniquely identifies past data. Field 126b is a field that indicates a viewpoint, and stores a value that indicates one of the viewpoints defined in field 112a in Fig. 5. Field 126c is a similarity, and stores a similarity value calculated using a predetermined method for the feature amount of past data having the data ID shown in Fig. 6 and the feature amount of input data shown in Fig. 8.

[0041] Fig. 10 is a diagram showing an example of the configuration of data related to search result information (by viewpoint) 131-1, which summarizes the rankings by viewpoint, and is included in the search result information storage unit 131. The search result information (by viewpoint) 131-1 in Fig. 10 has fields 131-1a to 131-1c. Field 131-1a is a data ID that uniquely identifies past data. Field 131-1b is a field that indicates the viewpoint, and stores a value that indicates one of the viewpoints defined in field 112a in Fig. 5. Field 131-1c stores a ranking, which is an arbitrary integer.

[0042] 11 is a diagram showing an example of the configuration of data related to search result information (general) 131-2, which is stored in the search result information storage unit 131 and which summarizes results by viewpoint and ranks them by data ID. The search result information (general) 131-2 in FIG. 11 has fields 131-2a to 131-2b. Field 131-2a indicates the ranking, which is an arbitrary integer. Field 131-2b stores a data ID that uniquely identifies past data.

[0043] 12 is a flowchart for explaining an example of the operation of the feature amount calculation processing unit 112 executed by the data search server 100. If "past data" in FIG. 12 is replaced with "input data," it can also be regarded as a flowchart explaining an example of the operation of the feature amount calculation processing unit 123.

[0044] The feature calculation processing unit 112 first acquires past data from the past data information storage unit 111 (step S101). At this time, if all past data has been processed, the process ends ("Yes" in step S102). If there is unprocessed past data, the past data is selected (step S103). Next, it is confirmed whether all aspects included in the past data have been processed, and if there are no unprocessed aspects, the process returns to step S102 and continues ("Yes" in step S104). If there is an unprocessed aspect, the aspect is selected (step S105), and the feature of that aspect is calculated according to the calculation method described in the feature calculation method 112-1 shown in FIG. 5, and stored in the past data feature information storage unit 113, after which the process returns to step S102 (step S106).

[0045] 13 is a flowchart illustrating an example of the operation of the search execution unit 120 and the search result presentation unit 130 of the data search server 100. The search condition reception processing unit 121 executed by the search execution unit 120 first receives a search query input from the user terminal 200 and stores it in the search query information storage unit 122 (step S201).

[0046] Next, the feature calculation processing unit 123 calculates the feature amounts of the input data included in the search query for each aspect using a predetermined calculation method, and stores the calculated result in the input data feature amount information storage unit 124 (step S202). The similarity calculation processing unit 125 acquires the feature amounts from the input data feature amount information storage unit 124 and the past data feature amount information storage unit 113, calculates the similarity for each aspect, and stores the calculated result in the similarity calculation result storage unit 126 (step S203). The search condition application processing unit 127 extracts and ranks data that meets the search conditions from the past data information storage unit 111 based on the search conditions and similarity included in the search query information, and stores the extracted data in the search result information storage unit 131 (step S204).

[0047] The search result display unit 130 reads the search result information from the search result information storage unit 131 using the search result display processing unit 132, and displays the results on the search screen of the user terminal 200 in descending order of ranking (step S205).

[0048] 14 shows an example of a search condition input screen 301. The search condition input screen 301 may be displayed or output on the user terminal 200 such as a personal computer.

[0049] 14 is an example of a screen that accepts a search query input from a user. The user inputs an input data name and file name in field 301b of the search condition input screen 301 and presses the register button. The input data name and file contents registered in field 301a are then displayed, allowing the user to confirm the contents of the data provided as input data.

[0050] There are multiple fields 301b for registering input data and multiple fields 301a for displaying its contents, and the user may register, for example, video data and sensor data simultaneously. Field 301a has a seek bar for selecting the time of video data or time-series data, and the user can specify part of the input data as a search query by operating the seek bar with a mouse click or the like.

[0051] As shown in field 301c, the user may specify partial time-series data that constitutes part of the sensor data as a search query by selecting the range of the sensor data with a mouse. When the delete button attached to field 301a is pressed, the registered input data name and file contents are deleted.

[0052] 14 accepts search criteria input from the user, including a viewpoint, an input data name, and search criteria. By pressing the add button in field 301e, the input field in field 301d can be expanded to accept multiple search criteria. After specifying the input data name, file, and search criteria, the user can click the execute search button in field 301f to pass the values ​​to the search criteria acceptance processing unit 121 and execute a series of search processes by the data search server 100.

[0053] In Figure 14, video data and sensor data are input as input data, with the video data showing a sedan-type vehicle traveling under clear skies, and the sensor data showing the vehicle's own speed increasing and then decreasing.

[0054] As search conditions, viewpoint: diversification with respect to weather, viewpoint: similarity with respect to vehicles, and viewpoint: dissimilarity with respect to vehicle speed are selected.

[0055] In other words, an example of data that satisfies this search condition is data that indicates "weather of any type, a vehicle similar to a sedan, and vehicle speed that is stable or changes at a constant gradient."

[0056] 15 shows an example of a search result display screen 401. The search result display screen 401 may be displayed or output on a user terminal 200 such as a personal computer. The search result display screen 401 has a field 401a, and presents information such as the search ranking of past data, measurement date and time, measurement location, subject, vehicle speed, and similarity to the user.

[0057] The results shown in Fig. 15 are checked against the search conditions in Fig. 14. The data extracted as the top ranking indicates that the weather is snow, the vehicle is a sedan, and the vehicle speed is low and accelerating at a constant gradient, which means that the data satisfies the search conditions shown in Fig. 14 to a high standard.

[0058] The second and third ranked data also show a variety of weather conditions, and the lower the ranking, the lower the similarity for the vehicle, while the lower the similarity for the vehicle speed. Therefore, it can be said that the displayed search results satisfy the search conditions shown in Figure 14.

[0059] 16 shows an example of a data selection screen 501. The data selection screen 501 may be displayed or output on a user terminal 200 such as a personal computer. The data selection screen 501 includes fields 501a to 501g, and visualizes the distribution of past data by plotting past data as points on a feature space according to a viewpoint, as shown in field 501a.

[0060] The user can select a plot point with the mouse, and the contents of the selected data can be displayed and confirmed as shown in fields 501b and 501c. Furthermore, to distinguish between past data and input data, the input data may be highlighted by changing the shape of the plot point as shown in field 501c. The user can select additional data as shown in 501d by clicking the mouse on the plot point, enter search conditions in field 501e, and then click the search execution button in field 501g to perform a search again using the additional data.

[0061] 16, vehicles are selected as the additional search perspective and dissimilarity as the additional search condition. In this case, the previously selected search condition similarity for the perspective vehicle is overwritten, and scenes showing vehicles dissimilar to sedan types are extracted as the top search results.

[0062] The above-described embodiment of the present invention provides the following advantageous effects.

[0063] (1) A data search device according to the present invention includes a past data management unit that accumulates past data together with features calculated based on characteristic information contained in the past data; a reception process that receives a search query including user input data and search conditions; a feature calculation process that calculates the features based on the characteristic information contained in the user input data; a similarity calculation process that compares the features related to the user input data with the features related to the past data to calculate similarity; and a search condition application process that extracts data that satisfies the search conditions from the past data based on the calculated similarity, and the search condition can be selected from a plurality of conditions that have at least three types of relationship with the user input data.

[0064] With the above configuration, search conditions can be set according to the results you want to obtain, making it possible to search for and obtain desired scenes from driving data that includes a variety of information such as video and sensor data.

[0065] (2) The search conditions include a similarity extraction condition that prioritizes data extraction based on high similarity. It is preferable that the search conditions include this condition.

[0066] (3) The search conditions include a dissimilar extraction condition that prioritizes data extraction with low similarity. It is preferable that the search conditions include this condition.

[0067] (4) The search conditions include a representative point extraction condition for extracting representative data for each region obtained by dividing the feature space by clustering. It is preferable that the search conditions include this condition.

[0068] (5) The search condition is selected from a similarity extraction condition that prioritizes data extraction of high similarity, a dissimilarity extraction condition that prioritizes data extraction of low similarity, and a representative point extraction condition that extracts representative data of each region obtained by dividing the feature space by clustering. This form, which includes all of the search conditions (2) to (4) and allows selection of one, is the most preferable.

[0069] (6) The feature is an embedding vector calculated by a neural network based on the feature information contained in the user-input data, and the similarity is calculated as a cosine similarity, which makes it easy to grasp the feature and calculate the similarity.

[0070] (7) The search execution unit receives an additional search query containing one or more selected data items selected from past data and additional search criteria, calculates feature quantities based on characteristic information of the selected data items included in the additional search query, compares the feature quantities related to the selected data items with the feature quantities related to the search results, and extracts data from the search results that satisfy the additional search criteria for the selected data. This makes it possible to obtain new search results after confirming the search without having to start the search again from the beginning.

[0071] (8) The past data is data that is classified by a machine learning model, and the selected data is selected from the past data based on a predetermined threshold value for the classification probability of the past data. This makes it possible to easily obtain the desired selected data by appropriately setting the threshold value.

[0072] (9) The search execution unit further includes an inquiry information unit including the number of inquiries about the past data, and the selected data is selected from the past data with a large number of inquiries. This allows the data with a large number of inquiries to be considered as important data and to be selected with priority.

[0073] (10) The selected data is selected by the user through interactive operations from a diagram in which the user-input data and past data are plotted as points on the feature space. In this way, it is preferable to set up the system so that the user can arbitrarily select data.

[0074] (11) User-entered data has multiple perspectives, and search conditions can be selected for each of the perspectives. This allows different search conditions to be set for each perspective, significantly expanding the range of data that can be obtained in a single search.

[0075] It should be noted that the present invention is not limited to the above-described embodiments, and various modifications are possible. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to embodiments that include all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment. It is also possible to add the configuration of another embodiment to the configuration of one embodiment. It is also possible to delete part of the configuration of each embodiment, or to add or replace other configurations.

[0076] 100: Data search server (data search device), 110: Past data management unit, 112: Feature amount calculation processing unit, 120: Search execution unit, 121: Search condition reception processing unit, 123: Feature amount calculation processing unit, 125: Similarity calculation processing unit, 127: Search condition application processing unit

Claims

1. A data search device comprising: a past data management unit that accumulates past data together with features calculated based on characteristic information contained in the past data; and a search execution unit that executes the following steps: a reception process that receives a search query including user input data and search conditions; a feature calculation process that calculates the feature based on the characteristic information contained in the user input data; a similarity calculation process that compares the feature related to the user input data with the feature related to the past data to calculate a similarity; and a search condition application process that extracts data that satisfies the search condition from the past data based on the calculated similarity; and wherein the search condition can be selected from a plurality of conditions having at least three types of relationship with the user input data.

2. A data search device according to claim 1, wherein the search conditions include similarity extraction conditions for extracting data with a high degree of similarity as a priority.

3. A data search device according to claim 1, wherein the search conditions include dissimilar extraction conditions for extracting data with a higher priority than that for data with a lower similarity.

4. A data search device according to claim 1, wherein the search conditions include a representative point extraction condition for extracting representative data for each area obtained by dividing a feature space by clustering.

5. A data search device as claimed in claim 1, characterized in that the search condition is one selected from a similarity extraction condition which gives priority to extracting data with a high degree of similarity, a dissimilarity extraction condition which gives priority to extracting data with a low degree of similarity, and a representative point extraction condition which extracts representative data of each area obtained by dividing a feature space by clustering.

6. A data search device as claimed in claim 1, characterized in that the feature amount is an embedding vector calculated by a neural network based on characteristic information contained in the user input data, and the similarity is calculated as a cosine similarity.

7. A data search device as described in claim 1, wherein the search execution unit accepts an additional search query including one or more selected data selected from the past data and additional search conditions, calculates a feature amount based on characteristic information of the selected data included in the additional search query, compares the feature amount related to the selected data with the feature amount related to the search results, and extracts data that satisfies the additional search conditions for the selected data from the search results.

8. A data search device as described in claim 7, wherein the past data is data that is subjected to classification processing by a machine learning model, and the selected data is selected from the past data based on a predetermined threshold value for the classification probability of the past data.

9. A data search device according to claim 7, wherein the search execution unit further comprises an inquiry information unit including a number of inquiries for the past data, and the selected data is data having a large number of inquiry information items selected from the past data.

10. A data search device according to claim 7, characterized in that the selected data is selected by an interactive operation by a user from a diagram in which the user-input data and the past data are plotted as points on a feature space.

11. A data search device according to claim 1, wherein the user input data has a plurality of viewpoints, and the search condition is selectable for each of the viewpoints.

Citation Information

Patent Citations

  • Device and method for retrieving image

    JP1994168277A

  • Similar image retrieval device

    JP2018165926A

  • Data processing system, and data processing device

    WO2012073526A1