Data retrieval device
The data search device efficiently retrieves desired scenes from driving data by calculating feature amounts and similarities, addressing the inefficiencies of manual sorting and limited keyword searches, and enhancing AI testing and data evaluation processes.
Patent Information
- Application Number
- JP2023198517
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-06-03
AI Technical Summary
Existing methods for searching driving data, such as those used for AI testing in autonomous driving, are inefficient as they require manual visual checking and sorting of large datasets, and are limited by the tags or keywords used, which cannot handle complex search queries like 'green trucks' or diverse data retrieval.
A data search device that accumulates past driving data with calculated feature amounts, receives user input data and search conditions, calculates feature amounts and similarities, and extracts data satisfying the search conditions based on these similarities, allowing for selection of search conditions that prioritize similarity, dissimilarity, or representative data.
Enables efficient and accurate retrieval of desired scenes from driving data, reducing the need for manual sorting and improving the ability to handle complex search queries, thereby enhancing the efficiency of AI testing and data evaluation.
Smart Images

Figure 2025084541000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a data search device that extracts data satisfying search conditions from past data.
Background Art
[0002] For the new development and defect correction of artificial intelligence (AI) for autonomous driving, analysis using driving data collected from vehicles is essential. For example, if it is suspected that the failure rate of lane detection is high from the operating conditions of an AI product, past driving data including scenes measured under various conditions such as driving in rainy weather, when the lane is blurred, or during high-speed driving is retrieved from a database, and then the data included in those scenes is used to comprehensively test the lane detection AI.
[0003] However, driving data is a combination of various data such as the video of a drive recorder, time-series numerical data from an acceleration sensor, etc., and operation logs obtained from an ECU (Electronic Control Unit). As it is, it is not possible to search for scenes that satisfy a specific condition such as "when the lane is blurred". Therefore, conventionally, in order to collect scenes that satisfy the conditions required for AI testing, a huge amount of accumulated past driving data has to be visually checked and sorted, which is time-consuming and has the problem of missing a lot of data.
[0004] There are roughly two types of solutions to this problem. The first solution is a method of making it searchable using tags newly added to driving data. For example, Non-Patent Document 1 discloses a method of detecting scenes such as overtaking operations and lane changes with AI and searching using the added tags.
[0005] The second solution is similar search. For example, Patent Document 1 discloses a method of searching for a desired scene regardless of tags or keywords by searching general videos based on the similarity of partial images.
Prior Art Documents
Patent Document
[0006]
Patent Document 1
Non-Patent Document
[0007]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0008] However, in the method disclosed in Non-Patent Document 1, since tags are assigned using AI, it is only possible to search using the tags defined during the development of the AI. For example, even if an object detection AI is tagged with "truck" to enable searching for scenes where a truck exists ahead, it cannot handle the use case of only searching for "green trucks", and there are cases where the amount of data that needs to be visually confirmed cannot be sufficiently reduced. Also, when creating test data using the tags assigned by another object detection AI to test a defect of a certain object detection AI, there is a problem that the evaluation result of the former object detection AI is affected by the technical limitations of the latter object detection AI and cannot be correctly evaluated.
[0009] On the other hand, in the method disclosed in Patent Document 1, if an image of "green truck" is given as a search query, an image similar to it can be obtained, so the problems of Non-Patent Document 1 can be solved. However, there are also problems in cases where it is desired to obtain an image that is not similar (dissimilar) to the image given as the search query depending on the application, or in cases where it is desired to obtain a wide variety of image data comprehensively regardless of similarity or dissimilarity.
[0010] Therefore, in view of the above problems, an object of the present invention is to provide a data search device that searches for a desired scene from driving data including various information such as video and sensor data and makes it acquirable.
Means for Solving the Problems
[0011] In order to solve the above problems, a data search device according to the present invention includes a past data management unit that accumulates past data together with feature amounts calculated based on feature information included in the past data, a reception process that receives a search query including user input data and search conditions, a feature amount calculation process that calculates a feature amount based on feature information included in the user input data, a similarity calculation process that calculates a similarity by comparing the feature amount related to the user input data and the feature amount related to the past data, and a search condition application process that extracts data satisfying the search conditions from the past data based on the calculated similarity, and a search execution unit that executes the processes, and the search conditions can be selected from among a plurality of conditions in which at least three types of relationships with the user input data are set.
Effects of the Invention
[0012] According to the present invention, it is possible to provide a data search device that searches for a desired scene from driving data including various information such as video and sensor data and makes it acquirable. Further features related to the present invention will become apparent from the description in this specification and the accompanying drawings. In addition, problems, configurations, and effects other than those described above will be clarified by the description of the following embodiments.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Modes for Carrying Out the Invention
[0014] Specific embodiments of the present invention will be described below with reference to the drawings. (System Configuration) FIG. 1 is a diagram showing a configuration example of a data search system including a data search server (data search device) 100 according to an embodiment of the present invention. In this embodiment, the data search system has a data search server 100 and a user terminal 200. The data search server 100 is communicably connected to a user terminal 200 used by a user of the data search system.
[0015] The data search server 100 includes, as its functional units, a past data management unit 110, a search execution unit 120, and a search result presentation unit 130.
[0016] The past data management unit 110 of the data search server 100 accumulates driving data composed of, for example, images collected in advance from a vehicle's drive recorder, sensor data such as speed, latitude, and longitude collected from an ECU, and system logs such as an accelerator operation state collected from an ECU, in a past data information storage unit 111. The past data management unit 110 reads the past data information storage unit 111 and records, in a past data feature amount information storage unit 113, a feature amount calculated by executing a predetermined feature amount calculation process based on the characteristic information included in the past data. This feature amount calculation process is executed by a feature amount calculation processing unit 112. The feature amount calculation process executes a process defined in advance according to the viewpoint of expressing it as a feature amount and the calculation target such as images and sensor data included in the past data information storage unit 111.
[0017] Here, the viewpoint in this embodiment means a concept expressed by a part of the driving data. For example, for video data, upper-level viewpoints such as a subject, and lower-level viewpoints such as a vehicle, a pedestrian, a building, a train, a two-wheeler, an animal, a lost item, a road surface condition, wetness, a crosswalk, a road marking, a lane, unevenness, a parking lot, a tunnel, a signal, a road sign, a utility pole, a streetlight, a guardrail, a sidewalk, a bridge, an intersection, a railway, a highway, a branch, a toll gate, a service area, a cloverleaf interchange, a time zone, weather, a light source, and lens dirt of a camera can be defined.
[0018] In addition, for sensor data, viewpoints can be defined from aspects such as the types of sensors such as speed, acceleration, latitude, and longitude, and from aspects consisting of operation records such as deceleration, acceleration, left turn, right turn, the state of the direction indicator, warning sound, and the operating state of the collision damage mitigation braking device.
[0019] Furthermore, it is also possible to define a higher-level viewpoint of a situation defined by a combination of sensor data and video data, and lower-level viewpoints such as traffic jams, interruptions, merging, traffic accidents, lane changes, and overtaking, which are types of situations.
[0020] According to such viewpoints, for example, for the viewpoint of a vehicle among the subjects included in the video data, an object detection AI that performs object detection on the video data is applied, a partial image determined to be a vehicle is extracted, and the embedding vector output by the neural network with the partial image as the input can be treated as a feature amount. Also, after extracting frame images from the video data, feature points in the image can be calculated using an OpenCV program or the like and treated as feature amounts. In the case where the data is time-series numerical data such as speed included in the sensor data, in addition to the embedding vector obtained by processing partial time-series data extracted with a predetermined window size by a neural network, vectors obtained by normalizing the partial time-series data and vectors composed of statistical quantities such as maximum values and minimum values can also be treated as feature amounts.
[0021] The past data information storage unit 111 may store not only the data accumulated in advance but also newly add data during the operation of the data search server 100 of the present embodiment. Regarding the newly added driving data, the feature amount obtained by executing the processing by the feature amount calculation processing unit 112 when the data is added to the past data information storage unit 111 may be recorded in the past data feature amount information storage unit 113.
[0022] When a new calculation method for feature quantities with respect to the above viewpoints is added to the feature quantity calculation processing unit 112, for all the running data already accumulated in the past data information storage unit 111, the added calculation method may be executed by the feature quantity calculation processing unit 112, and the feature quantities may be recorded in the past data feature quantity information storage unit 113.
[0023] The search execution unit 120 in the present embodiment executes a reception process for receiving a search query including user input data and search conditions, a feature quantity calculation process for calculating feature quantities based on the characteristic information included in the user input data, a similarity calculation process for calculating a similarity by comparing the feature quantities related to the user input data and the feature quantities related to past data, and a search condition application process for extracting data that satisfies the search conditions from the past data based on the calculated similarity.
[0024] The functions of the search execution unit 120 will be described in more detail. The search execution unit 120 of the data search server 100 executes the reception process by the search condition reception processing unit 121 according to a request from the user terminal 200. The search condition reception processing unit 121 receives a set of user input data and search conditions as a search query and stores it in the search query information storage unit 122. Here, the user input data is the same type of data as the running data stored in the past data information storage unit 111, and may be a part of the running data, such as only including video data. The search conditions can be selected one from a plurality of conditions in which at least three types of relationships with data are set. In the present embodiment, the values are selected from three types: "similar", "diversified", and "dissimilar", and are set for each viewpoint. This search condition is used by the search condition application processing unit 127 described later.
[0025] In the feature quantity calculation processing unit 123 of the search execution unit 120, the user input data included in the search query information is read out, and in the same manner as the feature quantity calculation processing unit 112 of the past data management unit 110, feature quantities are calculated by a predetermined method defined for each viewpoint based on the characteristic information included in the input data, and are stored in the input data feature quantity information storage unit 124.
[0026] The similarity calculation unit 125 of the search execution unit 120 reads out the feature amounts stored in the input data feature amount information storage unit 124 and the past data feature amount information storage unit 113 respectively, compares them, calculates the similarity between the feature amounts belonging to the same perspective, and stores the result in the similarity calculation result storage unit 126. At this time, known calculation methods such as cosine similarity, Euclidean distance, and dynamic time warping method can be used for the calculation method of similarity.
[0027] The search condition application unit 127 of the search execution unit 120 reads out the similarity stored in the similarity calculation result storage unit 126, extracts the data that meets the search conditions described in the search query information storage unit 122 based on the read similarity, and ranks them. Specifically, for the perspective where "similar" is specified as the search condition, those with higher similarity are prioritized, and for the perspective where "dissimilar" is specified as the search condition, those with lower similarity are prioritized for ranking. Also, for the perspective where "diversification" is specified as the search condition, representative data in each region obtained by clustering the feature space is extracted and ranked. Specifically, ranking is performed so that the distance between the data displayed as search results is maximized, and the ranking result is stored in the search result information storage unit 131 of the search result presentation unit 130. At this time, any known calculation method such as Euclidean distance, Mahalanobis distance, or Manhattan distance can be used for the calculation method of the distance between data. Also, as a method for maximizing the distance between data, for example, the K-means clustering can be applied to the feature amounts related to the perspective, and the center point of each cluster can be selected, but the method is not limited to this.
[0028] Finally, the search result display processing unit 132 of the search result presentation unit 130 reads out the search results stored in the search result information storage unit 131 and displays them on the user terminal 200.
[0029] At this time, the data search server 100 may receive an additional search query in response to a request from the user terminal 200. This additional search query includes one or more additional data selected from the past data and additional search conditions. After storing the additional search query received in the search condition receiving processing unit 121 of the search execution unit 120 in the search query information storage unit 122, the feature amount calculation processing unit 123 calculates a feature amount based on the characteristic information of the additional data included in the additional search query. Then, the similarity calculation processing unit 125 calculates the similarity between the feature amount related to the additional data and the feature amount related to the search result, the search condition application processing unit 127 updates the search result information storage unit 131, and updates the information to be displayed on the user terminal 200 through the search result display processing unit 132 of the search result presentation unit 130.
[0030] As shown in the search result display screen of FIG. 15, the additional data included in the additional search query may be selected by an interactive operation by the user from a diagram in which the user input data and the past data are plotted as points in the feature amount space. Further, when the past data is data for which classification processing is performed by a machine learning model, it may be selected from the past data by a predetermined threshold value for the classification probability. Further, the search execution unit 120 may further include an inquiry information unit including the number of inquiries for the past data, and the additional data may be preferentially selected from the past data with a large number of inquiry information.
[0031] (Hardware Configuration) FIG. 2 is a configuration diagram showing an example of the hardware configuration of the data search server 100. As shown in FIG. 2, the data search server 100 includes a storage device 101, an arithmetic device 103, a memory 104, and a communication device 105, and each device is connected to each other via a bus 106.
[0032] The storage device 101 is composed of non-volatile memory elements such as an SSD (Solid State Drive) and a hard disk drive. The storage device 101 stores the program 102 that defines the operations of the arithmetic unit 103 and various information used or generated by the arithmetic unit 103. The memory 104 is composed of a volatile memory element such as a RAM (Random Access Memory).
[0033] The arithmetic unit 103 is composed of a processor such as a CPU (Central Processing Unit). The arithmetic unit 103 has the functions of each functional unit described in FIG. 1, reads the program 102 stored in the storage device 101 into the memory 104, and executes it. Specifically, the arithmetic unit 103 reads out the feature amount calculation processing programs 112', 123', the search condition reception processing program 121', the similarity calculation processing program 125', the search condition application processing program 127', and the search result display processing program 132' to realize the processing by each functional unit shown in FIG. 1. The communication device 105 communicates with an external device such as the user terminal 200 shown in FIG. 1 via the network 107.
[0034] (Example of data structure) Hereinafter, an example of the data structure of various data handled in this embodiment will be described. FIG. 3 is a diagram showing a configuration example of the data stored in the past data information storage unit 111. The past data information storage unit 111 shown in FIG. 3 has fields 111a to 111c. The field 111a stores a data ID which is identification information for identifying past data. The field 111b stores the video file name of the data. The field 111c stores the sensor data file name of the data.
[0035] FIG. 4 is a diagram showing a configuration example of sensor data information 111-2, and shows an example of the content of sensor data defined by the sensor data file name included in field 111c of the past data information storage unit 111 shown in FIG. 3. The sensor data information 111-2 shown in FIG. 4 has fields 111-2a to 111-2f. Field 111-2a stores a Timestamp indicating the time when the data was recorded. Fields 111-2b to 111-2f store the values of various data measured at the time when the data was recorded. Specifically, field 111-2b stores the measured value of the vehicle speed. Field 111-2c stores the measured value of the accelerator operation state. Field 111-2d stores the measured value of the brake operation state. Field 111-2e stores the measured value of the latitude. Field 111-2f stores the measured value of the longitude.
[0036] FIG. 5 is a diagram showing a configuration example of the feature quantity calculation method 112-1, and is information stored as a set value in the feature quantity calculation processing unit 112 and the feature quantity calculation processing unit 123. The feature quantity calculation method 112-1 has fields 112-1a to 112-1c. Field 112-1a stores a value indicating the type of viewpoint. Field 112-1b stores a value indicating the type of target data. Field 112-1c is a field that stores a calculation method of a feature quantity according to the viewpoint defined by field 112-1a and the target data defined by field 112-1b, and in the configuration example of FIG. 5, shows the execution method of the program.
[0037] FIG. 6 is a diagram showing a configuration example of the data stored in the past data feature quantity information storage unit 113. The past data feature quantity information storage unit 113 shown in FIG. 6 has fields 113a to 113c. Field 113a is a data ID for uniquely identifying the data, and stores the same value as field 111a in FIG. 3. Field 113b is a field indicating the viewpoint, and stores one of the viewpoints defined by field 112-1a in FIG. 5. Field 113c is a field that stores the feature quantity, and stores a vector obtained as a result of executing the calculation method defined by field 112-1c in FIG. 5.
[0038] FIG. 7 is a diagram showing a configuration example of data stored in the search query information storage unit 122. The search query information storage unit 122 shown in FIG. 7 has fields 122a to 122d. Field 122a is a field indicating the name of the input data, and stores an arbitrary character string input through the user terminal 200. Field 122b is a field indicating the input file name, and stores the file name uploaded through the user terminal 200. Field 122c is a field indicating the perspective, and stores a value indicating the type of perspective specified by the user terminal 200. Field 122d is the search condition, and stores a value indicating the condition specified by the user terminal 200 among "similar", "diversified", and "dissimilar".
[0039] FIG. 8 is a diagram showing a configuration example of data stored in the input data feature amount information storage unit 124. The input data feature amount information storage unit 124 in FIG. 8 has fields 124a to 124b. Field 124a is a field indicating the perspective, and stores a value indicating one of the perspectives defined by field 112-1a in FIG. 5. Field 124b is a field for storing feature amounts, and stores a vector obtained as a result of executing the calculation method defined by field 112-1c in FIG. 5.
[0040] FIG. 9 is a diagram showing a configuration example of data stored in the similarity calculation result storage unit 126. The similarity calculation result storage unit 126 in FIG. 9 has fields 126a to 126c. Field 126a is a data ID that uniquely indicates past data. Field 126b is a field indicating the perspective, and stores a value indicating one of the perspectives defined by field 112a in FIG. 5. Field 126c is the similarity, and stores a similarity value calculated by a predetermined method for the feature amounts of the past data having the data ID shown in FIG. 6 and the feature amounts of the input data shown in FIG. 8.
[0041] FIG. 10 is a diagram showing a configuration example of data related to search result information (by perspective) 131-1 that summarizes the rankings by perspective included in the search result information storage unit 131. The search result information (by perspective) 131-1 in FIG. 10 has fields 131-1a to 131-1c. Field 131-1a is a data ID that uniquely indicates past data. Field 131-1b is a field indicating a perspective and stores a value indicating one of the perspectives defined by field 112a in FIG. 5. Field 131-1c stores a rank that is an arbitrary integer.
[0042] FIG. 11 is a diagram showing a configuration example of data related to search result information (comprehensive) 131-2 in which the results by perspective included in the search result information storage unit 131 are ranked for each data ID. The search result information (comprehensive) 131-2 in FIG. 11 has fields 131-2a to 131-2b. Field 131-2a is a rank that is an arbitrary integer. Field 131-2b stores a data ID that uniquely indicates past data.
[0043] (Flow example) FIG. 12 is a flowchart for explaining an example of the operation of the feature calculation processing unit 112 executed by the data search server 100. If "past data" in FIG. 12 is replaced with "input data", it can also be regarded as a flowchart for explaining an example of the operation of the feature calculation processing unit 123.
[0044] In the feature quantity calculation processing unit 112, first, past data is acquired from the past data information storage unit 111 (step S101). At this time, if all past data has been processed, the process ends (Yes in step S102). If there is still unprocessed past data, the unprocessed past data is selected (step S103). Next, it is confirmed whether all viewpoints included in the past data have been processed. If there is no unprocessed viewpoint, the process returns to step S102 to continue the process (Yes in step S104). If there is an unprocessed viewpoint, the viewpoint is selected (step S105), the feature quantity of the viewpoint is calculated according to the calculation method described in the feature quantity calculation method 112-1 shown in FIG. 5, stored in the past data feature quantity information storage unit 113, and then the process returns to step S102 (step S106).
[0045] FIG. 13 is a flowchart for explaining an example of the operations of the search execution unit 120 and the search result presentation unit 130 of the data search server 100. The search condition reception processing unit 121 executed by the search execution unit 120 first receives an input of a search query from the user terminal 200 and stores it in the search query information storage unit 122 (step S201).
[0046] Next, in the feature quantity calculation processing unit 123, the feature quantity of the input data included in the search query is calculated for each viewpoint by a predetermined calculation method and stored in the input data feature quantity information storage unit 124 (step S202). In the similarity calculation processing unit 125, feature quantities are acquired from the input data feature quantity information storage unit 124 and the past data feature quantity information storage unit 113 respectively, the similarity is calculated for each viewpoint, and the result is stored in the similarity calculation result storage unit 126 (step S203). In the search condition application processing unit 127, based on the search conditions and similarities included in the search query information, data that satisfies the search conditions is extracted from the past data information storage unit 111, ranked, and stored in the search result information storage unit 131 (step S204).
[0047] The search result presentation unit 130 reads the search result information storage unit 131 by the search result display processing unit 132 and displays it on the search screen of the user terminal 200 from the top in terms of ranking (step S205).
[0048] (Screen example) FIG. 14 shows an example of a search condition input screen 301. The search condition input screen 301 may be displayed and output on a user terminal 200 such as a personal computer, for example.
[0049] The search condition input screen 301 shown in FIG. 14 is an example of a screen that receives input of a search query from a user. The user inputs an input data name and a file name in the field 301b of the search condition input screen 301 and presses the registration button. Thereafter, the input data name registered in the field 301a and the content of the file are displayed, so that the content of the data given as the input data can be confirmed.
[0050] There are a plurality of fields 301b for registering input data and fields 301a for displaying the content thereof, and the user may register video data and sensor data simultaneously, for example. The field 301a is provided with a seek bar for selecting the time of video data or time-series data, and the user can specify a part of the input data as a search query by operating the seek bar by mouse click or the like.
[0051] Also, for sensor data, as shown as the field 301c, the user may specify part of the sensor data, i.e., partial time-series data, as a search query by selecting the range of the sensor data by mouse operation. When the delete button attached to the field 301a is pressed, the registered input data name and the content of the file are deleted.
[0052] The field 301d in FIG. 14 receives input of search conditions consisting of a viewpoint, an input data name, and search conditions from the user. Also, by pressing the add button in the field 301e, the input field in the field 301d can be increased to accept a plurality of search conditions. After the user specifies the input data name, file, and search conditions, by clicking the search execution button in the field 301f, values can be passed to the search condition reception processing unit 121 to execute a series of search processes by the data search server 100.
[0053] In FIG. 14, video data and sensor data are input as input data. The video data is a situation where a sedan-type vehicle is running under clear weather, and the sensor data is data in which the vehicle's own speed increases and then decreases.
[0054] And as search conditions, diversification with respect to the viewpoint of weather, similarity with respect to the viewpoint of vehicle, and dissimilarity with respect to the viewpoint of own vehicle speed are selected.
[0055] That is, as data satisfying this search condition, data indicating "regardless of the type of weather, the vehicle is a vehicle similar to the sedan type, and the own vehicle speed is stable or changing at a certain gradient" can be cited as an example.
[0056] FIG. 15 shows an example of a search result display screen 401. The search result display screen 401 may be displayed and output on a user terminal 200 such as a personal computer. The search result display screen 401 includes a field 401a and presents information such as the search ranking of past data, measurement date and time, measurement location, subject, own vehicle speed, and similarity to the user.
[0057] Compare and confirm the results shown in FIG. 15 with the search conditions in FIG. 14. The data extracted as the first rank is snow for the weather, sedan type for the vehicle, and in a state of accelerating at a low speed and a certain gradient for the own vehicle speed, and it can be said that it satisfies the search conditions shown in FIG. 14 to a high standard.
[0058] Regarding the data of the second and third ranks, the weather is also diverse respectively. For the vehicle, the similarity decreases as the rank goes down, and for the own vehicle speed, the similarity increases as the rank goes down. Therefore, it can be said that the displayed search results satisfy the search conditions shown in FIG. 14.
[0059] FIG. 16 shows an example of a data selection screen 501. The data selection screen 501 may be displayed and output on a user terminal 200 such as a personal computer. The data selection screen 501 includes fields 501a to 501g, and visualizes the distribution of past data by plotting the past data as points on a feature space according to the viewpoint as shown in field 501a.
[0060] The user can select a plotted point by a mouse operation and display and check the content of the selected data as shown in fields 501b and 501c. Also, in order to distinguish past data from input data, the input data may be highlighted by changing the shape of the plotted point as in field 501c. The user can select additional data as in 501d by clicking the mouse on the plotted point, and perform a re-search using the additional data by entering search conditions in field 501e and then clicking the search execution button in field 501g.
[0061] In FIG. 16, vehicle is selected as the additional search viewpoint and non-similar is selected as the additional search condition. In this case, the previously selected search condition: similar for the viewpoint: vehicle is overwritten, and a scene where vehicles non-similar to the sedan type are shown is extracted as the top of the search results.
[0062] According to the embodiments of the present invention described above, the following operational effects can be obtained.
[0063] (1) The data search device according to the present invention includes a past data management unit that accumulates past data together with feature amounts calculated based on characteristic information included in the past data, a reception process that receives a search query including user input data and search conditions, a feature amount calculation process that calculates a feature amount based on the characteristic information included in the user input data, a similarity calculation process that calculates a similarity by comparing the feature amount related to the user input data and the feature amount related to the past data, and a search condition application process that extracts data satisfying the search conditions from the past data based on the calculated similarity. The search execution unit that executes the above processes is provided, and the search conditions can be selected from among a plurality of conditions in which at least three types of relationships with the user input data are set.
[0064] With the above configuration, since the search conditions can be set according to the desired result, it is possible to search for and acquire a desired scene from the driving data including various information such as images and sensor data.
[0065] (2) The search conditions include a similar extraction condition that preferentially extracts data with a high similarity. It is preferable that this condition is included in the search conditions.
[0066] (3) The search conditions include a dissimilar extraction condition that preferentially extracts data with a low similarity. It is preferable that this condition is included in the search conditions.
[0067] (4) The search conditions include a representative point extraction condition that extracts representative data of each region obtained by clustering the feature amount space. It is preferable that this condition is included in the search conditions.
[0068] (5) The search conditions are selected from among a similar extraction condition that preferentially extracts data with a high similarity, a dissimilar extraction condition that preferentially extracts data with a low similarity, and a representative point extraction condition that extracts representative data of each region obtained by clustering the feature amount space. This form that includes all of the search conditions in (2) to (4) and allows one to be selected is most preferable.
[0069] (6) The feature quantity is an embedding vector calculated by a neural network based on the characteristic information included in the user input data, and the similarity is calculated as the cosine similarity. This makes it possible to easily grasp the feature quantity and calculate the similarity.
[0070] (7) The search execution unit receives an additional search query including one or more selected data selected from past data and an additional search condition, calculates a feature quantity based on the characteristic information of the selected data included in the additional search query, compares the feature quantity related to the selected data with the feature quantity related to the search result, and extracts data that satisfies the additional search condition for the selected data from the search result. This makes it possible to obtain new search results without having to start the search from the beginning after confirming the search.
[0071] (8) The past data is data for which classification processing is performed by a machine learning model, and the selected data is selected from the past data by a predetermined threshold value for the classification probability of the past data. This makes it possible to easily obtain desired selected data by appropriately setting the threshold value.
[0072] (9) The search execution unit further includes an inquiry information unit including the number of inquiries for the past data, and the selected data is selected from the past data as data with a large number of inquiry information. This makes it possible to regard data with many inquiries as data with high importance and preferentially select them.
[0073] (10) The selected data is selected by a dialogical operation by the user from a diagram in which the user input data and the past data are plotted as points on the feature quantity space. Thus, it is preferable to set so that the user can arbitrarily select data.
[0074] (11) The user input data has a plurality of viewpoints, and the search condition is selectable for each of the viewpoints. This makes it possible to set different search conditions for each viewpoint, so that the range of data that can be obtained in one search is greatly widened.
[0075] Note that the present invention is not limited to the above-described embodiments, and various modifications are possible. For example, the above embodiments have been described in detail for easy understanding of the present invention, and the present invention is not necessarily limited to the aspect having all the configurations described. Also, it is possible to replace a part of the configuration of one embodiment with the configuration of another embodiment. Further, it is possible to add the configuration of another embodiment to the configuration of one embodiment. Also, it is possible to delete a part of the configuration of each embodiment, or add or replace other configurations.
Description of Reference Numerals
[0076] 100: Data search server (data search device), 110: Past data management unit, 112: Feature quantity calculation processing unit, 120: Search execution unit, 121: Search condition reception processing unit, 123: Feature quantity calculation processing unit, 125: Similarity calculation processing unit, 127: Search condition application processing unit
Claims
1. A data search device, a past data management unit that accumulates past data together with feature amounts calculated based on characteristic information included in the past data; a reception process that receives a search query including user input data and search conditions, a feature amount calculation process that calculates a feature amount based on the characteristic information included in the user input data, a feature amount related to the user input data, and a feature amount related to the past data, and a similarity calculation process that calculates a similarity by comparing them, and a search condition application process that extracts data that satisfies the search conditions from the past data based on the calculated similarity, and a search execution unit that executes the process, wherein the search condition can be selected from among a plurality of conditions in which at least three types of relationships with the user input data are set. A data search device characterized by this.
2. The data search device according to claim 1, wherein the search condition includes a similarity extraction condition for preferentially extracting data having a high similarity. A data search device characterized by this.
3. The data search device according to claim 1, wherein the search condition includes a dissimilarity extraction condition for preferentially extracting data having a low similarity. A data search device characterized by this.
4. The data search device according to claim 1, wherein the search condition includes a representative point extraction condition for extracting representative data of each region obtained by clustering a feature amount space. A data search device characterized by this.
5. The data search device according to claim 1, wherein the search condition is selected from among a similarity extraction condition for preferentially extracting data having a high similarity, a dissimilarity extraction condition for preferentially extracting data having a low similarity, and a representative point extraction condition for extracting representative data of each region obtained by clustering a feature amount space. A data search device characterized by this.
6. The data search device according to claim 1, wherein the feature amount is an embedding vector calculated by a neural network based on the characteristic information included in the user input data, and the similarity is calculated as a cosine similarity. A data search device characterized by this.
7. The data search device according to claim 1, The search execution unit receives an additional search query including one or more selected data selected from the past data and additional search conditions, calculates a feature amount based on the characteristic information of the selected data included in the additional search query, compares the feature amount related to the selected data with the feature amount related to the search result, and extracts data that satisfies the additional search conditions for the selected data from the search result. A data search device characterized by the above.
8. The data search device according to claim 7, wherein the past data is data for which classification processing is performed by a machine learning model, and the selected data is selected from the past data according to a predetermined threshold with respect to the classification probability of the past data. A data search device characterized by the above.
9. The data search device according to claim 7, wherein the search execution unit further includes an inquiry information unit including the number of inquiries for the past data, and the selected data is selected from the past data as data with a large number of inquiry information. A data search device characterized by the above.
10. The data search device according to claim 7, wherein the selected data is selected by an interactive operation by the user from a diagram in which the user input data and the past data are plotted as points in a feature amount space. A data search device characterized by the above.
11. The data search device according to claim 1, wherein the user input data has a plurality of viewpoints, and the search conditions are selectable for each viewpoint. A data search device characterized by the above.
Citation Information
Patent Citations
Similarity retrieval processing method and device and program
JP2007323319A