Data management device and data management method

WO2026176811A1PCT designated stage Publication Date: 2026-08-27ASTEMO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/000332
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-21
Filing Date
2026-01-08
Publication Date
2026-08-27

Smart Images

  • Figure JP2026000332_27082026_PF_FP_ABST
    Figure JP2026000332_27082026_PF_FP_ABST
Patent Text Reader

Abstract

A data management device according to the present invention comprises: a data collection unit that collects, by means of wireless communication lines, a plurality of pieces of data acquired by a sensor mounted on a vehicle; a feature amount calculation unit that calculates a feature amount of each of the plurality of pieces of data; a similarity calculation unit that calculates the similarity of each of the plurality of pieces of data on the basis of the feature amount of each of the plurality of pieces of data; a grouping processing unit that classifies each of the plurality of pieces of data into a plurality of groups on the basis of the similarity of each of the plurality of pieces of data; a degree-of-overlap calculation unit that calculates the degree of overlap of data belonging to each of the plurality of groups; a storage unit that stores the plurality of pieces of data; and a deletion determination unit that sets, as deletion candidates, data stored in the storage unit in descending order of the degree of overlap for each of the plurality of groups on the basis of the calculated degree of overlap. The storage unit deletes the data set as the deletion candidates by the deletion determination unit.
Need to check novelty before this filing date? Find Prior Art

Description

Data Management Device and Data Management Method

[0001] The present invention relates to a data management device and a data management method.

[0002] In the development of automatic driving (AD: Automatic Driving) and advanced driver-assistance systems (ADAS: Advanced driver-assistance systems), it is essential to utilize data (such as images, CAN signals, and time-series numerical data from sensors other than images) acquired by sensors during driving. Among the data during driving, in particular, images in the form of videos have a large capacity. If all of them are stored in the cloud, it will waste storage and increase the storage cost. Therefore, for example, in Patent Document 1, a technique is disclosed in which a sensing device outputs image data including an image of a landscape including a detention target and an image of an intruder mixed in the image, and the cloud performs a process of generating a peripheral landscape browsing image based on the image data. In the technique of Patent Document Ⅰ, only the feature amount of the detention target is included in the data transmitted to the cloud, and the image from which the intruder has been removed is transmitted, and the cloud stores only the feature amount of the detention target. As a result, the data capacity can be reduced compared to the case of storing video data as it is.

[0003] [[ID=*9]] Japanese Patent No. 6943187

[0004] By the way, in a technique for processing data including only feature amounts, such as the technique of Patent Document 1, there is a possibility that it cannot be applied except for the assumed use cases.

[0005] The present invention has been made in view of the above problems, and an object thereof is to provide a data management device and a data management method capable of reducing the capacity of data to be stored while storing more useful data.

[0006] The data management device according to the present invention is a data management device that communicates with a vehicle via a wireless communication line and manages a plurality of data acquired by sensors mounted on the vehicle, comprising: a data acquisition unit that collects a plurality of data acquired by sensors via a wireless communication line; a feature calculation unit that calculates the feature quantities of each of the plurality of data collected by the data acquisition unit; a similarity calculation unit that calculates the similarity of each of the plurality of data based on the feature quantities of each of the plurality of data calculated by the feature calculation unit; a grouping processing unit that classifies each of the plurality of data into a plurality of groups based on the similarity of each of the plurality of data calculated by the similarity calculation unit; a duplication calculation unit that calculates the duplication of data belonging to each of the plurality of groups classified by the grouping processing unit; a storage unit that stores the plurality of data for which the duplication has been calculated by the duplication calculation unit; and a deletion determination unit that, based on the duplication calculated by the duplication calculation unit, designates the data stored in the storage unit as deletion candidates in order of the data with the highest duplication in each of the plurality of groups, wherein the storage unit deletes the data designated as deletion candidates by the deletion determination unit.

[0007] According to the present invention, it is possible to reduce the amount of data to be stored while storing more useful data. Further features related to the present invention will become apparent from the description herein and the accompanying drawings. In addition, problems, configurations, and effects other than those described above will become apparent from the description of the following embodiments.

[0008] A diagram showing a data management system according to the first embodiment. A diagram showing a data management device according to the first embodiment. A diagram showing an example of data related to images. A diagram showing an example of data other than that related to images. A diagram showing a method for calculating feature quantities for data related to images and time-series numerical data that is not related to images. A diagram showing an example of each feature quantity of the data. A diagram showing a detailed example of each feature quantity of the data. A diagram showing an example of the similarity of each data. A diagram showing an example of each classified group of data. A diagram showing an example of the degree of overlap of data belonging to each group. A diagram showing an example of each classified group of data and the presence or absence of a deletion flag. A flowchart showing the overall operation of the data management device according to the first embodiment. A flowchart showing the operation of feature quantity calculation, similarity calculation, and grouping processing in Figure 12. A diagram showing a data management device according to the second embodiment.

[0009] The following describes embodiments with reference to the attached drawings. In the attached drawings, functionally identical elements are indicated by the same numbers. The attached drawings show embodiments in accordance with the principles of this disclosure, but they are for the purpose of understanding this disclosure and are not to be used in any way to restrict the interpretation of this disclosure. The descriptions in this specification are typical examples and do not limit the claims or applications of this disclosure in any way.

[0010] [First Embodiment] The first embodiment will be described below. As shown in Figure 1, the data management system 200 of this embodiment comprises a data management device 1A and a plurality of vehicles 100A, 100B, and 100C. The data management system 200 of this embodiment acquires and stores data (images, CAN signals, and time-series numerical data from sensors other than images) acquired by sensors while the vehicles 100A, 100B, and 100C are in motion, for the development of autonomous driving (AD) and advanced driver-assistance systems (ADAS).

[0011] Vehicle 100A is equipped with sensors 101A, 102A, and 103A. Vehicle 100B is equipped with sensors 101B, 102B, and 103B. Vehicle 100C is equipped with sensors 101C, 102C, and 103C. Sensors 101A, 101B, and 101C are, for example, forward cameras that photograph the area in front of vehicles 100A, 100B, and 100C. Sensors 102A, 102B, and 102C are, for example, vehicle speed sensors that detect the vehicle speed of vehicles 100A, 100B, and 100C. Sensors 103A, 103B, and 103C are, for example, rear cameras that photograph the area behind vehicles 100A, 100B, and 100C. Vehicles 100A, 100B, and 100C are equipped with sensors (not shown) that detect the position, acceleration, deceleration, and steering angle of vehicles 100A, 100B, and 100C.

[0012] The data management device 1A communicates with vehicles 100A, 100B, and 100C via a wireless communication line and manages multiple data acquired by sensors 101A to 103C mounted on vehicles 100A, 100B, and 100C. The data management device 1A also communicates with the dealer 300 via a wireless communication line and acquires and stores data related to the inspection and repair history of vehicles 100A, 100B, and 100C. The data management device 1A also communicates with an organization that monitors traffic conditions (not shown) via a wireless communication line and acquires and stores data related to the traffic conditions around vehicles 100A, 100B, and 100C. The data management device 1A also communicates with an organization that observes weather conditions (not shown) via a wireless communication line and acquires and stores data related to the weather around vehicles 100A, 100B, and 100C.

[0013] As shown in Figure 2, the data management device 1A comprises a data collection unit 11, a feature calculation unit 12, a similarity calculation unit 13, a grouping processing unit 14, a duplication calculation unit 15, a storage unit 16, a deletion determination unit 17, and a control input receiving unit 18. The data management device 1A is configured as a cloud server that vehicles 100A, 100B, and 100C can access via a wireless communication line. The data management device 1A's hardware configuration includes a processor, RAM (Random Access Memory), ROM (Read Only Memory), and an auxiliary storage device. The processor consists of a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a DSP (Digital Signal Processor), etc.

[0014] The ROM stores a computer program capable of executing the processing of the data management device 1A described below. The computer program stored in the ROM is loaded into the RAM. The processor executes predetermined arithmetic processing according to the computer program loaded into the RAM. This executes the processing of the data management device 1A described below. The auxiliary storage device consists of an HDD (Hard Disk Drive) and an SSD (Solid State Drive). Various types of data calculated by the processor are recorded in the auxiliary storage device.

[0015] Although the data management device 1A is capable of processing as a server, the term "server" here refers to a processing device that can send and receive information with communication devices such as vehicles 100A, 100B, and 100C via communication lines, and does not refer to hardware such as personal computers or embedded information devices.

[0016] The data collection unit 11 collects multiple data acquired by sensors 101A, etc., via a wireless communication line. The data collection unit 11 also acquires data related to the inspection and repair history of vehicles 100A, 100B, and 100C. In addition, the data collection unit acquires data related to traffic conditions and weather around vehicles 100A, 100B, and 100C.

[0017] The data collected by the data acquisition unit 11 includes image-related data and time-series numerical data that is not image-related. Image-related data is data captured by sensors 101A such as a camera. Time-series numerical data that is not image-related is CAN (Controller Area Network) data, which is a collection of time-series numerical data measured by sensors 102A such as a vehicle speed sensor. As shown in Figure 3, image-related data includes, for example, a data ID (identification) for identifying the data, an image file name, a sensor file name, and the date the data was acquired. As shown in Figure 4, time-series numerical data that is not image-related includes a timestamp (date and time), the speed of vehicles 100A, 100B, and 100C, whether the accelerator is pressed or not, whether the brakes are pressed or not, and the latitude and longitude (positions of vehicles 100A, 100B, and 100C).

[0018] As shown in Figure 2, the feature calculation unit 12 calculates the features of each of the multiple data collected by the data collection unit 11. As shown in Figure 5, the method by which the feature calculation unit 12 calculates features differs between data related to images and time-series numerical data that is not related to images. For example, the feature calculation unit 12 calculates the features shown in Figure 6 for each of the multiple data.

[0019] As shown in Figure 7, the features calculated by the feature calculation unit 12 include, for example, the model number of the vehicle 100A, the model numbers of the components constituting the vehicle 100A, the model number of the sensor 101A, the version number of the software executed on the vehicle 100A, the measured values ​​of the sensor 101A, values ​​such as the inter-vehicle distance obtained by processing the measured values ​​of the sensor 101A, the caption attached to the image captured by the sensor 101A, the inspection history of the vehicle 100A, and The data acquired by sensor 101A, etc. includes the date and time, the position of vehicle 100A, the speed of vehicle 100A, the acceleration of vehicle 100A, the deceleration of vehicle 100A, the steering angle of vehicle 100A, the type of road on which vehicle 100A is traveling, the curvature of the road on which vehicle 100A is traveling, the gradient of the road on which vehicle 100A is traveling, the width of the road on which vehicle 100A is traveling, the lane of the road on which vehicle 100A is traveling, the traffic conditions around vehicle 100A, and the weather around vehicle 100A.

[0020] The features calculated by the feature calculation unit 12 include, for example, a combination of features such as the brightness of each pixel in the image data acquired by the sensor 101A, and features such as the speed of the vehicle 100A in time-series numerical data that is not image data but was acquired by the sensor 101B at the same time as the image data acquired by the sensor 101A.

[0021] The features calculated by the feature calculation unit 12 include, for example, a combination of features such as the brightness of each pixel in the image data acquired by sensor 101A, etc., and features such as the brightness of each pixel in the image data acquired by sensors 101B and 101C at the same time as the image data acquired by sensor 101A, etc.

[0022] This includes a combination of feature quantities such as the speed of the vehicle 100A in time-series numerical data that is not image-related data acquired by sensor 101B, etc., and feature quantities such as the acceleration, deceleration, and steering angle of the vehicle 100A in time-series numerical data that is not image-related data acquired by sensors other than sensor 101B.

[0023] As shown in Figure 2, the feature calculation unit 12 includes generative artificial intelligence 21. The feature calculation unit 12 uses the generative artificial intelligence 21 to generate features for each of the multiple data sets collected by the data collection unit 11. For example, the generative artificial intelligence 21 generates image-specific features for image-related data. Image-specific features refer to features for recognizing objects such as vehicles, people, signs, and obstacles that need to be recognized from the image. When generating image-specific features, the generative artificial intelligence 21 may first use a multimodal large-scale language model such as GPT-4o (Generative Pre-trained Transformer 4 Omni) as the generative artificial intelligence 21 to generate text such as words, sentences, and JSON (JavaScript Object Notation) from the image according to a predetermined viewpoint, and store them as image captions in Figure 7. Next, the generative artificial intelligence 21 may use a base model for language such as BERT (Bidirectional Encoder Representations from Transformers) to calculate fixed-length embedding vectors, i.e., feature quantities in the format shown in Figure 6, from image words, sentences, and JSON text.

[0024] When multiple data points are image-related data, the feature calculation unit 12 represents each of the feature quantities of the multiple data points for each specified viewpoint as an embedding vector. Specifically, the feature quantities for object recognition are generated, for example, as follows: The feature calculation unit 12 inputs the specified viewpoint, such as vehicles, people, signs, and obstacles, into an open vocabulary object detection model, such as Grounded Segment Anything, and extracts images of the specified viewpoint, such as vehicles, people, signs, and obstacles. The feature calculation unit 12 inputs the extracted images of the specified viewpoint, such as vehicles, people, signs, and obstacles, into a base model for images, such as ViT (Vision Transformer), and considers the values ​​of the output layer of ViT as fixed-length embedding vectors, i.e., feature quantities.

[0025] When multiple data points are time-series numerical data that are not image-related data, the feature calculation unit 12 represents the features of the multiple data points as embedding vectors for each specified viewpoint. Specifically, the features of time-series numerical data that are not image-related data are generated, for example, as follows: The feature calculation unit 12 converts the CAN data, which is a collection of time-series numerical data measured by the sensor 102A, etc., from binary format to tabular data format. The feature calculation unit 12 extracts columns such as velocity, which are the specified viewpoints, from the tabular data. The feature calculation unit 12 considers the extracted columns as time-series numerical data and considers the normalized vectors as embedding vectors, i.e., features. At this time, the feature calculation unit 12 applies a movement window to the existing data according to the data length of the search query.

[0026] As shown in Figure 2, the similarity calculation unit 13 calculates the similarity of each of the multiple data based on the respective features of the multiple data calculated by the feature calculation unit 12. When the multiple data are image-related data, the similarity calculation unit 13 calculates the similarity as the cosine similarity of the respective embedding vectors of the multiple data. When the multiple data are time-series numerical data that are not image-related data, the similarity calculation unit 13 calculates the similarity for each embedding vector of the multiple data for each specified viewpoint using Dynamic Time Warping (DTW). The similarity calculation unit 13 calculates the similarity between data with an arbitrary P-side data ID and data with an arbitrary Q-side data ID, for example, as shown in Figure 8.

[0027] As shown in Figure 2, the grouping processing unit 14 classifies each of the multiple data into multiple groups based on the similarity of each data calculated by the similarity calculation unit 13. When the multiple data are image data, the grouping processing unit 14 classifies data with a similarity threshold x such that the cosine similarity is similarity threshold x ≤ cosine similarity ≤ 1 into the same group. The similarity threshold x can be arbitrarily determined, for example, in the range of 0 ≤ similarity threshold x ≤ 1. In addition to cosine similarity, any distance measure such as Euclidean distance, Manhattan distance, or Mahalanobis distance may be used as the similarity measure. Furthermore, in addition to similarity threshold determination, groups may be classified by clustering algorithms such as k-means method, k-nearest neighbors method, or DBSCAN (Density-Based Spatial Clustering of Applications with Noise). The grouping process can also be repeated for data within a group, or hierarchical grouping can be performed using hierarchical clustering algorithms such as HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise).

[0028] When multiple data are time-series numerical data that are not image-related data, the grouping processing unit 14 classifies data whose distance indicating similarity between multiple time-series numerical data is less than or equal to the similarity threshold y, given a similarity threshold y, into the same group. The similarity threshold y can be arbitrarily determined, for example, within the range of 0 ≤ similarity threshold y. In addition to similarity threshold determination, groups may also be classified using clustering algorithms such as k-means, k-nearest neighbors, or DBSCAN (Density-Based Spatial Clustering of Applications with Noise). The grouping process can be repeated for data within a group, or hierarchical grouping can be performed using a hierarchical clustering algorithm such as HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise). If hierarchical grouping is not performed, the grouping processing unit 14 classifies each of the data with data IDs 1, 2, 3... into groups with group IDs 1, 1, 2..., for example, as shown in Figure 9. When hierarchical, group IDs are classified into groups such as 1-1, 1-2, 2-1, etc.

[0029] As shown in Figure 2, the overlap calculation unit 15 calculates the overlap of data belonging to each of the multiple groups classified by the grouping processing unit 14. The overlap calculation unit 15 calculates the overlap based on the number of data belonging to each of the multiple groups. Specifically, the overlap calculation unit 15 can use the number of data belonging to each of the multiple groups as the overlap. For example, as shown in Figure 10, the overlap calculation unit 15 calculates the overlap as the number of data belonging to each of the multiple groups with group IDs 1, 2, ...

[0030] As shown in Figure 2, the storage unit 16 stores multiple data whose degree of redundancy has been calculated by the redundancy calculation unit 15. The storage unit 16 includes a driving data storage unit 31, a feature quantity storage unit 32, a group information storage unit 33, and a redundancy storage unit 34. The driving data storage unit 31 stores data related to images, as shown in Figure 3, and time-series numerical data that is not related to images, as shown in Figure 4.

[0031] The feature memory unit 32 stores the feature quantities of the data as shown in Figures 6 and 7. The group information memory unit 33 stores information about the group to which each of the multiple data sets is classified, as shown in Figure 9. The redundancy memory unit 34 stores the redundancy of the data belonging to each of the multiple groups, as shown in Figure 10. The memory unit 16 performs processing to reduce the capacity of the stored data as appropriate, using methods such as compression, thumbnailing, frame rate reduction, and saving to a tape drive, etc.

[0032] As shown in Figure 2, the deletion determination unit 17, based on the degree of duplication calculated by the duplication calculation unit 15, designates data stored in the storage unit 16 as deletion candidates in order of the degree of duplication in each of the multiple groups. The deletion determination unit 17 then designates data stored in the storage unit 16 as deletion candidates in order of the degree of duplication in each of the multiple groups, and in order of the size of the data stored in the storage unit 16 after the capacity reduction process, and assigns a deletion flag to the data. As shown in Figure 11, the deletion determination unit 17 assigns a deletion flag to the data designated as deletion candidates. Information regarding the deletion flag is stored, for example, in the group information storage unit 33 with the deletion flag assigned. The storage unit 16 shown in Figure 2 deletes the data designated as deletion candidates by the deletion determination unit 17 and assigned a deletion flag.

[0033] The storage unit 16 does not necessarily need to store data for which the degree of redundancy has been calculated from the beginning. For example, the storage unit 16 may provisionally store the data collected by the data collection unit 11 in the driving data storage unit 31. The feature calculation unit 12 may calculate the features of the data stored in the driving data storage unit 31 and store them in the feature calculation unit 32. The similarity calculation unit 13 may calculate the similarity from the data stored in the driving data storage unit 31 and the features stored in the feature calculation unit 32.

[0034] The grouping processing unit 14 may classify each of the multiple data into multiple groups based on the similarity of each data calculated by the similarity calculation unit 13, and store information about the groups in the group information storage unit 33. The duplication calculation unit 15 may calculate the degree of duplication of the data belonging to each of the multiple groups stored in the group information storage unit 33 and store it in the duplication storage unit 34. The deletion determination unit 17 may, based on the duplication degrees stored in the duplication storage unit 34, select the data stored in the storage unit 16 as deletion candidates in order of the degree of duplication in each of the multiple groups, and attach a deletion flag to the information in the group information storage unit 33.

[0035] Furthermore, the storage unit 16 does not necessarily delete data that has been flagged for deletion immediately. The storage unit 16 deletes data that has been designated as a candidate for deletion by the deletion determination unit 17 if either the remaining capacity of the storage unit 16 falls below a predetermined remaining capacity threshold or the time the data has been stored in the storage unit 16 exceeds a predetermined storage time threshold. The remaining capacity threshold can be set, for example, to 5 to 20% of the total capacity of the storage unit 16. The storage time threshold can be set, for example, to several days to several months. The storage unit 16 can delete some or all of the data classified into a group. The storage unit 16 may also determine the number of data to delete from a group according to the remaining capacity of the storage unit 16 and the time the data has been stored in the storage unit.

[0036] Furthermore, as shown in Figure 2, the control input receiving unit 18 receives control input to the storage unit 16. Based on the control input received by the control input receiving unit 18, the storage unit 16 cancels the deletion of the data designated as a deletion candidate if either the remaining capacity of the storage unit 16 falls below the remaining capacity threshold or the time the data has been stored in the storage unit 16 exceeds the storage time threshold.

[0037] The operation of the data management device 1A in this embodiment will now be described. As shown in Figure 12, the data acquisition unit 11 performs the step of collecting multiple data acquired by the sensor 101A, etc. (S101). The feature calculation unit 12 performs the step of calculating the feature of each of the data collected by the data acquisition unit 11 (S102). The similarity calculation unit 13 performs the step of calculating the similarity of each of the multiple data based on the feature of each of the multiple data calculated by the feature calculation unit 12 (S103). The grouping processing unit 14 performs the step of classifying each of the multiple data into multiple groups based on the similarity of each of the multiple data calculated by the similarity calculation unit 13 (S104).

[0038] The duplication calculation unit 15 performs the step of calculating the degree of duplication of data belonging to each of the multiple groups classified by the grouping processing unit 14 (S105). The storage unit 16 performs the step of storing the multiple data for which the degree of duplication has been calculated by the duplication calculation unit 15 (S106). The storage unit 16 performs the step of reducing the capacity of the data stored in the storage unit 16 (S107). The deletion determination unit 17 performs the step of selecting the stored data as deletion candidates in order of the degree of duplication with the highest degree of duplication for each of the multiple groups, based on the degree of duplication calculated by the duplication calculation unit 15, and assigning a deletion flag to it (S108).

[0039] The storage unit 16 determines whether the remaining capacity of the storage unit 16 is below the remaining capacity threshold, or whether the time the data was stored in the storage unit 16 exceeds the storage time threshold (S109). If the remaining capacity of the storage unit 16 is not below the remaining capacity threshold, and the time the data was stored in the storage unit 16 does not exceed the storage time threshold (S109), the data management device 1A terminates processing.

[0040] If the remaining capacity of the storage unit 16 is below the remaining capacity threshold, or if the time the data has been stored in the storage unit 16 exceeds the storage time threshold (S109), the storage unit 16 determines whether the control input received by the control input receiving unit 18 is an input to cancel the deletion (S110). If the control input received by the control input receiving unit 18 is not an input to cancel the deletion (S110), the storage unit 16 executes the process of deleting the data that was designated as a deletion candidate (S111). If the control input received by the control input receiving unit 18 is an input to cancel the deletion (S110), the storage unit 16 cancels the deletion of the data that was designated as a deletion candidate and removes the deletion flag from the data (S112).

[0041] The following describes in detail the process of calculating features, calculating similarity, and classifying into groups. As shown in Figure 13, the feature calculation unit 12 determines whether the data is related to an image (S201).

[0042] If the data is image data (S201), the feature calculation unit 12 represents each of the feature quantities of multiple data for each specified viewpoint as an embedding vector (S202). The feature calculation unit 12 inputs the specified viewpoint, such as vehicles, people, signs, and obstacles, into Grounded Segment Anything, an open vocabulary object detection model, and extracts images of the specified viewpoint, such as vehicles, people, signs, and obstacles. The feature calculation unit 12 inputs the extracted images of the specified viewpoint, such as vehicles, people, signs, and obstacles, into ViT, an image-oriented base model, and considers the values ​​of the output layer of ViT as fixed-length embedding vectors, i.e., feature quantities.

[0043] The similarity calculation unit 13 calculates the similarity as the cosine similarity of the embedding vectors of each of the plurality of data (S203). The grouping processing unit 14 determines whether the similarity as the cosine similarity satisfies the similarity threshold x ≤ cosine similarity ≤ 1 when the similarity threshold is x (S204). When x ≤ cosine similarity ≤ 1 (S204), the grouping processing unit 14 classifies the data into the same group (S205). When x ≤ cosine similarity ≤ 1 does not hold (S204), the grouping processing unit 14 classifies the data into different groups (S206).

[0044] In S201, when the data is time-series numerical data rather than data related to an image, the feature quantity calculation unit 12 represents the feature quantities of the plurality of data for each specified viewpoint as embedding vectors (S207). The feature quantity calculation unit 12 converts CAN data, which is a set of time-series numerical data measured by the sensor 102A or the like, from binary format to table data format. The feature quantity calculation unit 12 extracts columns such as speed, which is the specified viewpoint, from the table data. The feature quantity calculation unit 12 regards the extracted column as time-series numerical data and regards the normalized vector as an embedding vector, that is, a feature quantity. At this time, the feature quantity calculation unit 12 applies a moving window according to the data length of the search query to the existing data.

[0045] The similarity calculation unit 13 calculates the similarity for the embedding vectors of each of the plurality of data for each specified viewpoint by the dynamic time warping method (S208). It is determined whether the distance indicating the similarity of each of the plurality of data, which is time-series numerical data, is less than or equal to the similarity threshold y (S209). When the distance indicating the similarity of each of the plurality of data is less than or equal to the similarity threshold y (S209), the grouping processing unit 14 classifies the data into the same group (S210). When the distance indicating the similarity of each of the plurality of data is not less than or equal to the similarity threshold y (S209), the grouping processing unit 14 classifies the data into different groups (S211).

[0046] According to this embodiment, the feature amount of each of the plurality of collected data is calculated, the similarity of each of the plurality of data is calculated based on the feature amount of each of the calculated plurality of data, and each of the plurality of data is classified into a plurality of groups based on the calculated similarity of each of the plurality of data. The duplication degree of the data belonging to each of the classified plurality of groups is calculated, and based on the calculated duplication degree, the stored data is set as a deletion candidate in the order of the data with a high duplication degree for each of the plurality of groups, and the data set as the deletion candidate is deleted. Since the data with a high duplication degree can be easily recollected, there is no problem even if it is deleted. Therefore, according to this embodiment, it is possible to reduce the capacity of the data to be stored while storing data with a low duplication degree and more useful data.

[0047] Further, according to this embodiment, the feature amount calculation unit 12 generates, by the generative artificial intelligence 21, the feature amount to be calculated for each of the plurality of data collected by the data collection unit 11. Therefore, the feature amount can be dynamically changed according to the collected data, and more useful data can be stored.

[0048] Further, according to this embodiment, since the feature amount calculated by the feature amount calculation unit 12 varies from the model number of the vehicle 100A to the weather around the vehicle 100A, it is possible to store useful data according to a more detailed situation.

[0049] Furthermore, according to this embodiment, when multiple data are image data, each of the feature quantities of the multiple data for each specified viewpoint is represented as an embedding vector, and for each embedding vector of the multiple data, the similarity is calculated as one of the following: cosine similarity, Euclidean distance, Manhattan distance, or Mahalanobis distance. Based on the similarity, each of the multiple data is classified into multiple groups by one of the following methods: similarity thresholding, k-means method, k-nearest neighbors method, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), or HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise). For example, according to this embodiment, when multiple data are image data, each of the feature quantities of the multiple data for each specified viewpoint is represented as an embedding vector, the similarity is calculated as the cosine similarity of each embedding vector of the multiple data, and data with a similarity of cosine similarity of similarity threshold x ≤ cosine similarity ≤ 1 are classified into the same group, thus allowing image data to be classified into groups using a simple method.

[0050] Furthermore, when multiple data are time-series numerical data that are not image-related data, the feature quantities of the multiple data are represented as embedding vectors for each specified viewpoint, and the similarity is calculated for each embedding vector of the multiple data using dynamic time stretching. Based on the similarity, each of the multiple data is classified into multiple groups using one of the following methods: similarity thresholding, k-means method, k-nearest neighbor method, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), or HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise). For example, according to this embodiment, when multiple data are time-series numerical data that are not image-related data, the feature quantities of the multiple data are represented as embedding vectors for each specified viewpoint, and the similarity is calculated for each embedding vector of the multiple data for each specified viewpoint using dynamic time stretching. Since data whose distance from each of the multiple time-series numerical data is less than or equal to a similarity threshold y is classified into the same group, time-series numerical data that are not image-related data can be classified into groups using a simple method.

[0051] Furthermore, according to this embodiment, the overlap calculation unit 15 calculates the overlap based on the number of data points belonging to each of the multiple groups, so the overlap can be calculated using a simple method.

[0052] Furthermore, according to this embodiment, the deletion determination unit 17 selects data stored in the storage unit 16 as deletion candidates in order of the degree of duplication within each of the multiple groups, and in order of the size of the data stored in the storage unit 16 after the capacity reduction process. Therefore, even after the capacity reduction process, data with a large size can be selected as deletion candidates, and the amount of data to be stored can be reduced efficiently.

[0053] Furthermore, according to this embodiment, the storage unit 16 deletes data designated as a deletion candidate by the deletion determination unit 17 when either the remaining capacity of the storage unit 16 falls below the remaining capacity threshold or the time the data has been stored in the storage unit 16 exceeds the storage time threshold. Therefore, by deleting data designated as a deletion candidate when there is insufficient remaining capacity or when the data is old, it is possible to reduce the amount of data to be stored while storing more useful data.

[0054] Furthermore, according to this embodiment, the storage unit 16 cancels the deletion of data designated as a deletion candidate based on the control input received by the control input receiving unit 18, and can retain the deletion candidate data as necessary.

[0055] [Second Embodiment] A second embodiment will be described below. As shown in Figure 14, the data management device 1B of this embodiment is mounted on a vehicle 100A and manages multiple data acquired by sensors 101A, 102A, and 103A mounted on the vehicle 100A. Thus, in this embodiment, even though the data management device 1B is mounted on a vehicle, it can perform the same functions as the data management device 1A on the cloud server described above for each individual vehicle 100A.

[0056] Although several embodiments of the present invention have been described above, the present invention is not limited to the embodiments described above and can be realized in various configurations without departing from its spirit. For example, configurations that arbitrarily combine the configurations of Embodiments 1 and 2 can easily be conceivable. These variations are included in the scope of the invention and its equivalents as described in the claims.

[0057] 1A, 1B Data Management Device 11 Data Acquisition Unit 12 Feature Calculation Unit 13 Similarity Calculation Unit 14 Grouping Processing Unit 15 Repetition Calculation Unit 16 Storage Unit 17 Deletion Judgment Unit 18 Control Input Reception Unit 21 Generative Artificial Intelligence 31 Driving Data Storage Unit 32 Feature Storage Unit 33 Group Information Storage Unit 34 Repetition Storage Unit 100A, 100B, 100C Vehicle 101A, 102A, 103A, 101B, 102B, 103B, 101C, 102C, 103C Sensor 200 Data Management System 300 Dealer

Claims

1. A data management device that communicates with a vehicle via a wireless communication line and manages a plurality of data acquired by sensors mounted on the vehicle, comprising: a data acquisition unit that collects the plurality of data acquired by the sensors via a wireless communication line; a feature calculation unit that calculates the feature quantities of each of the plurality of data collected by the data acquisition unit; a similarity calculation unit that calculates the similarity of each of the plurality of data based on the feature quantities of each of the plurality of data calculated by the feature calculation unit; a grouping processing unit that classifies each of the plurality of data into a plurality of groups based on the similarity of each of the plurality of data calculated by the similarity calculation unit; a duplication calculation unit that calculates the duplication of the data belonging to each of the plurality of groups classified by the grouping processing unit; a storage unit that stores the plurality of data for which the duplication has been calculated by the duplication calculation unit; and a deletion determination unit that, based on the duplication calculated by the duplication calculation unit, designates the data stored in the storage unit as deletion candidates in order of the data with the highest duplication in each of the plurality of groups, wherein the storage unit deletes the data designated as deletion candidates by the deletion determination unit.

2. The data management device according to claim 1, characterized in that the feature calculation unit generates the feature quantities to be calculated for each of the multiple data sets collected by the data collection unit using generative artificial intelligence.

3. The data management device according to claim 1, characterized in that the feature quantities calculated by the feature quantity calculation unit are any of the following: the model number of the vehicle, the model numbers of the parts constituting the vehicle, the model number of the sensor, the version number of the software executed in the vehicle, the measured value of the sensor, the value obtained by processing the measured value of the sensor, the caption attached to the image captured by the sensor, the inspection history of the vehicle, the date and time when the data was acquired by the sensor, the position of the vehicle, the speed of the vehicle, the acceleration of the vehicle, the deceleration of the vehicle, the steering angle of the vehicle, the type of road on which the vehicle travels, the curvature of the road on which the vehicle travels, the gradient of the road on which the vehicle travels, the width of the road on which the vehicle travels, the lane of the road on which the vehicle travels, the traffic conditions around the vehicle, and the weather around the vehicle.

4. When the multiple data are image data, the feature calculation unit represents each of the feature quantities of the multiple data for each specified viewpoint as an embedding vector; the similarity calculation unit calculates the similarity for each of the embedding vectors of the multiple data as one of cosine similarity, Euclidean distance, Manhattan distance, or Mahalanobis distance; and the grouping processing unit classifies each of the multiple data into multiple groups based on the similarity by one of similarity thresholding, k-means method, k-nearest neighbor method, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), or HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise).

5. When the multiple data are time-series numerical data that are not image data, the feature calculation unit represents the features of the multiple data as embedding vectors for each specified viewpoint, the similarity calculation unit calculates the similarity for each of the embedding vectors of the multiple data using a dynamic time stretching method, and the grouping processing unit classifies each of the multiple data into multiple groups based on the similarity by any of the following methods: threshold determination of similarity, k-means method, k-nearest neighbor method, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), and HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise).

6. The data management device according to claim 1, characterized in that the overlap calculation unit calculates the overlap based on the number of data items belonging to each of the multiple groups.

7. The data management device according to claim 1, characterized in that the deletion determination unit selects the data stored in the storage unit as deletion candidates in order of the order of the data with the highest degree of duplication for each of the plurality of groups, and in order of the amount of data stored in the storage unit after processing to reduce capacity.

8. The data management device according to claim 1, wherein the storage unit deletes the data designated as a candidate for deletion by the deletion determination unit when either the remaining capacity of the storage unit falls below a remaining capacity threshold or the time the data was stored in the storage unit exceeds a storage time threshold.

9. The data management device according to claim 8, further comprising a control input receiving unit for receiving control input to the storage unit, wherein the storage unit cancels the deletion of the data designated as a deletion candidate when the remaining capacity of the storage unit falls below a remaining capacity threshold or when the time the data has been stored in the storage unit exceeds a storage time threshold, based on the control input received by the control input receiving unit.

10. A data management device mounted on a vehicle and performing the management of a plurality of data acquired by a sensor mounted on the vehicle, comprising: a data acquisition unit for collecting the plurality of data acquired by the sensor; a feature calculation unit for calculating the feature quantities of each of the plurality of data collected by the data acquisition unit; a similarity calculation unit for calculating the similarity of each of the plurality of data based on the feature quantities of each of the plurality of data calculated by the feature calculation unit; a grouping processing unit for classifying each of the plurality of data into a plurality of groups based on the similarity of each of the plurality of data calculated by the similarity calculation unit; a duplication calculation unit for calculating the duplication of the data belonging to each of the plurality of groups classified by the grouping processing unit; a storage unit for storing the plurality of data for which the duplication has been calculated by the duplication calculation unit; and a deletion determination unit for designating the data stored in the storage unit as deletion candidates in order of the highest duplication in each of the plurality of groups based on the duplication calculated by the duplication calculation unit, wherein the storage unit deletes the data designated as deletion candidates by the deletion determination unit.

11. A data management method comprising: a step of collecting multiple data acquired by a sensor; a step of calculating the feature quantity of each of the collected multiple data; a step of calculating the similarity of each of the multiple data based on the calculated feature quantity of each of the multiple data; a step of classifying each of the multiple data into multiple groups based on the calculated similarity of each of the multiple data; a step of calculating the degree of overlap of the data belonging to each of the classified multiple groups; a step of storing the multiple data for which the degree of overlap has been calculated; a step of designating the stored data as candidates for deletion in order of the degree of overlap in each of the multiple groups based on the calculated degree of overlap; and a step of deleting the data designated as candidates for deletion.