Data analysis device, driving assistance system, data analysis method, and data analysis program

JPWO2025041218A5Inactive Publication Date: 2025-07-30
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023580605
Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2023-12-27
Publication Date
2025-07-30
Estimated Expiration
Not applicable · inactive patent
Patent Text Reader

Abstract

A data analysis device (200) included in a driving assistance system (90) comprises a comparison unit (240). When the classification of each of a first driver and a second driver is estimated to be a target classification, the comparison unit (240) determines, on the basis of a first unused data feature amount, which corresponds to both the first driver and an unused item, and a second unused data feature amount which corresponds to both the second driver and an unused item, whether the unused data items have the ability to classify the first driver and the second driver into mutually different classifications. The unused data items are not used when estimating the classification of each of the first driver and the second driver, and correspond to a combination of a target driving behavior and a target sensor provided to respective automobiles.
Need to check novelty before this filing date? Find Prior Art

Description

Data analysis device, driving assistance system, data analysis method, and data analysis program

[0001] The present disclosure relates to a data analysis device, a driving assistance system, a data analysis method, and a data analysis program.

[0002] In the field of automobile driving assistance technology, there is a method for estimating the classification of a driver's driving characteristics using data acquired by sensors mounted on the automobile. Cited Document 1 discloses, as a specific example, a technology for estimating the classification of a driver's driving characteristics using the timing at which the accelerator is released before an intersection.

[0003] JP 2011-096061 A

[0004] In data-driven development, which develops classification inference rules (algorithms) based on data, there is a method in which, in addition to experimental data acquired during development, data (in-operation data) is collected while the user is operating the car, such as when the user drives the car on a daily basis, and the collected in-operation data is used to continuously improve the classification inference rules. This method has the following challenges.

[0005] As a specific example, driving characteristics are thought to differ depending on the region or country. Therefore, it is necessary to adjust the classification estimation rules to address the decrease in classification accuracy in locations other than the region or country where the experimental data was acquired. In this case, it is necessary to analyze multiple sensor data for each condition (driving straight, turning, passing through a specific intersection, etc.). Therefore, the burden of understanding the difference between the classification ability of experimental data and the classification ability of operational data is very high. Therefore, a method to support the analysis of the classification ability of operational data is required.

[0006] Furthermore, storing all operational data requires a huge amount of storage capacity. Therefore, methods for reducing the amount of data are being implemented, such as by storing statistical quantities such as the mean and variance. However, storing only statistical quantities such as the mean and variance does not preserve the distribution of the original operational data. Therefore, data analysts may not be able to properly grasp the shape of the stored operational data. Therefore, there is a need for innovative methods for reducing the amount of data and storing operational data.

[0007] The present disclosure aims to support analysis of the classification capabilities of operational data while reducing the amount of data and preserving operational data in a manner that allows data analysts to understand the shape of the operational data.

[0008] The data analysis device according to the present disclosure is a data analysis device provided in a driving assistance system that estimates the classification of each driver who drove each vehicle based on time series data consisting of data acquired by sensors equipped in each vehicle and determines driving assistance content based on the estimated classification, and the data analysis device includes: a comparison unit that, when the classification of each of a first driver and a second driver is estimated to be a target classification, and when an item not used when estimating the classification of each of the first driver and the second driver is defined as an unused item, determines whether the unused item has the ability to classify the first driver and the second driver into different classifications based on features of first unused data, which is time series data corresponding to both the first driver and the unused item, and features of second unused data, which is time series data corresponding to both the second driver and the unused item, and the unused item is an item that corresponds to a combination of a target sensor equipped in each vehicle and a target driving behavior, which is the driving behavior while driving each vehicle.

[0009] According to the present disclosure, the comparison unit determines whether or not classification ability exists for unused items based on feature quantities of time-series data corresponding to unused items. Here, the time-series data corresponding to unused items may be operational data. Therefore, the data analysis device can support analysis of classification ability of operational data. Furthermore, since it is sufficient for supporting analysis of classification ability to store feature quantities, the operational data may be stored with a reduced data volume. Furthermore, the feature quantities may be useful for understanding the shape of the operational data. Therefore, according to the present disclosure, analysis of classification ability of operational data can be supported while reducing the data volume and storing the operational data in a manner that allows a data analyst to understand the shape of the operational data.

[0010] FIG. 1 is a diagram showing an example of the configuration of a driving assistance system 90 according to the first embodiment. FIG. 2 is a diagram showing a specific example of data stored in a data storage unit 290 according to the first embodiment. FIG. 3 is a diagram explaining an overview of the operation of the driving assistance system 90 according to the first embodiment. FIG. 4 is a diagram explaining the processing of a driving characteristics determination unit 220 according to the first embodiment. FIG. 5 is a diagram explaining the processing of a comparison unit 240 according to the first embodiment. FIG. 6 is a diagram explaining the processing of a comparison unit 240 according to the first embodiment. FIG. 7 is a diagram explaining the processing of a comparison unit 240 according to the first embodiment. FIG. 8 is a diagram showing an example of the hardware configuration of a data analysis apparatus 200 according to the first embodiment.

[0011] In the description of the embodiments and the drawings, the same elements and corresponding elements are given the same reference numerals. The description of elements given the same reference numerals will be omitted or simplified as appropriate. Arrows in the drawings mainly indicate the flow of data or the flow of processing. Furthermore, "unit" may be read as "circuit," "step," "procedure," "process," or "circuitry" as appropriate.

[0012] First Embodiment Hereinafter, the present embodiment will be described in detail with reference to the drawings.

[0013] ***Description of Configuration*** FIG. 1 shows an example configuration of a driving assistance system 90 according to the first embodiment. As shown in FIG. 1, the driving assistance system 90 includes an automobile 100 and a data analysis device 200. The driving assistance system 90 may also include a plurality of other automobiles 100. The driving assistance system 90 estimates the classification of each driver who drove each automobile 100 based on time-series data made up of data acquired by sensors provided in each automobile 100, and determines the driving assistance content based on the estimated classification. Each automobile 100 is also referred to as a target automobile. The automobile 100 and the data analysis device 200 are communicatively connected via an exterior communication channel 20. A specific example of the exterior communication channel 20 is the Internet.

[0014] The automobile 100 includes a driving operation detection unit 110, a driver information acquisition unit 120, an on-board control device 130, a driving assistance unit 140, and a temporary data storage unit 190. The other configurations of the automobile 100 are the same as the configuration of the automobile 100.

[0015] The driving operation detection unit 110 includes sensors such as an accelerator sensor 111, a brake sensor 112, a vehicle speed sensor 113, an acceleration sensor 114, a steering angle sensor 115, and a GPS (Global Positioning System) sensor 116. The driving operation detection unit 110 detects the driving operation of the driver of the automobile 100 using the sensors included in the driving operation detection unit 110.

[0016] The driver information acquisition unit 120 includes an in-vehicle camera 121. The in-vehicle camera 121 is a camera mounted on the automobile 100.

[0017] The on-board control device 130 includes a driving behavior determination unit 131, a feature calculation unit 132, and a communication function unit 133. The data acquired by the on-board control device 130 from the driving operation detection unit 110 and the driver information acquisition unit 120 is typically time-series data, and may be sensor data, on-board sensor data, or sensor signal data. The driving behavior determination unit 131 determines the driver's driving behavior based on the data acquired by each sensor. A rule-based automatic determination method for driving behavior based on values ​​such as vehicle acceleration has been studied in [Reference 1] and elsewhere. The driving behavior determination unit 131 determines the driving behavior using, as a specific example, the method described in "2.1 Development of a description method for driving behavior (classification and definition of driving behavior)" in "1.3. Driving behavior database" of "Part II: Project research and development results" in [Reference 1], and assigns a label based on the determination result. The feature calculation unit 132 calculates feature values ​​corresponding to the data acquired by each sensor. The communication function unit 133 has a function of communicating with the outside.

[0018] [Reference 1] "Human Behavior-Compatible Living Environment Creation System Technology Project," [online], March 2004, Human Life Engineering Research Center, [searched July 11, 2023], Internet <URL: https: / / www.hql.jp / database / wp-content / uploads / kodo_pro1999-2003.pdf>

[0019] The driving assistance unit 140 performs driving assistance for the automobile 100 and the driver based on the output of the on-board control device 130 and the data received from the data analysis device 200 .

[0020] The temporary data storage unit 190 is a storage area for temporarily storing various types of data.

[0021] The data analysis device 200 includes a communication function unit 210, a driving characteristic determination unit 220, a driving characteristic utilization unit 230, a comparison unit 240, a data storage unit 290, and a driving pattern DB (Database) 291. The data analysis device 200 is specifically realized by a server system.

[0022] The communication function unit 210 has a function of communicating with the outside.

[0023] The driving characteristics determination unit 220 is a determiner that determines the driving characteristics of the driver based on the data stored in the data storage unit 290 .

[0024] The driving characteristic utilization unit 230 includes a driving assistance determination unit 231 and a driving characteristic notification unit 232. The driving assistance determination unit 231 determines a driving assistance technique based on the driving characteristics determined by the driving characteristic determination unit 220. The driving characteristic notification unit 232 notifies the analysis result user of the analysis result of the driving characteristic determination unit 220. The analysis result user is a user who uses the analysis result of the driving characteristics, and specific examples include an insurance company or a bus management company.

[0025] The comparison unit 240 corresponds to a comparator, compares the data stored in the data storage unit 290, and outputs data indicating the comparison result to a data analyst. As a specific example, the comparison unit 240 determines whether multiple drivers classified in the same category can be classified into different categories based on the data stored in the data storage unit 290. The comparison unit 240 is also called a feature selection support unit. That is, when the categories of the first driver and the second driver are estimated to be target categories, the comparison unit 240 determines whether the unused items have the ability to classify the first driver and the second driver into different categories based on the feature amounts of the first unused data and the feature amounts of the second unused data. The target category indicates a certain category. The first unused data is time-series data corresponding to both the first driver and the unused items. The second unused data is time-series data corresponding to both the second driver and the unused items. The unused items are items that are not used when estimating the classification of each of the first driver and the second driver, and correspond to a combination of a target sensor provided in each vehicle 100 and a target driving behavior, which is a driving behavior performed while each vehicle 100 is in operation. The unused items indicate a combination of a type of driving behavior and a type of sensor signal. The feature amount of the first unused data may include information indicating a distribution characteristic of the first unused data. The feature amount of the second unused data may include information indicating a distribution characteristic of the second unused data. The comparison unit 240 may determine whether the unused items are capable of classifying the first driver and the second driver into different classifications based on an overlapping section between the distribution of the first unused data and the distribution of the second unused data. The comparison unit 240 may determine whether the unused items are capable of classifying the first driver and the second driver into different classifications based on an overlapping section between the first quartile to the third quartile in the first unused data and the first quartile to the third quartile in the second unused data. The feature amount of the first unused data and the feature amount of the second unused data may each include information used to draw a box plot. The comparison unit 240 may output a result of determining whether the unused items have the ability to classify the first driver and the second driver into different categories.The feature quantities of the unused data correspond to the feature components of the time-series data acquired from the sensor. A specific example of the feature components is a statistical quantity.

[0026] The data storage unit 290 stores data acquired from the automobile 100. The data stored in the data storage unit 290 is data corresponding to sensor signals. Fig. 2 shows a specific example of data stored in the data storage unit 290. As shown in Fig. 2, the data storage unit 290 stores, for each driver, data indicating each driver, the date and time when the automobile 100 was driven, a label indicating each driver's driving behavior, and feature amounts of data acquired by each sensor equipped in the automobile 100.

[0027] The driving pattern DB 291 stores data indicating each driving pattern, which is data for determining driving characteristics. The driving pattern DB 291 may store data indicating data used by the driving characteristic determination unit 220. The driving pattern DB 291 may also store data input by a data analyst.

[0028] FIG. 3 is a diagram illustrating an overview of the operation of the driving assistance system 90. The operation of the driving assistance system 90 will be described using FIG. 3. First, the driving characteristic determination unit 220 classifies each of the drivers A and B based on some of the data stored in the data storage unit 290, which is data corresponding to the driving characteristics of each of the drivers A and B. As a result, each of the drivers A and B is classified into a safe driving cluster. A specific example of the data used is data corresponding to the timing of releasing the accelerator before an intersection. Note that since the driving characteristic classifications are the same, the driving assistance for the drivers A and B will be the same. Next, the comparison unit 240 confirms that the driving characteristic determination unit 220 has classified the drivers A and B into the same classification. Thereafter, the comparison unit 240 determines whether or not it is possible to classify the drivers A and B into different classifications by using data stored in the data storage unit 290 that the driving characteristic determination unit 220 did not use when classifying the drivers A and B. A specific example of the data used is data indicating the acceleration or speed of the automobile 100 or the driver's line of sight. Next, the comparison unit 240 presents to the data analyst data indicating the determination result of the driving characteristics determination unit 220 and data indicating that the determination result of the driving characteristics determination unit 220 may be incorrect.

[0029] FIG. 4 is a diagram illustrating the processing of the driving characteristics determination unit 220. The driving characteristics determination unit 220 executes a classification estimation method. The processing of the driving characteristics determination unit 220 will be described with reference to FIG. 6. The driving characteristics determination unit 220 estimates the classification of each of drivers A to E according to a classification estimation rule. Specifically, the driving characteristics determination unit 220 estimates the classification of each of drivers A to E using on-board sensor data corresponding to each of drivers A to E, that is, on-board sensor data corresponding to items I1 and I2. Here, the classification estimation rule is a rule for estimating the classification of each driver based on their driving characteristics. A specific example of the classification estimation rule is a rule indicating a clustering method such as the k-means method and each data used in the clustering method. The on-board sensor data is data acquired when each driver drives the automobile 100 six times. Data indicating feature amounts extracted from the on-board sensor data is sometimes referred to as on-board sensor data. A specific example of the feature amount is quartiles. Each item corresponds to a combination of one of the driving behavior labels and one of the sensors equipped in the automobile 100. When two items are different from each other, at least one of the driving behavior labels corresponding to the items and the sensors corresponding to the items differ between the two items. Each driver corresponds to a subject. The driving characteristic determination unit 220 infers that Driver A, Driver B, and Driver C are in the same classification with respect to the driving characteristics corresponding to items I1 and I2. Here, multiple drivers belonging to the same cluster are estimated to belong to a classification having the same driving characteristics. The clustering results are utilized for driving assistance. The comparison unit 240 then performs processing on multiple drivers belonging to each cluster for each cluster. That is, the comparison unit 240 performs processing on multiple drivers whose driving characteristics are inferred to be in the same classification. It may be inferred that the classifications corresponding to each of the multiple clusters are the same.

[0030] 5 to 7 are diagrams illustrating the processing of the comparison unit 240. The processing of the comparison unit 240 will be described using FIGS. 5 to 7. Each of FIGS. 5 to 7 corresponds to FIG. 4. The comparison unit 240 selects, from the data stored in the data storage unit 290, data that the driving characteristics determination unit 220 did not use when classifying each driver as unused data. Specifically, the unused data is selected based on the corresponding driving behavior label and the corresponding sensor. The unused data corresponds to the actual data identified by the "driver" and the "unused item." The data is time-series data acquired from the sensor. Here, the unused data is assumed to be data corresponding to item I3. The comparison unit 240 determines whether the corresponding unused data for multiple drivers determined by the driving characteristics determination unit 220 to have the same driving characteristics tend to be similar. Here, when each driver belongs to the same classification, the unused data for each driver is typically expected to be concentrated in the same cluster as shown in FIG. 5, or to vary randomly and uncorrelated as shown in FIG. 6.

[0031] That is, when the classification estimation rule used by the driving characteristics determination unit 220 is a rule corresponding to the final stage where further refinement is not possible, or when the classification estimation rule is appropriate, the distribution of unused data is considered to be as shown in Fig. 5 or 6. However, when the unused data has additional classification capability, it is considered that an event will be observed in which the data corresponding to each driver are distributed at positions that are not the same as each other on the axis corresponding to item I3, as shown in Fig. 7.

[0032] The comparison unit 240 checks whether or not all of the sensor data stored in the data storage unit 290 has classification ability, and presents the check result to the data analyst. Note that the comparison unit 240 typically targets sensor data corresponding to one sensor in one check.

[0033] 8 shows an example of the hardware configuration of data analysis apparatus 200 according to this embodiment. Data analysis apparatus 200 is composed of a computer. Data analysis apparatus 200 may be composed of multiple computers.

[0034] As shown in the figure, the data analysis device 200 is a computer that includes hardware such as a processor 11, a memory 12, an auxiliary storage device 13, an input interface 14, an output interface 15, and a communication device 16. These pieces of hardware are appropriately connected via signal lines.

[0035] The processor 11 is an integrated circuit (IC) that performs arithmetic processing and controls the hardware of a computer. Specific examples of the processor 11 include a central processing unit (CPU), a digital signal processor (DSP), or a graphics processing unit (GPU). The data analysis apparatus 200 may include multiple processors that replace the processor 11. The multiple processors share the role of the processor 11.

[0036] The memory 12 is typically a volatile storage device, specifically a random access memory (RAM). The memory 12 is also called a primary storage device or a main memory. Data stored in the memory 12 is saved in the secondary storage device 13 as needed.

[0037] The auxiliary storage device 13 is typically a non-volatile storage device, and specific examples thereof include a ROM (Read Only Memory), an HDD (Hard Disk Drive), or a flash memory. Data stored in the auxiliary storage device 13 is loaded into the memory 12 as needed. The memory 12 and the auxiliary storage device 13 may be configured integrally.

[0038] The input interface 14 is a port to which an input device is connected. A specific example of the input interface 14 is a USB (Universal Serial Bus) terminal. A specific example of the input device is a keyboard and a mouse.

[0039] The output interface 15 is a port to which an output device is connected. A specific example of the output interface 15 is a USB (Universal Serial Bus) terminal. A specific example of the output device is a display.

[0040] The communication device 16 is a receiver and a transmitter, and is specifically a communication chip or a NIC (Network Interface Card).

[0041] Each unit of the data analysis device 200 may use the input interface 14, the output interface 15, and the communication device 16 as appropriate when communicating with other devices.

[0042] Auxiliary storage device 13 stores a data analysis program. The data analysis program is a program that causes a computer to realize the functions of each unit included in data analysis device 200. The data analysis program is loaded into memory 12 and executed by processor 11. The functions of each unit included in data analysis device 200 are realized by software.

[0043] Data used when executing the data analysis program and data obtained by executing the data analysis program are stored in a storage device as appropriate. Each part of the data analysis device 200 uses a storage device as appropriate. Specific examples of the storage device include at least one of the memory 12, the auxiliary storage device 13, a register in the processor 11, and a cache memory in the processor 11. Note that the terms "data" and "information" may have the same meaning. The storage device may be independent of the computer. The functions of the memory 12 and the auxiliary storage device 13 may be realized by other storage devices.

[0044] The data analysis program may be recorded on a computer-readable non-volatile recording medium. Specific examples of the non-volatile recording medium include an optical disk and a flash memory. The data analysis program may be provided as a program product. The hardware configuration of the on-board control device 130 may be the same as the hardware configuration of the data analysis device 200.

[0045] ***Description of Operation*** The operation procedures of the devices provided in the driving assistance system 90 correspond to a data analysis method. Also, the programs that realize the operations of the devices provided in the driving assistance system 90 correspond to data analysis programs.

[0046] 9 is a flowchart showing an example of the operation of the comparison section 240. The operation of the comparison section 240 will be described with reference to FIG.

[0047] (Step S101) The comparison unit 240 acquires data indicating the determination result of the driving characteristics by the driving characteristics determination unit 220 and data indicating the feature amount of each data for each driving behavior stored in the data storage unit 290.

[0048] (Step S102) The comparison unit 240 identifies, from the data indicating the characteristic amounts acquired in step S101, each piece of data corresponding to each unused item in the determination of the driving characteristics by the driving characteristics determination unit 220 as unused data.

[0049] (Step S103) The comparison unit 240 executes the process of this step for each driving characteristic category and each unused item. Here, each unused data corresponding to each unused item is defined as target data. The comparison unit 240 determines whether or not the Q1-Q3 sections in the feature quantities of the target data corresponding to each driver overlap between multiple drivers who are determined to have the same driving characteristics. Here, Q1 indicates the first quartile, and Q3 indicates the third quartile. The determination result of the comparison unit 240 in this step is referred to as the similarity determination result. Note that, with regard to the determination of similarity, the comparison unit 240 may determine whether or not there is similarity using any criteria, such as determining that there is similarity when the feature quantities of the target data corresponding to each driver are in an inclusive relationship.

[0050] (Step S104) The comparison unit 240 generates an evaluation result indicating the result of evaluating the classification ability of each unused item based on the similarity determination result, and presents the generated evaluation result to the data analyst.

[0051] FIG. 10 is a diagram illustrating a specific example of the processing of steps S103 and S104. The processing of this step will be described using FIG. 10. Here, the target data is data indicating the median vehicle speed when the automobile 100 is traveling straight. First, in step S103, the comparison unit 240 determines whether the Q1-Q3 sections in the feature quantities of the target data corresponding to each driver overlap. Specifically, the Q1-Q3 section corresponding to driver A overlaps with the Q1-Q3 section corresponding to driver C. Therefore, the comparison unit 240 determines that the target data corresponding to driver A and the target data corresponding to driver C are similar. Furthermore, the Q1-Q3 section corresponding to driver B does not overlap with the Q1-Q3 sections corresponding to any of the other drivers. Therefore, the comparison unit 240 determines that the target data corresponding to driver B is not similar to the target data corresponding to each of the other drivers. Next, in step S104, the comparison unit 240 generates a similarity determination table based on the similarity determination results. Here, "1" indicates that there is a similarity between the two drivers in terms of driving characteristics, and "0" indicates that there is no similarity between the two drivers in terms of driving characteristics. Then, the comparison unit 240 evaluates the classification ability of the target data based on the generated similarity determination table. Specifically, the comparison unit 240 determines that the target data has the ability to classify into two types, that is, the target data has the ability to classify multiple drivers into two types in terms of driving characteristics. Here, based on the target data, it is estimated that Driver A and Driver C have the same driving characteristics, and Driver B has different driving characteristics.

[0052] FIG. 11 shows a specific example of data presented by the comparison unit 240 in step S104. In this example, for each unused item, the presence or absence of classification ability and a specific driver classification are shown. In this example, each unused item is an item determined by a combination of the driving action shown in the "driving action label" column and the type of sensor signal shown in the "sensor signal name" column. Each driving action shown in the "driving action label" column corresponds to a target driving action. Each sensor that acquired the data shown in the "sensor signal name" column corresponds to a target sensor. Any of driver A, driver B, and driver C may be considered the first driver described above, and any of them may be considered the second driver described above.

[0053] ***Description of Effects of First Embodiment*** According to this embodiment, the comparison unit 240 presents the data analyst with the analysis results of the classification ability for each unused item. Therefore, the data analyst can determine whether the driving characteristic determination results based on the existing classification estimation rules are appropriate based on the presented analysis results. As a specific example, when a certain unused item is determined to have classification ability for two or more categories, it may be more appropriate to classify multiple drivers determined to belong to the same category into multiple categories. In other words, the data analyst can determine that determining the multiple drivers as belonging to the same category may be an erroneous determination. Therefore, this embodiment can provide the data analyst with an opportunity to adjust classification thresholds, etc. Furthermore, this embodiment can provide the data analyst with an opportunity to consider further subdividing the driving characteristic classifications, such as further dividing multiple drivers actually determined to belong to the same category into multiple categories.

[0054] As a specific example, driving characteristics are considered to differ from region to region (e.g., the Kanto region and the Kansai region, etc.). As a specific example, if experimental data is acquired in the Kanto region when developing the classification estimation rules, the classification ability of each unused item corresponding to the operational data in the Kanto region is considered to be low. On the other hand, if the same classification estimation rules are used in the Kansai region, where no operational data is acquired when developing the classification estimation rules, the classification ability of each unused item may be high. According to this embodiment, it is possible to relatively easily adjust the classification estimation rules between multiple regions with different driving habits.

[0055] ***Other Configurations*** <Modification 1> Fig. 12 shows an example of the hardware configuration of a data analysis apparatus 200 according to this modification. The data analysis apparatus 200 includes a processing circuitry 18 instead of the processor 11, the processor 11 and memory 12, the processor 11 and auxiliary storage device 13, or the processor 11, memory 12, and auxiliary storage device 13. The processing circuitry 18 is hardware that realizes at least a portion of the components included in the data analysis apparatus 200. The processing circuitry 18 may be dedicated hardware, or may be a processor that executes a program stored in the memory 12.

[0056] When processing circuitry 18 is dedicated hardware, processing circuitry 18 may be, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or a combination thereof. Data analysis apparatus 200 may include multiple processing circuits that replace processing circuitry 18. The multiple processing circuits share the role of processing circuitry 18.

[0057] In data analysis apparatus 200, some functions may be realized by dedicated hardware, and the remaining functions may be realized by software or firmware.

[0058] The processing circuitry 18 is realized by, for example, hardware, software, firmware, or a combination of these. The processor 11, memory 12, auxiliary storage device 13, and processing circuitry 18 are collectively referred to as "processing circuitry." In other words, the functions of the functional components of the data analysis device 200 are realized by the processing circuitry. Data analysis devices 200 according to other embodiments may also have a configuration similar to that of this modification.

[0059] Second Embodiment The following mainly describes the differences from the above-described embodiment with reference to the drawings.

[0060] ***Description of Configuration*** The configuration of the driving assistance system 90 according to the second embodiment is the same as the configuration of the driving assistance system 90 according to the first embodiment. The feature calculation unit 132 according to the present embodiment determines whether time-series data formed from data acquired by sensors provided in the automobiles 100 equipped with the feature calculation unit 132 has unimodalities, and generates information indicating the determination result. The communication function unit 210 according to the present embodiment receives first unimodal information and second unimodal information. The first unimodal information is information indicating whether the first unused data has unimodalities, and is information indicating the result of the determination made in the automobiles 100 from which the first unused data was acquired. The second unimodal information is information indicating whether the second unused data has unimodalities, and is information indicating the result of the determination made in the automobiles 100 from which the second unused data was acquired. Each of the first unimodal information and the second unimodal information is information calculated by the feature calculation unit 132 provided in one of the automobiles 100. The first unimodal information and the second unimodal information may be information generated in different automobiles 100. The comparison unit 240 according to this embodiment outputs the first unimodal information and the second unimodal information.

[0061] When analyzing driver characteristics, the distribution shape of each sensor data is considered to be important. Statistical quantities such as mean and variance generally used as feature quantities are parameters that assume a normal distribution. Therefore, if the distribution shape of each sensor data differs from the assumed distribution shape, it is not appropriate to extract these feature quantities from each sensor data. Here, operational data may contain unknown data distributions that were not obtained when the experimental data was acquired.

[0062] In this embodiment, information indicating the relative position (min / Q1 / median / Q3 / max) in the distribution is extracted to retain information indicating the shape of the distribution. A specific example of the information indicating the relative position is information for drawing a box-and-whisker plot, which is often used to confirm the overall shape of a distribution. Here, using a box-and-whisker plot is not particularly suitable when the shape of the distribution is multimodal. Therefore, in this embodiment, the automobile 100 determines whether the shape of the distribution is unimodal, and adds a flag indicating the determination result to the data transmitted from the automobile 100 to the data analysis device 200. Here, the Silberman test and the like are known as algorithms for determining unimodality. Note that, if there is sufficient storage capacity or communication capacity, the information indicating the relative position may include information indicating other statistics, information indicating the data distribution, such as a histogram.

[0063] FIG. 13 is a diagram illustrating the process of extracting information indicating relative positions. The notations "stop" and "go straight" in the graph correspond to driving behavior labels. In the example shown in FIG. 13, relative position information of the data distribution when the automobile 100 is going straight is extracted. In this example, the relative position information indicates the minimum value, first quartile, median, third quartile, maximum value, and outlier as each relative position. Furthermore, since the data distribution when the automobile 100 is going straight is unimodal, data indicating a flag indicating the presence of unimodality is transmitted to the data analysis device 200.

[0064] ***Explanation of Operation*** Fig. 14 is a flowchart showing an example of processing by the automobile 100. The processing by the automobile 100 will be described with reference to Fig. 14 .

[0065] (Step S201) The on-board controller 130 acquires sensor data from each of the driving performance detection unit 110 and the driver information acquisition unit 120. Note that step S201 and step S202 are repeatedly executed while the automobile 100 is traveling.

[0066] (Step S202) The on-board controller 130 stores the sensor data acquired in step S201 in the temporary data storage unit 190.

[0067] (Step S203) The driving behavior determination unit 131 assigns a corresponding driving behavior label and corresponding driver information to each piece of sensor data stored in the temporary data storage unit 190.

[0068] (Step S204) The feature amount calculation unit 132 determines whether each piece of sensor data is unimodal for each driving behavior, and extracts a feature amount (relative position) of each piece of sensor data.

[0069] (Step S205) The driving support unit 140 transmits data indicating the unimodal determination result by the feature amount calculation unit 132 and the feature amounts extracted by the feature amount calculation unit 132 to the data analysis device 200.

[0070] ***Explanation of Effect of Second Embodiment*** According to this embodiment, information indicating the relative position of the distribution shape is stored in the automobile 100, so that a box-and-whisker plot can be drawn based on the relative position indicated by the stored information. That is, according to this embodiment, it is possible to preserve the distribution shape that would be lost if only the mean, variance, etc. were stored. Therefore, according to this embodiment, data analysts can obtain data that is easier to analyze. Furthermore, according to this embodiment, the results of determining whether each piece of sensor data is unimodal are stored. Therefore, data analysts can determine whether it is appropriate to use each feature. Note that even when using statistics that assume a normal distribution, such as the mean and variance, the results of determining whether a piece of data is unimodal are useful in determining the validity of the statistics.

[0071] ***Other Embodiments*** The above-described embodiments can be freely combined, or any of the components of each embodiment can be modified, or any of the components can be omitted from each embodiment. Furthermore, the embodiments are not limited to those shown in embodiments 1 and 2, and various modifications are possible as needed. The procedures described using flowcharts, etc., can be modified as appropriate.

[0072] REFERENCE SIGNS LIST 11 Processor, 12 Memory, 13 Auxiliary storage device, 14 Input interface, 15 Output interface, 16 Communication device, 18 Processing circuit, 20 Exterior communication path, 90 Driving assistance system, 100 Automobile, 110 Driving operation detection unit, 111 Accelerator sensor, 112 Brake sensor, 113 Vehicle speed sensor, 114 Acceleration sensor, 115 Steering angle sensor, 116 GPS sensor, 120 Driver information acquisition unit, 121 In-vehicle camera, 130 In-vehicle control device, 131 Driving behavior determination unit, 132 Feature calculation unit, 133 Communication function unit, 140 Driving assistance unit, 190 Temporary data storage unit, 200 Data analysis device, 210 Communication function unit, 220 Driving characteristics determination unit, 230 Driving characteristics utilization unit, 231 Driving assistance determination unit, 232 Driving characteristics notification unit, 240 Comparison unit, 290 Data storage unit, 291 driving pattern DB.

Claims

1. A data analysis device included in a driving support system that estimates the classification of each driver who drove each vehicle based on time-series data composed of data acquired by sensors included in each vehicle, and determines driving support content based on the estimated classification, When it is estimated that the classifications of the first driver and the second driver are the target classifications, when an item not used when estimating the classifications of the first driver and the second driver is defined as an unused item, based on the feature amount of the first unused data which is time-series data corresponding to both the first driver and the unused item, and the feature amount of the second unused data which is time-series data corresponding to both the second driver and the unused item, a comparison unit that determines whether the unused item has the ability to classify the first driver and the second driver into different classifications A data analysis device comprising: The unused item is an item corresponding to a combination of a target sensor included in each vehicle and a target driving action which is a driving action during driving of each vehicle, The unused item is a data analysis device used to adjust classification estimation rules in data-driven development.

2. The feature amount of the first unused data includes information indicating the feature of the distribution of the first unused data, The data analysis device according to claim 1, wherein the feature amount of the second unused data includes information indicating the feature of the distribution of the second unused data.

3. The comparison unit according to claim 2, which determines whether the unused item has the ability to classify the first driver and the second driver into different classifications according to the overlapping interval between the distribution of the first unused data and the distribution of the second unused data.

4. The comparison unit according to claim 3, which determines whether the unused item has the ability to classify the first driver and the second driver into different classifications according to the overlapping interval between the interval from the first quartile to the third quartile in the first unused data and the interval from the first quartile to the third quartile in the second unused data.

5. The data analysis device according to any one of claims 1 to 4, wherein each of the feature amount of the first unused data and the feature amount of the second unused data includes information used to draw a box plot.

6. The comparison unit presents to a data analyst the result of determining whether the unused item has the ability to classify the first driver and the second driver into different classifications. The data analysis device according to any one of claims 1 to 4.

7. A driving support system including the data analysis device according to claim 6, wherein when each motor vehicle included in the driving support system is a target vehicle, the target vehicle A feature quantity calculation unit that determines whether time series data composed of data acquired by a sensor included in the target vehicle has unimodality and generates information indicating the determination result is provided with The data analysis device further A communication function unit that receives first unimodality information, which is information indicating whether the first unused data has unimodality and is information indicating the result determined in the vehicle in which the first unused data was acquired, and second unimodality information, which is information indicating whether the second unused data has unimodality and is information indicating the result determined in the vehicle in which the second unused data was acquired is provided with The comparison unit outputs the first unimodality information and the second unimodality information. A driving support system.

8. A data analysis method executed by a computer of a data analysis device included in a driving support system that estimates the classification of each driver who drove each motor vehicle based on time series data composed of data acquired by sensors included in each motor vehicle and determines driving support content based on the estimated classification, wherein when the data analysis device estimates that the classifications of the first driver and the second driver are target classifications, when items not used when estimating the classifications of the first driver and the second driver are defined as unused items, based on the feature quantity of first unused data, which is time series data corresponding to both the first driver and the unused items, and the feature quantity of second unused data, which is time series data corresponding to both the second driver and the unused items, it is determined whether the unused item has the ability to classify the first driver and the second driver into different classifications The unused item is an item corresponding to a combination of a target sensor included in each motor vehicle and a target driving action, which is a driving action during driving of each motor vehicle The unused item is a data analysis method used to adjust classification estimation rules in data-driven development.

9. A data analysis program executed by a data analysis device, which is a computer included in a driving support system that estimates the classification of each driver who has driven each vehicle based on time-series data composed of data acquired by sensors provided in each vehicle, and determines driving support content based on the estimated classification. When it is estimated that the classifications of each of the first driver and the second driver are target classifications, when an item not used when estimating the classifications of each of the first driver and the second driver is defined as an unused item, based on the feature amount of first unused data, which is time-series data corresponding to both the first driver and the unused item, and the feature amount of second unused data, which is time-series data corresponding to both the second driver and the unused item, a comparison process for determining whether the unused item has the ability to classify the first driver and the second driver into different classifications from each other. A data analysis program that causes the data analysis device to execute the above. The unused item is an item corresponding to a combination of a target sensor provided in each vehicle and a target driving action that is a driving action during the driving of each vehicle. The unused item is a data analysis program used to adjust classification estimation rules in data-driven development.