Data collection device and data collection method
The data collection device and method address the challenge of identifying new siren sounds by using a sound detection and behavior detection system to label unlearned siren sounds, enhancing the training of learning models for emergency vehicle detection.
Patent Information
- Application Number
- PCT/JP2025/019842
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-01
- Filing Date
- 2025-06-02
- Publication Date
- 2026-01-08
AI Technical Summary
Existing methods for training learning models to detect siren sounds from emergency vehicles struggle with identifying new siren sounds, as they are not effectively labeled or recognized by existing systems.
A data collection device and method that includes a sound detection unit to identify candidate emergency sounds, a behavior detection unit to recognize driver evasive actions, and a label setting unit to associate learning labels with sound data, enabling accurate identification of unlearned siren sounds.
Enables the collection of sound data useful for training models to detect new siren sounds by leveraging driver evasive actions as training information, ensuring accurate identification of such sounds.
Smart Images

Figure JP2025019842_08012026_PF_FP_ABST
Abstract
Description
Data collection device and data collection method CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application is based on Patent Application No. 2024-106304 filed in Japan on July 1, 2024, and the contents of the original application are incorporated by reference in their entirety.
[0002] The disclosure herein relates to data collection techniques for collecting acoustic data that is used to train learning models.
[0003] Patent Literature 1 describes a method for automatically generating labeled sound data by an autonomous driving system. Specifically, in the method disclosed in Patent Literature 1, sound data emitted by an emergency vehicle is labeled, and a machine learning algorithm is trained using the labeled sound data.
[0004] JP 2022-58593 A
[0005] Patent Literature 1 does not disclose a method for identifying whether measured sound data is a siren sound emitted by an emergency vehicle when a new siren sound is introduced to an emergency vehicle. Therefore, even if a learning model can be trained to detect known siren sounds, it is difficult to train the learning model to detect new siren sounds.
[0006] The present disclosure aims to provide a data collection device and a data collection method capable of collecting sound data useful for training a learning model to detect new siren sounds.
[0007] In order to achieve the above object, one disclosed aspect is a data collection device that collects sound data used to train a learning model for detecting siren sounds of emergency vehicles, the data collection device comprising: a sound detection unit that detects candidate emergency sounds that are candidates for siren sounds that the learning model has not yet learned; a behavior detection unit that detects when the driver has taken evasive action to avoid the emergency vehicle in relation to the candidate emergency sound; and a label setting unit that, when the driver has taken evasive action, associates a learning label indicating that the sound data of the candidate emergency sound is an unlearned siren sound.
[0008] Another disclosed aspect is a data collection method for collecting sound data used to train a learning model for detecting the siren sound of an emergency vehicle, the data collection method including, in processing performed by at least one processor, the steps of: detecting a candidate emergency sound that is a candidate for an unlearned siren sound by the learning model; detecting that the driver has taken evasive action to avoid the emergency vehicle in relation to the candidate emergency sound; and, if the driver has taken evasive action, associating the sound data of the candidate emergency sound with a learning label indicating that it is an unlearned siren sound.
[0009] In these embodiments, when a driver takes an evacuation action to avoid an emergency vehicle, a learning label indicating that the sound data is an unlearned siren sound is associated with the sound data. In this way, if the driver's evacuation action is used as training information, even if a new siren sound is introduced to an emergency vehicle, it becomes possible to accurately identify the new siren sound. As a result, it becomes possible to collect sound data useful for training a learning model to detect new siren sounds.
[0010] It should be noted that the reference numbers in parentheses in the claims merely indicate an example of the correspondence with the specific configurations in the embodiments described below, and do not limit the technical scope in any way. Furthermore, claims not explicitly stated in the claims may be combined together if no particular problems arise in the combination.
[0011] 1 is a diagram showing an overall view of a model learning system for vehicles according to the present disclosure; FIG. 2 is a block diagram showing an overall view of an in-vehicle system including a sound recognition ECU according to a first embodiment; FIG. 3 is a diagram explaining details of an evacuation-related period for collecting unlearned siren sounds; FIG. 4 is a flowchart showing details of an upload process executed by an autonomous driving ECU; FIG. 5 is a flowchart showing details of an upload process for collecting information on inappropriate behavior control; FIG. 6 is a flowchart showing details of an upload process for collecting position information of a target of caution; FIG. 7 is a block diagram showing an overall view of a model learning system including an in-vehicle system of a second embodiment; FIG. 8 is a block diagram showing an overall view of a model learning system including a learning server of a third embodiment;
[0012] Hereinafter, several embodiments will be described with reference to the drawings. Note that corresponding components in each embodiment are given the same reference numerals, and redundant description may be omitted. When only a portion of the configuration is described in each embodiment, the configuration of another embodiment described previously can be applied to the remaining portion of the configuration. Furthermore, in addition to the combinations of configurations explicitly stated in the description of each embodiment, configurations of several embodiments can also be partially combined together even if not explicitly stated, as long as there is no particular problem with the combination.
[0013] First Embodiment The functions of a data collection device according to a first embodiment of the present disclosure are realized by a sound recognition ECU (Electronic Control Unit) 10 shown in FIGS. 1 and 2 . The sound recognition ECU 10 is mounted on a vehicle together with an autonomous driving ECU 20. By mounting the autonomous driving ECU 20, the vehicle becomes an autonomous driving vehicle (ADV) equipped with an autonomous driving function. The autonomous driving vehicle ADV can travel under autonomous driving control of autonomous driving level 2 (hereinafter referred to as driving assistance control) or autonomous driving control of autonomous driving level 3 or higher (hereinafter referred to as autonomous driving control). Note that the autonomous driving levels in this disclosure are based on standards defined by the Society of Automotive Engineers.
[0014] [Vehicle Model Learning System] The vehicle model learning system MLS shown in Fig. 1 is an ecosystem including a large number of autonomously driven vehicles ADV and at least one training server 110. In the model learning system MLS, training data TD collected by the large number of autonomously driven vehicles ADV is provided to the training server 110 via a network NW. The training server 110 trains the learning model 80 using the large amount of data TD collected from the large number of autonomously driven vehicles ADV.
[0015] The learning server 110 deploys the training-completed learning model 80 to each autonomous driving vehicle ADV. An update method such as OTA (Over the Air) is used to deploy the learning model 80. By deploying the learning model 80, the software of the sound recognition ECU 10, the autonomous driving ECU 20, etc. is automatically updated in the autonomous driving vehicle ADV. As a result, as the mileage of a large number of autonomous driving vehicles ADV increases, the functions of the sound recognition ECU 10 and the autonomous driving ECU 20 are improved.
[0016] The following describes in detail the learning server 110 used in the model learning system MLS and the in-vehicle system 1 installed in the autonomous driving vehicle ADV.
[0017] 1 and 2 is a virtual configuration provided on the cloud. The training server 110 is managed by an information technology company that provides cloud infrastructure, a vehicle manufacturer that provides an autonomous driving vehicle (ADV), or the like. The training server 110 may also be an on-premise server device that is physically managed by the vehicle manufacturer.
[0018] The training server 110 is configured to be able to communicate with the in-vehicle systems 1 of a large number of autonomously driving vehicles ADV via a network NW. The training server 110 includes a processor 111, a RAM 112, a storage unit 113, an input / output interface, and an internal bus connecting these components, and functions as a high-performance computer that performs arithmetic processing at high speed.
[0019] The processor 111 is hardware for arithmetic processing coupled to the RAM 112. The processor 111 accesses the RAM 112 to execute various processes (instructions) related to the management of the data TD and the provision of the learning model 80. The storage unit 113 is a memory for storing a data processing program that realizes functions related to model learning. By executing the data processing program by the processor 111, functional units such as a data acquisition unit 121, a model training unit 122, and a model provision unit 123 are constructed in the learning server 110 as functional units for training the learning model 80.
[0020] The data acquisition unit 121 acquires data TD uploaded by each in-vehicle system 1 by receiving the data via the network NW. The data acquisition unit 121 can collect various types of data TD from multiple in-vehicle systems 1. The data acquisition unit 121 performs preprocessing such as cleaning, noise removal, and outlier detection on the data TD, and formats the data TD into a format suitable for machine learning training. The data acquisition unit 121 sorts the large amount of collected data TD and stores data TD of the same type used for training one learning model 80 in a database in a state where they are associated with each other.
[0021] The model training unit 122 uses a large amount of data TD accumulated in a database to train the learning model 80. The learning model 80 is constructed based on techniques such as deep learning and reinforcement learning. The model training unit 122 starts re-learning the learning model 80 when a certain amount of training data TD has been accumulated, or when a certain amount of time has passed since the previous update.
[0022] The model providing unit 123 provides the learning model 80 generated by the model training unit 122 to the in-vehicle system 1 of each autonomous driving vehicle ADV via the network NW. The new learning model 80 that has completed training is acquired by each in-vehicle system 1 via the network NW. The new learning model 80 replaces the old learning model 80 and is installed in the sound recognition ECU 10, the autonomous driving ECU 20, etc.
[0023] 2 is mounted on each autonomously driving vehicle ADV that is communicatively connected to a training server 110. The in-vehicle system 1 is configured with a large number of components (nodes) connected to a communication line 99 of an in-vehicle LAN (Local Area Network). The in-vehicle LAN is constructed using communication protocols such as CAN (Controller Area Network, registered trademark) and Ethernet (registered trademark).
[0024] The camera unit 31, the exterior acoustic sensor 32, the locator 35, the in-vehicle communication device 39, the cruise control ECU 40, the HMI (Human Machine Interface) system 50, the autonomous driving ECU 20, the sound recognition ECU 10, etc. are connected to the communication line 99. These nodes connected to the communication line 99 can communicate with each other. Some of the nodes may be electrically connected directly to each other and can communicate without going through the communication line 99.
[0025] The camera unit 31 and the exterior acoustic sensor 32 are mounted on the autonomous vehicle ADV and are perimeter monitoring sensors (autonomous sensors) that monitor the environment surrounding the autonomous vehicle ADV (hereinafter referred to as the host vehicle Am). The autonomous vehicle ADV may also be equipped with perimeter monitoring sensors such as millimeter-wave radar, lidar, and sonar. The perimeter monitoring sensors provide detection information of targets present around the host vehicle to the sound recognition ECU 10, the autonomous driving ECU 20, etc.
[0026] The camera unit 31 includes a front camera module, a rear camera module, and front, rear, left, and right surround view camera modules. The camera unit 31 includes multiple camera modules, enabling it to capture images of the entire surroundings of the vehicle Am. The camera unit 31 sequentially outputs image data captured by each camera module or analysis information of the image data as detection information to the communication line 99.
[0027] The exterior acoustic sensor 32 is mainly composed of a microphone element that converts sound into an electrical signal. The microphone element functions as a capacitor microphone that outputs an electrical signal based on a change in capacitance caused by a thin diaphragm vibrating due to sound pressure. A MEMS (Micro Electro Mechanical Systems) microphone or the like can be used as the microphone element. Alternatively, the exterior acoustic sensor 32 may be configured as a piezoelectric microphone that uses a piezoelectric sensor instead of a capacitor microphone. The piezoelectric sensor converts sound into an electronic signal using the piezoelectric element.
[0028] The exterior acoustic sensors 32 are held on the external structure of the vehicle Am with the sound collection surfaces of the microphone elements facing the external structure. Multiple exterior acoustic sensors 32 are provided on the front, rear, left and right sides, and ceiling of the vehicle Am. The exterior acoustic sensors 32 provided at each location can detect sounds coming from all around the vehicle Am. The exterior acoustic sensors 32 sequentially output sound data SD generated by each microphone element to the sound recognition ECU 10 as detection information.
[0029] The locator 35 includes a GNSS (Global Navigation Satellite System) receiver, an inertial sensor, etc. The locator 35 sequentially determines the position and traveling direction of the host vehicle Am by combining positioning signals received from multiple positioning satellites by the GNSS receiver, measurement results from the inertial sensor, and vehicle speed information output to the communication line 99. The locator 35 sequentially outputs position information and direction information of the host vehicle Am based on the positioning results to the communication line 99 as locator information.
[0030] The in-vehicle communication device 39 is an external communication unit mounted on the host vehicle Am. The in-vehicle communication device 39 functions as a V2X (Vehicle to Everything) communication device. The in-vehicle communication device 39 has at least a V2N (Vehicle to Cellular Network) communication function conforming to communication standards such as LTE (Long Term Evolution) and 5G, and transmits and receives radio waves to and from base stations around the host vehicle Am. In addition to base stations installed on the ground, the in-vehicle communication device 39 may transmit and receive radio waves to and from unmanned aircraft and low-earth orbit satellites located in the stratosphere.
[0031] The in-vehicle communication device 39 enables collaboration between the cloud-based learning server 110 and the in-vehicle system 1 (Cloud to Car) via V2N communication. By installing the in-vehicle communication device 39, the vehicle Am becomes a connected car that can connect to the Internet. The in-vehicle communication device 39 transmits data TD collected by the in-vehicle system 1 to the learning server 110 and receives information provided by the learning server 110.
[0032] The cruise control ECU 40 is an electronic control device that mainly includes a microcontroller. The cruise control ECU 40 generates vehicle speed information indicating the current traveling speed of the host vehicle Am based on detection signals from wheel speed sensors provided at the hubs of each wheel, and sequentially outputs the generated vehicle speed information to a communication line 99. The cruise control ECU 40 has at least the functions of a brake control ECU, a drive control ECU, and a steering control ECU. The cruise control ECU 40 continuously controls the braking force of each wheel, the output of the powertrain, and the steering angle based on either an operation command based on the driver's driving operation or a control command from the automatic driving ECU 20. The cruise control ECU 40 sequentially outputs information indicating the content of the driver's operation command, in other words, information indicating the type of driving operation performed by the driver, to the sound recognition ECU 10 as driving operation information.
[0033] The HMI system 50 has an input interface function that accepts operations by an occupant (such as a driver) of the vehicle Am, and an output interface function that presents information to the occupant. The HMI system 50 is composed of a display device, an audio device, an operation device, an HMI control device, etc. The HMI control device comprehensively controls the presentation of information using the display device, the audio device, etc. The HMI control device understands the content of user operations input to the operation device and provides operation information indicating the content of the user operations to the autonomous driving ECU 20, etc.
[0034] 1 and 2 is a computer that mainly includes a control circuit that includes a processor 21, a RAM 22, a storage unit 23, an input / output interface, and an internal bus connecting these. The processor 21 accesses the RAM 22 to execute various processes (instructions) for implementing the autonomous driving control method according to the present disclosure. The storage unit 23 is a memory that stores various programs (autonomous driving control programs, etc.) executed by the processor 21. As the processor 21 executes the programs, the autonomous driving ECU 20 is configured with an environment recognition unit 71, a behavior control unit 72, an information linking unit 73, etc. as functional units for implementing the autonomous driving function.
[0035] The environment recognition unit 71 recognizes the driving environment of the host vehicle Am using detection information acquired from the camera unit 31 and the sound recognition ECU 10. The environment recognition unit 71 uses a learning model 80 (e.g., a model for each object sign) on the image data acquired from the camera unit 31 to identify dynamic objects such as other vehicles, emergency vehicles, and pedestrians, as well as static objects such as traffic lights and road signs. The environment recognition unit 71 grasps the type, size, relative position, movement direction, and relative speed of dynamic objects around the host vehicle, as well as the type, size, and relative position of static objects around the host vehicle. The environment recognition unit 71 sequentially provides the behavior control unit 72 with the recognition results of the driving environment, including information on dynamic and static objects.
[0036] The behavior control unit 72 cooperates with the cruise control ECU 40 to switch the control state of the host vehicle Am between automatic driving and manual driving. In addition, the behavior control unit 72 switches the automation level of the automatic driving control implemented by the automatic driving function. When the automatic driving function has control over driving operations, the behavior control unit 72 uses a learning model 80 (e.g., a behavior judgment model, etc.) to determine in real time the behavior of the host vehicle Am corresponding to the driving environment, specifically, the direction of travel, driving speed, whether or not to change lanes, etc.
[0037] The behavior control unit 72 generates a planned driving line along which the host vehicle Am will travel based on the behavior determination result, and sequentially outputs control commands based on the planned driving line to the cruise control ECU 40. The behavior control unit 72 cooperates with the cruise control ECU 40 to perform acceleration / deceleration control, steering control, and the like of the host vehicle Am according to the planned driving line. As an example, when an emergency vehicle approaching the host vehicle Am is detected, the behavior control unit 72 performs automatic avoidance control to avoid the emergency vehicle.
[0038] The information linking unit 73 links with an ECU linking unit 63 (described later) of the sound recognition ECU 10, enabling information sharing between the autonomous driving ECU 20 and the sound recognition ECU 10. The information linking unit 73 acquires detection information about the surroundings of the host vehicle based on sound data SD from the sound recognition ECU 10. When the autonomous driving function has control over driving operations, the information linking unit 73 notifies the sound recognition ECU 10 of the content of the behavior control of the host vehicle Am that the autonomous driving function will implement. When the environment recognition unit 71 recognizes an emergency vehicle based on imaging data, the information linking unit 73 provides the sound recognition ECU 10 with detection information about the emergency vehicle based on the imaging data.
[0039] The information linking unit 73 links with the learning server 110 and controls updates to the learning model 80 used in the autonomous driving ECU 20. When the latest learning model 80 is prepared on the learning server 110, the information linking unit 73 acquires the new learning model 80 with improved functionality via the network NW and the in-vehicle communication device 39. The information linking unit 73 automatically updates the learning model 80 when autonomous driving control is not being executed, for example.
[0040] <Configuration of the Sound Recognition ECU> The sound recognition ECU 10 is a computer that mainly includes a control circuit equipped with a processor 11, RAM 12, storage unit 13, an input / output interface, and an internal bus connecting these. The processor 11 accesses the RAM 12 to execute various processes (instructions) for implementing the sound recognition method and data collection method according to the present disclosure. The storage unit 13 is a memory that stores various programs (sound recognition program, data collection program, etc.) executed by the processor 11. As the processor 11 executes the programs, functional units such as a data analysis unit 61, a behavior detection unit 62, an ECU collaboration unit 63, a position recognition unit 64, a label setting unit 65, and a server collaboration unit 66 are configured in the sound recognition ECU 10.
[0041] The data analysis unit 61 sequentially acquires sound data SD measured by the exterior acoustic sensor 32. The data analysis unit 61 has a function of analyzing the sound data SD to detect targets present around the vehicle, and a function of extracting sound data SD to be used for training the learning model 80. The storage unit 13 pre-stores a plurality of learning models 80 (e.g., sound detection models) used by the data analysis unit 61 for target detection.
[0042] The data analysis unit 61 uses the learning model 80 on the sound data SD to detect siren sounds emitted by emergency vehicles. The data analysis unit 61 generates emergency vehicle detection information based on the detection of siren sounds. Here, emergency vehicles include police vehicles, fire engines, ambulances, etc. An emergency vehicle detected by the data analysis unit 61 is an emergency vehicle that is emitting a siren sound or the like for emergency purposes. The data analysis unit 61 identifies the type of siren being sounded by the emergency vehicle based on the sound data SD and the learning model 80. The data analysis unit 61 determines the tone of the siren, the length of one siren sound, whether or not a warning bell is sounding, etc., and determines the type of emergency vehicle approaching the host vehicle Am and the state of emergency use, etc.
[0043] The data analysis unit 61 detects horn sounds from other vehicles traveling around the host vehicle Am. The data analysis unit 61 detects the voices of children and the barks of animals (pets, etc.) such as dogs that are emitted around the host vehicle Am. The data analysis unit 61 sequentially provides the ECU linkage unit 63 with detection information on emergency vehicles, children, animals, etc., and detection information on horn sounds.
[0044] The data analysis unit 61 collects sound data SD used to train a learning model 80 (sound detection model) for detecting siren sounds from emergency vehicles. Siren sounds vary depending on the type of emergency vehicle, and may be replaced with new siren sounds. Furthermore, siren sounds may differ depending on the manufacturer of the siren sound generator. The data analysis unit 61 collects sound data SD (candidate emergency sounds, described below) that are not identified as siren sounds by the learning model 80 but may be siren sounds.
[0045] The behavior detection unit 62 acquires driving operation information from the cruise control ECU 40. The behavior detection unit 62 refers to the driving operation information and determines what driving operations have been performed by the driver of the host vehicle Am. Based on the driving operation information, the behavior detection unit 62 detects that the driver has performed an evacuation action to avoid an emergency vehicle. Driver operations related to the evacuation action include a braking operation to decelerate the host vehicle Am, a steering operation to pull the host vehicle Am over to the side of the road, and a button operation (user operation) to start flashing the emergency flasher lights (hazard lamps).
[0046] The ECU linking unit 63 links with the information linking unit 73 of the automatic driving ECU 20 to acquire information generated by the automatic driving ECU 20 and provide information generated by the sound recognition ECU 10 to the automatic driving ECU 20. Specifically, the ECU linking unit 63 acquires control information indicating the content of behavioral control of the host vehicle Am implemented by the automatic driving function from the automatic driving ECU 20. The ECU linking unit 63 acquires emergency vehicle detection information based on image data captured by the camera unit 31 from the information linking unit 73 and determines whether an emergency vehicle has been detected based on the image data.
[0047] The ECU linking unit 63 sequentially outputs detection information of emergency vehicles, children, animals, etc. based on the sound data SD of the exterior acoustic sensor 32, as well as detection information of horn sounds, etc. to the information linking unit 73. The behavior control unit 72 performs behavior control to decelerate the host vehicle Am based on the detection information of children, animals, etc. In addition, the behavior control unit 72 suspends the ongoing behavior control based on the detection information of horn sounds.
[0048] The position determination unit 64 can acquire locator information output to the communication line 99 by the locator 35. When the data analysis unit 61 detects a child's voice or an animal's cry, the position determination unit 64 refers to the locator information and determines the current position information and direction information of the host vehicle Am.
[0049] The label setting unit 65 associates a training label LB with the training data TD to be transmitted to the training server 110. The training label LB is information indicating the correct answer to the question "What is this?" for the training data TD. By associating the training label LB with the training data TD, the model training unit 122 can perform supervised learning or semi-supervised learning.
[0050] Specifically, when the driver takes an evacuation action, the label setting unit 65 associates the sound data SD extracted for learning the siren sound with a learning label LB indicating that the sound data is an unlearned siren sound. On the other hand, when the driver does not take an evacuation action, the label setting unit 65 associates the sound data SD extracted for learning the siren sound with a learning label LB indicating that the sound data is unrelated to a siren sound.
[0051] When the data analysis unit 61 detects a horn sound, the label setting unit 65 records the content of the behavioral control that was being executed at the timing when the horn was sounded, based on the control information acquired by the ECU linkage unit 63. The label setting unit 65 sets the behavioral control when the horn sound was detected in learning data TD, and associates a learning label LB that indicates that the control is inappropriate with this data TD.
[0052] When the data analysis unit 61 detects a child's voice or the cry of an animal such as a pet, the label setting unit 65 records the current location information, etc. acquired by the location grasping unit 64. The label setting unit 65 sets the location information, etc. when the child's voice or the animal's cry is detected in the learning data TD. The label setting unit 65 may include time information when the child's voice or the animal's cry is detected in the learning data TD together with the location information, etc. The label setting unit 65 associates a learning label LB indicating the presence of a child or an animal with the data TD mainly consisting of location information, etc.
[0053] The server cooperation unit 66 cooperates with the in-vehicle communication device 39 and transmits the learning data TD, to which the learning labels LB have been assigned by the label setting unit 65, to the learning server 110. In addition, the server cooperation unit 66 cooperates with the learning server 110 and controls updates to the learning model 80 (sound detection model) used in the sound recognition ECU 10. When the latest learning model 80 is prepared in the learning server 110, the server cooperation unit 66 acquires the new learning model 80 with improved functionality via the network NW and the in-vehicle communication device 39. The server cooperation unit 66 automatically updates the learning model 80, for example, when sound detection is not being performed by the data analysis unit 61.
[0054] [Details of the Process for Collecting Sound Data of a New Siren Sound] Next, the process for collecting sound data SD for learning a new siren sound will be further described in detail with reference to FIGS. 1 to 3. FIG.
[0055] The data analysis unit 61 detects candidate emergency sounds. A candidate emergency sound is a sound that is a candidate for a new siren sound and has not yet been learned by the learning model 80 (sound detection model) used to distinguish siren sounds. A candidate emergency sound is a sound that is not distinguished as a siren sound by the learning model 80 and is a sound similar to a siren sound.
[0056] Specifically, the data analysis unit 61 removes noise components such as engine noise and tire noise of the vehicle Am from the sound data SD measured by the exterior acoustic sensor 32. From the sound data SD after filtering out the noise components, the data analysis unit 61 detects, as candidate emergency sounds, sounds that are included in a specific frequency band that may be used as a siren sound and that are repeated in a certain pattern (have periodicity). The data analysis unit 61 may set extraction conditions for candidate emergency sounds such as sounds exceeding a predetermined volume and sounds whose tone periodically changes up and down.
[0057] The data analysis unit 61 detects candidate emergency sounds only during manual driving periods when the driver has control of the driving operation. Furthermore, the data analysis unit 61 excludes sound data SD measured during normal driving of the vehicle Am from sound data SD to be determined as unlearned siren sounds. The data analysis unit 61 extracts unlearned siren sounds by limiting the sound data SD measured during a period in which evacuation actions to avoid an emergency vehicle are performed based on the driver's judgment (hereinafter referred to as an evacuation-related period Pe, see FIG. 3 ). The evacuation-related period Pe is a period related to evacuation actions and includes at least a period in which the vehicle Am is moved to an evacuation position and waits in a stopped state for the emergency vehicle to pass.
[0058] Specifically, the data analysis unit 61 includes the time from when the evacuation action is started (hereinafter, start time t1) to when normal driving is resumed (hereinafter, normal return time t3) in the evacuation-related period Pe. Furthermore, the data analysis unit 61 may set the time from a predetermined time before the start time t1 (hereinafter, predetermined pre-start time t0) to the normal return time t3 as the evacuation-related period Pe so as to partially include the period of normal driving. In this embodiment, the time from the predetermined pre-start time t0 to the normal return time t3 is set as the evacuation-related period Pe.
[0059] The data analysis unit 61 divides the evacuation-related period Pe into multiple detection target periods (see FIG. 3 ). The data analysis unit 61 analyzes whether or not there is a candidate emergency sound in each of the sound data SD for the multiple detection target periods. As an example, the data analysis unit 61 divides the evacuation-related period Pe into three periods to be considered. The first detection target period (hereinafter, the first detection period Pe1) is the period from a predetermined pre-start timing t0 to the start timing t1. The next detection target period (hereinafter, the second detection period Pe2) is the period from the start timing t1 to the evacuation movement end timing t2. The next detection target period (hereinafter, the third detection period Pe3) is the period from the evacuation movement end timing t2 to the normal return timing t3.
[0060] The first detection period Pe1 is set based on the start timing t1 and is set to a period of, for example, 10 to 20 seconds before the start timing t1. The predetermined pre-start timing t0, which is the start time of the first detection period Pe1, is the timing when an emergency vehicle approaches the host vehicle Am to a degree that the siren sound emitted by the emergency vehicle can be measured by the exterior acoustic sensor 32. In other words, the data analysis unit 61 sets the length of the first detection period Pe1 so as to include the initial period when recording of the siren sound in the sound data SD begins. The data analysis unit 61 may adjust the length of the first detection period Pe1 to be shorter as the traveling speed of the host vehicle Am increases.
[0061] Start timing t1, which is the end time of the first detection period Pe1 and the start time of the second detection period Pe2, is the time when the driver starts an operation related to an evacuation behavior. The data analysis unit 61 determines the start timing t1 based on the driving operation information acquired by the behavior detection unit 62. During the second detection period Pe2 after start timing t1, the driver moves the host vehicle Am toward an avoidance position where the emergency vehicle can be avoided.
[0062] The evacuation movement end timing t2, which is the end time of the second detection period Pe2 and the start time of the third detection period Pe3, is the time when the movement of the host vehicle Am to the avoidance position is completed. The data analysis unit 61 determines the evacuation movement end timing t2 based on the vehicle speed information transmitted through the communication line 99. During the third detection period Pe3 after the evacuation movement end timing t2, the driver keeps the host vehicle Am stopped at the avoidance position and waits for the emergency vehicle to pass. Then, when the emergency vehicle moves away from the host vehicle Am, the driver starts the host vehicle Am from the avoidance position toward the center of the lane.
[0063] The normal return timing t3, which is the end time of the third detection period Pe3, is the time when the vehicle transitions from the evacuation behavior to normal driving. For example, the normal return timing t3 is the timing when the vehicle Am starts accelerating after returning to the center of the lane. The data analysis unit 61 determines the normal return timing t3 based on the driving operation information acquired by the behavior detection unit 62.
[0064] The data analysis unit 61 determines whether a candidate emergency sound is an unlearned siren sound based on its similarity to known siren sounds that have already been learned in the learning model 80. To evaluate the similarity to the learned siren sound, the data analysis unit 61 uses the learning model 80 to obtain an evaluation score for the candidate emergency sound. The evaluation score is a value between 0 and 1, for example, and the closer the evaluation score is to 1, the higher the likelihood that it is a siren sound. The data analysis unit 61 determines that a sound whose evaluation score exceeds a determination threshold is a siren sound. The data analysis unit 61 sets a threshold closer to 0 than the determination threshold as a threshold (hereinafter referred to as the candidate threshold) for determining whether or not the sound is an unlearned siren sound. The data analysis unit 61 determines that a sound whose evaluation score does not exceed the determination threshold but exceeds the candidate threshold is an unlearned siren sound. In other words, the data analysis unit 61 estimates that a sound whose evaluation score is slightly lower than the determination threshold is an unlearned siren sound.
[0065] The data analysis unit 61 verifies whether a candidate emergency sound is an unlearned siren sound based on the similarity between each candidate emergency sound detected in two consecutive detection target periods. The data analysis unit 61 identifies the frequency and repetition period of the candidate emergency sound detected in each detection target period. The data analysis unit 61 considers frequency changes due to the Doppler effect of the measured siren sound and verifies whether the evaluation score of the candidate detection sound continuously increases from the first detection period Pe1 to the second detection period Pe2. Furthermore, the data analysis unit 61 verifies whether a change in the evaluation score corresponding to the transition from approaching to moving away of the emergency vehicle occurs from the second detection period Pe2 to the third detection period Pe3. Furthermore, the data analysis unit 61 verifies whether a candidate emergency sound is an unlearned siren sound based on a comparison of each candidate emergency sound detected in multiple non-consecutive detection target periods. For example, the data analysis unit 61 verifies whether a corresponding candidate emergency sound is recorded in the first detection period Pe1 and the third detection period Pe3.
[0066] The data analysis unit 61 cooperates with the ECU linkage unit 63 to determine whether an emergency vehicle has been detected by the autonomous driving ECU 20 based on the image data captured by the camera unit 31. If an emergency vehicle is detected based on the image data after the driver has started to take an evacuation action, the data analysis unit 61 determines that the detected candidate emergency sound is an unlearned siren sound. On the other hand, if no emergency vehicle is detected based on the image data after the driver has started to take an evacuation action, the data analysis unit 61 determines that the detected candidate emergency sound is a sound unrelated to a siren sound.
[0067] If the evaluation score of the candidate emergency sound exceeds the candidate threshold, the data analysis unit 61 may determine that the candidate emergency sound is an unlearned siren sound even if no emergency vehicle is detected based on the imaging data. Furthermore, even if a sound is determined to be a siren sound based on the learning model 80, the data analysis unit 61 determines that a false positive siren sound has been detected if the driver does not take evacuation action and no emergency vehicle is detected based on the imaging data. The data analysis unit 61 determines that a false positive siren sound is a sound unrelated to a siren sound.
[0068] When a candidate emergency sound determined by the data analysis unit 61 to be an unlearned siren sound is generated, the label setting unit 65 links the sound data SD of the candidate emergency sound with a learning label LB indicating that the sound is an unlearned siren sound. When a candidate emergency sound determined by the data analysis unit 61 to be a sound unrelated to a siren sound or a false positive siren sound is generated, the label setting unit 65 links the sound data SD with a learning label LB indicating that the sound is unrelated to a siren sound. The server linking unit 66 transmits the sound data SD linked with the learning label LB to the training server 110.
[0069] <Sound Data Upload Process> Details of the upload process that realizes the collection of the sound data SD described above will be described below based on Fig. 4 and with reference to Figs. 1 to 3. The upload process shown in Fig. 4 is repeatedly started by the sound recognition ECU 10 during the manual driving period when the driver performs driving operations.
[0070] In S11 of the upload process, sound data SD measured by the exterior acoustic sensor 32 is read. The data analysis unit 61 may process the sound data SD substantially in real time, or may collectively process sound data SD accumulated for a predetermined period or a predetermined amount. Furthermore, the data analysis unit 61 may collectively process the sound data SD accumulated while the host vehicle Am is in use after the use of the host vehicle Am has ended (for example, during a parking period, etc.).
[0071] In S12, the data analysis unit 61 detects candidate emergency sounds that are candidates for unlearned siren sounds from the sound data SD loaded in S11. If no candidate emergency sounds are detected (S12: NO), the upload process is terminated. On the other hand, if a candidate emergency sound is detected (S12: YES), the behavior detection unit 62, in S13, refers to the driving operation information and determines whether the driver took evacuation action to avoid an emergency vehicle in relation to the candidate emergency sound. If the behavior detection unit 62 did not detect evacuation action (S13: NO), in S18, the label setting unit 65 associates the sound data SD of the candidate emergency sound with a learning label LB indicating that the sound data is unrelated to a siren sound.
[0072] On the other hand, if the behavior detection unit 62 detects an evacuation behavior (S13: YES), the data analysis unit 61 divides the evacuation-related period Pe related to this evacuation behavior into multiple detection target periods in S14. In S15, the data analysis unit 61 analyzes whether or not there is an unlearned siren sound (new siren sound) in each of the sound data SD for the multiple detection target periods.
[0073] Based on the analysis and verification in S15, the data analysis unit 61 determines in S16 whether the detected candidate emergency sound is an unlearned siren sound. The data analysis unit 61 determines whether the candidate emergency sound is an unlearned siren sound based on the similarity to a learned siren sound (known siren sound). Furthermore, if an emergency vehicle is detected based on the imaging data, the data analysis unit 61 determines that the candidate emergency sound is an unlearned siren sound.
[0074] If the data analysis unit 61 determines that the candidate emergency sound is not an unlearned siren sound (S16: NO), the label setting unit 65 associates the sound data SD of the candidate emergency sound with a learning label LB indicating that the sound is unrelated to a siren sound in S18. On the other hand, if the data analysis unit 61 determines that the candidate emergency sound is an unlearned siren sound (S16: YES), the label setting unit 65 associates the sound data SD of the candidate emergency sound with a learning label LB indicating that the sound is an unlearned siren sound in S17.
[0075] In S19, the server link unit 66 uploads the sound data SD to which the learning label LB has been assigned in S17 or S18 to the training server 110. The training server 110 updates the training model 80 (sound detection model) when a certain amount of sound data SD with the learning label LB has been accumulated.
[0076] [Details of the process for collecting other learning data] Next, the details of each process for uploading the behavior control information and location information to which the learning label LB has been assigned to the learning server 110 as learning data TD will be explained below based on Figures 5 and 6 and with reference to Figures 1 to 3.
[0077] 5 is a process for providing learning data TD indicating inappropriate behavioral control to the learning server 110. The upload process shown in Fig. 5 is repeatedly started by the sound recognition ECU 10 during an autonomous driving period in which the host vehicle Am is traveling under driving assistance control or autonomous driving control.
[0078] In S51, the data analysis unit 61 reads the sound data SD measured by the exterior acoustic sensor 32. The data analysis unit 61 may process the sound data SD substantially in real time, as in the upload process related to the siren sound (see FIG. 4), or may process the accumulated sound data SD collectively at a predetermined timing.
[0079] In S52, the data analysis unit 61 detects the sound of a horn sound from another vehicle traveling around the host vehicle Am from the sound data SD read in S51. If a horn sound is not detected (S52: NO), the upload process is terminated. On the other hand, if a horn sound is detected (S52: YES), in S53, the label setting unit 65 records the content of the behavioral control that was being executed at the time the horn was sounded, based on the control information acquired by the ECU linkage unit 63.
[0080] In S54, the label setting unit 65 sets the behavioral control to be performed when a horn sound is detected in the learning data TD, and associates a learning label LB indicating that the control is inappropriate with this data TD. In S55, the server linking unit 66 uploads the data TD to which the learning label LB has been assigned in S54 to the learning server 110. The learning server 110 updates the learning model 80 (behavior judgment model) when a certain amount of data TD with the learning label LB has been accumulated.
[0081] <Uploading process of location information where attention targets exist> The uploading process shown in Fig. 6 is a process of providing learning data TD indicating detection locations of animals such as children and pets (hereinafter referred to as attention targets) to the learning server 110. The uploading process shown in Fig. 6 is repeatedly started by the sound recognition ECU 10 while the host vehicle Am is traveling. The uploading process of collecting detection locations of attention targets is performed both during the manual driving period and the autonomous driving period.
[0082] In S61, the data analysis unit 61 reads the sound data SD measured by the exterior acoustic sensor 32. The data analysis unit 61 may process the sound data SD substantially in real time, as in the upload process of the other data TD described above (see FIGS. 4 and 5), or may process the accumulated sound data SD collectively at a predetermined timing.
[0083] In S62, the data analysis unit 61 detects sounds such as children's voices and animal cries emitted around the vehicle Am from the sound data SD read in S61. If no children's voices or animal cries are detected (S62: NO), the upload process is terminated. On the other hand, if a children's voice or an animal cry is detected (S62: YES), the label setting unit 65 records the detection location of the object of caution based on the current location information acquired by the location determination unit 64 in S63.
[0084] In S64, the label setting unit 65 sets location information indicating the detection location of the target object in the learning data TD, and associates a learning label LB indicating the presence of the target object with this data TD. In S65, the server linking unit 66 uploads the data TD to which the learning label LB has been assigned in S64 to the learning server 110.
[0085] The learning server 110 updates the learning model 80 (hereinafter referred to as the area discrimination model) when a certain amount of data TD with learning labels LB has been accumulated. The area discrimination model is a learning model 80 that discriminates attention target areas where there is a high possibility that an attention target exists. The area discrimination model is provided to the autonomous driving ECU 20 and is used by the behavior control unit 72 to discriminate attention target areas. When the behavior control unit 72 determines, based on the area discrimination model, that the host vehicle Am is traveling in an attention target area, it sets the traveling speed of the host vehicle Am to be lower than when the host vehicle Am is traveling outside the attention target area.
[0086] (Summary of First Embodiment) In the first embodiment described so far, when a driver takes an evacuation action to avoid an emergency vehicle, a learning label LB indicating that the sound data SD is an unlearned siren sound is associated with the sound data SD. In this way, if the driver's evacuation action is used as training information, even if a new siren sound is introduced to an emergency vehicle, it becomes possible to accurately distinguish the siren sound. As a result, it becomes possible to collect sound data SD that is useful for training the learning model 80 to detect new siren sounds.
[0087] Additionally, in the first embodiment, the sound data SD measured during the evacuation-related period Pe associated with the evacuation action is limited to determining whether the sound is an unlearned siren sound. By limiting learning of unlearned siren sounds to the evacuation-related period Pe in this way, candidate emergency sounds detected during normal driving periods are eliminated as noise. As a result, the accuracy of the learning labels LB is improved, enabling effective learning of unlearned siren sounds.
[0088] In the first embodiment, the evacuation-related period Pe is the period from a predetermined timing t0 before the start of the evacuation behavior or the start timing t1 to a normal return timing t3 at which the evacuation behavior transitions to normal driving. This limited learning period allows the data analysis unit 61 to accurately extract, as learning data TD, sound data SD that is likely to be an unlearned siren sound. As a result, the accuracy of the learning labels LB can be further improved.
[0089] Furthermore, in the first embodiment, the evacuation-related period Pe is divided into multiple detection target periods. The data analysis unit 61 then analyzes whether or not there is an unlearned siren sound in the sound data SD for each of the multiple detection target periods. As described above, the data analysis unit 61 can analyze the sound data SD while taking into account the measurement environment, such as the positional relationship between the host vehicle Am and the emergency vehicle, and whether or not the emergency vehicle is approaching the host vehicle Am.
[0090] Additionally, in the first embodiment, the data analysis unit 61 verifies whether a candidate emergency sound is an unlearned siren sound based on the similarity between each candidate emergency sound detected during two consecutive detection target periods. This makes it possible to analyze the sound data SD while taking into account the influence of the Doppler effect. Furthermore, the data analysis unit 61 can also identify new siren sound data SD measured when the emergency vehicle is farther away, based on new siren sound data SD measured when the emergency vehicle is closer.
[0091] In the first embodiment, the data analysis unit 61 verifies whether the candidate emergency sound is an unlearned siren sound by comparing each candidate emergency sound detected during multiple non-consecutive detection target periods. As described above, the data analysis unit 61 can identify new siren sounds measured during each detection target period, for example, based on sound data SD measured when the host vehicle Am is stopped and there is little noise associated with driving. As a result, the data analysis unit 61 can accurately extract sound data SD that records unlearned siren sounds emitted by emergency vehicles located far from the host vehicle Am.
[0092] Furthermore, in the first embodiment, the data analysis unit 61 determines whether a candidate emergency sound is an unlearned siren sound based on the similarity to siren sounds that have already been learned in the learning model 80. Even if the siren sound is new, the frequency and repetition pattern of the sound are unlikely to change significantly from existing siren sounds. Therefore, by using the similarity to learned siren sounds, the data analysis unit 61 can accurately determine whether the siren sound is an unlearned siren sound.
[0093] Additionally, the ECU linkage unit 63 in the first embodiment determines whether an emergency vehicle has been detected based on the imaging data. If an emergency vehicle has been detected based on the imaging data, the data analysis unit 61 determines that the candidate emergency sound is an unlearned siren sound. By further using the emergency vehicle detection result using the imaging data as training information in this way, the data analysis unit 61 can accurately identify unlearned siren sounds newly installed in emergency vehicles and collect them as learning data TD.
[0094] In the first embodiment, if no evacuation action is taken, the sound data SD of the candidate emergency sound is associated with a learning label LB indicating that the sound data SD is unrelated to a siren sound. This allows for effective collection of not only sound data SD representing unlearned siren sounds, but also sound data SD representing noise sounds similar to the unlearned sound data SD. As a result, the training server 110 can efficiently train the learning model 80 to detect new siren sounds.
[0095] Furthermore, in the first embodiment, if no evacuation action is taken and no emergency vehicle is detected based on the imaging data, the sound data SD of the candidate emergency sound is associated with a learning label LB indicating that the sound is unrelated to a siren. As described above, by using detection information based on the imaging data to assign the learning label LB indicating that the sound is unrelated to a siren, the accuracy of assigning the learning label LB can be further improved.
[0096] Additionally, the ECU linkage unit 63 of the first embodiment understands the content of the behavioral control of the host vehicle Am implemented by the autonomous driving function. Furthermore, the data analysis unit 61 detects the sound of a horn honking from another vehicle traveling around the host vehicle Am. Then, a learning label LB indicating inappropriate control is associated with the behavioral control when a horn sound is detected. As a result, it is possible to improve the function of the learning model 80 for controlling the behavior of the autonomous driving by utilizing the information on the horn sound.
[0097] In the first embodiment, the location information of the vehicle Am is obtained, and the sounds of children and animals around the vehicle Am are detected. A learning label LB indicating the presence of a child or animal is then associated with the location information when the sound of a child or animal is detected. As a result, the detection information of the sound of children and animals can be used to improve the function of the learning model 80 for controlling the behavior of the autonomous driving system.
[0098] In the first embodiment, the data analysis unit 61 corresponds to the “sound detection unit,” the ECU linkage unit 63 corresponds to the “image detection unit” and the “control understanding unit,” and the sound recognition ECU 10 corresponds to the “data collection device.” Furthermore, the first detection period Pe1, the second detection period Pe2, and the third detection period Pe3 each correspond to a “detection target period.”
[0099] Second Embodiment A second embodiment of the present disclosure shown in Figure 7 is a modified example of the first embodiment. In the in-vehicle system 1 of the second embodiment, the functions of the sound recognition ECU 10 (see Figure 2) are integrated into the autonomous driving ECU 20. In other words, the functions of the data collection device according to the second embodiment are realized by the autonomous driving ECU 20. The autonomous driving ECU 20 executes the upload process of each piece of data TD (see Figures 4 to 6).
[0100] In addition to an environment recognition unit 71, a behavior control unit 72, and an information linkage unit 73, the autonomous driving ECU 20 is configured with a data analysis unit 61, a behavior detection unit 62, and a label setting unit 65 that are substantially the same as those in the first embodiment.
[0101] The environment recognition unit 71 has the function of the position recognition unit 64 (see FIG. 2) of the first embodiment, and refers to locator information to recognize the current position information and direction information of the host vehicle Am (see FIG. 1). The environment recognition unit 71 determines whether an emergency vehicle has been detected based on the image data captured by the camera unit 31.
[0102] The behavior control unit 72 grasps the content of behavior control of the host vehicle Am that is implemented by the automatic driving function, and provides the grasped content of behavior control to the label setting unit 65 .
[0103] The information linking unit 73 has the functions of the server linking unit 66 of the first embodiment (see Figure 2), and performs the process of sending the learning data TD to the learning server 110, to which the learning label LB has been assigned, and the process of updating the learning model 80 in cooperation with the learning server 110.
[0104] The second embodiment described so far also achieves the same effect as the first embodiment, and when a new siren sound is introduced to an emergency vehicle, it becomes possible to accurately distinguish such a siren sound. As a result, it becomes possible to collect sound data SD that is useful for training the learning model 80 to detect the new siren sound. Note that in the second embodiment, the environment recognition unit 71 corresponds to the "position recognition unit" and the "image detection unit," the behavior control unit 72 corresponds to the "control recognition unit," and the autonomous driving ECU 20 corresponds to the "data collection device."
[0105] (Third Embodiment) A third embodiment of the present disclosure shown in Figure 8 is yet another modified example of the first embodiment. In the first embodiment, the main steps of the upload process (S11 to S18, S51 to S54, S61 to S64, see Figures 4 to 6) that were executed by the sound recognition ECU 10 are executed by the learning server 110 in the third embodiment. That is, the sound recognition ECU 10 of the third embodiment is configured to simply transmit sound data SD measured by the exterior acoustic sensor 32 to the learning server 110.
[0106] Meanwhile, the function of extracting sound data SD to be used for training the learning model 80 is implemented in the data acquisition unit 121 of the learning server 110. Specifically, the data acquisition unit 121 acquires, from each vehicle, information transmitted by the in-vehicle communication device 39, such as driving operation information, locator information, emergency vehicle detection information based on imaging data, and autonomous driving function control information, in addition to the sound data SD. Similar to the data analysis unit 61 (see FIG. 2 ), the data acquisition unit 121 detects candidate emergency sounds from the sound data SD and detects whether the driver of the vehicle that transmitted the sound data SD has taken an evacuation action. The data acquisition unit 121 determines whether the candidate emergency sound is an unlearned siren sound based on whether the driver has taken an evacuation action and whether an emergency vehicle has been detected based on imaging data.
[0107] In addition, the model training unit 122 of the training server 110 has a function of associating each learning data TD with a learning label LB. When the driver takes an evacuation action, the model training unit 122 associates a learning label LB, which indicates that the sound data SD is an unlearned siren sound, with the candidate emergency sound sound data SD and stores the label for learning. When the sound data SD with the learning label LB assigned is accumulated, the model training unit 122 trains the learning model 80.
[0108] The third embodiment described so far also achieves the same effect as the first embodiment, and when a new siren sound is introduced to an emergency vehicle, it becomes possible to accurately distinguish such a siren sound. As a result, it becomes possible to collect sound data SD that is useful for training the learning model 80 to detect the new siren sound. In the third embodiment, the data acquisition unit 121 corresponds to the "sound analysis unit," "behavior detection unit," "video detection unit," "control recognition unit," and "position recognition unit," and the model training unit 122 corresponds to the "label setting unit." The learning server 110 corresponds to the "data collection device."
[0109] (Other Embodiments) Although multiple embodiments of the present disclosure have been described above, the present disclosure should not be construed as being limited to the above-described embodiments, and can be applied to various embodiments and combinations within the scope that does not deviate from the gist of the present disclosure.
[0110] The start time and end time of the evacuation-related period Pe in the above embodiment may be changed as appropriate. For example, the timing of departure from the avoidance location may be set as the end time of the evacuation-related period Pe. The start time and end time of the multiple detection target periods may also be changed as appropriate. Furthermore, the process of dividing the evacuation-related period Pe into multiple detection target periods does not need to be implemented.
[0111] In the first modification of the above embodiment, the detection information of the emergency vehicle based on the image data is not used to determine whether the candidate emergency sound is an unlearned siren sound. In the second modification of the above embodiment, the process of linking the learning label LB, which indicates that the candidate emergency sound is unrelated to a siren sound, to the sound data is not performed.
[0112] The sounds recognized by the sound recognition ECU 10 of the first embodiment are not limited to sirens, horns, children's voices, etc. The sound recognition ECU 10 may be capable of recognizing various sounds. The sound recognition ECU 10 according to the third modification of the above embodiment recognizes sounds for detecting road surface conditions, specifically, friction noise between the tires and the road surface, road noise, splashing water, and the sound of stepping on snow. The sound recognition ECU 10 associates learning labels LB indicating the driving state or road surface condition with the sound data SD of such road-related sounds, and uploads the sound data SD to the learning server 110.
[0113] In the in-vehicle system 1 of the fourth modification of the above embodiment, the function of the data collection device is realized by cooperation between the sound recognition ECU 10 and the automatic driving ECU 20. In the fourth modification, the system including the sound recognition ECU 10 and the automatic driving ECU 20 corresponds to the "data collection device."
[0114] In the fifth modification of the above embodiment, the in-vehicle communication device 39 transmits and receives information via wireless communication (road-to-vehicle communication) with a roadside device installed at the side of the road. In addition, the in-vehicle communication device 39 transmits and receives information via wireless communication (vehicle-to-vehicle communication) with an external communication unit or the like installed in another vehicle. As described above, the in-vehicle communication device 39 provides received information acquired via V2X (Vehicle to Infrastructure) communication and V2V (Vehicle to Vehicle) communication to the sound recognition ECU 10 and the autonomous driving ECU 20. As an example, the in-vehicle communication device 39 acquires detection information of an emergency vehicle detected by a roadside device or another vehicle via road-to-vehicle communication or vehicle-to-vehicle communication, and provides the information to the sound recognition ECU 10 and the autonomous driving ECU 20.
[0115] In the sound recognition ECU 10, the ECU interworking unit 63 determines whether an emergency vehicle approaching the host vehicle Am has been detected based on the reception information provided from the in-vehicle communication device 39 and acquired via external communication. The data analysis unit 61, in cooperation with the ECU interworking unit 63, determines that the detected candidate emergency sound is an unlearned siren sound if an emergency vehicle is detected based on the reception information after the driver has started to take an evacuation action. On the other hand, if an emergency vehicle is not detected based on the reception information after the driver has started to take an evacuation action, the data analysis unit 61 determines that the detected candidate emergency sound is a sound unrelated to a siren sound. The data analysis unit 61 may determine that the detected candidate emergency sound is a sound unrelated to a siren sound if an emergency vehicle is not detected based on either the imaging data or the reception information.
[0116] In this modification 5, when an emergency vehicle is detected based on the received information, the data analysis unit 61 determines that the candidate emergency sound is an unlearned siren sound. By further using the emergency vehicle detection result using the received information as training information in this way, the data analysis unit 61 can accurately identify unlearned siren sounds newly installed in emergency vehicles and collect them as learning data TD. Note that in modification 5, the ECU linkage unit 63 corresponds to the "communication detection unit."
[0117] In this disclosure and claims, the term "processor" refers to one or more hardware processors configured to load computer program code (i.e., one or more instructions of a computer program) included in a computer program and execute the processing defined by the code. In other words, a "processor" is a hardware device that executes one or more programmed processes. Therefore, computer program code can also be considered software that can define the processing of the processor depending on its content. For example, a "processor" may be a general-purpose or special-purpose processor, such as a CPU, a microprocessor, a GPU, or a DFP (Data Flow Processor), but is not limited to these.
[0118] In this disclosure and in the claims, the term "memory" refers to one or more hardware memories that are non-transitory tangible recording media configured to store computer program code and / or data accessible to a processor. The "memory" may be implemented using memory technologies such as SRAM, SDRAM, non-volatile / flash-type memory, or other types of memory. Computer program code that constitutes a program may be stored in the memory and executed by a processor to cause the processor to perform the various functions described above.
[0119] Furthermore, the storage medium that stores the "data collection program" according to the present disclosure is not limited to a configuration provided on a circuit board, but may be provided in the form of a memory card or the like, inserted into a slot, and electrically connected to the control circuit of each ECU. The storage medium may also be an optical disk, hard disk drive, solid state drive, or the like that serves as a source from which the program is copied or distributed to the sound recognition ECU 10 or the autonomous driving ECU 20.
[0120] (Disclosure of Technical Ideas) This specification discloses multiple technical ideas described in the following multiple clauses. Some clauses may be described in a multiple dependent form, with the subsequent clause alternatively referring to the preceding clause. Furthermore, some clauses may be described in a multiple dependent form, with the subsequent clause referring to another multiple dependent clause. These multiple dependent clauses define multiple technical ideas.
[0121] (Technical Idea 1) A data collection device that collects sound data (SD) used to train a learning model (80) for detecting siren sounds of emergency vehicles, comprising: a sound detection unit (61) that detects candidate emergency sounds that are candidates for the siren sounds that the learning model has not yet learned, a behavior detection unit (62) that detects that the driver has taken an evacuation action to avoid the emergency vehicle in relation to the candidate emergency sounds, and a label setting unit (65) that, when the driver has taken the evacuation action, associates a learning label (LB) that indicates that the sound data of the candidate emergency sounds is the unlearned siren sound. (Technical Idea 2) The data collection device according to Technical Idea 1, in which the sound detection unit determines that the sound data is the unlearned siren sound by limiting the sound data measured during an evacuation-related period (Pe) related to the evacuation action. (Technical Idea 3) The data collection device according to Technical Idea 2, wherein the sound detection unit defines the evacuation-related period as a period from a predetermined timing (t0) before the start of the evacuation behavior or a timing (t1) when the evacuation behavior starts to a normal return timing (t3) when the evacuation behavior transitions to normal driving. (Technical Idea 4) The data collection device according to Technical Idea 2 or 3, wherein the sound detection unit divides the evacuation-related period into a plurality of detection target periods (Pe1 to Pe3) and analyzes the sound data for each of the plurality of detection target periods to determine whether an unlearned siren sound is present. (Technical Idea 5) The data collection device according to Technical Idea 4, wherein the sound detection unit verifies whether a candidate emergency sound is an unlearned siren sound based on the similarity between each of the candidate emergency sounds detected in two consecutive detection target periods. (Technical Idea 6) The data collection device according to Technical Idea 4 or 5, wherein the sound detection unit verifies whether a candidate emergency sound is an unlearned siren sound based on a comparison between each of the candidate emergency sounds detected in a plurality of non-consecutive detection target periods. (Technical Idea 7) The data collection device according to any one of Technical Ideas 1 to 6, wherein the sound detection unit determines whether the candidate emergency sound is an unlearned siren sound based on the similarity to the siren sound for which the learning model has already learned.(Technical Idea 8) The data collection device according to any one of Technical Ideas 1 to 7, further comprising: an image detection unit that determines whether the emergency vehicle has been detected based on imaging data, wherein the sound detection unit determines that the candidate emergency sound is the unlearned siren sound if the emergency vehicle has been detected based on the imaging data. (Technical Idea 9) The data collection device according to any one of Technical Ideas 1 to 7, further comprising: a communication detection unit that determines whether the emergency vehicle has been detected based on received information acquired through exterior communication, wherein the sound detection unit determines that the candidate emergency sound is the unlearned siren sound if the emergency vehicle has been detected based on the received information. (Technical Idea 10) The data collection device according to any one of Technical Ideas 1 to 9, wherein the label setting unit associates the learning label, indicating that the sound data of the candidate emergency sound is unrelated to the siren sound, with the sound data of the candidate emergency sound if the evacuation action has not been taken. (Technical Idea 11) The data collection device according to any one of Technical Ideas 1 to 10, further comprising: an image detection unit that determines whether the emergency vehicle has been detected based on image data, and the label setting unit, when the evacuation action is not taken and the emergency vehicle is not detected based on the image data, associates the learning label indicating that the sound data of the candidate emergency sound is unrelated to the siren sound. (Technical Idea 12) The data collection device according to any one of Technical Ideas 1 to 11, further comprising: a control understanding unit that understands content of behavioral control of the host vehicle (Am) implemented by an autonomous driving function, and the sound detection unit detects a horn sound sounded by another vehicle traveling around the host vehicle, and the label setting unit associates the learning label indicating that the behavioral control is inappropriate control with the behavioral control when the horn sound is detected by the sound detection unit.(Technical Idea 13) A data collection device as described in any one of Technical Ideas 1 to 12, further comprising a position grasping unit (64, 71) that grasps position information of the host vehicle (Am), wherein the sound detection unit detects at least one of a child's voice and an animal's cry emitted around the host vehicle, and the label setting unit links the learning label indicating the presence of the child or the animal to the position information when the voice or the cry is detected by the sound detection unit. (Technical Idea 14) A data collection program that collects sound data (SD) used to train a learning model (80) for detecting the siren sound of an emergency vehicle, the data collection program causing at least one processor (11, 21, 111) to execute processing including: detecting candidate emergency sounds that are candidates for the siren sound that the learning model has not yet learned (S12); detecting that the driver has taken evasive action to avoid the emergency vehicle in relation to the candidate emergency sound (S13); and, if the driver has taken the evasive action, associating the sound data of the candidate emergency sound with a learning label (LB) indicating that the sound data is the siren sound that has not yet been learned (S17).
Claims
1. A data collection device that collects sound data (SD) used to train a learning model (80) for detecting the siren sound of an emergency vehicle, comprising: a sound detection unit (61) that detects candidate emergency sounds that are candidates for the siren sound that the learning model has not yet learned; a behavior detection unit (62) that detects when a driver has taken evasive action to avoid the emergency vehicle in relation to the candidate emergency sound; and a label setting unit (65) that, when the driver has taken the evasive action, associates a learning label (LB) indicating that the sound data of the candidate emergency sound is the siren sound that has not yet been learned with the sound data of the candidate emergency sound.
2. The data collection device of claim 1, wherein the sound detection unit determines that the siren sound is an unlearned sound by limiting the sound data measured during an evacuation-related period (Pe) related to the evacuation behavior.
3. The data collection device of claim 2, wherein the sound detection unit defines the evacuation-related period as the period from a predetermined timing (t0) before the start of the evacuation behavior or the start timing (t1) to the normal return timing (t3) when the evacuation behavior transitions to normal driving.
4. A data collection device as described in claim 2 or 3, wherein the sound detection unit divides the evacuation-related period into a plurality of detection target periods (Pe1 to Pe3) and analyzes the sound data for each of the plurality of detection target periods to determine whether or not there is an unlearned siren sound.
5. The data collection device of claim 4, wherein the sound detection unit verifies whether a candidate emergency sound is an unlearned siren sound based on the similarity of each candidate emergency sound detected during two consecutive detection target periods.
6. The data collection device of claim 4, wherein the sound detection unit verifies whether the candidate emergency sound is an unlearned siren sound based on a comparison of each candidate emergency sound detected during multiple non-consecutive detection target periods.
7. The data collection device according to claim 1, wherein the sound detection unit determines whether the candidate emergency sound is an unlearned siren sound based on the similarity to the siren sound for which the learning model has already learned.
8. The data collection device of claim 1, further comprising an image detection unit that determines whether the emergency vehicle has been detected based on the imaging data, and the sound detection unit determines that the candidate emergency sound is an unlearned siren sound when the emergency vehicle has been detected based on the imaging data.
9. A data collection device as described in claim 1, further comprising a communication detection unit that determines whether the emergency vehicle has been detected based on received information obtained by external vehicle communication, and the sound detection unit determines that the candidate emergency sound is an unlearned siren sound when the emergency vehicle has been detected based on the received information.
10. The data collection device of claim 1, wherein the label setting unit associates the sound data of the candidate emergency sound with the learning label indicating that the sound data is unrelated to the siren sound if the evacuation action is not taken.
11. The data collection device of claim 1, further comprising an image detection unit that determines whether the emergency vehicle has been detected based on the image data, and wherein the label setting unit, if the evacuation action is not taken and the emergency vehicle is not detected based on the image data, associates the learning label, which indicates that the sound data of the candidate emergency sound is unrelated to the siren sound, with the sound data of the candidate emergency sound.
12. A data collection device as described in claim 1, further comprising a control understanding unit that understands the content of behavioral control of the vehicle (Am) implemented by an autonomous driving function, wherein the sound detection unit detects horn sounds emitted by other vehicles traveling around the vehicle, and the label setting unit associates the learning label, which indicates that the behavioral control is inappropriate, with the behavioral control when the horn sound is detected by the sound detection unit.
13. A data collection device as described in claim 1, further comprising a position grasping unit (64, 71) that grasps position information of the vehicle (Am), wherein the sound detection unit detects at least one of a child's voice and an animal's cry emitted around the vehicle, and the label setting unit links the learning label indicating the presence of the child or the animal to the position information when the voice or the cry is detected by the sound detection unit.
14. A data collection method for collecting sound data (SD) used to train a learning model (80) for detecting the siren sound of an emergency vehicle, the data collection method including the following steps in processing performed by at least one processor (11, 21, 111): detecting a candidate emergency sound that is a candidate for the siren sound that has not yet been learned by the learning model (S12); detecting that the driver has taken evasive action to avoid the emergency vehicle in relation to the candidate emergency sound (S13); and, if the driver has taken evasive action, associating the sound data of the candidate emergency sound with a learning label (LB) indicating that it is the siren sound that has not yet been learned (S17).
Citation Information
Patent Citations
Audio logging for model training and onboard validation utilizing autonomous driving vehicle
JP2022058556A
Post-media convergence in detection of audio and visual of emergency vehicle
JP2022058594A
Safe driving behavior for autonomous vehicle
JP2022069415A
Emergency siren detection in autonomous vehicles
US20230065647A1
Acoustic control method and acoustic control device
WO2023204076A1