System
The system addresses the risk of missing emergency vehicle sirens by using sound collection, analysis, detection, volume adjustment, and notification means to ensure safe responses in vehicles.
Patent Information
- Application Number
- JP2024131351
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-07
- Publication Date
- 2026-02-20
AI Technical Summary
The increasing number of people enjoying music or other loud sounds in vehicles poses a risk of missing emergency vehicle sirens, potentially delaying appropriate responses and leading to accidents.
A system with sound collection means at vehicle corners, analysis means for identifying emergency vehicle sirens, detection means for calculating direction and distance, volume adjustment means to lower sound levels, and notification means for informing occupants.
Quickly and reliably detects emergency vehicles, allowing occupants to take appropriate action by reducing sound volume and providing location and direction notifications.
Smart Images

Figure 2026028735000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In recent years, the number of people enjoying music in their vehicles has increased. However, when music or other sound sources are being played at high volume, there is a high possibility that the siren of an emergency vehicle will be missed. As a result, appropriate response to the emergency vehicle may be delayed, which may result in an accident or delay. The present invention aims to solve the above problem by providing a means for quickly and reliably detecting the approach of an emergency vehicle and notifying the user even while listening to music in the vehicle. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems with a system including sound collection means installed at the four corners of a vehicle, an analysis means that analyzes the collected sound data and monitors surrounding sounds, a detection means that detects the approach of an emergency vehicle based on the analysis results, a volume adjustment means that automatically lowers the volume of the sound source being played inside the vehicle when the detection means detects the approach of an emergency vehicle, and a notification means that notifies users inside the vehicle of the location and approach of the emergency vehicle. Specifically, the analysis means calculates the direction and distance of the sound source based on the sound data collected from the sound collection means, and when the approach of an emergency vehicle is detected, the volume adjustment means lowers the volume of the sound source inside the vehicle, and the notification means makes a specific announcement based on the direction and distance of the emergency vehicle. As a result, users can quickly recognize the presence of an emergency vehicle and take appropriate action.
[0006] The "sound collection means" refers to microphone devices installed at the four corners of the vehicle to collect surrounding sounds.
[0007] The "analysis means" is a device or program for analyzing the sound data collected from the sound collection means and monitoring surrounding sounds.
[0008] The "detection means" is a device or program for detecting the approach of an emergency vehicle based on the analysis results obtained by the analysis means.
[0009] The "volume adjustment means" is a device or program for automatically lowering the volume of the sound source being played inside the vehicle when the detection means detects that an emergency vehicle is approaching.
[0010] "Notification means" refers to a device or program for informing a user of the location and approach of an emergency vehicle.
[0011] An "in-car speaker" is a device installed inside a vehicle to reproduce sound from a sound source.
[0012] "Sound source" is a general term that refers to sounds such as music and navigation information played in the vehicle.
[0013] A "navigation system" is a device or program that displays the current location of a vehicle and provides route guidance to a destination.
[0014] The "difference in sound arrival time" refers to the time difference between when the same sound is detected by multiple microphones, and is used to identify the direction of the sound source.
[0015] A "sound source localization algorithm" is an analytical method for calculating the direction and distance of a sound source based on sound data from a sound collection means. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] As an example of the present invention, an emergency vehicle approach detection system for a vehicle will be described. This system is composed of a sound collection means, an analysis means, a detection means, a volume adjustment means, and a notification means, and quickly and accurately detects the approach of an emergency vehicle and notifies the user.
[0038] overview
[0039] Recording Method:
[0040] Microphone devices installed at the four corners of the vehicle collect ambient sounds, effectively capturing sounds coming from all directions in 360 degrees.
[0041] Analysis method:
[0042] The system analyzes sound data from sound collection devices in real time to identify the siren sounds specific to emergency vehicles, using machine learning algorithms to detect specific frequency bands and sound patterns of siren sounds.
[0043] Detection Methods:
[0044] Based on the analysis results obtained by the analysis means, the system detects the approach of an emergency vehicle and calculates the direction and distance of the sound source based on the difference in arrival time of the sounds from multiple microphones.
[0045] Volume adjustment means:
[0046] If the detection means detects the approach of an emergency vehicle, the volume of any audio source (e.g., music, audiobooks, etc.) being played in the vehicle will be automatically reduced via Bluetooth connection.
[0047] Means of notification:
[0048] Specific announcements will be made over the vehicle's speakers to passengers informing them of the location and approach of emergency vehicles.
[0049] Program processing
[0050] Sound pickup and data transmission
[0051] Terminal: Collects sound data in real time from microphones installed at the four corners of the vehicle and temporarily stores it in a buffer.
[0052] Server: Receives sound data periodically sent from the terminal.
[0053] Siren sound analysis
[0054] Server: Analyzes the received sound data and applies machine learning algorithms to detect frequency bands and sound patterns to identify the sound of an emergency vehicle siren.
[0055] Terminal: Receives the analysis results from the server and prepares any necessary next steps.
[0056] Detecting approaching emergency vehicles
[0057] Terminal: Using data from multiple microphones, a sound source localization algorithm calculates the difference in sound arrival time and determines the direction and distance of the emergency vehicle.
[0058] Automatic volume adjustment
[0059] Device: Based on the detection results, it sends a command via Bluetooth connection to reduce the volume of the sound source inside the car to 50%.
[0060] Generate and play notifications
[0061] Server: Obtains the current location and direction of the emergency vehicle based on data from the navigation system and generates an appropriate announcement.
[0062] Terminal: The generated voice announcement is played over the car's speakers.
[0063] Specific examples
[0064] 1. Sound collection means:
[0065] The device collects sound in real time using microphones installed at the four corners of the vehicle and stores it in a buffer.
[0066] 2. Analysis method:
[0067] The server analyzes the sound data in the buffer using a machine learning algorithm to identify the sound of a siren.
[0068] 3. Detection methods:
[0069] Based on the analysis results from the server, the device calculates the direction and distance of the sound source.
[0070] 4. Volume adjustment means:
[0071] The server detects the approach of an emergency vehicle and the device sends a command via Bluetooth connection to reduce the music volume to 50%.
[0072] 5. Means of notification:
[0073] The server uses data from the navigation system to obtain the location information of the emergency vehicle and generates specific instructions.
[0074] The terminal generates a voice announcement that is played over the car's speakers, prompting the user to take specific action (e.g., there is an emergency vehicle approaching 100 meters ahead from the right; move to the shoulder of the road).
[0075] summary
[0076] This system allows even passengers who are enjoying loud music in their cars to quickly and reliably detect approaching emergency vehicles and respond safely, thereby reducing the risk of accidents and supporting smooth traffic management.
[0077] The processing flow will be explained below.
[0078] Step 1:
[0079] Terminal: Sound is collected by multiple microphones installed at the four corners of the vehicle. The collected sound data is temporarily stored in a buffer.
[0080] Specific operation: Sound data is sampled through the microphone interface, for example, every second, and stored in a buffer. Data from each microphone is stored in a separate buffer.
[0081] Step 2:
[0082] Terminal: Collected sound data is packetized at regular intervals (e.g., every second) and sent to the server.
[0083] Specific operation: The sound data is divided into packets and sent to the server using a network protocol. After sending, the buffer is cleared and waits for the next sample.
[0084] Step 3:
[0085] Server: Analyzes the received sound data in real time, using AI to identify specific frequency bands and patterns of siren sounds.
[0086] How it works: A machine learning model (e.g., a convolutional neural network) is applied to analyze sound data and extract features. If a specific pattern is detected, it is identified as a siren sound.
[0087] Step 4:
[0088] Server: If a siren sound is detected, the result is fed back to the device.
[0089] Specific operation: A data packet containing the analysis results is generated and sent to the terminal via the network.
[0090] Step 5:
[0091] Terminal: Based on the analysis results received from the server, the terminal calculates the difference in sound arrival time from multiple microphones and determines the direction and distance of the sound source.
[0092] Specific operation: An algorithm is applied that uses the difference in sound arrival time to calculate the direction and distance of the sound source using techniques such as triangulation.
[0093] Step 6:
[0094] Terminal: If an approaching emergency vehicle is detected, a command to reduce the volume of the in-car speakers to a certain percentage (e.g., 50%) is sent via Bluetooth.
[0095] Specific behavior: Uses the Bluetooth speaker's volume control protocol to send a command to decrease the volume of the currently playing music or voice by the specified percentage.
[0096] Step 7:
[0097] Server: Works with the navigation system to obtain the current location and direction of emergency vehicles.
[0098] Specific operation: Calls the navigation system's API to obtain the current GPS location and direction of the emergency vehicle. Based on this information, it generates an announcement for the next step.
[0099] Step 8:
[0100] Terminal: Generates specific voice announcements based on location information received from the server.
[0101] What it does: Creates the appropriate announcement using a text generation engine and generates an audio file using a text-to-speech engine.
[0102] Step 9:
[0103] Terminal: The generated voice announcement is played over the vehicle's speakers, providing specific instructions to the user.
[0104] Specific operation: The created audio file is added to the playback queue of the car speakers and begins playing immediately.
[0105] Step 10:
[0106] User: Follow the voice announcement and take appropriate action promptly.
[0107] Specific actions: For example, "An ambulance is approaching from the right, 100 meters ahead. Please move to the shoulder," and follow the instructions to safely move the vehicle to the shoulder.
[0108] Example 1
[0109] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0110] Conventional vehicle emergency vehicle approach detection systems lack the functionality to quickly detect approaching emergency vehicles and notify the driver. In particular, drivers listening to loud music or audiobooks in the vehicle may not notice the approach of an emergency vehicle, increasing the risk of an accident. The objective of this invention is to solve this problem by providing a system that quickly and accurately detects the approach of an emergency vehicle and notifies the driver in an appropriate manner.
[0111] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0112] In this invention, the server includes a sound collection means, a sound data analysis means, an emergency vehicle detection means, an automatic volume adjustment means, a notification means, and a notification generation means. This makes it possible to analyze in real time the sound data collected by the sound collection means installed at the four corners of the vehicle and quickly detect the approach of an emergency vehicle. Furthermore, when an emergency vehicle approaches, the server can automatically lower the volume inside the vehicle and generate an appropriate announcement to notify the driver.
[0113] The "sound collection means" is a device installed at each of the four corners of the vehicle that collects surrounding sounds in real time.
[0114] The "sound data analysis means" is a function that analyzes collected sound data, detects specific frequency bands and patterns, and identifies the siren sounds of emergency vehicles.
[0115] The "emergency vehicle detection means" is a function that detects the approach of an emergency vehicle based on the analysis results obtained by the sound data analysis means.
[0116] The "automatic volume adjustment means" is a function that automatically reduces the volume of the sound source being played inside the vehicle when the emergency vehicle detection means detects the approach of an emergency vehicle.
[0117] The "notification means" is a function for informing the vehicle occupants of the location and approach of an emergency vehicle.
[0118] The "notification generating means" is a function that obtains the location of emergency vehicles based on data from the navigation system and generates an appropriate announcement.
[0119] The present invention relates to a system for detecting approaching emergency vehicles in a vehicle, and specific embodiments thereof are described below. This system is composed of a sound collection means, a sound data analysis means, an emergency vehicle detection means, an automatic volume adjustment means, a notification means, and a notification generation means. By combining these means, it is possible to quickly and accurately detect the approach of an emergency vehicle and notify occupants in the vehicle.
[0120] Sound collection means
[0121] Microphones installed at the four corners of the vehicle act as sound collectors. These microphones collect ambient sounds in real time and temporarily store the data in a buffer. For example, when the vehicle approaches an intersection, each microphone collects the surrounding sounds.
[0122] Sound data analysis method
[0123] The server receives audio data periodically sent from the device. It analyzes the received audio data using a machine learning algorithm to detect specific frequency bands and sound patterns and identify the sound of an emergency vehicle siren. The server then sends the analysis results to the device and prepares for the next process.
[0124] Emergency vehicle detection means
[0125] Based on the analysis results sent from the server, the device uses data from multiple microphones and a sound source localization algorithm to calculate the direction and distance of the emergency vehicle. For example, it can determine that an emergency vehicle is approaching 100 meters behind and to the left.
[0126] Automatic volume adjustment means
[0127] If the device detects an approaching emergency vehicle, it will send a command over the Bluetooth connection to reduce the volume of the in-car audio source (e.g., music) to 50%, allowing the in-car audio system to automatically reduce the current volume.
[0128] Notification means and notification generation means
[0129] The server obtains data on the current location and direction of the emergency vehicle from the navigation system and generates an appropriate announcement for the user. The device receives this announcement and plays it through the car's speakers. For example, it can provide specific instructions to the user, such as "Police vehicle approaching 100 meters behind you on the left, please be careful."
[0130] Prompt Sentence Examples
[0131] Here are some examples of prompts:
[0132] "Please explain the process by which microphones installed at the four corners of the vehicle collect ambient sounds in real time and store that data in a buffer."
[0133] "Please explain in detail how the server uses a machine learning algorithm to analyze the sound data received from the device and identify the sound of a siren."
[0134] "Please explain in detail the algorithm that allows the device to use data from multiple microphones to calculate the direction and distance of emergency vehicles."
[0135] "Please explain the steps to reduce the volume of in-car audio sources to 50% when the device detects an approaching emergency vehicle."
[0136] "Please explain the procedure for the server to generate an announcement to notify the user of the approach of an emergency vehicle and send it to the car speaker."
[0137] In this way, the system of the present invention can quickly detect the approach of an emergency vehicle, provide the user with the necessary information, and encourage appropriate action.
[0138] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0139] Step 1: Sound collection and data transmission
[0140] Terminal
[0141] Input: Microphones installed at the four corners of the vehicle collect surrounding sounds in real time.
[0142] Data processing / data calculation: Collected sound data is temporarily stored in a buffer and prepared to be sent to the server at regular intervals.
[0143] Output: Sends the sound data stored in the temporary buffer to the server.
[0144] Specific operation: For example, when a vehicle is driving and approaches an intersection, the microphone collects sounds of nearby pedestrians and other vehicles and sends the data to a server at regular intervals.
[0145] Step 2: Analyzing the siren sound
[0146] server
[0147] Input: Sound data sent periodically from the device.
[0148] Data processing / data calculation: Applying machine learning algorithms to analyze sound data and detect specific frequency bands and sound patterns to identify the sound of an emergency vehicle siren.
[0149] Output: As an analysis result, information is generated as to whether or not an emergency vehicle siren was detected.
[0150] Specific operation: The server analyzes the sound data received from the device, and if it detects a specific frequency pattern between 700 and 1500 Hz, it identifies it as the sound of an emergency vehicle siren. The analysis result is returned to the device.
[0151] Step 3: Detect approaching emergency vehicles
[0152] Terminal
[0153] Input: Analysis results sent from the server.
[0154] Data processing / data calculation: Based on the difference in arrival time of sound data obtained from multiple microphones, a sound source localization algorithm is used to calculate the direction and distance of the emergency vehicle.
[0155] Output: Information on the approaching direction and distance of the emergency vehicle.
[0156] Specific operation: For example, identify that an emergency vehicle is approaching 100 meters behind and to the left. This information is used in the next step.
[0157] Step 4: Automatically adjust the volume
[0158] Terminal
[0159] Input: Information about the approaching direction and distance of the emergency vehicle, and information about the audio source currently being played.
[0160] Data processing / data calculation: If an approaching emergency vehicle is detected, a command is generated to reduce the volume of the sound source (e.g., music) inside the vehicle to 50%.
[0161] Output: A command to decrease the volume over the Bluetooth connection.
[0162] Specific behavior: For example, if music is playing in the car, the volume will be automatically reduced to 50%. This operation is performed via Bluetooth connection.
[0163] Step 5: Generate and play notifications
[0164] server
[0165] Input: Emergency vehicle approach direction and distance information, and location data from the navigation system.
[0166] Data processing / data calculation: Based on data from the navigation system, appropriate announcements are generated for users.
[0167] Output: The generated announcement.
[0168] Specific actions: For example, generate specific instructions such as "A police vehicle is approaching 100 meters behind you on the left. Be careful."
[0169] Terminal
[0170] Input: Server-generated announcement.
[0171] Data processing / data calculation: Generates commands to play announcements over the car speakers.
[0172] Output: Announcement command to the car speaker.
[0173] Specific operation: The generated voice announcement is played over the car's speakers to notify the user of the approaching emergency vehicle and how to respond. For example, it may say, "Siren approaching, emergency vehicle on the left."
[0174] (Application example 1)
[0175] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0176] Conventional vehicle emergency vehicle notification systems alert users to approaching emergency vehicles by lowering the volume inside the vehicle, but lack the technology for use outside the vehicle or on mobile devices. Furthermore, while it is necessary to quickly and accurately notify users of the direction and distance of approaching emergency vehicles visually and audibly, no appropriate solution for this purpose exists. Therefore, there is a need for a comprehensive emergency vehicle notification system that allows users to respond safely both inside and outside the vehicle.
[0177] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0178] In this invention, the server includes a sound collection means, an analysis means, a detection means, a volume adjustment means, and a notification means. This allows the server to quickly and accurately detect the approach of an emergency vehicle by collecting surrounding sounds, even when using a vehicle or a mobile device, automatically lower the volume of the sound source being played, and effectively notify the user of the location and approach of the emergency vehicle. Specifically, the analysis means calculates the direction and distance of the approaching emergency vehicle, and the information is provided via voice and screen display, allowing the user to quickly take appropriate action.
[0179] A "sound collection means" is a device installed to collect surrounding sounds.
[0180] An "analysis means" is a device or algorithm for analyzing collected sound data and identifying specific sound patterns.
[0181] The "detection means" is a device or function for detecting the approach of a specific sound source, for example, an emergency vehicle, based on the analysis result of the analysis means.
[0182] The "volume adjustment means" is a device or function that automatically adjusts the volume of the sound source being played back when the detection means detects the approach of an emergency vehicle.
[0183] "Notification means" refers to a device or function for informing users of the location and approach of an emergency vehicle.
[0184] A "mobile terminal" is a communication device or computer device that can be carried by a user.
[0185] The "audio output means" is a device for conveying the audio information generated by the notification means to the user.
[0186] "Screen display" refers to the function of displaying textual or graphic information on a mobile terminal or other display device.
[0187] The present invention provides a system that uses a vehicle or a mobile terminal to quickly detect the approach of an emergency vehicle and automatically reduce the volume of the sound source being played. Specific embodiments for carrying out the present invention are described below.
[0188] overview
[0189] This system consists of a sound collection means, an analysis means, a detection means, a volume adjustment means, and a notification means. The sound collection means is a microphone installed in a vehicle or a mobile device that collects surrounding sounds. The analysis means analyzes the collected sound data and identifies specific sound patterns, such as siren sounds. This analysis uses a machine learning algorithm (e.g., a model built with Keras). The detection means detects the approach of an emergency vehicle based on the analysis results. The volume adjustment means automatically lowers the volume of the sound being played when the detection means detects the approach of an emergency vehicle. The notification means effectively notifies the user of the location and approach of the emergency vehicle.
[0190] Program processing
[0191] Sound pickup and data transmission
[0192] The device collects ambient sounds through a microphone, and the collected sound data is temporarily stored in a buffer and processed by an analysis means.
[0193] Siren sound analysis
[0194] The server receives the sound data in the buffer and applies a machine learning algorithm to identify siren sounds. Keras is used as an analysis tool to analyze the collected sound data.
[0195] Detecting approaching emergency vehicles
[0196] The detection means detects the approach of an emergency vehicle based on the results of the analysis means, and calculates the direction and distance of the sound source from the collected sound data.
[0197] Volume adjustment
[0198] The volume control means automatically reduces the volume of the audio source being played when detected, and sends a volume control command via Bluetooth or another communication means.
[0199] Generate and play notifications
[0200] The server generates an appropriate announcement based on the location information of the emergency vehicle, and the notification means notifies the user of this announcement both by voice and by display on the screen.
[0201] Specific examples
[0202] While the user is driving, the device (smartphone) collects surrounding sounds using a built-in microphone and sends them to a server. The server analyzes the sound data in real time, and if a siren sound is detected, it automatically lowers the volume of the smartphone via Bluetooth and notifies the user of the location of an emergency vehicle via a display and voice message, allowing the user to take prompt and appropriate action.
[0203] Prompt Sentence Examples
[0204] "We are developing an emergency vehicle approach detection system for vehicles as a smartphone application. Please generate a Python program that follows the requirements below:
[0205] Built-in microphone for sound capture
[0206] Uses machine learning models to analyze and detect siren sounds in real time
[0207] Volume adjustment and notification functions implemented
[0208] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0209] Step 1:
[0210] The device collects ambient sounds. While the user is driving the vehicle, the device's microphone collects audio data in real time and temporarily stores it in a buffer. The input here is the ambient sounds, and the output is the audio data in the buffer.
[0211] Step 2:
[0212] The device sends the sound data in the buffer to the server. The collected sound data is transferred to the server at regular intervals. The input here is the sound data in the buffer, and the output is the sound data sent to the server.
[0213] Step 3:
[0214] The server analyzes the received sound data. The server uses a machine learning model (a model using Keras) to analyze the sound data and detect siren sounds. The input here is the sound data sent to the server, and the output is the analysis result.
[0215] Step 4:
[0216] The server detects emergency vehicles based on the analysis results. If a siren sound is detected, the server detects the approach of an emergency vehicle and calculates its direction and distance. The input here is the analysis result, and the output is the direction and distance information of the emergency vehicle.
[0217] Step 5:
[0218] The server sends a volume adjustment command to the device. Based on the detection result, the server sends an instruction to the device to lower the volume of the sound being played. The input here is the direction and distance information of the emergency vehicle, and the output is the volume adjustment command to the device.
[0219] Step 6:
[0220] The device performs the volume adjustment. Upon receiving the volume adjustment command, the device automatically reduces the volume of the audio source using Bluetooth or other communication methods. The input here is the volume adjustment command, and the output is the reduced volume.
[0221] Step 7:
[0222] The server generates a notification and sends it to the device. Based on the location information of the emergency vehicle, the server generates a notification for voice and screen display and sends it to the device. The input here is the direction and distance information of the emergency vehicle, and the output is the generated notification data.
[0223] Step 8:
[0224] The terminal notifies the user. The terminal that receives the notification data notifies the user of the approach of an emergency vehicle through voice and screen display. The input here is the notification data, and the output is the notification to the user.
[0225] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0226] As one embodiment of the present invention, we will describe an in-vehicle emergency vehicle approach detection system and a notification system that takes into account the user's emotional state. This system is composed of a sound collection means, an analysis means, a detection means, a volume adjustment means, a notification means, and an emotion engine, and quickly and accurately detects the approach of an emergency vehicle and notifies the user appropriately according to the user's emotional state.
[0227] overview
[0228] Recording Method:
[0229] Microphone devices installed at the four corners of the vehicle collect surrounding sounds, effectively capturing sounds from all directions in 360 degrees.
[0230] Analysis method:
[0231] This is an analysis method that analyzes sound data from sound collection means in real time to identify the siren sounds specific to emergency vehicles. This analysis uses machine learning algorithms to detect specific frequency bands and sound patterns of siren sounds.
[0232] Detection Methods:
[0233] Based on the analysis results obtained by the analysis means, the system detects the approach of an emergency vehicle and calculates the direction and distance of the sound source based on the difference in sound arrival time between multiple microphones.
[0234] Volume adjustment means:
[0235] If the detection means detects the approach of an emergency vehicle, the volume of any audio playing in the car (e.g., music, audiobooks, etc.) will be automatically reduced. This is achieved using a Bluetooth connection.
[0236] Means of notification:
[0237] A specific announcement is generated for approaching emergency vehicles and notified to the user through the in-car speakers, including the direction and distance of the emergency vehicle.
[0238] Emotion Engine:
[0239] It is an engine that analyzes the user's emotional state. It analyzes facial recognition data or voice tone to understand the user's emotional state, such as whether they are stressed or relaxed.
[0240] Program processing
[0241] Sound pickup and data transmission
[0242] Terminal: Microphones installed at the four corners of the vehicle collect sound in real time and temporarily store it in a buffer. The contents of the buffer are packetized at regular intervals and sent to the server.
[0243] Siren sound analysis
[0244] Server: Apply machine learning algorithms to the received sound data to identify specific frequency bands and patterns of siren sounds.
[0245] Detecting approaching emergency vehicles
[0246] Server: If a siren sound is detected, the result is fed back to the device.
[0247] Terminal: Based on the analysis results fed back, the direction and distance of the sound source are determined using the difference in sound arrival time from multiple microphones.
[0248] Automatic volume adjustment
[0249] Terminal: Based on the detection results, it sends an instruction via Bluetooth to reduce the volume of the sound source inside the car to a certain percentage (e.g., 50%).
[0250] Emotional state analysis
[0251] Device: The emotion engine uses facial recognition data captured by the in-car camera and the tone of the user's voice collected by the microphone to identify the user's emotional state.
[0252] Generate and play notifications
[0253] Server: Obtains the location and direction of emergency vehicles from the navigation system and generates specific announcement text.
[0254] Terminal: Based on the analysis results of the emotion engine, the content and tone of the announcement are adjusted according to the user's emotional state.
[0255] Emotion-based notifications
[0256] Terminal: The generated voice announcement is played over the car's speakers. Depending on the user's emotional state, the announcement can be calm or emphasize urgency.
[0257] Specific examples
[0258] 1. Sound collection means:
[0259] The device collects sound in real time using microphones installed at the four corners of the vehicle and stores it in a buffer.
[0260] 2. Analysis method:
[0261] The server analyzes the sound data in the buffer using a machine learning algorithm to identify the sound of a siren.
[0262] 3. Detection methods:
[0263] Based on the analysis results from the server, the device calculates the direction and distance of the sound source.
[0264] 4. Volume adjustment means:
[0265] The server detects the approach of an emergency vehicle and the device sends a command via Bluetooth connection to reduce the music volume to 50%.
[0266] 5. Emotion Engine:
[0267] The device recognizes the user's face using the onboard camera, and the emotion engine analyzes the data to understand the user's emotional state.
[0268] 6. Means of notification:
[0269] The server uses data from the navigation system to obtain the location information of the emergency vehicle and generates specific instructions.
[0270] The tone and content of the announcements are adjusted based on the results of the emotion engine. For example, if the user is feeling stressed, the announcements will be made in a calmer tone.
[0271] The terminal generates a voice announcement that is played over the car's speakers, prompting the user to take specific action (e.g., "There is an ambulance approaching 100 meters ahead from the right, so pull over to the side of the road").
[0272] summary
[0273] This system allows even those listening to loud music in their cars to quickly and reliably detect the approach of an emergency vehicle and provide appropriate notifications based on the user's emotional state, allowing them to safely take appropriate action and contributing to traffic management and accident prevention.
[0274] The processing flow will be explained below.
[0275] Step 1:
[0276] Terminal: Sound is collected by multiple microphones installed at the four corners of the vehicle. The collected sound data is temporarily stored in a buffer.
[0277] Specific operation: Sound data is sampled through the microphone interface, for example, every second, and stored in a buffer. Data from each microphone is stored in a separate buffer.
[0278] Step 2:
[0279] Terminal: Collected sound data is packetized at regular intervals (e.g., every second) and sent to the server.
[0280] Specific operation: The sound data is divided into packets and sent to the server using a network protocol. After sending, the buffer is cleared and waits for the next sample.
[0281] Step 3:
[0282] Server: Analyzes the received sound data in real time, using AI to identify specific frequency bands and patterns of siren sounds.
[0283] How it works: A machine learning model (e.g., a convolutional neural network) is applied to analyze sound data and extract features. If a specific pattern is detected, it is identified as a siren sound.
[0284] Step 4:
[0285] Server: If a siren sound is detected, the result is fed back to the device.
[0286] Specific operation: A data packet containing the analysis results is generated and sent to the terminal via the network.
[0287] Step 5:
[0288] Terminal: Based on the analysis results received from the server, the terminal calculates the difference in sound arrival time from multiple microphones and determines the direction and distance of the sound source.
[0289] Specific operation: An algorithm is applied that uses the difference in sound arrival time to calculate the direction and distance of the sound source using techniques such as triangulation.
[0290] Step 6:
[0291] Terminal: If an approaching emergency vehicle is detected, a command to reduce the volume of the in-car speakers to a certain percentage (e.g., 50%) is sent via Bluetooth.
[0292] Specific behavior: Uses the Bluetooth speaker's volume control protocol to send a command to decrease the volume of the currently playing music or voice by the specified percentage.
[0293] Step 7:
[0294] On the device: The emotion engine analyzes the user's facial recognition data from the in-car camera or the user's tone of voice collected by the microphone.
[0295] Specific behavior: Apply facial recognition and voice analysis algorithms to identify the user's emotional state (e.g., stress, relaxed).
[0296] Step 8:
[0297] Server: Works with the navigation system to obtain the current location and direction of emergency vehicles.
[0298] Specific operation: Calls the navigation system's API to obtain the current GPS location and direction of the emergency vehicle. Based on this information, it generates an announcement for the next step.
[0299] Step 9:
[0300] Terminal: Generates specific voice announcements based on the emergency vehicle location information received from the server and the analysis results of the emotion engine.
[0301] What it does: It uses a text generation engine to create an appropriate announcement, then uses a text-to-speech engine to generate an audio file, adjusting the tone depending on the user's emotional state.
[0302] Step 10:
[0303] Terminal: The generated voice announcement is played over the vehicle's speakers, providing specific instructions to the user.
[0304] Specific behavior: The created audio file is added to the playback queue of the car speakers and begins playing immediately. Based on the user's emotional state, the announcement is made in a calm manner or with an emphasis on urgency.
[0305] Step 11:
[0306] User: Follow the voice announcement and take appropriate action promptly.
[0307] Specific actions: For example, "An ambulance is approaching from the right, 100 meters ahead. Please move to the shoulder," and follow the instructions to safely move the vehicle to the shoulder.
[0308] Example 2
[0309] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0310] Conventional vehicle systems have the problem of not being able to take into account the emotional state of the user when notifying them of an approaching emergency vehicle. This can make it difficult for users to take appropriate action in stressful situations, potentially delaying their response to the emergency vehicle. Another issue is that the volume of the audio played in the vehicle must be adjusted manually, which requires the user to divert their attention.
[0311] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server is a processing device for analyzing sound data collected from a sound collection device, and includes an analysis device that monitors surrounding sounds based on the generated analysis results, a detection device that detects the approach of an emergency vehicle based on the analysis results, and an emotion recognition device that collects and analyzes the emotional state of the user. This makes it possible to quickly and accurately detect the approach of an emergency vehicle and to provide an appropriate notification according to the emotional state of the user.
[0312] A "sound collection device" is a device installed at the four corners of a vehicle to collect surrounding sounds.
[0313] The "analysis device" is a device that analyzes collected sound data and monitors surrounding sounds based on the generated analysis results.
[0314] The "detection device" is a device for detecting the approach of an emergency vehicle based on the analysis results generated by the analysis device.
[0315] The "volume adjustment device" is a device that automatically reduces the volume of the sound source being played inside the vehicle when the approach of an emergency vehicle is detected by the detection device.
[0316] A "notification device" is a device that notifies in-vehicle users of the location and approach of an emergency vehicle.
[0317] An "emotion recognition device" is a device for collecting and analyzing a user's emotional state.
[0318] The present invention relates to an in-vehicle emergency vehicle approach detection system and a notification system that takes into account the emotional state of the user. The system is composed of sound collection devices installed at the four corners of the vehicle, an analysis device that analyzes the collected sound data, a detection device that detects the approach of an emergency vehicle based on the analysis results, a volume adjustment device that automatically lowers the volume of the sound source being played in the vehicle when the approach of the emergency vehicle is detected, a notification device that notifies the in-vehicle user of the location and approach of the emergency vehicle, and an emotion recognition device that collects and analyzes the user's emotional state.
[0319] System configuration
[0320] Sound pickup device
[0321] The device includes microphones installed at the four corners of the vehicle, which collect ambient sounds in real time and temporarily store them in a buffer. At regular intervals, the contents of the buffer are sent to a server as data packets.
[0322] analysis device
[0323] The server analyzes the sound data received from the device using a generative AI model (e.g., TensorFlow) to identify specific frequency bands and sound patterns and identify the sound of a siren.
[0324] Detection device
[0325] When the server detects a siren, it sends the result back to the device, which then uses the difference in sound arrival times from multiple microphones to calculate the direction and distance of the sound source.
[0326] volume adjustment device
[0327] Based on the detection results, the device will automatically reduce the volume of in-car audio sources (e.g., radio, music player) to a certain percentage (e.g., 50%) via Bluetooth.
[0328] emotion recognition device
[0329] The device uses an onboard camera and microphone to collect facial recognition data and voice tone, and then uses an emotion engine to analyze the user's emotional state, determining whether the user is stressed or relaxed.
[0330] notification device
[0331] The server obtains the location and heading of the emergency vehicle from the navigation system and uses a generative AI model to generate a specific announcement. This announcement includes tone and content based on the user's emotional state. The device plays the generated voice announcement over the car's speakers, encouraging the user to take specific action.
[0332] Specific examples
[0333] Here is a concrete example of how the system works:
[0334] 1. Sound pickup devices are installed at the four corners of the vehicle to collect sound in real time, and this data is periodically sent to a server.
[0335] 2. The server analyzes the sound data and uses a generative AI model to identify specific frequency bands and detect siren sounds.
[0336] 3. The server sends the analysis results back to the device, which then calculates the direction and distance of the sound source.
[0337] 4. If an approaching emergency vehicle is detected, the device will send a command via Bluetooth to halve the volume of sound sources inside the vehicle.
[0338] 5. The terminal analyzes the user's emotional state using an emotion engine, and the server generates a specific announcement based on the navigation data.
[0339] 6. The generated announcement is played by the terminal through the in-car speaker, and an announcement such as "There is an ambulance approaching 100 meters ahead from the right, so please pull over to the side of the road" is made.
[0340] Prompt Sentence Examples
[0341] Please generate a natural language document that describes an in-vehicle emergency vehicle approach detection system and a notification system that takes into account the user's emotional state. Please provide a detailed description of the processing flow, including specific hardware and software, including sound collection devices, analysis devices, detection devices, volume control devices, notification devices, and emotion recognition devices.
[0342] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0343] Step 1: Recording and transmitting data
[0344] Terminal: Microphones installed at the four corners of the vehicle collect surrounding sounds in real time and temporarily store the data in a buffer. The sound data in the buffer is divided into packets at regular intervals (e.g., every 10 seconds) and sent to the server.
[0345] Input: Sound data collected from microphones at the four corners of the vehicle
[0346] Output: Sound data sent to the server in data packet format
[0347] Step 2: Analyzing the siren sound
[0348] Server: Applying a generative AI model (e.g., TensorFlow) to the received sound data to identify specific frequency bands and patterns of siren sounds. The results of this analysis are passed on to the next processing step.
[0349] Input: Sound data sent from the device
[0350] Output: Siren sound analysis results (specific frequency bands and patterns)
[0351] Step 3: Detect approaching emergency vehicles
[0352] Server: When the siren sound is identified, the server feeds back the result to the device. Based on the feedback, the device calculates the difference in sound arrival time from multiple microphones and specifically calculates the direction and distance of the sound source.
[0353] Input: Analysis result of siren sound
[0354] Output: direction and distance of sound source
[0355] Step 4: Automatically adjust the volume
[0356] Device: When an approaching emergency vehicle is detected, the device sends a command via Bluetooth to automatically lower the volume of the currently playing audio source (e.g., radio, music player), for example, by 50%.
[0357] Input: Proximity information based on the direction and distance of the sound source
[0358] Output: Volume down command via Bluetooth
[0359] Step 5: Analyze your emotional state
[0360] Device: Analyzes the user's facial recognition data acquired by the in-car camera and the tone of voice collected by the microphone, and uses a generative AI model to determine the user's emotional state (e.g., stressed, relaxed).
[0361] Input: User's facial recognition data and tone of voice
[0362] Output: User's emotional state
[0363] Step 6: Generate and play notifications
[0364] Server: Obtains the location and heading of the emergency vehicle from the navigation system, and uses a generative AI model to create a specific announcement based on that information. This announcement is adjusted according to the user's emotional state.
[0365] Input: Emergency vehicle location information, user emotional state
[0366] Output: Specific announcement text
[0367] Step 7: Emotional Notifications
[0368] Terminal: Receives announcements generated by the server and plays them in an appropriate tone through the car's speakers. The content of the notification is adjusted according to the user's emotional state, encouraging specific countermeasures.
[0369] Input: Specific announcement text
[0370] Output: Voice notification through car speakers
[0371] (Application example 2)
[0372] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0373] For self-driving vehicles, it is important to quickly and accurately detect the approach of an emergency vehicle and notify the user appropriately, but conventional systems have difficulty accurately detecting the direction and distance of an emergency vehicle and lack the ingenuity to enable the user to respond calmly according to the situation.In addition, there is no notification system that takes into account the user's emotional state, which makes it difficult for the user to respond appropriately if they are feeling stressed.
[0374] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes sound collection means installed at the four corners of the vehicle, analysis means for analyzing collected sound data and monitoring surrounding sounds, detection means for detecting the approach of an emergency vehicle based on the analysis results, volume adjustment means for automatically lowering the volume of the sound source being played in the vehicle when the detection means detects the approach of an emergency vehicle, notification means for notifying users in the vehicle of the location and approach of the emergency vehicle, an emotion engine for analyzing the emotional state of the user, and means for adjusting the content and tone of the notification based on the analysis results of the emotion engine. This makes it possible to quickly and accurately detect the approach of an emergency vehicle and to notify the user appropriately according to their emotional state.
[0375] The "sound collection means" is a device installed at the four corners of the vehicle to collect surrounding sounds.
[0376] The "analysis means" is a means for analyzing collected sound data and identifying the siren sound specific to emergency vehicles.
[0377] The "detection means" is a means for detecting the approach of an emergency vehicle based on the analysis results obtained by the analysis means.
[0378] The "volume adjustment means" is a means for automatically lowering the volume of the sound source being reproduced inside the vehicle when the detection means detects the approach of an emergency vehicle.
[0379] The "notification means" is a means for informing the vehicle occupants of the location and approach of an emergency vehicle.
[0380] "Emotion Engine" is an engine for analyzing the user's emotional state, such as analyzing facial recognition data or voice tone.
[0381] "Calculation by the analysis means" refers to calculation of the approaching direction and distance of the emergency vehicle based on the sound data collected from the sound collection means.
[0382] A "user's emotional state" is the emotional state that a user is feeling, identified based on a facial image or tone of voice.
[0383] The "means for adjusting the content and tone of the notification based on the analysis results" refers to means for changing the content and tone of the voice of the generated notification according to the analysis results of the emotional state by the emotion engine.
[0384] A system for realizing this invention includes sound collection means installed at the four corners of a vehicle, analysis means for analyzing collected sound data and monitoring surrounding sounds, detection means for detecting the approach of an emergency vehicle based on the analysis results, volume adjustment means for automatically lowering the volume of the sound source being played inside the vehicle when the detection means detects the approach of an emergency vehicle, notification means for notifying users inside the vehicle of the location and approach of the emergency vehicle, an emotion engine for analyzing the emotional state of the user, and means for adjusting the content and tone of the notification based on the analysis results of the emotion engine.
[0385] System Components
[0386] Sound collection means
[0387] The sound collection means consists of microphones installed at the four corners of the vehicle, which can effectively capture sounds from all directions (360 degrees).
[0388] Analysis means
[0389] The analysis means analyzes the sound data from the sound collection means in real time and identifies the siren sounds specific to emergency vehicles. This analysis uses machine learning algorithms to detect specific frequency bands and sound patterns of siren sounds. The main software used is Python's "librosa" and "keras."
[0390] Detection Method
[0391] The detection means detects the approach of an emergency vehicle based on the analysis results obtained from the analysis means, and calculates the direction and distance of the sound source based on the difference in sound arrival times detected by the multiple microphones.
[0392] Volume adjustment means
[0393] The volume adjustment means uses a Bluetooth connection to automatically reduce the volume of audio sources (such as music or audiobooks) in the car.
[0394] Notification means
[0395] The notification means generates a specific announcement when an emergency vehicle is approaching and notifies the user through an in-vehicle speaker, the announcement including the direction and distance of the emergency vehicle.
[0396] Emotion Engine
[0397] The emotion engine analyzes the user's emotional state by analyzing facial recognition data and voice tone. The main software used is Python's "opencv" and "keras."
[0398] System Operation
[0399] The server analyzes the sound data sent from the sound collection means and identifies the siren sound. After the analysis, it calculates the direction and distance of the emergency vehicle and feeds the results back to the device. The device then captures an image of the user's face using the in-car camera and analyzes it using the emotion engine. Based on the analysis results, the notification means announces the approach of an emergency vehicle.
[0400] The notification means obtains emergency vehicle location information from the navigation system and generates a specific announcement, the tone and content of which are adjusted according to the user's emotional state, and is played through the vehicle speakers.
[0401] Specific examples
[0402] As a specific example, the user collects sound in real time using microphone devices installed at the four corners of the vehicle and sends the data to a server. The server analyzes the siren sound and calculates the direction and distance of the emergency vehicle. The results are fed back to the device, which generates an announcement based on the emergency vehicle's location information. Furthermore, the device captures an image of the user's face using an in-car camera, analyzes it using an emotion engine, and the notification means adjusts the content and tone of the announcement.
[0403] Prompt Sentence Examples
[0404] "When providing voice notifications for approaching emergency vehicles in an autonomous vehicle, please generate notification messages that take into account the emotional state of the occupants. For example, if an emergency vehicle is approaching from the right direction and 100 meters ahead, the notification message should be delivered in a calm tone if the occupant is feeling stressed, and in a standard tone if the occupant is relaxed."
[0405] In this way, it is possible to quickly and accurately detect the approach of an emergency vehicle and notify the user according to their emotional state.
[0406] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0407] Step 1:
[0408] The device collects sound in real time using microphones installed at the four corners of the vehicle. The sound data is temporarily stored in a buffer. This makes it possible to effectively capture sound from all directions (360 degrees). The input is the sound picked up from around the vehicle, and the output is the sound data stored in the buffer.
[0409] Step 2:
[0410] The terminal packetizes the contents of the buffer at regular intervals and sends the sound data to the server. The server receives the sound data and passes it to the analysis means. The input is the sound data stored in the buffer, and the output is the sound data uploaded to the server.
[0411] Step 3:
[0412] The server applies machine learning algorithms to the received sound data to identify specific frequency bands and patterns of siren sounds. The software used is the Python libraries "librosa" and "keras." The input is the sound data uploaded to the server, and the output is the determination of whether or not a siren is being heard.
[0413] Step 4:
[0414] If the server detects a siren, it feeds the result back to the device. It then calculates the direction and distance of the emergency vehicle based on the analysis results. The input is the result of the siren detection, and the output is data on the direction and distance of the emergency vehicle.
[0415] Step 5:
[0416] Based on the detection results, the device sends an instruction via Bluetooth to reduce the volume of the sound source inside the car by a certain percentage (e.g., 50%). The input is data about the direction and distance of the emergency vehicle, and the output is a volume adjustment instruction. Specifically, the device sends an instruction to the car's audio system using Bluetooth.
[0417] Step 6:
[0418] The device captures the user's facial image using an onboard camera and analyzes it with an emotion engine. The software used is Python's "opencv" and "keras." The input is the captured facial image, and the output is the user's emotional state.
[0419] Step 7:
[0420] The server obtains the location and direction of the emergency vehicle from the navigation system and generates a specific announcement. The input is the data on the direction and distance of the emergency vehicle and the user's emotional state, and the output is the generated announcement.
[0421] Step 8:
[0422] The device adjusts the content and tone of the announcement based on the analysis results of the emotion engine according to the user's emotional state. For example, if the user is feeling stressed, the announcement will be made in a calmer tone. The input is the generated announcement text, and the output is the adjusted announcement text.
[0423] Step 9:
[0424] The terminal plays the generated voice announcement over the in-car speaker, prompting the user to take specific action (e.g., "There is an ambulance approaching 100 meters ahead from the right, please pull over to the side of the road.") The input is the adjusted announcement text, and the output is a voice notification from the in-car speaker.
[0425] In this way, through each step, it is possible to quickly and accurately detect the approach of an emergency vehicle and notify the user appropriately according to their emotional state.
[0426] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0427] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0428] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0429] [Second embodiment]
[0430] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0431] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0432] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0433] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0434] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0435] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0436] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0437] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0438] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0439] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0440] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0441] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0442] As an example of the present invention, an emergency vehicle approach detection system for a vehicle will be described. This system is composed of a sound collection means, an analysis means, a detection means, a volume adjustment means, and a notification means, and quickly and accurately detects the approach of an emergency vehicle and notifies the user.
[0443] overview
[0444] Recording Method:
[0445] Microphone devices installed at the four corners of the vehicle collect ambient sounds, effectively capturing sounds coming from all directions in 360 degrees.
[0446] Analysis method:
[0447] The system analyzes sound data from sound collection devices in real time to identify the siren sounds specific to emergency vehicles, using machine learning algorithms to detect specific frequency bands and sound patterns of siren sounds.
[0448] Detection Methods:
[0449] Based on the analysis results obtained by the analysis means, the system detects the approach of an emergency vehicle and calculates the direction and distance of the sound source based on the difference in arrival time of the sounds from multiple microphones.
[0450] Volume adjustment means:
[0451] If the detection means detects the approach of an emergency vehicle, the volume of any audio source (e.g., music, audiobooks, etc.) being played in the vehicle will be automatically reduced via Bluetooth connection.
[0452] Means of notification:
[0453] Specific announcements will be made over the vehicle's speakers to passengers informing them of the location and approach of emergency vehicles.
[0454] Program processing
[0455] Sound pickup and data transmission
[0456] Terminal: Collects sound data in real time from microphones installed at the four corners of the vehicle and temporarily stores it in a buffer.
[0457] Server: Receives sound data periodically sent from the terminal.
[0458] Siren sound analysis
[0459] Server: Analyzes the received sound data and applies machine learning algorithms to detect frequency bands and sound patterns to identify the sound of an emergency vehicle siren.
[0460] Terminal: Receives the analysis results from the server and prepares any necessary next steps.
[0461] Detecting approaching emergency vehicles
[0462] Terminal: Using data from multiple microphones, a sound source localization algorithm calculates the difference in sound arrival time and determines the direction and distance of the emergency vehicle.
[0463] Automatic volume adjustment
[0464] Device: Based on the detection results, it sends a command via Bluetooth connection to reduce the volume of the sound source inside the car to 50%.
[0465] Generate and play notifications
[0466] Server: Obtains the current location and direction of the emergency vehicle based on data from the navigation system and generates an appropriate announcement.
[0467] Terminal: The generated voice announcement is played over the car's speakers.
[0468] Specific examples
[0469] 1. Sound collection means:
[0470] The device collects sound in real time using microphones installed at the four corners of the vehicle and stores it in a buffer.
[0471] 2. Analysis method:
[0472] The server analyzes the sound data in the buffer using a machine learning algorithm to identify the sound of a siren.
[0473] 3. Detection methods:
[0474] Based on the analysis results from the server, the device calculates the direction and distance of the sound source.
[0475] 4. Volume adjustment means:
[0476] The server detects the approach of an emergency vehicle and the device sends a command via Bluetooth connection to reduce the music volume to 50%.
[0477] 5. Means of notification:
[0478] The server uses data from the navigation system to obtain the location information of the emergency vehicle and generates specific instructions.
[0479] The terminal generates a voice announcement that is played over the car's speakers, prompting the user to take specific action (e.g., there is an emergency vehicle approaching 100 meters ahead from the right; move to the shoulder of the road).
[0480] summary
[0481] This system allows even passengers who are enjoying loud music in their cars to quickly and reliably detect approaching emergency vehicles and respond safely, thereby reducing the risk of accidents and supporting smooth traffic management.
[0482] The processing flow will be explained below.
[0483] Step 1:
[0484] Terminal: Sound is collected by multiple microphones installed at the four corners of the vehicle. The collected sound data is temporarily stored in a buffer.
[0485] Specific operation: Sound data is sampled through the microphone interface, for example, every second, and stored in a buffer. Data from each microphone is stored in a separate buffer.
[0486] Step 2:
[0487] Terminal: Collected sound data is packetized at regular intervals (e.g., every second) and sent to the server.
[0488] Specific operation: The sound data is divided into packets and sent to the server using a network protocol. After sending, the buffer is cleared and waits for the next sample.
[0489] Step 3:
[0490] Server: Analyzes the received sound data in real time, using AI to identify specific frequency bands and patterns of siren sounds.
[0491] How it works: A machine learning model (e.g., a convolutional neural network) is applied to analyze sound data and extract features. If a specific pattern is detected, it is identified as a siren sound.
[0492] Step 4:
[0493] Server: If a siren sound is detected, the result is fed back to the device.
[0494] Specific operation: A data packet containing the analysis results is generated and sent to the terminal via the network.
[0495] Step 5:
[0496] Terminal: Based on the analysis results received from the server, the terminal calculates the difference in sound arrival time from multiple microphones and determines the direction and distance of the sound source.
[0497] Specific operation: An algorithm is applied that uses the difference in sound arrival time to calculate the direction and distance of the sound source using techniques such as triangulation.
[0498] Step 6:
[0499] Terminal: If an approaching emergency vehicle is detected, a command to reduce the volume of the in-car speakers to a certain percentage (e.g., 50%) is sent via Bluetooth.
[0500] Specific behavior: Uses the Bluetooth speaker's volume control protocol to send a command to decrease the volume of the currently playing music or voice by the specified percentage.
[0501] Step 7:
[0502] Server: Works with the navigation system to obtain the current location and direction of emergency vehicles.
[0503] Specific operation: Calls the navigation system's API to obtain the current GPS location and direction of the emergency vehicle. Based on this information, it generates an announcement for the next step.
[0504] Step 8:
[0505] Terminal: Generates specific voice announcements based on location information received from the server.
[0506] What it does: Creates the appropriate announcement using a text generation engine and generates an audio file using a text-to-speech engine.
[0507] Step 9:
[0508] Terminal: The generated voice announcement is played over the vehicle's speakers, providing specific instructions to the user.
[0509] Specific operation: The created audio file is added to the playback queue of the car speakers and begins playing immediately.
[0510] Step 10:
[0511] User: Follow the voice announcement and take appropriate action promptly.
[0512] Specific actions: For example, "An ambulance is approaching from the right, 100 meters ahead. Please move to the shoulder," and follow the instructions to safely move the vehicle to the shoulder.
[0513] Example 1
[0514] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0515] Conventional vehicle emergency vehicle approach detection systems lack the functionality to quickly detect approaching emergency vehicles and notify the driver. In particular, drivers listening to loud music or audiobooks in the vehicle may not notice the approach of an emergency vehicle, increasing the risk of an accident. The objective of this invention is to solve this problem by providing a system that quickly and accurately detects the approach of an emergency vehicle and notifies the driver in an appropriate manner.
[0516] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0517] In this invention, the server includes a sound collection means, a sound data analysis means, an emergency vehicle detection means, an automatic volume adjustment means, a notification means, and a notification generation means. This makes it possible to analyze in real time the sound data collected by the sound collection means installed at the four corners of the vehicle and quickly detect the approach of an emergency vehicle. Furthermore, when an emergency vehicle approaches, the server can automatically lower the volume inside the vehicle and generate an appropriate announcement to notify the driver.
[0518] The "sound collection means" is a device installed at each of the four corners of the vehicle that collects surrounding sounds in real time.
[0519] The "sound data analysis means" is a function that analyzes collected sound data, detects specific frequency bands and patterns, and identifies the siren sounds of emergency vehicles.
[0520] The "emergency vehicle detection means" is a function that detects the approach of an emergency vehicle based on the analysis results obtained by the sound data analysis means.
[0521] The "automatic volume adjustment means" is a function that automatically reduces the volume of the sound source being played inside the vehicle when the emergency vehicle detection means detects the approach of an emergency vehicle.
[0522] The "notification means" is a function for informing the vehicle occupants of the location and approach of an emergency vehicle.
[0523] The "notification generating means" is a function that obtains the location of emergency vehicles based on data from the navigation system and generates an appropriate announcement.
[0524] The present invention relates to a system for detecting approaching emergency vehicles in a vehicle, and specific embodiments thereof are described below. This system is composed of a sound collection means, a sound data analysis means, an emergency vehicle detection means, an automatic volume adjustment means, a notification means, and a notification generation means. By combining these means, it is possible to quickly and accurately detect the approach of an emergency vehicle and notify occupants in the vehicle.
[0525] Sound collection means
[0526] Microphones installed at the four corners of the vehicle act as sound collectors. These microphones collect ambient sounds in real time and temporarily store the data in a buffer. For example, when the vehicle approaches an intersection, each microphone collects the surrounding sounds.
[0527] Sound data analysis method
[0528] The server receives audio data periodically sent from the device. It analyzes the received audio data using a machine learning algorithm to detect specific frequency bands and sound patterns and identify the sound of an emergency vehicle siren. The server then sends the analysis results to the device and prepares for the next process.
[0529] Emergency vehicle detection means
[0530] Based on the analysis results sent from the server, the device uses data from multiple microphones and a sound source localization algorithm to calculate the direction and distance of the emergency vehicle. For example, it can determine that an emergency vehicle is approaching 100 meters behind and to the left.
[0531] Automatic volume adjustment means
[0532] If the device detects an approaching emergency vehicle, it will send a command over the Bluetooth connection to reduce the volume of the in-car audio source (e.g., music) to 50%, allowing the in-car audio system to automatically reduce the current volume.
[0533] Notification means and notification generation means
[0534] The server obtains data on the current location and direction of the emergency vehicle from the navigation system and generates an appropriate announcement for the user. The device receives this announcement and plays it through the car's speakers. For example, it can provide specific instructions to the user, such as "Police vehicle approaching 100 meters behind you on the left, please be careful."
[0535] Prompt Sentence Examples
[0536] Here are some examples of prompts:
[0537] "Please explain the process by which microphones installed at the four corners of the vehicle collect ambient sounds in real time and store that data in a buffer."
[0538] "Please explain in detail how the server uses a machine learning algorithm to analyze the sound data received from the device and identify the sound of a siren."
[0539] "Please explain in detail the algorithm that allows the device to use data from multiple microphones to calculate the direction and distance of emergency vehicles."
[0540] "Please explain the steps to reduce the volume of in-car audio sources to 50% when the device detects an approaching emergency vehicle."
[0541] "Please explain the procedure for the server to generate an announcement to notify the user of the approach of an emergency vehicle and send it to the car speaker."
[0542] In this way, the system of the present invention can quickly detect the approach of an emergency vehicle, provide the user with the necessary information, and encourage appropriate action.
[0543] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0544] Step 1: Sound collection and data transmission
[0545] Terminal
[0546] Input: Microphones installed at the four corners of the vehicle collect surrounding sounds in real time.
[0547] Data processing / data calculation: Collected sound data is temporarily stored in a buffer and prepared to be sent to the server at regular intervals.
[0548] Output: Sends the sound data stored in the temporary buffer to the server.
[0549] Specific operation: For example, when a vehicle is driving and approaches an intersection, the microphone collects sounds of nearby pedestrians and other vehicles and sends the data to a server at regular intervals.
[0550] Step 2: Analyzing the siren sound
[0551] server
[0552] Input: Sound data sent periodically from the device.
[0553] Data processing / data calculation: Applying machine learning algorithms to analyze sound data and detect specific frequency bands and sound patterns to identify the sound of an emergency vehicle siren.
[0554] Output: As an analysis result, information is generated as to whether or not an emergency vehicle siren was detected.
[0555] Specific operation: The server analyzes the sound data received from the device, and if it detects a specific frequency pattern between 700 and 1500 Hz, it identifies it as the sound of an emergency vehicle siren. The analysis result is returned to the device.
[0556] Step 3: Detect approaching emergency vehicles
[0557] Terminal
[0558] Input: Analysis results sent from the server.
[0559] Data processing / data calculation: Based on the difference in arrival time of sound data obtained from multiple microphones, a sound source localization algorithm is used to calculate the direction and distance of the emergency vehicle.
[0560] Output: Information on the approaching direction and distance of the emergency vehicle.
[0561] Specific operation: For example, identify that an emergency vehicle is approaching 100 meters behind and to the left. This information is used in the next step.
[0562] Step 4: Automatically adjust the volume
[0563] Terminal
[0564] Input: Information about the approaching direction and distance of the emergency vehicle, and information about the audio source currently being played.
[0565] Data processing / data calculation: If an approaching emergency vehicle is detected, a command is generated to reduce the volume of the sound source (e.g., music) inside the vehicle to 50%.
[0566] Output: A command to decrease the volume over the Bluetooth connection.
[0567] Specific behavior: For example, if music is playing in the car, the volume will be automatically reduced to 50%. This operation is performed via Bluetooth connection.
[0568] Step 5: Generate and play notifications
[0569] server
[0570] Input: Emergency vehicle approach direction and distance information, and location data from the navigation system.
[0571] Data processing / data calculation: Based on data from the navigation system, appropriate announcements are generated for users.
[0572] Output: The generated announcement.
[0573] Specific actions: For example, generate specific instructions such as "A police vehicle is approaching 100 meters behind you on the left. Be careful."
[0574] Terminal
[0575] Input: Server-generated announcement.
[0576] Data processing / data calculation: Generates commands to play announcements over the car speakers.
[0577] Output: Announcement command to the car speaker.
[0578] Specific operation: The generated voice announcement is played over the car's speakers to notify the user of the approaching emergency vehicle and how to respond. For example, it may say, "Siren approaching, emergency vehicle on the left."
[0579] (Application example 1)
[0580] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0581] Conventional vehicle emergency vehicle notification systems alert users to approaching emergency vehicles by lowering the volume inside the vehicle, but lack the technology for use outside the vehicle or on mobile devices. Furthermore, while it is necessary to quickly and accurately notify users of the direction and distance of approaching emergency vehicles visually and audibly, no appropriate solution for this purpose exists. Therefore, there is a need for a comprehensive emergency vehicle notification system that allows users to respond safely both inside and outside the vehicle.
[0582] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0583] In this invention, the server includes a sound collection means, an analysis means, a detection means, a volume adjustment means, and a notification means. This allows the server to quickly and accurately detect the approach of an emergency vehicle by collecting surrounding sounds, even when using a vehicle or a mobile device, automatically lower the volume of the sound source being played, and effectively notify the user of the location and approach of the emergency vehicle. Specifically, the analysis means calculates the direction and distance of the approaching emergency vehicle, and the information is provided via voice and screen display, allowing the user to quickly take appropriate action.
[0584] A "sound collection means" is a device installed to collect surrounding sounds.
[0585] An "analysis means" is a device or algorithm for analyzing collected sound data and identifying specific sound patterns.
[0586] The "detection means" is a device or function for detecting the approach of a specific sound source, for example, an emergency vehicle, based on the analysis result of the analysis means.
[0587] The "volume adjustment means" is a device or function that automatically adjusts the volume of the sound source being played back when the detection means detects the approach of an emergency vehicle.
[0588] "Notification means" refers to a device or function for informing users of the location and approach of an emergency vehicle.
[0589] A "mobile terminal" is a communication device or computer device that can be carried by a user.
[0590] The "audio output means" is a device for conveying the audio information generated by the notification means to the user.
[0591] "Screen display" refers to the function of displaying textual or graphic information on a mobile terminal or other display device.
[0592] The present invention provides a system that uses a vehicle or a mobile terminal to quickly detect the approach of an emergency vehicle and automatically reduce the volume of the sound source being played. Specific embodiments for carrying out the present invention are described below.
[0593] overview
[0594] This system consists of a sound collection means, an analysis means, a detection means, a volume adjustment means, and a notification means. The sound collection means is a microphone installed in a vehicle or a mobile device that collects surrounding sounds. The analysis means analyzes the collected sound data and identifies specific sound patterns, such as siren sounds. This analysis uses a machine learning algorithm (e.g., a model built with Keras). The detection means detects the approach of an emergency vehicle based on the analysis results. The volume adjustment means automatically lowers the volume of the sound being played when the detection means detects the approach of an emergency vehicle. The notification means effectively notifies the user of the location and approach of the emergency vehicle.
[0595] Program processing
[0596] Sound pickup and data transmission
[0597] The device collects ambient sounds through a microphone, and the collected sound data is temporarily stored in a buffer and processed by an analysis means.
[0598] Siren sound analysis
[0599] The server receives the sound data in the buffer and applies a machine learning algorithm to identify siren sounds. Keras is used as an analysis tool to analyze the collected sound data.
[0600] Detecting approaching emergency vehicles
[0601] The detection means detects the approach of an emergency vehicle based on the results of the analysis means, and calculates the direction and distance of the sound source from the collected sound data.
[0602] Volume adjustment
[0603] The volume control means automatically reduces the volume of the audio source being played when detected, and sends a volume control command via Bluetooth or another communication means.
[0604] Generate and play notifications
[0605] The server generates an appropriate announcement based on the location information of the emergency vehicle, and the notification means notifies the user of this announcement both by voice and by display on the screen.
[0606] Specific examples
[0607] While the user is driving, the device (smartphone) collects surrounding sounds using a built-in microphone and sends them to a server. The server analyzes the sound data in real time, and if a siren sound is detected, it automatically lowers the volume of the smartphone via Bluetooth and notifies the user of the location of an emergency vehicle via a display and voice message, allowing the user to take prompt and appropriate action.
[0608] Prompt Sentence Examples
[0609] "We are developing an emergency vehicle approach detection system for vehicles as a smartphone application. Please generate a Python program that follows the requirements below:
[0610] Built-in microphone for sound capture
[0611] Uses machine learning models to analyze and detect siren sounds in real time
[0612] Volume adjustment and notification functions implemented
[0613] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0614] Step 1:
[0615] The device collects ambient sounds. While the user is driving the vehicle, the device's microphone collects audio data in real time and temporarily stores it in a buffer. The input here is the ambient sounds, and the output is the audio data in the buffer.
[0616] Step 2:
[0617] The device sends the sound data in the buffer to the server. The collected sound data is transferred to the server at regular intervals. The input here is the sound data in the buffer, and the output is the sound data sent to the server.
[0618] Step 3:
[0619] The server analyzes the received sound data. The server uses a machine learning model (a model using Keras) to analyze the sound data and detect siren sounds. The input here is the sound data sent to the server, and the output is the analysis result.
[0620] Step 4:
[0621] The server detects emergency vehicles based on the analysis results. If a siren sound is detected, the server detects the approach of an emergency vehicle and calculates its direction and distance. The input here is the analysis result, and the output is the direction and distance information of the emergency vehicle.
[0622] Step 5:
[0623] The server sends a volume adjustment command to the device. Based on the detection result, the server sends an instruction to the device to lower the volume of the sound being played. The input here is the direction and distance information of the emergency vehicle, and the output is the volume adjustment command to the device.
[0624] Step 6:
[0625] The device performs the volume adjustment. Upon receiving the volume adjustment command, the device automatically reduces the volume of the audio source using Bluetooth or other communication methods. The input here is the volume adjustment command, and the output is the reduced volume.
[0626] Step 7:
[0627] The server generates a notification and sends it to the device. Based on the location information of the emergency vehicle, the server generates a notification for voice and screen display and sends it to the device. The input here is the direction and distance information of the emergency vehicle, and the output is the generated notification data.
[0628] Step 8:
[0629] The terminal notifies the user. The terminal that receives the notification data notifies the user of the approach of an emergency vehicle through voice and screen display. The input here is the notification data, and the output is the notification to the user.
[0630] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0631] As one embodiment of the present invention, we will describe an in-vehicle emergency vehicle approach detection system and a notification system that takes into account the user's emotional state. This system is composed of a sound collection means, an analysis means, a detection means, a volume adjustment means, a notification means, and an emotion engine, and quickly and accurately detects the approach of an emergency vehicle and notifies the user appropriately according to the user's emotional state.
[0632] overview
[0633] Recording Method:
[0634] Microphone devices installed at the four corners of the vehicle collect surrounding sounds, effectively capturing sounds from all directions in 360 degrees.
[0635] Analysis method:
[0636] This is an analysis method that analyzes sound data from sound collection means in real time to identify the siren sounds specific to emergency vehicles. This analysis uses machine learning algorithms to detect specific frequency bands and sound patterns of siren sounds.
[0637] Detection Methods:
[0638] Based on the analysis results obtained by the analysis means, the system detects the approach of an emergency vehicle and calculates the direction and distance of the sound source based on the difference in sound arrival time between multiple microphones.
[0639] Volume adjustment means:
[0640] If the detection means detects the approach of an emergency vehicle, the volume of any audio playing in the car (e.g., music, audiobooks, etc.) will be automatically reduced. This is achieved using a Bluetooth connection.
[0641] Means of notification:
[0642] A specific announcement is generated for approaching emergency vehicles and notified to the user through the in-car speakers, including the direction and distance of the emergency vehicle.
[0643] Emotion Engine:
[0644] It is an engine that analyzes the user's emotional state. It analyzes facial recognition data or voice tone to understand the user's emotional state, such as whether they are stressed or relaxed.
[0645] Program processing
[0646] Sound pickup and data transmission
[0647] Terminal: Microphones installed at the four corners of the vehicle collect sound in real time and temporarily store it in a buffer. The contents of the buffer are packetized at regular intervals and sent to the server.
[0648] Siren sound analysis
[0649] Server: Apply machine learning algorithms to the received sound data to identify specific frequency bands and patterns of siren sounds.
[0650] Detecting approaching emergency vehicles
[0651] Server: If a siren sound is detected, the result is fed back to the device.
[0652] Terminal: Based on the analysis results fed back, the direction and distance of the sound source are determined using the difference in sound arrival time from multiple microphones.
[0653] Automatic volume adjustment
[0654] Terminal: Based on the detection results, it sends an instruction via Bluetooth to reduce the volume of the sound source inside the car to a certain percentage (e.g., 50%).
[0655] Emotional state analysis
[0656] Device: The emotion engine uses facial recognition data captured by the in-car camera and the tone of the user's voice collected by the microphone to identify the user's emotional state.
[0657] Generate and play notifications
[0658] Server: Obtains the location and direction of emergency vehicles from the navigation system and generates specific announcement text.
[0659] Terminal: Based on the analysis results of the emotion engine, the content and tone of the announcement are adjusted according to the user's emotional state.
[0660] Emotion-based notifications
[0661] Terminal: The generated voice announcement is played over the car's speakers. Depending on the user's emotional state, the announcement can be calm or emphasize urgency.
[0662] Specific examples
[0663] 1. Sound collection means:
[0664] The device collects sound in real time using microphones installed at the four corners of the vehicle and stores it in a buffer.
[0665] 2. Analysis method:
[0666] The server analyzes the sound data in the buffer using a machine learning algorithm to identify the sound of a siren.
[0667] 3. Detection methods:
[0668] Based on the analysis results from the server, the device calculates the direction and distance of the sound source.
[0669] 4. Volume adjustment means:
[0670] The server detects the approach of an emergency vehicle and the device sends a command via Bluetooth connection to reduce the music volume to 50%.
[0671] 5. Emotion Engine:
[0672] The device recognizes the user's face using the onboard camera, and the emotion engine analyzes the data to understand the user's emotional state.
[0673] 6. Means of notification:
[0674] The server uses data from the navigation system to obtain the location information of the emergency vehicle and generates specific instructions.
[0675] The tone and content of the announcements are adjusted based on the results of the emotion engine. For example, if the user is feeling stressed, the announcements will be made in a calmer tone.
[0676] The terminal generates a voice announcement that is played over the car's speakers, prompting the user to take specific action (e.g., "There is an ambulance approaching 100 meters ahead from the right, so pull over to the side of the road").
[0677] summary
[0678] This system allows even those listening to loud music in their cars to quickly and reliably detect the approach of an emergency vehicle and provide appropriate notifications based on the user's emotional state, allowing them to safely take appropriate action and contributing to traffic management and accident prevention.
[0679] The processing flow will be explained below.
[0680] Step 1:
[0681] Terminal: Sound is collected by multiple microphones installed at the four corners of the vehicle. The collected sound data is temporarily stored in a buffer.
[0682] Specific operation: Sound data is sampled through the microphone interface, for example, every second, and stored in a buffer. Data from each microphone is stored in a separate buffer.
[0683] Step 2:
[0684] Terminal: Collected sound data is packetized at regular intervals (e.g., every second) and sent to the server.
[0685] Specific operation: The sound data is divided into packets and sent to the server using a network protocol. After sending, the buffer is cleared and waits for the next sample.
[0686] Step 3:
[0687] Server: Analyzes the received sound data in real time, using AI to identify specific frequency bands and patterns of siren sounds.
[0688] How it works: A machine learning model (e.g., a convolutional neural network) is applied to analyze sound data and extract features. If a specific pattern is detected, it is identified as a siren sound.
[0689] Step 4:
[0690] Server: If a siren sound is detected, the result is fed back to the device.
[0691] Specific operation: A data packet containing the analysis results is generated and sent to the terminal via the network.
[0692] Step 5:
[0693] Terminal: Based on the analysis results received from the server, the terminal calculates the difference in sound arrival time from multiple microphones and determines the direction and distance of the sound source.
[0694] Specific operation: An algorithm is applied that uses the difference in sound arrival time to calculate the direction and distance of the sound source using techniques such as triangulation.
[0695] Step 6:
[0696] Terminal: If an approaching emergency vehicle is detected, a command to reduce the volume of the in-car speakers to a certain percentage (e.g., 50%) is sent via Bluetooth.
[0697] Specific behavior: Uses the Bluetooth speaker's volume control protocol to send a command to decrease the volume of the currently playing music or voice by the specified percentage.
[0698] Step 7:
[0699] On the device: The emotion engine analyzes the user's facial recognition data from the in-car camera or the user's tone of voice collected by the microphone.
[0700] Specific behavior: Apply facial recognition and voice analysis algorithms to identify the user's emotional state (e.g., stress, relaxed).
[0701] Step 8:
[0702] Server: Works with the navigation system to obtain the current location and direction of emergency vehicles.
[0703] Specific operation: Calls the navigation system's API to obtain the current GPS location and direction of the emergency vehicle. Based on this information, it generates an announcement for the next step.
[0704] Step 9:
[0705] Terminal: Generates specific voice announcements based on the emergency vehicle location information received from the server and the analysis results of the emotion engine.
[0706] What it does: It uses a text generation engine to create an appropriate announcement, then uses a text-to-speech engine to generate an audio file, adjusting the tone depending on the user's emotional state.
[0707] Step 10:
[0708] Terminal: The generated voice announcement is played over the vehicle's speakers, providing specific instructions to the user.
[0709] Specific behavior: The created audio file is added to the playback queue of the car speakers and begins playing immediately. Based on the user's emotional state, the announcement is made in a calm manner or with an emphasis on urgency.
[0710] Step 11:
[0711] User: Follow the voice announcement and take appropriate action promptly.
[0712] Specific actions: For example, "An ambulance is approaching from the right, 100 meters ahead. Please move to the shoulder," and follow the instructions to safely move the vehicle to the shoulder.
[0713] Example 2
[0714] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0715] Conventional vehicle systems have the problem of not being able to take into account the emotional state of the user when notifying them of an approaching emergency vehicle. This can make it difficult for users to take appropriate action in stressful situations, potentially delaying their response to the emergency vehicle. Another issue is that the volume of the audio played in the vehicle must be adjusted manually, which requires the user to divert their attention.
[0716] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server is a processing device for analyzing sound data collected from a sound collection device, and includes an analysis device that monitors surrounding sounds based on the generated analysis results, a detection device that detects the approach of an emergency vehicle based on the analysis results, and an emotion recognition device that collects and analyzes the emotional state of the user. This makes it possible to quickly and accurately detect the approach of an emergency vehicle and to provide an appropriate notification according to the emotional state of the user.
[0717] A "sound collection device" is a device installed at the four corners of a vehicle to collect surrounding sounds.
[0718] The "analysis device" is a device that analyzes collected sound data and monitors surrounding sounds based on the generated analysis results.
[0719] The "detection device" is a device for detecting the approach of an emergency vehicle based on the analysis results generated by the analysis device.
[0720] The "volume adjustment device" is a device that automatically reduces the volume of the sound source being played inside the vehicle when the approach of an emergency vehicle is detected by the detection device.
[0721] A "notification device" is a device that notifies in-vehicle users of the location and approach of an emergency vehicle.
[0722] An "emotion recognition device" is a device for collecting and analyzing a user's emotional state.
[0723] The present invention relates to an in-vehicle emergency vehicle approach detection system and a notification system that takes into account the emotional state of the user. The system is composed of sound collection devices installed at the four corners of the vehicle, an analysis device that analyzes the collected sound data, a detection device that detects the approach of an emergency vehicle based on the analysis results, a volume adjustment device that automatically lowers the volume of the sound source being played in the vehicle when the approach of the emergency vehicle is detected, a notification device that notifies the in-vehicle user of the location and approach of the emergency vehicle, and an emotion recognition device that collects and analyzes the user's emotional state.
[0724] System configuration
[0725] Sound pickup device
[0726] The device includes microphones installed at the four corners of the vehicle, which collect ambient sounds in real time and temporarily store them in a buffer. At regular intervals, the contents of the buffer are sent to a server as data packets.
[0727] analysis device
[0728] The server analyzes the sound data received from the device using a generative AI model (e.g., TensorFlow) to identify specific frequency bands and sound patterns and identify the sound of a siren.
[0729] Detection device
[0730] When the server detects a siren, it sends the result back to the device, which then uses the difference in sound arrival times from multiple microphones to calculate the direction and distance of the sound source.
[0731] volume adjustment device
[0732] Based on the detection results, the device will automatically reduce the volume of in-car audio sources (e.g., radio, music player) to a certain percentage (e.g., 50%) via Bluetooth.
[0733] emotion recognition device
[0734] The device uses an onboard camera and microphone to collect facial recognition data and voice tone, and then uses an emotion engine to analyze the user's emotional state, determining whether the user is stressed or relaxed.
[0735] notification device
[0736] The server obtains the location and heading of the emergency vehicle from the navigation system and uses a generative AI model to generate a specific announcement. This announcement includes tone and content based on the user's emotional state. The device plays the generated voice announcement over the car's speakers, encouraging the user to take specific action.
[0737] Specific examples
[0738] Here is a concrete example of how the system works:
[0739] 1. Sound pickup devices are installed at the four corners of the vehicle to collect sound in real time, and this data is periodically sent to a server.
[0740] 2. The server analyzes the sound data and uses a generative AI model to identify specific frequency bands and detect siren sounds.
[0741] 3. The server sends the analysis results back to the device, which then calculates the direction and distance of the sound source.
[0742] 4. If an approaching emergency vehicle is detected, the device will send a command via Bluetooth to halve the volume of sound sources inside the vehicle.
[0743] 5. The terminal analyzes the user's emotional state using an emotion engine, and the server generates a specific announcement based on the navigation data.
[0744] 6. The generated announcement is played by the terminal through the in-car speaker, and an announcement such as "There is an ambulance approaching 100 meters ahead from the right, so please pull over to the side of the road" is made.
[0745] Prompt Sentence Examples
[0746] Please generate a natural language document that describes an in-vehicle emergency vehicle approach detection system and a notification system that takes into account the user's emotional state. Please provide a detailed description of the processing flow, including specific hardware and software, including sound collection devices, analysis devices, detection devices, volume control devices, notification devices, and emotion recognition devices.
[0747] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0748] Step 1: Recording and transmitting data
[0749] Terminal: Microphones installed at the four corners of the vehicle collect surrounding sounds in real time and temporarily store the data in a buffer. The sound data in the buffer is divided into packets at regular intervals (e.g., every 10 seconds) and sent to the server.
[0750] Input: Sound data collected from microphones at the four corners of the vehicle
[0751] Output: Sound data sent to the server in data packet format
[0752] Step 2: Analyzing the siren sound
[0753] Server: Applying a generative AI model (e.g., TensorFlow) to the received sound data to identify specific frequency bands and patterns of siren sounds. The results of this analysis are passed on to the next processing step.
[0754] Input: Sound data sent from the device
[0755] Output: Siren sound analysis results (specific frequency bands and patterns)
[0756] Step 3: Detect approaching emergency vehicles
[0757] Server: When the siren sound is identified, the server feeds back the result to the device. Based on the feedback, the device calculates the difference in sound arrival time from multiple microphones and specifically calculates the direction and distance of the sound source.
[0758] Input: Analysis result of siren sound
[0759] Output: direction and distance of sound source
[0760] Step 4: Automatically adjust the volume
[0761] Device: When an approaching emergency vehicle is detected, the device sends a command via Bluetooth to automatically lower the volume of the currently playing audio source (e.g., radio, music player), for example, by 50%.
[0762] Input: Proximity information based on the direction and distance of the sound source
[0763] Output: Volume down command via Bluetooth
[0764] Step 5: Analyze your emotional state
[0765] Device: Analyzes the user's facial recognition data acquired by the in-car camera and the tone of voice collected by the microphone, and uses a generative AI model to determine the user's emotional state (e.g., stressed, relaxed).
[0766] Input: User's facial recognition data and tone of voice
[0767] Output: User's emotional state
[0768] Step 6: Generate and play notifications
[0769] Server: Obtains the location and heading of the emergency vehicle from the navigation system, and uses a generative AI model to create a specific announcement based on that information. This announcement is adjusted according to the user's emotional state.
[0770] Input: Emergency vehicle location information, user emotional state
[0771] Output: Specific announcement text
[0772] Step 7: Emotional Notifications
[0773] Terminal: Receives announcements generated by the server and plays them in an appropriate tone through the car's speakers. The content of the notification is adjusted according to the user's emotional state, encouraging specific countermeasures.
[0774] Input: Specific announcement text
[0775] Output: Voice notification through car speakers
[0776] (Application example 2)
[0777] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0778] For self-driving vehicles, it is important to quickly and accurately detect the approach of an emergency vehicle and notify the user appropriately, but conventional systems have difficulty accurately detecting the direction and distance of an emergency vehicle and lack the ingenuity to enable the user to respond calmly according to the situation.In addition, there is no notification system that takes into account the user's emotional state, which makes it difficult for the user to respond appropriately if they are feeling stressed.
[0779] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes sound collection means installed at the four corners of the vehicle, analysis means for analyzing collected sound data and monitoring surrounding sounds, detection means for detecting the approach of an emergency vehicle based on the analysis results, volume adjustment means for automatically lowering the volume of the sound source being played in the vehicle when the detection means detects the approach of an emergency vehicle, notification means for notifying users in the vehicle of the location and approach of the emergency vehicle, an emotion engine for analyzing the emotional state of the user, and means for adjusting the content and tone of the notification based on the analysis results of the emotion engine. This makes it possible to quickly and accurately detect the approach of an emergency vehicle and to notify the user appropriately according to their emotional state.
[0780] The "sound collection means" is a device installed at the four corners of the vehicle to collect surrounding sounds.
[0781] The "analysis means" is a means for analyzing collected sound data and identifying the siren sound specific to emergency vehicles.
[0782] The "detection means" is a means for detecting the approach of an emergency vehicle based on the analysis results obtained by the analysis means.
[0783] The "volume adjustment means" is a means for automatically lowering the volume of the sound source being reproduced inside the vehicle when the detection means detects the approach of an emergency vehicle.
[0784] The "notification means" is a means for informing the vehicle occupants of the location and approach of an emergency vehicle.
[0785] "Emotion Engine" is an engine for analyzing the user's emotional state, such as analyzing facial recognition data or voice tone.
[0786] "Calculation by the analysis means" refers to calculation of the approaching direction and distance of the emergency vehicle based on the sound data collected from the sound collection means.
[0787] A "user's emotional state" is the emotional state that a user is feeling, identified based on a facial image or tone of voice.
[0788] The "means for adjusting the content and tone of the notification based on the analysis results" refers to means for changing the content and tone of the voice of the generated notification according to the analysis results of the emotional state by the emotion engine.
[0789] A system for realizing this invention includes sound collection means installed at the four corners of a vehicle, analysis means for analyzing collected sound data and monitoring surrounding sounds, detection means for detecting the approach of an emergency vehicle based on the analysis results, volume adjustment means for automatically lowering the volume of the sound source being played inside the vehicle when the detection means detects the approach of an emergency vehicle, notification means for notifying users inside the vehicle of the location and approach of the emergency vehicle, an emotion engine for analyzing the emotional state of the user, and means for adjusting the content and tone of the notification based on the analysis results of the emotion engine.
[0790] System Components
[0791] Sound collection means
[0792] The sound collection means consists of microphones installed at the four corners of the vehicle, which can effectively capture sounds from all directions (360 degrees).
[0793] Analysis means
[0794] The analysis means analyzes the sound data from the sound collection means in real time and identifies the siren sounds specific to emergency vehicles. This analysis uses machine learning algorithms to detect specific frequency bands and sound patterns of siren sounds. The main software used is Python's "librosa" and "keras."
[0795] Detection Method
[0796] The detection means detects the approach of an emergency vehicle based on the analysis results obtained from the analysis means, and calculates the direction and distance of the sound source based on the difference in sound arrival times detected by the multiple microphones.
[0797] Volume adjustment means
[0798] The volume adjustment means uses a Bluetooth connection to automatically reduce the volume of audio sources (such as music or audiobooks) in the car.
[0799] Notification means
[0800] The notification means generates a specific announcement when an emergency vehicle is approaching and notifies the user through an in-vehicle speaker, the announcement including the direction and distance of the emergency vehicle.
[0801] Emotion Engine
[0802] The emotion engine analyzes the user's emotional state by analyzing facial recognition data and voice tone. The main software used is Python's "opencv" and "keras."
[0803] System Operation
[0804] The server analyzes the sound data sent from the sound collection means and identifies the siren sound. After the analysis, it calculates the direction and distance of the emergency vehicle and feeds the results back to the device. The device then captures an image of the user's face using the in-car camera and analyzes it using the emotion engine. Based on the analysis results, the notification means announces the approach of an emergency vehicle.
[0805] The notification means obtains emergency vehicle location information from the navigation system and generates a specific announcement, the tone and content of which are adjusted according to the user's emotional state, and is played through the vehicle speakers.
[0806] Specific examples
[0807] As a specific example, the user collects sound in real time using microphone devices installed at the four corners of the vehicle and sends the data to a server. The server analyzes the siren sound and calculates the direction and distance of the emergency vehicle. The results are fed back to the device, which generates an announcement based on the emergency vehicle's location information. Furthermore, the device captures an image of the user's face using an in-car camera, analyzes it using an emotion engine, and the notification means adjusts the content and tone of the announcement.
[0808] Prompt Sentence Examples
[0809] "When providing voice notifications for approaching emergency vehicles in an autonomous vehicle, please generate notification messages that take into account the emotional state of the occupants. For example, if an emergency vehicle is approaching from the right direction and 100 meters ahead, the notification message should be delivered in a calm tone if the occupant is feeling stressed, and in a standard tone if the occupant is relaxed."
[0810] In this way, it is possible to quickly and accurately detect the approach of an emergency vehicle and notify the user according to their emotional state.
[0811] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0812] Step 1:
[0813] The device collects sound in real time using microphones installed at the four corners of the vehicle. The sound data is temporarily stored in a buffer. This makes it possible to effectively capture sound from all directions (360 degrees). The input is the sound picked up from around the vehicle, and the output is the sound data stored in the buffer.
[0814] Step 2:
[0815] The terminal packetizes the contents of the buffer at regular intervals and sends the sound data to the server. The server receives the sound data and passes it to the analysis means. The input is the sound data stored in the buffer, and the output is the sound data uploaded to the server.
[0816] Step 3:
[0817] The server applies machine learning algorithms to the received sound data to identify specific frequency bands and patterns of siren sounds. The software used is the Python libraries "librosa" and "keras." The input is the sound data uploaded to the server, and the output is the determination of whether or not a siren is being heard.
[0818] Step 4:
[0819] If the server detects a siren, it feeds the result back to the device. It then calculates the direction and distance of the emergency vehicle based on the analysis results. The input is the result of the siren detection, and the output is data on the direction and distance of the emergency vehicle.
[0820] Step 5:
[0821] Based on the detection results, the device sends an instruction via Bluetooth to reduce the volume of the sound source inside the car by a certain percentage (e.g., 50%). The input is data about the direction and distance of the emergency vehicle, and the output is a volume adjustment instruction. Specifically, the device sends an instruction to the car's audio system using Bluetooth.
[0822] Step 6:
[0823] The device captures the user's facial image using an onboard camera and analyzes it with an emotion engine. The software used is Python's "opencv" and "keras." The input is the captured facial image, and the output is the user's emotional state.
[0824] Step 7:
[0825] The server obtains the location and direction of the emergency vehicle from the navigation system and generates a specific announcement. The input is the data on the direction and distance of the emergency vehicle and the user's emotional state, and the output is the generated announcement.
[0826] Step 8:
[0827] The device adjusts the content and tone of the announcement based on the analysis results of the emotion engine according to the user's emotional state. For example, if the user is feeling stressed, the announcement will be made in a calmer tone. The input is the generated announcement text, and the output is the adjusted announcement text.
[0828] Step 9:
[0829] The terminal plays the generated voice announcement over the in-car speaker, prompting the user to take specific action (e.g., "There is an ambulance approaching 100 meters ahead from the right, please pull over to the side of the road.") The input is the adjusted announcement text, and the output is a voice notification from the in-car speaker.
[0830] In this way, through each step, it is possible to quickly and accurately detect the approach of an emergency vehicle and notify the user appropriately according to their emotional state.
[0831] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0832] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0833] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0834] [Third embodiment]
[0835] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0836] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0837] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0838] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0839] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0840] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0841] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0842] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0843] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0844] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0845] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0846] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0847] As an example of the present invention, an emergency vehicle approach detection system for a vehicle will be described. This system is composed of a sound collection means, an analysis means, a detection means, a volume adjustment means, and a notification means, and quickly and accurately detects the approach of an emergency vehicle and notifies the user.
[0848] overview
[0849] Recording Method:
[0850] Microphone devices installed at the four corners of the vehicle collect ambient sounds, effectively capturing sounds coming from all directions in 360 degrees.
[0851] Analysis method:
[0852] The system analyzes sound data from sound collection devices in real time to identify the siren sounds specific to emergency vehicles, using machine learning algorithms to detect specific frequency bands and sound patterns of siren sounds.
[0853] Detection Methods:
[0854] Based on the analysis results obtained by the analysis means, the system detects the approach of an emergency vehicle and calculates the direction and distance of the sound source based on the difference in arrival time of the sounds from multiple microphones.
[0855] Volume adjustment means:
[0856] If the detection means detects the approach of an emergency vehicle, the volume of any audio source (e.g., music, audiobooks, etc.) being played in the vehicle will be automatically reduced via Bluetooth connection.
[0857] Means of notification:
[0858] Specific announcements will be made over the vehicle's speakers to passengers informing them of the location and approach of emergency vehicles.
[0859] Program processing
[0860] Sound pickup and data transmission
[0861] Terminal: Collects sound data in real time from microphones installed at the four corners of the vehicle and temporarily stores it in a buffer.
[0862] Server: Receives sound data periodically sent from the terminal.
[0863] Siren sound analysis
[0864] Server: Analyzes the received sound data and applies machine learning algorithms to detect frequency bands and sound patterns to identify the sound of an emergency vehicle siren.
[0865] Terminal: Receives the analysis results from the server and prepares any necessary next steps.
[0866] Detecting approaching emergency vehicles
[0867] Terminal: Using data from multiple microphones, a sound source localization algorithm calculates the difference in sound arrival time and determines the direction and distance of the emergency vehicle.
[0868] Automatic volume adjustment
[0869] Device: Based on the detection results, it sends a command via Bluetooth connection to reduce the volume of the sound source inside the car to 50%.
[0870] Generate and play notifications
[0871] Server: Obtains the current location and direction of the emergency vehicle based on data from the navigation system and generates an appropriate announcement.
[0872] Terminal: The generated voice announcement is played over the car's speakers.
[0873] Specific examples
[0874] 1. Sound collection means:
[0875] The device collects sound in real time using microphones installed at the four corners of the vehicle and stores it in a buffer.
[0876] 2. Analysis method:
[0877] The server analyzes the sound data in the buffer using a machine learning algorithm to identify the sound of a siren.
[0878] 3. Detection methods:
[0879] Based on the analysis results from the server, the device calculates the direction and distance of the sound source.
[0880] 4. Volume adjustment means:
[0881] The server detects the approach of an emergency vehicle and the device sends a command via Bluetooth connection to reduce the music volume to 50%.
[0882] 5. Means of notification:
[0883] The server uses data from the navigation system to obtain the location information of the emergency vehicle and generates specific instructions.
[0884] The terminal generates a voice announcement that is played over the car's speakers, prompting the user to take specific action (e.g., there is an emergency vehicle approaching 100 meters ahead from the right; move to the shoulder of the road).
[0885] summary
[0886] This system allows even passengers who are enjoying loud music in their cars to quickly and reliably detect approaching emergency vehicles and respond safely, thereby reducing the risk of accidents and supporting smooth traffic management.
[0887] The processing flow will be explained below.
[0888] Step 1:
[0889] Terminal: Sound is collected by multiple microphones installed at the four corners of the vehicle. The collected sound data is temporarily stored in a buffer.
[0890] Specific operation: Sound data is sampled through the microphone interface, for example, every second, and stored in a buffer. Data from each microphone is stored in a separate buffer.
[0891] Step 2:
[0892] Terminal: Collected sound data is packetized at regular intervals (e.g., every second) and sent to the server.
[0893] Specific operation: The sound data is divided into packets and sent to the server using a network protocol. After sending, the buffer is cleared and waits for the next sample.
[0894] Step 3:
[0895] Server: Analyzes the received sound data in real time, using AI to identify specific frequency bands and patterns of siren sounds.
[0896] How it works: A machine learning model (e.g., a convolutional neural network) is applied to analyze sound data and extract features. If a specific pattern is detected, it is identified as a siren sound.
[0897] Step 4:
[0898] Server: If a siren sound is detected, the result is fed back to the device.
[0899] Specific operation: A data packet containing the analysis results is generated and sent to the terminal via the network.
[0900] Step 5:
[0901] Terminal: Based on the analysis results received from the server, the terminal calculates the difference in sound arrival time from multiple microphones and determines the direction and distance of the sound source.
[0902] Specific operation: An algorithm is applied that uses the difference in sound arrival time to calculate the direction and distance of the sound source using techniques such as triangulation.
[0903] Step 6:
[0904] Terminal: If an approaching emergency vehicle is detected, a command to reduce the volume of the in-car speakers to a certain percentage (e.g., 50%) is sent via Bluetooth.
[0905] Specific behavior: Uses the Bluetooth speaker's volume control protocol to send a command to decrease the volume of the currently playing music or voice by the specified percentage.
[0906] Step 7:
[0907] Server: Works with the navigation system to obtain the current location and direction of emergency vehicles.
[0908] Specific operation: Calls the navigation system's API to obtain the current GPS location and direction of the emergency vehicle. Based on this information, it generates an announcement for the next step.
[0909] Step 8:
[0910] Terminal: Generates specific voice announcements based on location information received from the server.
[0911] What it does: Creates the appropriate announcement using a text generation engine and generates an audio file using a text-to-speech engine.
[0912] Step 9:
[0913] Terminal: The generated voice announcement is played over the vehicle's speakers, providing specific instructions to the user.
[0914] Specific operation: The created audio file is added to the playback queue of the car speakers and begins playing immediately.
[0915] Step 10:
[0916] User: Follow the voice announcement and take appropriate action promptly.
[0917] Specific actions: For example, "An ambulance is approaching from the right, 100 meters ahead. Please move to the shoulder," and follow the instructions to safely move the vehicle to the shoulder.
[0918] Example 1
[0919] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0920] Conventional vehicle emergency vehicle approach detection systems lack the functionality to quickly detect approaching emergency vehicles and notify the driver. In particular, drivers listening to loud music or audiobooks in the vehicle may not notice the approach of an emergency vehicle, increasing the risk of an accident. The objective of this invention is to solve this problem by providing a system that quickly and accurately detects the approach of an emergency vehicle and notifies the driver in an appropriate manner.
[0921] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0922] In this invention, the server includes a sound collection means, a sound data analysis means, an emergency vehicle detection means, an automatic volume adjustment means, a notification means, and a notification generation means. This makes it possible to analyze in real time the sound data collected by the sound collection means installed at the four corners of the vehicle and quickly detect the approach of an emergency vehicle. Furthermore, when an emergency vehicle approaches, the server can automatically lower the volume inside the vehicle and generate an appropriate announcement to notify the driver.
[0923] The "sound collection means" is a device installed at each of the four corners of the vehicle that collects surrounding sounds in real time.
[0924] The "sound data analysis means" is a function that analyzes collected sound data, detects specific frequency bands and patterns, and identifies the siren sounds of emergency vehicles.
[0925] The "emergency vehicle detection means" is a function that detects the approach of an emergency vehicle based on the analysis results obtained by the sound data analysis means.
[0926] The "automatic volume adjustment means" is a function that automatically reduces the volume of the sound source being played inside the vehicle when the emergency vehicle detection means detects the approach of an emergency vehicle.
[0927] The "notification means" is a function for informing the vehicle occupants of the location and approach of an emergency vehicle.
[0928] The "notification generating means" is a function that obtains the location of emergency vehicles based on data from the navigation system and generates an appropriate announcement.
[0929] The present invention relates to a system for detecting approaching emergency vehicles in a vehicle, and specific embodiments thereof are described below. This system is composed of a sound collection means, a sound data analysis means, an emergency vehicle detection means, an automatic volume adjustment means, a notification means, and a notification generation means. By combining these means, it is possible to quickly and accurately detect the approach of an emergency vehicle and notify occupants in the vehicle.
[0930] Sound collection means
[0931] Microphones installed at the four corners of the vehicle act as sound collectors. These microphones collect ambient sounds in real time and temporarily store the data in a buffer. For example, when the vehicle approaches an intersection, each microphone collects the surrounding sounds.
[0932] Sound data analysis method
[0933] The server receives audio data periodically sent from the device. It analyzes the received audio data using a machine learning algorithm to detect specific frequency bands and sound patterns and identify the sound of an emergency vehicle siren. The server then sends the analysis results to the device and prepares for the next process.
[0934] Emergency vehicle detection means
[0935] Based on the analysis results sent from the server, the device uses data from multiple microphones and a sound source localization algorithm to calculate the direction and distance of the emergency vehicle. For example, it can determine that an emergency vehicle is approaching 100 meters behind and to the left.
[0936] Automatic volume adjustment means
[0937] If the device detects an approaching emergency vehicle, it will send a command over the Bluetooth connection to reduce the volume of the in-car audio source (e.g., music) to 50%, allowing the in-car audio system to automatically reduce the current volume.
[0938] Notification means and notification generation means
[0939] The server obtains data on the current location and direction of the emergency vehicle from the navigation system and generates an appropriate announcement for the user. The device receives this announcement and plays it through the car's speakers. For example, it can provide specific instructions to the user, such as "Police vehicle approaching 100 meters behind you on the left, please be careful."
[0940] Prompt Sentence Examples
[0941] Here are some examples of prompts:
[0942] "Please explain the process by which microphones installed at the four corners of the vehicle collect ambient sounds in real time and store that data in a buffer."
[0943] "Please explain in detail how the server uses a machine learning algorithm to analyze the sound data received from the device and identify the sound of a siren."
[0944] "Please explain in detail the algorithm that allows the device to use data from multiple microphones to calculate the direction and distance of emergency vehicles."
[0945] "Please explain the steps to reduce the volume of in-car audio sources to 50% when the device detects an approaching emergency vehicle."
[0946] "Please explain the procedure for the server to generate an announcement to notify the user of the approach of an emergency vehicle and send it to the car speaker."
[0947] In this way, the system of the present invention can quickly detect the approach of an emergency vehicle, provide the user with the necessary information, and encourage appropriate action.
[0948] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0949] Step 1: Sound collection and data transmission
[0950] Terminal
[0951] Input: Microphones installed at the four corners of the vehicle collect surrounding sounds in real time.
[0952] Data processing / data calculation: Collected sound data is temporarily stored in a buffer and prepared to be sent to the server at regular intervals.
[0953] Output: Sends the sound data stored in the temporary buffer to the server.
[0954] Specific operation: For example, when a vehicle is driving and approaches an intersection, the microphone collects sounds of nearby pedestrians and other vehicles and sends the data to a server at regular intervals.
[0955] Step 2: Analyzing the siren sound
[0956] server
[0957] Input: Sound data sent periodically from the device.
[0958] Data processing / data calculation: Applying machine learning algorithms to analyze sound data and detect specific frequency bands and sound patterns to identify the sound of an emergency vehicle siren.
[0959] Output: As an analysis result, information is generated as to whether or not an emergency vehicle siren was detected.
[0960] Specific operation: The server analyzes the sound data received from the device, and if it detects a specific frequency pattern between 700 and 1500 Hz, it identifies it as the sound of an emergency vehicle siren. The analysis result is returned to the device.
[0961] Step 3: Detect approaching emergency vehicles
[0962] Terminal
[0963] Input: Analysis results sent from the server.
[0964] Data processing / data calculation: Based on the difference in arrival time of sound data obtained from multiple microphones, a sound source localization algorithm is used to calculate the direction and distance of the emergency vehicle.
[0965] Output: Information on the approaching direction and distance of the emergency vehicle.
[0966] Specific operation: For example, identify that an emergency vehicle is approaching 100 meters behind and to the left. This information is used in the next step.
[0967] Step 4: Automatically adjust the volume
[0968] Terminal
[0969] Input: Information about the approaching direction and distance of the emergency vehicle, and information about the audio source currently being played.
[0970] Data processing / data calculation: If an approaching emergency vehicle is detected, a command is generated to reduce the volume of the sound source (e.g., music) inside the vehicle to 50%.
[0971] Output: A command to decrease the volume over the Bluetooth connection.
[0972] Specific behavior: For example, if music is playing in the car, the volume will be automatically reduced to 50%. This operation is performed via Bluetooth connection.
[0973] Step 5: Generate and play notifications
[0974] server
[0975] Input: Emergency vehicle approach direction and distance information, and location data from the navigation system.
[0976] Data processing / data calculation: Based on data from the navigation system, appropriate announcements are generated for users.
[0977] Output: The generated announcement.
[0978] Specific actions: For example, generate specific instructions such as "A police vehicle is approaching 100 meters behind you on the left. Be careful."
[0979] Terminal
[0980] Input: Server-generated announcement.
[0981] Data processing / data calculation: Generates commands to play announcements over the car speakers.
[0982] Output: Announcement command to the car speaker.
[0983] Specific operation: The generated voice announcement is played over the car's speakers to notify the user of the approaching emergency vehicle and how to respond. For example, it may say, "Siren approaching, emergency vehicle on the left."
[0984] (Application example 1)
[0985] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0986] Conventional vehicle emergency vehicle notification systems alert users to approaching emergency vehicles by lowering the volume inside the vehicle, but lack the technology for use outside the vehicle or on mobile devices. Furthermore, while it is necessary to quickly and accurately notify users of the direction and distance of approaching emergency vehicles visually and audibly, no appropriate solution for this purpose exists. Therefore, there is a need for a comprehensive emergency vehicle notification system that allows users to respond safely both inside and outside the vehicle.
[0987] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0988] In this invention, the server includes a sound collection means, an analysis means, a detection means, a volume adjustment means, and a notification means. This allows the server to quickly and accurately detect the approach of an emergency vehicle by collecting surrounding sounds, even when using a vehicle or a mobile device, automatically lower the volume of the sound source being played, and effectively notify the user of the location and approach of the emergency vehicle. Specifically, the analysis means calculates the direction and distance of the approaching emergency vehicle, and the information is provided via voice and screen display, allowing the user to quickly take appropriate action.
[0989] A "sound collection means" is a device installed to collect surrounding sounds.
[0990] An "analysis means" is a device or algorithm for analyzing collected sound data and identifying specific sound patterns.
[0991] The "detection means" is a device or function for detecting the approach of a specific sound source, for example, an emergency vehicle, based on the analysis result of the analysis means.
[0992] The "volume adjustment means" is a device or function that automatically adjusts the volume of the sound source being played back when the detection means detects the approach of an emergency vehicle.
[0993] "Notification means" refers to a device or function for informing users of the location and approach of an emergency vehicle.
[0994] A "mobile terminal" is a communication device or computer device that can be carried by a user.
[0995] The "audio output means" is a device for conveying the audio information generated by the notification means to the user.
[0996] "Screen display" refers to the function of displaying textual or graphic information on a mobile terminal or other display device.
[0997] The present invention provides a system that uses a vehicle or a mobile terminal to quickly detect the approach of an emergency vehicle and automatically reduce the volume of the sound source being played. Specific embodiments for carrying out the present invention are described below.
[0998] overview
[0999] This system consists of a sound collection means, an analysis means, a detection means, a volume adjustment means, and a notification means. The sound collection means is a microphone installed in a vehicle or a mobile device that collects surrounding sounds. The analysis means analyzes the collected sound data and identifies specific sound patterns, such as siren sounds. This analysis uses a machine learning algorithm (e.g., a model built with Keras). The detection means detects the approach of an emergency vehicle based on the analysis results. The volume adjustment means automatically lowers the volume of the sound being played when the detection means detects the approach of an emergency vehicle. The notification means effectively notifies the user of the location and approach of the emergency vehicle.
[1000] Program processing
[1001] Sound pickup and data transmission
[1002] The device collects ambient sounds through a microphone, and the collected sound data is temporarily stored in a buffer and processed by an analysis means.
[1003] Siren sound analysis
[1004] The server receives the sound data in the buffer and applies a machine learning algorithm to identify siren sounds. Keras is used as an analysis tool to analyze the collected sound data.
[1005] Detecting approaching emergency vehicles
[1006] The detection means detects the approach of an emergency vehicle based on the results of the analysis means, and calculates the direction and distance of the sound source from the collected sound data.
[1007] Volume adjustment
[1008] The volume control means automatically reduces the volume of the audio source being played when detected, and sends a volume control command via Bluetooth or another communication means.
[1009] Generate and play notifications
[1010] The server generates an appropriate announcement based on the location information of the emergency vehicle, and the notification means notifies the user of this announcement both by voice and by display on the screen.
[1011] Specific examples
[1012] While the user is driving, the device (smartphone) collects surrounding sounds using a built-in microphone and sends them to a server. The server analyzes the sound data in real time, and if a siren sound is detected, it automatically lowers the volume of the smartphone via Bluetooth and notifies the user of the location of an emergency vehicle via a display and voice message, allowing the user to take prompt and appropriate action.
[1013] Prompt Sentence Examples
[1014] "We are developing an emergency vehicle approach detection system for vehicles as a smartphone application. Please generate a Python program that follows the requirements below:
[1015] Built-in microphone for sound capture
[1016] Uses machine learning models to analyze and detect siren sounds in real time
[1017] Volume adjustment and notification functions implemented
[1018] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1019] Step 1:
[1020] The device collects ambient sounds. While the user is driving the vehicle, the device's microphone collects audio data in real time and temporarily stores it in a buffer. The input here is the ambient sounds, and the output is the audio data in the buffer.
[1021] Step 2:
[1022] The device sends the sound data in the buffer to the server. The collected sound data is transferred to the server at regular intervals. The input here is the sound data in the buffer, and the output is the sound data sent to the server.
[1023] Step 3:
[1024] The server analyzes the received sound data. The server uses a machine learning model (a model using Keras) to analyze the sound data and detect siren sounds. The input here is the sound data sent to the server, and the output is the analysis result.
[1025] Step 4:
[1026] The server detects emergency vehicles based on the analysis results. If a siren sound is detected, the server detects the approach of an emergency vehicle and calculates its direction and distance. The input here is the analysis result, and the output is the direction and distance information of the emergency vehicle.
[1027] Step 5:
[1028] The server sends a volume adjustment command to the device. Based on the detection result, the server sends an instruction to the device to lower the volume of the sound being played. The input here is the direction and distance information of the emergency vehicle, and the output is the volume adjustment command to the device.
[1029] Step 6:
[1030] The device performs the volume adjustment. Upon receiving the volume adjustment command, the device automatically reduces the volume of the audio source using Bluetooth or other communication methods. The input here is the volume adjustment command, and the output is the reduced volume.
[1031] Step 7:
[1032] The server generates a notification and sends it to the device. Based on the location information of the emergency vehicle, the server generates a notification for voice and screen display and sends it to the device. The input here is the direction and distance information of the emergency vehicle, and the output is the generated notification data.
[1033] Step 8:
[1034] The terminal notifies the user. The terminal that receives the notification data notifies the user of the approach of an emergency vehicle through voice and screen display. The input here is the notification data, and the output is the notification to the user.
[1035] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1036] As one embodiment of the present invention, we will describe an in-vehicle emergency vehicle approach detection system and a notification system that takes into account the user's emotional state. This system is composed of a sound collection means, an analysis means, a detection means, a volume adjustment means, a notification means, and an emotion engine, and quickly and accurately detects the approach of an emergency vehicle and notifies the user appropriately according to the user's emotional state.
[1037] overview
[1038] Recording Method:
[1039] Microphone devices installed at the four corners of the vehicle collect surrounding sounds, effectively capturing sounds from all directions in 360 degrees.
[1040] Analysis method:
[1041] This is an analysis method that analyzes sound data from sound collection means in real time to identify the siren sounds specific to emergency vehicles. This analysis uses machine learning algorithms to detect specific frequency bands and sound patterns of siren sounds.
[1042] Detection Methods:
[1043] Based on the analysis results obtained by the analysis means, the system detects the approach of an emergency vehicle and calculates the direction and distance of the sound source based on the difference in sound arrival time between multiple microphones.
[1044] Volume adjustment means:
[1045] If the detection means detects the approach of an emergency vehicle, the volume of any audio playing in the car (e.g., music, audiobooks, etc.) will be automatically reduced. This is achieved using a Bluetooth connection.
[1046] Means of notification:
[1047] A specific announcement is generated for approaching emergency vehicles and notified to the user through the in-car speakers, including the direction and distance of the emergency vehicle.
[1048] Emotion Engine:
[1049] It is an engine that analyzes the user's emotional state. It analyzes facial recognition data or voice tone to understand the user's emotional state, such as whether they are stressed or relaxed.
[1050] Program processing
[1051] Sound pickup and data transmission
[1052] Terminal: Microphones installed at the four corners of the vehicle collect sound in real time and temporarily store it in a buffer. The contents of the buffer are packetized at regular intervals and sent to the server.
[1053] Siren sound analysis
[1054] Server: Apply machine learning algorithms to the received sound data to identify specific frequency bands and patterns of siren sounds.
[1055] Detecting approaching emergency vehicles
[1056] Server: If a siren sound is detected, the result is fed back to the device.
[1057] Terminal: Based on the analysis results fed back, the direction and distance of the sound source are determined using the difference in sound arrival time from multiple microphones.
[1058] Automatic volume adjustment
[1059] Terminal: Based on the detection results, it sends an instruction via Bluetooth to reduce the volume of the sound source inside the car to a certain percentage (e.g., 50%).
[1060] Emotional state analysis
[1061] Device: The emotion engine uses facial recognition data captured by the in-car camera and the tone of the user's voice collected by the microphone to identify the user's emotional state.
[1062] Generate and play notifications
[1063] Server: Obtains the location and direction of emergency vehicles from the navigation system and generates specific announcement text.
[1064] Terminal: Based on the analysis results of the emotion engine, the content and tone of the announcement are adjusted according to the user's emotional state.
[1065] Emotion-based notifications
[1066] Terminal: The generated voice announcement is played over the car's speakers. Depending on the user's emotional state, the announcement can be calm or emphasize urgency.
[1067] Specific examples
[1068] 1. Sound collection means:
[1069] The device collects sound in real time using microphones installed at the four corners of the vehicle and stores it in a buffer.
[1070] 2. Analysis method:
[1071] The server analyzes the sound data in the buffer using a machine learning algorithm to identify the sound of a siren.
[1072] 3. Detection methods:
[1073] Based on the analysis results from the server, the device calculates the direction and distance of the sound source.
[1074] 4. Volume adjustment means:
[1075] The server detects the approach of an emergency vehicle and the device sends a command via Bluetooth connection to reduce the music volume to 50%.
[1076] 5. Emotion Engine:
[1077] The device recognizes the user's face using the onboard camera, and the emotion engine analyzes the data to understand the user's emotional state.
[1078] 6. Means of notification:
[1079] The server uses data from the navigation system to obtain the location information of the emergency vehicle and generates specific instructions.
[1080] The tone and content of the announcements are adjusted based on the results of the emotion engine. For example, if the user is feeling stressed, the announcements will be made in a calmer tone.
[1081] The terminal generates a voice announcement that is played over the car's speakers, prompting the user to take specific action (e.g., "There is an ambulance approaching 100 meters ahead from the right, so pull over to the side of the road").
[1082] summary
[1083] This system allows even those listening to loud music in their cars to quickly and reliably detect the approach of an emergency vehicle and provide appropriate notifications based on the user's emotional state, allowing them to safely take appropriate action and contributing to traffic management and accident prevention.
[1084] The processing flow will be explained below.
[1085] Step 1:
[1086] Terminal: Sound is collected by multiple microphones installed at the four corners of the vehicle. The collected sound data is temporarily stored in a buffer.
[1087] Specific operation: Sound data is sampled through the microphone interface, for example, every second, and stored in a buffer. Data from each microphone is stored in a separate buffer.
[1088] Step 2:
[1089] Terminal: Collected sound data is packetized at regular intervals (e.g., every second) and sent to the server.
[1090] Specific operation: The sound data is divided into packets and sent to the server using a network protocol. After sending, the buffer is cleared and waits for the next sample.
[1091] Step 3:
[1092] Server: Analyzes the received sound data in real time, using AI to identify specific frequency bands and patterns of siren sounds.
[1093] How it works: A machine learning model (e.g., a convolutional neural network) is applied to analyze sound data and extract features. If a specific pattern is detected, it is identified as a siren sound.
[1094] Step 4:
[1095] Server: If a siren sound is detected, the result is fed back to the device.
[1096] Specific operation: A data packet containing the analysis results is generated and sent to the terminal via the network.
[1097] Step 5:
[1098] Terminal: Based on the analysis results received from the server, the terminal calculates the difference in sound arrival time from multiple microphones and determines the direction and distance of the sound source.
[1099] Specific operation: An algorithm is applied that uses the difference in sound arrival time to calculate the direction and distance of the sound source using techniques such as triangulation.
[1100] Step 6:
[1101] Terminal: If an approaching emergency vehicle is detected, a command to reduce the volume of the in-car speakers to a certain percentage (e.g., 50%) is sent via Bluetooth.
[1102] Specific behavior: Uses the Bluetooth speaker's volume control protocol to send a command to decrease the volume of the currently playing music or voice by the specified percentage.
[1103] Step 7:
[1104] On the device: The emotion engine analyzes the user's facial recognition data from the in-car camera or the user's tone of voice collected by the microphone.
[1105] Specific behavior: Apply facial recognition and voice analysis algorithms to identify the user's emotional state (e.g., stress, relaxed).
[1106] Step 8:
[1107] Server: Works with the navigation system to obtain the current location and direction of emergency vehicles.
[1108] Specific operation: Calls the navigation system's API to obtain the current GPS location and direction of the emergency vehicle. Based on this information, it generates an announcement for the next step.
[1109] Step 9:
[1110] Terminal: Generates specific voice announcements based on the emergency vehicle location information received from the server and the analysis results of the emotion engine.
[1111] What it does: It uses a text generation engine to create an appropriate announcement, then uses a text-to-speech engine to generate an audio file, adjusting the tone depending on the user's emotional state.
[1112] Step 10:
[1113] Terminal: The generated voice announcement is played over the vehicle's speakers, providing specific instructions to the user.
[1114] Specific behavior: The created audio file is added to the playback queue of the car speakers and begins playing immediately. Based on the user's emotional state, the announcement is made in a calm manner or with an emphasis on urgency.
[1115] Step 11:
[1116] User: Follow the voice announcement and take appropriate action promptly.
[1117] Specific actions: For example, "An ambulance is approaching from the right, 100 meters ahead. Please move to the shoulder," and follow the instructions to safely move the vehicle to the shoulder.
[1118] Example 2
[1119] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1120] Conventional vehicle systems have the problem of not being able to take into account the emotional state of the user when notifying them of an approaching emergency vehicle. This can make it difficult for users to take appropriate action in stressful situations, potentially delaying their response to the emergency vehicle. Another issue is that the volume of the audio played in the vehicle must be adjusted manually, which requires the user to divert their attention.
[1121] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server is a processing device for analyzing sound data collected from a sound collection device, and includes an analysis device that monitors surrounding sounds based on the generated analysis results, a detection device that detects the approach of an emergency vehicle based on the analysis results, and an emotion recognition device that collects and analyzes the emotional state of the user. This makes it possible to quickly and accurately detect the approach of an emergency vehicle and to provide an appropriate notification according to the emotional state of the user.
[1122] A "sound collection device" is a device installed at the four corners of a vehicle to collect surrounding sounds.
[1123] The "analysis device" is a device that analyzes collected sound data and monitors surrounding sounds based on the generated analysis results.
[1124] The "detection device" is a device for detecting the approach of an emergency vehicle based on the analysis results generated by the analysis device.
[1125] The "volume adjustment device" is a device that automatically reduces the volume of the sound source being played inside the vehicle when the approach of an emergency vehicle is detected by the detection device.
[1126] A "notification device" is a device that notifies in-vehicle users of the location and approach of an emergency vehicle.
[1127] An "emotion recognition device" is a device for collecting and analyzing a user's emotional state.
[1128] The present invention relates to an in-vehicle emergency vehicle approach detection system and a notification system that takes into account the emotional state of the user. The system is composed of sound collection devices installed at the four corners of the vehicle, an analysis device that analyzes the collected sound data, a detection device that detects the approach of an emergency vehicle based on the analysis results, a volume adjustment device that automatically lowers the volume of the sound source being played in the vehicle when the approach of the emergency vehicle is detected, a notification device that notifies the in-vehicle user of the location and approach of the emergency vehicle, and an emotion recognition device that collects and analyzes the user's emotional state.
[1129] System configuration
[1130] Sound pickup device
[1131] The device includes microphones installed at the four corners of the vehicle, which collect ambient sounds in real time and temporarily store them in a buffer. At regular intervals, the contents of the buffer are sent to a server as data packets.
[1132] analysis device
[1133] The server analyzes the sound data received from the device using a generative AI model (e.g., TensorFlow) to identify specific frequency bands and sound patterns and identify the sound of a siren.
[1134] Detection device
[1135] When the server detects a siren, it sends the result back to the device, which then uses the difference in sound arrival times from multiple microphones to calculate the direction and distance of the sound source.
[1136] volume adjustment device
[1137] Based on the detection results, the device will automatically reduce the volume of in-car audio sources (e.g., radio, music player) to a certain percentage (e.g., 50%) via Bluetooth.
[1138] emotion recognition device
[1139] The device uses an onboard camera and microphone to collect facial recognition data and voice tone, and then uses an emotion engine to analyze the user's emotional state, determining whether the user is stressed or relaxed.
[1140] notification device
[1141] The server obtains the location and heading of the emergency vehicle from the navigation system and uses a generative AI model to generate a specific announcement. This announcement includes tone and content based on the user's emotional state. The device plays the generated voice announcement over the car's speakers, encouraging the user to take specific action.
[1142] Specific examples
[1143] Here is a concrete example of how the system works:
[1144] 1. Sound pickup devices are installed at the four corners of the vehicle to collect sound in real time, and this data is periodically sent to a server.
[1145] 2. The server analyzes the sound data and uses a generative AI model to identify specific frequency bands and detect siren sounds.
[1146] 3. The server sends the analysis results back to the device, which then calculates the direction and distance of the sound source.
[1147] 4. If an approaching emergency vehicle is detected, the device will send a command via Bluetooth to halve the volume of sound sources inside the vehicle.
[1148] 5. The terminal analyzes the user's emotional state using an emotion engine, and the server generates a specific announcement based on the navigation data.
[1149] 6. The generated announcement is played by the terminal through the in-car speaker, and an announcement such as "There is an ambulance approaching 100 meters ahead from the right, so please pull over to the side of the road" is made.
[1150] Prompt Sentence Examples
[1151] Please generate a natural language document that describes an in-vehicle emergency vehicle approach detection system and a notification system that takes into account the user's emotional state. Please provide a detailed description of the processing flow, including specific hardware and software, including sound collection devices, analysis devices, detection devices, volume control devices, notification devices, and emotion recognition devices.
[1152] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1153] Step 1: Recording and transmitting data
[1154] Terminal: Microphones installed at the four corners of the vehicle collect surrounding sounds in real time and temporarily store the data in a buffer. The sound data in the buffer is divided into packets at regular intervals (e.g., every 10 seconds) and sent to the server.
[1155] Input: Sound data collected from microphones at the four corners of the vehicle
[1156] Output: Sound data sent to the server in data packet format
[1157] Step 2: Analyzing the siren sound
[1158] Server: Applying a generative AI model (e.g., TensorFlow) to the received sound data to identify specific frequency bands and patterns of siren sounds. The results of this analysis are passed on to the next processing step.
[1159] Input: Sound data sent from the device
[1160] Output: Siren sound analysis results (specific frequency bands and patterns)
[1161] Step 3: Detect approaching emergency vehicles
[1162] Server: When the siren sound is identified, the server feeds back the result to the device. Based on the feedback, the device calculates the difference in sound arrival time from multiple microphones and specifically calculates the direction and distance of the sound source.
[1163] Input: Analysis result of siren sound
[1164] Output: direction and distance of sound source
[1165] Step 4: Automatically adjust the volume
[1166] Device: When an approaching emergency vehicle is detected, the device sends a command via Bluetooth to automatically lower the volume of the currently playing audio source (e.g., radio, music player), for example, by 50%.
[1167] Input: Proximity information based on the direction and distance of the sound source
[1168] Output: Volume down command via Bluetooth
[1169] Step 5: Analyze your emotional state
[1170] Device: Analyzes the user's facial recognition data acquired by the in-car camera and the tone of voice collected by the microphone, and uses a generative AI model to determine the user's emotional state (e.g., stressed, relaxed).
[1171] Input: User's facial recognition data and tone of voice
[1172] Output: User's emotional state
[1173] Step 6: Generate and play notifications
[1174] Server: Obtains the location and heading of the emergency vehicle from the navigation system, and uses a generative AI model to create a specific announcement based on that information. This announcement is adjusted according to the user's emotional state.
[1175] Input: Emergency vehicle location information, user emotional state
[1176] Output: Specific announcement text
[1177] Step 7: Emotional Notifications
[1178] Terminal: Receives announcements generated by the server and plays them in an appropriate tone through the car's speakers. The content of the notification is adjusted according to the user's emotional state, encouraging specific countermeasures.
[1179] Input: Specific announcement text
[1180] Output: Voice notification through car speakers
[1181] (Application example 2)
[1182] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1183] For self-driving vehicles, it is important to quickly and accurately detect the approach of an emergency vehicle and notify the user appropriately, but conventional systems have difficulty accurately detecting the direction and distance of an emergency vehicle and lack the ingenuity to enable the user to respond calmly according to the situation.In addition, there is no notification system that takes into account the user's emotional state, which makes it difficult for the user to respond appropriately if they are feeling stressed.
[1184] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes sound collection means installed at the four corners of the vehicle, analysis means for analyzing collected sound data and monitoring surrounding sounds, detection means for detecting the approach of an emergency vehicle based on the analysis results, volume adjustment means for automatically lowering the volume of the sound source being played in the vehicle when the detection means detects the approach of an emergency vehicle, notification means for notifying users in the vehicle of the location and approach of the emergency vehicle, an emotion engine for analyzing the emotional state of the user, and means for adjusting the content and tone of the notification based on the analysis results of the emotion engine. This makes it possible to quickly and accurately detect the approach of an emergency vehicle and to notify the user appropriately according to their emotional state.
[1185] The "sound collection means" is a device installed at the four corners of the vehicle to collect surrounding sounds.
[1186] The "analysis means" is a means for analyzing collected sound data and identifying the siren sound specific to emergency vehicles.
[1187] The "detection means" is a means for detecting the approach of an emergency vehicle based on the analysis results obtained by the analysis means.
[1188] The "volume adjustment means" is a means for automatically lowering the volume of the sound source being reproduced inside the vehicle when the detection means detects the approach of an emergency vehicle.
[1189] The "notification means" is a means for informing the vehicle occupants of the location and approach of an emergency vehicle.
[1190] "Emotion Engine" is an engine for analyzing the user's emotional state, such as analyzing facial recognition data or voice tone.
[1191] "Calculation by the analysis means" refers to calculation of the approaching direction and distance of the emergency vehicle based on the sound data collected from the sound collection means.
[1192] A "user's emotional state" is the emotional state that a user is feeling, identified based on a facial image or tone of voice.
[1193] The "means for adjusting the content and tone of the notification based on the analysis results" refers to means for changing the content and tone of the voice of the generated notification according to the analysis results of the emotional state by the emotion engine.
[1194] A system for realizing this invention includes sound collection means installed at the four corners of a vehicle, analysis means for analyzing collected sound data and monitoring surrounding sounds, detection means for detecting the approach of an emergency vehicle based on the analysis results, volume adjustment means for automatically lowering the volume of the sound source being played inside the vehicle when the detection means detects the approach of an emergency vehicle, notification means for notifying users inside the vehicle of the location and approach of the emergency vehicle, an emotion engine for analyzing the emotional state of the user, and means for adjusting the content and tone of the notification based on the analysis results of the emotion engine.
[1195] System Components
[1196] Sound collection means
[1197] The sound collection means consists of microphones installed at the four corners of the vehicle, which can effectively capture sounds from all directions (360 degrees).
[1198] Analysis means
[1199] The analysis means analyzes the sound data from the sound collection means in real time and identifies the siren sounds specific to emergency vehicles. This analysis uses machine learning algorithms to detect specific frequency bands and sound patterns of siren sounds. The main software used is Python's "librosa" and "keras."
[1200] Detection Method
[1201] The detection means detects the approach of an emergency vehicle based on the analysis results obtained from the analysis means, and calculates the direction and distance of the sound source based on the difference in sound arrival times detected by the multiple microphones.
[1202] Volume adjustment means
[1203] The volume adjustment means uses a Bluetooth connection to automatically reduce the volume of audio sources (such as music or audiobooks) in the car.
[1204] Notification means
[1205] The notification means generates a specific announcement when an emergency vehicle is approaching and notifies the user through an in-vehicle speaker, the announcement including the direction and distance of the emergency vehicle.
[1206] Emotion Engine
[1207] The emotion engine analyzes the user's emotional state by analyzing facial recognition data and voice tone. The main software used is Python's "opencv" and "keras."
[1208] System Operation
[1209] The server analyzes the sound data sent from the sound collection means and identifies the siren sound. After the analysis, it calculates the direction and distance of the emergency vehicle and feeds the results back to the device. The device then captures an image of the user's face using the in-car camera and analyzes it using the emotion engine. Based on the analysis results, the notification means announces the approach of an emergency vehicle.
[1210] The notification means obtains emergency vehicle location information from the navigation system and generates a specific announcement, the tone and content of which are adjusted according to the user's emotional state, and is played through the vehicle speakers.
[1211] Specific examples
[1212] As a specific example, the user collects sound in real time using microphone devices installed at the four corners of the vehicle and sends the data to a server. The server analyzes the siren sound and calculates the direction and distance of the emergency vehicle. The results are fed back to the device, which generates an announcement based on the emergency vehicle's location information. Furthermore, the device captures an image of the user's face using an in-car camera, analyzes it using an emotion engine, and the notification means adjusts the content and tone of the announcement.
[1213] Prompt Sentence Examples
[1214] "When providing voice notifications for approaching emergency vehicles in an autonomous vehicle, please generate notification messages that take into account the emotional state of the occupants. For example, if an emergency vehicle is approaching from the right direction and 100 meters ahead, the notification message should be delivered in a calm tone if the occupant is feeling stressed, and in a standard tone if the occupant is relaxed."
[1215] In this way, it is possible to quickly and accurately detect the approach of an emergency vehicle and notify the user according to their emotional state.
[1216] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1217] Step 1:
[1218] The device collects sound in real time using microphones installed at the four corners of the vehicle. The sound data is temporarily stored in a buffer. This makes it possible to effectively capture sound from all directions (360 degrees). The input is the sound picked up from around the vehicle, and the output is the sound data stored in the buffer.
[1219] Step 2:
[1220] The terminal packetizes the contents of the buffer at regular intervals and sends the sound data to the server. The server receives the sound data and passes it to the analysis means. The input is the sound data stored in the buffer, and the output is the sound data uploaded to the server.
[1221] Step 3:
[1222] The server applies machine learning algorithms to the received sound data to identify specific frequency bands and patterns of siren sounds. The software used is the Python libraries "librosa" and "keras." The input is the sound data uploaded to the server, and the output is the determination of whether or not a siren is being heard.
[1223] Step 4:
[1224] If the server detects a siren, it feeds the result back to the device. It then calculates the direction and distance of the emergency vehicle based on the analysis results. The input is the result of the siren detection, and the output is data on the direction and distance of the emergency vehicle.
[1225] Step 5:
[1226] Based on the detection results, the device sends an instruction via Bluetooth to reduce the volume of the sound source inside the car by a certain percentage (e.g., 50%). The input is data about the direction and distance of the emergency vehicle, and the output is a volume adjustment instruction. Specifically, the device sends an instruction to the car's audio system using Bluetooth.
[1227] Step 6:
[1228] The device captures the user's facial image using an onboard camera and analyzes it with an emotion engine. The software used is Python's "opencv" and "keras." The input is the captured facial image, and the output is the user's emotional state.
[1229] Step 7:
[1230] The server obtains the location and direction of the emergency vehicle from the navigation system and generates a specific announcement. The input is the data on the direction and distance of the emergency vehicle and the user's emotional state, and the output is the generated announcement.
[1231] Step 8:
[1232] The device adjusts the content and tone of the announcement based on the analysis results of the emotion engine according to the user's emotional state. For example, if the user is feeling stressed, the announcement will be made in a calmer tone. The input is the generated announcement text, and the output is the adjusted announcement text.
[1233] Step 9:
[1234] The terminal plays the generated voice announcement over the in-car speaker, prompting the user to take specific action (e.g., "There is an ambulance approaching 100 meters ahead from the right, please pull over to the side of the road.") The input is the adjusted announcement text, and the output is a voice notification from the in-car speaker.
[1235] In this way, through each step, it is possible to quickly and accurately detect the approach of an emergency vehicle and notify the user appropriately according to their emotional state.
[1236] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1237] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1238] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1239] [Fourth embodiment]
[1240] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1241] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1242] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1243] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1244] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1245] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1246] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1247] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1248] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1249] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1250] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1251] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1252] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1253] As an example of the present invention, an emergency vehicle approach detection system for a vehicle will be described. This system is composed of a sound collection means, an analysis means, a detection means, a volume adjustment means, and a notification means, and quickly and accurately detects the approach of an emergency vehicle and notifies the user.
[1254] overview
[1255] Recording Method:
[1256] Microphone devices installed at the four corners of the vehicle collect ambient sounds, effectively capturing sounds coming from all directions in 360 degrees.
[1257] Analysis method:
[1258] The system analyzes sound data from sound collection devices in real time to identify the siren sounds specific to emergency vehicles, using machine learning algorithms to detect specific frequency bands and sound patterns of siren sounds.
[1259] Detection Methods:
[1260] Based on the analysis results obtained by the analysis means, the system detects the approach of an emergency vehicle and calculates the direction and distance of the sound source based on the difference in arrival time of the sounds from multiple microphones.
[1261] Volume adjustment means:
[1262] If the detection means detects the approach of an emergency vehicle, the volume of any audio source (e.g., music, audiobooks, etc.) being played in the vehicle will be automatically reduced via Bluetooth connection.
[1263] Means of notification:
[1264] Specific announcements will be made over the vehicle's speakers to passengers informing them of the location and approach of emergency vehicles.
[1265] Program processing
[1266] Sound pickup and data transmission
[1267] Terminal: Collects sound data in real time from microphones installed at the four corners of the vehicle and temporarily stores it in a buffer.
[1268] Server: Receives sound data periodically sent from the terminal.
[1269] Siren sound analysis
[1270] Server: Analyzes the received sound data and applies machine learning algorithms to detect frequency bands and sound patterns to identify the sound of an emergency vehicle siren.
[1271] Terminal: Receives the analysis results from the server and prepares any necessary next steps.
[1272] Detecting approaching emergency vehicles
[1273] Terminal: Using data from multiple microphones, a sound source localization algorithm calculates the difference in sound arrival time and determines the direction and distance of the emergency vehicle.
[1274] Automatic volume adjustment
[1275] Device: Based on the detection results, it sends a command via Bluetooth connection to reduce the volume of the sound source inside the car to 50%.
[1276] Generate and play notifications
[1277] Server: Obtains the current location and direction of the emergency vehicle based on data from the navigation system and generates an appropriate announcement.
[1278] Terminal: The generated voice announcement is played over the car's speakers.
[1279] Specific examples
[1280] 1. Sound collection means:
[1281] The device collects sound in real time using microphones installed at the four corners of the vehicle and stores it in a buffer.
[1282] 2. Analysis method:
[1283] The server analyzes the sound data in the buffer using a machine learning algorithm to identify the sound of a siren.
[1284] 3. Detection methods:
[1285] Based on the analysis results from the server, the device calculates the direction and distance of the sound source.
[1286] 4. Volume adjustment means:
[1287] The server detects the approach of an emergency vehicle and the device sends a command via Bluetooth connection to reduce the music volume to 50%.
[1288] 5. Means of notification:
[1289] The server uses data from the navigation system to obtain the location information of the emergency vehicle and generates specific instructions.
[1290] The terminal generates a voice announcement that is played over the car's speakers, prompting the user to take specific action (e.g., there is an emergency vehicle approaching 100 meters ahead from the right; move to the shoulder of the road).
[1291] summary
[1292] This system allows even passengers who are enjoying loud music in their cars to quickly and reliably detect approaching emergency vehicles and respond safely, thereby reducing the risk of accidents and supporting smooth traffic management.
[1293] The processing flow will be explained below.
[1294] Step 1:
[1295] Terminal: Sound is collected by multiple microphones installed at the four corners of the vehicle. The collected sound data is temporarily stored in a buffer.
[1296] Specific operation: Sound data is sampled through the microphone interface, for example, every second, and stored in a buffer. Data from each microphone is stored in a separate buffer.
[1297] Step 2:
[1298] Terminal: Collected sound data is packetized at regular intervals (e.g., every second) and sent to the server.
[1299] Specific operation: The sound data is divided into packets and sent to the server using a network protocol. After sending, the buffer is cleared and waits for the next sample.
[1300] Step 3:
[1301] Server: Analyzes the received sound data in real time, using AI to identify specific frequency bands and patterns of siren sounds.
[1302] How it works: A machine learning model (e.g., a convolutional neural network) is applied to analyze sound data and extract features. If a specific pattern is detected, it is identified as a siren sound.
[1303] Step 4:
[1304] Server: If a siren sound is detected, the result is fed back to the device.
[1305] Specific operation: A data packet containing the analysis results is generated and sent to the terminal via the network.
[1306] Step 5:
[1307] Terminal: Based on the analysis results received from the server, the terminal calculates the difference in sound arrival time from multiple microphones and determines the direction and distance of the sound source.
[1308] Specific operation: An algorithm is applied that uses the difference in sound arrival time to calculate the direction and distance of the sound source using techniques such as triangulation.
[1309] Step 6:
[1310] Terminal: If an approaching emergency vehicle is detected, a command to reduce the volume of the in-car speakers to a certain percentage (e.g., 50%) is sent via Bluetooth.
[1311] Specific behavior: Uses the Bluetooth speaker's volume control protocol to send a command to decrease the volume of the currently playing music or voice by the specified percentage.
[1312] Step 7:
[1313] Server: Works with the navigation system to obtain the current location and direction of emergency vehicles.
[1314] Specific operation: Calls the navigation system's API to obtain the current GPS location and direction of the emergency vehicle. Based on this information, it generates an announcement for the next step.
[1315] Step 8:
[1316] Terminal: Generates specific voice announcements based on location information received from the server.
[1317] What it does: Creates the appropriate announcement using a text generation engine and generates an audio file using a text-to-speech engine.
[1318] Step 9:
[1319] Terminal: The generated voice announcement is played over the vehicle's speakers, providing specific instructions to the user.
[1320] Specific operation: The created audio file is added to the playback queue of the car speakers and begins playing immediately.
[1321] Step 10:
[1322] User: Follow the voice announcement and take appropriate action promptly.
[1323] Specific actions: For example, "An ambulance is approaching from the right, 100 meters ahead. Please move to the shoulder," and follow the instructions to safely move the vehicle to the shoulder.
[1324] Example 1
[1325] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1326] Conventional vehicle emergency vehicle approach detection systems lack the functionality to quickly detect approaching emergency vehicles and notify the driver. In particular, drivers listening to loud music or audiobooks in the vehicle may not notice the approach of an emergency vehicle, increasing the risk of an accident. The objective of this invention is to solve this problem by providing a system that quickly and accurately detects the approach of an emergency vehicle and notifies the driver in an appropriate manner.
[1327] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1328] In this invention, the server includes a sound collection means, a sound data analysis means, an emergency vehicle detection means, an automatic volume adjustment means, a notification means, and a notification generation means. This makes it possible to analyze in real time the sound data collected by the sound collection means installed at the four corners of the vehicle and quickly detect the approach of an emergency vehicle. Furthermore, when an emergency vehicle approaches, the server can automatically lower the volume inside the vehicle and generate an appropriate announcement to notify the driver.
[1329] The "sound collection means" is a device installed at each of the four corners of the vehicle that collects surrounding sounds in real time.
[1330] The "sound data analysis means" is a function that analyzes collected sound data, detects specific frequency bands and patterns, and identifies the siren sounds of emergency vehicles.
[1331] The "emergency vehicle detection means" is a function that detects the approach of an emergency vehicle based on the analysis results obtained by the sound data analysis means.
[1332] The "automatic volume adjustment means" is a function that automatically reduces the volume of the sound source being played inside the vehicle when the emergency vehicle detection means detects the approach of an emergency vehicle.
[1333] The "notification means" is a function for informing the vehicle occupants of the location and approach of an emergency vehicle.
[1334] The "notification generating means" is a function that obtains the location of emergency vehicles based on data from the navigation system and generates an appropriate announcement.
[1335] The present invention relates to a system for detecting approaching emergency vehicles in a vehicle, and specific embodiments thereof are described below. This system is composed of a sound collection means, a sound data analysis means, an emergency vehicle detection means, an automatic volume adjustment means, a notification means, and a notification generation means. By combining these means, it is possible to quickly and accurately detect the approach of an emergency vehicle and notify occupants in the vehicle.
[1336] Sound collection means
[1337] Microphones installed at the four corners of the vehicle act as sound collectors. These microphones collect ambient sounds in real time and temporarily store the data in a buffer. For example, when the vehicle approaches an intersection, each microphone collects the surrounding sounds.
[1338] Sound data analysis method
[1339] The server receives audio data periodically sent from the device. It analyzes the received audio data using a machine learning algorithm to detect specific frequency bands and sound patterns and identify the sound of an emergency vehicle siren. The server then sends the analysis results to the device and prepares for the next process.
[1340] Emergency vehicle detection means
[1341] Based on the analysis results sent from the server, the device uses data from multiple microphones and a sound source localization algorithm to calculate the direction and distance of the emergency vehicle. For example, it can determine that an emergency vehicle is approaching 100 meters behind and to the left.
[1342] Automatic volume adjustment means
[1343] If the device detects an approaching emergency vehicle, it will send a command over the Bluetooth connection to reduce the volume of the in-car audio source (e.g., music) to 50%, allowing the in-car audio system to automatically reduce the current volume.
[1344] Notification means and notification generation means
[1345] The server obtains data on the current location and direction of the emergency vehicle from the navigation system and generates an appropriate announcement for the user. The device receives this announcement and plays it through the car's speakers. For example, it can provide specific instructions to the user, such as "Police vehicle approaching 100 meters behind you on the left, please be careful."
[1346] Prompt Sentence Examples
[1347] Here are some examples of prompts:
[1348] "Please explain the process by which microphones installed at the four corners of the vehicle collect ambient sounds in real time and store that data in a buffer."
[1349] "Please explain in detail how the server uses a machine learning algorithm to analyze the sound data received from the device and identify the sound of a siren."
[1350] "Please explain in detail the algorithm that allows the device to use data from multiple microphones to calculate the direction and distance of emergency vehicles."
[1351] "Please explain the steps to reduce the volume of in-car audio sources to 50% when the device detects an approaching emergency vehicle."
[1352] "Please explain the procedure for the server to generate an announcement to notify the user of the approach of an emergency vehicle and send it to the car speaker."
[1353] In this way, the system of the present invention can quickly detect the approach of an emergency vehicle, provide the user with the necessary information, and encourage appropriate action.
[1354] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1355] Step 1: Sound collection and data transmission
[1356] Terminal
[1357] Input: Microphones installed at the four corners of the vehicle collect surrounding sounds in real time.
[1358] Data processing / data calculation: Collected sound data is temporarily stored in a buffer and prepared to be sent to the server at regular intervals.
[1359] Output: Sends the sound data stored in the temporary buffer to the server.
[1360] Specific operation: For example, when a vehicle is driving and approaches an intersection, the microphone collects sounds of nearby pedestrians and other vehicles and sends the data to a server at regular intervals.
[1361] Step 2: Analyzing the siren sound
[1362] server
[1363] Input: Sound data sent periodically from the device.
[1364] Data processing / data calculation: Applying machine learning algorithms to analyze sound data and detect specific frequency bands and sound patterns to identify the sound of an emergency vehicle siren.
[1365] Output: As an analysis result, information is generated as to whether or not an emergency vehicle siren was detected.
[1366] Specific operation: The server analyzes the sound data received from the device, and if it detects a specific frequency pattern between 700 and 1500 Hz, it identifies it as the sound of an emergency vehicle siren. The analysis result is returned to the device.
[1367] Step 3: Detect approaching emergency vehicles
[1368] Terminal
[1369] Input: Analysis results sent from the server.
[1370] Data processing / data calculation: Based on the difference in arrival time of sound data obtained from multiple microphones, a sound source localization algorithm is used to calculate the direction and distance of the emergency vehicle.
[1371] Output: Information on the approaching direction and distance of the emergency vehicle.
[1372] Specific operation: For example, identify that an emergency vehicle is approaching 100 meters behind and to the left. This information is used in the next step.
[1373] Step 4: Automatically adjust the volume
[1374] Terminal
[1375] Input: Information about the approaching direction and distance of the emergency vehicle, and information about the audio source currently being played.
[1376] Data processing / data calculation: If an approaching emergency vehicle is detected, a command is generated to reduce the volume of the sound source (e.g., music) inside the vehicle to 50%.
[1377] Output: A command to decrease the volume over the Bluetooth connection.
[1378] Specific behavior: For example, if music is playing in the car, the volume will be automatically reduced to 50%. This operation is performed via Bluetooth connection.
[1379] Step 5: Generate and play notifications
[1380] server
[1381] Input: Emergency vehicle approach direction and distance information, and location data from the navigation system.
[1382] Data processing / data calculation: Based on data from the navigation system, appropriate announcements are generated for users.
[1383] Output: The generated announcement.
[1384] Specific actions: For example, generate specific instructions such as "A police vehicle is approaching 100 meters behind you on the left. Be careful."
[1385] Terminal
[1386] Input: Server-generated announcement.
[1387] Data processing / data calculation: Generates commands to play announcements over the car speakers.
[1388] Output: Announcement command to the car speaker.
[1389] Specific operation: The generated voice announcement is played over the car's speakers to notify the user of the approaching emergency vehicle and how to respond. For example, it may say, "Siren approaching, emergency vehicle on the left."
[1390] (Application example 1)
[1391] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1392] Conventional vehicle emergency vehicle notification systems alert users to approaching emergency vehicles by lowering the volume inside the vehicle, but lack the technology for use outside the vehicle or on mobile devices. Furthermore, while it is necessary to quickly and accurately notify users of the direction and distance of approaching emergency vehicles visually and audibly, no appropriate solution for this purpose exists. Therefore, there is a need for a comprehensive emergency vehicle notification system that allows users to respond safely both inside and outside the vehicle.
[1393] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1394] In this invention, the server includes a sound collection means, an analysis means, a detection means, a volume adjustment means, and a notification means. This allows the server to quickly and accurately detect the approach of an emergency vehicle by collecting surrounding sounds, even when using a vehicle or a mobile device, automatically lower the volume of the sound source being played, and effectively notify the user of the location and approach of the emergency vehicle. Specifically, the analysis means calculates the direction and distance of the approaching emergency vehicle, and the information is provided via voice and screen display, allowing the user to quickly take appropriate action.
[1395] A "sound collection means" is a device installed to collect surrounding sounds.
[1396] An "analysis means" is a device or algorithm for analyzing collected sound data and identifying specific sound patterns.
[1397] The "detection means" is a device or function for detecting the approach of a specific sound source, for example, an emergency vehicle, based on the analysis result of the analysis means.
[1398] The "volume adjustment means" is a device or function that automatically adjusts the volume of the sound source being played back when the detection means detects the approach of an emergency vehicle.
[1399] "Notification means" refers to a device or function for informing users of the location and approach of an emergency vehicle.
[1400] A "mobile terminal" is a communication device or computer device that can be carried by a user.
[1401] The "audio output means" is a device for conveying the audio information generated by the notification means to the user.
[1402] "Screen display" refers to the function of displaying textual or graphic information on a mobile terminal or other display device.
[1403] The present invention provides a system that uses a vehicle or a mobile terminal to quickly detect the approach of an emergency vehicle and automatically reduce the volume of the sound source being played. Specific embodiments for carrying out the present invention are described below.
[1404] overview
[1405] This system consists of a sound collection means, an analysis means, a detection means, a volume adjustment means, and a notification means. The sound collection means is a microphone installed in a vehicle or a mobile device that collects surrounding sounds. The analysis means analyzes the collected sound data and identifies specific sound patterns, such as siren sounds. This analysis uses a machine learning algorithm (e.g., a model built with Keras). The detection means detects the approach of an emergency vehicle based on the analysis results. The volume adjustment means automatically lowers the volume of the sound being played when the detection means detects the approach of an emergency vehicle. The notification means effectively notifies the user of the location and approach of the emergency vehicle.
[1406] Program processing
[1407] Sound pickup and data transmission
[1408] The device collects ambient sounds through a microphone, and the collected sound data is temporarily stored in a buffer and processed by an analysis means.
[1409] Siren sound analysis
[1410] The server receives the sound data in the buffer and applies a machine learning algorithm to identify siren sounds. Keras is used as an analysis tool to analyze the collected sound data.
[1411] Detecting approaching emergency vehicles
[1412] The detection means detects the approach of an emergency vehicle based on the results of the analysis means, and calculates the direction and distance of the sound source from the collected sound data.
[1413] Volume adjustment
[1414] The volume control means automatically reduces the volume of the audio source being played when detected, and sends a volume control command via Bluetooth or another communication means.
[1415] Generate and play notifications
[1416] The server generates an appropriate announcement based on the location information of the emergency vehicle, and the notification means notifies the user of this announcement both by voice and by display on the screen.
[1417] Specific examples
[1418] While the user is driving, the device (smartphone) collects surrounding sounds using a built-in microphone and sends them to a server. The server analyzes the sound data in real time, and if a siren sound is detected, it automatically lowers the volume of the smartphone via Bluetooth and notifies the user of the location of an emergency vehicle via a display and voice message, allowing the user to take prompt and appropriate action.
[1419] Prompt Sentence Examples
[1420] "We are developing an emergency vehicle approach detection system for vehicles as a smartphone application. Please generate a Python program that follows the requirements below:
[1421] Built-in microphone for sound capture
[1422] Uses machine learning models to analyze and detect siren sounds in real time
[1423] Volume adjustment and notification functions implemented
[1424] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1425] Step 1:
[1426] The device collects ambient sounds. While the user is driving the vehicle, the device's microphone collects audio data in real time and temporarily stores it in a buffer. The input here is the ambient sounds, and the output is the audio data in the buffer.
[1427] Step 2:
[1428] The device sends the sound data in the buffer to the server. The collected sound data is transferred to the server at regular intervals. The input here is the sound data in the buffer, and the output is the sound data sent to the server.
[1429] Step 3:
[1430] The server analyzes the received sound data. The server uses a machine learning model (a model using Keras) to analyze the sound data and detect siren sounds. The input here is the sound data sent to the server, and the output is the analysis result.
[1431] Step 4:
[1432] The server detects emergency vehicles based on the analysis results. If a siren sound is detected, the server detects the approach of an emergency vehicle and calculates its direction and distance. The input here is the analysis result, and the output is the direction and distance information of the emergency vehicle.
[1433] Step 5:
[1434] The server sends a volume adjustment command to the device. Based on the detection result, the server sends an instruction to the device to lower the volume of the sound being played. The input here is the direction and distance information of the emergency vehicle, and the output is the volume adjustment command to the device.
[1435] Step 6:
[1436] The device performs the volume adjustment. Upon receiving the volume adjustment command, the device automatically reduces the volume of the audio source using Bluetooth or other communication methods. The input here is the volume adjustment command, and the output is the reduced volume.
[1437] Step 7:
[1438] The server generates a notification and sends it to the device. Based on the location information of the emergency vehicle, the server generates a notification for voice and screen display and sends it to the device. The input here is the direction and distance information of the emergency vehicle, and the output is the generated notification data.
[1439] Step 8:
[1440] The terminal notifies the user. The terminal that receives the notification data notifies the user of the approach of an emergency vehicle through voice and screen display. The input here is the notification data, and the output is the notification to the user.
[1441] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1442] As one embodiment of the present invention, we will describe an in-vehicle emergency vehicle approach detection system and a notification system that takes into account the user's emotional state. This system is composed of a sound collection means, an analysis means, a detection means, a volume adjustment means, a notification means, and an emotion engine, and quickly and accurately detects the approach of an emergency vehicle and notifies the user appropriately according to the user's emotional state.
[1443] overview
[1444] Recording Method:
[1445] Microphone devices installed at the four corners of the vehicle collect surrounding sounds, effectively capturing sounds from all directions in 360 degrees.
[1446] Analysis method:
[1447] This is an analysis method that analyzes sound data from sound collection means in real time to identify the siren sounds specific to emergency vehicles. This analysis uses machine learning algorithms to detect specific frequency bands and sound patterns of siren sounds.
[1448] Detection Methods:
[1449] Based on the analysis results obtained by the analysis means, the system detects the approach of an emergency vehicle and calculates the direction and distance of the sound source based on the difference in sound arrival time between multiple microphones.
[1450] Volume adjustment means:
[1451] If the detection means detects the approach of an emergency vehicle, the volume of any audio playing in the car (e.g., music, audiobooks, etc.) will be automatically reduced. This is achieved using a Bluetooth connection.
[1452] Means of notification:
[1453] A specific announcement is generated for approaching emergency vehicles and notified to the user through the in-car speakers, including the direction and distance of the emergency vehicle.
[1454] Emotion Engine:
[1455] It is an engine that analyzes the user's emotional state. It analyzes facial recognition data or voice tone to understand the user's emotional state, such as whether they are stressed or relaxed.
[1456] Program processing
[1457] Sound pickup and data transmission
[1458] Terminal: Microphones installed at the four corners of the vehicle collect sound in real time and temporarily store it in a buffer. The contents of the buffer are packetized at regular intervals and sent to the server.
[1459] Siren sound analysis
[1460] Server: Apply machine learning algorithms to the received sound data to identify specific frequency bands and patterns of siren sounds.
[1461] Detecting approaching emergency vehicles
[1462] Server: If a siren sound is detected, the result is fed back to the device.
[1463] Terminal: Based on the analysis results fed back, the direction and distance of the sound source are determined using the difference in sound arrival time from multiple microphones.
[1464] Automatic volume adjustment
[1465] Terminal: Based on the detection results, it sends an instruction via Bluetooth to reduce the volume of the sound source inside the car to a certain percentage (e.g., 50%).
[1466] Emotional state analysis
[1467] Device: The emotion engine uses facial recognition data captured by the in-car camera and the tone of the user's voice collected by the microphone to identify the user's emotional state.
[1468] Generate and play notifications
[1469] Server: Obtains the location and direction of emergency vehicles from the navigation system and generates specific announcement text.
[1470] Terminal: Based on the analysis results of the emotion engine, the content and tone of the announcement are adjusted according to the user's emotional state.
[1471] Emotion-based notifications
[1472] Terminal: The generated voice announcement is played over the car's speakers. Depending on the user's emotional state, the announcement can be calm or emphasize urgency.
[1473] Specific examples
[1474] 1. Sound collection means:
[1475] The device collects sound in real time using microphones installed at the four corners of the vehicle and stores it in a buffer.
[1476] 2. Analysis method:
[1477] The server analyzes the sound data in the buffer using a machine learning algorithm to identify the sound of a siren.
[1478] 3. Detection methods:
[1479] Based on the analysis results from the server, the device calculates the direction and distance of the sound source.
[1480] 4. Volume adjustment means:
[1481] The server detects the approach of an emergency vehicle and the device sends a command via Bluetooth connection to reduce the music volume to 50%.
[1482] 5. Emotion Engine:
[1483] The device recognizes the user's face using the onboard camera, and the emotion engine analyzes the data to understand the user's emotional state.
[1484] 6. Means of notification:
[1485] The server uses data from the navigation system to obtain the location information of the emergency vehicle and generates specific instructions.
[1486] The tone and content of the announcements are adjusted based on the results of the emotion engine. For example, if the user is feeling stressed, the announcements will be made in a calmer tone.
[1487] The terminal generates a voice announcement that is played over the car's speakers, prompting the user to take specific action (e.g., "There is an ambulance approaching 100 meters ahead from the right, so pull over to the side of the road").
[1488] summary
[1489] This system allows even those listening to loud music in their cars to quickly and reliably detect the approach of an emergency vehicle and provide appropriate notifications based on the user's emotional state, allowing them to safely take appropriate action and contributing to traffic management and accident prevention.
[1490] The processing flow will be explained below.
[1491] Step 1:
[1492] Terminal: Sound is collected by multiple microphones installed at the four corners of the vehicle. The collected sound data is temporarily stored in a buffer.
[1493] Specific operation: Sound data is sampled through the microphone interface, for example, every second, and stored in a buffer. Data from each microphone is stored in a separate buffer.
[1494] Step 2:
[1495] Terminal: Collected sound data is packetized at regular intervals (e.g., every second) and sent to the server.
[1496] Specific operation: The sound data is divided into packets and sent to the server using a network protocol. After sending, the buffer is cleared and waits for the next sample.
[1497] Step 3:
[1498] Server: Analyzes the received sound data in real time, using AI to identify specific frequency bands and patterns of siren sounds.
[1499] How it works: A machine learning model (e.g., a convolutional neural network) is applied to analyze sound data and extract features. If a specific pattern is detected, it is identified as a siren sound.
[1500] Step 4:
[1501] Server: If a siren sound is detected, the result is fed back to the device.
[1502] Specific operation: A data packet containing the analysis results is generated and sent to the terminal via the network.
[1503] Step 5:
[1504] Terminal: Based on the analysis results received from the server, the terminal calculates the difference in sound arrival time from multiple microphones and determines the direction and distance of the sound source.
[1505] Specific operation: An algorithm is applied that uses the difference in sound arrival time to calculate the direction and distance of the sound source using techniques such as triangulation.
[1506] Step 6:
[1507] Terminal: If an approaching emergency vehicle is detected, a command to reduce the volume of the in-car speakers to a certain percentage (e.g., 50%) is sent via Bluetooth.
[1508] Specific behavior: Uses the Bluetooth speaker's volume control protocol to send a command to decrease the volume of the currently playing music or voice by the specified percentage.
[1509] Step 7:
[1510] On the device: The emotion engine analyzes the user's facial recognition data from the in-car camera or the user's tone of voice collected by the microphone.
[1511] Specific behavior: Apply facial recognition and voice analysis algorithms to identify the user's emotional state (e.g., stress, relaxed).
[1512] Step 8:
[1513] Server: Works with the navigation system to obtain the current location and direction of emergency vehicles.
[1514] Specific operation: Calls the navigation system's API to obtain the current GPS location and direction of the emergency vehicle. Based on this information, it generates an announcement for the next step.
[1515] Step 9:
[1516] Terminal: Generates specific voice announcements based on the emergency vehicle location information received from the server and the analysis results of the emotion engine.
[1517] What it does: It uses a text generation engine to create an appropriate announcement, then uses a text-to-speech engine to generate an audio file, adjusting the tone depending on the user's emotional state.
[1518] Step 10:
[1519] Terminal: The generated voice announcement is played over the vehicle's speakers, providing specific instructions to the user.
[1520] Specific behavior: The created audio file is added to the playback queue of the car speakers and begins playing immediately. Based on the user's emotional state, the announcement is made in a calm manner or with an emphasis on urgency.
[1521] Step 11:
[1522] User: Follow the voice announcement and take appropriate action promptly.
[1523] Specific actions: For example, "An ambulance is approaching from the right, 100 meters ahead. Please move to the shoulder," and follow the instructions to safely move the vehicle to the shoulder.
[1524] Example 2
[1525] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1526] Conventional vehicle systems have the problem of not being able to take into account the emotional state of the user when notifying them of an approaching emergency vehicle. This can make it difficult for users to take appropriate action in stressful situations, potentially delaying their response to the emergency vehicle. Another issue is that the volume of the audio played in the vehicle must be adjusted manually, which requires the user to divert their attention.
[1527] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server is a processing device for analyzing sound data collected from a sound collection device, and includes an analysis device that monitors surrounding sounds based on the generated analysis results, a detection device that detects the approach of an emergency vehicle based on the analysis results, and an emotion recognition device that collects and analyzes the emotional state of the user. This makes it possible to quickly and accurately detect the approach of an emergency vehicle and to provide an appropriate notification according to the emotional state of the user.
[1528] A "sound collection device" is a device installed at the four corners of a vehicle to collect surrounding sounds.
[1529] The "analysis device" is a device that analyzes collected sound data and monitors surrounding sounds based on the generated analysis results.
[1530] The "detection device" is a device for detecting the approach of an emergency vehicle based on the analysis results generated by the analysis device.
[1531] The "volume adjustment device" is a device that automatically reduces the volume of the sound source being played inside the vehicle when the approach of an emergency vehicle is detected by the detection device.
[1532] A "notification device" is a device that notifies in-vehicle users of the location and approach of an emergency vehicle.
[1533] An "emotion recognition device" is a device for collecting and analyzing a user's emotional state.
[1534] The present invention relates to an in-vehicle emergency vehicle approach detection system and a notification system that takes into account the emotional state of the user. The system is composed of sound collection devices installed at the four corners of the vehicle, an analysis device that analyzes the collected sound data, a detection device that detects the approach of an emergency vehicle based on the analysis results, a volume adjustment device that automatically lowers the volume of the sound source being played in the vehicle when the approach of the emergency vehicle is detected, a notification device that notifies the in-vehicle user of the location and approach of the emergency vehicle, and an emotion recognition device that collects and analyzes the user's emotional state.
[1535] System configuration
[1536] Sound pickup device
[1537] The device includes microphones installed at the four corners of the vehicle, which collect ambient sounds in real time and temporarily store them in a buffer. At regular intervals, the contents of the buffer are sent to a server as data packets.
[1538] analysis device
[1539] The server analyzes the sound data received from the device using a generative AI model (e.g., TensorFlow) to identify specific frequency bands and sound patterns and identify the sound of a siren.
[1540] Detection device
[1541] When the server detects a siren, it sends the result back to the device, which then uses the difference in sound arrival times from multiple microphones to calculate the direction and distance of the sound source.
[1542] volume adjustment device
[1543] Based on the detection results, the device will automatically reduce the volume of in-car audio sources (e.g., radio, music player) to a certain percentage (e.g., 50%) via Bluetooth.
[1544] emotion recognition device
[1545] The device uses an onboard camera and microphone to collect facial recognition data and voice tone, and then uses an emotion engine to analyze the user's emotional state, determining whether the user is stressed or relaxed.
[1546] notification device
[1547] The server obtains the location and heading of the emergency vehicle from the navigation system and uses a generative AI model to generate a specific announcement. This announcement includes tone and content based on the user's emotional state. The device plays the generated voice announcement over the car's speakers, encouraging the user to take specific action.
[1548] Specific examples
[1549] Here is a concrete example of how the system works:
[1550] 1. Sound pickup devices are installed at the four corners of the vehicle to collect sound in real time, and this data is periodically sent to a server.
[1551] 2. The server analyzes the sound data and uses a generative AI model to identify specific frequency bands and detect siren sounds.
[1552] 3. The server sends the analysis results back to the device, which then calculates the direction and distance of the sound source.
[1553] 4. If an approaching emergency vehicle is detected, the device will send a command via Bluetooth to halve the volume of sound sources inside the vehicle.
[1554] 5. The terminal analyzes the user's emotional state using an emotion engine, and the server generates a specific announcement based on the navigation data.
[1555] 6. The generated announcement is played by the terminal through the in-car speaker, and an announcement such as "There is an ambulance approaching 100 meters ahead from the right, so please pull over to the side of the road" is made.
[1556] Prompt Sentence Examples
[1557] Please generate a natural language document that describes an in-vehicle emergency vehicle approach detection system and a notification system that takes into account the user's emotional state. Please provide a detailed description of the processing flow, including specific hardware and software, including sound collection devices, analysis devices, detection devices, volume control devices, notification devices, and emotion recognition devices.
[1558] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1559] Step 1: Recording and transmitting data
[1560] Terminal: Microphones installed at the four corners of the vehicle collect surrounding sounds in real time and temporarily store the data in a buffer. The sound data in the buffer is divided into packets at regular intervals (e.g., every 10 seconds) and sent to the server.
[1561] Input: Sound data collected from microphones at the four corners of the vehicle
[1562] Output: Sound data sent to the server in data packet format
[1563] Step 2: Analyzing the siren sound
[1564] Server: Applying a generative AI model (e.g., TensorFlow) to the received sound data to identify specific frequency bands and patterns of siren sounds. The results of this analysis are passed on to the next processing step.
[1565] Input: Sound data sent from the device
[1566] Output: Siren sound analysis results (specific frequency bands and patterns)
[1567] Step 3: Detect approaching emergency vehicles
[1568] Server: When the siren sound is identified, the server feeds back the result to the device. Based on the feedback, the device calculates the difference in sound arrival time from multiple microphones and specifically calculates the direction and distance of the sound source.
[1569] Input: Analysis result of siren sound
[1570] Output: direction and distance of sound source
[1571] Step 4: Automatically adjust the volume
[1572] Device: When an approaching emergency vehicle is detected, the device sends a command via Bluetooth to automatically lower the volume of the currently playing audio source (e.g., radio, music player), for example, by 50%.
[1573] Input: Proximity information based on the direction and distance of the sound source
[1574] Output: Volume down command via Bluetooth
[1575] Step 5: Analyze your emotional state
[1576] Device: Analyzes the user's facial recognition data acquired by the in-car camera and the tone of voice collected by the microphone, and uses a generative AI model to determine the user's emotional state (e.g., stressed, relaxed).
[1577] Input: User's facial recognition data and tone of voice
[1578] Output: User's emotional state
[1579] Step 6: Generate and play notifications
[1580] Server: Obtains the location and heading of the emergency vehicle from the navigation system, and uses a generative AI model to create a specific announcement based on that information. This announcement is adjusted according to the user's emotional state.
[1581] Input: Emergency vehicle location information, user emotional state
[1582] Output: Specific announcement text
[1583] Step 7: Emotional Notifications
[1584] Terminal: Receives announcements generated by the server and plays them in an appropriate tone through the car's speakers. The content of the notification is adjusted according to the user's emotional state, encouraging specific countermeasures.
[1585] Input: Specific announcement text
[1586] Output: Voice notification through car speakers
[1587] (Application example 2)
[1588] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1589] For self-driving vehicles, it is important to quickly and accurately detect the approach of an emergency vehicle and notify the user appropriately, but conventional systems have difficulty accurately detecting the direction and distance of an emergency vehicle and lack the ingenuity to enable the user to respond calmly according to the situation.In addition, there is no notification system that takes into account the user's emotional state, which makes it difficult for the user to respond appropriately if they are feeling stressed.
[1590] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes sound collection means installed at the four corners of the vehicle, analysis means for analyzing collected sound data and monitoring surrounding sounds, detection means for detecting the approach of an emergency vehicle based on the analysis results, volume adjustment means for automatically lowering the volume of the sound source being played in the vehicle when the detection means detects the approach of an emergency vehicle, notification means for notifying users in the vehicle of the location and approach of the emergency vehicle, an emotion engine for analyzing the emotional state of the user, and means for adjusting the content and tone of the notification based on the analysis results of the emotion engine. This makes it possible to quickly and accurately detect the approach of an emergency vehicle and to notify the user appropriately according to their emotional state.
[1591] The "sound collection means" is a device installed at the four corners of the vehicle to collect surrounding sounds.
[1592] The "analysis means" is a means for analyzing collected sound data and identifying the siren sound specific to emergency vehicles.
[1593] The "detection means" is a means for detecting the approach of an emergency vehicle based on the analysis results obtained by the analysis means.
[1594] The "volume adjustment means" is a means for automatically lowering the volume of the sound source being reproduced inside the vehicle when the detection means detects the approach of an emergency vehicle.
[1595] The "notification means" is a means for informing the vehicle occupants of the location and approach of an emergency vehicle.
[1596] "Emotion Engine" is an engine for analyzing the user's emotional state, such as analyzing facial recognition data or voice tone.
[1597] "Calculation by the analysis means" refers to calculation of the approaching direction and distance of the emergency vehicle based on the sound data collected from the sound collection means.
[1598] A "user's emotional state" is the emotional state that a user is feeling, identified based on a facial image or tone of voice.
[1599] The "means for adjusting the content and tone of the notification based on the analysis results" refers to means for changing the content and tone of the voice of the generated notification according to the analysis results of the emotional state by the emotion engine.
[1600] A system for realizing this invention includes sound collection means installed at the four corners of a vehicle, analysis means for analyzing collected sound data and monitoring surrounding sounds, detection means for detecting the approach of an emergency vehicle based on the analysis results, volume adjustment means for automatically lowering the volume of the sound source being played inside the vehicle when the detection means detects the approach of an emergency vehicle, notification means for notifying users inside the vehicle of the location and approach of the emergency vehicle, an emotion engine for analyzing the emotional state of the user, and means for adjusting the content and tone of the notification based on the analysis results of the emotion engine.
[1601] System Components
[1602] Sound collection means
[1603] The sound collection means consists of microphones installed at the four corners of the vehicle, which can effectively capture sounds from all directions (360 degrees).
[1604] Analysis means
[1605] The analysis means analyzes the sound data from the sound collection means in real time and identifies the siren sounds specific to emergency vehicles. This analysis uses machine learning algorithms to detect specific frequency bands and sound patterns of siren sounds. The main software used is Python's "librosa" and "keras."
[1606] Detection Method
[1607] The detection means detects the approach of an emergency vehicle based on the analysis results obtained from the analysis means, and calculates the direction and distance of the sound source based on the difference in sound arrival times detected by the multiple microphones.
[1608] Volume adjustment means
[1609] The volume adjustment means uses a Bluetooth connection to automatically reduce the volume of audio sources (such as music or audiobooks) in the car.
[1610] Notification means
[1611] The notification means generates a specific announcement when an emergency vehicle is approaching and notifies the user through an in-vehicle speaker, the announcement including the direction and distance of the emergency vehicle.
[1612] Emotion Engine
[1613] The emotion engine analyzes the user's emotional state by analyzing facial recognition data and voice tone. The main software used is Python's "opencv" and "keras."
[1614] System Operation
[1615] The server analyzes the sound data sent from the sound collection means and identifies the siren sound. After the analysis, it calculates the direction and distance of the emergency vehicle and feeds the results back to the device. The device then captures an image of the user's face using the in-car camera and analyzes it using the emotion engine. Based on the analysis results, the notification means announces the approach of an emergency vehicle.
[1616] The notification means obtains emergency vehicle location information from the navigation system and generates a specific announcement, the tone and content of which are adjusted according to the user's emotional state, and is played through the vehicle speakers.
[1617] Specific examples
[1618] As a specific example, the user collects sound in real time using microphone devices installed at the four corners of the vehicle and sends the data to a server. The server analyzes the siren sound and calculates the direction and distance of the emergency vehicle. The results are fed back to the device, which generates an announcement based on the emergency vehicle's location information. Furthermore, the device captures an image of the user's face using an in-car camera, analyzes it using an emotion engine, and the notification means adjusts the content and tone of the announcement.
[1619] Prompt Sentence Examples
[1620] "When providing voice notifications for approaching emergency vehicles in an autonomous vehicle, please generate notification messages that take into account the emotional state of the occupants. For example, if an emergency vehicle is approaching from the right direction and 100 meters ahead, the notification message should be delivered in a calm tone if the occupant is feeling stressed, and in a standard tone if the occupant is relaxed."
[1621] In this way, it is possible to quickly and accurately detect the approach of an emergency vehicle and notify the user according to their emotional state.
[1622] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1623] Step 1:
[1624] The device collects sound in real time using microphones installed at the four corners of the vehicle. The sound data is temporarily stored in a buffer. This makes it possible to effectively capture sound from all directions (360 degrees). The input is the sound picked up from around the vehicle, and the output is the sound data stored in the buffer.
[1625] Step 2:
[1626] The terminal packetizes the contents of the buffer at regular intervals and sends the sound data to the server. The server receives the sound data and passes it to the analysis means. The input is the sound data stored in the buffer, and the output is the sound data uploaded to the server.
[1627] Step 3:
[1628] The server applies machine learning algorithms to the received sound data to identify specific frequency bands and patterns of siren sounds. The software used is the Python libraries "librosa" and "keras." The input is the sound data uploaded to the server, and the output is the determination of whether or not a siren is being heard.
[1629] Step 4:
[1630] If the server detects a siren, it feeds the result back to the device. It then calculates the direction and distance of the emergency vehicle based on the analysis results. The input is the result of the siren detection, and the output is data on the direction and distance of the emergency vehicle.
[1631] Step 5:
[1632] Based on the detection results, the device sends an instruction via Bluetooth to reduce the volume of the sound source inside the car by a certain percentage (e.g., 50%). The input is data about the direction and distance of the emergency vehicle, and the output is a volume adjustment instruction. Specifically, the device sends an instruction to the car's audio system using Bluetooth.
[1633] Step 6:
[1634] The device captures the user's facial image using an onboard camera and analyzes it with an emotion engine. The software used is Python's "opencv" and "keras." The input is the captured facial image, and the output is the user's emotional state.
[1635] Step 7:
[1636] The server obtains the location and direction of the emergency vehicle from the navigation system and generates a specific announcement. The input is the data on the direction and distance of the emergency vehicle and the user's emotional state, and the output is the generated announcement.
[1637] Step 8:
[1638] The device adjusts the content and tone of the announcement based on the analysis results of the emotion engine according to the user's emotional state. For example, if the user is feeling stressed, the announcement will be made in a calmer tone. The input is the generated announcement text, and the output is the adjusted announcement text.
[1639] Step 9:
[1640] The terminal plays the generated voice announcement over the in-car speaker, prompting the user to take specific action (e.g., "There is an ambulance approaching 100 meters ahead from the right, please pull over to the side of the road.") The input is the adjusted announcement text, and the output is a voice notification from the in-car speaker.
[1641] In this way, through each step, it is possible to quickly and accurately detect the approach of an emergency vehicle and notify the user appropriately according to their emotional state.
[1642] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1643] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1644] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1645] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1646] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1647] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1648] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1649] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1650] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1651] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1652] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1653] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1654] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1655] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1656] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1657] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1658] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1659] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1660] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1661] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1662] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1663] The following is further disclosed regarding the above embodiment.
[1664] (Claim 1)
[1665] Sound collection means installed at the four corners of the vehicle;
[1666] an analysis means for analyzing the collected sound data and monitoring surrounding sounds;
[1667] a detection means for detecting the approach of an emergency vehicle based on the analysis result;
[1668] a volume adjusting means for automatically lowering the volume of the sound source being reproduced in the vehicle when the detection means detects the approach of an emergency vehicle;
[1669] A system that includes a notification means for informing in-vehicle users of the location and approach of emergency vehicles.
[1670] (Claim 2)
[1671] 2. The system according to claim 1, wherein the analyzing means calculates the approaching direction and distance of an emergency vehicle based on the sound data collected from the sound collecting means.
[1672] (Claim 3)
[1673] 2. The system according to claim 1, wherein the notification means plays a specific announcement including the approaching direction and estimated distance of the emergency vehicle through an in-vehicle speaker.
[1674] "Example 1"
[1675] (Claim 1)
[1676] sound collecting means installed at the four corners of the vehicle;
[1677] a sound data analysis means for analyzing collected sound data and monitoring surrounding sounds;
[1678] emergency vehicle detection means for detecting the approach of an emergency vehicle based on the analysis result;
[1679] an automatic volume adjustment means for automatically lowering the volume of the sound source being reproduced in the vehicle when the detection means detects the approach of an emergency vehicle;
[1680] a notification means for notifying an in-vehicle user of the location and approach of an emergency vehicle;
[1681] The system includes a notification generation means for obtaining the location of an emergency vehicle based on data from a navigation system and generating an appropriate announcement.
[1682] (Claim 2)
[1683] 2. The system according to claim 1, wherein the sound data analysis means calculates the approaching direction and distance of an emergency vehicle based on the collected sound data.
[1684] (Claim 3)
[1685] 10. The system of claim 1, wherein the notification generating means plays a specific announcement through an in-vehicle speaker, the announcement including the direction and estimated distance of the approaching emergency vehicle.
[1686] "Application Example 1"
[1687] (Claim 1)
[1688] A sound collection means installed in a vehicle or a mobile terminal;
[1689] an analysis means for analyzing the collected sound data and monitoring surrounding sounds;
[1690] a detection means for detecting the approach of an emergency vehicle based on the analysis result;
[1691] a volume adjusting means for automatically lowering the volume of the sound source being reproduced when the detection means detects the approach of an emergency vehicle;
[1692] A system including a notification means for informing users of the location and approach of emergency vehicles.
[1693] (Claim 2)
[1694] 2. The system according to claim 1, wherein the analysis means calculates the approaching direction and distance of an emergency vehicle based on the sound data collected from the sound collection means, and displays the analysis results on the mobile terminal.
[1695] (Claim 3)
[1696] 2. The system according to claim 1, wherein the notification means reproduces a specific announcement including the approaching direction and estimated distance of the emergency vehicle through the audio output means, and also notifies the user by displaying the announcement on the screen in addition to the audio.
[1697] "Example 2: Combining Emotion Engines"
[1698] (Claim 1)
[1699] Sound collection devices installed at the four corners of the vehicle,
[1700] a processing device for analyzing the collected sound data, the processing device being an analysis device that monitors surrounding sounds based on the generated analysis results;
[1701] a detection device for detecting the approach of an emergency vehicle based on the analysis result;
[1702] a volume control device that automatically reduces the volume of the sound source being played in the vehicle via Bluetooth communication when the approach of an emergency vehicle is detected by the detection device;
[1703] a notification device that notifies in-vehicle users of the location and approach of an emergency vehicle;
[1704] A system including an emotion recognizer for collecting and analyzing a user's emotional state.
[1705] (Claim 2)
[1706] 2. The system according to claim 1, wherein the analysis device calculates the approaching direction and distance of an emergency vehicle based on sound data collected from the sound collection device.
[1707] (Claim 3)
[1708] 2. The system of claim 1, wherein the notification device plays a specific announcement including the approaching direction and estimated distance of the emergency vehicle through an in-vehicle speaker in a tone that corresponds to the user's emotional state.
[1709] "Application example 2 when combining emotion engines"
[1710] (Claim 1)
[1711] Sound collection means installed at the four corners of the vehicle;
[1712] an analysis means for analyzing the collected sound data and monitoring surrounding sounds;
[1713] a detection means for detecting the approach of an emergency vehicle based on the analysis result;
[1714] a volume adjusting means for automatically lowering the volume of the sound source being reproduced in the vehicle when the detection means detects the approach of an emergency vehicle;
[1715] a notification means for notifying an in-vehicle user of the location and approach of an emergency vehicle;
[1716] an emotion engine that analyzes the user's emotional state;
[1717] The system includes means for adjusting the content and tone of notifications based on the analysis results of the sentiment engine.
[1718] (Claim 2)
[1719] The analysis means calculates the approaching direction and distance of the emergency vehicle based on the sound data collected from the sound collection means,
[1720] 2. The system according to claim 1, wherein the notification means generates a specific announcement including the direction and estimated distance of the approaching emergency vehicle in a tone that corresponds to the emotional state of the user.
[1721] (Claim 3)
[1722] an emotion engine for analyzing a user's facial image or tone of voice to identify their emotional state;
[1723] 10. The system of claim 1, wherein the notification means adjusts the tone and content of the announcement based on the results of the emotion engine. [Explanation of symbols]
[1724] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. Sound collection means installed at the four corners of the vehicle; an analysis means for analyzing the collected sound data and monitoring surrounding sounds; a detection means for detecting the approach of an emergency vehicle based on the analysis result; a volume adjusting means for automatically lowering the volume of the sound source being reproduced in the vehicle when the detection means detects the approach of an emergency vehicle; A system that includes a notification means for informing in-vehicle users of the location and approach of emergency vehicles.
2. 2. The system according to claim 1, wherein the analyzing means calculates the approaching direction and distance of the emergency vehicle based on the sound data collected from the sound collecting means.
3. 2. The system according to claim 1, wherein the notification means plays a specific announcement including the approaching direction and estimated distance of the emergency vehicle through an in-vehicle speaker.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A