Voice processing device, its operating method, and vehicle control system including the voice processing device

The audio processing device in vehicles addresses the challenge of assessing passenger conditions during emergencies by processing voices to determine sound source locations and injuries, enabling rapid and accurate communication to emergency services.

JP2026513301APending Publication Date: 2026-04-23AMOSENSE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
AMOSENSE CO LTD
Filing Date
2024-03-25
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing vehicle systems lack the ability to accurately assess the condition and location of passengers during emergencies, such as collisions, and to effectively communicate this information to emergency responders.

Method used

An audio processing device installed in vehicles that processes passenger voices to determine sound source locations, assess injuries, and output appropriate conversation attempts based on injury estimation and passenger identity, transmitting this information to emergency services.

Benefits of technology

Enables accurate assessment of passenger conditions and rapid communication of this information to emergency responders, facilitating quicker and more targeted rescue efforts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026513301000001_ABST
    Figure 2026513301000001_ABST
Patent Text Reader

Abstract

The present invention provides an audio processing device that can output the content of a conversation with an occupant in the event of an emergency situation such as a vehicle collision, a method for operating the device, and a vehicle control system including such an audio processing device. [Solution] The disclosed audio processing device is an audio processing device installed in a vehicle and includes a microphone, a speaker, configured to generate audio signals related to the voice of an occupant in response to the voice of an occupant in the vehicle, and a processor that, upon receiving a collision signal from a vehicle controller configured to control the operation of the vehicle, outputs a conversation attempt message to the occupant through the speaker, processes the occupant's response audio signal to the conversation attempt message received from the microphone to generate sound source location information indicating the sound source location of the response audio signal, groups the response audio signals according to the sound source location information, and outputs the grouped response audio signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an audio processing device, an operating method thereof, and a vehicle control system including the audio processing device. More specifically, the present invention relates to an audio processing device capable of outputting information corresponding to the occurrence of an emergency situation for a vehicle, an operating method thereof, and a vehicle control system including such an audio processing device.

Background Art

[0002] An e-call service system is a service system of a vehicle ICT infrastructure that automatically transmits accident location, accident information, etc. to an emergency rescue agency when a vehicle collision and a serious accident occur, requests emergency rescue, and enables rapid life-saving. That is, the e-call service system transmits specific information such as the accident occurrence location, vehicle type, running direction, and the number of seat belts operated at the time of the accident to the nearest emergency call response center (PSAP: Public-Safety Answering Point) in cooperation with a satellite position transmission system and a communication company. The emergency call response center grasps the transmitted content and transmits accident-related information to the nearest rescue agency at the accident site. Through such a rapid communication system, the accident is quickly dealt with and life-saving is carried out. The matters described in the above background art are for helping to understand the background of the invention and may include matters that are not publicly known prior arts.

Summary of the Invention

Problems to be Solved by the Invention

[0003] The present invention has been proposed in consideration of the above-described conventional circumstances, and an object thereof is to provide an audio processing device capable of outputting the content of a conversation with a passenger when an emergency situation such as a vehicle collision occurs, an operating method thereof, and a vehicle control system including such an audio processing device. ]

Means for Solving the Problems

[0004] To achieve the above objective, a preferred embodiment of the present invention is an audio processing device installed in a vehicle, comprising: a microphone configured to generate an audio signal related to the voice of a passenger in response to the voice of a passenger in the vehicle; a speaker configured to output audio to the passenger in the vehicle; a memory; and a processor configured to load commands stored in the memory and perform one or more operations by executing the commands, wherein the processor, upon receiving a collision signal from a vehicle controller configured to control the operation of the vehicle, outputs a conversation attempt message to the passenger through the speaker; processes the passenger's response audio signal to the conversation attempt message received from the microphone to generate sound source location information indicating the sound source location of the response audio signal; groups the response audio signals according to the sound source location information; and outputs the grouped response audio signals.

[0005] The memory stores a plurality of candidate conversation attempt messages, and the processor can output a first candidate conversation attempt message from among the plurality of candidate conversation attempt messages, and based on the response regarding the passenger's response voice signal to the first candidate conversation attempt message, it can output a second candidate conversation attempt message that follows the first candidate conversation attempt message.

[0006] The memory stores a plurality of candidate conversation attempt messages, the processor transmits the sound source location information to the vehicle controller, the vehicle controller receives injury estimation information indicating the degree and location of injury of the occupant at the sound source location estimated based on the sound source location information and sensor sensing information, and can output a candidate conversation attempt message corresponding to the injury estimation information from among the plurality of candidate conversation attempt messages.

[0007] The processor analyzes the response audio signal at the sound source location to determine the pitch and timbre of the response audio signal, determines whether the passenger at the sound source location is a child passenger based on the determined pitch and timbre, and if it is determined that the passenger is a child passenger, it can output identity information indicating that the passenger at the sound source location is a child passenger.

[0008] The memory stores multiple candidate conversation attempt messages, and the processor can output a candidate conversation attempt message corresponding to the child passenger from among the multiple candidate conversation attempt messages.

[0009] The processor can output the grouped response voice signals to an emergency rescue request device installed in the vehicle.

[0010] The processor can output the grouped response audio signals to a control server and output sound source location information corresponding to the grouped response audio signals to an emergency rescue request device installed in the vehicle.

[0011] The processor can receive vehicle location information from the vehicle controller and output the vehicle location information at the time the collision occurrence signal was received, along with the grouped response audio signals.

[0012] On the other hand, a preferred embodiment of the present invention is a method for operating an audio processing device provided in a vehicle, characterized in that, upon receiving a collision signal from a vehicle controller configured to control the operation of the vehicle, the method includes the steps of: outputting a conversation attempt message to the occupants of the vehicle; processing the occupants' response audio signals to the conversation attempt message to generate sound source location information indicating the sound source location of the response audio signals; grouping the response audio signals according to the sound source location information; and outputting the grouped response audio signals.

[0013] A preferred embodiment of the present invention is a vehicle control system comprising a vehicle controller configured to control the operation of a vehicle, and an audio processing device configured to receive a collision occurrence signal from the vehicle controller, output a conversation attempt message to the occupants of the vehicle, process the occupants' response audio signals to the conversation attempt message to generate sound source location information indicating the sound source location of the response audio signals, group the response audio signals according to the sound source location information, and output the grouped response audio signals. [Effects of the Invention]

[0014] According to the present invention with this configuration, when an accident (e.g., a collision) occurs in a vehicle, the content of the conversation with the occupants inside the vehicle, sound source location information, and vehicle location information can be transmitted to the control server, so that the control server can more accurately grasp the details of the accident (e.g., the extent of injuries to the occupants, the location of the injuries, etc.). This allows the control server to direct quicker and more accurate responses. Even in cases of emergencies where passengers cannot report them directly, reports can be accepted, allowing passengers in critical situations to receive emergency treatment quickly and reducing the time required for rescuers to respond. Furthermore, according to the present invention, rescuers who are not present at the accident scene can more accurately understand the situation inside the vehicle at the time of the accident, can appropriately formulate a rescue plan for accident response, and can proceed quickly with accident response and emergency patient transport. [Brief explanation of the drawing]

[0015] [Figure 1] This diagram shows the configuration of a system including a vehicle employing an audio processing device according to an embodiment of the present invention and a control server that networks with the vehicle. [Figure 2] This is a diagram illustrating the operation of an audio processing device according to an embodiment of the present invention. [Figure 3] This diagram illustrates the data transmitted from a vehicle to a control server in the event of an accident. [Figure 4]Figure 1 is an internal configuration diagram of the audio processing device. [Figure 5] This is a flowchart illustrating the operation of the audio processing device according to an embodiment of the present invention. [Modes for carrying out the invention]

[0016] The present invention is subject to various modifications and can be implemented in many ways; therefore, specific embodiments are illustrated and described in detail in the drawings. However, this should not be understood as an attempt to limit the present invention to any particular embodiment, but rather as including all modifications, equivalents, or substitutes that fall within the spirit and technical scope of the present invention. The terms used in this application are used solely to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this application, terms such as “includes” or “has” should be understood as intending to specify the existence of features, figures, steps, actions, components, parts, or combinations thereof described in the specifications, and not as preemptively excluding the existence or possibility of adding one or more other features, figures, steps, actions, components, parts, or combinations thereof. Unless otherwise defined, all terms used herein, including technical and scientific terms, have the same meaning as those generally understood by a person of ordinary skill in the art to which this invention pertains. Terms as defined in commonly used dictionaries should be interpreted as having the meaning consistent with their meaning in the context of the relevant art, and not as ideal or overly formal unless expressly defined herein.

[0017] Preferred embodiments of the present invention will be described in more detail below with reference to the attached drawings. In describing the present invention, the same reference numerals are used for identical components in the drawings to facilitate overall understanding, and redundant descriptions of identical components are omitted.

[0018] FIG. 1 is a configuration diagram of a system including a vehicle in which a voice processing device according to an embodiment of the present invention is employed and a control server that networks with the vehicle, FIG. 2 is a diagram for explaining the operation of the voice processing device according to an embodiment of the present invention, and FIG. 3 is a diagram illustrating data transmitted from the vehicle to the control server at the time of an accident.

[0019] As shown in FIG. 1, the vehicle 10 can be defined as a transportation or conveyance means that travels on roads, sea routes, railway lines, and air routes, such as automobiles, trains, motorcycles, ships, and airplanes. Depending on the embodiment, the vehicle 10 may be a concept that includes any of an internal combustion engine vehicle equipped with an engine as a power source, a hybrid vehicle equipped with an engine and an electric motor as power sources, and an electric vehicle equipped with an electric motor as a power source. The vehicle 10 may be provided with a sensor 20, a vehicle controller 30, a voice processing device 40, an emergency rescue request device 50, and the like. Although not shown in the drawings, in addition to these, the vehicle 10 is provided with a steering device, a driving device, and the like.

[0020] The sensor 20 can include a number of sensors and can sense the state of the vehicle 10 to generate vehicle state information. Here, the vehicle state information can include information regarding the current position of the vehicle 10, the presence or absence of a collision, and the like. If necessary, the vehicle state information may further include information regarding the vehicle speed of the vehicle 10, surrounding obstacles, and the like. The sensor 20 can transmit the vehicle state information to the vehicle controller 30. Here, the vehicle state information may be an analog signal or a digital signal depending on the characteristics of the sensor. Thereby, when the received vehicle state information is an analog signal, the vehicle controller 30 can include an AD converter (not shown) that converts the vehicle state information, which is an analog signal, into a digital signal, and when the received vehicle state information is a digital signal, a digital input buffer (not shown) that buffers the vehicle state information, which is a digital signal.

[0021] The sensor 20 can include a GPS sensor that can sense the current position of the vehicle 10 in real time, a collision sensor that can detect a collision between the vehicle 10 and other objects (e.g., other vehicles, obstacles, etc.), an obstacle sensor that can sense obstacles around the vehicle 10, a vehicle speed sensor that can sense the vehicle speed of the vehicle 10, a shock sensor, an acceleration sensor, etc. Of course, if necessary, the sensor 20 can sense the internal / external information of the vehicle 10. Accordingly, incidentally, the sensor 20 can sense the temperature and humidity inside / outside the vehicle 10, the longitudinal / lateral acceleration, etc.

[0022] The vehicle controller 30 may be configured to control the operation of the vehicle 10. The vehicle controller 30 can receive the input of vehicle state information from the sensor 20 and can receive the input of vehicle information through the OBD-II port of the vehicle 10. Here, the vehicle information can mean information regarding the state or abnormality of components mounted or installed in the vehicle 10. When a vehicle collision occurs, the vehicle controller 30 will receive vehicle state information (especially, a collision occurrence signal) from the sensor 20. Thereby, the vehicle controller 30 can transmit the received vehicle state information (especially, the collision occurrence signal) to the audio processing device 40. The vehicle controller 30 may be configured to include, for example, an ECU (Electronic Control Unit).

[0023] According to an embodiment, the vehicle controller 30 can receive the sound source position information from the audio processing device 40 and can receive various sensing information at the time of the vehicle collision from the sensor 20. Here, the sound source position information can mean the sound source position information indicating the sound source position of the passenger's response voice signal to the conversation attempt message. The various sensing information at the time of the vehicle collision can include, for example, the vehicle speed, the longitudinal / lateral acceleration value, the shock value, etc. The vehicle controller 30 can estimate the degree and location of injury of the passenger at the sound source position based on the received sound source position information and the sensing information of the sensor.

[0024] For example, the vehicle controller 30 pre-stores information on the degree and location of injuries to passengers at different sound source locations (i.e., passenger positions) under various collision conditions. Therefore, when the vehicle controller 30 receives sound source location information and sensing information at the time of the vehicle collision, it can estimate the degree and location of injuries to passengers at the sound source location corresponding to the received sound source location information and sensor sensing information. The vehicle controller 30 can then transmit injury estimation information, indicating the estimated degree and location of injuries for the passenger at that sound source location, to the voice processing device 40. This allows the voice processing device 40 to output a candidate conversation attempt message corresponding to the injury estimation information from among several already stored candidate conversation attempt messages. In this way, the voice processing device 40 can obtain responses (answers) regarding the degree and location of injuries to passengers more quickly and accurately.

[0025] On the other hand, the vehicle controller 30 can estimate the degree and location of injuries to occupants at different occupant positions (i.e., sound source locations) based on various sensing information from the sensor 20 at the time of the vehicle collision. The various sensing information from the sensor 20 at the time of the vehicle collision may include, for example, vehicle speed, longitudinal / lateral acceleration values, and impact values. The vehicle controller 30 has pre-stored information on the degree and location of injuries to occupants at different sound source locations (i.e., occupant positions) under various collision conditions. As a result, when the vehicle controller 30 receives sensing information at the time of the vehicle collision, it can estimate the degree and location of injuries to occupants at the sound source location corresponding to the sensing information from the received sensor.

[0026] Thereafter, the vehicle controller 30 receives sound source location information and passenger identification information (information indicating whether the passenger is an adult or a child) from the voice processing device 40. Here, a child can mean a child from one year old to six years old. Based on the received information (i.e., sound source location information and identification information, or sound source location information, identification information and sensing information from sensor 20), the vehicle controller 30 can predict the degree and location of injuries of the passenger for each sound source location (i.e., passenger position). Here, the estimation of the degree and location of injuries of the passenger for each passenger position only considers the case where the passenger is an adult, and the estimation may be somewhat inaccurate when the passenger is a child. In contrast, the prediction of the degree and location of injuries of the passenger for each passenger position considers both adult and child passengers and predicts the degree and location of injuries of the passenger depending on whether the passenger is an adult or a child. Even in vehicle collision accidents under the same conditions, the degree and location of injuries may differ between adults and children, so the aforementioned prediction operation may be necessary.

[0027] The vehicle controller 30 can transmit injury prediction information to the voice processing device 40, indicating the predicted degree and location of injury for each passenger at the sound source location (i.e., boarding position). As a result, the voice processing device 40 can receive the injury prediction information and output a conversation attempt message corresponding to the injury prediction information. In particular, since the voice processing device 40 can output a conversation attempt message corresponding to the injury prediction information for child passengers, responses (answers) regarding the degree and location of injuries can be obtained more quickly and accurately not only from adult passengers but also from child passengers.

[0028] The voice processing device 40 can generate an audio signal related to the voice of each passenger in the vehicle 10 in response to the voice of each passenger. Here, the audio signal is a signal related to the voice spoken for a specific time, and may be a signal indicating the voice of each of multiple passengers. The voice processing device 40 can separate and recognize the voices of each passenger individually. When multiple passengers speak simultaneously, the audio contains the voices of all the passengers who spoke. In order to accurately process the voice of each passenger, it is necessary to separate the voice of each passenger from the audio containing the voices of all the passengers.

[0029] The voice processing device 40 according to an embodiment of the present invention can determine the sound source location of each passenger's voice from the voice signal related to the voices of multiple passengers, and extract (or generate) a separated voice signal related to each passenger's voice from the voice signal by performing sound source separation based on the sound source location. In other words, the voice processing device 40 can generate a separated voice signal related to the voice at each sound source location based on the sound source location of the voice (i.e., which can be the passenger's seating position). In the embodiment, the voice processing device 40 can classify the components of the voice signal by sound source location and generate a separated voice signal related to the voice spoken at each sound source location using the classified components corresponding to each sound source location. For example, the voice processing device 40 can generate a first separated voice signal related to the voice of the first passenger spoken at the first sound source location based on the voice signal. In this case, the first separated voice signal may have the highest degree of association with the voice of the first passenger among the passengers' voices. In other words, the proportion of the voice component of the first passenger may be the highest among the voice components included in the first separated voice signal.

[0030] According to the embodiment, the voice processing device 40 can determine the sound source location of each passenger's voice using the time delay (or phase delay) between voice signals related to the voices of passengers in the vehicle 10, and generate a separated voice signal corresponding only to the sound source at a specific location. For example, the voice processing device 40 can generate a separated voice signal related to a voice spoken at a specific location (or direction). This allows the voice processing device 40 to generate a separated voice signal related to each passenger's voice. Through the operation up to the generation of such separated voice signals, the voice processing device 40 can sufficiently generate the passenger's response voice signal at a specific sound source location to a conversational attempt message, and sound source location information indicating the sound source location of the response voice signal.

[0031] For example, as shown in Figure 2, let's assume that four passengers SPK1 to SPK4 are in the vehicle 10 and are able to pronounce speech. The first passenger SPK1 may be located in the front row left area FL of the vehicle 10, the second passenger SPK2 may be located in the front row right area FR of the vehicle 10, the third passenger SPK3 may be located in the rear row left area BL of the vehicle 10, and the fourth passenger SPK4 may be located in the rear row right area BR of the vehicle 10, but the embodiments of the present invention are not limited to these. Speakers S1 to S4 can output speech corresponding to the speech signal.

[0032] When the first to fourth passengers SPK1 to SPK4 utter predetermined sounds within the vehicle 10, the voice processing device 40 can generate voice signals related to the sounds of each of the first to fourth passengers SPK1 to SPK4 in response to their respective voices. The voice signals are signals related to the sounds uttered for a specific time and may be signals indicating the voices of multiple passengers.

[0033] Thereafter, the voice processing device 40 uses the time delay (or phase delay) between the voice signals related to the voices of the first to fourth passengers SPK1 to SPK4 to determine the sound source position of each of the voices of the first to fourth passengers SPK1 to SPK4, and extracts (or generates) a separated voice signal corresponding only to the sound source at a specific position. As illustrated in Figure 2, it is assumed that the first passenger SPK1 utters the voice "AAA", the second passenger SPK2 utters the voice "BBB", the third passenger SPK3 utters the voice "CCC", and the fourth passenger SPK4 utters the voice "DDD". The voice processing device 40 generates voice signals in response to the voices "AAA", "BBB", "CCC", and "DDD", and can use the generated voice signals to generate separated voice signals related to the voices of each passenger SPK1 to SPK4.

[0034] For example, the audio processing device 40 can separate the audio signal according to the location of the sound source, and then match and store the first separated audio signal associated with the voice "AAA" of the first passenger SPK1 with the first sound source location information, which indicates the sound source location of the voice "AAA" (i.e., the boarding position of the first passenger SPK1), namely the front row left FL. Similarly, the audio processing device 40 can match and store the second separated audio signal associated with the voice "BBB" of the second passenger SPK2 with the second sound source location information, which indicates the sound source location of the voice "BBB" (i.e., the boarding position of the second passenger SPK2), namely the front row right FR. Furthermore, the audio processing device 40 can match and store the third separated audio signal associated with the voice "CCC" of the third passenger SPK3 with the third sound source location information, which indicates the sound source location of the voice "CCC" (i.e., the boarding position of the third passenger SPK3), namely the rear row left BL. Furthermore, the audio processing device 40 can match and store the fourth separated audio signal associated with the voice "DDD" of the fourth passenger SPK4 and the fourth sound source position information indicating the rear right side BR, which is the sound source position of the voice "DDD" (i.e., the boarding position of the fourth passenger SPK4).

[0035] According to the embodiment, when the voice processing device 40 receives a collision signal from the vehicle controller 30, it can output a predetermined conversation attempt message to the occupant. There are various detailed methods for outputting the conversation attempt message. One example is as follows: The voice processing device 40 may store multiple candidate conversation attempt messages in advance. The multiple candidate conversation attempt messages may be linked in a hierarchical question structure. As a result, when the voice processing device 40 receives a collision signal from the vehicle controller 30, it can output a first candidate conversation attempt message from among the multiple candidate conversation attempt messages in order to converse with the occupant in the vehicle 10.

[0036] The first candidate conversation attempt message may be a message pre-specified at the beginning, or it may be a message selected randomly. Thereafter, the voice processing device 40 can output a second candidate conversation attempt message that follows the first candidate conversation attempt message based on the passenger's response to the first candidate conversation attempt message. Here, the response to the response to the voice signal may include whether or not the passenger replies, the time taken to reply, the volume (strength) of the reply, and the content of the reply (e.g., an affirmative or indefinite reply). In other words, the second candidate conversation attempt message that follows the first candidate conversation attempt message is not unconditionally determined, but rather the second candidate conversation attempt message is selected and output based on the passenger's response (e.g., reply) to the first candidate conversation attempt message (e.g., a question).

[0037] Among the detailed methods for outputting the aforementioned attempt at conversation messages, one method different from the example illustrated above is one that takes into account the injuries of the occupants. That is, when a vehicle collision occurs, the occupants in the vehicle 10 may be injured. Therefore, the voice processing device 40 can output the attempt at conversation messages while taking into account the extent and location of the occupants' injuries. For example, the voice processing device 40 may store multiple candidate attempt at conversation messages in advance. Multiple candidate attempt at conversation messages may be linked in a hierarchical question structure. The voice processing device 40 transmits sound source location information to the vehicle controller 30. Here, the sound source location information may be generated by processing the occupants' response voice signals to the attempt at conversation messages for a while after the vehicle collision occurred. Of course, if necessary, the sound source location information may be generated in advance by determining the sound source location of each occupant's voice from the voice signals related to the occupants' voices received from the microphone 41 before the vehicle collision.

[0038] The vehicle controller 30 estimates the degree and location of injuries to the occupant at the sound source location based on the sound source location information and the sensing information from the sensor 20. The vehicle controller 30 generates injury estimation information indicating the estimated degree and location of injuries to the occupant at the sound source location and transmits it to the voice processing device 40. Thereafter, the voice processing device 40 receives the injury estimation information and can output a candidate conversation attempt message corresponding to the injury estimation information from among a plurality of candidate conversation attempt messages. The sensing information from the sensor 20 can include, for example, vehicle speed, longitudinal / lateral acceleration values, impact values, etc., and can be useful in estimating the degree and location of injuries to the occupant at the sound source location.

[0039] On the other hand, the vehicle controller 30 can receive sound source location information and passenger identity information (information indicating whether the person is an adult or a child; identity information will be described later) from the voice processing device 40, and can receive various sensing information from the sensor 20 at the time of the vehicle collision. Based on the received information, the vehicle controller 30 can predict the degree and location of injuries of passengers at each sound source location (i.e., passenger position). For example, the vehicle controller 30 may have a table in which information such as the degree and location of injuries of passengers is matched with the content of the sound source location information, identity information, and sensing information from the sensors.

[0040] The vehicle controller 30 can easily predict the degree and location of injuries of passengers based on the sound source location by utilizing such a table. In this case, the vehicle controller 30 uses passenger identity information in addition to the estimation method described above, so it can predict the degree and location of injuries of passengers based on the sound source location more accurately. The vehicle controller 30 can transmit injury prediction information, which indicates the predicted degree and location of injuries for passengers based on the sound source location (i.e., passenger position), to the voice processing device 40. As a result, the voice processing device 40 receives the injury prediction information and can output a candidate conversation attempt message corresponding to the injury prediction information from among multiple candidate conversation attempt messages, based on the sound source location. In particular, in the case of child passengers, the vehicle controller 30 can generate appropriate injury prediction information and transmit it to the voice processing device 40, so that the voice processing device 40 can output a conversation attempt message corresponding to the injury prediction information for child passengers. This allows for appropriate questioning (inquiries) that take into account the degree and location of injuries, even for child passengers.

[0041] According to the embodiment, the voice processing device 40 can receive the passenger's response voice signal to a conversation attempt message, process the received response voice signal, and generate sound source location information indicating the sound source location of the response voice signal. Since the voice processing device 40 can classify the components of the passenger's voice signal according to the sound source location, it is considered that it can sufficiently generate sound source location information indicating the sound source location. The sound source location information may be used to output the aforementioned conversation attempt message as voice, or it may be used when grouping the passenger's response voice signals.

[0042] According to the embodiment, the voice processing device 40 can group the response voice signals of each passenger based on sound source location information. For example, the voice processing device 40 can generate a first group of response voice signals consisting only of the response voice signals of the first passenger SPK1 at the first sound source location, but the response voice signals of the first group may be arranged in chronological order.

[0043] According to the embodiment, the voice processing device 40 can transmit grouped response voice signals and sound source location information corresponding to the grouped response voice signals to the emergency rescue request device 50. The emergency rescue request device 50 can then transmit the grouped response voice signals and sound source location information corresponding to the grouped response voice signals to the control server 70. Of course, if necessary, the voice processing device 40 can also transmit the grouped response voice signals directly to the control server 70, and the sound source location information corresponding to the grouped response voice signals can be transmitted to the control server 70 via the emergency rescue request device 50.

[0044] The grouped response audio signals mentioned above may be a collection of replies to predetermined questions, but the first group of response audio signals given as an example above may consist only of the response audio signals of the first passenger SPK1 at the first sound source location. When such a first group of response audio signals is transmitted to the control server 70, the control server 70 will be able to more accurately grasp the situation of the passenger at the sound source location in the vehicle (i.e., the first passenger SPK1) (severity of injury, location of injury, etc.) through the first group of response audio signals.

[0045] On the other hand, a predetermined time is required to group the passengers' response audio signals. Therefore, when the audio processing device 40 receives a collision signal, it can group the passengers' response audio signals according to their sound source location within a pre-set time. In this embodiment, the audio processing device 40 can store the grouped response audio signals in its internal memory and then output the grouped response audio signals stored in its internal memory.

[0046] According to the embodiment, the voice processing device 40 analyzes the response voice signal for each sound source location to determine the pitch and timbre of the response voice signal, and can determine whether the passenger at the sound source location is a child passenger based on the determined pitch and timbre. If the voice processing device 40 determines that the passenger is a child passenger, it can output identity information indicating that the passenger at the sound source location is a child passenger. Of course, if the voice processing device 40 determines that the passenger at the sound source location is an adult passenger, it can output identity information indicating that the passenger is an adult passenger. For example, when comparing the case of an adult passenger and a child passenger, the pitch of the response voice signal of an adult passenger (i.e., the pitch to the utterance position) is higher than the pitch of the response voice signal of a child passenger (i.e., the pitch to the utterance position). Also, the timbre of an adult and the timbre of a child should be different. In this way, the voice processing device 40 can determine whether the passenger at the sound source location is a child passenger through the pitch and timbre of the response voice signal. In other words, since the pitch of the response voice signal alone can be somewhat inaccurate in determining whether the passenger is an adult or a child, the tone of the response voice signal was used as another criterion for judgment.

[0047] Here, identity information indicating that the passenger is a child may be transmitted to the control server 70 along with the grouped response voice signals. For example, the voice processing device 40 can transmit the grouped response voice signals, the sound source location information corresponding to the grouped response voice signals, and the identity information indicating that the passenger is a child to the control server 70 via the emergency rescue request device 50. Of course, if necessary, the voice processing device 40 can also transmit the grouped response voice signals directly to the control server 70, and the sound source location information corresponding to the grouped response voice signals and the identity information indicating that the passenger is a child may be transmitted to the control server 70 via the emergency rescue request device 50. If necessary, the voice processing device 40 may include a communication unit (not shown) configured to send and receive data (i.e., grouped response voice signals) with the control server 70 using radio waves of various frequencies in order to transmit the grouped response voice signals directly to the control server 70. Here, the communication unit (not shown) in the voice processing device 40 can send and receive data with the control server 70 by at least one wireless communication method from short-range wireless communication, medium-range wireless communication, and long-range wireless communication.

[0048] By transmitting identification information indicating that a passenger is a child to the control server 70 in this way, the control server 70 can determine through the identification information that there are children as well as adults among the passengers. This allows the control server 70 to issue more accurate response orders for rescuing children as well as adults.

[0049] On the other hand, when a vehicle collision occurs, the type and location of injuries may differ between adult and child passengers, so the attempt at conversation should also differ. Therefore, the voice processing device 40 can output a candidate attempt at conversation message corresponding to the child passenger from among several candidate attempt at conversation messages stored in its internal memory. In other words, in order to more accurately inquire about the type and location of injuries sustained by the child passenger, it is preferable to output an attempt at conversation message appropriate for the child passenger (i.e., a candidate attempt at conversation message) in the case of a child passenger. Furthermore, the voice processing device 40 can receive vehicle location information from the vehicle controller 30 in real time. As a result, when the voice processing device 40 outputs grouped response voice signals and sound source location information, it can output the vehicle location information at the time the collision occurrence signal was received.

[0050] On the other hand, the voice processing device 40 does not receive vehicle position information from the vehicle controller 30 in real time, but rather receives the vehicle position information at the time of the collision from the vehicle controller 30 when a vehicle collision occurs. For example, when the vehicle controller 30 receives a collision occurrence signal from the sensor 20, it can also receive the vehicle position information at that time and provide the voice processing device 40 with the collision occurrence signal and the vehicle position information at the time of the collision.

[0051] In Figure 1, the emergency rescue request device 50 may be configured to receive grouped response voice signals and send the received grouped response voice signals to the control server 70. Of course, the emergency rescue request device 50 can also transmit the current location information of the vehicle 10 and the sound source location information of the occupants to the control server 70 along with the grouped response voice signals. The emergency rescue request device 50 may include a communication unit (not shown) configured to send and receive data with the control server 70 using radio waves of various frequencies. Here, the communication unit in the emergency rescue request device 50 can send and receive data with the control server 70 using at least one wireless communication method from short-range wireless communication, medium-range wireless communication, and long-range wireless communication.

[0052] Therefore, as illustrated in Figure 3, when a vehicle 10 equipped with an audio processing device 40 and an emergency rescue request device 50 occurs, it can transmit grouped response audio signals, the current location information of the vehicle 10, and the sound source location information of the occupants to the control server 70. In Figure 1, the network 60 can represent a connecting structure that enables information exchange between the vehicle 10 and the control server 70. The network 60 can have any structure as long as it enables information exchange between the vehicle 10 and the control server 70.

[0053] In Figure 1, the control server 70 can receive grouped response audio signals, the current location information of the vehicle 10, and the sound source location information of the occupants from the vehicle 10 in the event of an accident such as a collision. Of course, the control server 70 can also receive the grouped response audio signals directly from the voice processing device 40 if necessary. Based on the received grouped response audio signals, the current location information of the vehicle 10, and the sound source location information of the occupants, the control server 70 can more accurately grasp the number of occupants inside the vehicle and the current state of the occupants at the time of the vehicle collision. This allows the control server 70 to give instructions for a quicker and more accurate response to the vehicle collision. In Figure 1, the sensor 20, vehicle controller 30, voice processing device 40, and emergency rescue request device 50 can be collectively referred to as the vehicle control system.

[0054] Figure 4 is an internal configuration diagram of the audio processing device 40 shown in Figure 1. The audio processing device 40 may include a microphone 41, a memory 42, a communication unit 43, a processor 44, and a speaker 45. The microphone 41 can generate an audio signal in response to generated sound. In some embodiments, the microphone 41 can detect air vibrations caused by sound and generate an audio signal, which is an electrical signal corresponding to the vibrations, based on the detection result. For example, the microphone 41 can receive the voice of an occupant located inside the vehicle 10 and convert the occupant's voice into an audio signal, which is an electrical signal. For example, the microphone 41 may include a plurality of microphones arranged to form an array, and each of the plurality of microphones can generate an audio signal in response to sound. In this case, since the positions in which each of the plurality of microphones is located may differ from one another, the audio signals generated from each of the plurality of microphones may have a phase difference (or time delay) from one another.

[0055] For example, the microphone 41 may be located in the center fascia of the vehicle 10. Here, the center fascia can refer to the control panel area in the dashboard between the driver's seat and the passenger seat. On the other hand, although this specification has described the voice processing device 40 as including the microphone 41 and directly generating voice signals related to the occupant's voice using the microphone 41, in some embodiments the microphone may be configured separately from the voice processing device 40. That is, the voice processing device 40 can also receive voice signals from a separately configured microphone and process or utilize the received voice signals. For example, the voice processing device 40 can also generate a separated voice signal from the voice signal received from a separated microphone.

[0056] For the sake of explanation, unless otherwise specified, it will be assumed that the audio processing device 40 includes a microphone 41. The memory 42 can store data necessary for the operation of the audio processing device 40. For example, the memory 42 may include at least one of non-volatile memory and volatile memory. In some embodiments, the memory 42 can store identifiers corresponding to each sound source location in the vehicle 10 (i.e., the passenger's seating position). The identifiers may be data for distinguishing sound source locations. Since each sound source location corresponds to each passenger, each passenger can be distinguished using the identifier corresponding to the sound source location. For example, the first identifier indicating the first sound source location may indicate the first passenger. In this view, the identifiers corresponding to each sound source location in the vehicle 10 can also function as passenger identifiers for identifying each passenger.

[0057] If necessary, the identifier may be input through an input device (e.g., a touchpad, not shown) of the voice processing device 40. In some embodiments, the memory 42 can store grouped response voice signals and sound source location information relating to the sound source locations of each occupant. When the grouped response voice signals and occupant sound source location information are stored in the memory 42, the position information of the vehicle 10 at the time the collision signal was received may also be stored. The communication unit 43 is configured to send and receive data with the vehicle controller 30 and / or the emergency rescue request device 50.

[0058] The processor 44 can control the overall operation of the audio processing device 40. Depending on the embodiment, the processor 44 may include a processor having arithmetic processing capabilities. For example, the processor 44 may include, but is not limited to, a CPU (central processing unit), an MCU (microcontroller unit), a GPU (graphics processing unit), a DSP (digital signal processor), an ADC converter (analog to digital converter), or a DAC converter (digital to analog converter).

[0059] Unless otherwise noted, the operation of the audio processing device 40 described herein can be understood as the operation of the processor 44. The processor 44 can process the audio signal generated by the microphone 41. For example, the processor 44 can convert an analog audio signal generated by the microphone 41 into a digital audio signal and process the converted digital audio signal. In this case, since the type of signal (analog or digital) changes, digital audio signals and analog audio signals will be used interchangeably in the description of embodiments of the present invention.

[0060] In one embodiment, the processor 44 can process the audio signal generated by the microphone 41 and extract (or generate) an audio signal associated with the voice of each passenger. In another embodiment, the processor 44 can generate separated audio signals associated with the voices of passengers located at each sound source location. The separated audio signals may be in the form of audio data or text data.

[0061] The processor 44 can determine the sound source location of a passenger's voice using the time delay (or phase delay) between the separated voice signals. For example, the processor 44 can determine the relative location of the sound source (i.e., the relative location of the passenger). In other words, the processor 44 can classify the components of the voice signal by sound source location and generate separated voice signals related to the voice spoken at each sound source location using the classified components corresponding to each sound source location. For example, the processor 44 can generate a first separated voice signal related to the voice of a first passenger based on the sound source location of the voice.

[0062] In one embodiment, the processor 44 can store sound source location information indicating the determined sound source location by matching it with the separated audio signal. For example, the processor 44 can store in memory 42 a match between a first separated audio signal related to the voice of the first passenger and first sound source location information indicating the sound source location of the voice of the first passenger by matching them. That is, since the location of the sound source corresponds to the respective boarding positions of the passengers, the sound source location information can function as passenger boarding position information for identifying the respective boarding positions of the passengers.

[0063] In this embodiment, when the processor 44 receives a collision signal from the vehicle controller 30, it can output a predetermined conversation attempt message to the occupant. For example, when the processor 44 receives a collision signal from the vehicle controller 30, it can output a first candidate conversation attempt message from among several candidate conversation attempt messages in order to converse with the occupant in the vehicle 10. Here, it is assumed that the multiple candidate conversation attempt messages are linked in a hierarchical question structure and are pre-stored in the memory 42. Subsequently, the processor 44 can output a second candidate conversation attempt message that follows the first candidate conversation attempt message, based on the occupant's response to the first candidate conversation attempt message (for example, whether the occupant responded, the time taken to respond, the volume (strength) of the response, whether the response was affirmative or indefinite, etc.).

[0064] In another example, the processor 44 can process the passenger's response audio signal to a conversation attempt message received from the microphone 41 to generate sound source location information indicating the location of the sound source of the response audio signal, and transmit the generated sound source location information to the vehicle controller 30. Here, the sound source location information may be generated by processing the passenger's response audio signal to conversation attempt messages for a period of time after the vehicle collision occurred.

[0065] On the other hand, if necessary, pre-generated sound source location information may be used by determining the location of each sound source of the occupant's voice from the audio signal related to the occupant's voice received from the microphone 41 before the vehicle collision. The vehicle controller 30 estimates the degree and location of the occupant's injury at the sound source location based on the sound source location information and the sensing information from the sensor 20. The vehicle controller 30 generates injury estimation information indicating the estimated degree and location of the occupant's injury at the sound source location and transmits it to the processor 44. Thereafter, the processor 44 receives the injury estimation information and can output a candidate conversation attempt message corresponding to the injury estimation information from among a plurality of candidate conversation attempt messages. The sensing information from the sensor 20 can include, for example, vehicle speed, longitudinal / lateral acceleration values, impact values, etc., and can be useful in estimating the degree and location of the occupant's injury at the sound source location. Multiple candidate conversation attempt messages are pre-stored in the memory 42.

[0066] Another example is that the processor 44 can receive injury prediction information for each sound source location (i.e., passenger position) from the vehicle controller 30, based on the passenger's sound source location information, passenger's identity information, and sensing information from the sensor 20 at the time of the vehicle collision. As a result, the processor 44 can output a candidate conversation attempt message corresponding to the injury prediction information from among multiple candidate conversation attempt messages, categorized by sound source location.

[0067] The processor 44 can group the response audio signals of each occupant based on sound source location information generated after receiving the collision signal. For example, the processor 44 can generate a first group of response audio signals consisting only of the response audio signals of the first occupant SPK1 at the first sound source location, but the response audio signals of the first group may be arranged in chronological order.

[0068] On the other hand, a predetermined time is required to group the passengers' response voice signals. Therefore, when the processor 44 receives a collision signal, it groups the passengers' response voice signals according to their sound source location within a pre-set time. In this embodiment, the processor 44 can store the grouped response voice signals in the memory 42 and then output the grouped response voice signals stored in the memory 42. Subsequently, the processor 44 can transmit the grouped response voice signals and the sound source location information corresponding to the grouped response voice signals to the emergency rescue request device 50. Of course, if necessary, the processor 44 can also directly transmit the grouped response voice signals to the control server 70, and transmit the sound source location information corresponding to the grouped response voice signals to the control server 70 via the emergency rescue request device 50. Since the sound source location of a passenger corresponds to the passenger's boarding position, the sound source location information can be said to be the passenger's boarding position information.

[0069] In this embodiment, the processor 44 analyzes the response audio signal for each sound source location to determine the pitch and timbre of the response audio signal, and can determine whether the passenger at the sound source location is a child passenger based on the determined pitch and timbre. If the processor 44 determines that the passenger is a child passenger, it can output identity information indicating that the passenger at the sound source location is a child passenger. Of course, if the processor 44 determines that the passenger at the sound source location is an adult passenger, it can output identity information indicating that the passenger is an adult passenger.

[0070] The processor 44 can transmit identity information indicating that the passenger is a child, along with grouped response voice signals, to the control server 70. For example, the processor 44 can transmit the grouped response voice signals, the sound source location information corresponding to the grouped response voice signals, and the identity information indicating that the passenger is a child to the control server 70 via the emergency rescue request device 50. Of course, if necessary, the processor 44 can also transmit the grouped response voice signals directly to the control server 70, and transmit the sound source location information corresponding to the grouped response voice signals and the identity information indicating that the passenger is a child to the control server 70 via the emergency rescue request device 50.

[0071] When identification information indicating that a passenger is a child is transmitted to the control server 70, the control server 70 will know that there is a child among the passengers, and will be able to issue more precise response orders for child rescue as well as adult rescue.

[0072] When a vehicle collision occurs, the type and location of injuries may differ between adult and child passengers, and therefore the conversational attempt messages should also differ. For this reason, the processor 44 can output a candidate conversational attempt message corresponding to the child passenger from among several candidate conversational attempt messages stored in its internal memory. In other words, in order to more accurately inquire about the type and location of injuries sustained by the child passenger, it is preferable to output a conversational attempt message appropriate for the child passenger (i.e., a candidate conversational attempt message) in the case of a child passenger.

[0073] Furthermore, the processor 44 can receive vehicle position information from the vehicle controller 30 at the time of the vehicle collision (i.e., at the time the collision signal was received). As a result, when the processor 44 outputs grouped response audio signals and sound source position information, it can output the vehicle position information at the time the collision signal was received. The operation of the processor 44 or audio processing device 40 described herein can be implemented in the form of a program executable by a computer. For example, the processor 44 can execute an application stored in memory 42 and perform operations corresponding to commands that instruct specific operations through the execution of the application.

[0074] The speaker 45 can vibrate under the control of the processor 44, and can generate sound through vibration. In some embodiments, the speaker 45 can reproduce sound related to an audio signal by forming vibrations corresponding to the audio signal. In some embodiments, the speaker 45 may include a number of speakers (S1 to S4, see Figure 2). Speakers S1 to S4 can generate vibrations based on an audio signal, and can reproduce sound through the vibrations of speakers S1 to S4. Speakers S1 to S4 may be placed at the respective positions of passengers SPK1 to SPK4. For example, each of speakers S1 to S4 may be a speaker placed on the headrest of the seat where passengers SPK1 to SPK4 are located, but embodiments of the present invention are not limited thereto.

[0075] On the other hand, although this specification has described the voice processing device 40 as including a speaker 45 and using the speaker 45 to output a conversation attempt message, depending on the embodiment, the speaker 45 may be configured separately from the voice processing device 40 and externally. That is, the voice processing device 40 can output a conversation attempt message through the speaker 45 included inside, or it can output a conversation attempt message using externally configured speakers S1 to S4, as illustrated in Figure 2.

[0076] Figure 5 is a flowchart illustrating the operation of an audio processing device according to an embodiment of the present invention. First, the vehicle 10 is in motion (S10). When the audio processing device 40 receives a collision signal from the vehicle controller 30 while the vehicle is in motion (S20, "Yes"), the audio processing device 40 outputs a predetermined conversation attempt message to the occupants (S30). This conversation attempt message may be output through speakers (S1-S4, see Figure 2) placed at each occupant's seating position. As a result, the occupants in the vehicle 10 can respond to the conversation attempt message with their own voices. The microphone 41 of the audio processing device 40 receives the occupants' response voice signals to the conversation attempt message (S40).

[0077] As a result, the voice processing device 40 processes the received passenger's response voice signal and generates sound source location information indicating the sound source location of the response voice signal (S50). For example, the voice processing device 40 can classify the components of the voice signal according to the sound source location. By employing such a method, the voice processing device 40 can sufficiently generate sound source location information from the passenger's response voice signal.

[0078] After generating the sound source location information of the passengers in this manner, the voice processing device 40 uses the sound source location information to group the response voice signals for each sound source location (i.e., voice signals associated with the voices of the passengers responding to the attempted conversation message at each passenger location) (S60). Assume that there are four passengers SPK1 to SPK4 in the vehicle 10, and that all four passengers SPK1 to SPK4 have sustained minor injuries. In this case, the voice processing device 40 can generate a first group of response voice signals consisting only of the response voice signal of the first passenger SPK1, a second group of response voice signals consisting only of the response voice signal of the second passenger SPK2, a third group of response voice signals consisting only of the response voice signal of the third passenger SPK3, and a fourth group of response voice signals consisting only of the response voice signal of the fourth passenger SPK4.

[0079] From this point onward, the voice processing device 40 outputs the grouped response voice signals (S70). At this time, the voice processing device 40 can transmit the grouped response voice signals, the sound source location information corresponding to each grouped response voice signal, and the vehicle's location information at the time the collision signal was received to the emergency rescue request device 50. The emergency rescue request device 50 can then transmit the grouped response voice signals, the sound source location information corresponding to each grouped response voice signal, and the vehicle's location information at the time the collision signal was received to the control server 70. Of course, if necessary, the voice processing device 40 can also directly transmit the grouped response voice signals to the control server 70, and transmit the sound source location information corresponding to each grouped response voice signal and the vehicle's location information at the time the collision signal was received to the control server 70 via the emergency rescue request device 50.

[0080] As a result, the control server 70 can more accurately determine the number of occupants inside the vehicle at the time of the accident, the current condition of the occupants (severity of injuries, location of injuries), etc., based on the grouped response audio signals, the corresponding sound source location information for each grouped response audio signal, and the vehicle's location information at the time the collision signal was received. This allows the control server 70 to issue instructions for a quicker and more accurate response to the occurrence of a vehicle accident.

[0081] Furthermore, the operation method of the audio processing device of the present invention described above can be implemented as a computer-readable code on a computer-readable recording medium. A computer-readable recording medium includes all types of recording devices on which data read by a computer system is stored. Examples of computer-readable recording media include ROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. Moreover, the computer-readable recording media can be distributed across a network of connected computer systems, and the computer-readable code can be stored and executed in a distributed manner. Functional programs, code, and code segments for implementing the above method can be easily inferred by programmers in the technical field to which the present invention belongs.

[0082] The above description is merely illustrative of the technical concept of the present invention, and a person with ordinary skill in the art to which the present invention pertains can make various modifications and variations within the bounds of the essential characteristics of the present invention. Therefore, the embodiments disclosed herein are for illustrative purposes only, not to limit the technical concept of the present invention, and the scope of the technical concept of the present invention is not limited by such embodiments. The scope of protection of the present invention shall be interpreted in accordance with the following claims, and all technical concepts within an equivalent scope shall be interpreted as being included within the scope of the rights of the present invention.

Claims

1. A sound processing device installed in a vehicle, A microphone configured to generate an audio signal related to the voice of a passenger in a vehicle in response to the passenger's voice. A speaker configured to output sound to the occupants inside the vehicle, memory, and A processor configured to load commands stored in the memory and perform one or more operations by executing the commands, The aforementioned processor, Upon receiving a collision signal from a vehicle controller configured to control the vehicle's operation, a conversation attempt message is output to the occupant via the speaker. The system processes the passenger's response audio signal to the attempted conversation message received from the microphone to generate sound source location information indicating the sound source location of the response audio signal. The response audio signals are grouped according to the sound source location information, A voice processing device characterized by outputting the grouped response voice signals.

2. The memory stores multiple candidate conversation attempt messages, The aforementioned processor, The audio processing device according to claim 1, wherein a first candidate conversation attempt message is output as audio from among the plurality of candidate conversation attempt messages, and a second candidate conversation attempt message following the first candidate conversation attempt message is output as audio based on the passenger's response to the first candidate conversation attempt message.

3. The memory stores multiple candidate conversation attempt messages, The aforementioned processor, The sound source location information is transmitted to the vehicle controller. The vehicle controller receives injury estimation information indicating the degree and location of injury of the occupant at the sound source location, which is estimated based on the sound source location information and sensor sensing information. The voice processing device according to claim 1, which outputs a candidate conversation attempt message corresponding to the injury estimation information from among the plurality of candidate conversation attempt messages.

4. The aforementioned processor, The response audio signal at the aforementioned sound source location is analyzed to determine the pitch and timbre of the response audio signal. Based on the determined pitch and tone, it is determined whether the passenger at the sound source location is a child passenger. The sound processing device according to claim 1, which, when it is determined that the passenger at the sound source location is a child passenger, outputs identity information indicating that the passenger is a child passenger.

5. The memory stores multiple candidate conversation attempt messages, The aforementioned processor, The voice processing device according to claim 4, which outputs a candidate conversation attempt message corresponding to the child passenger from among the plurality of candidate conversation attempt messages.

6. The aforementioned processor, The voice processing device according to claim 1, which outputs the grouped response voice signals to an emergency rescue request device provided in the vehicle.

7. The aforementioned processor, The audio processing device according to claim 1, which outputs the grouped response audio signals to a control server and outputs sound source location information corresponding to the grouped response audio signals to an emergency rescue request device installed in the vehicle.

8. The aforementioned processor, The vehicle's location information is received from the aforementioned vehicle controller. The audio processing device according to claim 1, which outputs vehicle position information at the time the collision occurrence signal is received, along with the grouped response audio signals.

9. A method for operating an audio processing device installed in a vehicle, Upon receiving a collision signal from a vehicle controller configured to control the operation of the vehicle, the step of outputting a voice message attempting to communicate to the occupants of the vehicle, A step of processing the passenger's response audio signal to the attempted conversation message and generating sound source location information indicating the sound source location of the response audio signal, The steps of grouping the response audio signals according to the sound source location information, and A method for operating an audio processing device, characterized by including the step of outputting the grouped response audio signals.

10. A vehicle controller configured to control the operation of a vehicle, and A vehicle control system comprising an audio processing device configured to receive a collision occurrence signal from the vehicle controller, output a conversation attempt message to the occupants of the vehicle, process the occupants' response audio signals to the conversation attempt message to generate sound source location information indicating the sound source location of the response audio signals, group the response audio signals according to the sound source location information, and output the grouped response audio signals.

11. The aforementioned voice processing device stores multiple candidate conversation attempt messages, The aforementioned audio processing device is The sound source location information is transmitted to the vehicle controller. The vehicle controller receives injury estimation information indicating the degree and location of injury of the occupant at the sound source location, which is estimated based on the sound source location information and sensor sensing information. The vehicle control system according to claim 10, wherein a candidate conversation attempt message corresponding to the injury estimation information is output as voice from among the plurality of candidate conversation attempt messages.

12. The aforementioned audio processing device is The timbre is determined by analyzing the response audio signal at the aforementioned sound source location. Based on the determined tone, it is determined whether the passenger at the sound source location is a child passenger. The vehicle control system according to claim 10, which, when it is determined that the passenger at the sound source location is a child passenger, outputs identity information indicating that the passenger is a child passenger.

13. The aforementioned vehicle controller is The sound source location information, the identity information, and the sensor sensing information are received. Based on the sound source location information, the identity information, and the sensor sensing information, the extent and location of injuries to the child passenger are predicted. The vehicle control system according to claim 12, wherein injury prediction information based on the above prediction is output to the voice processing device.

14. The aforementioned audio processing device is The vehicle control system according to claim 13, which receives the injury prediction information and outputs a voice message to the child passenger that attempts to communicate with the injury prediction information.

15. The aforementioned voice processing device stores multiple candidate conversation attempt messages, The aforementioned audio processing device is From among the aforementioned multiple candidate conversation attempt messages, the first candidate conversation attempt message is output as audio. The vehicle control system according to claim 10, wherein a second candidate conversation attempt message following the first candidate conversation attempt message is output as audio based on the passenger's response to the first candidate conversation attempt message.