Acoustic analysis device, acoustic analysis method, and program
The acoustic analysis device identifies silent periods and sound sources in emergency calls to predict the incident scene, enhancing communication and response efficiency in emergency situations.
Patent Information
- Application Number
- JP2023567663
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-16
- Filing Date
- 2022-11-30
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-11-30
AI Technical Summary
Existing emergency call systems struggle to effectively communicate with callers who are confused or unable to speak, leading to difficulties in quickly and accurately responding to incidents.
An acoustic analysis device that identifies silent periods in an acoustic signal, separates sound sources, and predicts the acoustic scene at the incident location based on identified sound sources, using machine learning and noise removal techniques.
Enables rapid and accurate understanding of the incident scene, allowing operators to provide prompt and accurate instructions to callers.
Smart Images

Figure 0007732520000001 
Figure 0007732520000002 
Figure 0007732520000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an acoustic analysis device, an acoustic analysis method, and a program, and more particularly to an acoustic analysis device, an acoustic analysis method, and a program that analyze an acoustic signal from a reporter when reporting an incident, for example. [Background technology]
[0002] Each country and region has its own designated emergency telephone number, such as 110 or 119 in Japan, 911 in the United States and Canada, 000 in Australia, 999 in the United Kingdom, and 112 or 110 in Germany. When an emergency call (hereinafter simply referred to as a call) is received from a caller, an operator at a command center, i.e., a call receiver, confirms with the caller the type of incident (whether it is a crime or an accident), the location of the incident, the time of the incident, etc., and asks the caller about the situation and environment at the scene of the incident. The receiver then inputs information about the incident obtained from the caller into a command system using a terminal or the like at the command center that issues commands to emergency personnel. Patent Document 1 discloses an emergency response support system that supports emergency response activities.
[0003] The ambulance operation support system described in Patent Document 1 converts the voice contained in the acoustic signal into text data. The ambulance operation support system then records the text data and displays the sentences corresponding to the text data on a terminal. This allows the exchange between the call receiver and the caller to be saved without error. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent Publication No. 2021-93228 [Patent Document 2] International Publication No. 2021 / 014649 [Patent Document 3] International Publication No. 2020 / 184631 Summary of the Invention [Problem to be solved by the invention]
[0005] When the incident is complicated or the caller is confused, the caller may not be able to communicate effectively with the caller. The caller may also be unable to speak. In such cases, it is difficult for the caller to respond quickly and accurately to the incident, such as issuing commands to emergency personnel, based solely on the conversation with the caller.
[0006] The present invention has been made in view of the above-mentioned problems, and its purpose is to help the receiver of a report deal with the case promptly and accurately. [Means for solving the problem]
[0007] An acoustic analysis device according to one aspect of the present invention includes an identification means for identifying silent periods in an input acoustic signal when a reporter reporting an incident is not speaking, an identification means for identifying a sound source of a sound contained in the acoustic signal during the silent periods, and a prediction means for predicting an acoustic scene at the scene of the incident based on the identified sound source.
[0008] An acoustic analysis method according to one aspect of the present invention identifies silent periods in an input acoustic signal when a reporter reporting an incident is not speaking, identifies the source of the sound contained in the acoustic signal during the silent periods, and predicts the acoustic scene at the scene where the incident occurred based on the identified sound source.
[0009] A recording medium according to one aspect of the present invention stores a program for causing a computer to execute the following steps: identify silent periods in an input acoustic signal during which a reporter reporting an incident is not speaking; identify the sound source of the sound contained in the acoustic signal during the silent periods; and predict the acoustic scene at the scene of the incident based on the identified sound source. [Effects of the Invention]
[0010] According to one aspect of the present invention, it is possible to help the recipient of a report deal with the incident quickly and accurately. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a diagram schematically illustrating an example of the configuration of a command system to which an acoustic analysis device according to any one of the first to fourth embodiments can be applied. [Figure 2] 1 is a block diagram showing the configuration of an acoustic analysis device according to a first embodiment. [Figure 3] 4 is a flowchart showing the operation of the acoustic analysis device according to the first embodiment. [Figure 4] FIG. 10 is a block diagram showing the configuration of an acoustic analysis device according to a second embodiment. [Figure 5] An example of information showing the situation, environment, and actions of people at the scene of an incident is shown below. [Figure 6] 10 is a flowchart showing the operation of the acoustic analysis device according to the second embodiment. [Figure 7] FIG. 10 is a block diagram showing the configuration of an acoustic analysis device according to a third embodiment. [Figure 8] 10 is a flowchart showing the operation of the acoustic analysis device according to the third embodiment. [Figure 9] FIG. 10 is a block diagram showing the configuration of an acoustic analysis device according to a fourth embodiment. [Figure 10] 10 is a flowchart showing the operation of the acoustic analysis device according to the fourth embodiment. [Figure 11] 1 is a diagram illustrating an example of a hardware configuration of an acoustic analysis device according to first to fourth embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0012] Embodiments of the present invention will be described below with reference to the drawings.
[0013] (Command System 1) A command system 1 to which any of the acoustic analysis devices 10, 20, 30, and 40 according to the first to fourth embodiments described below can be applied will be described with reference to Fig. 1. The command system 1 is used by a commander to receive reports, issue commands to the scene, and support emergency activities in emergency activities such as firefighting, rescue, relief, accident handling, and maintaining public order. Fig. 1 is a diagram schematically illustrating an example of the configuration of the command system 1.
[0014] 1, the command system 1 includes an acoustic analysis device 10 (20, 30, 40) and an OA (Office Automation) terminal 100 used by an operator (receiver). Here, the "acoustic analysis device 10 (20, 30, 40)" refers to any one of the acoustic analysis devices 10, 20, 30, 40 according to the first to fourth embodiments described below.
[0015] The OA terminal 100 includes a telephone, an input device, a speaker, a personal computer, a display, a monitor, etc. The OA terminal 100 is connected to the acoustic analysis device 10 (20, 30, 40) via a LAN (Local Area Network) of the command system 1.
[0016] The OA terminal 100 is also configured to enable a call between a caller reporting an incident and an operator (receiver) via the acoustic analysis device 10 (20, 30, 40). Incidents include accidents such as traffic accidents and emergency medical emergencies, as well as fires, floods, power outages, other disasters, wild animal sightings, and crimes. Generally, the incidents referred to here are those handled by emergency services, fire departments, or police.
[0017] In FIG. 1, examples of questions that an operator (recipient) may ask the reporter are listed on the right side of the OA terminal 100. For example, the questions to the reporter include the type of incident. The questions to the reporter also include when and where the incident occurred, whether there were any witnesses to the incident, the name of the reporter, and the situation at the scene. Depending on the details of the type of incident (for example, whether it was an accident between two cars or a pedestrian accident), the questions to the reporter may differ from those shown in FIG. 1.
[0018] When the command system 1 receives a call from a caller, the acoustic analysis device 10 (20, 30, 40) receives, via a telephone line or an IP (Internet Protocol) line, an acoustic signal input to the communication device used by the caller to make the call. The acoustic signal may contain background sounds in addition to the caller's voice.
[0019] For example, background sounds include information about sounds emitted from sources present at or in the scene of the incident. Examples of sound sources include people other than the caller, animals, trains, automobiles, machines, speakers, and alarms. Background sounds may also include information about the geography (e.g., urban areas, industrial areas, roadsides, mountains, seaside) and weather (e.g., rain, wind, thunderstorms) of the scene of the incident.
[0020] The acoustic analysis device 10 (20, 30, 40) performs acoustic analysis on the received acoustic signal. The acoustic analysis device 10 (20, 30, 40) also transfers the received acoustic signal to the OA terminal 100 used by the operator (receiver). This allows the acoustic analysis device 10 (20, 30, 40) to perform acoustic analysis on the acoustic signal without interrupting the call between the operator (receiver) and the caller.
[0021] The acoustic analysis device 10 (20, 30, 40) may be a part of a command control device that controls the command line of the command system 1 and realizes the functions of the command system 1.
[0022] The functions of the acoustic analysis device 10 (20, 30, 40) will be described in detail in the following first to fourth embodiments.
[0023] [Embodiment 1] A first embodiment will be described with reference to FIGS.
[0024] (Acoustic analyzer 10) The configuration of the acoustic analysis device 10 according to the first embodiment will be described with reference to Fig. 2. Fig. 2 is a block diagram showing the configuration of the acoustic analysis device 10.
[0025] As shown in FIG. 2, the acoustic analysis device 10 includes a specifying unit 11, a classifying unit 12, and a predicting unit 13.
[0026] The identification unit 11 identifies, in the input acoustic signal, silent periods during which the caller reporting the occurrence of an incident is not speaking. The identification unit 11 is an example of an identification means. Incidents include accidents such as traffic accidents and emergency medical emergencies, as well as fires, floods, power outages, other disasters, wild animal sightings, and crimes. Generally, incidents here are those handled by emergency services, fire departments, or police.
[0027] In one example, when a call is made from a caller to a telephone line (e.g., 119) of the command system 1 (FIG. 1), the identification unit 11 receives an acoustic signal from the caller's communication device via the telephone line or IP line. The acoustic signal includes background sounds in addition to the caller's voice. For example, if it is raining at the scene of the incident, the acoustic signal may include the sound of rain as background sounds.
[0028] First, the determination unit 11 uses a noise removal technique such as a digital filter or a well-known noise removal algorithm to remove components whose frequencies do not change significantly over time from the audio signal, thereby enabling the determination unit 11 to remove noise from the audio signal.
[0029] Next, the identification unit 11 applies sound source separation technology in the field of machine learning to the audio signal from which noise has been removed to separate the caller's voice contained in the audio signal from other sounds (i.e., background sounds). This enables the identification unit 11 to distinguish between time periods in which the caller's voice is present and time periods in which the caller's voice is not present in the audio signal.
[0030] The identification unit 11 identifies a time period in which there is no voice from the reporter as a silent time period in which the reporter is not speaking. As described above, the acoustic signal in the silent time period may contain background sound.
[0031] The identification unit 11 outputs the acoustic signal during the silent time to the classification unit 12. Alternatively, the identification unit 11 may output information specifying the silent time together with the acoustic signal to the classification unit 12. In this case, the classification unit 12, which will be described later, uses the information specifying the silent time to extract the acoustic signal during the silent time from the entire acoustic signal.
[0032] The identification unit 12 identifies a sound source at the scene of an incident by analyzing the acoustic signal during silent time. The identification unit 12 is an example of an identification means.
[0033] In one example, the identification unit 12 receives an acoustic signal during a non-speech time from the identification unit 11. The identification unit 12 determines whether the acoustic signal during a non-speech time contains strong reverberation. If the acoustic signal during a non-speech time contains strong reverberation, the identification unit 12 identifies the scene of the incident as a closed space (e.g., indoors). On the other hand, if the acoustic signal during a non-speech time does not contain reverberation or contains weak reverberation, the identification unit 12 identifies the scene of the incident as a semi-open space or an open space (e.g., outdoors).
[0034] The classification unit 12 also uses the machine-learned model to determine whether the acoustic signal during silent periods contains a characteristic sound. A characteristic sound is a sound whose source can be identified, and includes, for example, the sound of a train or car running, an announcement on a station platform, the sound of a traffic light for the visually impaired, voices or music repeatedly played in chain stores such as electronics retailers or grocery stores, and the commotion or screams of a crowd.
[0035] The identification unit 12 identifies the sound source based on characteristic sounds contained in the acoustic signal. The identification unit 12 may identify the sound source of a sound associated with the acceptance procedure for the incident that has occurred. The acceptance procedure specifies the basic procedure for accepting an incident from a report, etc. The acceptance procedure may differ depending on the type of incident. For example, the acceptance procedure when the incident is an emergency is different from the acceptance procedure when the incident is a fire. Therefore, the sound source identified by the identification unit 12 may differ depending on the type of incident.
[0036] The identification unit 12 outputs the sound source identification result to the prediction unit 13. The sound source identification result includes information indicating the sound source identified from the acoustic signal in the non-speech time.
[0037] The prediction unit 13 predicts an acoustic scene at the scene of the incident based on the identified sound source. The prediction unit 13 is an example of a prediction means. The acoustic scene refers to a scene or situation implied by the acoustic signal. The acoustic scene includes the situation, environment, and human behavior at the scene of the incident.
[0038] In one example, the prediction unit 13 receives a sound source identification result from the identification unit 12. The prediction unit 13 extracts information indicating the sound source identified from the acoustic signal during the silent time from the sound source identification result. The prediction unit 13 refers to a database (not shown) that stores a table linking sound sources with acoustic scenes. The prediction unit 13 then compares the sound sources listed in the table with the sound sources identified from the acoustic signal during the silent time to predict the acoustic scene at the scene where the incident occurred.
[0039] Thereafter, the prediction unit 13 may display information based on the predicted acoustic scene on the OA terminal 100 (FIG. 1) (Embodiment 2). Alternatively, the prediction unit 13 may record information indicating the predicted acoustic scene in a recording medium such as the ROM 902 (FIG. 11).
[0040] (Operation of the acoustic analysis device 10) The operation of the acoustic analysis device 10 according to the first embodiment will be described with reference to Fig. 3. Fig. 3 is a flowchart showing the flow of processing executed by each unit of the acoustic analysis device 10.
[0041] As shown in FIG. 3, the identifying unit 11 identifies a silent period in the acoustic signal input to the command system 1 (FIG. 1) during which the reporter reporting the occurrence of an incident is not speaking (S101).
[0042] The specifying unit 11 outputs the acoustic signal during the silent time to the classifying unit 12 .
[0043] Next, the identification unit 12 identifies the sound source of the sound contained in the acoustic signal during the silent time (S102).
[0044] The identification unit 12 outputs the sound source identification result to the prediction unit 13. The sound source identification result includes information indicating the sound source identified from the acoustic signal in the non-speech time.
[0045] Thereafter, the prediction unit 13 predicts an acoustic scene based on the identified sound source (S103).
[0046] Furthermore, the prediction unit 13 may display information based on the predicted acoustic scene on the OA terminal 100 (FIG. 1) (Embodiment 2). Alternatively, the prediction unit 13 may record information indicating the predicted acoustic scene in a recording medium such as the ROM 902 (FIG. 11).
[0047] After step S101, the identification unit 11 may output information identifying the silent time together with the acoustic signal to the classification unit 12. In this case, in step S102, the classification unit 12 extracts the acoustic signal in the silent time from the entire acoustic signal using the information identifying the silent time.
[0048] This completes the operation of the acoustic analysis device 10 according to the first embodiment.
[0049] (Effects of this embodiment) According to the configuration of this embodiment, the identification unit 11 identifies silent periods in the input acoustic signal during which the reporter reporting the occurrence of an incident is not speaking. The identification unit 12 identifies sound sources at the scene of the incident by analyzing the acoustic signal during silent periods. The prediction unit 13 predicts the acoustic scene at the scene of the incident based on the identified sound sources. In this way, the acoustic scene at the scene of the incident is predicted from the input acoustic signal. The recipient of the report can understand the situation, state, scene, environment, etc. at the scene of the incident from the predicted acoustic scene. This allows the recipient of the report to respond to the incident quickly and accurately.
[0050] [Embodiment 2] A second embodiment will be described with reference to Figs. 4 to 6. In the second embodiment, a configuration will be described in which information indicating the situation, environment, and human behavior corresponding to the predicted acoustic scene is provided to the receiver of the message. In the second embodiment, the same reference numerals as those in the first embodiment will be used for the components described in the first embodiment, and the description thereof will be omitted.
[0051] (Acoustic analyzer 20) The configuration of the acoustic analysis device 20 according to the second embodiment will be described with reference to Fig. 4. Fig. 4 is a block diagram showing the configuration of the acoustic analysis device 20.
[0052] 4, the acoustic analysis device 20 includes a specifying unit 11, a classifying unit 12, and a predicting unit 13. The acoustic analysis device 20 further includes an output unit .
[0053] The output unit 24 outputs information indicating the situation, environment, and human behavior corresponding to the predicted acoustic scene. The output unit 24 is an example of an output means.
[0054] In one example, the output unit 24 receives information indicating an acoustic scene at the scene where an incident occurs from the prediction unit 13. Alternatively, the output unit 24 obtains information indicating the predicted acoustic scene from a recording medium such as the ROM 902 (FIG. 11).
[0055] The output unit 24 generates information indicating the situation, environment, and behavior of people at the scene of the incident based on the information indicating the sound scene at the scene of the incident. For example, the output unit 24 refers to a database (not shown) that stores a table linking the sound scene with the situation, environment, and behavior of people. The output unit 24 then compares the sound scene recorded in the table with the sound scene at the scene of the incident to determine the situation, environment, and behavior of people at the scene of the incident.
[0056] Then, the output unit 24 outputs information indicating the situation, environment, and behavior of people at the scene of the incident to the OA terminal 100 used by the operator (receiver) (FIG. 1).
[0057] The operator (receiver) can estimate the situation, environment, and human behavior corresponding to the predicted acoustic scene by checking the information output to the OA terminal 100. This allows the operator (receiver) to more smoothly progress the conversation with the caller (FIG. 1).
[0058] (An example of the progression of a conversation between an operator and a caller) With reference to Fig. 5, information output by the output unit 24 according to the second embodiment to the OA terminal 100 or the like will be described. Fig. 5 shows an example of the progress of a conversation between an operator (recipient) and a reporter when a fire (one example of an incident) occurs. Fig. 5 also shows an example of information indicating the situation, environment, and behavior of people at the scene of the incident. The direction from left to right in Fig. 5 corresponds to the direction in which time progresses.
[0059] In Figure 5, the dashed-line box in the center shows an example of information indicating the situation, environment, and behavior of people at the scene of the incident. The top shows an example of a caller's statement, and the bottom shows an example of an operator's (receiver's) statement. The arrow extending upward from the box surrounding the operator's statement indicates a question or confirmation from the operator (receiver) to the caller. On the other hand, the arrow extending downward from the box surrounding the caller's statement indicates a reply from the caller to the operator (receiver).
[0060] 5, first, as information indicating the environment at the scene of the incident, "indoors" is presented to the operator (recipient) by the output unit 24. Furthermore, as information indicating the behavior of a person at the scene of the incident, "stop" is presented to the operator (recipient) by the output unit 24.
[0061] By checking the information output to the OA terminal 100, the operator (receiver) can know that the environment at the scene of the incident is "indoors." Therefore, the operator (receiver) does not need to ask, "Where are you now?" to find out the current location of the caller. The operator (receiver) can omit the question, "Where are you now?"
[0062] For example, the operator (receiver) may ask the caller, "Are you still indoors? Please quickly get outside," without asking about the caller's current location. The caller can quickly begin evacuation actions following instructions from the operator (receiver).
[0063] Second, "walking" is presented to the operator (receiver) by the output unit 24 as information indicating the behavior of a person at the scene of an incident. Therefore, the operator (reporter) does not need to ask the question "Can you walk?" to find out whether the reporter can walk. The operator (reporter) can omit the question "Can you walk?"
[0064] For example, the operator (receiver) may ask the caller, "Do you know where the exit is?" The caller can quickly begin evacuation actions based on instructions from the operator (receiver).
[0065] Third, "outdoors" is presented to the operator (receiver) by the output unit 24 as information indicating the environment at the scene of the incident. Furthermore, "no rain" and "strong wind" are presented to the operator (receiver) by the output unit 24 as information indicating the weather at the scene of the incident. Therefore, the operator (reporter) does not need to ask questions such as "What's it like outside? Is it windy?" in order to know the weather at the scene of the incident. The operator (receiver) can omit asking questions such as "What's it like outside? Is it windy?"
[0066] For example, the operator (receiver) may ask the caller, "It looks like the wind is strong. Is the house next door OK?" The caller responds to the question from the operator (receiver).
[0067] In this way, the operator (receiver) can omit some of the questions to the caller by referring to the information presented by the output unit 24. This allows the operator to quickly advance the conversation with the caller and give accurate instructions to the caller.
[0068] (Operation of the acoustic analysis device 20) The operation of the acoustic analysis device 20 according to the second embodiment will be described with reference to Fig. 6. Fig. 6 is a flowchart showing the flow of processing executed by each unit of the acoustic analysis device 20.
[0069] As shown in FIG. 6, the identifying unit 11 identifies a silent period in the acoustic signal input to the command system 1 (FIG. 1) during which the reporter reporting the occurrence of an incident is not speaking (S201).
[0070] The specifying unit 11 outputs the acoustic signal during the silent time to the classifying unit 12 .
[0071] Next, the identification unit 12 identifies the sound source of the sound contained in the acoustic signal during the silent time (S202).
[0072] The identification unit 12 outputs the sound source identification result to the prediction unit 13. The sound source identification result includes information indicating the sound source identified from the acoustic signal in the non-speech time.
[0073] Thereafter, the prediction unit 13 predicts an acoustic scene based on the identified sound source (S203).
[0074] Furthermore, the prediction unit 13 may record information indicating the predicted acoustic scene in a recording medium such as the ROM 902 (FIG. 11).
[0075] After step S201, the identification unit 11 may output information identifying the silent time together with the audio signal to the classification unit 12. In this case, in step S202, the classification unit 12 extracts the audio signal in the silent time from the entire audio signal using the information identifying the silent time.
[0076] The prediction unit 13 outputs information indicating the predicted acoustic scene to the output unit 24.
[0077] The output unit 24 outputs information indicating the situation, environment, and human behavior corresponding to the predicted acoustic scene (S204).
[0078] This completes the operation of the acoustic analysis device 20 according to the second embodiment.
[0079] (Effects of this embodiment) According to the configuration of this embodiment, the identification unit 11 identifies silent periods in the input acoustic signal during which the reporter reporting the occurrence of an incident is not speaking. The identification unit 12 identifies sound sources at the scene of the incident by analyzing the acoustic signal during silent periods. The prediction unit 13 predicts the acoustic scene at the scene of the incident based on the identified sound sources. In this way, the acoustic scene at the scene of the incident is predicted from the input acoustic signal. The recipient of the report can understand the situation, state, scene, environment, etc. at the scene of the incident from the predicted acoustic scene. This allows the recipient of the report to respond to the incident quickly and accurately.
[0080] Furthermore, according to the configuration of this embodiment, the output unit 24 outputs information indicating the situation, environment, and human behavior corresponding to the predicted acoustic scene, thereby making it possible to provide the recipient of the message with information indicating the situation, environment, and human behavior corresponding to the predicted acoustic scene.
[0081] [Embodiment 3] A third embodiment will be described with reference to Figs. 7 and 8. In the third embodiment, a configuration will be described in which an acoustic scene is predicted using the results of speech recognition of the caller's voice in addition to the identified sound source. In the third embodiment, the components described in the first or second embodiment will be assigned the same reference numerals as those in the first or second embodiment, and the description thereof will be omitted.
[0082] (Acoustic analyzer 30) The configuration of an acoustic analysis device 30 according to the third embodiment will be described with reference to Fig. 7. Fig. 7 is a block diagram showing the configuration of the acoustic analysis device 30.
[0083] 7, the acoustic analysis device 30 includes a specifying unit 11, a classifying unit 12, and a predicting unit 13. The acoustic analysis device 30 further includes a voice recognition unit .
[0084] The speech recognition unit 34 performs speech recognition on speech in a predetermined language from the speech time excluding the non-speech time of the inputted speech signal. The speech recognition unit 34 is an example of a speech recognition means.
[0085] In one example, the speech recognition unit 34 receives an acoustic signal during speech time excluding non-speech time from the identification unit 11. The speech recognition unit 34 analyzes the acoustic signal during speech time using speech recognition techniques such as pattern matching and a language model (e.g., a recurrent neural network language model). As a result of the analysis, the speech recognition unit 34 obtains text data converted from the acoustic signal during speech time. The text data is sentence information expressed in a predetermined language.
[0086] Although the subject of the voice recognized by the voice recognition unit 34 is usually the caller, the possibility of it being someone other than the caller cannot be excluded, because the acoustic signal during the voice time may contain the voice of a person other than the caller.
[0087] The speech recognition unit 34 outputs text data converted from the acoustic signal in the speech time to the prediction unit 13.
[0088] In the third embodiment, the prediction unit 13 predicts an acoustic scene at the scene of an incident based on the sound source identified from the acoustic signal as well as the result of the speech recognition unit 34 recognizing the speech included in the acoustic signal. For example, the prediction unit 13 extracts keywords (e.g., fire, rain, train, etc.) from the result of recognizing the speech included in the acoustic signal. Then, the prediction unit 13 refers to a table that associates preset keywords with situations, environments, or human behavior, and identifies the situation, environment, or human behavior that corresponds to the extracted keyword. The prediction unit 13 includes the identified situation, environment, or human behavior in the elements for predicting the acoustic scene at the scene of an incident. This allows the prediction unit 13 to more accurately predict the acoustic scene at the scene of an incident.
[0089] (Operation of the acoustic analysis device 30) The operation of the acoustic analysis device 30 according to the third embodiment will be described with reference to Fig. 8. Fig. 8 is a flowchart showing the flow of processing executed by each unit of the acoustic analysis device 30.
[0090] As shown in FIG. 8, the identifying unit 11 identifies a silent period in the acoustic signal input to the command system 1 (FIG. 1) during which the reporter reporting the occurrence of an incident is not speaking (S301).
[0091] The identification unit 11 outputs the acoustic signal during the silent time to the identification unit 12. The identification unit 11 also outputs the acoustic signal during the speech time other than the silent time to the speech recognition unit .
[0092] Next, the identification unit 12 identifies the sound source of the sound contained in the acoustic signal during the silent time (S302).
[0093] The identification unit 12 outputs the sound source identification result to the prediction unit 13. The sound source identification result includes information indicating the sound source identified from the acoustic signal in the non-speech time.
[0094] The speech recognition unit 34 performs speech recognition on speech in a predetermined language from the speech signal during speech time excluding non-speech time in the inputted speech signal (S303).
[0095] The speech recognition unit 34 outputs text data converted from the acoustic signal in the speech time to the prediction unit 13.
[0096] Thereafter, the prediction unit 13 predicts an acoustic scene based on the identified sound source and the speech recognition result (S304).
[0097] Furthermore, the prediction unit 13 may display information based on the predicted acoustic scene on the OA terminal 100 (FIG. 1) (Embodiment 2). Alternatively, the prediction unit 13 may record information indicating the predicted acoustic scene in a recording medium such as the ROM 902 (FIG. 11).
[0098] After step S301, the identification unit 11 may output information identifying the silent time together with the audio signal to the classification unit 12. In this case, in step S302, the classification unit 12 extracts the audio signal in the silent time from the entire audio signal using the information identifying the silent time.
[0099] This completes the operation of the acoustic analysis device 30 according to the third embodiment.
[0100] (Variation) In one variation, the acoustic analysis device 30 may further include the output unit 24 (FIG. 4) described in the second embodiment. As in the second embodiment, the output unit 24 outputs information indicating the situation, environment, and human behavior corresponding to the predicted acoustic scene. In addition, in the third embodiment, the output unit 24 may further output the result of recognition of the voice contained in the acoustic signal by the voice recognition unit 34. For example, the output unit 24 receives text data converted from the acoustic signal in the voice time from the voice recognition unit 34. Then, the output unit 24 converts the received text data into character image data and displays the character image data on the screen of the OA terminal 100 (FIG. 1) or the like.
[0101] According to the configuration of this modified example, the operator (receiver) can visually check the message of the reporter using the OA terminal 100, which can prevent misrecognition of the incident due to mishearing.
[0102] (Effects of this embodiment) According to the configuration of this embodiment, the identification unit 11 identifies silent periods in the input acoustic signal during which the reporter reporting the occurrence of an incident is not speaking. The identification unit 12 identifies sound sources at the scene of the incident by analyzing the acoustic signal during silent periods. The prediction unit 13 predicts the acoustic scene at the scene of the incident based on the identified sound sources. In this way, the acoustic scene at the scene of the incident is predicted from the input acoustic signal. The recipient of the report can understand the situation, state, scene, environment, etc. at the scene of the incident from the predicted acoustic scene. This allows the recipient of the report to respond to the incident quickly and accurately.
[0103] Furthermore, according to the configuration of this embodiment, the speech recognition unit 34 recognizes speech in a predetermined language from the inputted speech signal during speech time excluding non-speech time. The prediction unit 13 predicts the acoustic scene at the scene of the incident based on the results of the speech recognition unit 34 recognizing the speech contained in the speech signal, in addition to the sound source identified from the speech signal. The prediction unit 13 includes the situation, environment, or human behavior identified from the speech recognition result as factors for predicting the acoustic scene at the scene of the incident. This allows the prediction unit 13 to more accurately predict the acoustic scene at the scene of the incident.
[0104] [Embodiment 4] A fourth embodiment will be described with reference to Figs. 9 and 10. In this fourth embodiment, a configuration will be described in which an acoustic scene is predicted by using the result of emotion recognition in addition to the identified sound source. In this fourth embodiment, the configurations described in at least one of the first to third embodiments will be assigned the same reference numerals as those in the first to third embodiments, and the description thereof will be omitted.
[0105] (Acoustic analyzer 40) The configuration of an acoustic analysis device 40 according to the fourth embodiment will be described with reference to Fig. 9. Fig. 9 is a block diagram showing the configuration of the acoustic analysis device 40.
[0106] 9, the acoustic analysis device 40 includes a specification unit 11, a classification unit 12, and a prediction unit 13. The acoustic analysis device 40 further includes an emotion recognition unit 44.
[0107] The emotion recognition unit 44 recognizes emotions from the audio signal during the audio time excluding the non-audio time in the input audio signal. The emotion recognition unit 44 is an example of emotion recognition means.
[0108] In one example, the emotion recognition unit 44 receives an audio signal during audio time excluding non-audio time from the identification unit 11. The emotion recognition unit 44 analyzes the audio signal during audio time using an emotion recognition technique such as emotion learning using a DNN (Deep Neural Network). As a result of the analysis, the emotion recognition unit 44 obtains information indicating an emotion recognized from the audio signal during audio time. The information indicating an emotion represents an emotional pattern such as "joy," "sadness," or "anger."
[0109] Although the subject of the emotion recognized by the emotion recognition unit 44 is usually the caller, the possibility of it being someone other than the caller cannot be excluded, because the acoustic signal during the audio time may contain the voice of someone other than the caller.
[0110] The emotion recognition unit 44 outputs information indicating the emotion recognized from the audio signal in the speech time to the prediction unit 13.
[0111] In the fourth embodiment, the prediction unit 13 predicts the acoustic scene at the scene of the incident based on the result of emotion recognition by the emotion recognition unit 44 in addition to the sound source identified from the acoustic signal. For example, the prediction unit 13 extracts an emotion pattern from the result of emotion recognition. Then, the prediction unit 13 refers to a table that associates preset emotion patterns with situations, environments, or human behavior, and identifies the situation, environment, or human behavior that corresponds to the extracted emotion pattern. The prediction unit 13 includes the identified situation, environment, or human behavior in the elements for predicting the acoustic scene at the scene of the incident. This allows the prediction unit 13 to more accurately predict the acoustic scene at the scene of the incident.
[0112] (Operation of the acoustic analysis device 40) The operation of the acoustic analysis device 40 according to the fourth embodiment will be described with reference to Fig. 10. Fig. 10 is a flowchart showing the flow of processing executed by each unit of the acoustic analysis device 40.
[0113] As shown in FIG. 10, the identifying unit 11 identifies a silent period in the acoustic signal input to the command system 1 (FIG. 1) during which the reporter reporting the occurrence of an incident is not speaking (S401).
[0114] The specification unit 11 outputs the acoustic signal during the silent time to the classification unit 12. The specification unit 11 also outputs the acoustic signal during the speech time other than the silent time to the emotion recognition unit 44.
[0115] Next, the identification unit 12 identifies the sound source of the sound contained in the acoustic signal during the silent time (S402).
[0116] The identification unit 12 outputs the sound source identification result to the prediction unit 13. The sound source identification result includes information indicating the sound source identified from the acoustic signal in the non-speech time.
[0117] The emotion recognition unit 44 recognizes emotions from the audio signal during the audio time excluding the non-audio time in the input audio signal, targeting audio in a predetermined language (S403).
[0118] The emotion recognition unit 44 outputs information indicating the emotion recognized from the audio signal in the speech time to the prediction unit 13.
[0119] Thereafter, the prediction unit 13 predicts an acoustic scene based on the identified sound source and emotion recognition result (S404).
[0120] Furthermore, the prediction unit 13 may display information based on the predicted acoustic scene on the OA terminal 100 (FIG. 1) (Embodiment 2). Alternatively, the prediction unit 13 may record information indicating the predicted acoustic scene in a recording medium such as the ROM 902 (FIG. 11).
[0121] After step S401, the identification unit 11 may output information identifying the silent time together with the audio signal to the classification unit 12. In this case, in step S402, the classification unit 12 extracts the audio signal in the silent time from the entire audio signal using the information identifying the silent time.
[0122] This completes the operation of the acoustic analysis device 40 according to the fourth embodiment.
[0123] (Variation) In one variation, the acoustic analysis device 40 may further include the output unit 24 (FIG. 4) described in the second embodiment. As in the second embodiment, the output unit 24 outputs information indicating the situation, environment, and human behavior corresponding to the predicted acoustic scene. In addition, in the fourth embodiment, the output unit 24 may further output the result of emotion recognition by the emotion recognition unit 44. For example, the output unit 24 receives information indicating the recognized emotional state from the emotion recognition unit 44. Then, the output unit 24 converts the information received from the emotion recognition unit 44 into image data of symbols indicating the emotional state, and displays the image data of the symbols on the screen of the OA terminal 100 (FIG. 1) or the like.
[0124] According to the configuration of this modified example, the operator (receiver) can use the OA terminal 100 to visually confirm the emotional state recognized by the emotion recognition unit 44, thereby enabling better communication with the caller and enabling the conversation with the caller to proceed more smoothly.
[0125] (Effects of this embodiment) According to the configuration of this embodiment, the identification unit 11 identifies silent periods in the input acoustic signal during which the reporter reporting the occurrence of an incident is not speaking. The identification unit 12 identifies sound sources at the scene of the incident by analyzing the acoustic signal during silent periods. The prediction unit 13 predicts the acoustic scene at the scene of the incident based on the identified sound sources. In this way, the acoustic scene at the scene of the incident is predicted from the input acoustic signal. The recipient of the report can understand the situation, state, scene, environment, etc. at the scene of the incident from the predicted acoustic scene. This allows the recipient of the report to respond to the incident quickly and accurately.
[0126] Furthermore, according to the configuration of this embodiment, the emotion recognition unit 44 recognizes emotions from the acoustic signal during the voice time when the caller is speaking. The prediction unit 13 predicts the acoustic scene at the scene of the incident based on the result of emotion recognition by the emotion recognition unit 44 in addition to the sound source identified from the acoustic signal. The prediction unit 13 includes the situation, environment, or human behavior identified from the emotion recognition result as factors for predicting the acoustic scene at the scene of the incident. This allows the prediction unit 13 to more accurately predict the acoustic scene at the scene of the incident.
[0127] (About hardware configuration) Each of the components of the acoustic analysis devices 10, 20, 30, and 40 described in the first to fourth embodiments represents a functional block. Some or all of these components are realized by an information processing device 900 as shown in Fig. 11. Fig. 11 is a block diagram showing an example of the hardware configuration of the information processing device 900.
[0128] As shown in FIG. 11, the information processing device 900 includes, for example, the following configuration.
[0129] ·CPU(Central Processing Unit)901 ROM (Read Only Memory) 902 ·RAM(Random Access Memory)903 Program 904 loaded into RAM 903 A storage device 905 for storing a program 904 A drive device 907 for reading and writing data from and to the recording medium 906 A communication interface 908 for connecting to a communication network 909 An incoming prediction interface 910 for performing incoming prediction of data Bus 911 connecting each component Each of the components of the acoustic analysis devices 10, 20, 30, and 40 described in the first to fourth embodiments is realized by the CPU 901 reading and executing a program 904 that realizes the functions of the components. The program 904 that realizes the functions of the components is stored in advance in, for example, the storage device 905 or the ROM 902, and is loaded into the RAM 903 and executed by the CPU 901 as needed. The program 904 may be supplied to the CPU 901 via the communication network 909, or may be stored in advance in the recording medium 906, and the drive device 907 may read out the program and supply it to the CPU 901.
[0130] According to the above configuration, the acoustic analysis devices 10, 20, 30, and 40 described in the first to fourth embodiments are realized as hardware, and therefore the same effects as those described in any of the first to fourth embodiments can be achieved.
[0131] (Addendum) One embodiment of the present invention is also described as in the following supplementary notes, but is not limited to the following.
[0132] (Appendix 1) a means for identifying a silent period in the input acoustic signal during which a reporter reporting the occurrence of an incident is not speaking; an identification means for identifying a sound source at the scene of the incident by analyzing the acoustic signal during the silent time; a prediction means for predicting an acoustic scene at the scene of the incident based on the identified sound source; An acoustic analysis device equipped with:
[0133] (Appendix 2) The acoustic scene includes the situation, environment, and actions of people at the scene of the incident. 2. The acoustic analysis device according to claim 1,
[0134] (Appendix 3) The identification means identifies the source of the sound associated with the receipt of the incident that occurred. 3. The acoustic analysis device according to claim 1 or 2.
[0135] (Appendix 4) The apparatus further includes an output unit for outputting information indicating a situation, an environment, and a person's behavior corresponding to the identified acoustic scene. 4. The acoustic analysis device according to claim 1, wherein:
[0136] (Appendix 5) The audio signal processing device further includes a voice recognition unit for recognizing voices included in the audio signal during the voice time excluding the non-voice time in the input audio signal. 4. The acoustic analysis device according to claim 1, wherein:
[0137] (Appendix 6) The prediction means predicts an acoustic scene at the scene of the incident based on the sound source identified from the acoustic signal during the silent time as well as the result of recognition by the speech recognition means of the sound included in the acoustic signal during the speech time. 6. The acoustic analysis device according to claim 5,
[0138] (Appendix 7) The audio signal processing device further includes emotion recognition means for recognizing emotions from the audio signal during the audio time excluding the non-audio time. 4. The acoustic analysis device according to claim 1, wherein:
[0139] (Appendix 8) The prediction means predicts an acoustic scene at the scene of the incident based on the emotion recognized by the emotion recognition means from the acoustic signal during the voice time in addition to the sound source identified from the acoustic signal during the non-voice time. 8. The acoustic analysis device according to claim 7,
[0140] (Appendix 9) Identifying a non-voice period in which a reporter reporting an incident is not speaking in the input acoustic signal; identifying sound sources at the incident scene by analyzing acoustic signals during the non-speech time; Predicting an acoustic scene at the scene of the incident based on the identified sound source. Acoustic analysis methods.
[0141] (Appendix 10) Identifying a non-voice period in which a reporter reporting the occurrence of an incident is not speaking in the input acoustic signal; identifying sound sources at the incident scene by analyzing acoustic signals during the non-speech time; predicting an acoustic scene at the scene of the incident based on the identified sound sources; A program that causes a computer to execute the following.
[0142] Although the present invention has been described above with reference to the embodiments (and examples), the present invention is not limited to the above-described embodiments (and examples). Various modifications that can be understood by those skilled in the art can be made to the configurations and details of the embodiments (and examples) within the scope of the present invention.
[0143] This application claims priority based on Japanese Patent Application No. 2021-204070, filed on December 16, 2021, the entire disclosure of which is incorporated herein by reference. [Industrial Applicability]
[0144] The present invention can be used, for example, in an emergency command system, to provide information indicating the location of an incident by analyzing the acoustic signal from the caller when reporting an incident. [Explanation of symbols]
[0145] 1. Command System 10 Acoustic analysis device 11 Specific section 12 Identification unit 13 Prediction Department 20 Acoustic analysis device 24 Output section 30 Acoustic analysis device 34 Voice Recognition Unit 40 Acoustic analysis device 44 Emotion Recognition Department 900 Information Processing Equipment 901 CPU 902 ROM 903 RAM 904 Program 905 Storage device 906 Recording Media 907 Drive unit 908 Communication Interface 909 Communication Network
Claims
1. a means for identifying a silent period in the input acoustic signal during which a reporter reporting the occurrence of an incident is not speaking; an identification means for identifying a sound source at the scene of the incident by analyzing the acoustic signal during the silent time; a prediction means for predicting an acoustic scene at the scene of the incident based on the identified sound source; Equipped with The identification means identifies the source of a sound associated with a receipt for the incident that occurred; The sound source identified by the identification means varies depending on the type of the incident. Acoustic analysis device.
2. The acoustic scene includes the situation, environment, and actions of people at the scene of the incident.
2. The acoustic analysis device according to claim 1.
3. The audio system further includes an output unit for outputting information indicating a situation, an environment, and a person's behavior corresponding to the identified acoustic scene.
3. The acoustic analysis device according to claim 1 or 2.
4. The audio signal processing device further includes a voice recognition unit for recognizing voices contained in the audio signal during a voice time excluding the non-voice time in the input audio signal.
3. The acoustic analysis device according to claim 1 or 2.
5. The prediction means predicts the acoustic scene at the scene of the incident based on the sound source identified from the acoustic signal during the silent time as well as the result of recognition by the speech recognition means of the sound included in the acoustic signal during the speech time.
5. The acoustic analysis device according to claim 4.
6. The audio signal processing device further includes emotion recognition means for recognizing emotions from the audio signal during the audio time excluding the non-audio time.
3. The acoustic analysis device according to claim 1 or 2.
7. The prediction means predicts the acoustic scene at the scene of the incident based on the emotion recognized by the emotion recognition means from the acoustic signal during the voice time in addition to the sound source identified from the acoustic signal during the non-voice time.
7. The acoustic analysis device according to claim 6.
8. A computer comprising: Identifying a non-voice period in which a reporter reporting an incident is not speaking in the input acoustic signal; identifying sound sources at the incident scene by analyzing acoustic signals during the non-speech time; An acoustic analysis method for predicting an acoustic scene at the scene of the incident based on the identified sound source, comprising: The computer identifies the source of the sound associated with the receipt of the incident; The sound source identified by the computer varies depending on the type of the case. Acoustic analysis methods.
9. Identifying a non-voice period in which a reporter reporting the occurrence of an incident is not speaking in the input acoustic signal; identifying sound sources at the incident scene by analyzing acoustic signals during the non-speech time; and predicting an acoustic scene at the scene of the incident based on the identified sound source, causing the computer to identify the source of a sound associated with an admission of the incident; The sound source that the computer is made to identify varies depending on the type of the case. program.
Citation Information
Patent Citations
Voice processing device, voice processing method, and program
JP2016042132A
Emergency notification listening support system and emergency notification listening support method
JP2020013234A
Information processing device and information processing method
JP2020066339A
Rescue dispatch support system
JP2021093228A
Terminal to provide user interface and method
US20120046942A1