Call processing device, call processing program, call processing method, and call processing system

The call processing device addresses emotional interference in emergency communications by using biometric sensors and AI to correct speech imperfections, ensuring clear and accurate information transmission.

JP2025151855APending Publication Date: 2025-10-09DENSO TEN LTD

Patent Information

Application Number
JP2024053463
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-28
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

In emergency situations like traffic accidents, speakers may become confused or excited, leading to inaccurate information transmission due to psychological shock or strong emotional stimulation, making it difficult for receivers to obtain accurate information.

Method used

A call processing device that processes the speaker's voice, detects their emotional state through biometric sensors and cameras, and uses generative AI to correct speech imperfections, generating a clear and accurate utterance sentence for transmission.

Benefits of technology

The device generates a corrected utterance sentence that is easily understood by the listener, ensuring accurate information conveyance by correcting speech disturbances caused by emotions, thereby improving communication clarity and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025151855000001_ABST
    Figure 2025151855000001_ABST
Patent Text Reader

Abstract

To correct an utterance disorder based on an emotion and a condition of a call originating party to perform correct information transmission to a call receiving party.SOLUTION: A call processing device that processes the voice of a call originating party acquires the voice of the call originating party, extracts utterance content from the acquired voice, detects an utterance situation of the call originating party at the time of utterance, and generates a corrected utterance sentence obtained by correcting the extracted utterance content on the basis of the detected utterance situation.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a call processing device, a call processing program, a call processing method, and a call processing system. [Background technology]

[0002] BACKGROUND ART Conventionally, an emergency call system is known that allows a caller and an operator at a call destination to directly exchange information while mutually checking a video screen in an emergency (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2007-116476 Summary of the Invention [Problem to be solved by the invention]

[0004] When encountering an emergency such as a traffic accident, the parties involved may experience psychological shock or strong stimulation of the sympathetic nervous system. Under such circumstances, when transmitting information to the call destination, the parties may become confused or excited, causing them to speak fluently or utter words that are inaccurate from the facts. In other words, the speaker (the party) may not be able to organize and convey the information, making it difficult to convey the information accurately. This presents a problem in that it is difficult for the receiver to receive accurate information. Furthermore, question and answer sessions between the speaker and receiver may become chaotic, and it may take a long time to obtain accurate information.

[0005] In view of the above-mentioned problems, an object of the present invention is to provide a technology that can correct speech imperfections based on the speaker's emotions and state, and can accurately convey information to the listener. [Means for solving the problem]

[0006] An exemplary call processing device of the present invention is a call processing device that processes the speaker's voice, acquires the speaker's voice, extracts the speech content from the acquired voice, detects the speaker's speech situation (e.g., emotion) at the time of speaking, and creates a corrected speech sentence by correcting the extracted speech content based on the detected speech situation. [Effects of the Invention]

[0007] According to the present invention, a corrected utterance sentence is generated that is appropriately processed according to the speaker's speech situation (e.g., emotion) at the time of utterance so that the speaker's voice can be easily understood, i.e., so that disturbances in the speaker's speech due to emotions such as impatience are corrected. Therefore, by notifying the listener of the corrected utterance sentence by voice or the like, it becomes possible to accurately convey information to the listener. [Brief explanation of the drawings]

[0008] [Figure 1] Overall configuration diagram of a call processing system according to the present embodiment [Figure 2] FIG. 2 is a block diagram showing the configuration of the call processing device and terminal device of FIG. 1. [Figure 3] FIG. 3 is an explanatory diagram showing an outline of a call processing method in the call processing device of FIG. 2; [Figure 4] FIG. 10 is a diagram showing an example of an utterance situation data table. [Figure 5] A diagram showing an example of a speech data table [Figure 6] 3 is a flowchart showing the call processing executed by the call processing device of FIG. 2. [Figure 7] FIG. 10 is an explanatory diagram showing an outline of a call processing method in a modified call processing device. [Figure 8] A block diagram showing the configuration of the call processing device of FIG. [Figure 9] A flowchart showing the call processing executed by the call processing device of FIG. 8. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the drawings. However, the present invention is not limited to the contents of the embodiments shown below.

[0010] <1. Call processing system> FIG. 1 is an overall configuration diagram of a call processing system 1 of this embodiment. The call processing system 1 is a system that transmits the voice of a speaker U1 to a receiver P1. In this embodiment, the call processing system 1 is assumed to be used when the speaker U1 encounters an emergency such as a traffic accident and transmits (reports) information about the situation to a receiver (operator) P1 who is the contact point for the police, fire department, insurance company, etc. The call processing system 1 includes a call processing device (receiver device) 10 configured as a server device or the like, and a terminal device (transmitter device) 20 owned by the speaker U1.

[0011] The call processing device 10 and the terminal device 20 are connected to each other so as to be able to communicate bidirectionally via a communication network N such as a mobile communication network. The terminal device 20 may be a drive recorder installed in a vehicle driven by the speaker U1, or may be a mobile terminal such as a smartphone or tablet terminal owned by the speaker U1.

[0012] Fig. 2 is a block diagram showing the configuration of the call processing device 10 and the terminal device 20 of Fig. 1. Fig. 2 shows components necessary for explaining the features of this embodiment, and omits the description of general components.

[0013] In this embodiment, the terminal device 20 will be described as a drive recorder installed in a vehicle. The drive recorder (terminal device 20) includes a microphone 21, a speaker 22, a camera 23, and a biometric sensor 24. The microphone 21 collects the voice of the speaker U1. The speaker 22 outputs the voice. The camera 23 captures an image of the speaker U1's appearance. The biometric sensor 24 is a sensor that detects biometric information such as the heartbeat and brain waves of the speaker U1, and detects, for example, the heartbeat of the speaker U1 in a non-contact manner (for example, a millimeter-wave vital (heartbeat) sensor, etc.).

[0014] <2. Call processing device> <2-1. Overview of call processing device> The call processing device 10 is a server device installed in a management center or the like that receives emergency reports. The server device may be a physical server or a virtual server such as a cloud server. The call processing device 10 processes the voice of a speaker U1 and transmits it to a receiver (operator) P1. The call processing device 10 includes a controller 11, an operation unit 12, a display unit 13, a microphone 14, a speaker 15, a communication unit 16, and a storage unit 17.

[0015] The controller 11 is configured with a processor that performs arithmetic processing and the like, and controls various operations in the call processing device 10. The processor is configured to include, for example, a CPU (Central Processing Unit). The controller 11 executes a call processing program stored in a call processing program storage unit 171 of the storage unit 17, which will be described later, and performs call processing on the voice of the speaker U1. The call processing program includes various programs that realize various functions of the call processing device 10.

[0016] The operation unit 12 is composed of an input device such as a keyboard that is operated by the user. The display unit 13 is composed of an output device such as a display. The display unit 13 is, for example, a liquid crystal display panel, and may be equipped with an operation unit 12 of a touch panel type or the like. The display unit 13 displays various information related to emergency notifications. Furthermore, the call processing device 10 is equipped with a microphone 14 through which the recipient (operator) P1 makes a call, and a speaker 15 through which the voice of the speaker U1, etc., can be heard.

[0017] The communication unit 16 is an interface for making calls and performing data communications with other devices (terminal devices 20) via the communication network N. The communication unit 16 includes a wireless communication device for performing wireless communications with other devices, and is configured, for example, by a transmission / reception device of a mobile telephone network for 5G communication (fifth generation mobile communication system).

[0018] The storage unit 17 includes a volatile memory and a nonvolatile memory, and stores various information necessary for call processing. The volatile memory is, for example, a random access memory (RAM). The nonvolatile memory is, for example, a read only memory (ROM), a flash memory, or a hard disk drive. The nonvolatile memory stores programs and data that can be read by the controller 11. At least some of the programs and data stored in the nonvolatile memory may be obtained from another computer device connected by wire or wirelessly, or from a portable recording medium.

[0019] The storage unit 17 includes a call processing program storage unit 171, an utterance correction program storage unit 172, an emotion estimation model storage unit 173, an utterance situation data table storage unit 174, a reading data table storage unit 175, and a prompt template storage unit 176. The contents of the programs, data tables, etc. stored in each of these storage units will be described separately. Furthermore, the storage unit 17 stores data tables (not shown) for various processes.

[0020] <2-2. Overview of call processing methods> Fig. 3 is an explanatory diagram showing an outline of a call processing method in the call processing device 10 of Fig. 2. The call processing method in the call processing device 10 shown in Fig. 3 is used to accurately convey the content of an utterance by a speaker U1 to a listener P1.

[0021] First, the speech processing device 10 acquires (receives) speech (speech information) from the speaker U1 via the terminal device 20 (not shown in FIG. 3). The speech processing device 10 extracts the speech content from the speech information acquired from the speaker U1. Specifically, the speech information is converted into text by speech analysis, the text data is converted into words using dictionary data or the like, and the words are converted into sentences by linguistic analysis.

[0022] Furthermore, the speech processing device 10 detects (estimates) the speaker state, such as the emotion of the speaker U1, at the time of speaking from the voice information acquired from the speaker U1. To detect (estimate) the speaker state, such as the emotion of the speaker U1, in addition to the voice information of the speaker U1, an image (appearance information) of the speaker U1 captured by a camera, biometric information of the speaker U1 detected by a biometric sensor, etc. may be used.

[0023] Next, the speech processing device 10 corrects the extracted utterance content (utterance correction) to create a corrected utterance sentence. For the utterance correction, an utterance correction program is used, but in this embodiment, a generative artificial intelligence (AI) configured by a large-scale language model or the like is used.

[0024] Such a generation AI can be generated by training with training data including, for example, a large amount of speech data (in the case of a traffic accident reporting system, traffic accident report text) from a wide variety of speakers in a wide variety of speech situations, preferably training data tagged with the speech situation (the speaker's emotions, biological state, etc.). It is also preferable to train with training data tagged with surrounding environments that (adversely) affect the speech, such as the accident situation (information on the urgency of the report, the scale of the accident (e.g., estimated from the acceleration at the time of the accident), road conditions, weather, etc.). It is also preferable to train with training data tagged with standardized speech data whose meaning can be accurately understood. A corrected speech generation method using a large-scale correction table for correcting ambiguous language (e.g., target data of precise language that is estimated to be ambiguous language) can also be applied.

[0025] Then, the speech processing device 10 reads out (outputs as voice) the created corrected utterance sentence to the listener P1 using speech synthesis processing etc. The speech processing device 10 will be further described in detail below.

[0026] <2-3. Details of the call processing device> 2, the controller 11 includes, as its functions, an acquisition unit 111, an utterance extraction unit 112, an utterance situation detection unit 113, an utterance correction unit 114, and a provision unit 115. In this embodiment, the functions of the controller 11 are realized by a processor executing arithmetic processing in accordance with a call processing program stored in the storage unit 17.

[0027] The acquisition unit 111 acquires (receives) the voice (voice information) of the speaker U1 collected by the microphone 21 of the drive recorder (terminal device 20) owned by the speaker U1 via the communication unit 16. The acquisition unit 111 also acquires (receives) via the communication unit 16 appearance information of the speaker U1 captured by the camera 23 of the drive recorder (terminal device 20) and biometric information of the speaker U1 detected by the biometric sensor 24.

[0028] That is, the acquisition unit 111 acquires (receives) speech information including the voice and biometric information of the speaker U1 from the terminal device 20 via the communication unit 16. It is also possible to acquire surrounding environment information that affects the speech, such as the location of the terminal device 20, road conditions at that location (congestion level and area information such as urban or suburban), and the scale of the accident (acceleration at the time of the accident).

[0029] The acquisition unit 111 stores the acquired voice information and various information as speech information in a data table formed in the storage unit 17 as necessary for subsequent processing. The acquisition unit 111 acquires the voice information and various information at substantially the same time and stores them in the data table as one data set. These data are then used to detect (estimate) the speech situation at that time, for example, the speaker state such as the emotion of the speaker U1.

[0030] The speech extraction unit 112 extracts the speech content by performing speech recognition processing on the speech information of the speaker U1 acquired by the acquisition unit 111. More specifically, the speech extraction unit 112 performs analysis processing on the speech information of the speaker U1, such as acoustic analysis, phoneme identification, and text sentence conversion. Note that, for text sentence conversion, it is preferable to appropriately express the length of long sounds and silent parts and the length of silent parts based on predetermined rules (such as inserting "-" or "blank" as appropriate). In this way, the speech extraction unit 112 extracts the speech content converted into text sentences from the speech (speech information) of the speaker U1.

[0031] The speech situation detection unit 113 detects (estimates) the speaker state (speech situation), such as the emotion of the speaker U1 when speaking, by performing analysis processing on the voice information, appearance information (image information), and biometric information of the speaker U1 acquired by the acquisition unit 111. The speaker state (speech situation) detected (estimated) by the speech situation detection unit 113 is information corresponding to the type of data input to the speech correction unit 114. Therefore, the speech situation detection unit 113 performs various processes according to the input specifications of the speech correction unit 114, for example, a process of detecting (estimating) and outputting an emotion type (impatience, calm, etc.) that the speech correction unit 114 can handle.

[0032] For emotion estimation, the heart rate and brain wave state are used as they are. For example, if the heart rate is below a threshold, it is estimated as a normal state, and if it is above the threshold, it is estimated as an abnormal state. In addition to emotion, the speech situation detection unit 113 also estimates the speaker's predetermined surrounding environment (surrounding environment that affects speech) in addition to emotion. For example, in the case of an accident report, the accident level (estimated from the acceleration at the time of the accident, etc.) and road conditions are considered to be surrounding environments (speech situations) that affect speech.

[0033] Regarding audio information including pronunciation of the voice of speaker U1, speech situation detection unit 113 operates as an emotional state estimation unit, and performs linguistic analysis of the speech content based on speaker U1's voice volume, frequency, speaking rate, speech frequency, and speech recognition. Based on the results of the linguistic analysis, such as the appearance of words or sentences related to joy, anger, anxiety, etc., or voice changes (tempo, range, etc.), speech situation detection unit 113 estimates speaker U1's emotion related to the speech of speaker U1.

[0034] The emotions of the speaker U1 can also be estimated based on signals from the camera 23 and biometric sensor 24. However, if the terminal device 20 does not include the camera 23 and biometric sensor 24, the speech situation detection unit 113 estimates (detects) the speaker state (speech situation), such as the emotions of the speaker U1, based solely on the voice information of the speaker U1 acquired by the acquisition unit 111. Meanwhile, by estimating the emotions of the speaker U1 using a combination of the voice information, appearance information, and biometric information of the speaker U1, it is possible to improve the accuracy of emotion estimation. This makes it easier to grasp the content of the speaker U1's utterance, as well as the degree of confusion and urgency. Therefore, when creating a corrected utterance sentence, as described below, it is possible to improve the accuracy of meeting the requirements of the speaker U1.

[0035] When estimating emotions (speech situation, biological state) based on appearance information including the facial expression and complexion of speaker U1, utterance situation detection unit 113 performs analysis processing such as feature calculation, shape discrimination, and complexion discrimination using the face of speaker U1 as the detection target from an image including the face of speaker U1. Based on the results of the analysis processing, such as the angle of the corners of the mouth, the angle of the eyebrows, the degree of eye opening, line of sight, and complexion, utterance situation detection unit 113 estimates the emotion of speaker U1 using a data table or the like showing the facial expression, complexion, and emotion of speaker U1.

[0036] When estimating emotions (utterance situation, biological state) based on the biometric information of speaker U1, utterance situation detection unit 113 performs an analysis process on, for example, the heartbeat, which is a biometric signal of speaker U1. More specifically, utterance situation detection unit 113 calculates an emotion index value of autonomic nervous system activity (hereinafter referred to as activity level), which is an index showing the emotional state of the speaker U1, from the heartbeat signal. The emotion index value of activity level can be calculated, for example, as the standard deviation of the heartbeat LF (Low Frequency) component (the low frequency component of the heartbeat waveform signal).

[0037] The speech situation detection unit 113 estimates whether the speaker U1 is sympathetic (strong emotion) or parasympathetic (weak emotion) based on the emotion index value of the activity level of the speaker U1. If the speaker U1 is sympathetic (strong emotion), it is estimated that the speaker U1 is in an emotional state such as "excitement, tension, anxiety, danger." If the speaker U1 is parasympathetic (weak emotion), it is estimated that the speaker U1 is in an emotional state such as "melancholy, calm."

[0038] Furthermore, the speech situation detection unit 113 may rank the state and degree of confusion of the speaker U1 based on the estimated emotion and state of the speaker U1. The ranking of the state and degree of confusion is performed by using a threshold for rank classification, etc., and replacing it with a numerical value or a word expressing the degree. In detail, the state and degree of confusion is ranked using, for example, a Likert scale. Furthermore, the speech situation detection unit 113 estimates the degree of influence on the speech of the speaker U1 based on the ranked state and degree of confusion of the speaker U1.

[0039] The speech situation detection unit 113 also operates as a surrounding environment estimation unit that estimates the surrounding environment of the speaker U1. Specifically, it estimates circumstances that affect the speech situation of the speaker U1, such as the situation of the accident, etc. (acceleration at the time of the accident, injuries to vehicle occupants, etc., and damage to the vehicle, etc., based on images from the camera 23, etc.), road conditions, etc. The surrounding environment (speech situation) is a major factor in the urgency of information notification of the accident, etc., and the urgency has a large impact on the speech of the speaker U1, so in the processing of this example, it is treated as urgency information. The urgency is calculated based on information such as the situation of the accident, etc., road conditions, etc., using these information as parameters and a data table in which an urgency value is set for each parameter value, for example.

[0040] 4 is a diagram showing an example of an utterance situation data table. As shown in FIG. 4, the utterance situation data table includes items such as "utterance situation data ID," "emotional state type," "urgency," "utterance correction instruction information," "utterance content," and "corrected utterance sentence."

[0041] When processing in real time (when there is no need to store information), it is sufficient to store only the most recent data to be processed in the speech situation data table, and the speech situation data table acts as a buffer that temporarily stores the data to be processed during processing.

[0042] The "utterance situation data ID" is identification information for identifying a data set of utterance situation data. The utterance situation data ID data is also the primary key of the data record in the utterance situation data table. In other words, in the utterance situation data table, a data record is configured for each utterance situation data ID, and data for each item linked to the utterance situation data ID is stored in that data record.

[0043] "Emotional state type" and "urgency" are data items corresponding to the emotion and surrounding environment, which are the speech situation of speaker U1, estimated by the speech situation detection unit 113. When generating corrected utterance sentence data, the speech situation detection unit 113 extracts the emotional state type and urgency corresponding to the speech content data to be processed.

[0044] "Utterance correction instruction information" is a data item that indicates the correction content to be made to the utterance content data by the utterance correction unit 114. The utterance correction instruction information data is data that is determined based on the emotional state type data and the urgency data.

[0045] "Utterance content" is a data item corresponding to the utterance content extracted by the utterance extraction unit 112 to be processed. Also, "corrected utterance sentence" is a data item corresponding to the corrected utterance sentence that is the result of the correction process performed by the utterance correction unit 114 on the utterance content data.

[0046] The speech correction unit 114 creates a corrected utterance sentence by correcting the speech content of the speaker U1 extracted by the speech extraction unit 112 based on the speech situation (emotions and surrounding environment) of the speaker U1 estimated by the speech situation detection unit 113. The speech correction unit 114 converts the corrected utterance sentence into audio data and stores it in the "corrected utterance sentence" field of the speech situation data table. The function of the speech correction unit 114 is realized by executing an utterance correction program stored in the speech correction program storage unit 172 of the storage unit 17. In this embodiment, a generation AI is used as the speech correction program, but the program is not limited to a generation AI.

[0047] When the speech correction unit 114 is configured to use a generation AI, it will have the following configuration: The generation AI model may be installed within the call processing device 10 or on an external server (connected via the communication unit 16), but since this increases the system scale, it is preferable to use an AI model installed on an external server.

[0048] The speech correction unit 114 creates a speech correction prompt to be sent to the generation AI. The prompt template storage unit 176 of the storage unit 17 stores templates of speech correction prompts. The template is composed of a standard instruction phrase for speech correction and parameter items to be inserted into the standard phrase, and is composed, for example, of "(standard phrase) + (parameter: speech content data) + (parameter: emotion data) + (parameter: risk level (surrounding environment) data)".

[0049] The utterance correction unit 114 sets each parameter value (values ​​in the data table in FIG. 4) corresponding to the template and generates an utterance correction prompt. Specifically, the utterance correction unit 114 creates a prompt such as, "Please generate a corrected utterance sentence that clarifies the following utterance content according to the speaker's emotion type and danger level. The utterance content is [There has been an accident... The location is about 3 on the Osaka side of Kyoto Higashi Inn on the Shin Expressway]. The speaker's emotion is [impatience]. The danger level is [level 4 of maximum danger level 5]."

[0050] The utterance correction unit 114 then applies (sends) the prompt created in this way to the generation AI, and acquires (receives) the corrected utterance created by the generation AI, for example, a corrected utterance clarified as "An accident has occurred. The location is about 3 km on the Osaka side of the Kyoto East Interchange on the Meishin Expressway." The utterance correction unit 114 then converts the corrected utterance created in this way into audio data and stores it in the "corrected utterance" field of the utterance situation data table.

[0051] If a generation AI is not used, a method of creating corrected utterances using a large-scale correction table (target data of accurate language that is estimated to be ambiguous language, etc.) that corrects ambiguous language as described above can also be applied. However, because the correction table is large and difficult to generate, and it is difficult to handle data that is not in the correction table, it is preferable to configure the utterance correction unit 114 with a generation AI.

[0052] The providing unit 115 outputs (reads aloud) the corrected utterance sentence created by the utterance correction unit 114 to the listener P1. In detail, the providing unit 115 performs speech synthesis on the corrected utterance sentence, adjusts the speech expression based on the reading data table shown in FIG. 5, and reads it aloud. The read-out corrected utterance sentence is output aloud to the listener P1 via the speaker 15. The providing unit 115 may also display the corrected utterance sentence created by the utterance correction unit 114 on the display unit 13.

[0053] 5 is a diagram showing an example of a reading data table. As shown in Fig. 5, the reading data table includes items such as "reading data ID," "emotional state type," "urgency," "volume," "intonation," and "repetition."

[0054] The "reading data ID" is identification information for identifying a data set of read-aloud data. The read-aloud data ID data is also the primary key of the data record in the read-aloud data table. In other words, in the read-aloud data table, a data record is configured for each read-aloud data ID, and data for each item linked to the read-aloud data ID is stored in that data record.

[0055] The reading data ID data is linked to the utterance situation data ID data in the utterance situation data table shown in Figure 4, and for data for the same utterance, the reading data ID data and the utterance situation data ID are stored in the same data record.

[0056] In addition, when processing in real time (when there is no need to save information), just like the speech situation data table, it is sufficient to store only the most recent data to be processed in the reading data table, and the reading data table acts like a buffer that temporarily stores the data to be processed during processing.

[0057] "Emotional state type" and "urgency" are data items corresponding to the emotion and surrounding environment, which are the speech situation of speaker U1, estimated by speech situation detection unit 113. The "emotional state type" data and "urgency" data are recorded (set) when data is generated in the speech situation data table shown in Fig. 4. The "emotional state type" data and "urgency" data may be extracted from a data record in which the speech situation data ID data in the speech situation data table shown in Fig. 4 is the same as the read-aloud data ID data.

[0058] When outputting the corrected utterance sentence aloud to the listener P1, the providing unit 115 extracts the emotional state type data and urgency data corresponding to the utterance content data to be processed. Then, the providing unit 115 searches a reading method data table (not shown) that stores reading method data (in this example, data on volume, intonation, and repetition) that has been separately generated in advance and whose parameter values ​​are the emotional state type data and urgency data, using the extracted emotional state type data and urgency data, and extracts the corresponding reading method data. The providing unit 115 then stores the extracted reading method data in a data record corresponding to the reading data table.

[0059] The reading method data, "volume," "intonation," and "repetition," represent the volume during reading, the intonation of the corrected utterance sentence, and the repetition of words or sentences. The providing unit 115 then outputs the corrected utterance sentence aloud from the speaker 15 using the volume, intonation, and number of repetitions based on the volume data, intonation data, and repetition data extracted and stored in this manner.

[0060] According to the above configuration, a corrected utterance sentence is generated that is appropriately processed to make the speech of the speaker U1 easier to understand in accordance with the speech situation (e.g., emotion) of the speaker U1 at the time of speaking, that is, to correct for speech impatience or other disturbances in the speaker U1's speech due to such emotions as impatience. Therefore, by notifying the listener of the corrected utterance sentence by voice or the like, it becomes possible to accurately convey information to the listener P1.

[0061] Furthermore, by estimating the emotion and state of speaker U1, the degree of confusion and urgency of speaker U1 can be grasped. Then, the content of speaker U1's utterance can be corrected based on the degree of confusion of speaker U1, and a corrected utterance can be created that takes the urgency into account. In this way, it is possible to correct the speech disturbance based on speaker U1's emotion and state, and to accurately convey information to listener P1.

[0062] Although clear information can be transmitted by simply synthesizing the corrected utterance and outputting it as voice, in this embodiment, the voice is output using a reading method that corresponds to the speaker's emotions and the level of danger (surrounding environment) when speaking. This allows for voice output that expresses the realism of the accident scene, etc., making it possible to transmit information more effectively.

[0063] <2-4. Example of operation of call processing device> Fig. 6 is a flowchart showing the call processing executed by the call processing device 10 of Fig. 2. The operation according to this flowchart is realized by computer programs (call processing program and speech correction program) executed by the controller 11 (the computer constituting the controller 11).

[0064] A computer program that causes a computer device to implement the call processing method according to this embodiment is installed in a computer device such as the call processing device 10 to implement the various functions described above. Such a computer program is provided to the computer device via a computer-readable non-volatile recording medium. For example, an optical disk or the like on which the computer program is recorded may be distributed or sold, or a computer program stored on a hard disk or the like of a server device may be distributed or sold via an Internet environment. The computer program that causes a computer device to implement the call processing method according to this embodiment may consist of only one program, or may consist of multiple programs.

[0065] The process shown in FIG. 6 starts when the call processing device 10 is running, the receiver P1 is waiting for a call, and receives a report from the speaker U1 and receives the voice of the speaker U1 (by detecting voice reception).

[0066] In step S101, the controller 11 (acquisition unit 111) acquires the voice (voice information) of the speaker U1, and proceeds to step S102. In detail, the controller 11 (acquisition unit 111) acquires the voice (voice information) of the speaker U1 collected by the microphone 21 of the drive recorder (terminal device 20) owned by the speaker U1 via the communication unit 16. The acquired voice information is stored in a data table in the storage unit 17 as necessary.

[0067] In step S102, the controller 11 (acquisition unit 111) acquires appearance information and biometric information of the speaker U1, and proceeds to step S103. In detail, the controller 11 (acquisition unit 111) acquires, via the communication unit 16, appearance information of the speaker U1 captured by the camera 23 of the drive recorder (terminal device 20) owned by the speaker U1, and biometric information of the speaker U1 detected by the biometric sensor 24. The acquired appearance information and biometric information are stored in a data table in the storage unit 17 as necessary.

[0068] In step S103, the controller 11 (utterance extraction unit 112) extracts the content of the utterance from the voice information of the speaker U1, and proceeds to step S104. Specifically, the utterance content data is generated by converting the voice information into text through voice analysis, converting the text data into words using dictionary data or the like, and converting the words into text sentences through linguistic analysis, etc.

[0069] In step S104, the controller 11 (utterance situation detection unit 113) detects (estimates) the emotion and surrounding environment (level of danger) of the speaker U1 based on the voice information, appearance information (image information), and biometric information of the speaker U1, and proceeds to step S105. In detail, the controller 11 (utterance situation detection unit 113) refers to the utterance situation data table and acquires utterance correction instruction information for the utterance correction unit 114 based on the estimated emotion and surrounding environment (level of danger) of the speaker U1. When a generation AI is used, a prompt corresponding to the utterance correction instruction information is generated by the method described above.

[0070] In step S105, the controller 11 (utterance correction unit 114) creates a corrected utterance sentence by correcting the utterance content (data) generated in step S103 based on the above-mentioned utterance correction instruction information, and proceeds to step S106. In detail, the controller 11 (utterance correction unit 114) refers to the utterance situation data table, corrects the utterance content of the speaker U1 using utterance correction instruction information corresponding to the emotion of the speaker U1 and the surrounding environment (level of danger), and creates a corrected utterance sentence. When a generation AI is used, the generated prompt is set (sent) to the generation AI as a question sentence, and the answer sentence from the generation AI to the question sentence is set as the corrected utterance sentence.

[0071] In step S106, the controller 11 (providing unit 115) performs speech synthesis on the corrected utterance sentence, and the process proceeds to step S107.

[0072] In step S107, the controller 11 (providing unit 115) ends the reading of the voice-synthesized corrected utterance sentence and the processing related to Fig. 6. In detail, the controller 11 (providing unit 115) adjusts the voice expression (volume, intonation, etc.) based on the reading data table, and outputs the corrected utterance sentence aloud from the speaker 15.

[0073] <3. Modifications> Next, a modified example of the call processing device 10 will be described. The configuration of the modified example of the call processing device 10 (see FIG. 8) and the call processing method (see FIG. 7) are basically similar to the configuration of the call processing device 10 and the call processing method previously explained using FIGS. 2 and 3. Therefore, in the following, explanations of the configuration and processing steps common to them will be omitted.

[0074] <3-1. Modified call processing method> 7 is an explanatory diagram showing an outline of a call processing method in a modified example of the call processing device 10. The call processing device 10 organizes information collected from the speaker U1 based on the utterance content extracted from the voice information of the speaker U1 and the corrected utterance sentence of the utterance content.

[0075] To organize the information, the call processing device 10 generates missing information that is missing in the speech information of the speaker U1 (expressed mathematically as [corrected utterance sentence-utterance information], which is information that would have been included in the speech information if the speaker U1 had been in a calm state), and contradictory information that is contradictory (information that is contradictory in [corrected utterance sentence-utterance information]), based on the speech content extracted from the speech of the speaker U1 and the corrected corrected utterance sentence, and outputs the missing information and the contradictory information to the listener P1 as text information (displayed on the display unit 13).

[0076] Furthermore, the call processing device 10 uses the utterance information, corrected utterance sentence, missing information, contradictory information, and data on the utterance situation (emotions, etc.) generated as a result of the information organization to create data for generating a response sentence for the speaker U1. Specifically, a prompt for causing the generation AI to generate a response sentence is created using the utterance information, corrected utterance sentence, missing information, contradictory information, and data on the utterance situation.

[0077] Then, the call processing device 10 sets the created prompt to the generation AI (sends it as a question sentence), and uses the answer sentence from the generation AI as a response sentence to the speaker U1 (creates a response sentence based on the answer sentence from the generation AI). Note that a method of generating a response sentence that does not use a generation AI, such as a method of generating a response sentence using a template that generates a response sentence by inserting data on utterance information, corrected utterance sentences, missing information, contradictory information, and utterance situation, can also be applied.

[0078] The speech processing device 10 then performs speech synthesis on the created response sentence to convert it into speech information, and outputs (transmits) it to the terminal device 20 of the speaker U1. In the terminal device 20, speech based on the received speech information is output from the speaker 22. The speech processing device 10 of the modified example will now be described in detail.

[0079] <3-2. Modified speech processing device> Fig. 8 is a block diagram showing the configuration of the call processing device 10 of Fig. 7. Fig. 8 shows components necessary for explaining the features of this embodiment, and omits the description of general components. Also, Fig. 8 omits the description of the terminal device 20.

[0080] The controller 11 includes, as its functions, an acquisition unit 111, an utterance extraction unit 112, an utterance situation detection unit 113, an utterance correction unit 114, an information organization unit 116, a response generation unit 117, and a provision unit 115. In this embodiment, the functions of the controller 11 are realized by a processor executing arithmetic processing in accordance with a call processing program stored in the storage unit 17.

[0081] The information organizing unit 116 generates utterance information based on the utterance content of the speaker U1. More specifically, the information organizing unit 116 organizes the information collected from the speaker U1 based on the utterance content extracted from the audio information of the speaker U1 by the utterance extraction unit 112, and generates utterance information that constitutes the utterance content. The utterance information includes words that constitute the utterance content. The information organizing unit 116 also organizes information on the corrected utterance sentence corrected and generated by the utterance correction unit 114 in the same way, and generates corrected utterance sentence information that constitutes the corrected utterance sentence.

[0082] Furthermore, the information organizing unit 116 compares the above-mentioned utterance information with necessary information required for the utterance information to be valid (for example, information required for accident occurrence notification). The necessary information is the so-called "5W1H" used for information collection and analysis, and corresponds to "When," "Where," "Who," "What," "Why," and "How." The user of the call processing device 10 (for example, the administrator of the accident notification system) sets an appropriate information type for the necessary information.

[0083] The information organizing unit 116 judges whether the utterance information is sufficient, for example, whether necessary information is included, whether the information type "5W1H" is established, etc. If there is missing information when comparing the utterance information with the necessary information, the information organizing unit 116 generates missing information based on the difference between them (for example, if the occurring incident (accident, etc.), which is the necessary information, is unknown, missing information such as "occurrence matter") and stores it in a data table formed in the storage unit 17. That is, the information organizing unit 116 generates missing information that is missing in the information notification regarding the utterance based on the content of the utterance of the speaker U1, and stores (adds) the generated missing information in the storage unit 17.

[0084] Furthermore, when there is contradictory information when comparing the above-mentioned utterance information with the corrected utterance sentence information, the information organizing unit 116 generates contradictory information that suggests the contradictory information. For example, when the utterance information collected from the speaker U1 includes information corresponding to multiple different "Where," the information organizing unit 116 regards the different information as contradictory information. For example, if the location extracted from the utterance (utterance information) of the speaker U1 is "Tomei Expressway" and the location extracted from the corrected utterance sentence is "Meishin Expressway," the contradictory information is information such as the contradictory data type "location" and the contradiction point "Tomei Expressway (utterance information)" and "Meishin Expressway (corrected utterance sentence)." The information organizing unit 116 stores (adds) the generated contradictory information to a data table formed in the storage unit 17.

[0085] The utterance information, corrected utterance sentence information, missing information, and contradictory information of the speaker U1 are output as text information by the providing unit 115 to the display unit 13 and provided to the listener P1. For example, text such as "The incident is unknown. It is unclear whether the location is the [Tomei Expressway (utterance information)] or the [Meishin Expressway (corrected utterance)]" is displayed on the display unit 13. This configuration allows the listener P1 to improve the accuracy of the information by capturing various pieces of information as text. This makes it easier for the listener P1 to understand the requirements of the speaker U1 regarding the content of the utterance made by the speaker U1.

[0086] Furthermore, the speech uttered by the speaker U1 may be output to the listener P1 as the information to be provided to the listener P1. With this configuration, the listener P1 can check the speech uttered by the speaker U1 and compare it with the corrected utterance created by the utterance correction unit 114. This makes it easier to grasp and understand the missing information and contradictory information.

[0087] The response generation unit 117 generates a response to the speaker U1 based on the utterance information, corrected utterance sentence information, missing information, and contradictory information generated as a result of information organization by the information organization unit 116. The response sentence is, for example, a sentence such as "What happened? Did it happen on the Tomei Expressway? Or the Meishin Expressway? (Extracted from the analysis result (corrected utterance sentence information) by the AI)." In detail, the response generation unit 117 generates a response such as a repetition to confirm the utterance content, providing missing information, reconfirming contradictory information, showing consideration for the speaker U1, or guiding the speaker U1 to calm down his / her confusion. The response generation unit 117 stores the generated response sequence in a data table formed in the storage unit 17.

[0088] The response generation unit 117 can be realized by a configuration in which a prompt is created using a template based on the utterance information, corrected utterance sentence information, missing information, and contradiction information, and the prompt is set in the generation AI. Also, various templates for creating a response sentence are prepared (stored in the storage unit 17), and an appropriate template is selected based on the content type of the utterance information, corrected utterance sentence information, missing information, and contradiction information. Then, a response sentence can be created by applying each piece of information to the selected template.

[0089] The response sentence is subjected to speech synthesis by the providing unit 115 and converted into speech information, which is provided (output) to the terminal device 20 of the speaker U1.

[0090] <3-3. Operational example of the speech processing device of the modified example> Figure 9 is a flowchart showing the call processing executed by the call processing device 10 of Figure 8. The operation according to this flowchart is realized by computer programs (call processing program and speech correction program) executed by the controller 11 (the computer constituting the controller 11). Like the processing of Figure 6, this processing is started when the call processing device 10 is running, the receiver P1 is waiting for a call, and the voice of the speaker U1 is received upon receiving a report from the speaker U1 (by detecting voice reception).

[0091] In step S201, the same processes as those in steps S101 to S107 in the operational flow shown in FIG. 6 of the embodiment described above are performed, and then the process proceeds to step S202.

[0092] In step S202, the controller 11 (information organizing unit 116) organizes the utterance content extracted by the utterance extraction unit 112 and the corrected utterance sentence to generate utterance information and corrected utterance sentence information, and then proceeds to step S203.

[0093] In step S203, the controller 11 (response generation unit 117) generates missing information and contradictory information based on the utterance information and corrected utterance sentence information, and then proceeds to step S204.

[0094] In step S204, the controller 11 (response generation unit 117) creates a response sentence based on the missing information, contradictory information, utterance information, and corrected utterance sentence information, and then proceeds to step S205.

[0095] In step S205, the controller 11 (providing unit 115) performs speech synthesis on the response sentence to convert it into voice information, and then the process proceeds to step S206.

[0096] In step S206, the controller 11 (providing unit 115) provides (outputs) the voice information of the response sentence to the terminal device 20 of the speaker U1, and the process proceeds to step S207.

[0097] In step S207, the controller 11 (providing unit 115) generates display data based on the missing information and contradictory information, as well as the utterance information and corrected utterance sentence information, outputs the display data to the display unit 13, and terminates the processing related to Figure 9.

[0098] According to the above configuration, by organizing the information on the content of the utterance of speaker U1, it becomes easier to grasp missing information regarding the content of the utterance or unclear information (information that contradicts the analysis results (meaning estimation of the utterance content) using AI, etc. and can be said to be unclear). In addition, for matters that require confirmation from speaker U1, an appropriate response sentence is created for speaker U1, and questions are automatically asked to speaker U1. Therefore, information is supplemented and corrected, and as a result, the accuracy of the information is improved. Furthermore, by inserting words that show consideration for speaker U1 and calm speaker U1's confusion, it becomes possible to alleviate speaker U1's mental stress.

[0099] In the above embodiment, the controller 11 of the call processing device 10 realizes the functions of the acquisition unit 111, the utterance extraction unit 112, the utterance situation detection unit 113, the utterance correction unit 114, the information organization unit 116, the response organization unit, and the provision unit 115, and the data required for these operations is stored in the storage unit 17 (the call processing program storage unit 171, the utterance correction program storage unit 172, the emotion estimation model storage unit 173, the utterance situation data table storage unit 174, the reading data table storage unit 175, and the prompt template storage unit 176). However, this data can also be stored in the terminal device 20 connected to the call processing device 10 for communication, and the controller of the terminal device 20 can perform the above processing. Also, a configuration in which each process is shared between the call processing device 10 and the terminal device 20 as appropriate is possible.

[0100] Furthermore, the call processing system 1 of this embodiment may be used in emergency calls, etc., as well as in call centers, etc., to handle conversations in situations where the speaker is confused or excited. By having the receiver simplify the content of the speaker's speech received by the speaker and repeat it back to the speaker, the speaker in the call center, etc., can understand whether the requirements were properly conveyed to the receiver. This may alleviate the speaker's confusion or excitement.

[0101] Furthermore, when using the call processing system 1 to respond to complaints, it can be installed and used not only in call centers and other places that handle voice calls, but also in stores such as convenience stores. It can also be used as an information transmission tool to support conversations for people who are not accustomed to speaking or who are not good at conversation.

[0102] <5. Things to keep in mind> Various technical features disclosed as embodiments in this specification may be modified in various ways without departing from the spirit of the technical creation. In other words, the above-described embodiments are illustrative in all respects and are not limiting. The technical scope of the present invention is defined by the claims, not by the description of the above-described embodiments, and includes all modifications that fall within the meaning and scope of the claims. Furthermore, the multiple embodiments described in this specification may be combined as appropriate to the extent possible.

[0103] In the above embodiment, various functions are realized by software through the arithmetic processing of a CPU in accordance with a program, but at least some of these functions may be realized by electrical hardware resources. All or part of the hardware resources may be realized by, for example, an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). Conversely, at least some of the functions realized by hardware resources may be realized by software.

[0104] It may also include a computer program that causes a processor (computer) to realize at least some of the functions of the call processing system 1 (call processing device 10, terminal device 20). Such a computer program can be stored in a computer-readable nonvolatile recording medium (for example, the above-mentioned nonvolatile memory, as well as an optical recording medium (for example, an optical disk), a magneto-optical recording medium (for example, a magneto-optical disk), a USB memory, or an SD card) and provided (sold, etc.), or can be provided from a server device via a communication line such as the Internet, i.e., by downloading. [Explanation of symbols]

[0105] 1. Call Processing System 10. Call processing device (receiver) 11 Controller 12 Control section 13 Display section 14. Mike 15 Speaker 16 Communications Department 17 Memory section 20 Terminal equipment (transmitting equipment) 21. Mike 22 Speaker 23 Camera 24 Biometric Sensors 111 Acquisition Department 112 Speech Extraction Unit 113 Speech situation detection unit 114 Speech Correction Unit 115 Provision Department 116 Information Organizing Department 117 Response Generation Unit 171 Call processing program storage unit 172 Speech correction program memory unit 173 Emotion estimation model memory section 174 Speech situation data table storage unit 175 Reading data table storage unit 176 Prompt template memory section P1 Listener U1 Speaker

Claims

1. A call processing device that processes a speaker's voice, acquiring the speech of the speaker; Extracting speech content from the acquired voice; Detecting a speech situation of the speaker when speaking; creating a corrected utterance sentence by correcting the extracted utterance content based on the detected utterance situation; Call processing equipment.

2. The speech situation is a biological state of the speaker. The call processing device of claim 1 .

3. The biological state of the speaker is an emotion of the speaker. The call processing device according to claim 2 .

4. acquiring appearance information of the speaker, and estimating the emotion of the speaker based on the appearance information; The call processing device according to claim 3 .

5. The speech situation is the surrounding environment of the speaker. The speech processing device according to any one of claims 1 to 4.

6. generating missing information that is missing in the information notification of the utterance based on the content of the utterance of the speaker; Adding the generated missing information; The speech processing device according to any one of claims 1 to 4.

7. converting the corrected utterance sentence into voice data; The speech processing device according to any one of claims 1 to 4.

8. A call processing program for processing a speaker's voice, acquiring the speech of the speaker; Extracting speech content from the acquired voice; Detecting a speech situation of the speaker when speaking; causing a computer to perform a process of correcting the extracted utterance content based on the detected utterance situation to create a corrected utterance sentence; Call processing program.

9. A call processing method for processing a speaker's voice, comprising: acquiring the speech of the speaker; Extracting speech content from the acquired voice; Detecting a speech situation of the speaker when speaking; a computer performs a process of correcting the extracted utterance content based on the detected utterance situation to create a corrected utterance sentence; Call handling methods.

10. A call processing system that transmits a speaker's voice to a receiver, a transmitter and a receiver, The transmitter device acquiring the speech of the speaker; Transmitting speech information including the acquired voice to the receiver device; The receiving device receiving the speech information of the speaker from the transmitting device; Extracting the speech content from the received speech information; Detecting the speech state of the speaker at the time of speaking based on the received speech information; creating a corrected utterance sentence by correcting the extracted utterance content based on the detected utterance situation; outputting the corrected utterance sentence by voice; Call processing system.

11. A call processing system that transmits a speaker's voice to a receiver, A transmitter and a receiver are included, The transmitter device acquiring the speech of the speaker; Extracting speech content from the acquired voice; Detecting a speech situation of the speaker when speaking; creating a corrected utterance sentence by correcting the extracted utterance content based on the detected utterance situation; transmitting the created corrected utterance sentence to the receiver; The receiving device receiving the corrected utterance sentence from the transmitter; outputting the received corrected utterance sentence by voice; Call processing system.

Citation Information

Patent Citations

  • Emergency call system

    JP2007116476A

Cited By

  • Information processing systems, information processing methods, and programs

    JP7845734B1

  • Information processing system, information processing method, and program

    JP7845735B1