Information processing device and information processing method
The information processing device addresses unnatural speech rate adjustments by detecting topic changes and adjusting the system's speech rate to match user speech patterns, enhancing dialogue comfort.
Patent Information
- Application Number
- JP2024053702
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-28
- Publication Date
- 2025-10-09
AI Technical Summary
Existing systems that adjust speech rates to match user speech rates can result in unnatural or uncomfortable responses, particularly when the user stutters, leading to discomfort.
An information processing device that includes a topic change detection unit to identify shifts in conversation topics and adjusts the system's speech rate before and after the topic change, using user speech rate averages and anxiety/filler frequency to set appropriate speech rates.
Prevents user discomfort by ensuring the system's speech rate adapts naturally to topic changes, maintaining a smooth and natural dialogue experience.
Smart Images

Figure 2025152012000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device and an information processing method. [Background technology]
[0002] There is a system that conducts a dialogue using user utterances, which are utterances made by a user, and system utterances, which are utterances made by the system's voice. Also, a technique is known that controls the speaking rate of a response voice so that it corresponds to the speaking rate of the speaker (see, for example, Patent Document 1 below). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2008-26463 Summary of the Invention [Problem to be solved by the invention]
[0004] If the speech rate of the system utterances is set to the same rate as the user's speech rate, the system utterances will sound more robotic, resulting in unnatural responses. Also, if the speech rate of the system utterances is set according to the user's speech rate, the speed of the system utterances will change unnaturally if the user stutters, causing discomfort to the user.
[0005] Therefore, an object of the present disclosure is to prevent the user from feeling uncomfortable by having the system speak at a suitable speech rate. [Means for solving the problem]
[0006] In order to solve the above problem, an information processing device according to one aspect of the present disclosure includes a topic change detection unit that detects a change in topic in a dialogue based on at least a user utterance in a dialogue between a user utterance, which is speech spoken by a user, and a system utterance, which is speech generated by the system, and a speech rate setting unit that sets a system speech rate, which is the speed of system speech after a topic change, to a speech rate that is different from the system speech rate before the topic change.
[0007] According to the above aspect, when a change in topic is detected in a dialogue between a user and a system, different system speech rates are set before and after the topic change. By setting the system speech rate to an appropriate speech rate when the topic change occurs, it is possible to prevent the user from feeling uncomfortable with the system speech. [Effects of the Invention]
[0008] According to the present disclosure, by having the system speak at a suitable speech rate, it is possible to prevent the user from feeling uncomfortable. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 2 is a block diagram showing the functional configuration of the information processing apparatus according to the present embodiment. [Figure 2] FIG. 10 is a diagram illustrating a process of calculating a user's speaking rate. [Figure 3] FIG. 10 is a diagram showing a predetermined number (N) of recent user utterances and their speaking rates stored in a user profile storage unit. [Figure 4] 10 is a flowchart illustrating an example of processing content of an information processing method in the information processing device of the dialogue system. [Figure 5] FIG. 2 is a diagram showing a configuration of an information processing program. [Figure 6] FIG. 2 is a hardware block diagram of the information processing device. DETAILED DESCRIPTION OF THE INVENTION
[0010] An information processing device according to an embodiment of the present invention will be described with reference to the drawings. Whenever possible, the same components are designated by the same reference numerals, and redundant description will be omitted.
[0011] 1 is a block diagram showing the configuration of a dialogue system including an information processing device according to this embodiment and the functional configuration of the information processing device. The dialogue system 1 is a system that conducts a dialogue between a user utterance, which is an utterance made by a user, and a system utterance, which is an utterance made by voice from the system. The dialogue system 1 is configured to include an information processing device 10, as an example.
[0012] 1, the information processing device 10 functionally comprises a user utterance acquisition unit 11, a system utterance output unit 12, a topic change detection unit 13, a speech rate calculation unit 14, an anxiety level measurement unit 15, a filler frequency acquisition unit 16, and an utterance rate setting unit 17. In this embodiment, the functional units 11 to 17 may be configured in one device as exemplified in FIG. 1, or may be distributed across multiple devices.
[0013] Each functional unit of the information processing device 10 is configured to be able to access storage means such as a dialogue information storage unit 21 and a user profile storage unit 22. The dialogue information storage unit 21 is a storage unit (storage) that stores dialogue information that is referenced or used to realize a dialogue with a user by outputting a system utterance in response to a user utterance. The dialogue information may be information consisting of a dialogue scenario. The dialogue information may also be a language model for generating an appropriate system utterance in response to a user utterance. The user profile storage unit 22 is a storage unit (storage) that stores at least a predetermined number of recent user utterances. The dialogue information storage unit 21 and the user profile storage unit 22 may be configured in the information processing device 10, as exemplified in FIG. 1, or may be configured in a separate device configured to be accessible from the information processing device 10.
[0014] Next, each functional unit of the information processing device 10 will be described. The user utterance acquisition unit 11 acquires user utterances, which are utterances made by a user. When the information processing device 10 constitutes a robot that interacts with a user, the user utterance acquisition unit 11 may acquire the user utterances via an input device such as a microphone provided in the information processing device 10. Furthermore, when the information processing device 10 is constituted by a server computer that communicates with a user terminal, the user utterance acquisition unit 11 may acquire user utterances transmitted from the user terminal via a network.
[0015] The system utterance output unit 12 outputs system utterances, which are speech utterances by a dialogue system that dialogues with a user. The system utterance output unit 12 of this embodiment outputs the system utterances at the speech rate set by the speech rate setting unit 17. The system utterance output unit 12 generates suitable system utterances in response to the user utterances by referring to a dialogue scenario, which is dialogue information stored in the dialogue information storage unit 21, or by using a language model, and outputs the generated system utterances.
[0016] When the information processing device 10 constitutes a robot that interacts with a user, the system utterance output unit 12 may output the system utterance via an output device such as a speaker provided in the information processing device 10. When the information processing device 10 is constituted by a server computer that communicates with a user's terminal, the user utterance acquisition unit 11 may transmit the system utterance to the user's terminal via a network.
[0017] The topic change detection unit 13 detects a change in topic (a change in topic) in a dialogue between a user utterance and a system utterance, based on at least a user utterance. Specifically, for example, when a dialogue is being carried out based on a dialogue scenario, the topic change detection unit 13 may detect the time when a response to a question from the user is completed as the time of the topic change. Also, when a dialogue is being carried out to support a user in some kind of work, the topic change detection unit 13 may detect the time when the work is completed through a dialogue based on a scenario or the like as the time of the topic change.
[0018] In addition, when a conversation such as a casual chat is being conducted with a user, the topic change detection unit 13 may vectorize (embedding) the words in the conversation using a technique such as word2vec, and detect a change in topic when the similarity between the vectors of a predetermined number of words before and after a certain point in time becomes less than a predetermined level.
[0019] The speech rate calculation unit 14 calculates a reference speech rate that serves as a reference for the system speech rate based on the user speech rate, which is the speed at which the user speaks. The process of acquiring the user speech rate will be described with reference to FIG. 2. The user speech rate may be expressed, for example, by the number of moras per unit time. However, the user speech rate is not limited to being expressed by the number of moras.
[0020] 2, the speech rate calculation unit 14 acquires a unit of user utterance speech via the user utterance acquisition unit 11, and acquires a user utterance us1 as a result of speech recognition of the acquired speech. The speech rate calculation unit 14 removes fillers f from the user utterance us1 to acquire a filler-removed user utterance us2. Fillers are words and sounds that have no semantic content in speech and are used to buy time, fill gaps in thought, connect words, etc.
[0021] Fillers can be removed by applying known techniques. For example, the speech rate calculation unit 14 may detect and remove fillers from the user utterance us1 by referring to a pre-registered filler dictionary. Alternatively, the speech rate calculation unit 14 may detect fillers from the user utterance us1 using a filler detection model for detecting fillers that is configured by machine learning.
[0022] The speech rate calculation unit 14 obtains the number of moras nm and the speech length ls of the user utterance us2 after the fillers have been removed. A mora is a unit (syllable) in prosody or phonology.
[0023] The speech rate calculation unit 14 acquires the number of moras nm, which is the number of moras included in the user utterance us2. The speech rate calculation unit 14 also acquires the speaking time of the user utterance us as the speech length ls. Then, the speech rate calculation unit 14 calculates the user's speaking rate ss by dividing the number of moras nm by the speech length ls. The user's speaking rate ss may be the number of moras per second.
[0024] The speech rate calculation unit 14 may calculate the average of the speech rates of a predetermined number of recent user utterances as the reference speech rate. In this embodiment, the speech rate calculation unit 14 may calculate the reference speech rate based on the most recent N user utterances stored in the user profile storage unit 22. FIG. 3 is a diagram showing the most recent predetermined number (N) of user utterances and their speech rates stored in the user profile storage unit 22. The user profile storage unit 22 stores the most recent predetermined number (N) of user utterances in association with each user.
[0025] The user utterances sp_a1, sp_a2, sp_a3, ..., sp_an stored in the user profile storage unit 22 may be voice data of user utterances acquired by the user utterance acquisition unit 11, or may be voice recognition results by the speech rate calculation unit 14. Every time a user utterance is acquired, the user utterances stored in the user profile storage unit 22 are updated to a predetermined number (N) of user utterances most recently acquired.
[0026] The user profile storage unit 22 may store the speech rate of each user utterance. That is, the speech rate calculation unit 14 may store the user speech rates ms_a1, ms_a2, ms_a3, . . . , ms_an calculated based on each of the user utterances sp_a1, sp_a2, sp_a3, . . . , sp_an in the user profile storage unit 22.
[0027] The speech rate calculation unit 14 obtains a predetermined number (N) of the most recent user speech rates ms_a1, ms_a2, ms_a3, . . . , ms_an from the user profile storage unit 22, and calculates the average of the obtained user speech rates to obtain the reference speech rate.
[0028] Referring again to FIG. 1 , the anxiety level measurement unit 15 measures the anxiety level of the user based on the user utterance. The anxiety level can be measured by applying known techniques. For example, the anxiety level measurement unit 15 may measure the anxiety level based on the user utterance using a machine learning model trained using a speech corpus in which speech voices are annotated with anxiety levels as training data. The anxiety level measurement unit 15 may also estimate the anxiety level of the user based on the user utterance using other known techniques.
[0029] The filler frequency acquiring unit 16 acquires the frequency of fillers in a user utterance. Specifically, as described with reference to Fig. 2, fillers are extracted from the user utterance, and therefore the filler frequency acquiring unit 16 can calculate the filler frequency, which is the number of fillers per unit time. Note that the filler frequency may be the number of fillers included in one unit of user utterance, or the number of fillers included in a predetermined number (N) of recent user utterances.
[0030] The speech rate setting unit 17 sets the system speech rate, which is the speech rate of the system utterance. Specifically, the speech rate setting unit 17 sets the speech rate of the system utterance output by the system utterance output unit 12.
[0031] The speech rate setting unit 17 may set the standard speech rate calculated by the speech rate calculation unit 14 as the system speech rate. Setting the system speech rate to a speech rate that corresponds to the user's speech rate can prevent the user from feeling uncomfortable during a conversation. Furthermore, applying the average of the speech rates of a predetermined number of recent user utterances as the standard speech rate can prevent sudden fluctuations in the system speech rate.
[0032] The speech rate setting unit 17 sets the system speech rate after the topic change detected by the topic change detection unit 13 to a speech rate that is different from the system speech rate before the topic change. In this way, when different system speech rates are set before and after the topic change, the system speech rate is set to an appropriate speech rate when the topic change occurs, thereby preventing the user from feeling uncomfortable with the system speech.
[0033] The speech rate setting unit 17 may set the system speech rate after the topic change to a slower rate than the system speech rate before the topic change. Setting the system speech rate in this way makes it easier for the user to recognize the new topic, and reduces the unnaturalness of the conversation between the user and the user.
[0034] More specifically, the speech rate setting unit 17 may set the speech rate obtained by adding a given rate value to the reference speech rate as the system speech rate before the topic change, and may set the speech rate obtained by subtracting the given rate value from the reference speech rate as the system speech rate after the topic change.
[0035] That is, the speech rate setting unit 17 may set the system speech rate according to the following formula.
[0036] Before topic change: System speech rate = Reference speech rate + Given speech rate After topic change: System speech rate = Reference speech rate - Given speech rate In this way, the system speech rate is set by adding or subtracting a given constant value based on the reference speech rate, so that the system speech rate can be set to a rate close to the user speech rate, which reduces the unnaturalness of the system speech before and after topic changes.
[0037] If the calculated system speech rate is greater than a given maximum value, the speech rate setting unit 17 sets the system speech rate to the maximum value. If the calculated system speech rate is less than a given minimum value, the speech rate setting unit 17 sets the system speech rate to the minimum value.
[0038] Furthermore, the speech rate setting unit 17 may set the system speech rate to a slower speech rate as the user's anxiety level acquired by the anxiety level measurement unit 15 increases. By setting the system speech rate to a slower speech rate as the user's anxiety level increases, the system will produce speech that calms the user, thereby reducing the user's sense of anxiety.
[0039] Specifically, the speech rate setting unit 17 may set the system speech rate by adding or subtracting a speed value according to the level of anxiety to or from the reference speech rate. That is, the speech rate setting unit 17 may set the system speech rate according to the following formula:
[0040] System speech rate = Reference speech rate - (Current anxiety level - Average anxiety level) x (Constant 1) (Constant 1 > 0) In this way, the system speech rate is set by adding or subtracting a speed value corresponding to the level of anxiety to or from the reference speech rate, so that the system speech rate can be set to a speech rate that is close to the user speech rate and corresponds to the level of anxiety, thereby providing a dialogue that feels natural to the user according to the level of anxiety.
[0041] If the system speech rate calculated according to the anxiety level is greater than a given maximum value, the speech rate setting unit 17 sets the system speech rate to the maximum value. If the calculated system speech rate is less than a given minimum value, the speech rate setting unit 17 sets the system speech rate to the minimum value.
[0042] Furthermore, the speech rate setting unit 17 may set a slower speech rate as the system speech rate, the more frequently fillers there are in at least one recent user utterance. Since a high frequency of fillers in a user utterance is likely to reflect tension or the like in the user, by setting the system speech rate to a slower speech rate the more frequently fillers there are in the user utterance, the user's tension can be alleviated.
[0043] Specifically, the speech rate setting unit 17 may set the system speech rate by adding or subtracting a speed value according to the frequency of fillers to or from the reference speech rate. That is, the speech rate setting unit 17 may set the system speech rate according to the following formula:
[0044] System speech rate = Reference speech rate - (Filler frequency x Constant 2) (Constant 2 > 0) In this way, the system speech rate is set by adding or subtracting a speed value according to the frequency of fillers to or from the reference speech rate, so that the system speech rate can be set to a speech rate that is close to the user's speech rate and that corresponds to the user's level of nervousness, thereby providing a dialogue that feels natural to the user according to the level of nervousness.
[0045] If the system speech rate calculated according to the frequency of fillers is greater than a given maximum value, the speech rate setting unit 17 sets the system speech rate to the maximum value. If the calculated system speech rate is less than a given minimum value, the speech rate setting unit 17 sets the system speech rate to the minimum value.
[0046] Alternatively, the speech rate setting unit 17 may set the system speech rate based on both the anxiety level and the frequency of fillers, as in the following formula:
[0047] System speech rate = Baseline speech rate - (Current anxiety level - Average anxiety level) x (Constant 3) - (Filler frequency x Constant 4) (Constant 3 > 0, Constant 4 > 0) Furthermore, the speech rate setting unit 17 may set the system speech rate based on the detection of a change in topic, the anxiety level of the user, and the frequency of fillers, as in the following formula.
[0048] Before a topic change: System speech rate = Baseline speech rate + Given speech rate - (Current anxiety level - Average anxiety level) x (Constant 1) - (Filler frequency x Constant 2) (Constant 1 > 0, Constant 2 > 0) After the topic change: System speech rate = Baseline speech rate - Given speech rate - (Current anxiety level - Average anxiety level) x (Constant 1) - (Filler frequency x Constant 2) (Constant 1 > 0, Constant 2 > 0) If the system speech rate calculated based on the detection of a change in topic, the user's anxiety level, and the frequency of fillers is greater than a given maximum value, the speech rate setting unit 17 sets the system speech rate to the maximum value. If the calculated system speech rate is less than a given minimum value, the speech rate setting unit 17 sets the system speech rate to the minimum value.
[0049] FIG. 4 is a flowchart showing the processing content of the information processing method in the dialogue system 1.
[0050] In step S1, the user utterance acquisition unit 11 acquires a user utterance. In the following step S2, the topic change detection unit 13 detects whether or not there is a change in the topic in the dialogue based on at least the user utterance in the dialogue between the user utterance and the system utterance.
[0051] In step S3, the speech rate setting unit 17 determines whether a change in topic has been detected by the topic change detection unit 13. If it is determined that a change in topic has been detected, the process proceeds to step S4. On the other hand, if it is determined that a change in topic has not been detected, the process returns to step S1.
[0052] In step S4, the speech rate setting unit 17 sets the system speech rate after the topic change detected by the topic change detection unit 13 to a speech rate different from the system speech rate before the topic change. Note that in step S4, the speech rate setting unit 17 may set the system speech rate based on at least one of the user's anxiety level acquired by the anxiety level measurement unit 15 and the frequency of fillers in the user's utterance acquired by the filler frequency acquisition unit 16, in addition to the topic change.
[0053] Next, an information processing program for causing a computer to function as the information processing device 10 of this embodiment will be described with reference to Fig. 5. Fig. 5 is a diagram showing the configuration of the information processing program. The information processing program P1 is configured to include a main module m10 that comprehensively controls information processing in the information processing device 10, a user utterance acquisition module m11, a system utterance output module m12, a topic change detection module m13, a speaking rate calculation module m14, an anxiety level measurement module m15, a filler frequency acquisition module m16, and a speaking rate setting module m17. Each of the modules m11 to m17 realizes a function for each of the functional units 11 to 17.
[0054] The information processing program P1 may be transmitted via a transmission medium such as a communication line, or may be stored in a recording medium M1 as shown in FIG.
[0055] According to the information processing device 10, information processing method, and information processing program P1 of the present embodiment described above, when a change in topic is detected in a dialogue between a user utterance and a system utterance, different system utterance rates are set before and after the topic change. By setting the system utterance rate to an appropriate rate when the topic changes, it is possible to prevent the user from feeling uncomfortable with the system utterance.
[0056] The information processing device according to the present disclosure may have the following configurations: The actions and effects of each configuration will be described as follows.
[0057] An information processing device according to one aspect of the present disclosure includes a topic change detection unit that detects a change in topic in a dialogue based on at least a user utterance in a dialogue between a user utterance, which is speech spoken by a user, and a system utterance, which is speech generated by the system, and a speech rate setting unit that sets a system speech rate, which is the speed of system speech after a topic change, to a speech rate that is different from the system speech rate before the topic change.
[0058] An information processing method according to one aspect of the present disclosure includes a topic change detection step executed by a processor, which detects a change in topic in a dialogue based on at least a user utterance in a dialogue between a user utterance, which is speech spoken by a user, and a system utterance, which is speech generated by the system; and a speech rate setting step, which sets a system speech rate, which is the speed of system speech after a topic change, to a speech rate different from the system speech rate before the topic change.
[0059] According to the above aspect, when a change in topic is detected in a dialogue between a user and a system, different system speech rates are set before and after the topic change. By setting the system speech rate to an appropriate speech rate when the topic change occurs, it is possible to prevent the user from feeling uncomfortable with the system speech.
[0060] In an information processing device according to another aspect, the speech rate setting unit may set the system speech rate after the topic change to a slower speech rate than the system speech rate before the topic change.
[0061] According to the above aspect, when a topic change occurs, the system speech rate after the topic change is set to a slower speech rate than the system speech rate before the topic change, which makes it easier for the user to recognize the new topic and reduces the unnaturalness of the user's dialogue.
[0062] In addition, an information processing device according to another aspect may further include a speech rate calculation unit that calculates a reference speech rate that serves as a standard for the system speech rate based on a user speech rate, which is the speed at which the user speaks. The speech rate setting unit may set the system speech rate to a speech rate obtained by adding a given rate value to the reference speech rate before the topic change, and may set the system speech rate to a speech rate obtained by subtracting the given rate value from the reference speech rate after the topic change.
[0063] According to the above aspect, the reference speech rate is calculated based on the user speech rate, and the system speech rate before and after the topic change is set based on the reference speech rate, so that the system speech rate can be set to a speed close to the user speech rate, thereby reducing the unnaturalness of the system speech before and after the topic change.
[0064] In addition, in the information processing device according to another aspect, the speech rate calculation unit may calculate an average of speech rates of a predetermined number of recent user utterances as the reference speech rate.
[0065] According to the above aspect, the average of the speech rates of a predetermined number of recent user utterances is applied to the reference speech rate, thereby preventing sudden fluctuations in the system speech rate.
[0066] In addition, in the information processing device according to another aspect, the speech rate calculation unit may calculate the number of moras per unit time of the user utterance from which the fillers have been removed as the speech rate of the user utterance.
[0067] According to the above aspect, the speaking rate expressed by the number of moras is calculated based on the user's utterance from which fillers that have no semantic content have been removed, thereby making it possible to appropriately obtain the actual speaking rate of the user's utterance.
[0068] In addition, an information processing device according to another aspect may further include an anxiety level measurement unit that measures the user's anxiety level based on the user's speech, and the speech rate setting unit may set the system speech rate to a slower speech rate as the anxiety level increases.
[0069] According to the above aspect, the system speech rate is set to a slower rate as the user's level of anxiety increases, so the system produces speech that calms the user, thereby reducing the user's sense of anxiety.
[0070] In addition, in an information processing device according to another aspect, the speech rate setting unit may set the system speech rate by adding or subtracting a rate value corresponding to the level of anxiety to or from the reference speech rate.
[0071] According to the above aspect, the system speech rate is set by adding or subtracting a speed value corresponding to the level of anxiety to or from the reference speech rate, so that the system speech rate can be set to a speech rate that is close to the user speech rate and that corresponds to the level of anxiety, thereby providing a dialogue that feels natural to the user according to the level of anxiety.
[0072] In addition, an information processing device according to another aspect may further include a filler frequency acquisition unit that acquires the frequency of fillers in user utterances, and the speech rate setting unit may set a slower speech rate as the system speech rate the more frequently fillers there are in at least one recent user utterance.
[0073] According to the above aspect, the frequency of fillers in a user's speech is likely to reflect tension or the like in the user, and the system speech rate is set to a slower rate the more fillers there are in the user's speech, thereby reducing the user's sense of tension.
[0074] In addition, in an information processing device according to another aspect, the speech rate setting unit may set the system speech rate by adding or subtracting a speed value according to the frequency of fillers to or from the reference speech rate.
[0075] According to the above aspect, the system speech rate is set by adding or subtracting a speed value corresponding to the frequency of fillers from the reference speech rate, so that the system speech rate can be set to a speech rate that is close to the user speech rate and corresponds to the user's level of tension. Therefore, a dialogue that feels natural to the user can be provided according to the level of tension.
[0076] The block diagram shown in FIG. 1 shows functional blocks. These functional blocks (components) are realized by any combination of at least one of hardware and software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using one device that is physically or logically coupled, or may be realized using two or more devices that are physically or logically separated and connected directly or indirectly (for example, by wire, wirelessly, etc.). The functional block may also be realized by combining software with the one device or the multiple devices.
[0077] Functions include, but are not limited to, judgment, determination, judgment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, election, establishment, comparison, assumption, expectation, consideration, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocation, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.
[0078] For example, the information processing device 10 according to an embodiment of the present invention may function as a computer. Fig. 6 is a diagram showing an example of the hardware configuration of the information processing device 10 according to this embodiment. The information processing device 10 may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, etc.
[0079] In the following description, the term "apparatus" can be interpreted as a circuit, a device, a unit, etc. The hardware configuration of the information processing device 10 may be configured to include one or more of the apparatuses shown in FIG. 6, or may be configured to exclude some of the apparatuses.
[0080] Each function of the information processing device 10 is realized by loading specified software (programs) onto hardware such as the processor 1001, memory 1002, etc., so that the processor 1001 performs calculations and controls communication via the communication device 1004 and the reading and / or writing of data in the memory 1002 and storage 1003.
[0081] The processor 1001 controls the entire computer by running, for example, an operating system. The processor 1001 may be configured as a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. For example, the functional units 11 to 17 shown in FIG. 1 may be realized by the processor 1001.
[0082] Furthermore, the processor 1001 reads programs (program codes), software modules, and data from the storage 1003 and / or the communication device 1004 into the memory 1002, and executes various processes in accordance with these. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, the functional units 11 to 17 of the information processing device 10 may be implemented by a control program stored in the memory 1002 and executed by the processor 1001. While the above-described various processes have been described as being executed by one processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented on one or more chips. The programs may also be transmitted from a network via a telecommunications line.
[0083] The memory 1002 is a computer-readable recording medium and may be composed of at least one of, for example, a read-only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), and a random access memory (RAM). The memory 1002 may also be called a register, a cache, a main memory (primary storage device), or the like. The memory 1002 can store executable programs (program codes), software modules, and the like for implementing an information processing method according to one embodiment of the present invention.
[0084] Storage 1003 is a computer-readable recording medium, and may be, for example, at least one of an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray disc), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other suitable medium including memory 1002 and / or storage 1003.
[0085] The communication device 1004 is hardware (transmission / reception device) for performing communication between computers via a wired and / or wireless network, and is also called, for example, a network device, a network controller, a network card, or a communication module.
[0086] The input device 1005 is an input device (for example, a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives input from the outside. The output device 1006 is an output device (for example, a display, a speaker, an LED lamp, etc.) that outputs to the outside. The input device 1005 and the output device 1006 may be integrated into one device (for example, a touch panel).
[0087] Furthermore, each device such as the processor 1001 and the memory 1002 is connected by a bus 1007 for communicating information. The bus 1007 may be configured as a single bus, or may be configured as different buses between the devices.
[0088] The information processing device 10 may also be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented by at least one of these pieces of hardware.
[0089] The notification of information is not limited to the aspects / embodiments described in the present disclosure and may be performed using other methods. For example, the notification of information may be performed by physical layer signaling (e.g., Downlink Control Information (DCI) and Uplink Control Information (UCI)), higher layer signaling (e.g., Radio Resource Control (RRC) signaling, Medium Access Control (MAC) signaling, and broadcast information (Master Information Block (MIB) and System Information Block (SIB))), other signals, or a combination thereof. Furthermore, the RRC signaling may be referred to as an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.
[0090] Each aspect / embodiment described in the present disclosure may be applied to at least one of systems using LTE (Long Term Evolution), LTE-Advanced (LTE-A), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), FRA (Future Radio Access), NR (New Radio), W-CDMA (registered trademark), GSM (registered trademark), CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark), IEEE 802.20, UWB (Ultra-Wideband), Bluetooth (registered trademark), or other appropriate systems, and next-generation systems extended based on these. Furthermore, a combination of multiple systems (e.g., a combination of at least one of LTE and LTE-A with 5G, etc.) may also be applied.
[0091] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.
[0092] In the present disclosure, a specific operation described as being performed by a base station may be performed by its upper node in some cases. In a network consisting of one or more network nodes having a base station, it is clear that various operations performed for communication with a terminal may be performed by at least one of the base station and another network node other than the base station (for example, but not limited to, an MME or an S-GW). Although the above example illustrates a case where there is one other network node other than the base station, a combination of multiple other network nodes (for example, an MME and an S-GW) may also be used.
[0093] Information etc. may be output from a higher layer (or a lower layer) to a lower layer (or a higher layer), or may be input / output via multiple network nodes.
[0094] Input and output information may be stored in a specific location (for example, memory) or managed in a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.
[0095] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).
[0096] Each aspect / embodiment described in this disclosure may be used alone, in combination, or switched depending on the implementation. Furthermore, notification of predetermined information (e.g., notification that "X is true") is not limited to being done explicitly, but may be done implicitly (e.g., by not notifying the predetermined information).
[0097] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.
[0098] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0099] Software, instructions, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies such as coaxial cable, fiber optic cable, twisted pair, and Digital Subscriber Line (DSL), and / or wireless technologies such as infrared, radio, and microwave, these wired and / or wireless technologies are included within the definition of transmission media.
[0100] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0101] It should be noted that terms explained in this disclosure and / or terms necessary for understanding this specification may be replaced with terms having the same or similar meanings.
[0102] As used in this disclosure, the terms "system" and "network" are used interchangeably.
[0103] Furthermore, the information, parameters, etc. described in the present disclosure may be expressed as absolute values, relative values from a predetermined value, or other corresponding information. For example, a radio resource may be indicated by an index.
[0104] The names used for the above-described parameters are not intended to be limiting in any way. Furthermore, the mathematical expressions using these parameters may differ from those explicitly disclosed in this disclosure. The various channels (e.g., PUCCH, PDCCH, etc.) and information elements may be identified by any suitable names, and therefore the various names assigned to these various channels and information elements are not intended to be limiting in any way.
[0105] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.
[0106] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly specified otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."
[0107] When designations such as "first," "second," etc. are used in this disclosure, any reference to an element does not generally limit the quantity or order of those elements. These designations may be used herein as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed therein or that the first element must precede the second element in some way.
[0108] The "means" in the configuration of each of the above devices may be replaced with "part," "circuit," "device," etc.
[0109] To the extent that the terms "include," "including," and variations thereof are used herein or in the claims, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, the term "or," as used herein or in the claims, is not intended to be an exclusive or.
[0110] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.
[0111] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."
[0112] The information processing device 10 of the present disclosure may have the following configuration.
[0113] [1] a topic change detection unit that detects a change in topic in a dialogue based on at least a user utterance in a dialogue between a user utterance, which is an utterance by a user, and a system utterance, which is an utterance by voice generated by the system; a speech rate setting unit that sets a system speech rate, which is a speed of system speech after a topic change, to a speech rate different from the system speech rate before the topic change; An information processing device comprising: [2] the speech rate setting unit sets the system speech rate after the topic change to a slower speech rate than the system speech rate before the topic change; [1] The information processing device according to [1]. [3] a speech rate calculation unit that calculates a reference speech rate as a reference for the system speech rate based on a user speech rate that is a speed of user speech, the speech rate setting unit sets a speech rate obtained by adding a given rate value to a reference speech rate as the system speech rate before the topic change, and sets a speech rate obtained by subtracting the given rate value from the reference speech rate as the system speech rate after the topic change; [1] or [2]. [4] the speech rate calculation unit calculates an average of the speech rates of a predetermined number of recent user utterances as the reference speech rate; [3] The information processing device according to [3]. [5] the speech rate calculation unit calculates the number of moras per unit time of the user utterance from which the fillers have been removed as the speech rate of the user utterance; [4] The information processing device according to [4]. [6] An anxiety level measurement unit that measures a user's anxiety level based on the user's utterance, the speech rate setting unit sets a slower speech rate as the system speech rate as the anxiety level increases; The information processing device according to any one of [3] to [5]. [7] the speech rate setting unit sets the system speech rate by adding or subtracting a rate value corresponding to the degree of anxiety to or from the reference speech rate; [6] The information processing device according to [6]. [8] a filler frequency acquisition unit that acquires a frequency of fillers in a user utterance, the speech rate setting unit sets a slower speech rate as the system speech rate as the frequency of fillers in at least one recent user utterance increases; The information processing device according to any one of [3] to [7]. [9] the speech rate setting unit sets the system speech rate by adding or subtracting a speed value according to the frequency of the filler to or from the reference speech rate; [8] The information processing device according to [8].
[10] a topic change detection step of detecting a change in topic in the dialogue based on at least a user utterance in the dialogue between a user utterance, which is an utterance by a user, and a system utterance, which is an utterance by voice generated by the system; a speech rate setting step of setting a system speech rate, which is a speed of system speech after a topic change, to a speech rate different from a system speech rate before the topic change; An information processing method executed by a processor, comprising: [Explanation of symbols]
[0114] 1...dialogue system, 10...information processing device, 11...user speech acquisition unit, 12...system speech output unit, 13...topic change detection unit, 14...speech rate calculation unit, 15...anxiety level measurement unit, 16...filler frequency acquisition unit, 17...speech rate setting unit, 21...dialogue information storage unit, 22...user profile storage unit, M1...recording medium, m11...user speech acquisition module, m12...system speech output module, m13...topic change detection module, m14...speech rate calculation module, m15...anxiety level measurement module, m16...filler frequency acquisition module, m17...speech rate setting module, P1...information processing program.
Claims
1. a topic change detection unit that detects a change in topic in a dialogue based on at least a user utterance in a dialogue between a user utterance, which is an utterance by a user, and a system utterance, which is an utterance by voice generated by a system; a speech rate setting unit that sets a system speech rate, which is a speed of the system speech after a topic change, to a speech rate different from the system speech rate before the topic change; An information processing device comprising:
2. the speech rate setting unit sets the system speech rate after the topic change to a speech rate slower than the system speech rate before the topic change. The information processing device according to claim 1 .
3. a speech rate calculation unit that calculates a reference speech rate that is a reference for the system speech rate, based on a user speech rate that is the speed of the user's speech; the speech rate setting unit sets the system speech rate to a speech rate obtained by adding a given rate value to the reference speech rate before the topic change, and sets the system speech rate to a speech rate obtained by subtracting the given rate value from the reference speech rate after the topic change. The information processing device according to claim 1 .
4. the speech rate calculation unit calculates an average of speech rates of a predetermined number of recent user utterances as the reference speech rate; The information processing device according to claim 3 .
5. The information processing device according to claim 4 , wherein the speech rate calculation unit calculates the number of moras per unit time of the user utterance from which the fillers have been removed as the speech rate of the user utterance.
6. an anxiety level measurement unit that measures a user's anxiety level based on the user utterance, the speech rate setting unit sets a slower speech rate as the system speech rate as the anxiety level increases; The information processing device according to claim 3 .
7. the speech rate setting unit sets the system speech rate by adding or subtracting a rate value corresponding to the level of anxiety to or from the reference speech rate. The information processing device according to claim 6 .
8. a filler frequency acquisition unit that acquires a frequency of fillers in the user utterance, the speech rate setting unit sets a slower speech rate as the system speech rate as the frequency of fillers in at least one recent user utterance increases; The information processing device according to claim 3 .
9. the speech rate setting unit sets the system speech rate by adding or subtracting a speed value according to the frequency of the filler to or from the reference speech rate. The information processing device according to claim 8 .
10. a topic change detection step of detecting a change in topic in a dialogue based on at least a user utterance in a dialogue between a user utterance, which is an utterance by a user, and a system utterance, which is an utterance by voice generated by a system; a speech rate setting step of setting a system speech rate, which is a speed of the system speech after a topic change, to a speech rate different from the system speech rate before the topic change; An information processing method executed by a processor, comprising:
Citation Information
Patent Citations
Voice interaction apparatus
JP2008026463A