Information processing device, information processing method, and program

The information processing device addresses the challenge of varying speech rates by estimating speech rate and detecting silent periods to adjust backchannel response timing, ensuring smooth dialogue and reducing processing load.

WO2025249074A1PCT designated stage Publication Date: 2025-12-04SONY GROUP CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/016360
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-29
Filing Date
2025-04-30
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing technologies struggle to maintain natural and smooth dialogue by accurately determining the timing for backchannel responses, which can lead to user confusion when speech rates vary.

Method used

An information processing device that estimates speech rate, detects silent periods, and sets a time threshold for outputting responses based on the speech rate, using spectral envelope estimation to quickly and efficiently detect silent periods and adjust the timing of backchannel responses.

Benefits of technology

Enables natural and smooth dialogue by ensuring backchannel responses are timely and appropriate to the speaker's pace, maintaining the speaker's tempo and reducing processing load, suitable for mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025016360_04122025_PF_FP_ABST
    Figure JP2025016360_04122025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device according to one embodiment of the present technology comprises an estimation unit, a detection unit, and an output unit. The estimation unit estimates a speech rate, which is the rate of speech. The detection unit detects a non-speech time, which is a duration of time without speech. The output unit sets a time threshold according to the speech rate, and outputs a response to the speech if the non-speech time exceeds the time threshold. The information processing device executes the estimation of the speech rate, the detection of the non-speech time, and the setting of a threshold according to the speech rate, and outputs a response to the speech if the non-speech time exceeds the threshold. This makes it possible to achieve a natural and smooth conversation.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and program

[0001] The present technology relates to an information processing device, an information processing method, and a program that can be applied to a dialogue agent or the like.

[0002] Patent Literature 1 discloses a voice dialogue device. This voice dialogue device estimates the timing of a backchannel response based on the speaker's voice features. It also determines whether or not to output a backchannel response sound depending on the voice power immediately before that timing. This makes it possible to realize a natural and smooth dialogue.

[0003] Japanese Patent Application Laid-Open No. 2009-003040

[0004] Thus, there is a demand for technology that enables natural and smooth dialogue.

[0005] In view of the above circumstances, an object of the present technology is to provide an information processing device, an information processing method, and a program that enable natural and smooth dialogue to be realized.

[0006] In order to achieve the above object, an information processing device according to an embodiment of the present technology includes an estimation unit, a detection unit, and an output unit. The estimation unit estimates a speech rate, which is a speed of speech. The detection unit detects a silent period, which is a time duration during which no speech is being made. The output unit sets a time threshold according to the speech rate, and outputs a response to the speech when the silent period exceeds the threshold.

[0007] This information processing device estimates the speech rate, detects silent periods, and sets a threshold value according to the speech rate, and when the silent periods exceed the threshold value, it outputs a response to the speech, thereby enabling natural and smooth dialogue.

[0008] The detection unit may detect the silent period by estimating a spectral envelope of the speech.

[0009] The output unit may set the threshold value so that the threshold value decreases as the speech rate increases.

[0010] The estimation unit may estimate the speech rate based on an average of the silent periods in the past.

[0011] The estimation unit may estimate the speech rate such that the speech rate increases as the average of the past silent periods decreases.

[0012] The response may be a vocal response.

[0013] The response may be a nod of approval.

[0014] The output unit may change the content of the response depending on the magnitude of the threshold.

[0015] The output unit may change the magnitude of the threshold value depending on the clarity of the content of the utterance.

[0016] The information processing device may further include a generation unit that generates an avatar, which is a virtual character that makes the response.

[0017] When there are a plurality of avatars, the output unit may set the threshold for each of the avatars to a different value.

[0018] The output unit may change the threshold or change the frequency of the response according to the content of the character setting of the avatar.

[0019] The output unit may output the response when a user listening to the utterance understands the content of the utterance.

[0020] An information processing method according to an embodiment of the present technology includes estimating a speech rate, detecting a silent period, which is a time period during which no speech is being made, setting a time threshold according to the speech rate, and outputting a response to the speech when the silent period exceeds the threshold.

[0021] A program according to an embodiment of the present invention causes a computer system to execute the following steps: estimating a speech rate, which is the speed of speech; detecting a silent period, which is a period of time during which no speech is being made; setting a time threshold according to the speech rate, and outputting a response to the speech when the silent period exceeds the threshold.

[0022] FIG. 1 is a schematic diagram showing a configuration example of an information processing device according to an embodiment of the present technology; FIG. 2 is a flowchart showing details of processing according to the present technology; FIG. 3 is a graph showing an example of a result of F0 estimation; FIG. 4 is a graph showing an example of a result of spectrum envelope estimation; FIG. 5 is a table showing the results of an experiment regarding the relationship between speaking rate and backchannel response time; FIG. 6 is a table showing the results of an experiment regarding the relationship between speaking rate and backchannel response time; FIG. 7 is a graph showing the relationship between speaking rate and average pause length; FIG. 8 is a graph showing the relationship between backchannel response time and average pause length; FIG. 9 is a schematic diagram showing a configuration example of an information processing device applicable to dialogue between people; FIG. 10 is a flowchart related to processing according to the present embodiment; FIG. 11 is a block diagram showing an example of the hardware configuration of a computer capable of realizing an information processing device.

[0023] First Embodiment Hereinafter, an embodiment according to the present technology will be described with reference to the drawings.

[0024] 1 is a schematic diagram showing a configuration example of an information processing device 1 according to an embodiment of the present technology. The information processing device 1 is applicable to, for example, a dialogue agent. In this embodiment, while a user 2 is speaking, the information processing device 1 outputs audible responses.

[0025] The information processing device 1 has a display unit 3, an operation unit 4, a communication unit 5, a storage unit 6, a microphone 7, a speaker 8, and a controller 9. These are connected to each other via a bus 10. Instead of the bus 10, the blocks may be connected using a communication network or a non-standardized proprietary communication method.

[0026] The display unit 3 is a display device using, for example, a liquid crystal display (LCD) or an EL (Electro-Luminescence) display, and displays various images, a GUI (Graphical User Interface), etc. The operation unit 4 is, for example, a keyboard or a pointing device. The operation unit 4 may also be a touch panel, in which case it is configured integrally with the display unit 3.

[0027] The communication unit 5 is a communication module for communicating with other devices via a network such as a local area network (LAN) or a wide area network (WAN). A wireless LAN module such as Wi-Fi or a communication module for short-range wireless communication such as Bluetooth (registered trademark) may be provided. Communication devices such as a modem or a router may also be used. The display unit 3, the operation unit 4, and the communication unit 5 do not necessarily have to be provided in the information processing device 1.

[0028] The storage unit 6 is a storage device such as a non-volatile memory, and may be, for example, a hard disk drive (HDD) or a solid state drive (SSD). Any other computer-readable non-transitory storage medium may be used. The storage unit 6 stores, for example, a control program for controlling the overall operation of the information processing device 1. The method for installing the control program in the storage unit 6 is not limited. The storage unit 6 also stores voice data for backchannel responses, etc.

[0029] The microphone 7 acquires sounds generated around the information processing device 1. In this embodiment, the microphone 7 acquires speech by the user 2. The microphone 7 can also acquire environmental sounds around the user 2.

[0030] The speaker 8 outputs audio. In this embodiment, audio responses are output by the speaker 8. The configuration of the speaker 8 is not limited, and a speaker 8 capable of outputting stereo audio, monaural audio, or the like may be used as appropriate.

[0031] The controller 9 controls the operation of each block included in the information processing device 1. The controller 9 includes hardware circuits necessary for a computer, such as a CPU and memory (RAM, ROM). The CPU executes programs according to the present technology stored in the storage unit 6 or the memory, thereby performing various processes. The controller 9 may be, for example, a programmable logic device (PLD) such as a field programmable gate array (FPGA), or another device such as an application specific integrated circuit (ASIC).

[0032] In this embodiment, the CPU of the controller 9 executes a program (e.g., an application program) according to the present technology, thereby realizing functional blocks including a voice input unit 11, a feature extraction unit 12, a silent time detection unit 13, a speech rate estimation unit 14, a backchannel response control unit 15, and a voice output unit 16. These functional blocks then execute the information processing method according to the present embodiment. Note that dedicated hardware such as an IC (integrated circuit) may be used as appropriate to realize each functional block.

[0033] The voice input unit 11 receives voice acquired by the microphone 7. Specifically, voice related to the speech of the user 2 may be input. For example, voice with relatively short content such as "Hello" or "I see," or voice such as a long sentence of several seconds to several tens of seconds may be input. In addition, various types of voice may be input to the voice input unit 11, such as minute environmental sounds and noise.

[0034] The feature extractor 12 extracts features of speech. In this embodiment, the feature extractor 12 performs F0 estimation and spectral envelope estimation.

[0035] The silent time detection unit 13 detects silent time. Silent time is a time span during which no speech is made by the user 2. For example, if the user 2 speaks two sentences in succession, the time span that is a break in speech between when the user 2 finishes speaking the first sentence and when the user 2 starts speaking the second sentence is detected as silent time. Furthermore, even if the user 2 speaks only one sentence, if a pause occurs in the middle of the sentence due to taking a breath or the like, this time span can also be detected as silent time.

[0036] The speech rate estimation unit 14 estimates the speech rate, which is the speed at which the user 2 speaks. The speech rate is estimated in units of, for example, mora per second (mora / s). A mora is a sound segmentation unit having a certain time length, and is sometimes called a beat. In Japanese, one kana character basically corresponds to one mora.

[0037] The backchannel response control unit 15 determines whether or not to cause the audio output unit 16 to output a backchannel response based on a predetermined condition.

[0038] The audio output unit 16 outputs a response to the utterance via the speaker 8. In particular, in this embodiment, the audio output unit 16 outputs a vocal interjection as a response. A response is a word that means a reply or reaction to an utterance. For example, a response such as "My hobby is reading" to the utterance "What are your hobbies?" is an example of a response.

[0039] A backchannel is a word that means a short reply or reaction to an utterance. For example, in response to an utterance such as "My hobby is reading," short utterances such as "Um," "Yes," or "I see" are included in audible backchannels. Relatively short questions such as "Really?" may also be included in audible backchannels. The specific content of audible backchannels is not limited.

[0040] Other examples of backchannels include nodding, changes in facial expression, eye contact, etc. Backchannels are also a concept included in responses. The information processing device 1 may output any type of backchannel or any type of response other than backchannels.

[0041] Other than that, the specific configuration of the information processing device 1 is not limited. The speech rate estimation unit 14 corresponds to an embodiment of an estimation unit related to the present technology. The feature extraction unit 12 and the silent time detection unit 13 correspond to an embodiment of a detection unit related to the present technology. The backchannel response control unit 15 and the audio output unit 16 correspond to an embodiment of an output unit related to the present technology.

[0042] "Processing Flow" Fig. 2 is a flowchart showing the details of the processing according to the present technology. In this embodiment, the processing according to this flowchart is executed at a predetermined cycle. The predetermined cycle may be any cycle, such as once every 100 milliseconds.

[0043] Speech captured by the microphone 7 is input to the speech input unit 11 (step 101). For example, 100 milliseconds of speech is input. Next, the feature extraction unit 12 estimates the F0 of the captured speech (step 102). F0 is also called the fundamental frequency or pitch.

[0044] 3 is a graph showing an example of the results of F0 estimation. The horizontal axis of the graph represents time (seconds) and the vertical axis represents the frequency (Hz) of the speech. In this example, for ease of understanding, the estimation results are shown for a speech with a duration of approximately 20 seconds.

[0045] In the graph, frequency bands where pitch was present are shown in gray. On the other hand, frequency bands where pitch was not present are shown in white and nothing is shown. The frequency bands where pitch was present vary depending on the time, but pitch was mostly distributed between 200 and 300 Hz.

[0046] The feature extraction unit 12 estimates the spectral envelope (step 103). The spectral envelope indicates the intensity and distribution of the frequency components of the speech. The feature extraction unit 12 estimates the spectral envelope using the result of the F0 estimation.

[0047] Figure 4 is a graph showing an example of the results of spectral envelope estimation. The horizontal axis of the graph represents time (seconds), and the vertical axis represents the frequency of the audio (Hz). The graph also shows the intensity (power) of the audio using shades of gray. The darker the gray, the lower the power, and the lighter the gray, the higher the power. As the graph shows, power tends to increase as the frequency decreases over time.

[0048] There are no particular limitations on the specific method for performing F0 estimation and spectral envelope estimation, and any known method may be used.

[0049] It is determined for each time whether or not there is silence (step 104). That is, the processes up to step 103 are performed for the entire acquired speech of, for example, 100 milliseconds, but the processes from step 105 onwards are performed once for each of the speech segments obtained by dividing the speech of, for example, 100 milliseconds, in chronological order. Hereinafter, the divided speech will be referred to as divided speech. For example, 100 milliseconds of speech is divided into 100 1-millisecond divided speech segments, and the processes from step 105 onwards are repeated a total of 100 times. Note that the number of divisions and division widths described above are merely examples and are not limited to these.

[0050] It is determined whether the divided speech is silent (step 105). In this embodiment, the silent time detection unit 13 uses the results of the spectral envelope estimation to detect the silent time. Specifically, the silent time detection unit 13 determines a power threshold, and if the power is less than the threshold at any frequency of the divided speech, the divided speech is determined to be silent. On the other hand, if there is a frequency where the power is equal to or greater than the threshold, the divided speech is determined to be not silent (to be speech).

[0051] In the example of Figure 4, there are certain time ranges (i.e., ranges that are darker gray than other times) where the power is lower than other times at low frequencies of about 100 Hz around 3.5 seconds, 7 seconds, 14 seconds, etc. In fact, no speech is being produced during these time ranges.

[0052] An appropriate threshold is set so that the divided audio included in such a time width can be detected as silent. In this example, the power threshold is set to 200. As a result, for example, if the divided audio is a 5-second portion of audio, the power is below the threshold in most frequency bands, but the power is above the threshold around 100 Hz, so the divided audio is determined to be silent. On the other hand, if the divided audio is 7 seconds long, the power is below the threshold even around 100 Hz, so the divided audio is determined to be silent.

[0053] In this embodiment, both F0 estimation and spectral envelope estimation are used, but within the scope of feasibility of this technology, only one of these may be used, or silent periods may be detected using another method that does not use either of these.

[0054] If the divided speech is silent (Yes in step 105), a process of increasing the silent time is executed (step 106). Specifically, for example, the silent time is stored as a variable in the storage unit 6, and the value of the time width of the divided speech is added to this value. For example, if the time width is 1 millisecond, 1 millisecond is added to the silent time.

[0055] When this process is repeated, for example, if silent divided audio is detected consecutively, the silent time will increase by 1 millisecond, 2 milliseconds, etc. On the other hand, if speech-containing divided audio is detected (No in step 105), the silent time will be reset to 0 by the process in step 115, which will be described later. Then, if silent divided audio is detected consecutively again, the silent time will increase again.

[0056] The backchannel response control unit 15 determines whether the backchannel response flag is off (step 107). In this embodiment, if the silent period exceeds a predetermined threshold, a backchannel response is output only once. No backchannel response is output after the first time until the next speech is detected and the silent period is reset to 0.

[0057] The backchannel response flag is a flag that indicates whether or not the backchannel response, which is output only once, has been output. If a backchannel response has been output, the backchannel response flag is turned on, and if no backchannel response has been output, the flag is turned off. Even if the backchannel response flag is turned on, if speech is detected, the backchannel response flag is set to off by the processing in step 114, which will be described later, and the backchannel response can be output again.

[0058] If the backchannel response flag is off (Yes in step 107), it is determined whether the silent time exceeds the backchannel response waiting time (step 108). The backchannel response waiting time is a value indicating how much the silent time must increase before a backchannel response is output. The backchannel response waiting time is stored as a variable in the storage unit 6. The backchannel response waiting time corresponds to an embodiment of a time threshold value according to the present technology.

[0059] If the silent time exceeds the backchannel wait time (Yes in step 108), a backchannel is spoken (output) (step 109). Specifically, the backchannel voice stored in the storage unit 6 is output by the audio output unit 16 via the speaker 8. Thereafter, the backchannel control unit 15 sets the backchannel completed flag to ON (step 110).

[0060] In Figure 3, when there is no pitch, i.e., when all frequencies are white, the power does not exceed the threshold even in spectral envelope estimation, and is detected as silence. Also in Figure 3, the timing at which backchannel responses are output is indicated by solid black vertical lines. A backchannel response is output when silent divided audio continues for a period of time equal to or greater than a predetermined threshold, and this predetermined threshold corresponds to the backchannel response waiting time.

[0061] For example, the t shown in the figure 1 ~t 3 The time range of t is a silent time. When the divided voices in the time range are processed sequentially, the silent time ranges from 0 to t 3 -t 1 It increases until t 2 -t 1 corresponds to the waiting time for a response, and the time is t2 When the silent time reaches t 2 -t 1 When this is reached, a response is output.

[0062] This turns on the backchannel response flag, and then at time t 3 No response is output until time t 3 The speech is detected and the backchannel flag is turned off again.

[0063] If the divided audio is speech (No in step 105), it is determined whether the silent time is 0.2 seconds or more and less than 1 second (step 111). If the immediately previous divided audio processed was speech, the silent time was reset to 0 by the immediately previous processing in step 115, and therefore, in the processing of step 111 for the current divided audio, it is determined that the silent time is not 0.2 seconds or more and less than 1 second (No in step 111). Therefore, the processing of step 111 becomes Yes only when the immediately previous divided audio was speechless (and the silent time lasted for 0.2 seconds or more).

[0064] If the silent time is 0.2 seconds or more and less than 1 second (Yes in step 111), the current silent time is added to the silent time history (step 112). For example, a silent time such as "0.3 seconds" or "0.7 seconds" is stored in the storage unit 6 as the silent time history.

[0065] The speech rate estimation unit 14 estimates the backchannel response waiting time (step 113). Fig. 5 is a table showing the results of an experiment on the relationship between the speech rate and the backchannel response waiting time. The inventor conducted an experiment to consider what timing of a backchannel response would be perceived as natural by user 2 when information processing device 1 returns a backchannel response to an utterance by user 2.

[0066] As shown in the table, in this experiment, the same sentence was read aloud by the same person at three different speaking rates (average speaking rates): 10.50 moras per second, 8.76 moras per second, and 7.26 moras per second. We then investigated the backchannel latency (the length of the pause preceding a vocal backchannel) that the three subjects found most natural.

[0067] In the table, the experimental results for Subject 1 (white), Subject 2 (polka dots), and Subject 3 (black) are shown by the position of the triangle marks. For example, the center position of the "200 ms" block indicates that the backchannel wait time is 200 ms. For example, for Subject 1, the backchannel wait time for 10.50 mora per second is approximately 210 ms. In other words, when a sentence is read aloud at 10.50 mora per second, Subject 1 feels that it is most natural to make a backchannel when there is 210 ms of silence.

[0068] The experiment showed that, although there were differences between each subject, the faster the speaking rate, the shorter the waiting time for natural backchannel responses, as shown in the table.

[0069] FIG. 6 is a table showing the results of an experiment on the relationship between speech rate and backchannel response waiting time. Based on the results of the experiment shown in FIG. 5, the inventor conducted an experiment to investigate the relationship between speech rate and backchannel response waiting time in more detail. FIGS. 6A to 6C show the results of a repeat experiment similar to that shown in FIG. 5. However, in this experiment, each subject read the sentences aloud themselves. Therefore, there were differences in speech rate between each subject.

[0070] For example, the backchannel latency for subject 1 at 10.49 mora per second was approximately 200 ms. As shown in Figures 6A to 6C, this experiment also showed that the faster the speech rate, the shorter the backchannel latency became.

[0071] FIG. 7 is a graph showing the relationship between speaking rate and average pause length. The inventors newly investigated the relationship between speaking rate and average pause length for each of the speech sounds read aloud in the experiment shown in FIG. 6. The average pause length is the average value of silent periods. The vertical axis of the graph represents speaking rate, and the horizontal axis represents average pause length. The graph plots the values ​​of subjects 1 to 3 using the same colors and patterns as in FIG. 6. For example, the average pause length for the speech sound read aloud by subject 1 (white) in FIG. 6 at 10.49 mora per second is slightly less than 0.4 seconds.

[0072] As shown in the graph, the result shows that there is a negative correlation between speaking rate and average pause length. In other words, the faster the speaking rate, the shorter the average pause length. Furthermore, when a straight line is calculated to approximate each point, as shown in the graph, the result is y = -6.7772x + 13.475, and the root mean square error R 2 A straight line with ρ = 0.9299 was obtained.

[0073] Fig. 8 is a graph showing the relationship between backchannel response waiting time and average pause length. Fig. 6 shows the relationship between backchannel response waiting time and speech rate, and Fig. 7 shows the relationship between speech rate and average pause length. Therefore, based on these, the relationship between backchannel response waiting time and average pause length can be obtained. Fig. 8 shows this relationship in a graph. The vertical axis of the graph shows backchannel response waiting time, and the horizontal axis shows average pause length.

[0074] For example, for subject 1, the results obtained in Figure 6 are "backchannel wait time of 200 ms, speech rate of 10.49 mora per second," and in Figure 7 are "speech rate of 10.49 mora per second, average pause length of just under 0.4 seconds." Therefore, the relationship obtained is "backchannel wait time of 200 ms, average pause length of just under 0.4 seconds," and values ​​close to this are actually plotted in Figure 8 (although there is some error due to the approximation made in Figure 7). The values ​​in other respects are also close to the actual relationship.

[0075] In step 113, the speech rate estimation unit 14 estimates the backchannel response waiting time based on the average of past silent periods. Specifically, it acquires a plurality of silent period histories stored in the storage unit 6 and calculates the average value of these. Here, the silent period corresponds to the pause length, so the average value of the silent period histories corresponds to the average pause length in Fig. 8. The speech rate estimation unit 14 estimates the backchannel response waiting time based on the average pause length, for example, using the relationship in Fig. 8.

[0076] 8 is merely an example, and the actual relationship between the average pause length and the backchannel response waiting time is not limited to this. For example, a nonlinear relationship may hold.

[0077] The speech rate estimation unit 14 then changes the backchannel response waiting time to the new estimated value. As a result, in subsequent processing, a backchannel response will be output only if the silent period exceeds the new backchannel response waiting time.

[0078] The speech rate estimation unit 14 estimates the backchannel response waiting time based on the average of past silent periods, and there is a relationship between the average of past silent periods and the speech rate shown in FIG. 7. There is also a relationship between the speech rate and the backchannel response waiting time shown in FIG. 6. Therefore, it can be said that the speech rate estimation unit 14 estimates the speech rate based on the average of past silent periods and sets the backchannel response waiting time according to the estimated speech rate. However, the speech rate may be estimated by a method that does not use the average of past silent periods.

[0079] 7, there is a negative correlation between the average past silent time and the speaking rate. Therefore, it can be said that the speaking rate estimation unit 14 estimates the speaking rate so that the shorter the average past silent time, the faster the speaking rate. Similarly, it can be said that the longer the average past silent time, the slower the speaking rate.

[0080] 6, there is a negative correlation between the speech rate and the backchannel response waiting time. Therefore, it can be said that the speech rate estimation unit 14 sets the backchannel response waiting time so that the faster the speech rate, the shorter the backchannel response waiting time. Similarly, it can be said that the speech rate estimation unit 14 sets the backchannel response waiting time so that the slower the speech rate, the longer the backchannel response waiting time.

[0081] 8, there is a positive correlation between the average past silent time and the backchannel wait time. Therefore, it can be said that the speech rate estimation unit 14 sets the backchannel wait time so that the longer the average past silent time, the longer the backchannel wait time. Similarly, it can be said that the shorter the average past silent time, the shorter the backchannel wait time.

[0082] The specific relationship between the average past silent time, speech rate, and backchannel wait time may be arbitrary. For example, the graphs of Figures 7 and 8 may include a flat portion without decreasing.

[0083] After the processing of step 113, the backchannel response flag is set to OFF (step 114). Furthermore, if the silent time is less than 0.2 seconds or greater than or equal to 1 second (No in step 111), the processing of step 114 is also executed. In this case, the processing of steps 112 and 113 is not executed. That is, if the silent time is extremely short or long, the silent time is discarded without being added to the silent time history. Note that, within the scope of feasibility of the present technology, the range of values ​​related to the processing of step 111 may be different, or the processing may not be included in the flowchart.

[0084] The silent time is reset to 0 (step 115). Then, it is determined whether processing has been completed for all the segmented speech (step 116). If processing has been completed for all the segmented speech, the processing ends (Yes in step 116). If processing has not yet been completed for some segmented speech (No in step 116), processing begins for the next segmented speech (step 104 and subsequent steps).

[0085] Also, if the backchannel response flag is on in step 107 (No in step 107), the process of step 116 is executed. That is, when a continuous silent period is defined as a continuous silent period, if a backchannel response has already been output once in the continuous silent period to which the current divided audio belongs, the process of outputting the backchannel response (step 109) is not executed.

[0086] Similarly, if the silent time does not exceed the backchannel response waiting time in step 108 (No in step 108), no backchannel response is output and the process proceeds to step 116. The process also proceeds to step 116 after a backchannel response is output in step 110.

[0087] As described above, the information processing device 1 according to this embodiment estimates the speech rate, detects silent periods, and sets a backchannel response waiting time according to the speech rate. When the silent period exceeds the backchannel response waiting time, a backchannel response is output in response to the speech. This makes it possible to realize natural and smooth dialogue.

[0088] A backchannel is a response that occurs at the end of a sentence or at the end of an utterance, and is usually given 100 to 400 milliseconds after the break. However, if a dialogue agent or other such system sets a fixed timing for outputting a backchannel, the user will feel confused if the speech rate is fast, and will feel rushed if the speech rate is slow. Therefore, the delay time for outputting a backchannel must be determined according to the speech rate.

[0089] This technology sets the backchannel wait time according to the speaking speed, making it possible to output backchannels at appropriate times while the speaker is speaking, thereby enabling natural and smooth dialogue.

[0090] This technology also uses spectral envelope estimation to detect periods of silence. Although there is a technique called VAD (Voice Activity Detection) for detecting the end of speech, this technique is designed to detect entire speech segments, not short breaks within speech.

[0091] Therefore, a technology that can quickly detect breaks in speech is required. While methods such as detecting breaks from sound amplitude can be considered, they lack robustness, such as noise resistance. Furthermore, using speech recognition imposes a high processing load, and performing it on the server side increases communication time, making it difficult to implement.

[0092] On the other hand, this technology uses spectral envelope estimation, which enables it to output backchannel responses quickly and with a small processing load. As a result, it is lightweight and can be implemented on mobile devices with limitations on processing speed, power consumption, and communication capabilities.

[0093] Furthermore, with this technology, the faster the speaking speed, the shorter the waiting time for backchannel responses is set, which means that backchannel responses are output earlier when the speaking speed is fast and later when the speaking speed is slow, making it possible to maintain the speaker's speaking tempo.

[0094] This technology estimates the speaking rate based on the average of past silent periods. It is difficult to measure the speaking rate (mora rate) while the speech is being output. This technology does not directly measure the speaking rate, but estimates it using silent periods, making it possible to estimate the speaking rate with high accuracy.

[0095] With this technology, the shorter the average silent time, the faster the speech rate is estimated to be, which makes it possible to more reliably maintain the speaker's speech tempo.

[0096] This technology outputs audible responses, giving the speaker the impression that the system is listening attentively.

[0097] <Other Embodiments> The present technology is not limited to the above-described embodiments, and various other embodiments can be realized. [Application to dialogue between people] Fig. 9 is a schematic diagram showing configuration examples of information processing devices 1a and 1b that can be applied to dialogue between people.

[0098] The information processing device 1a is a device located on the side of user 2a, who is the speaker. The information processing device 1b is a device located on the side of user 2b, who is the listener who listens to the speech of user 2a. Each of the information processing devices 1a and 1b has a display unit, an operation unit, a communication unit, a storage unit, a microphone, a speaker, and a camera, all of which are not shown, and communication between the information processing devices 1a and 1b is carried out via each of the communication units. Note that the information processing devices 1a and 1b may be configured as an integrated device.

[0099] The information processing device 1a has a controller 9a, and as functional blocks, comprises a voice input unit 11, a feature extraction unit 12, a silent time detection unit 13, a speech rate estimation unit 14, a backchannel response control unit 15, and a voice output unit 16, which are similar to those in Fig. 1. Among these, the description of the contents having the same configuration or function as the embodiment in Fig. 1 may be omitted.

[0100] Furthermore, new functional blocks are configured, including a video input unit 18, a video transmission unit 19, an audio transmission unit 20, a user status receiving unit 21, an audio generation unit 22, an audio receiving unit 23, an audio synthesis unit 24, a video receiving unit 25, a video generation unit 26, and a video output unit 27.

[0101] The information processing device 1b has a controller 9b, and is configured with functional blocks including a video receiving unit 28, a video output unit 29, an audio receiving unit 30, an audio output unit 31, a video input unit 32, an audio input unit 33, a user state estimation unit 34, a user state transmission unit 35, an audio transmission unit 36, and a video transmission unit 37.

[0102] A video (moving image) of the speaking user 2a is input to the video input unit 18 of the information processing device 1a via a camera. For example, a video of the face or body of the user 2a is input to the video input unit 18. The video transmission unit 19 transmits the video of the user 2a to the information processing device 1b via a communication unit. The audio input unit 11 receives audio from the user 2a via a microphone. The audio transmission unit 20 transmits the audio from the user 2a to the information processing device 1b via the communication unit.

[0103] The video receiving unit 28 of the information processing device 1b receives a video of the user 2a via the communication unit. The video output unit 29 controls the output of the video of the user 2a to the display unit. The audio receiving unit 30 receives audio from the user 2a via the communication unit. The audio output unit 31 outputs audio from the user 2a via a speaker.

[0104] As a result, the image of the speaking user 2a is displayed to the listening user 2b, and the voice of the speaking user 2a is output.

[0105] An image of the listener user 2b is input to the image input unit 32 of the information processing device 1b via a camera. The image transmission unit 37 transmits the image of the user 2b to the information processing device 1a via a communication unit. The audio input unit 33 receives the audio of the user 2b via a microphone. The audio transmission unit 36 ​​transmits the audio of the user 2b to the information processing device 1a via the communication unit.

[0106] The user state estimation unit 34 also estimates the state of user 2b. For example, the user state estimation unit 34 estimates whether user 2b is looking at the display unit or whether user 2b is nodding in response to the speech of user 2a. These estimates are realized by the user state estimation unit 34 acquiring a video of user 2b from the video input unit 32 and acquiring the voice of user 2b from the voice input unit 33. There are no other limitations on the specific content of the estimated state of user 2b or the specific estimation method. The user state transmission unit 35 transmits the state of user 2b to the information processing device 1a via the communication unit.

[0107] The video receiving unit 25 of the information processing device 1a receives a video of the user 2b via the communication unit. The audio receiving unit 23 receives audio from the user 2b via the communication unit. The user status receiving unit 21 receives the status of the user 2b via the communication unit.

[0108] The feature extraction unit 12, the silent time detection unit 13, the speech rate estimation unit 14, and the backchannel response control unit 15 determine whether or not to output a backchannel response voice in a manner generally similar to that of the embodiment in Fig. 1. However, in this embodiment, the state of the user 2b received by the user state receiving unit 21 is also taken into consideration when determining whether or not to output a backchannel response voice.

[0109] The voice generation unit 22 generates a backchannel voice when the backchannel control unit 15 determines to output a backchannel voice. The voice synthesis unit 24 synthesizes the backchannel voice generated by the voice generation unit 22 and the voice of the user 2b received by the voice receiving unit 23. The voice output unit 16 outputs the synthesized voice. The voice generation unit 22 and the voice synthesis unit 24 correspond to an embodiment of an output unit according to the present technology.

[0110] The image generation unit 26 generates an avatar, which is a virtual character that responds with a response. In particular, in this embodiment, the image generation unit 26 generates an avatar representing the user 2b listening to a speech. For example, an image of a human-shaped avatar or an animal-shaped avatar that resembles the user 2b is generated. The image generation unit 26 acquires an image of the user 2b and generates an image of the avatar so that the avatar moves in accordance with the movements of the user 2b. For example, if the user 2b turns to the right, an image of the avatar facing right is generated. The image generation unit 26 corresponds to an embodiment of a generation unit according to the present technology.

[0111] Furthermore, in this embodiment, the backstory control unit 15 controls the output of backstory by nodding. For example, when a backstory voice is output, a backstory by nodding is output at the same time. In other words, when the silent time exceeds the backstory waiting time, it is determined that a backstory by nodding will be output. As a result, the image generation unit 26 generates an image of the avatar nodding. Specifically, an image of the avatar lightly nodding its head is generated. Furthermore, in this embodiment, the state of the user 2b received by the user state receiving unit 21 is also taken into consideration when determining whether or not to output a backstory by nodding.

[0112] The video output unit 27 outputs the video of the avatar via the display unit. As a result, the video of the avatar of the listener user 2b is displayed to the speaker user 2a. That is, the avatar is displayed to move in accordance with the movements of the user 2b, and the system also automatically displays a nod as a response. Furthermore, audio such as a response from the user 2b to the speech of the user 2a is output to the user 2a, and the system also automatically outputs a response audio.

[0113] Fig. 10 is a flowchart relating to the processing of this embodiment. In outputting backchannel responses in this embodiment, processing is generally similar to that in the flowchart shown in Fig. 2, but the processing portion for backchannel response utterance (step 109) in Fig. 2 is replaced with processing of steps 109-1 to 109-5 shown in Fig. 10. The content of the processing related to the replaced portion will be described below.

[0114] If the silent time exceeds the backchannel wait time (Yes in step 108), the backchannel control unit 15 acquires the state of the listener, user 2b (step 109-1). Next, it is determined whether user 2b is looking at the screen (display unit) (step 109-2). If user 2b is looking at the screen (Yes in step 109-2), it is determined whether user 2b is not making a backchannel response (step 109-3).

[0115] If user 2b does not nod (Yes in step 109-3), the video generation unit 26 generates a video of the avatar nodding and outputs it from the video output unit 27 (step 109-4). The audio generation unit 22 generates an audio response, which is synthesized with the audio of user 2b by the audio synthesis unit 24 and then output from the audio output unit 16 (step 109-5).

[0116] On the other hand, if user 2b is not looking at the screen (No in step 109-2), the system does not output a backslide response by voice or nodding. If a backslide response is output even though user 2b is not listening to the speech, a mismatch occurs. To prevent this, if user 2b is not looking at the screen, it is determined that user 2b is not listening to the content of the speech by user 2a, and no backslide response is output. On the other hand, if user 2b is looking at the screen, it is determined that user 2b is paying attention to the content of the speech, and a backslide response is output. In other words, the audio output unit 16 outputs a backslide response if user 2b understands the content of the speech.

[0117] Furthermore, even if user 2b is nodding (No in step 109-3), the system does not output a backspeaker response by voice or nodding. In other words, if user 2b nods, the action of user 2b takes priority, and the system does not output a backspeaker response. On the other hand, if user 2b does not nod, the system outputs a backspeaker response instead.

[0118] In this embodiment, a nod is used as a response, which allows for smoother non-verbal communication. Note that the action may be started in advance so that the lowest point of the nod coincides with the timing of the response.

[0119] Only audio backslashes may be output, and no nodding backslashes may be output. In this case, an avatar need not be generated and displayed. Conversely, only nodding backslashes may be output, and no audio backslashes may be output. For example, if it is clear that the output of backslashes will be delayed due to a communication delay or the like, outputting only nodding backslashes, which are relatively weak, can reduce the sense of discomfort. Furthermore, nodding backslashes do not hinder speech as much as audio backslashes, and therefore enable even smoother dialogue. Additionally, backslashes such as blinking may be output.

[0120] [Other Variations] In the embodiment of the person-to-person system of FIG. 1, an avatar may also be generated on the display unit.

[0121] The audio output unit 16 in Fig. 1 and the audio output unit 31 in Fig. 9 may change the content of the backspeech depending on the length of the backspeech waiting time. For example, if the backspeech waiting time is relatively long, such as 400 ms, a long backspeech such as "I see" may be output. Conversely, if the backspeech waiting time is short, a short backspeech such as "Yeah" may be output. The depth of the nod may also be changed.

[0122] Multiple types and contents of backchannel responses may be output randomly. Also, the output probability may be weighted, so that, for example, short backchannel responses are output more frequently. In addition, how the content of backchannel responses is changed may be arbitrary.

[0123] As a translation or assistance function, the conversation may be converted into text and displayed on the display unit. In this case, the timing of outputting backchannel responses may be slightly delayed to take into account the time it takes for user 2b to read the text.

[0124] 9, the backchannel response is generated by the information processing device 1a on the speaking side, but the backchannel response may be generated by the information processing device 1b on the listening side. Also, the state of the user 2b may be estimated by the information processing device 1a on the speaking side.

[0125] When users 2a and 2b use different types of devices (smartphones, VR devices, earphones, etc.), backchannel responses may be output only in modalities that can be expressed on those devices. For example, when a device without a display unit, such as earphones, is used, only audio backchannel responses are output. Furthermore, when output is possible using multiple modalities, backchannel responses may be output with priority assigned to each modal.

[0126] When the backchannel response output by the speaker differs from the state of the avatar, video, or audio output by the listener, the output content of the speaker may be fed back to the listener to match the output content of both. Conversely, the output content of the listener may be fed back to the speaker to match the output content. Matching the output content in this way makes it more susceptible to delays, etc., but these methods are also effective if the effects are tolerable. Furthermore, whether or not to match the output content may be switched depending on the situation, such as delays.

[0127] In the embodiment of the person-to-person system of Fig. 1, the backchannel response control unit 15 may change the backchannel response waiting time depending on the clarity of the utterance content. For example, if the system is unable to fully grasp the gist of the utterance, the backchannel response waiting time may be significantly changed. This causes the backchannel response to be output at a later timing, making it possible to non-verbally indicate to user 2 that the gist of the utterance is unclear.

[0128] On the other hand, if the gist of the utterance can be estimated and the utterance content on the system side has been confirmed, the backchannel wait time can be reduced, allowing the backchannel to be output at an earlier timing, making it possible to non-verbally indicate that the system has grasped the gist of the utterance.

[0129] Depending on the clarity, the speech rate and speech expression of the voice backchannel, the rate, depth, number of backchannels by nodding, etc. may be changed. Alternatively, if the clarity is low, the backchannel may not be output at all. There are no other limitations on how the content of the backchannel may be specifically changed depending on the clarity. Furthermore, the clarity may be calculated based on acoustic features and linguistic information.

[0130] When there are multiple avatars, the backchannel response control unit 15 may set the backchannel response waiting time for each avatar to a different value. This prevents multiple avatars from outputting backchannel responses at exactly the same time, realizing natural communication. Alternatively, the backchannel response waiting time may be changed randomly to create fluctuations in the timing of backchannel responses. Alternatively, a backchannel response may be output by only one of the multiple avatars.

[0131] The avatar may be assigned a character setting, such as the character's personality (impatient, laid-back, etc.), culture (Western, Asian), and age.

[0132] In this case, the backchannel response control unit 15 may change the backchannel response waiting time depending on the content of the avatar's character settings. For example, if the personality is impatient, the backchannel response waiting time is set to be short, and backchannel responses are output quickly. Conversely, if the personality is laid-back, the backchannel response waiting time is set to be long, and backchannel responses are output slowly.

[0133] Furthermore, the backchannel response control unit 15 may change the frequency of backchannel responses according to the content of the avatar's character settings. For example, if the culture is Western, the backchannel response frequency may be set low. Specifically, control may be performed so that backchannel responses are not output with a certain probability even if the silent time exceeds the backchannel response waiting time. Furthermore, if the culture is Asian, the backchannel response frequency may be set high.

[0134] When there are multiple speakers, it is possible to handle each of the multiple speakers by executing control such as identifying the current speaker using sound source separation, etc. Furthermore, based on an image of the speaker obtained via a camera, mouth movements may be detected, etc., and used as supplementary information for separation.

[0135] By using speech recognition, semantic analysis, and LLMs (Large Language Models) in combination, the present technology may be used for answering telephone calls (when conducting interviews at call centers, etc.) In this case, for example, for normal responses to spoken utterances, a response generation system such as speech recognition and LLMs is used in combination.

[0136] If a backchannel response is output by mistake due to background noise or the like, a voice message such as "I made a mistake" may be output to prompt the speaker to speak.

[0137] If the microphone is of high performance, it may be possible to detect the speaker's state of breathing. In this case, if a voice is detected without breathing, it may be determined to be a non-speech voice, and control may be exercised so that no backchannel response is output.

[0138] The information processing device 1 may be realized by a plurality of computers or by a single computer.

[0139] 11 is a block diagram showing an example of the hardware configuration of a computer 500 that can realize the information processing devices 1, 1a, and 1b. The computer 500 includes a CPU 501, a ROM 502, a RAM 503, an input / output interface 505, and a bus 504 that interconnects these components. A display unit 506, an input unit 507, a storage unit 508, a communication unit 509, a drive unit 510, and the like are connected to the input / output interface 505.

[0140] The display unit 506 is a display device using, for example, a liquid crystal display (LCD), an electroluminescence (EL), or the like. The input unit 507 is, for example, a keyboard, a pointing device, a touch panel, or other operating device. If the input unit 507 includes a touch panel, the touch panel may be integrated with the display unit 506. The storage unit 508 is a non-volatile storage device, such as a HDD, a flash memory, or other solid-state memory. The drive unit 510 is a device capable of driving a removable storage medium 511, such as an optical storage medium or a magnetic recording tape. The communication unit 509 is a modem, a router, or other communication equipment that can be connected to a LAN, a WAN, or the like and is used to communicate with other devices. The communication unit 509 may be wired or wireless. The communication unit 509 is often used separately from the computer 500.

[0141] Information processing by the computer 500 having the above-described hardware configuration is realized by cooperation between software stored in the storage unit 508 or the ROM 502, etc., and the hardware resources of the computer 500. Specifically, the information processing method according to the present technology is realized by loading a program constituting the software stored in the ROM 502, etc., into the RAM 503 and executing it.

[0142] The program is installed in the computer 500 via, for example, a removable recording medium 511. Alternatively, the program may be installed in the computer 500 via a global network or the like. Alternatively, any non-transitory storage medium that can be read by the computer 500 may be used.

[0143] In this disclosure, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0144] Execution of the information processing method according to the present technology by a computer system includes both cases where, for example, spectral envelope estimation, silent time detection, speech rate estimation, backchannel response waiting time estimation, backchannel response utterance, etc. are performed by a single computer, and cases where each process is performed by a different computer. Furthermore, execution of each process by a specific computer also includes having another computer perform part or all of the process and obtaining the results. In other words, the information processing method according to the present technology can also be applied to a cloud computing configuration in which a single function is shared and processed collaboratively by multiple devices via a network.

[0145] The information processing device and each processing flow described with reference to the drawings are merely one embodiment, and can be arbitrarily modified without departing from the spirit of the present technology. In other words, any other configuration, algorithm, etc. for implementing the present technology may be adopted.

[0146] In the present disclosure, when the term "approximately" is used, this is used merely to facilitate understanding of the description, and the use or non-use of the term "approximately" does not have any special meaning. That is, in the present disclosure, concepts defining shape, size, positional relationship, state, etc., such as "center," "central," "uniform," "equal," "same," "orthogonal," "parallel," "symmetrical," "extended," "axial direction," "cylindrical," "cylindrical," "ring-shaped," "annular," etc., are concepts that include "substantially center," "substantially central," "substantially uniform," "substantially equal," "substantially the same," "substantially orthogonal," "substantially parallel," "substantially symmetrical," "substantially extended," "substantially axial direction," "substantially cylindrical," "substantially cylindrical," "substantially ring-shaped," "substantially annular," etc. For example, states that fall within a predetermined range (e.g., a range of ±10%) based on "completely centered," "completely central," "completely uniform," "completely equal," "completely the same," "completely orthogonal," "completely parallel," "completely symmetrical," "completely extended," "completely axial direction," "completely cylindrical," "completely cylindrical," "completely ring-shaped," "completely annular," etc., are also included. Therefore, even if the word "abbreviated" is not added, the concept expressed by adding "abbreviated" may be included. Conversely, a state expressed by adding "abbreviated" does not exclude a complete state.

[0147] In the present disclosure, expressions using "than", such as "greater than A" and "smaller than A", are expressions that comprehensively include both concepts that include the case where it is equivalent to A and concepts that do not include the case where it is equivalent to A. For example, "greater than A" is not limited to cases that do not include equivalent to A, but also includes "A or greater". Furthermore, "smaller than A" is not limited to "less than A" but also includes "A or less". When implementing the present technology, specific settings and the like can be appropriately adopted from the concepts included in "greater than A" and "smaller than A" so that the effects described above can be achieved.

[0148] It is also possible to combine at least two of the features of the present technology described above. That is, the various features described in each embodiment may be arbitrarily combined without distinguishing between the embodiments. Furthermore, the various effects described above are merely examples and are not intended to be limiting, and other effects may also be achieved.

[0149] Note that the present technology can also adopt the following configurations. (1) An information processing device comprising: an estimation unit that estimates a speech rate, which is the speed of speech; a detection unit that detects a silent period, which is a time duration during which no speech is made; and an output unit that sets a time threshold according to the speech rate, and outputs a response to the speech when the silent period exceeds the threshold. (2) The information processing device according to (1), in which the detection unit detects the silent period by estimating a spectral envelope of sound related to the speech. (3) The information processing device according to (1) or (2), in which the output unit sets the threshold so that the threshold decreases as the speech rate increases. (4) The information processing device according to any one of (1) to (3), in which the estimation unit estimates the speech rate based on an average of the silent periods in the past. (5) The information processing device according to (4), wherein the estimation unit estimates the speaking rate such that the speaking rate increases as the average of the past silent periods decreases. (6) The information processing device according to any one of (1) to (5), wherein the response is a verbal backchannel. (7) The information processing device according to any one of (1) to (6), wherein the response is a nod backchannel. (8) The information processing device according to any one of (1) to (7), wherein the output unit changes the content of the response according to the magnitude of the threshold. (9) The information processing device according to any one of (1) to (8), wherein the output unit changes the magnitude of the threshold according to clarity of the content of the utterance. (10) The information processing device according to any one of (1) to (9), further comprising a generation unit that generates an avatar that is a virtual character that makes the response. (11) The information processing device according to (10), wherein, when there are a plurality of avatars, the output unit sets the threshold for each of the avatars to a different value.(12) The information processing device according to (10) or (11), wherein the output unit changes the threshold or changes the frequency of the response according to content of character settings of the avatar. (13) The information processing device according to any one of (1) to (12), wherein the output unit outputs the response when a user hearing the utterance understands the content of the utterance. (14) An information processing method comprising: estimating a speech rate that is the speed of utterance; detecting a silent period that is a duration when the utterance is not made; setting a time threshold according to the speech rate; and outputting a response to the utterance when the silent period exceeds the threshold. (15) A program causing a computer system to execute the steps of estimating a speech rate that is the speed of utterance; detecting a silent period that is a duration when the utterance is not made; and setting a time threshold according to the speech rate; and outputting a response to the utterance when the silent period exceeds the threshold.

[0150] DESCRIPTION OF SYMBOLS 1, 1a, 1b... Information processing device 2, 2a, 2b... User 12... Feature extraction unit 13... Silent speech time detection unit 14... Speech rate estimation unit 15... Backchannel response control unit 16, 31... Audio output unit 22... Audio generation unit 24... Audio synthesis unit 26... Video generation unit

Claims

1. An information processing device comprising: an estimation unit that estimates a speech rate, which is the speed of speech; a detection unit that detects silent periods, which are periods of time when no speech is being made; and an output unit that sets a time threshold according to the speech rate, and outputs a response to the speech when the silent periods exceed the threshold.

2. An information processing device according to claim 1, wherein the detection unit detects the silent period by estimating a spectral envelope of the speech related to the speech.

3. An information processing device according to claim 1, wherein the output unit sets the threshold value so that the threshold value decreases as the speech rate increases.

4. An information processing device according to claim 1, wherein the estimation unit estimates the speech rate based on an average of the silent periods in the past.

5. An information processing device according to claim 4, wherein the estimation unit estimates the speech rate such that the speech rate increases as the average of the past silent periods decreases.

6. An information processing device according to claim 1, wherein the response is a voice response.

7. An information processing device according to claim 1, wherein the response is a nod of approval.

8. An information processing device according to claim 1, wherein the output unit changes the content of the response depending on the magnitude of the threshold value.

9. An information processing device according to claim 1, wherein the output unit changes the magnitude of the threshold value depending on the clarity of the content of the utterance.

10. An information processing device according to claim 1, further comprising: a generation unit that generates an avatar, which is a virtual character that makes the response.

11. An information processing device according to claim 10, wherein, when there are a plurality of avatars, the output unit sets the threshold for each of the avatars to a different value.

12. An information processing device according to claim 10, wherein the output unit changes the threshold or the frequency of the response according to the content of the character setting of the avatar.

13. An information processing device according to claim 1, wherein the output unit outputs the response when a user listening to the utterance understands the content of the utterance.

14. An information processing method comprising: estimating a speech rate, which is the speed of speech; detecting a silent period, which is a period of time during which no speech is being made; setting a time threshold according to the speech rate; and outputting a response to the speech when the silent period exceeds the threshold.

15. A program that causes a computer system to execute the steps of: estimating a speech rate, which is the speed of speech; detecting a silent period, which is a period of time during which no speech is being made; and setting a time threshold according to the speech rate, and outputting a response to the speech when the silent period exceeds the threshold.

Citation Information

Patent Citations

  • Method of judging generating / Resting section of voice wave and device therefor

    JP1999024692A

  • Voice recognition type equipment control device

    JP2006209215A

  • Voice interaction apparatus

    JP2008026463A

  • Voice processor, voice processing method and voice processing program

    JP2020134545A

  • Interactive reception device

    JP2022045277A