Method and apparatus for interactive information

By acquiring the playback duration of voice information in smart devices and playing prompting voice messages when user silence is detected, the problem of inaccurate interaction of smart devices in silent states is solved, thus improving the accuracy and effectiveness of interaction.

CN110457437BActive Publication Date: 2025-12-30BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201910720040.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-08-06
Publication Date
2025-12-30
Estimated Expiration
2039-08-06

AI Technical Summary

Technical Problem

When intelligent robots or devices interact with users, they have difficulty accurately detecting the user's silent state, resulting in low accuracy and effectiveness of information exchange.

Method used

By obtaining the playback duration of the currently playing voice information and marking the start of the timing when the voice information stops, if no user information is detected within a set time threshold, a prompt voice message is played to improve the accuracy of the interaction.

Benefits of technology

It improves the accuracy and effectiveness of information interaction, reduces the interference of environmental noise on interaction, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110457437B_ABST
    Figure CN110457437B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method and device for interactive information, and relate to the field of cloud computing. A specific embodiment of the method includes: obtaining a playing duration of first voice information currently playing, the playing duration being a time required for playing the first voice information; in response to the first voice information playing stopping, marking a first time corresponding to the playing stopping, and starting timing from the first time as a first timing starting point; and in response to a first timing time corresponding to the first timing starting point exceeding a first set time threshold without detecting first to-be-processed information, playing second voice information. The embodiment improves the accuracy and effectiveness of information interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of data processing technology, and more specifically to methods and apparatus for interactive information. Background Technology

[0002] Chatbots or conversational programs on smart devices can interact with users in various ways (such as voice, text, and gestures) to perform functions such as retrieving information for users. When interacting with users, chatbots or conversational programs on smart devices can receive user information in real time. Summary of the Invention

[0003] Embodiments of this disclosure provide methods and apparatus for interactive information.

[0004] In a first aspect, embodiments of this disclosure provide a method for interactive information, the method comprising: obtaining the playback duration of currently playing first voice information, the playback duration being the time required to complete the playback of the first voice information; in response to the stopping of playback of the first voice information, marking a first time corresponding to the stopping of playback, and starting timing with the first time as a first timing start point; in response to the first timing time corresponding to the first timing start point exceeding a first set time threshold and no first pending information being detected, playing second voice information.

[0005] In some embodiments, the above-mentioned response to the stopping of the first voice information playback, marking the first time corresponding to the stopping of playback, and starting timing with the first time as the first timing start point, includes: in response to the completion of the first voice information playback, marking the first time corresponding to the completion of playback, and starting timing with the first time as the first timing start point.

[0006] In some embodiments, the above-mentioned response to the cessation of playback of the first voice information, marking the first time corresponding to the cessation of playback, and starting timing with the first time as the first timing start point, includes: in response to the detection of second pending information while playing the first voice information, stopping the playback of the first voice information and playing third voice information corresponding to the second pending information; in response to the completion of playback of the third voice information, marking the first time corresponding to the completion of playback, and starting timing with the first time as the first timing start point.

[0007] In some embodiments, the method further includes: in response to detecting third pending information when the first timing time corresponding to the first timing start point does not exceed the first set time threshold, playing fourth voice information corresponding to the third pending information.

[0008] In some embodiments, the method further includes: in response to the completion of the playback of the second voice information, marking a second time corresponding to the completion of playback, and starting timing with the second time as a second timing start point; in response to the second timing time corresponding to the second timing start point exceeding a second set time threshold and no fourth pending information being detected, stopping the interaction.

[0009] In some embodiments, the above-mentioned response to the fact that no fourth pending information is detected when the second timing time corresponding to the second timing start point exceeds the second set time threshold, and the stopping of interaction includes: setting a time marker for the second timing time, wherein the time marker is a time point within the second timing time corresponding to the second set time threshold; and stopping interaction in response to the existence of a continuously set number of the time markers within the second timing time.

[0010] In some embodiments, the above-mentioned response to the fact that no fourth pending information is detected when the second timing time corresponding to the second timing start point exceeds the second set time threshold, stopping the interaction, includes: obtaining the silence duration of the second voice information, the silence duration being used to mark the pending information response time after the second voice information is played; and stopping the interaction in response to the fact that the second timing time exceeds the silence duration.

[0011] In some embodiments, the method further includes: recording the duration from playing the first voice information to playing the second voice information as the interaction duration; and in response to the interaction duration exceeding a set time, playing the fifth voice information and stopping the interaction.

[0012] Secondly, embodiments of this disclosure provide an apparatus for interactive information, the apparatus comprising: a playback duration acquisition unit configured to acquire the playback duration of currently playing first voice information, wherein the playback duration is the time required to complete the playback of the first voice information; a first time marking unit configured to mark a first time corresponding to the playback stop in response to the first voice information playback stopping, and to start timing with the first time as a first timing start point; and a first voice information playback unit configured to play second voice information in response to the first timing time corresponding to the first timing start point exceeding a first set time threshold when no first pending information is detected.

[0013] In some embodiments, the first time marker unit includes a first time marker subunit, which, in response to the completion of the playback of the first voice information, is configured to mark the first time when the playback is completed, and to start timing with the first time as the first timing start point.

[0014] In some embodiments, the first time-marking unit includes: a voice information playback subunit, which, in response to detecting second pending information while playing the first voice information, stops playing the first voice information and plays third voice information corresponding to the second pending information; and a second time-marking subunit, which, in response to the completion of the playback of the third voice information, is configured to mark a first time when the playback is completed and starts timing with the first time as a first timing start point.

[0015] In some embodiments, the above-described apparatus further includes: a second voice information playback unit, which is configured to play fourth voice information corresponding to the third voice information when a third pending information is detected when a first timing time corresponding to the first timing start point does not exceed a first set time threshold.

[0016] In some embodiments, the device further includes: a second time stamping unit, configured to mark a second time corresponding to the completion of playback of the second voice information, and to start timing with the second time as a second timing start point; and a first interaction stop unit, configured to stop interaction when no fourth pending information is detected when the second timing time corresponding to the second timing start point exceeds a second set time threshold.

[0017] In some embodiments, the first interaction stop unit includes: a time marker setting subunit configured to set a time marker for the second timing period, wherein the time marker is a time point corresponding to the second set time threshold within the second timing period; and a first interaction stop subunit configured to stop interaction in response to the existence of a continuously set number of the time markers within the second timing period.

[0018] In some embodiments, the first interaction stop unit includes: a silence duration acquisition subunit, configured to acquire the silence duration of the second voice information, wherein the silence duration is used to mark the response time of pending information after the second voice information has been played; and a second interaction stop subunit, which stops the interaction in response to the second timing time exceeding the silence duration.

[0019] In some embodiments, the apparatus further includes: an interaction duration recording unit configured to record the duration from the playback of the first voice information to the playback of the second voice information as the interaction duration; and a second interaction stop unit configured to play the fifth voice information and stop the interaction in response to the interaction duration exceeding a set time.

[0020] Thirdly, embodiments of this disclosure provide an electronic device, including: one or more processors; and a memory having stored one or more programs thereon, which, when executed by the one or more processors, cause the one or more processors to perform the method for interactive information described in the first aspect.

[0021] Fourthly, embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, characterized in that the program, when executed by a processor, implements the method for interactive information described in the first aspect.

[0022] The method and apparatus for interactive information provided in the embodiments of this disclosure first obtain the playback duration of the currently playing first voice information; then, when the playback of the first voice information stops, mark the first time corresponding to the playback stop time, and start timing with the first time as the first timing start point; finally, when no first information to be processed is detected when the first timing time from the first timing start point exceeds a first set time threshold, play the second voice information. The technical solution of this application improves the accuracy and effectiveness of information interaction. Attached Figure Description

[0023] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0024] Figure 1 This is an exemplary system architecture diagram to which one embodiment of this disclosure can be applied;

[0025] Figure 2 This is a flowchart of one embodiment of a method for interactive information according to the present disclosure;

[0026] Figure 3 This is a schematic diagram of an application scenario of the method for interactive information according to this disclosure;

[0027] Figure 4 This is a flowchart of yet another embodiment of the method for interactive information according to the present disclosure;

[0028] Figure 5 This is a schematic diagram of a structure of an embodiment of a device for interactive information according to the present disclosure;

[0029] Figure 6 This is a schematic diagram of an electronic device structure suitable for implementing embodiments of the present disclosure. Detailed Implementation

[0030] The present disclosure will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.

[0031] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0032] Figure 1 An exemplary system architecture 100 for a method or apparatus for interactive information, to which embodiments of the present disclosure may be applied, is shown.

[0033] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0034] Users can interact with terminal devices 101, 102, and 103. Terminal devices 101, 102, and 103 can interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as voice acquisition applications, voice recognition applications, information search applications, voice selection applications, and voice playback applications.

[0035] Terminal devices 101, 102, and 103 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with voice playback and information processing capabilities, including but not limited to intelligent robots, smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices. They can be implemented as multiple software programs or software modules (e.g., to provide distributed services) or as a single software program or software module; no specific limitations are made here.

[0036] Server 105 can be a server that provides various services, such as a server that processes information sent by terminal devices 101, 102, and 103. The server can process the received information and send the processing results back to the terminal devices.

[0037] It should be noted that the method for interactive information provided in the embodiments of this disclosure is generally executed by terminal devices 101, 102, and 103, and correspondingly, the device for interactive information is generally disposed in terminal devices 101, 102, and 103.

[0038] It should be noted that a server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software programs or software modules (for example, used to provide distributed services), or as a single software program or software module; no specific limitations are made here.

[0039] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0040] Continue to refer to Figure 2 The diagram illustrates a flow 200 of an embodiment of a method for interactive information according to the present disclosure. The method for interactive information includes the following steps:

[0041] Step 201: Obtain the playback duration of the currently playing first audio message.

[0042] In this embodiment, the execution subject of the method for exchanging information (e.g.) Figure 1 The terminal devices 101, 102, and 103 shown can actively play the first voice information, or they can play the first voice information after receiving other signals (such as voice commands issued by the user). Accordingly, the first voice information can be a pre-set voice message that asks a question, or it can be a voice message that responds to other signals received.

[0043] In existing technologies, intelligent robots or devices typically complete one interaction with a user before initiating another based on information provided by the user. However, sometimes the user does not provide any information (i.e., the user is silent). In this case, the intelligent robot or device should be able to detect this and initiate the next interaction (e.g., ask a follow-up question). In practice, due to various environmental noises, the audio signal detected by the conversational program on the intelligent robot or device may not be the audio information emitted by the user. In this situation, the conversational program on the intelligent robot or device will still interact based on the detected audio information, resulting in low accuracy and effectiveness of information interaction.

[0044] Therefore, the executing entity of this application first obtains the playback duration of the currently playing first audio information. The playback duration is the time required to play the first audio information completely.

[0045] Step 202: In response to the stopping of the playback of the first voice information, mark the first time when the playback stops, and start timing with the first time as the first timing start point.

[0046] In practice, the playback of the first voice message may stop for various reasons (such as the ambient noise volume being higher than the volume at which the first voice message was played). In this case, the executing entity can mark the time when playback stops as the first time. That is, the first time is a point in time. To accurately determine the state of information interaction with the user, the executing entity can start timing from the first time as the first timing starting point.

[0047] In some optional implementations of this embodiment, the above-mentioned response to the stopping of the first voice information playback, marking the first time corresponding to the stopping of playback, and starting the timing with the first time as the first timing start point, may include: in response to the completion of the first voice information playback, marking the first time corresponding to the completion of playback, and starting the timing with the first time as the first timing start point.

[0048] The playback of the first voice message may stop because the executing entity stopped playing the first voice message normally. In this case, the executing entity can mark the moment corresponding to the completion of playback as the first time and start timing from the first time.

[0049] In some optional implementations of this embodiment, the above-mentioned response to the stopping of the first voice information playback, marking the first time when the playback stops, and starting the timing with the first time as the first timing start point, may further include the following steps:

[0050] The first step is to stop playing the first voice information and play the third voice information corresponding to the second voice information when the second voice information is detected during playback.

[0051] Besides the execution entity stopping playback of the first voice message after it has been played normally, user-initiated interruption can also cause the execution entity to stop playing the first voice message. That is, if the execution entity detects a second pending message while playing the first voice message, it can consider the playback of the first voice message to have been interrupted. In this case, the execution entity can stop playing the first voice message and play the third voice message corresponding to the second pending message. The second pending message can be a voice message issued by the user, text information corresponding to the user's voice message, action information, etc., depending on the specific circumstances; the third voice message is the response information corresponding to the second pending message.

[0052] The second step is to mark the first time when the third voice information is played, and start timing with the first time as the first timing start point.

[0053] When the third voice message finishes playing, the executing entity can mark the moment of completion as the first time and start timing from that first time. Similarly, the executing entity may be interrupted by the user while playing the third voice message; in this case, it will be processed in the same way as the second pending information described above.

[0054] Step 203: In response to the first timing time corresponding to the first timing start point exceeding the first set time threshold and no first pending information being detected, the second voice information is played.

[0055] If no first pending information is detected when the first timing period exceeds the first set time threshold, it indicates that the user has not sent any information within the time range corresponding to the first set time threshold. At this time, the executing entity can actively play the second voice information. This improves the accuracy and effectiveness of information interaction. The first set time threshold is used to mark the time range during which the user has not sent any information. For example, the first set time threshold can be 10 seconds. If no information is detected from the user within 10 seconds after the first voice information stops playing, the executing entity can play the second voice information. The second voice information is used to prompt the user whether to interact with information. For example, the second voice information could be: "Are you still there?" Here, the first pending information is similar to the second pending information; it can be voice information sent by the user, text information corresponding to the voice information sent by the user, action information, etc. When the first pending information detected by the executing entity is text information and / or action information, it has better accuracy and effectiveness because it avoids noise interference that may exist in voice information.

[0056] The above describes the stopping scenarios when the first voice information finishes playing normally and when playback is interrupted by the user. When the first voice information finishes playing normally, the first timing start point is the moment playback completes, and a first timing duration is obtained from this moment. This first timing duration is then compared with a first set time threshold to determine whether the first pending information has been detected. If no first pending information is detected, the executing entity can play the second voice information.

[0057] When the playback of the first voice message is interrupted, the first timing start point is the moment of interruption. A first timing period is then calculated from this moment and compared with a first set time threshold to determine whether the first pending information has been detected. For example, if the user utters the voice message "Stop" while the executing entity is playing the first voice message, the executing entity will stop playing the first voice message upon detecting the "Stop" voice message and will start timing from the moment playback stops. However, if the user does not utter any further voice messages, and no voice messages are detected even after the first timing period exceeds the first set time threshold, the executing entity can then play the second voice message. This improves the correctness and effectiveness of information interaction by the executing entity.

[0058] In some optional implementations of this embodiment, the method may further include: in response to detecting third pending information when the first timing time corresponding to the first timing start point does not exceed the first set time threshold, playing fourth voice information corresponding to the third pending information.

[0059] When the first timing period has not exceeded the first set time threshold, a third pending message is detected, indicating that the user is still in a state of information interaction with the executing entity. The third pending message is the user's response based on the first voice message. The third pending message is similar to the second pending message and will not be described again here. At this time, the executing entity can play the fourth voice message corresponding to the aforementioned third pending message to continue the interaction normally.

[0060] See also Figure 3 , Figure 3 This is a schematic diagram illustrating an application scenario of the method for interactive information according to this embodiment. Figure 3 In the application scenario, the terminal device 103 (intelligent robot) can obtain the playback duration of the first voice information when playing the first voice information; when the playback of the first voice information stops due to user interruption or other reasons, the terminal device 103 can mark the first time when the playback stops and start timing with the first time as the first timing start point; finally, when no first pending information is detected when the first timing time exceeds the first set time threshold, the terminal device 103 can play the second voice information.

[0061] The method provided in the above embodiments of this disclosure firstly obtains the playback duration of the currently playing first voice information; then, when the playback of the first voice information stops, it marks the first time corresponding to the playback stop time, and starts timing with the first time as the first timing start point; finally, when no first pending information is detected when the first timing time from the first timing start point exceeds a first set time threshold, it plays the second voice information. The technical solution of this application improves the accuracy and effectiveness of information interaction.

[0062] Further reference Figure 4 This illustrates a flow 400 of yet another embodiment of a method for interactive information. The flow 400 of this method for interactive information includes the following steps:

[0063] Step 401: Obtain the playback duration of the currently playing first audio message.

[0064] The content of step 401 is the same as that of step 201, and will not be repeated here.

[0065] Step 402: In response to the stopping of the playback of the first voice information, mark the first time when the playback stops, and start timing with the first time as the first timing start point.

[0066] The content of step 402 is the same as that of step 202, and will not be repeated here.

[0067] Step 403: In response to the first timing time corresponding to the first timing start point exceeding the first set time threshold and no first pending information being detected, the second voice information is played.

[0068] The content of step 403 is the same as that of step 203, and will not be repeated here.

[0069] Step 404: In response to the completion of the playback of the second voice information, mark the second time when the playback is completed, and start timing with the second time as the second timing start point.

[0070] As described above, the second voice message is used to prompt the user whether to engage in information interaction. To further determine whether the user is in an information interaction state, the executing entity can mark the time when the second voice message finishes playing as the second time, and start timing from the second time as the second timing start point.

[0071] Step 405: In response to the fact that no fourth pending information is detected when the second timing time corresponding to the second timing start point exceeds the second set time threshold, the interaction stops.

[0072] If the executing entity does not detect the fourth pending information when the second timeout exceeds the second set time threshold, it can be assumed that the user is in a non-interaction state (e.g., the user has left). The fourth pending information can be a further confirmation of whether the user is in a non-interaction state. For example, the fourth pending information could be: "If it's inconvenient for you to communicate now, let's talk another time." In this case, the executing entity can stop the interaction. If the executing entity detects the fourth pending information when the second timeout has not exceeded the second set time threshold, it processes the fourth pending information in the same way as the third pending information. The fourth pending information is similar to the second pending information and will not be described again here.

[0073] In some optional implementations of this embodiment, stopping the interaction when no fourth pending information is detected after the second timing time corresponding to the second timing start point exceeds the second set time threshold may include the following steps:

[0074] The first step is to set a time marker for the second timing period mentioned above.

[0075] In practice, users may fail to respond within the second set time threshold for various reasons. To improve the effectiveness of information interaction between the executing entity and the user, the executing entity can set time markers for the aforementioned second timing period. These time markers can be the time points within the second timing period corresponding to the aforementioned second set time threshold. For example, if the second set time threshold is 5 seconds and the second timing period is 16 seconds, then the time markers can be set at the 5th, 10th, and 15th seconds of the second timing period, respectively.

[0076] The second step is to stop the interaction if a set number of the above-mentioned time markers are continuously present within the second timing period.

[0077] The more time markers there are, the longer the user has not responded, and the greater the likelihood that the user is in a non-interactive state. The executing entity can count the number of time markers; when a set number of time markers appear consecutively, it is highly probable that the user is in a non-interactive state. At this point, the executing entity can stop the interaction.

[0078] In some optional implementations of this embodiment, stopping the interaction when no fourth pending information is detected after the second timing time corresponding to the second timing start point exceeds the second set time threshold may include the following steps:

[0079] The first step is to obtain the silence duration of the second voice information mentioned above.

[0080] As described above, the second voice message is used to prompt the user whether to engage in information interaction. The second voice message can be something like, "Are you still there?" or "What are your thoughts on what I just said?". When the second voice message is "Are you still there?", if the user is in an information interaction state, since the user doesn't need to think too much about the second voice message, the executing entity can set a shorter silence duration for "Are you still there?" (e.g., 3 seconds). When the second voice message is "What are your thoughts on what I just said?", if the user is in an information interaction state, since the user needs to think about the second voice message to reach a conclusion, the executing entity can set a longer silence duration for "What are your thoughts on what I just said?" (e.g., 10 seconds). That is, the executing entity can set different silence durations for different second voice messages. The aforementioned silence duration can be used to mark the response time for pending information after the second voice message has finished playing.

[0081] The second step is to stop the interaction when the second timing period exceeds the aforementioned silence duration.

[0082] When the second timing period exceeds the aforementioned silence duration, the user can be considered to be in a non-interactive state, and the executing entity can stop the interaction.

[0083] In some optional implementations of this embodiment, the above method may further include the following steps:

[0084] The first step is to record the duration from the playback of the first voice message to the playback of the second voice message as the interaction duration.

[0085] In practice, the duration of a dialogue between the executing entity and the user is uncertain. Generally, the longer the interaction lasts, the less effective it becomes (e.g., the user becomes impatient, or it consumes too much of the user's time). Therefore, the executing entity can record the time from playing the first voice message to playing the second voice message as the interaction duration. That is, the interaction duration is the total time from the start of the interaction to the prompt for further interaction.

[0086] The second step is to play the fifth voice message and stop the interaction if the interaction duration exceeds the set time.

[0087] When the interaction duration exceeds the set time, the executing entity can proactively play a fifth voice message and stop the interaction. This fifth voice message is the message from the executing entity that actively requests the end of the interaction. For example, the fifth voice message could be: "Thank you for your assistance. We've taken up so much of your time. We can talk again another time." This increases the flexibility of the interaction between the executing entity and the user, improving the user experience.

[0088] Further reference Figure 5 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a device for interactive information, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0089] like Figure 5 As shown, the device 500 for interactive information in this embodiment may include: a playback duration acquisition unit 501, a first time marker unit 502, and a first voice information playback unit 503. The playback duration acquisition unit 501 is configured to acquire the playback duration of the currently playing first voice information, where the playback duration is the time required to complete the playback of the first voice information. The first time marker unit 502, in response to the stopping of playback of the first voice information, is configured to mark a first time corresponding to the stopping of playback, and to start timing with the first time as a first timing start point. The first voice information playback unit 503, in response to the first timing time corresponding to the first timing start point exceeding a first set time threshold without detecting the first pending information, is configured to play second voice information.

[0090] In some optional implementations of this embodiment, the first time marker unit 502 may include: a first time marker subunit (not shown in the figure), which is configured to mark the first time when the first voice information is played and to start timing with the first time as the first timing start point in response to the completion of the first voice information playback.

[0091] In some optional implementations of this embodiment, the first time-marking unit 502 may include: a voice information playback subunit (not shown in the figure) and a second time-marking subunit (not shown in the figure). The voice information playback subunit, in response to detecting second pending information while playing the first voice information, stops playing the first voice information and plays third voice information corresponding to the second pending information. The second time-marking subunit, in response to the completion of playback of the third voice information, is configured to mark a first time when playback is complete, and starts timing with the first time as a first timing start point.

[0092] In some optional implementations of this embodiment, the device 500 for interactive information may further include: a second voice information playback unit (not shown in the figure), which is configured to play fourth voice information corresponding to the third voice information when a third pending information is detected when the first timing time corresponding to the first timing start point does not exceed a first set time threshold.

[0093] In some optional implementations of this embodiment, the device 500 for interactive information may further include: a second time stamping unit (not shown in the figure) and a first interaction stop unit (not shown in the figure). The second time stamping unit is configured to mark a second time corresponding to the completion of playback of the second voice information, and to start timing with the second time as a second timing start point; the first interaction stop unit is configured to stop interaction when no fourth pending information is detected when the second timing time corresponding to the second timing start point exceeds a second set time threshold.

[0094] In some optional implementations of this embodiment, the first interaction stop unit may include: a time marker setting subunit (not shown in the figure) and a first interaction stop subunit (not shown in the figure). The time marker setting subunit is configured to set a time marker for the second timing period, where the time marker is a time point corresponding to the second set time threshold within the second timing period. The first interaction stop subunit is configured to stop interaction in response to the existence of a consecutive set number of time markers within the second timing period.

[0095] In some optional implementations of this embodiment, the first interaction stop unit may include: a silence duration acquisition subunit (not shown in the figure) and a second interaction stop subunit (not shown in the figure). The silence duration acquisition subunit is configured to acquire the silence duration of the second voice information, which is used to mark the response time of pending information after the second voice information has been played. The second interaction stop subunit stops the interaction in response to the second timing time exceeding the silence duration.

[0096] In some optional implementations of this embodiment, the device 500 for interactive information may further include: an interaction duration recording unit (not shown in the figure) and a second interaction stop unit (not shown in the figure). The interaction duration recording unit is configured to record the duration from the playback of the first voice information to the playback of the second voice information as the interaction duration; the second interaction stop unit, in response to the interaction duration exceeding a set time, plays the fifth voice information and stops the interaction.

[0097] This embodiment also provides an electronic device, including: one or more processors; and a memory storing one or more programs thereon, which, when executed by the one or more processors, cause the one or more processors to perform the above-described method for interactive information.

[0098] This embodiment also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method for interactive information.

[0099] The following is for reference. Figure 6 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 1 A schematic diagram of the structure of a computer system 600 (server 105). Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0100] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0101] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 6 Each box shown can represent a device or multiple devices as needed.

[0102] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by a processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0103] It should be noted that the computer-readable medium described in the embodiments of this disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0104] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire the playback duration of currently playing first audio information, the playback duration being the time required to complete playing the first audio information; in response to the stopping of playback of the first audio information, mark a first time corresponding to the playback stopping, and start timing with the first time as a first timing start point; in response to no detection of first pending information when the first timing time corresponding to the first timing start point exceeds a first set time threshold, play second audio information.

[0105] Computer program code for performing the operations of embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0107] The units described in the embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a playback duration acquisition unit, a first time stamp unit, and a first voice information playback unit. The names of these units do not necessarily limit the specific unit; for example, the first voice information playback unit may also be described as "a unit for playing second voice information."

[0108] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

Claims

1. A method for interacting information, comprising: obtaining a playing duration of a first voice information currently playing, the playing duration being a time required for playing the first voice information completely; in response to the first voice information playing being stopped, marking a first time corresponding to the playing being stopped, and starting a first counting from the first time as a first counting start point; in response to no first to-be-processed information being detected when a first counting time corresponding to the first counting start point exceeds a first set time threshold, playing a second voice information; determining, according to the second voice information, a silence duration for representing a to-be-processed information response time after the second voice information is played; in response to no fourth to-be-processed information being detected when a second counting time from the second voice information playing being completed exceeds the silence duration, stopping the interaction; recording a duration from the first voice information being played to the second voice information being played as an interaction duration, wherein the interaction duration is a total duration of multiple rounds of interaction between starting to interact with a user and prompting the user whether to continue the interaction; in response to the interaction duration exceeding a set time, playing a fifth voice information, and stopping the interaction.

2. The method of claim 1, wherein, The response to the first voice information playing being stopped, the first time corresponding to the playing being stopped is marked, and the first time is taken as the first counting start point to start the counting, comprising: in response to the first voice information playing being completed, the first time corresponding to the playing being completed is marked, and the first time is taken as the first counting start point to start the counting.

3. The method of claim 1, wherein, The response to the first voice information playing being stopped, the first time corresponding to the playing being stopped is marked, and the first counting start point is started, comprising: in response to the second to-be-processed information being detected when the first voice information is played, the first voice information is stopped, and a third voice information corresponding to the second to-be-processed information is played; in response to the third voice information playing being completed, the first time corresponding to the playing being completed is marked, and the first counting start point is started.

4. The method of claim 1, wherein, The method further comprises: in response to the third to-be-processed information being detected when the first counting time corresponding to the first counting start point does not exceed the first set time threshold, a fourth voice information corresponding to the third to-be-processed information is played.

5. The method of claim 1, wherein, The method further comprises: in response to the second voice information playing being completed, a second time corresponding to the playing being completed is marked, and a second counting is started from the second time as a second counting start point; in response to no fourth to-be-processed information being detected when a second counting time corresponding to the second counting start point exceeds a second set time threshold, the interaction is stopped.

6. The method of claim 5, wherein, The response to no fourth to-be-processed information being detected when the second counting time corresponding to the second counting start point exceeds the second set time threshold, the interaction is stopped, comprising: setting a time mark point for the second counting time, the time mark point being a time point corresponding to the second set time threshold in the second counting time; in response to a continuous set number of the time mark points existing in the second counting time, the interaction is stopped. 7.A device for interacting information, comprising: The playing duration acquisition unit is configured to acquire a playing duration of the currently played first voice information, the playing duration being a time required for playing the first voice information completely; The first time marking unit is configured to, in response to the first voice information playing being stopped, mark a first time corresponding to the playing being stopped and start timing from the first time as a first timing starting point; The first voice information playing unit is configured to, in response to a first timing time corresponding to the first timing starting point not being detected to exceed a first set time threshold, play second voice information; The interactive duration recording unit is configured to record a duration from playing the first voice information to playing the second voice information as an interactive duration, wherein the interactive duration is a total duration of multiple rounds of interactions from starting the interaction with the user to prompting the user whether to continue the interaction; The second interactive stopping unit is configured to, in response to the interactive duration exceeding a set time, play fifth voice information and stop the interaction; The device further comprises a first interactive stopping unit, comprising: The silence duration acquisition sub-unit is configured to determine, according to the second voice information, a silence duration for representing a response time of the to-be-processed information after the second voice information is played; The second interactive stopping sub-unit is configured to, in response to a second timing time from the second voice information being played completely exceeding the silence duration without detecting fourth to-be-processed information, stop the interaction.

8. The apparatus of claim 7, wherein, The first time marking unit comprises: The first time marking sub-unit is configured to, in response to the first voice information being played completely, mark a first time corresponding to the playing being completed and start timing from the first time as a first timing starting point.

9. The apparatus of claim 7, wherein, The first time marking unit comprises: The voice information playing sub-unit is configured to, in response to detecting second to-be-processed information when the first voice information is played, stop playing the first voice information and play third voice information corresponding to the second to-be-processed information; The second time marking sub-unit is configured to, in response to the third voice information being played completely, mark a first time corresponding to the playing being completed and start timing from the first time as a first timing starting point.

10. The apparatus of claim 7, wherein, The device further comprises: The second voice information playing unit is configured to, in response to a first timing time corresponding to the first timing starting point exceeding a first set time threshold without detecting third to-be-processed information, play fourth voice information corresponding to the third to-be-processed information.

11. The apparatus of claim 7, wherein, The device further comprises: The second time marking unit is configured to, in response to the second voice information being played completely, mark a second time corresponding to the playing being completed and start timing from the second time as a second timing starting point; The first interactive stopping unit is configured to, in response to a second timing time corresponding to the second timing starting point exceeding a second set time threshold without detecting fourth to-be-processed information, stop the interaction.

12. The apparatus of claim 11, wherein, The first interactive stopping unit comprises: The time marking point setting sub-unit is configured to set a time marking point for the second timing time, the time marking point being a time point corresponding to the second set time threshold in the second timing time; A first interaction stopping subunit, in response to the existence of a continuous set number of the time mark points within the second timing time, is configured to stop interaction. 13.An electronic device, comprising: one or more processors; memory storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors perform the method of any one of claims 1-6.

14. A computer readable medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Man-machine voice interaction method and system

    CN105070290A

  • Reminding method and terminal

    CN107172297A

  • Voice interaction method and device based on external trigger

    CN109741740A