Dialog processing method and apparatus, and electronic device, storage medium and program product
By using the recovery model to analyze the timestamp information input by the user's voice in an intelligent voice dialogue system, the intent detection problem when the user interrupts the machine's voice is solved, and the smooth and coherent recovery of voice interaction is achieved, improving the user experience.
Patent Information
- Application Number
- PCT/CN2024/114902
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-20
- Filing Date
- 2024-08-27
- Publication Date
- 2025-07-10
AI Technical Summary
The intelligent voice dialogue system cannot effectively detect the true intention when the user interrupts the machine's voice, resulting in poor voice interaction effects, lag and unnatural playback phenomena.
Through the recovery model, the timestamp information of user voice input is analyzed, the recovery playback node of the machine voice is determined, the machine voice volume is stopped or reduced, and the playback is restored in a suitable location, combining the voice characteristics and context data to make user intention judgments.
It improves the fluency and anthropomorphic experience of voice interaction, reduces the abruptness and lag in the user experience, and ensures the coherence and nature of voice recovery.
Smart Images

Figure CN2024114902_10072025_PF_FP_ABST
Abstract
Description
Dialogue processing method, device, electronic device, storage medium and program product
[0001] This disclosure is based on the Chinese patent application with application number 202311369891.3, application date October 20, 2023, and invention name “Dialogue processing method, device, electronic device and computer storage medium”, and claims the priority of the Chinese patent application. The entire content of the Chinese patent application is hereby introduced into this disclosure as a reference. Technical Field
[0002] The present disclosure relates to the field of speech processing technology, and in particular to a conversation processing method, device, electronic device, computer-readable storage medium, and computer program product. Background Art
[0003] As intelligent voice dialogue systems become more and more widely used, the intelligent experience of intelligent voice dialogue is becoming an increasingly important indicator of system performance. Whether users can effectively interrupt the machine is one of the most effective ways to improve the experience.
[0004] In related technologies, after the intelligent voice dialogue system recognizes the user's input, it cannot detect whether the user has a real intention to interrupt the voice. The user may just agree casually, or there may be people talking in the background, and the system will stop broadcasting and there will be a noticeable abruptness. The system also has obvious lag. Moreover, the timing of the system resuming broadcasting is uncertain. It may start to resume from the middle of a word. For example, when saying "we", the system starts broadcasting from the word "we". In worse cases, it may start broadcasting from the second half of the syllable of "we", resulting in poor voice interaction effect and poor user experience.
[0005] Summary of the Invention
[0006] The present disclosure provides a conversation processing method, device, electronic device, computer-readable storage medium, and computer program product, which at least to some extent overcome the problem of poor voice interaction effect in related technologies.
[0007] According to one aspect of the present disclosure, a conversation processing method is provided, comprising: when a user voice input is received while machine voice is being played, and the user has no real intention to interrupt the machine voice, stopping the playing of the machine voice, and obtaining the timestamp information corresponding to the machine voice at the time of the user voice input; determining the resume playing node corresponding to the machine voice based on the timestamp information; and playing the machine voice from the resume playing node.
[0008] In one embodiment of the present disclosure, when a user's voice is input while the machine voice is playing, and the user has no real intention to interrupt the machine voice, the machine voice is stopped from playing, and the timestamp information corresponding to the machine voice at the time of the user's voice input is obtained, including: inputting the user voice and the machine voice into a recovery model, so as to determine through the recovery model whether the user has a real intention to interrupt the machine voice.
[0009] In one embodiment of the present disclosure, the training data of the recovery model includes: conversation audio feature data, conversation context data, conversation turn data, and / or conversation semantic data.
[0010] In one embodiment of the present disclosure, the inputting of the user voice and the machine voice into the recovery model so as to determine through the recovery model whether the user has a real intention to interrupt the machine voice includes: the recovery model determines whether the user has a real intention to interrupt the machine voice based on voice data, the user voice and the machine voice; wherein, the voice data includes at least one of the following: guiding speech content data, playback progress data of the machine voice, and voice confidence data.
[0011] In one embodiment of the present disclosure, the resuming playback node includes at least one of the following:
[0012] The starting position of the machine voice;
[0013] The position of the machine voice at the time of the user voice input;
[0014] The corresponding segmentation position of the machine voice when the user voice is input;
[0015] The machine voice corresponds to a preset pause position when the user voice is input.
[0016] In one embodiment of the present disclosure, the machine voice includes: voice audio, text corresponding to the voice audio, and / or timestamp information corresponding to the voice audio, wherein the timestamp information includes: punctuation marks, sentence positions, and / or preset pause positions.
[0017] In one embodiment of the present disclosure, when a user's voice input is received while the machine voice is playing, and the user has no real intention to interrupt the machine voice, the playing of the machine voice is stopped, and the timestamp information corresponding to the machine voice at the time of the user's voice input is obtained. The method includes: when a user's voice input is received while the machine voice is playing, and the user has no real intention to interrupt the machine voice, the playing of the machine voice is stopped immediately or the volume of the machine voice is slowly reduced.
[0018] In one embodiment of the present disclosure, the method further includes: when the pause of the user voice exceeds a first time threshold, indicating that the user voice input is completed.
[0019] In one embodiment of the present disclosure, the method further includes: when the user has a real intention to interrupt the machine voice, generating a reply voice corresponding to the user voice.
[0020] In one embodiment of the present disclosure, the method further includes filtering silence data and noise data in the user voice.
[0021] According to another aspect of the present disclosure, there is also provided a conversation processing device, comprising:
[0022] an acquisition module, which, when a user's voice input is received while the machine voice is being played and the user has no real intention to interrupt the machine voice, stops playing the machine voice and acquires the timestamp information of the machine voice corresponding to the time of the user's voice input;
[0023] A determination module, which determines a resume playback node corresponding to the machine voice according to the timestamp information;
[0024] A recovery module plays the machine voice from the recovery playback node.
[0025] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any one of the above-mentioned conversation processing methods by executing the executable instructions.
[0026] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the computer program implements any one of the above-mentioned dialogue processing methods.
[0027] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein the computer program implements the above-mentioned dialog processing method when executed by a processor.
[0028] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0030] FIG1 shows a flow chart of a method for processing a conversation in an embodiment of the present disclosure;
[0031] FIG2 shows a flow chart of a method for determining user intention in an embodiment of the present disclosure;
[0032] FIG3 shows a flow chart of another method for processing a conversation in an embodiment of the present disclosure;
[0033] FIG4 shows a schematic diagram of a conversation recovery according to an embodiment of the present disclosure;
[0034] FIG5 shows a schematic diagram of a conversation processing device according to an embodiment of the present disclosure;
[0035] FIG6 shows a schematic diagram of a conversation processing system according to an embodiment of the present disclosure;
[0036] FIG7 is a schematic diagram showing an exemplary system architecture that can be applied to the dialogue processing method or dialogue processing device of the embodiment of the present disclosure; and
[0037] FIG8 shows a structural block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0038] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0039] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0040] This exemplary implementation is described in detail below with reference to the accompanying drawings and examples.
[0041] First, an embodiment of the present disclosure provides a method for processing a conversation, which can be executed by any electronic device with computing and processing capabilities.
[0042] FIG1 shows a flow chart of a method for processing a conversation in an embodiment of the present disclosure. As shown in FIG1 , the method for processing a conversation provided in an embodiment of the present disclosure includes the following steps:
[0043] S102: When a user voice input is received while the machine voice is being played, and the user has no real intention to interrupt the machine voice, the machine voice is stopped from being played and the timestamp information corresponding to the user voice input is obtained.
[0044] In one embodiment, machine speech includes but is not limited to: speech audio, text corresponding to the speech audio, and / or timestamp information corresponding to the speech audio, etc., and the corresponding text, timestamp information, etc. are generated synchronously when synthesizing the speech audio; wherein, the timestamp information includes but is not limited to: the position of interrupting the playing speech, punctuation marks, sentence punctuation position, and / or preset pause position, the sentence punctuation position can be the position where the sentence punctuation is required determined according to semantics, etc., and the preset pause position is the position where the pause is required set manually or automatically.
[0045] In one embodiment, data such as silence data and noise data in the user's voice is filtered.
[0046] In one embodiment, the user voice and the machine voice are input into the recovery model so as to determine whether the user has a real intention to interrupt the machine voice through the recovery model; the lack of real intention to interrupt the machine voice includes but is not limited to: the user interrupts by mistake, the user does not change the content of the conversation, etc.
[0047] In one embodiment, the user voice and the machine voice after filtering out the silent data and the noise data are input into the recovery model so as to determine whether the user has a real intention to interrupt the machine voice through the recovery model.
[0048] In one embodiment, when there is user voice input while the machine voice is playing and the user has no real intention to interrupt the machine voice, the machine voice is stopped immediately, or the volume of the machine voice is slowly reduced until the pause of the user voice exceeds a first time threshold, indicating that the user voice input is completed.
[0049] In one embodiment, selection and configuration are made based on the usage scenario to immediately stop the machine voice broadcast behavior or reduce the volume of the machine voice broadcast.
[0050] In one embodiment, when the pause of the user's voice exceeds a first time threshold, it indicates that the user's voice input is completed. The first time threshold can be set automatically or manually according to the application scenario, for example, the first time threshold can be set to 600ms-800ms.
[0051] In one embodiment, the recovery model determines whether the user has a genuine intention to interrupt the machine speech based on the speech data, the user speech, and the machine speech;
[0052] Among them, voice data includes but is not limited to at least one of the following: guiding speech content data, machine voice playback progress data, and voice confidence data; guiding speech content data is the data guiding user interaction included in the machine voice; voice confidence data is the recognition confidence corresponding to the user voice and the machine voice.
[0053] S104: Determine the resume playback node corresponding to the machine voice according to the timestamp information.
[0054] In one embodiment, the resume playback node includes but is not limited to at least one of the following:
[0055] The starting position of the machine speech;
[0056] The position of the machine voice during user voice input;
[0057] The corresponding sentence segmentation position of the machine voice during user voice input;
[0058] The preset pause positions of machine speech corresponding to user voice input.
[0059] In one embodiment, the resumption playback node can be set automatically or manually according to the scenario, etc. For example, the resumption playback node can be set to the corresponding punctuation position of the machine voice when the user voice is input, and a mapping table between the timestamp and the resumption playback node is established, and the corresponding resumption playback node is determined to be the punctuation position according to the timestamp information; the resumption playback node can be set to the corresponding preset pause position of the machine voice when the user voice is input, and a mapping table between the timestamp and the resumption playback node is established, and the corresponding resumption playback node is determined to be the preset pause position B according to the timestamp information A.
[0060] S106, playing the machine voice from the resumed playing node.
[0061] In one embodiment, after the first round of user voice input is completed, the user voice of the next round may be further detected and collected to further verify the user voice input by the first round of users.
[0062] In one embodiment, when resuming playback of machine voice, the volume of the machine voice can be directly restored, or it can be smoothly restored using a fade-in and fade-out method to ensure smooth and coherent voice restoration. The method of restoring the volume and the smooth transition time can be automatically or manually configured according to needs and scenarios.
[0063] In one embodiment, when the user has a real intention to interrupt the machine voice, a reply voice corresponding to the user voice is generated, including but not limited to: "Hmm", "Got it", etc., and then the system dialogue corresponding to the real intention is broadcast according to the user voice.
[0064] In the above embodiment, based on the recovery model, the user intention is confirmed, and the serious problem of accidental interruption in the intelligent voice conversation is corrected. The timestamp is maintained as the context of the voice conversation. When processing the user interruption, the timestamp corresponding to the user's voice input when the machine voice is played is recorded. A variety of recovery strategies can be selected through the timestamp, and the appropriate recovery playback node can be determined to ensure the smoothness and coherence of the voice recovery. It can not only flexibly interrupt, but also detect accidental interruptions in time, and can restore the machine voice in a friendly manner, thereby improving the humanized experience of the intelligent voice conversation.
[0065] FIG2 shows a flow chart of a method for determining user intention in an embodiment of the present disclosure. As shown in FIG2 , the method for determining user intention provided in an embodiment of the present disclosure includes the following steps:
[0066] S202: Training the recovery model.
[0067] In one embodiment, the training data of the recovery model includes but is not limited to: conversation audio feature data, conversation context data, conversation turn data, voice data, and / or conversation semantic data; conversation audio features include but are not limited to short-time zero-crossing rate, short-time energy, short-time autocorrelation function, short-time average amplitude, spectral difference amplitude, spectral centroid, spectral width, Mel-frequency cepstral coefficients, etc.; conversation context is contextual information of the interaction between machine voice and user voice, etc.; conversation turn data includes but is not limited to: turn data such as sending a signal, confirming the end, and formally saying goodbye; conversation semantic data is the meaning of the concept represented by the real-world objects corresponding to the data, etc.
[0068] S204: Input the user voice and the machine voice into the recovery model.
[0069] In one embodiment, the user voice and the machine voice are input into the recovery model so as to determine whether the user has a real intention to interrupt the machine voice through the recovery model; the lack of real intention to interrupt the machine voice includes but is not limited to: the user interrupts by mistake, the user does not change the content of the conversation, etc.
[0070] In one embodiment, the user voice and the machine voice after filtering out the silent data and the noise data are input into the recovery model so as to determine whether the user has a real intention to interrupt the machine voice through the recovery model.
[0071] In one embodiment, the recovery model determines whether the user has a genuine intention to interrupt the machine voice based on voice data, user voice, and machine voice.
[0072] Among them, voice data includes but is not limited to at least one of the following: guiding speech content data, machine voice playback progress data, and voice confidence data; guiding speech content data is the data guiding user interaction included in the machine voice; voice confidence data is the recognition confidence corresponding to the user voice and the machine voice.
[0073] S206: Determine whether to resume playing the machine voice based on the recovery model judgment result.
[0074] One of the recovery model's judgment results is that the user's intention to interrupt is confirmed. At this time, the current round of conversation ends, and the user's voice is fully semantically analyzed and processed to generate new words for further interaction.
[0075] Another judgment result of the recovery model is that the user's intention to interrupt is denied. In this case, it is necessary to resume the current round of dialogue, scroll the machine voice forward to the appropriate resumption playback node, and resume the playback of the machine voice at the resumption playback node.
[0076] In the above embodiment, the user intention is confirmed based on the recovery model, and according to the confirmation result, it is determined whether to further perform intention recognition or resume the conversation, so as to correct the serious problem of accidental interruption in the intelligent voice conversation. When resuming the conversation, the method of adding timestamps is used to scroll forward from the interruption position to the appropriate resumption playback node. This strategy makes the sound smoother and more coherent, the expression clearer, and closer to the real human conversation experience.
[0077] FIG3 shows a flow chart of another method for processing a conversation in an embodiment of the present disclosure. As shown in FIG3 , the method for processing a conversation provided in an embodiment of the present disclosure includes the following steps:
[0078] S302, synthesize and play machine voice.
[0079] In one embodiment, the machine speech is synthesized by a speech synthesis engine, or it can be a pre-made recording; when the machine speech audio is generated, a timestamp of the corresponding text is generated synchronously, and is maintained as a context by the central control module.
[0080] S304: When it is detected that a user voice input is present during the machine voice playback, the timestamp information corresponding to the machine voice at the time of the user voice input is obtained.
[0081] The interruption function can be activated in combination with the voice endpoint detection module and / or noise reduction module in the central control module. After detecting the user's voice input, the central control module can immediately interrupt or lower the volume of the machine voice, and record the interruption timestamp corresponding to the user's voice input in the context. The interruption position is recorded by the timestamp, which is used to restore the interruption position during recovery.
[0082] In one embodiment, user voice data is processed in real time, and a voice endpoint detection module and a noise reduction module are called respectively to filter the user voice data. If the voice endpoint detection module detects that the user voice data contains sound, the noise reduction module will further detect it. If the sound contains noise, it will be filtered out. If it does not contain noise, it will be determined to be user input, thereby improving the accuracy of the user voice.
[0083] The voice endpoint detection module and noise reduction module are locally integrated lightweight modules with relatively fast response speed. Using the voice endpoint detection module and noise reduction module to handle interruptions can respond to users more quickly, but their accuracy is easily affected by background noise. Therefore, after the user finishes inputting user voice, sufficient information needs to be further collected to further verify the user input.
[0084] S306: Detect whether the user has a real intention to interrupt the machine voice through the recovery model.
[0085] In one embodiment, the recovery model combines contextual information, speech and semantic information to detect whether the user has a real intention to interrupt the machine speech.
[0086] S308: If yes, then this round of dialogue ends, and the central control module calls the semantic recognition engine and dialogue management service to perform complete semantic analysis and dialogue processing, and generate new words for further interaction.
[0087] S310, otherwise, resume the current round of dialogue, the central control module integrates the timestamp of the speech synthesis engine, scrolls the speech forward to a suitable pause position, and resumes the playback of the machine speech at the pause position.
[0088] In the above embodiment, a solution for rapid interruption, intention confirmation, and error recovery is provided. Rapid voice breakpoint detection is used to timely interrupt machine voice. After the user finishes speaking, a rapid recovery model is used for confirmation. Based on the confirmation result, it is determined whether to further perform intention recognition or resume the conversation. When resuming the conversation, the method of adding timestamps is used to scroll forward from the interruption position to the appropriate resumption playback node. This strategy produces a smoother and more coherent voice, clearer expression, and is closer to a real human-to-human conversation experience.
[0089] Figure 4 shows a schematic diagram of conversation recovery in an embodiment of the present disclosure, where the coordinate axis represents the direction of real-time generation of machine voice, and each square represents a time slice. When the machine voice asks the user's name, the user may immediately confirm when hearing his or her name. At this time, the machine voice will be immediately interrupted or the volume of the machine voice will be slowly lowered; for example, the machine voice may be: "Yes, are you Mr. Wang?" The user responds when the voice plays to "Wang", which corresponds to the user voice input position.
[0090] If the user replies "yes", the recovery model will combine the guidance content data, playback progress data, voice confidence and other parameters to determine that the user has the real intention to interrupt the machine voice.
[0091] If the user's reply is "Ah~", the recovery model will determine that the user has no real intention to interrupt the machine voice. At this time, the central control module rolls back the voice to the recovery playback node position "you" based on the timestamp information, and repeats the guidance from here.
[0092] In the above embodiment, scrolling forward from the interruption position to the appropriate resumption playback node, this strategy makes the sound smoother and more coherent, the expression clearer, and closer to the real person-to-person conversation experience, solving the problem of accidental interruption recovery in the intelligent voice dialogue system.
[0093] Based on the same inventive concept, the present disclosure also provides a conversation processing device, such as the following embodiment. Since the principle of solving the problem in the device embodiment is similar to that in the above method embodiment, the implementation of the device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.
[0094] FIG5 shows a schematic diagram of a conversation processing device according to an embodiment of the present disclosure. As shown in FIG5 , the conversation processing device 5 includes: an acquisition module 501 , a determination module 502 , and a recovery module 503 ;
[0095] Acquisition module 501, when a user voice input is received during machine voice playback and the user has no real intention to interrupt the machine voice, stops playing the machine voice and obtains the timestamp information corresponding to the machine voice input;
[0096] Determining module 502, determining the resume playback node corresponding to the machine voice according to the timestamp information;
[0097] The recovery module 503 plays the machine voice from the recovery playback node.
[0098] In the above embodiment, based on the recovery model, the user intention is confirmed, and the serious problem of accidental interruption in the intelligent voice conversation is corrected. The timestamp is maintained as the context of the voice conversation. When processing the user interruption, the timestamp corresponding to the user's voice input when the machine voice is played is recorded. A variety of recovery strategies can be selected through the timestamp, and the appropriate recovery playback node can be determined to ensure the smoothness and coherence of the voice recovery. It can not only flexibly interrupt, but also detect accidental interruptions in time, and can restore the machine voice in a friendly manner, thereby improving the humanized experience of the intelligent voice conversation.
[0099] Based on the same inventive concept, the present disclosure also provides a conversation processing system, such as the following embodiment. Since the principle of solving the problem in the system embodiment is similar to that in the above method embodiment, the implementation of the system embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.
[0100] FIG6 shows a schematic diagram of a dialogue processing system according to an embodiment of the present disclosure. As shown in FIG6 , the system includes: a central control module 601, a speech synthesis engine 602; a speech recognition engine 603; a semantic recognition engine 604; and a dialogue management service 605. The central control module 601 integrates various technologies to control the dialogue logic.
[0101] The speech synthesis engine 602 synthesizes machine speech, and a timestamp of the corresponding text is generated synchronously when the machine speech audio is generated; the user speech and machine speech are recognized by the speech recognition engine 603; the semantic recognition engine 604 and the dialogue management service 605 perform complete semantic analysis and dialogue processing, and generate new words for further interaction.
[0102] The central control module 601 includes a voice endpoint detection module 6011, a noise reduction module 6012, and a revocation recovery module 6013;
[0103] The voice endpoint detection module 6011 is used to identify the beginning and end of the silence period and the voice period from the sound signal. The application can only process the valid voice and filter out the redundant silence, saving the transmission bandwidth or computing power.
[0104] The noise reduction module 6012 is used to identify the beginning and end of noise and speech from the sound signal.
[0105] The undo recovery module 6013 continues to collect data after the machine voice is interrupted, and further checks whether the interruption is correct or reasonable based on more abundant data, and recovers incorrect interruptions.
[0106] The machine speech is synthesized by the speech synthesis engine 602 in the central control module 601. After detecting human voice input through the speech endpoint detection module 6011 or the noise reduction module 6012 in the central control module 601, the system can immediately stop the system's broadcast behavior, or reduce the system's broadcast volume to produce a fade-in and fade-out effect, giving the user feedback that the system is aware that he is speaking. At the same time, when the user continues to speak for more than the second time threshold, it means that the user is expressing some complete intentions. The system will stop broadcasting and wait for the user to complete the expression. When the user pauses for more than the first time threshold, the system will consider that the user's expression is complete, and will record the interruption timestamp corresponding to the user's voice input in the context. The timestamp records the interruption position for restoring the interruption position when recovering.
[0107] In one embodiment, the first time threshold and the second time threshold may be set automatically or manually according to the scenario and user habits. The second time threshold is generally about 1 second, and the first time threshold is generally 600ms-800ms.
[0108] In one embodiment, selection and configuration are made based on the usage scenario to immediately stop the system's machine voice broadcasting behavior, or to reduce the system's machine voice broadcasting volume.
[0109] The recovery model is trained using conversation audio feature data, conversation context data, conversation turn data, etc., and is used to determine whether the user has a real intention to interrupt. If the user has a real intention to interrupt the machine voice, the system will add an interjection, such as "hmm" or "got it", and then broadcast the system dialogue corresponding to the real intention based on the user's voice; if the user has no real intention to interrupt, the system will enter the natural recovery broadcast process.
[0110] Through the timestamp information of the speech synthesis engine 602, the system can calculate a resumption playback node of the system broadcast since the last interruption, and resume the broadcast from the resumption playback node, making the system more humanized.
[0111] In one embodiment, when broadcasting the system words corresponding to the real intention or in the process of naturally resuming the broadcast, the playback volume can be restored. The recovery process can be direct recovery or smooth recovery using fade-in and fade-out. The smooth transition time is automatically or manually configured according to needs and scenarios.
[0112] In the above embodiment, it is applied to an intelligent voice conversation scenario that supports interruption. One party of the conversation is a real person and the other party is a voice robot. The real person and the machine speak alternately. When the machine speaks, the real person is allowed to interrupt the machine and obtain the right to speak. At this time, the machine waits for the real person to finish speaking, and uses the recovery model to determine whether the previous interruption intention is correct or whether the real person's speech changes the content of the conversation. If the judgment result is a wrong interruption or does not change the content of the conversation, the conversation is restored. It can interrupt flexibly, detect wrong interruptions in time, and restore the machine voice in a friendly manner, ensuring the smoothness and coherence of voice recovery, improving the user experience of interruption scenarios in voice interaction, and improving the anthropomorphic experience of intelligent voice conversations. It ensures that the system can respond in time and correctly understand user intentions, and the intelligent experience is better.
[0113] FIG7 shows a schematic diagram of an exemplary system architecture that can be applied to the dialog processing method or the dialog processing device according to an embodiment of the present disclosure.
[0114] As shown in FIG. 7 , a system architecture 700 may include terminal devices 701 , 702 , and 703 , a network 704 , and a server 705 .
[0115] The network 704 is used as a medium to provide a communication link between the terminal devices 701 , 702 , 703 and the server 705 , and can be a wired network or a wireless network.
[0116] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is typically the Internet, but can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or any combination of a virtual private network). In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged over the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies.
[0117] Terminal devices 701, 702, and 703 can be various electronic devices, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, etc., and can be used to display the playback progress of machine voice, the judgment results of the recovery model, etc.
[0118] Optionally, the client of the application installed in different terminal devices 701, 702, and 703 is the same, or is a client of the same type of application based on different operating systems. Based on different terminal platforms, the specific form of the client of the application can also be different, for example, the application client can be a mobile phone client, a PC client, etc.
[0119] Server 705 can be a server that provides various services, such as a background management server that supports devices operated by users using terminal devices 701, 702, and 703. The background management server can analyze and process received data such as requests and feedback the processing results to the terminal device. For example, it can obtain the timestamp information corresponding to the machine speech at the time of user voice input; determine the corresponding resumption playback node for the machine speech based on the timestamp information; play the machine speech from the resumption playback node; and train a recovery model.
[0120] Optionally, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, etc., but is not limited to these. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application.
[0121] Those skilled in the art will appreciate that the number of terminal devices, networks, and servers in FIG7 is merely illustrative, and any number of terminal devices, networks, and servers may be provided based on actual needs, which is not limited in the present disclosure.
[0122] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods, or program products. Therefore, various aspects of the present disclosure may be implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits," "modules," or "systems."
[0123] The electronic device 800 according to this embodiment of the present disclosure is described below with reference to Figure 8. The electronic device 800 shown in Figure 8 is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0124] As shown in Figure 8, electronic device 800 is implemented as a general-purpose computing device. Components of electronic device 800 may include, but are not limited to, the aforementioned at least one processing unit 810, the aforementioned at least one storage unit 820, and a bus 830 connecting various system components (including storage unit 820 and processing unit 810).
[0125] The storage unit stores program codes, which can be executed by the processing unit 810, so that the processing unit 810 performs the steps described in the above “Exemplary Method” section of this specification according to various exemplary embodiments of the present disclosure.
[0126] For example, the processing unit 810 can execute the following steps of the above-mentioned method embodiment: when there is user voice input during machine voice playing, stop playing the machine voice, input the user voice and the machine voice into the recovery model, and when it is determined through the recovery model that the user has no real intention to interrupt the machine voice, obtain the timestamp information corresponding to the machine voice when the user voice is input, determine the resumption playback node corresponding to the machine voice according to the timestamp information, for example, set the resumption playback node at the preset pause position corresponding to the machine voice when the user voice is input, and play the machine voice from the resumption playback node.
[0127] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 8201 and / or a cache memory unit 8202 , and may further include a read-only memory unit (ROM) 8203 .
[0128] The storage unit 820 may also include a program / utility 8204 having a set (at least one) of program modules 8205, such program modules 8205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0129] Bus 830 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0130] The electronic device 800 can also communicate with one or more external devices 840 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 800, and / or any device that enables the electronic device 800 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication can occur via an input / output (I / O) interface 850. Furthermore, the electronic device 800 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 860. As shown, the network adapter 860 communicates with other modules of the electronic device 800 via a bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 800, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0131] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0132] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided. The computer-readable storage medium may be a readable signal medium or a readable storage medium. A program product capable of implementing the above-mentioned method of the present disclosure is stored thereon. In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Methods" section above of this specification.
[0133] For example, when the program product in the embodiment of the present disclosure is executed by a processor, the following steps are implemented: when there is user voice input during machine voice playback, the machine voice is stopped, the user voice and the machine voice are input into the recovery model, and when it is determined through the recovery model that the user has no real intention to interrupt the machine voice, the timestamp information corresponding to the machine voice when the user voice is input is obtained, and the resumption playback node corresponding to the machine voice is determined according to the timestamp information, for example, the resumption playback node is set at a preset pause position corresponding to the machine voice when the user voice is input, and the machine voice is played from the resumption playback node.
[0134] More specific examples of computer-readable storage media in the present disclosure may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0135] In the present disclosure, a computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0136] Alternatively, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.
[0137] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a standalone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0138] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0139] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0140] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0141] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.
Claims
1. A dialogue processing method, wherein, Including: When there is user voice input during the playback of machine voice and the user has no intention of truly interrupting the machine voice, stop playing the machine voice and obtain the timestamp information corresponding to the machine voice at the time of the user voice input; Determine the resume playback node corresponding to the machine voice according to the timestamp information; Play the machine voice from the resume playback node.
2. The dialogue processing method according to claim 1, wherein, The step of "When there is user voice input during the playback of machine voice and the user has no intention of truly interrupting the machine voice, stop playing the machine voice and obtain the timestamp information corresponding to the machine voice at the time of the user voice input" includes: Input the user voice and the machine voice into a recovery model to determine whether the user has the intention of truly interrupting the machine voice through the recovery model.
3. The conversation processing method according to claim 2, wherein, The training data of the recovery model includes: dialogue audio feature data, dialogue context data, dialogue turn data, and / or dialogue semantic data.
4. The dialogue processing method according to claim 2, wherein, The step of "Input the user voice and the machine voice into a recovery model to determine whether the user has the intention of truly interrupting the machine voice through the recovery model" includes: The recovery model determines whether the user has the intention of truly interrupting the machine voice according to the voice data, the user voice, and the machine voice; Wherein, the voice data includes at least one of the following: guiding script content data, playback progress data of the machine voice, voice confidence data.
5. The conversation processing method according to claim 1, wherein, The resume playback node includes at least one of the following: The starting position of the machine voice; The position of the machine voice at the time of the user voice input; The sentence-breaking position corresponding to the machine voice at the time of the user voice input; The preset pause position corresponding to the machine voice at the time of the user voice input.
6. The dialogue processing method according to claim 1, wherein, The machine voice includes: voice audio, the text corresponding to the voice audio, and / or the timestamp information corresponding to the voice audio, wherein the timestamp information includes: punctuation marks, sentence-breaking positions, and / or preset pause positions.
7. The dialogue processing method according to claim 2, wherein, The step of "When there is user voice input during the playback of machine voice and the user has no intention of truly interrupting the machine voice, stop playing the machine voice and obtain the timestamp information corresponding to the machine voice at the time of the user voice input" includes: When there is user voice input during the playback of machine voice and the user has no intention of truly interrupting the machine voice, immediately stop playing the machine voice or slowly reduce the volume of the machine voice.
8. The conversation processing method according to claim 1, wherein, Also including: When the pause of the user voice exceeds the first time threshold, it indicates that the user voice input is completed.
9. The dialogue processing method according to claim 1, wherein, Also including: When the user has the intention of truly interrupting the machine voice, generate a reply voice corresponding to the user voice.
10. The dialogue processing method according to claim 7 or 8, wherein, Also including: Filter the silent data and noise data in the user voice.
11. A dialogue processing device, wherein, Including: An acquisition module, when there is user voice input during the playback of machine voice and the user has no intention of truly interrupting the machine voice, stop playing the machine voice and obtain the timestamp information corresponding to the machine voice at the time of the user voice input; A determination module that determines a resume playback node corresponding to the machine voice according to the timestamp information; A resume module that plays the machine voice from the resume playback node.
12. An electronic device, wherein, It includes: A processor; And A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the dialogue processing method according to any one of claims 1 to 10 by executing the executable instructions.
13. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by the processor, it implements the dialogue processing method according to any one of claims 1 to 10.
14. A computer program product comprising a computer program, wherein, When the computer program is executed by the processor, it implements the dialogue processing method according to any one of claims 1 to 10.