Dialogue processing method and device and related product

Through the combination of streaming output and content review, the problem of users waiting for a long time to reply messages is solved, and the dialogue efficiency is improved.

CN120354941APending Publication Date: 2025-07-22DOUYIN VISION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510457816.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, users need to wait a long time to get a reply message when talking to a dialogue assistant in an e-book, resulting in low conversation efficiency.

Method used

Get the reply message of the conversation assistant through streaming output, and perform content review during the output process, withdraw the reply characters that have not passed the review, and output the preset reply message.

Benefits of technology

Improves the efficiency of conversation between users and dialogue assistants and reduces the time to wait for replying to messages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354941A_ABST
    Figure CN120354941A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a dialogue processing method and device and a related product, and the method comprises the steps: obtaining a dialogue message for a dialogue assistant, and inputting the dialogue message to a generative model corresponding to the dialogue assistant; the dialogue assistant is arranged in the electronic book and is used for carrying out dialogue aiming at contents in the electronic book; sequentially obtaining each original reply character output by the generative model for the dialogue message; the original reply characters are sequentially output by the generative model in a streaming output mode; according to the obtaining sequence of the original reply characters, content auditing is conducted on the obtained original reply characters, and the dialogue assistant is triggered to output the obtained original reply characters in a streaming mode; if the original reply character passing the content auditing is detected, content auditing continues to be carried out on subsequent characters of the original reply character, and if the original reply character not passing the content auditing is detected, the dialogue assistant is triggered to withdraw the output original reply character, and the dialogue assistant is triggered to output a preset reply message.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a dialogue processing method, apparatus, and related products. Background Art

[0002] Currently, during the process of reading an e-book, a user can have a conversation with a dialogue assistant in the e-reader. By having a conversation with the dialogue assistant, the user's needs such as understanding the plot development and interacting with the characters in the book can be met. In related technologies, after obtaining the dialogue message input by the user, the dialogue assistant generates a reply message corresponding to the dialogue message. After all the reply messages are generated, content detection is performed on all the reply messages. After all the reply messages pass the content detection, all the reply messages are output to the user at one time. This way of waiting for all the reply messages to be generated and then outputting all the reply messages to the user at one time can be referred to as non-streaming output. Since it takes a long time to generate all the reply messages and perform content detection on all the reply messages, in related technologies, the user needs to wait for a long time to obtain the reply message, reducing the dialogue efficiency between the user and the dialogue assistant. Summary of the Invention

[0003] Embodiments of the present disclosure provide a dialogue processing method, apparatus, and related products. In a scenario where a user has a conversation with a dialogue assistant in an e-book, the dialogue assistant can be triggered to stream output the reply messages of the dialogue messages sent by the user, and content review is performed on the reply messages during the process of streaming outputting the reply messages, reducing the time consumed by the user waiting for the reply messages and improving the dialogue efficiency between the user and the dialogue assistant.

[0004] In a first aspect, embodiments of the present disclosure provide a dialogue processing method, including: Obtaining a dialogue message for the dialogue assistant, and inputting the dialogue message into a generative model corresponding to the dialogue assistant; the dialogue assistant is set in the e-book and is used to have a conversation regarding the content in the e-book; Sequentially obtaining each original reply character output by the generative model for the dialogue message; the original reply characters are successively output by the generative model in a streaming output manner; According to the acquisition order of the original reply characters, performing content review on the obtained original reply characters and triggering the dialogue assistant to stream output the obtained original reply characters; If an original reply character that passes the content review is detected, continuing to perform content review on the subsequent characters of the original reply character; if an original reply character that fails to pass the content review is detected, triggering the dialogue assistant to withdraw the already output original reply characters and triggering the dialogue assistant to output a preset reply message.

[0005] In a second aspect, an embodiment of the present disclosure provides a dialogue processing device, including: A message acquisition unit, configured to acquire dialogue messages for a dialogue assistant, and input the dialogue messages into a generative model corresponding to the dialogue assistant; the dialogue assistant is set in an e-book and is used to conduct a dialogue on the content in the e-book; A reply acquisition unit, configured to sequentially acquire each original reply character output by the generative model for the dialogue message; the original reply characters are successively output by the generative model in a streaming output manner; A reply transmission unit, configured to, according to the acquisition order of the original reply characters, conduct content review on the acquired original reply characters and trigger the dialogue assistant to stream out the acquired original reply characters; A reply processing unit, configured to, if it detects an original reply character that passes the content review, continue to conduct content review on the subsequent characters of the original reply character, and if it detects an original reply character that fails to pass the content review, trigger the dialogue assistant to withdraw the output original reply character and trigger the dialogue assistant to output a preset reply message.

[0006] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor; and a memory configured to store computer-executable instructions, where the computer-executable instructions, when executed, cause the processor to implement the method described in the first aspect above.

[0007] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, where the computer-readable storage medium is used to store computer-executable instructions, and the computer-executable instructions, when executed by a processor, implement the method described in the first aspect above.

[0008] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, where the computer program product includes a computer program, and the computer program, when executed by a processor, implements the method described in the first aspect above.

[0009] Through this embodiment, during the process of the generative model corresponding to the dialogue assistant streaming out each original reply character corresponding to the dialogue message, each output original reply character can be obtained in sequence. According to the acquisition order of the original reply characters, content review is performed on the obtained original reply characters, and the dialogue assistant is triggered to stream out the obtained original reply characters. Moreover, in the case of detecting an original reply character that passes the content review, content review continues for the subsequent original reply characters. When detecting an original reply character that fails to pass the content review, the dialogue assistant is triggered to withdraw all the output original reply characters and output a preset reply message. Thus, on the one hand, the effect of triggering the dialogue assistant to stream out the reply message of the user's sent dialogue message is achieved, and on the other hand, the effect of content review on the reply message during the process of streaming out the reply message is achieved. It is not necessary to wait for all the original reply characters of the generative model to be output and then perform content review before the dialogue assistant outputs the reply message of the dialogue message at one time, reducing the time-consuming for the user to wait for the reply message and improving the dialogue efficiency between the user and the dialogue assistant. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in one or more embodiments of the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments recorded in the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Figure 1 It is a schematic flowchart of a dialogue processing method provided by an embodiment of the present disclosure; Figure 2 It is a schematic diagram of a generative model streaming various original reply characters to the server provided by an embodiment of the present disclosure; Figure 3 It is a schematic diagram of determining the character length that can be accommodated by a character truncation window provided by an embodiment of the present disclosure; Figure 4 It is a schematic flowchart of a dialogue processing method provided by another embodiment of the present disclosure; Figure 5 It is a schematic flowchart of outputting the audio data of the final reply message provided by an embodiment of the present disclosure; Figure 6a It is a schematic diagram of a dialogue assistant streaming out the final reply message provided by an embodiment of the present disclosure; Figure 6b It is a schematic diagram of a dialogue assistant streaming out the final reply message provided by another embodiment of the present disclosure; Figure 7 It is a schematic structural diagram of a dialogue processing device provided by an embodiment of the present disclosure; Figure 8 Schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0011] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of the present disclosure, the following will clearly and completely describe the technical solutions in one or more embodiments of the present disclosure with reference to the accompanying drawings in one or more embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on one or more embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.

[0012] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, usage scenarios, etc. of the information involved in the present disclosure should be informed to the relevant parties in an appropriate manner and the authorization of the relevant parties should be obtained in accordance with relevant laws and regulations.

[0013] For example, when a user's active request is received, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operations of the technical solutions of the present disclosure according to the prompt message.

[0014] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0015] It can be understood that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0016] The embodiments of the present disclosure provide a dialogue processing method, apparatus, and related products. In a scenario where a user has a dialogue with a dialogue assistant in an e-book, it is possible to trigger the dialogue assistant to stream out a reply message to the dialogue message sent by the user, and perform content review on the reply message during the process of streaming out the reply message, reducing the time-consuming for the user to wait for the reply message and improving the dialogue efficiency between the user and the dialogue assistant. Among them, the dialogue processing method can be applied to a server side and executed by the server side. The server side includes both a single-server form and a server cluster form.

[0017] To facilitate the understanding of various embodiments of the present disclosure, the nouns that may be involved are explained here.

[0018] Streaming output: Streaming output is a data output method, which means that data is continuously output from a data source (such as a program, device, etc.) to a target location (such as a screen, file, network, etc.) like a water flow, rather than outputting all the data at once after it is prepared. During the streaming output process, data is output one by one in smaller blocks or units. After the previous data unit is output, the next data unit is immediately output. This method allows data to be processed and displayed in a timely manner after it is generated, without waiting for all the data to be generated, improving the data response speed and user experience.

[0019] Streaming transmission: Streaming transmission is a technology and method for transmitting data. It divides the data into a series of data packets and then continuously transmits them from the sender to the receiver on the network like a water flow. After receiving the data packets, the receiver can start processing the received data packets without waiting for all the subsequent data packets to be transmitted.

[0020] Streaming sending: Streaming sending is a way for the sender to send data, which means that the sender sends the data in a streaming manner. The sender will split the data into data packets according to certain rules and protocols and then send these data packets out one by one in sequence.

[0021] Figure 1 The following is a schematic flowchart of a dialogue processing method provided by an embodiment of the present disclosure. As Figure 1 shown, the method includes: Step S102, obtain a dialogue message for the dialogue assistant and input the dialogue message into the generative model corresponding to the dialogue assistant; the dialogue assistant is set in the e-book and is used to conduct a dialogue on the content in the e-book; Step S104, sequentially obtain each original reply character output by the generative model for the dialogue message; the original reply characters are successively output by the generative model in a streaming output manner; Step S106, according to the acquisition order of the original reply characters, conduct content review on the obtained original reply characters and trigger the dialogue assistant to stream out the obtained original reply characters; Step S108, if an original reply character that passes the content review is detected, continue to conduct content review on the subsequent characters of the original reply character; if an original reply character that fails to pass the content review is detected, trigger the dialogue assistant to withdraw the already output original reply characters and trigger the dialogue assistant to output a preset reply message.

[0022] In the embodiments of the present disclosure, first, a dialogue message for the dialogue assistant is obtained, and the dialogue message is input into the generative model corresponding to the dialogue assistant. The dialogue assistant is set in the e-book and is used to conduct a dialogue regarding the content in the e-book. Then, each original reply character output by the generative model for the dialogue message is sequentially obtained. The original reply characters are successively output by the generative model in a streaming manner. Next, according to the obtaining order of the original reply characters, content review is performed on the obtained original reply characters, and the dialogue assistant is triggered to stream output the obtained original reply characters. Finally, if an original reply character that passes the content review is detected, content review is continued for the subsequent characters of the original reply character. If an original reply character that fails to pass the content review is detected, the dialogue assistant is triggered to withdraw the already output original reply characters, and the dialogue assistant is triggered to output a preset reply message. It can be seen that through this embodiment, during the process of the generative model corresponding to the dialogue assistant streaming outputting each original reply character corresponding to the dialogue message, each output original reply character can be sequentially obtained, and according to the obtaining order of the original reply characters, content review is performed on the obtained original reply characters and the dialogue assistant is triggered to stream output the obtained original reply characters. Also, in the case where an original reply character that passes the content review is detected, content review is continued for the subsequent original reply characters. When an original reply character that fails to pass the content review is detected, the dialogue assistant is triggered to withdraw each already output original reply character, and the dialogue assistant is triggered to output a preset reply message. Thus, on the one hand, the effect of triggering the dialogue assistant to stream output the reply message of the dialogue message sent by the user is achieved, and on the other hand, the effect of performing content review on the reply message during the process of streaming outputting the reply message is achieved. It is not necessary to wait for all the original reply characters of the generative model to be output and content review to be performed before the dialogue assistant outputs the reply message of the dialogue message at one time, reducing the time consumed by the user waiting for the reply message and improving the dialogue efficiency between the user and the dialogue assistant.

[0023] In step S102 above, a dialogue message for the dialogue assistant is obtained, and the dialogue message is input into the generative model corresponding to the dialogue assistant. The dialogue assistant is set in the e-book and is used to conduct a dialogue regarding the content in the e-book.

[0024] In this embodiment, the e-book can be located in a terminal device and is read through a reading application in the terminal device. The terminal device includes but is not limited to various types of user terminals such as laptop computers, tablet computers, desktop computers, mobile devices (such as mobile phones, personal digital assistants, dedicated messaging devices), smart phones, smart watches, smart TVs, in-vehicle terminals, etc.

[0025] The e - book has a corresponding dialogue assistant, which can be a dialogue robot pre - created for the characters in the e - book to represent that character in the e - book. For example, it represents the female protagonist in the e - book. By having a conversation with the dialogue assistant, users' needs such as understanding the plot development and interacting with the characters can be met. As an alternative implementation, the dialogue assistant is implemented through a generative model, which at least includes a large - language model (LLM), and may also include other content - generation models such as an AI (Artificial Intelligence) drawing model and an AI video - generation model. That is, the generative model can be a multi - modal content - generation model that can generate and output data such as text, images, and videos according to the information input by the user.

[0026] In this embodiment, the user can send a dialogue message to the dialogue assistant, such as "Why do you like the male protagonist of this book instead of the second male character?" The dialogue assistant can represent the character of the female protagonist in the e - book. The terminal device obtains this dialogue message and sends it to the server, so that the server obtains the dialogue message from the user for the dialogue assistant in the e - book. The server also inputs this dialogue message into the generative model corresponding to the dialogue assistant. The generative model corresponding to the dialogue assistant can analyze and process the user's dialogue message and generate a corresponding reply message.

[0027] In the above step S104, each original reply character output by the generative model for the dialogue message is obtained in sequence. Each original reply character is output successively by the generative model in a streaming output manner, and all the original reply characters together form the reply message of the generative model for the dialogue message.

[0028] In this embodiment, after processing the dialogue message, the generative model can generate one or more characters each time, use these one or more characters as the original reply characters, and perform the operation of generating the original reply characters one or more times to generate one or more original reply characters. The generative model outputs each original reply character successively in a streaming output manner. Streaming output means that the generative model outputs each original reply character as soon as it is generated, without considering whether the next original reply character is generated. The output order of each original reply character is the generation order of each original reply character. Thus, the server obtains each original reply character output by the generative model in sequence according to the output order of the generative model.

[0029] In one example, in the generative model, through the same process, each time an original reply character is generated, one or more of the generated original reply characters are output, and then the subsequent original reply characters are continuously generated and output. In another example, in the generative model, through the first process, each original reply character is continuously generated, and through the second process, one or more of the generated original reply characters are output each time an original reply character is generated, without considering whether the next original reply character is generated.

[0030] In this embodiment, all the original reply characters generated and output by the generative model together constitute the reply message of the generative model for the dialogue message. For example, the dialogue message is "Why do you like the male protagonist of this book rather than the second male protagonist?", and the generative model outputs the original reply characters three times, with multiple original reply characters output each time, namely "Because the male protagonist has a very good personality,", "The second male protagonist has a really bad temper,", and "Moreover, the male protagonist is also very good to me". All these original reply characters together constitute the reply message of the generative model for the dialogue message "Because the male protagonist has a very good personality, the second male protagonist has a really bad temper, and moreover, the male protagonist is also very good to me".

[0031] In one embodiment, successively obtaining each original reply character output by the generative model for the dialogue message includes: Using a specific remote call framework to create a first remote call channel with the generative model; the type of the first remote call channel is a streaming channel for streaming transmission; Through the first remote call channel, obtain each original reply character in sequence according to the order in which the generative model streams out the original reply characters.

[0032] In this embodiment, first use a specific remote call framework such as the KiteX framework to create a first remote call (RPC, Remote Procedure Cal) channel between the server and the generative model. The type of this channel is a streaming channel for streaming transmission. A streaming interface for streaming transmission can be created for the server and the generative model respectively through the specific remote call framework. Based on the streaming interface of the server and the streaming interface of the generative model, create a streaming channel between the server and the generative model as the first remote call channel. The streaming channel allows the generative model to return multiple original reply characters based on a single dialogue message sent by the server. Different from the streaming channel is the non-streaming channel, which only allows the generative model to return an original reply character once based on a single dialogue message sent by the server.

[0033] The server sends a conversation message to the generative model through the first remote call channel. The generative model uses the first remote call channel to obtain the conversation message, generates one or more original reply characters through the above process, generates one or more original reply characters each time, and based on the type of the first remote call channel, outputs each of the originally generated reply characters each time they are generated, and sends each of the originally generated reply characters to the server through the first remote call channel. Thus, the server obtains each of the originally generated reply characters in sequence through the first remote call channel according to the order in which the generative model streams and sends each of the originally generated reply characters. It can be understood that the order in which the generative model generates each of the originally generated reply characters, the order in which the generative model outputs each of the originally generated reply characters, and the order in which the generative model sends each of the originally generated reply characters to the server are the same.

[0034] Figure 2 FIG. is a schematic diagram of a generative model provided by an embodiment of the present disclosure for streaming and sending each original reply character to the server. As Figure 2 shown, a first remote call channel can be created between the server and the generative model. The server sends a conversation message to the generative model through the first remote call channel. The generative model uses the first remote call channel to obtain the conversation message. Based on the type of the first remote call channel, after generating one or more original reply characters each time, the one or more original reply characters are output and sent to the server through the first remote call channel. Thus, the server obtains each of the originally generated reply characters in sequence according to the order in which the generative model streams and sends each of the originally generated reply characters.

[0035] As Figure 2 shown, in streaming transmission, the server only needs to send a conversation message to the generative model once, and the generative model can return the original reply characters multiple times until the last original reply character is returned, that is, the conversation message reply is completed. The generative model can be deployed inside the server or outside the server.

[0036] It can be seen that through this embodiment, a first remote call channel with the generative model can be created by using a specific remote call framework; the type of the first remote call channel is a streaming channel for streaming transmission. Through the first remote call channel, each of the originally generated reply characters is obtained in sequence according to the order in which the generative model streams and outputs the original reply characters, thereby achieving streaming acquisition of each of the originally generated reply characters and preparing for triggering the streaming output of each of the originally generated reply characters by the conversation assistant.

[0037] In the above step S106, according to the acquisition order of the original reply characters, content review is performed on the obtained original reply characters and the conversation assistant is triggered to stream and output the obtained original reply characters.

[0038] In this embodiment, the server continuously obtains each original reply character, and according to the acquisition order of the original reply characters, conducts content review on the obtained original reply characters and triggers the dialogue assistant to stream out the obtained original reply characters.

[0039] In one embodiment, conducting content review on the obtained original reply characters and triggering the dialogue assistant to stream out the obtained original reply characters according to the acquisition order of the original reply characters includes: Generating at least one final reply message according to the acquisition order of the original reply characters; Conducting content review on the final reply message according to the generation order of the final reply message and triggering the dialogue assistant to stream out the final reply message.

[0040] In this embodiment, one or more final reply messages of the dialogue assistant are generated according to the obtained original reply characters. The final reply message is the reply message output and displayed by the dialogue assistant in the terminal device and is the reply message of the dialogue message browsed by the user. The server can also perform the operations of conducting content review on the final reply message and triggering the dialogue assistant to stream out the final reply message according to the generation order of each final reply message.

[0041] Since the final reply message is generated according to the original reply characters, and all the original reply characters together constitute the reply message of the generative model for the dialogue message, in order to ensure accurate transmission of the semantic information of the reply message of the generative model, in this embodiment, it is set that each final reply message jointly represents the semantic information of the reply message of the generative model. For example, it is set that the total number of characters of all the final reply messages is the same as the total number of characters of all the original reply characters, so as to convey the semantic information of the reply message of the generative model to the user.

[0042] Of course, the total number of characters of all the final reply messages can also be different from the total number of characters of all the original reply characters. For example, the redundant characters in the original reply characters are not included in all the final reply messages, or the error characters in the original reply characters are corrected in all the final reply messages, so as to obtain a more accurate final reply message on the basis of ensuring that all the final reply messages jointly represent the semantic information of the reply message of the generative model.

[0043] It can be seen that through this embodiment, at least one final reply message can be generated according to the acquisition order of the original reply characters, content review can be conducted on the final reply message according to the generation order of the final reply message, and the dialogue assistant can be triggered to stream out the final reply message, so as to achieve the effect of triggering the dialogue assistant to stream out the original reply characters, and during the process of generating the final reply message, the redundant characters or error characters in the original reply characters can be corrected to improve the accuracy of the reply message.

[0044] In one embodiment, at least one final reply message is generated according to the acquisition order of the original reply characters, including: Obtain a character truncation window; the character length that the character truncation window can accommodate is the character length corresponding to the final reply message; Using the character truncation window, truncate the obtained original reply characters one or more times in the acquisition order of the original reply characters until the last original reply character is truncated; wherein, the last character truncated by the previous character truncation window is the previous character of the first character truncated by the next character truncation window; Generate a final reply message using the characters accommodated in each character truncation window.

[0045] In this embodiment, first, obtain a pre-created character truncation window. The character length that the character truncation window can accommodate is the character length corresponding to the final reply message. For example, the character truncation window can accommodate 10 characters.

[0046] Then, during the process of receiving each original reply character, continuously use the character truncation window to truncate the obtained original reply characters one or more times in the acquisition order of the original reply characters until the characters in the last original reply character are truncated. And set the last character truncated by the previous character truncation window to be the previous character of the first character truncated by the next character truncation window. Finally, generate a final reply message using the characters accommodated in each character truncation window.

[0047] Considering that the generative model outputs one or more original reply characters each time, two cases are described here.

[0048] Case 1: The character length that the character truncation window can accommodate is greater than the character length of the original reply characters output by the generative model at one time. In this case, after the server obtains the original reply characters output for the first time, it is still unable to generate the first final reply message and needs to wait for the original reply characters output by the generative model subsequently. The first output of the original reply characters and the subsequent output of the original reply characters are truncated using the character truncation window until the character truncation window is filled, and the characters with the character length that the character truncation window can accommodate are truncated. This process can also be understood as accumulating the first output of the original reply characters and the subsequent output of the original reply characters through the character truncation window until the character truncation window is full. In this case, the characters in the character truncation window constitute a final reply message, or the incorrect characters in the character truncation window are corrected to obtain a final reply message. Repeat the above process to truncate the subsequently received original reply characters to obtain each final reply message. During the process of generating each final reply message, redundant characters such as redundant punctuation marks in the received original reply characters can also be deleted.

[0049] In an example, the character truncation window can accommodate 10 characters. The original reply characters output by the generative model for the first time are "Hello,", the original reply characters output for the second time are "My name is Tian Tian.", the original reply characters output for the third time are "Nice to meet you!", the original reply characters output for the fourth time are "Now let's start our", and the original reply characters output for the fifth time are " friendship!". Through the above process, the first final reply message "Hello, My name is Tian Tian. Nice to " can be generated. Then, starting from the character "You", continue to accumulate characters through the character truncation window to obtain the second final reply message " meet you! Now let's start our", and then, starting from the character "Are", continue to accumulate characters through the character truncation window to obtain the next final reply message.

[0050] Case 2: The length of characters that the character interception window can accommodate is less than or equal to the length of the original response characters output by the generative model at one time. In this case, after the server obtains the original response characters output for the first time, it uses the character interception window to intercept the characters in the original response characters output for the first time until the character interception window is filled up, and intercepts the characters of the character length that the character interception window can accommodate. This process can also be understood as accumulating the characters in the original response characters output for the first time through the character interception window until the character interception window is full. In this case, the characters in the character interception window constitute a final response message, or the wrong characters in the character interception window are corrected to obtain a final response message. Repeat the above process, intercept the original response characters received for subsequent output, and obtain each final response message. In the process of generating each final response message, redundant characters such as redundant punctuation marks in the received original response characters can also be deleted.

[0051] In an example, the character interception window can accommodate 10 characters, and the original reply characters of the first output of the generative model are "Hello, my name is Tiantian. Nice to meet you! From now on, we are good friends!". Through the above process, the first final reply message "Hello, my name is Tiantian. Nice to meet you!" can be generated. Then, starting from the word "you", characters are accumulated through the character interception window to obtain the second final reply message "you meet! From now on, we are!". Then, starting from the word "yes", characters are accumulated through the character interception window to obtain the next final reply message.

[0052] A foreseeable situation is that when the original reply characters are intercepted using the character interception window, the length of the last intercepted characters may be less than the character length that can be accommodated. In this case, the last intercepted characters are used to generate the last final reply message.

[0053] It can be seen that through this embodiment, the character interception window can be used to intercept the original reply characters once or multiple times in the order in which the original reply characters are obtained until the last original reply character is intercepted. The characters accommodated in each character interception window are used to generate a final reply message. The character interception window is used to control the length of the final reply message, thereby avoiding returning too long or too short messages to the conversation assistant, which affects the user experience.

[0054] In this embodiment, the length of characters that the character capture window can accommodate is the length of the final reply message. Considering the need to conduct content review on the final reply message, it is necessary to comprehensively consider the content review effect and the speed of returning the final reply message to the conversation assistant. Figure 3Schematic diagram for determining the character length that a character truncation window can accommodate provided by an embodiment of the present disclosure. As Figure 3 shown, the shorter the character length that the character truncation window can accommodate, the worse the content review effect, but the faster the speed of returning the final reply message to the dialogue assistant. The longer the character length that the character truncation window can accommodate, the better the content review effect, but the slower the speed of returning the final reply message to the dialogue assistant. The character length that the character truncation window can accommodate can be determined by comprehensively considering the content review effect and the speed of returning the final reply message to the dialogue assistant.

[0055] In one example, the character length that the character truncation window used to generate the first final reply message can accommodate can be set to be less than the character length that the character truncation window used to generate subsequent final reply messages can accommodate. That is, the size of the character truncation window used to generate the first final reply message is set to be less than the size of the character truncation window used to generate subsequent final reply messages, so that the length of the first final reply message is less than the length of subsequent final reply messages, thereby improving the generation speed of the first final reply message and reducing the delay for the user to view the first final reply message.

[0056] In one example, the size of the character truncation window corresponding to each final reply message can also be set to be greater than or equal to the average length of the original reply characters output by the generative model each time. That is, the length of each final reply message is made greater than or equal to the average length of the original reply characters output by the generative model each time. This average length can be an estimated value, so that one or more outputs of the original reply characters can be used to generate a final reply message, avoiding splitting the original reply characters, reducing the number of generated final reply messages, and reducing the data transmission cost between the server and the terminal device where the dialogue assistant is located.

[0057] In this embodiment, content review of the final reply message is also performed according to the generation order of the final reply message, and the dialogue assistant is triggered to stream output the final reply message. In one embodiment, performing content review of the final reply message according to the generation order of the final reply message and triggering the dialogue assistant to stream output the final reply message includes: For the generated first final reply message, content review is performed on the first final reply message. After the content review passes, the dialogue assistant is triggered to stream output the first final reply message; For the generated non-first final reply message, the dialogue assistant is triggered to stream output the non-first final reply message. After triggering the dialogue assistant, content review is performed on the non-first final reply message.

[0058] In this embodiment, for the first generated final reply message, the content of the first final reply message is first audited. After the content audit passes, the first final reply message is sent to the dialogue assistant, triggering the dialogue assistant to stream out the characters in the first final reply message. The dialogue assistant can output the characters in the first final reply message one by one on the screen, presenting the effect of a typewriter outputting one character at a time, improving the user's browsing experience.

[0059] For the generated non-first final reply message, the non-first final reply message is sent to the dialogue assistant, triggering the dialogue assistant to stream out the non-first final reply message. The dialogue assistant can output the characters in the non-first final reply message one by one on the screen, presenting the effect of a typewriter outputting one character at a time, improving the user's browsing experience. After the server triggers the dialogue assistant to stream out the non-first final reply message, that is, after sending the non-first final reply message to the dialogue assistant, the content of the non-first final reply message is audited. For example, after generating the first final reply message and the content audit of the first final reply message passes, each time a new final reply message is generated, the process of sending the final reply message to the dialogue assistant is first started, and after starting, the content of the final reply message is audited.

[0060] Considering that the final reply messages are generated one after another, in an optional approach, a first thread and a second thread are created, and the first thread and the second thread are set to run in parallel. After generating the first final reply message and the content audit of the first final reply message passes, each time a new final reply message is generated, the process of sending the final reply message to the dialogue assistant is started through the first thread. The first thread is used to send the generated non-first final reply messages one by one to the dialogue assistant. And, when the first thread starts the process of sending a final reply message, the content of the sent final reply message is audited through the second thread. The second thread is used to audit the non-first final reply messages that the first thread has started to send one by one in the sending order.

[0061] It can be seen that through this embodiment, considering the situation where unreasonable replies output by the generative model can be distinguished by the first few characters output by the generative model, the generated first final reply message and non-first final reply messages are distinguished. For the first final reply message, a processing strategy of auditing first and then sending is adopted, and for the non-first final reply message, a processing strategy of sending first and then auditing is adopted, taking into account both the user's need to quickly view the final reply message and the need to audit the content of the final reply message.

[0062] In one embodiment, auditing the content of the first final reply message includes: Through the content review model, detect whether the first final reply message contains specified characters, and detect whether the semantic information of the first final reply message meets the semantic requirements; If the first final reply message does not contain specified characters and the semantic information of the first final reply message meets the semantic requirements, it is determined that the content review of the first final reply message passes; otherwise, it is determined that the content review of the first final reply message fails.

[0063] In this embodiment, the first final reply message is input into a pre-trained content review model. Through the content review model, detect whether the first final reply message contains specified characters, and detect whether the semantic information of the first final reply message meets the semantic requirements. Among them, the specified characters include pre-determined illegal characters, and the semantic requirements refer to that the semantic information does not contain illegal semantic information. That the semantic information of the first final reply message meets the semantic requirements may mean that the semantic information of the first final reply message does not contain illegal semantic information. On the contrary, that the semantic information of the first final reply message does not meet the semantic requirements may mean that the semantic information of the first final reply message contains illegal semantic information.

[0064] If it is determined through the content review model that the first final reply message does not contain specified characters and the semantic information of the first final reply message meets the semantic requirements, it is determined that the content review of the first final reply message passes. If it is determined through the content review model that the first final reply message contains specified characters, or the semantic information of the first final reply message does not meet the semantic requirements, it is determined that the content review of the first final reply message fails.

[0065] It can be seen that through this embodiment, it is possible to accurately determine whether the content review of the first final reply message passes from two aspects: whether it contains specified characters and whether the semantic information meets the semantic requirements through the content review model.

[0066] In one embodiment, the content review of non-first final reply messages includes: Generate a target message based on the non-first final reply message; the target message includes the non-first final reply message and the final reply messages generated before the non-first final reply message; Through the content review model, detect whether the target message contains specified characters, and detect whether the semantic information of the target message meets the semantic requirements; If the target message does not contain specified characters and the semantic information of the target message meets the semantic requirements, it is determined that the content review of the non-first final reply message passes; otherwise, it is determined that the content review of the non-first final reply message fails.

[0067] As can be seen from the above description, the non-first final reply messages that have been started to be sent can be content-reviewed one by one in the order of being sent to the dialogue assistant. Therefore, in this embodiment, it is possible to perform content review on multiple non-first final reply messages. Here, any non-first final reply message is taken as an example for illustration.

[0068] For any non-first final reply message to be reviewed, based on this non-first final reply message, a target message is generated. The target message includes this non-first final reply message and the final reply messages generated before this non-first final reply message. In this embodiment, in the order of generation of the final reply messages, this non-first final reply message and each of the final reply messages generated before this non-first final reply message can be concatenated together to obtain the target message. Taking the third generated final reply message as an example, the target message includes the first, second, and third final reply messages, and the target message is the message obtained by concatenating the first, second, and third final reply messages in the order of generation of the final reply messages. Since the generation order of the final reply messages is the same as the character generation order of the generative model and is also the same as the character output order of the generative model, this step can be understood as traversing from the first character output by the generative model to the last character of this non-first final reply message to obtain the target message.

[0069] Next, the target message is input into a pre-trained content review model. Through the content review model, it is detected whether the target message contains specified characters, and whether the semantic information of the target message meets the semantic requirements. Among them, the specified characters include pre-determined illegal characters, and the semantic requirements mean that the semantic information does not contain illegal semantic information. That the semantic information of the target message meets the semantic requirements can be that the semantic information of the target message does not contain illegal semantic information. On the contrary, that the semantic information of the target message does not meet the semantic requirements can be that the semantic information of the target message contains illegal semantic information.

[0070] If it is determined through the content review model that the target message does not contain specified characters and that the semantic information of the target message meets the semantic requirements, then it is determined that the content review of this non-first final reply message passes. If it is determined through the content review model that the target message contains specified characters, or that the semantic information of the target message does not meet the semantic requirements, then it is determined that the content review of this non-first final reply message fails.

[0071] It can be seen that through this embodiment, it is possible to achieve the effect of traversing from the first character output by the generative model to the last character of the non-first final reply message to be reviewed, obtaining the target message, and through the content review model, accurately judging whether the content review of the target message passes from two aspects: whether it contains the specified character and whether the semantic information meets the semantic requirements. Since the target message can be a long text from the first character output by the generative model to the last character of the non-first final reply message to be reviewed, by reviewing the target message, the accuracy of semantic detection can be improved based on the semantic coherence between each final reply message.

[0072] In the above step S108, if an original reply character that passes the content review is detected, the content review continues for the subsequent characters of the original reply character. If an original reply character that fails the content review is detected, the dialogue assistant is triggered to withdraw the originally output original reply character and output a preset reply message.

[0073] In this embodiment, through the above process, the first final reply message and the non-first final reply message are detected. If an original reply character that passes the content review is detected, the content review continues for the subsequent characters of the original reply character. For example, if any final reply message passes the content review, it is determined that an original reply character that passes the content review is detected, and the content review continues for the subsequent final reply messages. If an original reply character that fails the content review is detected, the dialogue assistant is triggered to withdraw the originally output original reply character and output a preset reply message. For example, if a certain final reply message fails the content review, it is determined that an original reply character that fails the content review is detected, and the dialogue assistant is triggered to withdraw all the finally output reply messages and output a preset reply message.

[0074] In this embodiment, triggering the dialogue assistant to withdraw the originally output original reply character can be triggering the dialogue assistant to hide the originally displayed original reply character on the screen, such as hiding the finally displayed reply message.

[0075] In one embodiment, if an original reply character that fails the content review is detected, triggering the dialogue assistant to withdraw the originally output original reply character and output a preset reply message includes: If the original reply character that fails the content review is in the generated first final reply message, trigger the dialogue assistant to output a preset reply message; wherein, the dialogue assistant has not output the first final reply message. If the original reply characters that fail the content review are located in the generated non-first final reply message, the dialogue assistant is triggered to withdraw all the output final reply messages and output a preset reply message; among them, the dialogue assistant has been triggered to output a non-first final reply message.

[0076] In this embodiment, if the original reply characters that fail the content review are detected, such as illegal characters or characters with semantics not meeting the requirements, it is determined whether the characters are located in the generated first final reply message or in the generated non-first final reply message.

[0077] If it is located in the first final reply message, it is determined that the first final reply message fails the content review. In this case, the first final reply message is not sent to the dialogue assistant, the acquisition of the original reply characters output by the generative model is terminated, and a preset reply message is output to the dialogue assistant, triggering the dialogue assistant to output the preset reply message to the user.

[0078] In one example, if it is determined that the content review of the first final reply message fails, the server terminates the acquisition of the original reply characters output by the generative model and sends a preset reply message such as "How about we change the topic?" to the dialogue assistant, thereby ending the processing flow for the dialogue message and guiding the user to initiate a new topic. The preset reply message is output by the dialogue assistant, such as in a streaming manner.

[0079] If it is located in the non-first final reply message, it is determined that the non-first final reply message fails the content review. In this case, the acquisition of the original reply characters output by the generative model is terminated, the dialogue assistant is triggered to withdraw all the output final reply messages, and the dialogue assistant is triggered to output a preset reply message. Since the strategy for the non-first final reply message is to send first and then review, through the previous process, the dialogue assistant has been triggered to output the non-first final reply message that fails the content review. The non-first final reply message that fails the content review may have started to be output to the user by the dialogue assistant in a streaming manner, and may have been completely output or may not have been completely output.

[0080] In one example, if the content review of any non-first final reply message fails, the server terminates the acquisition of the original reply characters output by the generative model and triggers the dialogue assistant to withdraw all the displayed final reply messages (including the first final reply message and the non-first final reply messages), and sends a preset reply message such as "How about we change the topic?" to the dialogue assistant, thereby ending the processing flow for the dialogue message and guiding the user to initiate a new topic. The preset reply message is output by the dialogue assistant, such as in a streaming manner.

[0081] It can be seen that through this embodiment, if it is determined that the content review of the first final reply message fails, on the one hand, the first final reply message is not sent to the dialogue assistant to prevent the user from viewing illegal replies. On the other hand, the acquisition of the original reply characters output by the generative model is terminated, saving the processing resources of the server and the generative model. On the third hand, a preset reply message is sent to the dialogue assistant to guide the user to initiate a new topic and improve the dialogue experience between the user and the dialogue assistant. Moreover, in the case where the content review of any non-first final reply message fails, on the one hand, the acquisition of the original reply characters output by the generative model is terminated, saving the processing resources of the server and the generative model. On the other hand, the dialogue assistant is triggered to withdraw all the final reply messages that have been displayed to prevent the user from viewing illegal replies. On the third hand, a preset reply message is sent to the dialogue assistant to guide the user to initiate a new topic and improve the dialogue experience between the user and the dialogue assistant.

[0082] In this embodiment, if it is determined that the original reply characters with passed content review are located in any final reply message, that is, the content review of any first final reply message passes, then the next first final reply message of this message is jumped to as the processing object, and the operations of content review and triggering the dialogue assistant to output messages are repeatedly executed for the processing object until the last final reply message is sent to the dialogue assistant and the content review passes. In this case, the dialogue assistant can output the characters in each received final reply message one by one on the screen, presenting the effect of a typewriter outputting one character at a time and improving the user's browsing experience.

[0083] It can be seen that through this embodiment, when the content review of each final reply message passes, all the final reply messages can be sent to the dialogue assistant, so that the user can view the reply messages for the dialogue messages.

[0084] In a detailed example, in this embodiment, after the server sends the dialogue message to the generative model, it can successively receive each original reply character sent by the generative model, and starting from receiving the first original reply character, use the character truncation window to accumulate the received characters to generate the final reply message. Therefore, the server can continuously receive the original reply characters and continuously generate the final reply messages.

[0085] After the server generates the first final reply message, if it is determined that the content review of the first final reply message fails, and if the generative model is still outputting the original reply characters at this time, or the server has not generated the final reply message based on all the original reply characters, then the server terminates the acquisition of the original reply characters and terminates the generation of the final reply message, and outputs a preset reply message to the dialogue assistant to end the current dialogue between the dialogue assistant and the user, and the preset reply message is output by the dialogue assistant.

[0086] If it is determined that the content review of the first final reply message is passed, the server sends the first final reply message to the terminal device where the dialogue assistant is located, and the dialogue assistant displays each character in the first final reply message on the screen one by one, bringing the user an experience of streaming output of each character.

[0087] Next, after the server generates the second final reply message, it sends the second final reply message to the terminal device where the dialogue assistant is located, and the dialogue assistant displays each character in the second final reply message on the screen one by one, bringing the user an experience of streaming output of each character.

[0088] After the server starts to send the second final reply message, it also determines whether the content review of the second final reply message is passed. If it passes, after the server generates the third final reply message, it sends the third final reply message to the terminal device where the dialogue assistant is located, and the dialogue assistant displays each character in the third final reply message on the screen one by one. Also, after the server starts to send the third final reply message, it also determines whether the content review of the third final reply message is passed. The server loops this process until the last final reply message is sent to the dialogue assistant and the content review is passed.

[0089] If the server determines that the content review of any final reply message after the second (including the second) fails, and if the generative model is still outputting the original reply characters at this time, or the server has not yet generated the final reply message based on all the original reply characters, the server terminates obtaining the original reply characters and terminates generating the final reply message, triggering the dialogue assistant to withdraw all the final reply messages that have been displayed (including the first final reply message and non-first final reply messages), and the messages will be hidden on the screen after being withdrawn. Also, the server outputs a preset reply message to the dialogue assistant to end the current conversation between the dialogue assistant and the user, and the preset reply message is output by the dialogue assistant.

[0090] Considering that it takes a certain amount of time for the dialogue assistant to obtain the original reply characters sent by the server and display them in a streaming manner, all the final reply messages that the server has sent to the dialogue assistant can be withdrawn here, and the server can also trigger the dialogue assistant to withdraw all the final reply messages that have been displayed, so that the user no longer browses the final reply messages.

[0091] It can be seen that in this example, the four processes of obtaining the original response characters, generating the final response message, sending the final response message, and detecting the final response message are carried out in a streaming manner, without waiting for the previous process to end completely before starting the next process. As a result, the conversation assistant can receive each final response message sent by the server in a streaming manner, and then the conversation assistant can display each character in the final response message on the screen one by one. This can greatly shorten the speed at which the user browses the response to the conversation message, reduce the latency of the user browsing the first character in the response, improve the conversation fluency between the user and the conversation assistant, and improve the user's conversation experience.

[0092] In one embodiment, triggering the conversation assistant to stream and output the obtained original response characters includes: Using a specific remote call framework to create a second remote call channel with the conversation assistant; the type of the second remote call channel is a streaming channel for streaming transmission; Through the second remote call channel, stream and send the obtained original response characters to the conversation assistant to trigger the conversation assistant to stream and output the obtained original response characters.

[0093] In this embodiment, first, use a specific remote call framework such as the KiteX framework to create a second remote call channel between the server and the terminal device where the conversation assistant is located, such as a mobile phone. The type of this channel is a streaming channel for streaming transmission. Specific remote call frameworks can be used to create streaming interfaces for the server and the terminal device respectively. Based on the streaming interface of the server and the streaming interface of the terminal device, create a streaming channel between the server and the terminal device as the second remote call channel. The streaming channel allows the server to return multiple final response messages based on a single conversation message sent by the terminal device. Different from the streaming channel is the non-streaming channel, which only allows the server to return a single response message based on a single conversation message sent by the terminal device.

[0094] The terminal device sends a conversation message to the server through this second remote call channel. The server uses this second remote call channel to obtain the conversation message, generates one or more final response messages through the above process, and based on the type of the second remote call channel, sends each generated final response message to the conversation assistant through the second remote call channel after each final response message is generated, so as to achieve the effect of streaming and sending each final response message to the conversation assistant, that is, streaming and sending each original response character.

[0095] In this embodiment, in streaming transmission, the terminal device only needs to send a single conversation message to the server, and the server can return one or more final response messages until the last final response message is returned and the conversation message is replied.

[0096] The dialogue assistant in the terminal device can, under the trigger of the server, output each received original reply character one by one on the screen in sequence (i.e., stream output), presenting the effect of a typewriter outputting one character at a time, and improving the user's browsing experience.

[0097] It can be seen that through this embodiment, a second remote call channel with the dialogue assistant can be created by using a specific remote call framework. The type of the second remote call channel is a streaming channel for streaming transmission. Through the second remote call channel, the obtained original reply characters are sent to the dialogue assistant in a streaming manner, triggering the dialogue assistant to stream output the obtained original reply characters, without waiting for all the original reply characters of the generative model to be output and then the dialogue assistant outputting all the original reply characters at once, reducing the time for the user to wait for the reply message and improving the dialogue efficiency between the user and the dialogue assistant.

[0098] In some embodiments, the above method process further includes: For the original reply characters that have been streamed output by the dialogue assistant, if the content of the output original reply characters passes the content review, obtain the audio data corresponding to the output original reply characters, and send the audio data to the dialogue assistant.

[0099] In this embodiment, for the original reply characters that have been output by the dialogue assistant to the user, if the content of the original reply characters passes the content review, obtain the audio data corresponding to the original reply characters, and send the audio data to the dialogue assistant, so that the dialogue assistant can play the audio data during the process of displaying the original reply characters.

[0100] In one case, after the content review of each final reply message passes and triggers the dialogue assistant to stream output each final reply message, for example, after the content review of all the final reply messages passes and all the final reply messages are sent to the dialogue assistant in a streaming manner, or, after the content review of all the final reply messages passes and the dialogue assistant has displayed all the final reply messages on the screen, obtain the first audio data corresponding to all the final reply messages, and have the dialogue assistant play the first audio data, so that the user can listen to the audio data of the final reply messages.

[0101] Considering that the text length of all the final reply messages is relatively long, in one embodiment, the server determines whether the text length of all the final reply messages is greater than a length threshold. If it is greater, the server takes all the final reply messages as the processing object, divides the processing object into multiple message segments, and generates first audio data corresponding to each message segment respectively. The terminal device can, after determining that the content review of each final reply message has passed and the dialogue assistant has received each final reply message, request the first audio data corresponding to one message segment from the server each time, and obtain each first audio data through multiple requests.

[0102] If it is not greater, then generate the first audio segment corresponding to all the final reply messages without splitting each final reply message. After the server determines that the content review of each final reply message has passed and has streamed each final reply message to the dialogue assistant, the server sends the first audio data to the dialogue assistant, and there is no need for the terminal device to actively request the audio data.

[0103] In this embodiment, in another case, for the final reply messages in which the dialogue assistant has been triggered to stream output among each final reply message, for example, for the final reply messages that have been output to the dialogue assistant among each final reply message, if the content of the output final reply message passes the content review, the server generates second audio data corresponding to the output final reply message and sends the second audio data to the dialogue assistant, and the dialogue assistant can play the second audio data. In this case, it is not necessary to wait until all the final reply messages are sent to the dialogue assistant. Each time it is determined that a final reply message has been sent to the dialogue assistant and the content of the message passes the content review, the second audio data of the message is generated and the dialogue assistant plays the second audio data of the message, achieving the effect that the user can browse the final reply message while listening to the audio data of the final reply message.

[0104] It can be seen that through this embodiment, the audio data of the original reply characters can also be output to the dialogue assistant, improving the user experience of browsing the original reply characters.

[0105] Figure 4 The flowchart of the dialogue processing method provided by another embodiment of the present disclosure is as Figure 4 shown, and this process includes: Step S402, the server receives a dialogue message for the dialogue assistant in the e - book and inputs the dialogue message into the generative model corresponding to the dialogue assistant.

[0106] Step S404, the server returns a prompt message to the dialogue assistant, and the prompt message is used to prompt the user that the reply is being generated.

[0107] The prompt message can be, for example, a waiting animation, which is used to prompt the user that the dialogue assistant is generating a reply.

[0108] Step S406, the server sequentially obtains at least one original reply character streamed out by the generative model for the conversation message.

[0109] Step S408, the server uses the obtained original reply characters to stream-generate at least one final reply message.

[0110] Step S410, the server determines whether the first final reply message passes the content review.

[0111] If it passes, execute step S412; if not, execute step S420.

[0112] Step S412, the server sends the first final reply message to the conversation assistant, triggering the conversation assistant to stream-output the first final reply message to the screen.

[0113] Step S414, the server determines whether there are non-first final reply messages that have not been sent.

[0114] If there are, execute step S416; if not, end the process.

[0115] Step S416, the server streams and sends each generated non-first final reply message to the conversation assistant, triggering the conversation assistant to stream-output the non-first final reply message to the screen.

[0116] Step S418, the server determines whether the non-first final reply messages that have been started to be sent pass the content review.

[0117] If it passes, return to step S414 and repeat the execution; if not, execute step S422.

[0118] Step S420, the server terminates obtaining each original reply character and terminates generating the final reply message, and sends a preset reply message to the conversation assistant.

[0119] The conversation assistant outputs the preset reply message to the screen.

[0120] Step S422, the server terminates obtaining each original reply character and terminates generating the final reply message, triggers the conversation assistant to withdraw all the final reply messages that have been displayed, and sends a preset reply message to the conversation assistant.

[0121] The conversation assistant outputs the preset reply message to the screen. The withdrawn messages are hidden on the screen.

[0122] Pass Figure 4The process can achieve the effect of triggering each final reply message corresponding to the streaming output of the conversation assistant's conversation messages, without waiting for all the original reply characters of the generative model to be output and then the conversation assistant outputting all the original reply characters at once, reducing the time-consuming for the user to wait for the reply message and improving the conversation efficiency between the user and the conversation assistant.

[0123] Figure 5 This is a schematic flowchart of the process for outputting the audio data of the final reply message provided by an embodiment of the present disclosure. As Figure 5 shown, this process includes: Step S502, the server determines that the content review of all the final reply messages has passed and has streamed all the final reply messages to the conversation assistant.

[0124] Step S504, the server determines whether the text length of all the final reply messages is greater than the length threshold.

[0125] If it is greater, execute Step S506; if not, execute Step S512.

[0126] Step S506, the server takes all the final reply messages as the processing object and divides the processing object into multiple message segments.

[0127] Step S508, the server generates the first audio data corresponding to each message segment respectively.

[0128] Step S510, according to the polling request of the terminal device where the conversation assistant is located, the server returns the first audio data to the terminal device in multiple times, and returns the first audio data corresponding to one message segment each time.

[0129] The terminal device can start playing the first audio data immediately after receiving the first audio data corresponding to one message segment each time, realizing the streaming playback of the audio data.

[0130] Step S512, the server generates the first audio segment corresponding to all the final reply messages and sends the first audio data to the conversation assistant.

[0131] The terminal device can start playing the first audio data immediately after receiving the first audio data.

[0132] Through Figure 5 the process in, the audio data of the final reply message can be output to the conversation assistant, improving the user experience of browsing the final reply message.

[0133] Figure 6a This is a schematic diagram of the conversation assistant streaming outputting the final reply message provided by an embodiment of the present disclosure, Figure 6b This is a schematic diagram of the conversation assistant streaming outputting the final reply message provided by another embodiment of the present disclosure. AsFigure 6a and Figure 6b As shown in Figure 6b , for the reply generated for the conversation message "I wonder how to lift the curse on you?", it can be output to the user character by character in a streaming manner, so that the user can view the corresponding reply message "Well, first you need to open the gate to the underworld and get the treasure." Here, the reply message is divided into two final reply messages and sent to the conversation assistant, namely "Well, first you need to" and "open the gate to the underworld and get the treasure." The conversation assistant can receive each final reply message in a streaming manner and output the characters in the received final reply message one by one on the screen (i.e., streaming output), presenting the effect of a typewriter outputting one character at a time, reducing the time the user waits for the reply message, and improving the conversation efficiency between the user and the conversation assistant.

[0134] In summary, through the above embodiments, on the one hand, it achieves the effect of triggering the conversation assistant to stream and output the reply message of the conversation message sent by the user, and on the other hand, it achieves the effect of content auditing of the reply message during the process of streaming and outputting the reply message. It is not necessary to wait for all the original reply characters of the generative model to be output and content audited before the conversation assistant outputs the reply message of the conversation message at one time, reducing the time the user waits for the reply message and improving the conversation efficiency between the user and the conversation assistant.

[0135] Figure 7 The following is a schematic structural diagram of a conversation processing device provided by an embodiment of the present disclosure. As Figure 7 shown, the device includes: A message acquisition unit 71, configured to acquire a conversation message for the conversation assistant and input the conversation message into the generative model corresponding to the conversation assistant; the conversation assistant is set in the e-book and is used to conduct a conversation for the content in the e-book; A reply acquisition unit 72, configured to sequentially acquire each original reply character output by the generative model for the conversation message; the original reply characters are output by the generative model in a streaming output manner one after another; A reply transmission unit 73, configured to perform content auditing on the acquired original reply characters according to the acquisition order of the original reply characters and trigger the conversation assistant to stream and output the acquired original reply characters; A reply processing unit 74, configured to, if it detects an original reply character that passes the content auditing, continue to perform content auditing on the subsequent characters of the original reply character, and if it detects an original reply character that fails to pass the content auditing, trigger the conversation assistant to withdraw the output original reply character and trigger the conversation assistant to output a preset reply message.

[0136] Optionally, the reply obtaining unit 72 is specifically configured to: create a first remote call channel with the generative model by using a specific remote call framework; the type of the first remote call channel is a streaming channel for streaming transmission; and obtain each of the original reply characters in sequence through the first remote call channel according to the sequence of streaming output of the original reply characters by the generative model.

[0137] Optionally, the reply transmission unit 73 is specifically configured to: generate at least one final reply message according to the obtaining order of the original reply characters; and perform content review on the final reply message according to the generation order of the final reply message and trigger the dialogue assistant to stream output the final reply message.

[0138] Optionally, the reply transmission unit 73 is further specifically configured to: obtain a character truncation window; the character length that the character truncation window can accommodate is the character length corresponding to the final reply message; use the character truncation window to truncate the obtained original reply characters once or multiple times according to the obtaining order of the original reply characters until the last original reply character is truncated; wherein the last character truncated by the previous character truncation window is the character before the first character truncated by the next character truncation window; and generate one final reply message by using the characters accommodated in each character truncation window.

[0139] Optionally, the reply transmission unit 73 is further specifically configured to: perform content review on the first generated final reply message, and trigger the dialogue assistant to stream output the first final reply message after the content review passes; for the non-first generated final reply message, trigger the dialogue assistant to stream output the non-first final reply message, and perform content review on the non-first final reply message after triggering the dialogue assistant.

[0140] Optionally, the reply transmission unit 73 is further specifically configured to: detect whether the first final reply message contains a specified character through a content review model, and detect whether the semantic information of the first final reply message meets the semantic requirements; if the first final reply message does not contain the specified character and the semantic information of the first final reply message meets the semantic requirements, determine that the content review of the first final reply message passes, otherwise, determine that the content review of the first final reply message fails.

[0141] Optionally, the reply transmission unit 73 is further specifically configured to: generate a target message based on the non-first final reply message; the target message includes the non-first final reply message and the final reply messages generated before the non-first final reply message; detect whether the target message contains a specified character through a content review model, and detect whether the semantic information of the target message meets the semantic requirements; if the target message does not contain a specified character and the semantic information of the target message meets the semantic requirements, it is determined that the content review of the non-first final reply message passes, otherwise, it is determined that the content review of the non-first final reply message fails.

[0142] Optionally, the reply processing unit 74 is specifically configured to: if the original reply character that fails the content review is in the first generated final reply message, trigger the dialogue assistant to output the preset reply message; wherein, the dialogue assistant has not output the first final reply message; if the original reply character that fails the content review is in the non-first final reply message generated, trigger the dialogue assistant to withdraw all the final reply messages that have been output, and trigger the dialogue assistant to output the preset reply message; wherein, the dialogue assistant has been triggered to output the non-first final reply message.

[0143] Optionally, the reply transmission unit 73 is specifically configured to: create a second remote call channel with the dialogue assistant by using a specific remote call framework; the type of the second remote call channel is a streaming channel for streaming transmission; through the second remote call channel, stream the obtained original reply characters to the dialogue assistant, and trigger the dialogue assistant to stream output the obtained original reply characters.

[0144] Optionally, it further includes a voice sending unit, configured to: for the original reply characters that have been streamed output by the dialogue assistant, if the content review of the output original reply characters passes, obtain the audio data corresponding to the output original reply characters, and send the audio data to the dialogue assistant.

[0145] Through this embodiment, on the one hand, it achieves the effect of triggering the dialogue assistant to stream output the reply message of the dialogue message sent by the user, and on the other hand, it achieves the effect of performing content review on the reply message during the process of streaming outputting the reply message, without waiting for all the original reply characters of the generative model to be output and then performing content review and then the dialogue assistant outputting the reply message of the dialogue message at one time, reducing the time-consuming for the user to wait for the reply message and improving the dialogue efficiency between the user and the dialogue assistant.

[0146] The dialogue processing device in the embodiments of the present disclosure can implement each process of the above-mentioned dialogue processing method embodiments and achieve the same effects and functions, which will not be repeated here.

[0147] An embodiment of the present disclosure further provides an electronic device. Figure 8 The following is a schematic structural diagram of the electronic device provided by an embodiment of the present disclosure. As Figure 8 shown, the electronic device may have relatively large differences due to different configurations or performances, and may include one or more processors 801 and a memory 802. One or more applications or data may be stored in the memory 802. Among them, the memory 802 may be short-term storage or persistent storage. The applications stored in the memory 802 may include one or more modules (not shown in the figure), and each module may include a series of computer-executable instructions in the electronic device. Further, the processor 801 may be set to communicate with the memory 802 and execute a series of computer-executable instructions in the memory 802 on the electronic device. The electronic device may also include one or more power supplies 803, one or more wired or wireless network interfaces 804, one or more input or output interfaces 805, one or more keyboards 806, etc.

[0148] In a specific embodiment, the electronic device includes a processor; and a memory configured to store computer-executable instructions, and when the computer-executable instructions are executed, the processor implements the following processes: Obtain a dialogue message for the dialogue assistant, and input the dialogue message into the generative model corresponding to the dialogue assistant; the dialogue assistant is set in the e-book and is used for having a dialogue regarding the content in the e-book; Successively obtain each original reply character output by the generative model for the dialogue message; the original reply characters are successively output by the generative model in a streaming output manner; According to the acquisition order of the original reply characters, conduct content review on the obtained original reply characters and trigger the dialogue assistant to stream output the obtained original reply characters; If it is detected that the original reply character passes the content review, continue to conduct content review on the subsequent characters of the original reply character. If it is detected that the original reply character fails to pass the content review, trigger the dialogue assistant to withdraw the originally output original reply character and trigger the dialogue assistant to output a preset reply message.

[0149] Through this embodiment, on the one hand, it achieves the effect of triggering the dialogue assistant to stream out the reply message of the dialogue message sent by the user, and on the other hand, it achieves the effect of content auditing of the reply message during the process of streaming out the reply message. It is not necessary to wait for all the original reply characters of the generative model to be output and then undergo content auditing before the dialogue assistant outputs the reply message of the dialogue message at one time, reducing the time consumed by the user waiting for the reply message and improving the dialogue efficiency between the user and the dialogue assistant.

[0150] The electronic device in the embodiments of the present disclosure can implement each process of the above-mentioned dialogue processing method embodiment and achieve the same effects and functions, which will not be repeated here.

[0151] Another embodiment of the present disclosure also provides a computer-readable storage medium, which is used to store computer-executable instructions. When the computer-executable instructions are executed by a processor, the following processes are implemented: Obtain the dialogue message for the dialogue assistant, and input the dialogue message into the generative model corresponding to the dialogue assistant; the dialogue assistant is set in the e-book and is used to conduct a dialogue on the content in the e-book; Successively obtain each original reply character output by the generative model for the dialogue message; the original reply characters are successively output by the generative model in a streaming output manner; According to the acquisition order of the original reply characters, conduct content auditing on the obtained original reply characters and trigger the dialogue assistant to stream out the obtained original reply characters; If it is detected that the original reply character passes the content auditing, continue to conduct content auditing on the subsequent characters of the original reply character. If it is detected that the original reply character fails to pass the content auditing, trigger the dialogue assistant to withdraw the already output original reply character and trigger the dialogue assistant to output a preset reply message.

[0152] Through this embodiment, on the one hand, it achieves the effect of triggering the dialogue assistant to stream out the reply message of the dialogue message sent by the user, and on the other hand, it achieves the effect of content auditing of the reply message during the process of streaming out the reply message. It is not necessary to wait for all the original reply characters of the generative model to be output and then undergo content auditing before the dialogue assistant outputs the reply message of the dialogue message at one time, reducing the time consumed by the user waiting for the reply message and improving the dialogue efficiency between the user and the dialogue assistant.

[0153] The computer-readable storage medium in the embodiments of the present disclosure can implement each process of the above-mentioned dialogue processing method embodiment and achieve the same effects and functions, which will not be repeated here.

[0154] Another embodiment of the present disclosure further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the following processes are implemented: Obtain a conversation message for the conversation assistant, and input the conversation message into the generative model corresponding to the conversation assistant; the conversation assistant is set in the e-book and is used to conduct conversations on the content in the e-book; Successively obtain each original reply character output by the generative model for the conversation message; the original reply characters are successively output by the generative model in a streaming output manner; According to the acquisition order of the original reply characters, conduct content review on the already acquired original reply characters and trigger the conversation assistant to stream output the already acquired original reply characters; If it is detected that the original reply character passes the content review, continue to conduct content review on the subsequent characters of the original reply character. If it is detected that the original reply character fails to pass the content review, trigger the conversation assistant to withdraw the already output original reply character and trigger the conversation assistant to output a preset reply message.

[0155] Through this embodiment, on the one hand, it achieves the effect of triggering the conversation assistant to stream output the reply message of the conversation message sent by the user, and on the other hand, it achieves the effect of conducting content review on the reply message during the process of streaming outputting the reply message, without waiting for all the original reply characters of the generative model to be output and then conducting content review and then having the conversation assistant output the reply message of the conversation message at one time, reducing the time consumed by the user waiting for the reply message and improving the conversation efficiency between the user and the conversation assistant.

[0156] The computer program product in the embodiments of the present disclosure can implement each process of the above-mentioned conversation processing method embodiment and achieve the same effects and functions, which will not be repeated here.

[0157] In each embodiment of the present disclosure, the computer-readable storage medium includes a read-only memory (ROM for short), a random access memory (RAM for short), a magnetic disk, an optical disc, etc.

[0158] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structures of diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logical function is determined by a user's programming of the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, today, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one type of HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain a hardware circuit that implements the logical method flow.

[0159] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium that stores computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, application specific integrated circuit (ASIC), programmable logic controller, and embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0160] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0161] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the embodiments of the present disclosure, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0162] Those skilled in the art should understand that one or more embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.

[0163] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0164] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0165] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0166] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0167] One or more embodiments of the present disclosure may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present disclosure may also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media including storage devices.

[0168] Each embodiment in the present disclosure is described in a progressive manner. For the same or similar parts among the embodiments, reference may be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference may be made to the corresponding parts of the method embodiments for the relevant content.

[0169] The above description is only for the embodiments of the present disclosure and is not intended to limit the present disclosure. For those skilled in the art, various changes and modifications can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the scope of the claims of the present disclosure.

Claims

1. A method for dialogue processing, characterized in that, Including: Obtain a conversation message for the conversation assistant, and input the conversation message into the generative model corresponding to the conversation assistant; The conversation assistant is set in the e-book and is used to conduct a conversation regarding the content in the e-book; Sequentially obtain each original reply character output by the generative model for the conversation message; The original reply characters are successively output by the generative model in a streaming output manner; According to the acquisition order of the original reply characters, conduct content review on the obtained original reply characters and trigger the conversation assistant to stream output the obtained original reply characters; If it is detected that the original reply character passes the content review, continue to conduct content review on the subsequent characters of the original reply character. If it is detected that the original reply character fails to pass the content review, trigger the conversation assistant to withdraw the output original reply character and trigger the conversation assistant to output a preset reply message.

2. The method according to claim 1, characterized in that The sequentially obtaining each original reply character output by the generative model for the conversation message includes: Use a specific remote call framework to create a first remote call channel with the generative model; the type of the first remote call channel is a streaming channel for streaming transmission; Through the first remote call channel, sequentially obtain each of the original reply characters in the order in which the generative model streams out the original reply characters.

3. The method according to claim 1, characterized in that, The conducting content review on the obtained original reply characters according to the acquisition order of the original reply characters and triggering the conversation assistant to stream output the obtained original reply characters includes: Generate at least one final reply message according to the acquisition order of the original reply characters; According to the generation order of the final reply messages, conduct content review on the final reply messages and trigger the conversation assistant to stream output the final reply messages.

4. The method according to claim 3, characterized in that, The generating at least one final reply message according to the acquisition order of the original reply characters includes: Obtain a character truncation window; the character length that the character truncation window can accommodate is the character length corresponding to the final reply message; Use the character truncation window to truncate the obtained original reply characters one or more times according to the acquisition order of the original reply characters until the last original reply character is truncated; wherein, the last character truncated by the previous character truncation window is the character before the first character truncated by the next character truncation window; Generate one final reply message using the characters contained in each character truncation window.

5. The method according to claim 3, characterized in that, The conducting content review on the final reply messages according to the generation order of the final reply messages and triggering the conversation assistant to stream output the final reply messages includes: For the first generated final reply message, conduct content review on the first final reply message. After the content review passes, trigger the conversation assistant to stream output the first final reply message; For the generated non-first final reply message, trigger the dialogue assistant to stream output the non-first final reply message. After triggering the dialogue assistant, conduct content review on the non-first final reply message.

6. The method according to claim 5, characterized in that, The content review of the first final reply message includes: Using a content review model, detect whether the first final reply message contains specified characters, and detect whether the semantic information of the first final reply message meets the semantic requirements; If the first final reply message does not contain specified characters and the semantic information of the first final reply message meets the semantic requirements, determine that the content review of the first final reply message passes; otherwise, determine that the content review of the first final reply message fails.

7. The method according to claim 5, wherein The content review of the non-first final reply message includes: Based on the non-first final reply message, generate a target message; the target message includes the non-first final reply message and the final reply messages generated before the non-first final reply message; Using a content review model, detect whether the target message contains specified characters, and detect whether the semantic information of the target message meets the semantic requirements; If the target message does not contain specified characters and the semantic information of the target message meets the semantic requirements, determine that the content review of the non-first final reply message passes; otherwise, determine that the content review of the non-first final reply message fails.

8. The method according to claim 3, characterized in that If the detected original reply characters fail the content review, trigger the dialogue assistant to withdraw the output original reply characters and trigger the dialogue assistant to output a preset reply message, including: If the original reply characters that fail the content review are in the generated first final reply message, trigger the dialogue assistant to output the preset reply message; where the dialogue assistant has not output the first final reply message; If the original reply characters that fail the content review are in the generated non-first final reply message, trigger the dialogue assistant to withdraw all the output final reply messages and trigger the dialogue assistant to output the preset reply message; where the dialogue assistant has been triggered to output the non-first final reply message.

9. The method according to claim 1, wherein The triggering of the dialogue assistant to stream output the obtained original reply characters includes: Using a specific remote call framework, create a second remote call channel with the dialogue assistant; the type of the second remote call channel is a streaming channel for streaming transmission; Through the second remote call channel, stream send the obtained original reply characters to the dialogue assistant to trigger the dialogue assistant to stream output the obtained original reply characters.

10. The method according to claim 1, wherein It also includes: For the original reply characters that have been streamed output by the dialogue assistant, if the content review of the output original reply characters passes, obtain the audio data corresponding to the output original reply characters and send the audio data to the dialogue assistant.

11. A dialogue processing device, characterized in that, It includes: A message acquisition unit, configured to acquire a conversation message for a conversation assistant and input the conversation message into a generative model corresponding to the conversation assistant; The conversation assistant is set in an e-book and is configured to conduct a conversation on the content in the e-book; A reply acquisition unit, configured to sequentially acquire each original reply character output by the generative model for the conversation message; the original reply characters are successively output by the generative model in a streaming output manner; A reply transmission unit, configured to, according to the acquisition order of the original reply characters, conduct content review on the acquired original reply characters and trigger the conversation assistant to stream output the acquired original reply characters; A reply processing unit, configured to, if it detects an original reply character that passes the content review, continue to conduct content review on the subsequent characters of the original reply character, and if it detects an original reply character that fails to pass the content review, trigger the conversation assistant to withdraw the output original reply character and trigger the conversation assistant to output a preset reply message.

12. An electronic device, characterized in that, Comprising: A processor; And A memory configured to store computer-executable instructions, and when the computer-executable instructions are executed, the processor implements the method according to any one of claims 1-10 above.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store computer-executable instructions, and when the computer-executable instructions are executed by a processor, the method according to any one of claims 1-10 above is implemented.

14. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1-10 above is implemented.