Message processing method and apparatus, device, and storage medium

By receiving user input messages, generating streaming text, and splitting it into multiple text parts, the problem of single-content responses in traditional models is solved, enabling the generation of multiple response messages and improving user experience and efficiency.

WO2026102728A1PCT designated stage Publication Date: 2026-05-21BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-11-15
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Traditional model-based entity processing can only generate a single response, which is insufficient to meet users' dialogue needs.

Method used

It receives user input messages, generates streaming text, and then parses the streaming text to split it into multiple text parts, providing multiple reply messages.

Benefits of technology

It generates smoother and more natural response content, improves message processing efficiency and user experience, and meets users' dialogue needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024132470_21052026_PF_FP_ABST
    Figure CN2024132470_21052026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a message processing method and apparatus, a device, and a storage medium. The method provided herein comprises: receiving a first input message of a user; acquiring streaming text generated by a model on the basis of the first input message; splitting the streaming text into a plurality of text portions by parsing the streaming text; and providing, to the user, a plurality of reply messages corresponding to the plurality of text portions. By means of said method, the embodiments of the present disclosure can improve the efficiency of message processing.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, devices and storage media for message processing Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to a method, apparatus, device, and computer-readable storage medium for message processing. Background Technology

[0002] In recent years, with the development of the internet, more and more users are engaging in interactive activities on online platforms. For example, users can interact with model-based processing entities through dialogue. However, traditional model-based processing entities can only generate a single response to user messages, which is insufficient to meet users' dialogue needs. Summary of the Invention

[0003] In a first aspect of this disclosure, a message processing method is provided, comprising: receiving a first input message from a user; obtaining streaming text generated by a model based on the first input message; splitting the streaming text into multiple text parts by parsing the streaming text; and providing the user with multiple reply messages corresponding to the multiple text parts.

[0004] In a second aspect of this disclosure, an apparatus for message processing is provided. The apparatus includes: a receiving module configured to receive a first input message from a user; an acquiring module configured to acquire streaming text generated by a model based on the first input message; a parsing module configured to split the streaming text into multiple text parts by parsing the streaming text; and a providing module configured to provide the user with multiple reply messages corresponding to the multiple text parts.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0007] It should be understood that the content described in this text section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0009] Figure 1 shows a schematic diagram of an example environment capable of real-time implementation of some embodiments of this disclosure;

[0010] Figure 2 illustrates a flowchart of an example message processing procedure according to some embodiments of the present disclosure;

[0011] Figures 3A and 3B illustrate example interfaces according to some embodiments of the present disclosure;

[0012] Figure 4 shows a schematic structural block diagram of an example apparatus for message processing according to some embodiments of the present disclosure; and

[0013] Figure 5 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0016] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0017] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0018] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0019] As briefly mentioned above, with the development of the internet, more and more users are engaging in interactive activities on online platforms. For example, users can interact with model-based processing entities on online platforms, such as through dialogue. However, traditional model-based processing entities can only generate a single response to user messages, which is insufficient to meet users' dialogue needs.

[0020] Embodiments of this disclosure propose a scheme for message processing. According to this scheme, a user's first input message can be received; streaming text generated by a model based on the first input message can be obtained; the streaming text can be parsed and split into multiple text parts; and multiple reply messages corresponding to the multiple text parts can be provided to the user.

[0021] In this way, embodiments of this disclosure can generate streaming text based on received user input messages. Furthermore, embodiments of this disclosure can split the streaming text into multiple text parts to provide multiple reply messages to the user. Thus, embodiments of this disclosure can generate multiple reply messages for a single user input message, overcoming the problem of rigid and monotonous reply content in traditional model-based responses, allowing users to receive more fluent and natural content replies, and meeting the user's dialogue needs. In addition, embodiments of this disclosure can generate splittable streaming text, improving message processing efficiency.

[0022] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0023] Example Environment

[0024] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in Figure 1, the example environment 100 may include an electronic device 110.

[0025] In this example environment 100, electronic device 110 can run an application 120 that supports user interface interaction. Application 120 can be any suitable type of application for messaging, examples of which may include, but are not limited to, conversational applications, social applications, or other suitable applications. User 140 can interact with application 120 via electronic device 110 and / or its attached devices.

[0026] In environment 100 of Figure 1, if application 120 is active, electronic device 110 can use application 120 to present interface 150 for supporting interface interaction.

[0027] In some embodiments, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).

[0028] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for applications 120 that support user interface interaction in electronic devices 110.

[0029] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections; the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, server 130 and electronic device 110 can achieve signaling interaction through the communication connection between them.

[0030] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0031] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0032] Example process

[0033] Figure 2 illustrates a flowchart of an example process 200 for message processing according to some embodiments of the present disclosure. Process 200 may be implemented at electronic device 110 or server 130. Process 200 will now be described with reference to Figure 1.

[0034] As shown in the figure, in box 210, electronic device 110 or server 130 receives the user's first input message.

[0035] The following description of an example message processing procedure 200 according to an embodiment of the present disclosure will refer to the interactive interfaces shown in Figures 3A and 3B.

[0036] Figures 3A and 3B illustrate example interfaces 300A and 300B according to some embodiments of the present disclosure. Interfaces 300A and 300B may be provided, for example, by the electronic device 110 shown in Figure 1.

[0037] In some embodiments, as shown in FIG3A, the electronic device 110 may present an interactive interface 300A. As an example, the interactive interface 300A may be an interactive interface (e.g., a conversational interface) between a user and a model-based processing entity. As an example, the model-based processing entity may be implemented as a virtual object that supports interaction via prompts or similar means.

[0038] In some scenarios, virtual objects may also be referred to as virtual entities (or digital assistants, etc.), and examples may include, but are not limited to, agents, bots, or other objects that support interaction through dialogue (e.g., text or voice dialogue). Such virtual objects can be implemented using appropriate machine learning models to interact accordingly based on received prompts. This disclosure is not intended to limit the specific construction of virtual objects. It should be understood that, on the one hand, virtual objects can have a virtual avatar for interaction. On the other hand, virtual objects can also utilize appropriate models to drive their interactive behavior, such models can include any appropriate machine learning model, for example, a generative model. In this disclosure, the various interactive processes and / or generative operations performed by virtual objects can be actually performed using appropriate models corresponding to the virtual object.

[0039] In some embodiments, the electronic device 110 may present identification information 305 associated with the model-based processing entity in the interactive interface 300A. For example, the identification information 305 may include image information (e.g., an avatar) and / or text information (e.g., the name "XX Dialogue Model") associated with the model-based processing entity. Alternatively, the electronic device 110 may also present descriptive information associated with the model-based processing entity in the interactive interface 300A. The descriptive information may, for example, describe the processing entity's conversational abilities (e.g., areas of expertise in conversation) and / or interactive capabilities.

[0040] It is understandable that the processing capabilities of model-based processing entities are provided by the model. For ease of description, the following text will directly describe the message processing process based on the model.

[0041] As an example, electronic device 110 or server 130 may receive a first input message from a user via natural language. For example, the first input message may include text content, voice content, and / or image content entered by the user. Electronic device 110 or server 130 may send the first input message to a model. This disclosure is not intended to limit the specific form of the first input message.

[0042] For example, electronic device 110 can present the user's first input message 302 in the interactive interface 300A.

[0043] Referring again to Figure 2, in box 220, electronic device 110 or server 130 acquires the streaming text generated by the model based on the first input message. Streaming text refers to the model (e.g., a language model) outputting text content as a continuous stream, rather than waiting for the entire processing to complete before outputting it all at once. For example, a language model can predict each token in a token sequence one by one, thus achieving token-by-token output.

[0044] In some embodiments, the streaming text may be generated by the model based on a system prompt and a first input message. As an example, the system prompt may indicate constraint information regarding the model's output. For instance, the developer of the agent or machine program can configure constraints on the model's output using a system prompt. Such a system prompt can serve as global configuration information for the agent and / or machine program associated with the model.

[0045] As an example, constraint information can instruct the model to use preset delimiters to separate different text sections. For example, preset delimiters may include line breaks, etc. For example, constraint information can instruct the model to add line breaks to the generated streaming text. In this way, server 130 can split the model-generated streaming text into different text sections based on line breaks, in order to send multiple reply messages corresponding to the multiple text sections to the user.

[0046] As an example, constraint information may indicate the length of the text portions output by the model. For instance, the constraint information could instruct the model to add preset delimiters to the streaming text based on a preset text length. For example, if the preset text length is 10 characters, the model could add a newline character every 10 characters to divide the streaming text into different text portions. Thus, based on constraint information including field length, the model can divide the generated streaming text into text portions that meet the preset text length. For example, multiple text portions in the streaming text could correspond to multiple reply messages that are short sentences of no more than 10 characters, making the reply messages easier for users to accept and understand, thus meeting the user's dialogue needs.

[0047] Alternatively, the constraint information may include semantic constraints. For example, the constraint information may instruct the model to add preset delimiters to the streaming text based on preset semantic types. As an example, the preset semantic types may include transition types, so the model can add line breaks based on the semantics of the transition type (e.g., words indicating transition, such as "but" or "however," can indicate that the current semantic type is transition) to separate the streaming text into text parts corresponding to different semantics. As an example, the streaming text generated by the model is, for example, "I like reading, but I prefer watching movies." Based on the constraint information including semantic constraints, the model can add line breaks between "I like reading" and "but I prefer watching movies" in this streaming text. Therefore, multiple reply messages corresponding to this streaming text may include, for example, "I like reading" and "but I prefer watching movies" presented on two lines.

[0048] Optionally, the constraint information may include contextual constraints. For example, the constraint information may instruct the model to add preset delimiters to the streaming text based on a preset context. As an example, the context may include a scenario that creates suspense or atmosphere. As an example, to make the generated reply message more atmospheric, the model may split long sentences into short sentences or individual words based on preset delimiters. For example, if the streaming text generated by the model is "Guess what surprise I prepared for you today," the multiple text parts obtained by adding line breaks based on the model may include, for example, "guess," "I," "today," "for you," "prepared," "what," and "surprise." In this way, presenting the reply message through word segmentation can make the reply message more atmospheric, thereby increasing the interest of the reply message.

[0049] Alternatively or additionally, constraint information may indicate the number of text segments output by the model. As an example, constraint information may instruct the model to add a first number (e.g., 2) preset delimiters to the streaming text to divide it into a second number (e.g., 3) text segments. For example, the streaming text generated by the model might be "I think pets like cats and dogs can improve our happiness." For instance, to satisfy the requirement of generating 3 reply messages, the multiple text segments obtained by adding line breaks based on the model could include, for example, "I think," "pets like cats and dogs," "can improve our happiness." Similarly, to satisfy the requirement of generating 4 reply messages, the multiple text segments obtained by adding line breaks based on the model could include, for example, "I think," "pets like cats and dogs," "can improve our happiness."

[0050] In some embodiments, constraint information can instruct the format of the streaming text output by the model. For example, constraint information can instruct the streaming text not to use first-class punctuation marks (e.g., not to use commas and / or periods, etc.). For example, constraint information can instruct the streaming text to use second-class punctuation marks (e.g., using spaces instead of commas and / or periods, etc.). As an example, based on this constraint information, the streaming text generated by the model could be something like "Wow, what books do you usually read?" or "By the way, do you have any pets?" Such streaming text, using spaces instead of at least one class of punctuation marks, is closer to the user's everyday conversation, thereby improving the user experience and meeting user needs.

[0051] Referring again to Figure 2, in box 230, electronic device 110 or server 130 splits the streaming text into multiple text parts by parsing the streaming text.

[0052] As an example, electronic device 110 or server 130 can split the streaming text into multiple text parts by detecting preset delimiters in the streaming text. For example, electronic device 110 can split the streaming text into three text parts in response to detecting two newline characters in the streaming text.

[0053] Alternatively or additionally, electronic device 110 or server 130 may determine the order of multiple text portions in a streaming text. For example, electronic device 110 may determine, based on line breaks in the streaming text, that a first order corresponding to text portions preceding a line break takes precedence over a second order corresponding to text portions following a line break.

[0054] Referring again to Figure 2, in box 240, electronic device 110 provides the user with multiple reply messages corresponding to multiple text segments. As an example, electronic device 110 can present multiple reply messages in a target order within the interactive interface. The target order can correspond to the order of the multiple text segments in the streaming text.

[0055] For example, in Figure 3A, the multiple reply messages presented by the electronic device 110 in the interactive interface 300A may include message 304, message 306 and message 308.

[0056] In some embodiments, multiple response messages may include a second response message and a third response message. For example, message 306 may correspond to the second response message. Message 308 may correspond to the third response message. As an example, electronic device 110 may provide the second response message (e.g., message 306) in interactive interface 300A. Further, electronic device 110 may provide the third response message (e.g., message 308) in interactive interface 300A after providing the second response message for a target duration. For example, the second response message may include "What books do you like to read?" The third response message may include "Do you usually like to exercise?" Electronic device 110 may continue to send "Do you usually like to exercise?" to the user 3 seconds after sending "What books do you like to read?".

[0057] Alternatively or additionally, the target duration can be a preset duration (e.g., 1 second) or a random duration (e.g., 0.5-3 seconds). Thus, embodiments of this disclosure send messages to the user sequentially based on the target duration, unlike the message processing method in traditional dialogue systems where messages are sent immediately upon generation. This allows users to have a more realistic dialogue experience, thereby meeting their needs.

[0058] In some embodiments, electronic device 110 or server 130 may, in response to receiving additional input from the user before providing the first response message (e.g., receiving a new input message after receiving the user's first input message), present referenced content associated with the first response message (e.g., including text content corresponding to the first input message). The referenced content may indicate that the first response message is a response to the first input message.

[0059] As an example, the aforementioned streaming text for the first input message can be the first streaming text. As shown in Figure 3B, the electronic device 110 or the server 130 can receive the user's second input message (e.g., message 310 (e.g., "What books do you usually read?"). The electronic device 110 can obtain the second streaming text generated by the model based on the second input message. Furthermore, the electronic device 110 or the server 130 can detect whether a third input message from the user has been received before providing the second streaming text to the user.

[0060] Alternatively or additionally, electronic device 110 or server 130 may stop providing the user with a reply message corresponding to the second streamed text in response to receiving a third input message (e.g., message 312) before providing the user with the second streamed text.

[0061] For example, electronic device 110 or server 130 may, in response to receiving a third input message (e.g., message 312 (e.g., “I always make a book list but forget to read it” or “Recommend a movie for me”)) before providing the second streaming text to the user, trigger the model to generate third streaming text based on the second input message (e.g., message 310) and the third input message (e.g., message 312).

[0062] Optionally, electronic device 110 or server 130 can determine the relevance between the second and third input messages (e.g., based on server 130 or a preset model). As an example, relevance may include the semantic relevance between the second and third input messages. As an example, semantic relevance can be used to measure the semantic closeness of two texts. The determination of semantic relevance can be based on existing machine learning algorithms for semantic analysis (e.g., word vector models, neural network-based semantic recognition models, etc.). This disclosure is not intended to limit the specific process or algorithm for determining the semantic relevance between texts.

[0063] Furthermore, electronic device 110 or server 130 can trigger a model to generate third streaming text based on the second and third input messages in response to the correlation between the two being greater than a threshold. For example, the text content corresponding to the second input message is "What books do you usually read?", and the text content corresponding to the third input message is "I always make a reading list but forget to read it". For example, a machine learning algorithm based on semantic analysis can determine that the correlation between the second and third input messages is greater than a threshold. This disclosure is not intended to limit the specific value or implementation of the threshold. As an example, the model can concatenate the second and third input messages (e.g., "What books do you usually read? I always make a reading list but forget to read it"), and use the concatenated content as input to generate third streaming text. As an example, the third streaming text may include, for example, "I like reading literature, so I'll remind you to read it together next time." The corresponding reply message for the third streaming text may include, for example, "I like reading literature" and "I'll remind you to read it together next time."

[0064] Optionally, the electronic device 110 or server 130 may, in response to the correlation between the two being no greater than a threshold, trigger the model to generate the first text corresponding to the second input message and the second text corresponding to the third input message, respectively, to determine the third streaming text based on the first and second texts. That is, the third streaming text generated by the model may include at least the first text corresponding to the second input message and the second text corresponding to the third input message. The model may separate the first and second texts into different text parts based on a preset delimiter (e.g., a newline character). For example, the text content corresponding to the second input message might be "What books do you usually read?", and the text content corresponding to the third input message might be "Recommend a movie for me". For example, a machine learning algorithm based on semantic analysis can determine that the correlation between the second and third input messages is no greater than a threshold.

[0065] As an example, electronic device 110 can acquire third-stream text generated based on a model. For example, the third-stream text may include phrases like "I usually enjoy reading literature and sometimes biographies; I recommend the movie XX." Furthermore, electronic device 110 can provide the user with a fourth reply message corresponding to the first text and a fifth reply message corresponding to the second text.

[0066] As an example, the fourth response message corresponding to the first text may include one or more messages. As an example, the number of fourth response messages may be determined by preset conditions. Preset conditions may include, for example, the number of messages presented in the interactive interface (which may include user-input messages and / or model-generated response messages). For example, electronic device 110 may determine that the fourth response message is multiple if the number of messages in the interactive interface is not greater than a preset number (e.g., less than 10). For example, the fourth response message may include message 314-1 (e.g., "I usually like to read literature books") and message 314-2 (e.g., "Sometimes I also read biographies"). Further, electronic device 110 may determine that the fourth response message is one if the number of book templates in the interactive interface meets a preset number (e.g., 10 or more). For example, the fourth response message may include "I usually like to read literature books and sometimes I also read biographies."

[0067] As an example, the fifth reply message corresponding to the second text may include one or more messages. As an example, the method for determining the number of fifth reply messages can refer to the process for determining the number of fourth reply messages, and will not be repeated here. For example, the fifth reply message may include message 316 (e.g., "I recommend movie XX").

[0068] Alternatively or additionally, electronic device 110 may present at least the referenced content associated with the fifth reply message. For example, electronic device 110 may present referenced content 320 in the interactive interface 300B. Referenced content 320 may include at least a portion of the third input message (e.g., message 312) (e.g., a reply: "Recommend a movie for me"). In this way, the referenced content associated with the fifth reply message can indicate that the fifth reply message is a response to the third input message.

[0069] Optionally, the electronic device 110 may present reference content associated with the fourth response message. For example, the electronic device 110 may present reference content 322-1 and / or reference content 322-2 in the interactive interface 300B. Reference content 322-1 and / or reference content 322-2 may include at least a portion of the second input message (e.g., message 310) (e.g., a reply: "What books do you usually read?"). In this way, the reference content associated with the fourth response message can indicate that the fourth response message is a response to the second input message. Thus, embodiments of this disclosure can generate different response messages based on different input messages and present reference content associated with the corresponding input message in the response message, so that the response message can clearly indicate a response to the corresponding input message, avoiding user confusion about the response message, improving the presentation efficiency of the response message, and meeting the user's needs.

[0070] Based on the process described above, streaming text is generated based on the received user input messages. Furthermore, embodiments of this disclosure can split the streaming text into multiple text parts to provide the user with multiple reply messages. Thus, embodiments of this disclosure can generate multiple reply contents for a single user input message, overcoming the problem of rigid and monotonous reply content in traditional model-based replies, allowing users to obtain more fluent and natural content replies, and meeting the user's dialogue needs. In addition, embodiments of this disclosure can generate different reply messages for different user input messages. Moreover, embodiments of this disclosure can present different reply messages that reference the corresponding input messages, avoiding user confusion about reply messages, improving the efficiency of reply message presentation, and meeting user needs.

[0071] Example devices and equipment

[0072] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 4 shows a schematic structural block diagram of an example apparatus 400 for message processing according to certain embodiments of this disclosure. Apparatus 400 may be implemented as or included in an electronic device. The various modules / components in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0073] As shown in Figure 4, the device 400 includes a receiving module 410 configured to receive a first input message from a user; an acquisition module 420 configured to acquire streaming text generated by a model based on the first input message; a parsing module 430 configured to split the streaming text into multiple text parts by parsing the streaming text; and a providing module 440 configured to provide the user with multiple reply messages corresponding to the multiple text parts.

[0074] In some embodiments, the providing module 440 is further configured to: in response to receiving a viewing request associated with the work, present an interactive panel associated with the work, the interactive panel presenting a set of descriptive information, the set of descriptive information including at least first descriptive information; and in response to the selection of the first descriptive information in the set of descriptive information, present an input panel and an information display area.

[0075] In some embodiments, the plurality of reply messages include at least a first reply message, and the apparatus 400 further includes: in response to receiving additional input from the user before providing the first reply message, presenting reference content associated with the first reply message to indicate that the first reply message is a reply to the first input message.

[0076] In some embodiments, the multiple reply messages include a second reply message and a third reply message, and the providing module 440 is further configured to: provide the second reply message in the interactive interface; and provide the third reply message in the interactive interface after a target duration of providing the second reply message.

[0077] In some embodiments, the target duration is a preset duration or a random duration.

[0078] In some embodiments, the streaming text is generated by the model based on system prompts and a first input message, the system prompts indicating constraint information for the model's output content.

[0079] In some embodiments, the constraint information instructs the model to use a preset delimiter to separate different text parts, and the parsing module 430 is further configured to split the streaming text into multiple text parts by detecting the preset delimiter in the streaming text.

[0080] In some embodiments, the constraint information also indicates the length and / or number of text portions output by the model.

[0081] In some embodiments, the streaming text is a first streaming text, and the apparatus 400 further includes a first processing module configured to: receive a second input message from a user; obtain second streaming text generated by a model based on the second input message; detect whether a third input message from a user has been received before providing the second streaming text to the user; and, in response to receiving the third input message, trigger the model to generate third streaming text based on the second input message and the third input message.

[0082] In some embodiments, the device 400 further includes a stop module configured to stop providing the user with a reply message corresponding to the second streaming text.

[0083] In some embodiments, the first processing module is further configured to: determine the correlation between the second input message and the third input message; and, in response to the correlation being greater than a threshold, trigger the model to generate a third streaming text based on the second input message and the third input message.

[0084] In some embodiments, the apparatus 400 further includes a second processing module, which is configured to: in response to a correlation not exceeding a threshold, trigger a model to generate a first text corresponding to a second input message and a second text corresponding to a third input message, respectively, so as to determine a third streaming text based on the first text and the second text.

[0085] In some embodiments, the device 400 further includes a third processing module configured to: provide a user with a fourth reply message corresponding to the first text and a fifth reply message corresponding to the second text; and at least present reference content associated with the fifth reply message to indicate that the fifth reply message is a reply to the third input message.

[0086] The units included in device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units may be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 400 may be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.

[0087] Figure 5 illustrates a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 500 shown in Figure 5 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 500 shown in Figure 5 can be used in electronic devices.

[0088] As shown in Figure 5, the electronic device 500 is in the form of a general-purpose electronic device. Components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 500.

[0089] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.

[0090] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0091] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0092] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0093] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0094] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0095] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0096] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0098] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A message processing method, comprising: Receive the user's first input message; Obtain the streaming text generated by the model based on the first input message; By parsing the streaming text, the streaming text is split into multiple text parts; and Provide the user with multiple reply messages corresponding to the multiple text sections.

2. The method according to claim 1, wherein providing the user with multiple reply messages corresponding to the plurality of text portions includes: The multiple reply messages are presented in the interactive interface in a target order, which corresponds to the order of the multiple text portions in the streaming text.

3. The method according to claim 2, wherein the plurality of reply messages at least includes a first reply message, and the method further includes: In response to receiving additional input from the user before providing the first reply message, referenced content associated with the first reply message is presented to indicate that the first reply message is a response to the first input message.

4. The method according to claim 2, wherein the plurality of reply messages includes a second reply message and a third reply message, and presenting the plurality of reply messages in the interactive interface in a target order includes: The second response message is provided in the interactive interface; as well as After providing the target duration of the second response message, the third response message is provided in the interactive interface.

5. The method according to claim 4, wherein the target duration is a preset duration or a random duration.

6. The method of claim 1, wherein the streaming text is generated by the model based on a system prompt and the first input message, the system prompt indicating constraint information for the output content of the model.

7. The method of claim 6, wherein the constraint information instructs the model to use a preset delimiter to separate different text parts, and splitting the streaming text into multiple text parts by parsing the streaming text includes: The streaming text is split into multiple text parts by detecting the preset delimiter in the streaming text.

8. The method of claim 7, wherein the constraint information further indicates the length and / or number of text portions output by the model.

9. The method according to claim 1, wherein the streaming text is a first streaming text, and the method further comprises: Receive the user's second input message; Obtain the second streaming text generated by the model based on the second input message; Before providing the second streaming text to the user, detect whether the user's third input message has been received; as well as In response to receiving the third input message, the model is triggered to generate a third streaming text based on the second input message and the third input message.

10. The method of claim 9, further comprising: Stop providing the user with the reply message corresponding to the second streaming text.

11. The method of claim 9, wherein triggering the model to generate third streaming text based on the second input message and the third input message in response to receiving the third input message comprises: Determine the correlation between the second input message and the third input message; as well as In response to the correlation being greater than a threshold, the model is triggered to generate the third streaming text based on the second input message and the third input message.

12. The method according to claim 9, further comprising: In response to the correlation between the second input message and the third input message being less than or equal to a threshold, the model is triggered to generate a fourth streaming text based on the third input message; as well as Provide a first set of reply messages corresponding to the second streaming text and a second set of reply messages corresponding to the fourth streaming text.

13. The method of claim 12, further comprising: Present a first reference associated with the first set of response messages to indicate that the first set of response messages is a response to the second input message; and / or A second reference is presented in connection with the second set of response messages to indicate that the second set of response messages is a response to the third input message.

14. An apparatus for message processing, comprising: The receiving module is configured to receive the user's first input message; The acquisition module is configured to acquire the streaming text generated by the model based on the first input message; The parsing module is configured to split the streaming text into multiple text parts by parsing the streaming text; as well as The module is configured to provide the user with multiple reply messages corresponding to the multiple text portions.

15. An electronic device comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 13.

16. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 13.