Method for providing auxiliary data in man-machine conversation, electronic equipment and storage medium

By displaying the semantic content of dialogue flow data in human-computer dialogue, focusing on the key object and its auxiliary data, the problem of limited information transmission is solved, and more comprehensive information presentation and convenient data acquisition are achieved.

CN121809681APending Publication Date: 2026-04-07ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing human-computer dialogue systems have high limitations in information transmission, requiring users to perform additional operations to obtain external data in order to understand more about the focus of their attention.

Method used

The interactive page displays the semantic content of the dialogue flow data, focusing on the target object and its auxiliary data. The auxiliary data is provided by an external data source and is processed collaboratively by the terminal device and the service device.

Benefits of technology

It enhances the completeness and sufficiency of information during human-computer dialogue, reduces additional user operations, and improves the convenience of information access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121809681A_ABST
    Figure CN121809681A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method for providing auxiliary data in man-machine conversation, electronic equipment and a storage medium, the method is applied to terminal equipment, the method comprises the steps that an interaction page used for carrying out conversation with an intelligent agent is displayed, and the interaction page comprises a conversation area and an auxiliary area which are independent of each other and are displayed on the same screen; in the dialogue process of a user and the intelligent agent, dialogue flow data between the user and the intelligent agent are presented in the dialogue area, at least one focus object and corresponding auxiliary data are presented in the auxiliary area, the focus object is an object concerned by semantic content of the dialogue flow data, and the focus object is an object concerned by the semantic content of the dialogue flow data. The auxiliary data is provided by an external data source which is not controlled by the intelligent agent and is used for assisting the user in knowing external data associated with dialogue content in the dialogue process.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present specification relates to the technical field of artificial intelligence, and in particular, to a method for providing auxiliary data in human-computer conversation, an electronic device, and a storage medium. BACKGROUND

[0002] With the development of artificial intelligence technology, large language models have been applied to many scenarios to assist users in conveniently obtaining required information or completing to-do tasks.

[0003] A user can interact with an intelligent agent based on a large language model in a conversational manner. The user inputs a question in a conversation page, and the question is provided to the intelligent agent to obtain a corresponding answer. In this manner, the answer only contains core data directly corresponding to the user's question. The information transmission to the user in human-computer conversation has certain limitations, and how to reduce the limitations is a technical problem to be considered.

[0004] The content of the background section only represents the inventor's own knowledge and does not mean that the above information has entered the public domain before the filing date of the present disclosure, nor does it mean that it can be prior art of the present disclosure. SUMMARY

[0005] The present specification provides a method for providing auxiliary data in human-computer conversation, an electronic device, and a storage medium, which can reduce the limitations of information transmission to the user in human-computer conversation.

[0006] In a first aspect, the present specification provides a method for providing auxiliary data in human-computer conversation, applied to a terminal device, the method comprising: displaying an interaction page for conversing with an intelligent agent, the interaction page comprising a conversation area and an auxiliary area which are independently displayed on the same screen; and presenting conversation flow data between the user and the intelligent agent in the conversation area and at least one focus object and its corresponding auxiliary data in the auxiliary area during the conversation between the user and the intelligent agent, wherein the focus object is an object of interest of semantic content of the conversation flow data, and the auxiliary data is provided by an external data source not controlled by the intelligent agent and is used to assist the user in understanding external data associated with the conversation content during the conversation.

[0007] In a second aspect, the present specification also provides a method for providing auxiliary data in a human-computer conversation, applied to a service device, the method comprising: receiving conversation flow data from a terminal device, the conversation flow data being at least part of conversation content presented in a conversation area in an interactive page, the interactive page being a page in which a user has a conversation with an agent; performing semantic analysis on the conversation flow data to identify at least one focus object to which semantic content of the conversation flow data pays attention; obtaining auxiliary data corresponding to each of the at least one focus object from an external data source not controlled by the agent; and sending the at least one focus object and the auxiliary data corresponding thereto to the terminal device for presentation in an auxiliary area of the interactive page, wherein the auxiliary data is used to assist the user in understanding external data associated with the conversation content during the conversation.

[0008] In a third aspect, the present specification also provides an electronic device, comprising: at least one storage medium storing at least one instruction set; and at least one processor in communication connection with the at least one storage medium, wherein the at least one processor reads the at least one instruction set and performs the method of the first aspect or the second aspect according to the indication of the at least one instruction set.

[0009] In a fourth aspect, the present specification also provides a computer-readable nonvolatile storage medium, wherein the computer-readable nonvolatile storage medium stores at least one instruction set, and the at least one instruction set is executed by at least one processor to implement the method of the first aspect or the second aspect.

[0010] The method, electronic device and storage medium for providing auxiliary data in a human-computer conversation provided by the present specification, in the case of presenting conversation flow data of a user and an agent in an interactive page during a conversation between the user and the agent, also present a focus object to which semantic content of the conversation flow data pays attention and auxiliary data thereof in an auxiliary area in the interactive page. The auxiliary data is provided by an external data source not controlled by the agent, and can assist the user in understanding external data associated with the conversation content during the conversation. In this way, the user can be made to understand more data of the focus object of interest during the human-computer conversation, the information transmitted to the user in the human-computer conversation is enriched, the completeness and sufficiency of the information are improved, and the limitation of information transmission to the user in the human-computer conversation is reduced.

[0011] Other functions of the method, electronic device and storage medium for providing auxiliary data in a human-computer conversation provided by the present specification will be partially listed in the following description. The creative aspects of the method, electronic device and storage medium for providing auxiliary data in a human-computer conversation provided by the present specification can be fully explained by practicing or using the methods, devices and combinations described in the following detailed examples. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present specification, the following will briefly introduce the drawings needed to be used in the embodiments description. Obviously, the drawings in the following description only some embodiments of the present specification, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0013] Figure 1 A schematic diagram of a human-computer conversation scene provided by an embodiment of the present specification is shown; Figure 2 A hardware structure diagram of an electronic device provided by an embodiment of the present specification is shown; Figure 3 A flowchart of a method for providing auxiliary data in a human-computer conversation provided by an embodiment of the present specification is shown; Figure 4 A schematic diagram of an interactive page provided by an embodiment of the present specification is shown Figure 4 A simple block diagram of a strategy generation method provided by an embodiment of the present specification is shown; and Figure 5 An implementation block diagram of a method for providing auxiliary data in a human-computer conversation provided by an embodiment of the present specification is shown. DETAILED DESCRIPTION

[0014] The following description provides specific application scenarios and requirements of the present specification, which is to enable those skilled in the art to manufacture and use the contents of the present specification. Various local modifications of the disclosed embodiments are obvious to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of the present specification. Therefore, the present specification is not limited to the shown embodiments, but is consistent with the widest range of claims.

[0015] The terms used herein are only for the purpose of describing specific example embodiments, and are not limiting. For example, unless the context clearly indicates otherwise, as used herein, the singular forms "a", "an", and "the" can also include the plural forms. The term "a plurality of" refers to two or more, and "at least one" refers to one or more. The present specification can use the terms "first", "second", and the like to describe various information, but these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, "first" can also be referred to as "second" without departing from the scope of the embodiments of the present specification, and similarly, "second" can also be referred to as "first".

[0016] The term "at least one of A, B, or C" includes seven options: A only, B only, C only, both A and B, both A and C, both B and C, and both A, B, and C. Similarly, the statement "at least one of multiple items" refers to all possible combinations based on these items. The term "and / or" refers to any or all possible combinations of one or more associated listed items. For example, "A and / or B" includes three options: A only, B only, and both A and B. "A, B, and / or C" is equivalent to "at least one of A, B, or C".

[0017] The term "comprising" is an open-ended description and should be understood as "including but not limited to," potentially including other content beyond what has been described. When used in this specification, the terms "comprising," "including," and / or "containing" mean that the associated integers, steps, operations, elements, and / or components are present, but do not exclude the presence of one or more other features, integers, steps, operations, elements, components, and / or groups, or that other features, integers, steps, operations, elements, components, and / or groups may be added to the system / method.

[0018] Considering the following description, these and other features of this specification, as well as the operation and function of the related components of the structure, and the economy of assembly and manufacture of the parts, can be significantly improved. All of these form part of this specification with reference to the accompanying drawings. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.

[0019] The flowcharts used in this specification illustrate operations implemented according to some embodiments of this specification. It should be clearly understood that the operations in the flowcharts may not be implemented in a sequential order. Instead, the operations may be implemented in reverse order or simultaneously. Furthermore, one or more additional operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.

[0020] In this specification, the Large Language Model (LLM) may also be referred to simply as the Large Model. A Large Language Model is a natural language processing model based on deep learning techniques, typically with billions to hundreds of billions or even more parameters, possessing powerful language understanding and generation capabilities. Large Language Models can employ the Transformer architecture or its variants (such as GPT, BERT, etc.), which utilizes an attention mechanism to globally model sequential data, efficiently handling long-distance dependencies and thus performing exceptionally well in natural language tasks. Large Language Models learn the statistical features and semantic relationships of language through pre-training on large-scale corpora, giving them excellent generalization capabilities. The core capabilities of Large Language Models include, but are not limited to: understanding contextual semantics, generating coherent and grammatically correct text, performing logical reasoning, and handling multi-task scenarios. Its usage typically includes two modes: direct inference and fine-tuning. In direct inference mode, the user guides the Large Language Model to generate specific outputs by designing prompts. Cue words can be task descriptions or instructions in text form, used to stimulate the semantic understanding and generation capabilities of large language models. In fine-tuning mode, large language models are further trained on small-scale datasets in specific domains to optimize their performance on specific tasks. The powerful generalization ability and flexibility of large language models make them an important tool in the field of artificial intelligence, providing efficient and accurate solutions for automated text generation and understanding.

[0021] In some embodiments, large language models can also understand and generate data from other modalities (such as visual and audio data). In this case, large language models can also be called multimodal large language models (MLLMs). MLLMs provide a richer and more natural interactive experience by integrating multiple types of input and output, such as text, images, and sound. The core advantage of MLLMs lies in their ability to process and understand information from different modalities and fuse this information to complete complex tasks. For example, MLLMs can analyze an image and generate descriptive text, or generate a corresponding image based on a text description. This cross-modal understanding and generation capability makes MLLMs widely applicable across multiple fields.

[0022] It should be noted that the key technologies of large language models can be found in the detailed description in the paper "A Survey of Large Language Models" (paper number: arXiv:2303.18223v16, published on March 11, 2025, public link: https: / / doi.org / 10.48550 / arXiv.2303.18223), and will not be repeated here.

[0023] With the development of artificial intelligence technology, various AI models (such as large language models) are widely used. In many scenarios, such as financial analysis, e-commerce guidance, itinerary planning, or general dialogue, users can engage in conversational interaction with intelligent agents based on large language models (such as chatbots), i.e., human-computer dialogue. Some applications or web pages have human-computer dialogue functions. Users can launch the application or visit the web page using their terminal device and trigger the human-computer dialogue function to display an interactive page on the terminal device. Users and intelligent agents can converse on this interactive page; for example, the dialogue information between users and intelligent agents can be displayed on the interactive page as dialogue bubbles. Users can input dialogue information (such as questions) into the interactive page, and the intelligent agent generates corresponding response information based on the question information and displays it on the interactive screen, thus realizing human-computer dialogue.

[0024] Typically, intelligent agents only reason based on the user's input to obtain a response, a response model with significant limitations. In some scenarios, when performing tasks or making decisions, users need more than just dialogue with an intelligent agent to obtain relevant information or seek advice; they also require more external data. For example, in a financial analysis scenario, a user asks about the current day's stock price movement, and the intelligent agent outputs the percentage increase. After receiving this answer, the user still needs to manually switch to the stock chart page to view the candlestick chart or intraday chart for a more detailed understanding of the stock's situation. In this approach, the information delivery to the user during the human-computer dialogue is limited, and the user needs to perform cumbersome operations to obtain the data they are interested in.

[0025] This specification provides a method for providing auxiliary data in human-computer dialogue. During a dialogue between a user and an intelligent agent, in addition to displaying the dialogue flow data on the interactive page, an auxiliary area on the interactive page also displays the semantic focus objects of the dialogue flow data and their auxiliary data. This auxiliary data is provided by an external data source not controlled by the intelligent agent, and can help the user understand external data related to the dialogue content during the dialogue. This allows users to understand more data about their focus objects during human-computer dialogue, enriching the information conveyed to the user, improving the completeness and sufficiency of the information, and reducing the limitations of information delivery to users in human-computer dialogue. Correspondingly, users can more conveniently obtain the information they are interested in, reducing the need for additional operations and improving the ease of information access for users.

[0026] The method for providing auxiliary data in human-computer dialogue provided in this specification can be applied to various scenarios. The dialogue flow data between the user and the intelligent agent differs in different application scenarios, as do the relevant focus objects and external data sources, but the execution process of the method remains the same. For example, this method can be applied to financial analysis scenarios, where the dialogue flow data may involve financial industry terminology and expression habits. Relevant focus objects may include financial institutions, objects, and indicators, and external data sources may include data sources from financial institutions and official statistical data. As another example, this method can be applied to e-commerce shopping guide scenarios, where the dialogue flow data may involve products, purchase rules, or prices. Focus objects may include products, merchants, and channels, and external data sources may include data provided by e-commerce platforms. As yet another example, this method can be applied to traffic planning scenarios, where the dialogue flow data may involve geographical location, travel time, and traffic duration. Focus objects may include routes, waypoints, and congestion conditions, and external data sources may include satellite positioning data, real-time traffic data, and traffic event data.

[0027] The following example illustrates how to apply the method for providing auxiliary data in human-computer dialogue provided in this manual in a financial analysis scenario.

[0028] Figure 1 This diagram illustrates a human-computer dialogue scenario according to an embodiment of this specification, specifically an application scenario 100 of a method for providing auxiliary data in human-computer dialogue. For ease of description, the method for providing auxiliary data in human-computer dialogue will hereinafter be referred to as a dialogue enhancement method. Figure 1 As shown, the application scenario 100 may include a target user 110, a terminal device 120, a service device 130, a network 140, and an external data source 150.

[0029] The target user 110 can be a user who engages in dialogue with the intelligent agent. The target user 110 and the intelligent agent can converse using the human-computer dialogue function provided by the terminal device 120, such as through the interactive page 1201 displayed on the terminal device 120. The target user 110 can operate the terminal device 120 to display the interactive page 1201 and input questions on the interactive page 1201. The intelligent agent can generate corresponding response information based on the question information and display it on the interactive page 1201.

[0030] In some embodiments, at least some steps of the dialogue enhancement method provided in this specification can be executed on the terminal device 120. In this case, the terminal device 120 may store data or instructions for executing the dialogue enhancement method described in this specification, and may execute or be used to execute said data or instructions. In some embodiments, the terminal device 120 may include a hardware device with data processing capabilities and the necessary programs required to drive the hardware device to operate.

[0031] In some embodiments, the terminal device 120 may have one or more applications (APPs) installed. The interaction between the target user 110 and the intelligent agent may rely on the human-computer dialogue function provided by the application. For example, if the target user 110 needs to trigger the terminal device 120 to run the application and trigger the human-computer dialogue function, the terminal device 120 may display the interactive page 1201. In some embodiments, the terminal device 120 may implement the dialogue enhancement method described in this specification based on the application.

[0032] In some embodiments, terminal device 120 may include mobile devices, tablets, laptops, desktop computers, smart home devices, built-in devices in motor vehicles, or similar content, or any combination thereof. In some embodiments, the mobile device may include smart mobile devices, virtual reality devices, augmented reality devices, or similar devices, or any combination thereof. In some embodiments, the smart home device may include smart TVs, smart refrigerators, etc., or any combination thereof. In some embodiments, the smart mobile device may include smartphones, personal digital assistants, gaming devices, navigation devices, etc., or any combination thereof. In some embodiments, the virtual reality device or augmented reality device may include virtual reality headsets, virtual reality glasses, virtual reality patches, augmented reality headsets, augmented reality glasses, augmented reality patches, or similar content, or any combination thereof. For example, the virtual reality device or the augmented reality device may include Google Glass, head-mounted displays, VR, etc. In some embodiments, the built-in device in the motor vehicle may include an in-vehicle computer, an in-vehicle television, etc.

[0033] Service device 130 may be a backend server providing support for the human-computer dialogue function of terminal device 120. In some embodiments, at least some steps of the dialogue enhancement method may be executed on service device 130. In this case, service device 130 may store data or instructions for executing the dialogue enhancement method described herein, and may execute or be used to execute said data or instructions. In some embodiments, service device 130 may include hardware devices with data processing capabilities and necessary programs for driving the hardware devices. Service device 130 may be communicatively connected to multiple terminal devices 120, receiving data sent by terminal devices 120, and sending data to terminal devices 120.

[0034] In some embodiments, the intelligent agent that interacts with the user is deployed in the service device 130. The terminal device 120 can send the question information entered by the target user 110 into the interaction page 1201 to the service device 130, causing the intelligent agent therein to generate corresponding response information. The service device 130 then sends the response information to the terminal device 120 for display.

[0035] In some embodiments, terminal device 120 also sends dialogue flow data from the interactive page to service device 130. Service device 130 executes the dialogue enhancement method described herein to determine the focus object of the dialogue flow data. In some embodiments, service device 130 connects to external data source 150 to obtain auxiliary data corresponding to the focus object from external data source 150. For example, service device 130 sends a data acquisition request for auxiliary data corresponding to the focus object to external data source 150 and obtains the auxiliary data returned by external data source 150. Service device 130 sends the focus object and its auxiliary data to terminal device 120 for display to assist target user 110 in understanding external data associated with the dialogue content during the dialogue process.

[0036] Network 140 is a medium used to provide a communication connection between terminal device 120 and service device 130. Service device 130 and external data source 150 can also communicate via the network. Figure 1 Not illustrated in the diagram. Network 140 can facilitate the exchange of information or data. For example... Figure 1As shown, terminal device 120 and service device 130 can connect to network 140 and transmit information or data to each other through network 140. In some embodiments, network 140 can be any type of wired or wireless network, or a combination thereof. For example, network 140 may include cable networks, wired networks, fiber optic networks, telecommunications networks, intranets, the Internet, local area networks (LANs), wide area networks (WANs), wireless local area networks (WLANs), metropolitan area networks (MANs), public switched telephone networks (PSTNs), Bluetooth networks, ZigBee networks, near field communication (NFC) networks, or similar networks. In some embodiments, network 140 may include one or more network access points. For example, network 140 may include wired or wireless network access points, such as base stations or Internet switching points, through which one or more components of terminal device 120 and service device 130 can connect to network 140 to exchange data or information.

[0037] It should be understood that Figure 1 The number of terminal devices 120, service devices 130, networks 140, and external data sources 150 shown is merely illustrative. Depending on implementation needs, any number of terminal devices 120, service devices 130, networks 140, and external data sources 150 can be included.

[0038] It should be noted that the dialogue enhancement method can be executed entirely on the terminal device 120; or it can be executed partially on the terminal device 120 and partially on the service device 130. In this specification, the data processing and acquisition steps of the dialogue enhancement method are executed on the service device 130, while other steps are executed on the terminal device 120, as an example for illustration.

[0039] It should be noted that all user data obtained in this manual has been authorized by the user and does not involve user privacy.

[0040] Figure 2 A hardware structure diagram of an electronic device provided according to an embodiment of this specification is shown. Figure 2 Electronic device 200 in the middle can be used as Figure 1 The terminal device 120 executes the steps performed by the terminal device in the dialogue enhancement method described in this specification. The electronic device 200 can also be used as... Figure 1 The service device 130 in the specification executes the steps performed by the service device in the dialogue enhancement method described herein.

[0041] like Figure 2As shown, the electronic device 200 may include at least one storage medium 230 and at least one processor 220. In some embodiments, the electronic device 200 may also include a communication port 250 and an internal communication bus 210. The electronic device 200 may also include an input / output component (I / O component) 260.

[0042] The internal communication bus 210 can connect to different system components. For example, the internal communication bus 210 can connect to storage medium 230, processor 220, communication port 250, and I / O component 260, etc.

[0043] I / O component 260 supports input / output between electronic device 200 and other components.

[0044] Communication port 250 is used for data communication between electronic device 200 and the outside world. For example, communication port 250 can be used for data communication between electronic device 200 and a network. Communication port 250 can be a wired communication port or a wireless communication port.

[0045] Storage medium 230 may include a data storage device. The data storage device may be a non-transitory storage medium or a temporary storage medium. For example, the data storage device may include one or more of a disk 232, a read-only storage medium (ROM) 234, or a random access storage medium (RAM) 236. Storage medium 230 also includes at least one instruction set stored in the data storage device. The instruction set may include computer program code, which may include programs, routines, objects, components, data structures, procedures, modules, etc.

[0046] At least one processor 220 may be communicatively connected to at least one storage medium 230. When the electronic device 200 is running, at least one processor 220 reads the at least one instruction set and, according to the instructions of the at least one instruction set, executes the steps included in the dialogue enhancement method provided in this specification. Processor 220 may be in the form of one or more processors. In some embodiments, processor 220 may include one or more hardware processors, such as microcontrollers, microprocessors, reduced instruction set computers (RISC), application-specific integrated circuits (ASICs), application-specific instruction set processors (ASIPs), central processing units (CPUs), graphics processing units (GPUs), physical processing units (PPUs), microcontroller units, digital signal processors (DSPs), field-programmable gate arrays (FPGAs), advanced RISC machines (ARMs), programmable logic devices (PLDs), any circuit or processor capable of performing one or more functions, or any combination thereof.

[0047] For illustrative purposes only, the electronic device 200 in the accompanying drawings shows only one processor 220. However, it should be noted that the electronic device 200 in this specification may also include multiple processors. Therefore, the operation and / or method steps disclosed in this specification may be executed by one processor or by multiple processors in combination. For example, if the processor 220 of the electronic device 200 described in this specification executes steps A and B, it should be understood that steps A and B may also be executed jointly or separately by two different processors 220 (e.g., the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).

[0048] Figure 3 A flowchart is shown illustrating a method for providing auxiliary data in human-computer dialogue according to an embodiment of this specification. This method P300 for providing auxiliary data in human-computer dialogue is also the aforementioned dialogue enhancement method, and can be applied to… Figure 1 In the human-computer dialogue scenario shown, the method is executed by the cooperation of the terminal device and the service device in the scenario. This method is also an interaction method between the terminal device and the service device. This specification will not use separate flowcharts to illustrate the steps executed by the terminal device and the service device respectively. As mentioned above, the electronic device 200 can be used as… Figure 1 The terminal device 100 in the electronic device 200 executes the steps performed by the terminal device in method P300 of this specification. Specifically, the processor 220 in the electronic device 200 can read the instruction set stored in its local storage medium, and then execute the steps performed by the terminal device in method P300 of this specification according to the provisions of the instruction set. The electronic device 200 can also be used as... Figure 1 The service device 130 in the electronic device 200 executes the steps performed by the service device in method P300 of this specification. Specifically, the processor 220 in the electronic device 200 can read the instruction set stored in its local storage medium and then execute the steps performed by the service device in method P300 of this specification according to the instructions in the instruction set. Figure 3 As shown, method P300 may include steps 310 to 370 as described below.

[0049] Step 310: The terminal device displays an interactive page for dialogue with the intelligent agent. The interactive page includes a dialogue area and an auxiliary area that are independent of each other and displayed on the same screen.

[0050] The intelligent agent can be a conversational AI system based on a large language model. When a user needs to converse with the intelligent agent, they can operate the terminal device to trigger an interactive page for dialogue. For example, in the financial field, when making investments or adjusting financial products, users can converse with a financial analysis intelligent agent to seek relevant investment or financial advice. In some embodiments, the terminal device may have relevant applications installed, such as investment applications, financial management applications, or banking applications. This application may have human-computer dialogue capabilities. The user can trigger the terminal device to launch the application and activate the human-computer dialogue entry point within the application, causing the terminal device to display an interactive page for conversing with the intelligent agent.

[0051] Figure 4 A schematic diagram of an interactive page provided according to an embodiment of this specification is shown. Figure 4 As shown, the interactive page 400 may include a dialogue area 401 and an auxiliary area 402 that are independent of each other and displayed on the same screen. The dialogue area 401 is used to display the dialogue content between the user and the agent, which includes the question information 01 input by the user and the response information 02 output by the agent. In some embodiments, based on the question information 01 and the response information 02, the dialogue content may also include the thought process information 03 on which the agent generates the response information 02.

[0052] In some embodiments, such as Figure 4 As shown, the dialogue between the user and the agent can be displayed in the form of speech bubbles. The dialogue input by the user and the dialogue output by the agent are displayed in two different speech bubbles, which are located near the left and right sides of the interactive page 400, respectively. The dialogue input by the user is located near the right side of the interactive page 400, and the dialogue output by the agent is located near the left side of the interactive page 400. In some embodiments, only the dialogue input by the user may be displayed in the speech bubble.

[0053] In some embodiments, the auxiliary area is located above, to the left, or to the right of the dialog area in the interactive page. The spatial positions of the auxiliary area and the dialog area can be flexibly set according to user needs, and the position of the auxiliary area in the interactive page is not limited here. Figure 4Taking a vertical screen display of a terminal device as an example, where the vertical display space is ample, the auxiliary area 402 is positioned above the dialogue area 401. For a horizontal screen display terminal device, where the horizontal distance is ample, the auxiliary area 402 can be positioned to the left or right of the dialogue area 401 to facilitate user operation and viewing. In some embodiments, the display sizes of the auxiliary area and the dialogue area can also be adjusted to match user needs or data display requirements.

[0054] In some embodiments, the dialog area in the interactive page may belong to a conversational user interface (CUI). The auxiliary area may belong to a graphical user interface (GUI).

[0055] Step 320: During the dialogue between the user and the intelligent agent, the terminal device displays the dialogue content between the user and the intelligent agent in the dialogue area.

[0056] In this specification, the dialogue process between the user and the intelligent agent refers to a continuous interaction cycle between the user and the intelligent agent. This cycle begins when the user initiates or enters the interaction page and ends when the user actively closes the page or the session times out. This process includes, but is not limited to, the system initialization ready state, the user input state, the intelligent agent processing state, and the intelligent agent output state. At one stage of the dialogue process, when the user has opened the interaction page and the intelligent agent has completed initialization, but the user has not yet entered new dialogue content, and the intelligent agent has not output any dialogue content, the intelligent agent is in a dialogue standby state. In this state, the current dialogue context is maintained in the dialogue area of ​​the interaction page, and the terminal device can perform dialogue-related background tasks, such as preloading resources, updating real-time data to be displayed, and displaying dialogue history or suggested questions.

[0057] During a dialogue between the user and the intelligent agent, the user can input dialogue content into a dialogue area on the interactive page. Once the user has finished inputting, the dialogue content will be displayed in that area. For an example, please refer to [link / reference]. Figure 4 The dialogue area 401 includes an input control 04. Users can operate the input control 04 to input dialogue content into the dialogue area via voice input or by inputting characters via a virtual keyboard.

[0058] After a user inputs any dialogue content into the dialogue area, the terminal device can transmit the dialogue content to the service device. The intelligent agent deployed on the service device then reasons about the dialogue content, obtains the corresponding thought process information, and generates a corresponding response based on this thought process. The service device sequentially transmits the thought process information and the response back to the terminal device, which then presents them sequentially in the dialogue area. This thought process information and response information can be streamed back to the terminal device, allowing the terminal device to present the received information in order.

[0059] Step 330: The terminal device sends dialogue stream data to the service device. The dialogue stream data includes at least a portion of the dialogue content presented in the dialogue area of ​​the interactive page.

[0060] A terminal device can send the dialogue content between a user and an agent to a service device, enabling the service device to determine, based on the dialogue content, the data the user might want to know and present it to the user accordingly. In some cases, the dialogue content between the user and the agent (especially the dialogue content output by the agent) is long, or the terminal device streams the dialogue content generated by the agent; in such cases, the terminal device can stream this dialogue content to the service device. Accordingly, the service device can process the received dialogue content in a streaming manner. This reduces the transmission load on the terminal device and the instantaneous processing load on the service device, reduces the processing latency of the service device, and improves response speed. For example, the terminal device can send dialogue stream data to the service device based on the dialogue content in a dialogue area. This dialogue stream data includes at least a portion of the dialogue content presented in the dialogue area.

[0061] In some embodiments, a data volume threshold can be preset for the dialogue stream data, such as a data volume threshold that can be represented by text length. When the newly added dialogue content in the dialogue area of ​​the interactive page reaches the data volume threshold, the terminal device will send the newly added dialogue content as new dialogue stream data to the service device.

[0062] In some embodiments, the user's dialogue content and the agent's dialogue content can be sent to the service device as separate dialogue streams. Typically, user-input dialogue content is short; when new user dialogue content is added to the dialogue area, the terminal device can directly treat this user's dialogue content as a new dialogue stream. For the agent's dialogue content, the terminal device can determine whether to segment the dialogue content based on its length, sending it to the service device in separate dialogue streams. This data volume threshold can also be greater than a certain amount to reduce the negative impact on the display effect of the focus object caused by frequent changes in the focus object identified in a single dialogue round due to excessively short dialogue streams, and to reduce the waste of communication resources. In some embodiments, for the agent's dialogue content, the thought process information and response information can be sent to the service device as two separate dialogue streams, sequentially.

[0063] In some embodiments, a preset sending period may be defined for the dialogue stream data. After each sending period, the terminal device may send the newly added dialogue content in the dialogue area of ​​the interactive page during that sending period as new dialogue stream data to the service device.

[0064] Step 340: The service device performs semantic parsing on the dialogue stream data to identify at least one focus object of interest in the semantic content of the dialogue stream data.

[0065] Upon receiving any dialogue stream data, the service device can perform semantic parsing on that data. During the semantic parsing process, the semantic content of the dialogue stream data can be obtained first, and then reasoning can be performed based on that semantic content (such as combining it with other relevant industry knowledge) to determine at least one focus object of interest for that semantic content.

[0066] Whether the dialogue stream data includes user conversations or agent conversations, it is all related to the user's needs or dialogue intentions. The user's conversation content represents the user's direct need to engage in dialogue with the agent, while the agent's conversation content provides information to satisfy that direct need. Correspondingly, the focus of the semantic content of the dialogue stream data can represent the objects of interest in the dialogue between the user and the agent. These focus objects can include those mentioned in the user's direct needs, or objects that are not mentioned in the user's direct needs but are determined through inference that the user might be interested in.

[0067] In some embodiments, the service device utilizes a large language model to perform semantic parsing on the received dialogue stream data. Accordingly, the service device can input the dialogue stream data into the large language model and guide it to perform semantic understanding of the data, identifying the focus object of the dialogue based on the understood semantic content. This large language model can be deployed within the service device, or it can be deployed on other devices besides the service device, and it supports invocation by the service device.

[0068] Large language models are cue-based generative models whose output is highly dependent on the guidance and control of cue text. For each large language agent, specific cue text needs to be constructed to instruct the corresponding large language model to perform a specific task. Cue text is essentially a structured instruction or query information that can translate task intent into contextual information that the model can understand and execute, used to guide, constrain, and customize the output behavior of the large language model. In some embodiments, the sample generation device can generate the cue text required by each large language agent based on a static template filling method, such as filling a preset cue template with information related to the sample generation task.

[0069] For example, the prompt text may include one or more components such as role definition, task instructions, contextual information, input / output format constraints, and examples. The role definition section is used to assign a specific role to the model (e.g., "You are a dialogue analytics expert with basic knowledge of the financial industry"), enabling it to generate information from the perspective and expertise of that role. The task instructions section describes the specific task the model needs to perform (e.g., "Based on your knowledge, analyze the objects the user might be interested in based on the xx dialogue content"). The contextual information section is used to provide background knowledge, reference information, or specific data to be processed related to the task. The input / output format constraints section is used to explicitly specify the format requirements of the model's output (e.g., "Output in JSON format"), ensuring the structured and parsable nature of the output. The examples section is used to provide one or more example input-output pairs to demonstrate the desired task execution, enabling the model to output the required results more quickly and in a more standardized manner.

[0070] In some embodiments, the service device can generate corresponding prompt text based on the dialogue stream data, and input the prompt text into a large language model to guide the large language model in determining the focus object of the semantic content of the dialogue stream data. For example, the dialogue stream data can be input into a preset template to obtain the corresponding prompt text. The template may at least include task instructions, which instruct the input dialogue stream data to first undergo natural language understanding (NLU), and then identify the focus object of the dialogue based on the obtained semantic content. This natural language understanding is the aforementioned semantic understanding. In some embodiments, the semantic content can be structured semantic content conforming to a preset format.

[0071] In some embodiments, the service device may also utilize a dedicated NLU mini-model to perform semantic understanding on the received dialogue stream data to obtain structured semantic content. This semantic content is then input into a large language model to guide the model in identifying the focus object of the semantic content.

[0072] In some embodiments, during the process of identifying focus objects based on the semantic content of dialogue stream data, a first focus object contained in the semantic content can be identified, and the user intent corresponding to the semantic content can be inferred to determine a second focus object associated with the user intent. The first focus object and the second focus object can jointly serve as the focus objects of interest in the semantic content of the dialogue stream data.

[0073] In some embodiments, the user and the intelligent agent engage in dialogue regarding the financial sector. The aforementioned focus object can be a financial entity. This financial entity can include tradable assets within the financial sector. In some embodiments, the financial entity includes at least one of the following: stocks, funds, bonds, physical assets, monetary assets, or financial derivatives. Bonds include government bonds, local government bonds, corporate bonds, etc. Physical assets can include precious metals, crude oil, copper, soybeans, or wheat, etc. Monetary assets include treasury bills, cash, etc. Financial derivatives can include futures or options, etc.

[0074] Step 350: The service device obtains the auxiliary data corresponding to each of the at least one focus object from an external data source that is not controlled by the intelligent agent.

[0075] After identifying at least one focus object of interest in the semantic content of the dialogue stream data, the service device can acquire auxiliary data corresponding to each focus object. This auxiliary data is provided by an external data source not controlled by the agent and is used to help the user understand external data related to the dialogue content during the dialogue between the user and the agent.

[0076] In some embodiments, a focus object may correspond to multiple types of data, and the auxiliary data may include at least one of these types. The types of auxiliary data corresponding to different types of focus objects may differ. The type of auxiliary data corresponding to each focus object may be pre-defined or obtained by the service device based on dialogue stream data analysis. For example, in step 340 above, the service device performs semantic parsing on the received dialogue stream data. Besides identifying at least one focus object that the semantic content of the dialogue stream data focuses on, it also identifies the type of auxiliary data corresponding to each focus object. The service device can obtain the auxiliary data from the corresponding external data source based on the type of auxiliary data corresponding to each focus object that the user focuses on (i.e., that the semantic content of the dialogue stream data focuses on).

[0077] In some embodiments, data for different focus objects may need to be provided by different data sources, and data of different types for the same focus object may also need to be provided by different data sources. The service device can connect to multiple external data sources to obtain auxiliary data from the appropriate external data source based on the type of the focus object and its corresponding auxiliary data after determining any focus object.

[0078] A focus object can correspond to static data, such as the object's name, affiliated organization, or relationships with other objects. A focus object can also correspond to dynamically changing data (hereinafter referred to as dynamic data), such as data that changes dynamically over time or with events. When a user focuses on an object, the object's dynamic data usually has a greater impact on the user, and the user pays more attention to the object's dynamic data.

[0079] Accordingly, in some embodiments, the auxiliary data for each focus object includes dynamic data corresponding to that focus object, which can be characterized by at least one parameter. The auxiliary data for each focus object describes the dynamic changes of at least one parameter of that focus object over time or events. Thus, the service device obtains the corresponding auxiliary data for the focus object of user interest, obtaining the dynamic changes of at least one parameter of that focus object over time or events. This dynamic change information can be provided to the user, improving the matching degree between the information delivered to the user and the user's needs. Furthermore, the dynamic changes of the parameter can more intuitively and conveniently represent the changes in the dynamic data of the focus object, making it easier for the user to understand the information they are interested in based on this auxiliary data.

[0080] In some embodiments, during a dialogue between a user and an intelligent agent concerning the financial sector, the focus object identified by the service device based on the dialogue flow data can be a financial entity. Static data corresponding to a financial entity may include its name, affiliated institution or issuing institution, and associated industry. Dynamic data corresponding to a financial entity includes its market conditions. Accordingly, if the focus object is a financial entity, the auxiliary data corresponding to that focus object may come from a financial market data source and be used to describe changes in the financial entity's market condition parameters.

[0081] The financial market data source can be a single, independent source, such as an official website. The service device can retrieve market data related to the financial entity from this source and determine the changes in the financial entity's market parameters accordingly. Alternatively, the financial market data source can include multiple data sources, such as multiple websites or data platforms of multiple institutions. The service device can retrieve market data related to the financial entity from each of these multiple data sources and integrate it to obtain the changes in the financial entity's market parameters.

[0082] In some embodiments, market parameters for a financial entity may include price parameters and / or trading activity parameters. The price parameter may include at least one of the following: real-time price (e.g., real-time transaction price and / or average transaction price), stock price index, sector index of the stock's sector, price change, volume ratio, yield, interest rate, price-to-earnings ratio, net asset value, valuation, or volatility. The stock price index may include the Dow Jones Industrial Average, Shanghai Composite Index, CSI 300 Index, or CSI 300 Index, etc. The stock sector may include the ChiNext, technology sector, pharmaceutical sector, semiconductor sector, new energy sector, etc. The trading activity parameter reflects the trading activity of the financial entity. This trading activity parameter may include at least one of the following: trading volume, trading value, or open interest, etc.

[0083] For example, the financial entity is a stock, and the market data parameters for that stock may include at least one of the following: the stock's real-time trading price, the stock's average trading price, the stock's trading volume per minute, a stock price index related to the stock, or a sector index of the stock's sector. The corresponding auxiliary data for this stock can describe the changes in at least one of the above market data parameters. Taking the use of auxiliary data to describe changes in the stock's real-time trading price as an example, the auxiliary data could include the stock's real-time trading price at the current time and at multiple previous time points.

[0084] In some embodiments, market data parameters for different financial entities can overlap. For example, the market data parameters for stock a and stock b can both include the Shanghai Composite Index. Since stock a and stock b belong to the same sector C, their market data parameters can both include the sector index of sector C.

[0085] In some embodiments, the user and the intelligent agent can also engage in dialogues outside the financial sector. During these dialogues, the focus of the service device's identification of the dialogue stream data differs from that of a financial entity. For example, if the user and the intelligent agent are conversing about travel planning, the focus of the service device's identification of the dialogue stream data in this scenario could include a city. The city's dynamic data includes information such as weather, temperature, air quality, and traffic flow. Correspondingly, auxiliary data for the city describes the changes in at least one of these dynamic data. For instance, the auxiliary data could include the city's weather data for multiple dates within a given period, describing the weather changes during that time. As another example, if the user and the intelligent agent are conversing about e-commerce shopping guides, the focus of the service device's identification of the dialogue stream data in this scenario could include a product. The product's dynamic data includes information such as its price and any free gifts offered. Correspondingly, the auxiliary data for the product could include the product's price at various points in time within a given period, describing the price changes over that period.

[0086] In some embodiments, the service device can passively receive auxiliary data for the focus object via a subscription-push approach to improve the timeliness and reduce the latency of auxiliary data acquisition. Accordingly, after identifying at least one focus object that the semantic content of the dialogue stream data is interested in, the service device can send a subscription request for the at least one focus object to an external data source; and receive auxiliary data for the at least one focus object pushed by the external data source in response to the subscription request. In some embodiments, the service device can send the subscription request to the external data source via the WebSocket protocol.

[0087] For example, a subscription request sent by a service device to an external data source may carry the object identifier of at least one focus object. The object identifier may include the name or identifier of the focus object, enabling the external data source to locate the focus object based on the object identifier. The subscription request may also carry type indication information for the auxiliary data corresponding to the at least one focus object. After locating the corresponding focus object based on the object identifier carried in the subscription request, the external data source can obtain the auxiliary data corresponding to the focus object based on the type indication information and send it to the external data source.

[0088] If the auxiliary data is used to describe the changes in the dynamic data of the focus object, the external data source responds to the subscription request by sending the current auxiliary data of the focus object to the service device. This current auxiliary data may include at least one dynamic data point of the focus object within the current time and a period preceding it. For example, the dynamic data may include the values ​​of at least one of the aforementioned parameters of the focus object. The external data source also responds to the subscription request by monitoring the updates to the at least one dynamic data point corresponding to the focus object. When the at least one dynamic data point is updated, the external data source pushes the latest dynamic data to the service device so that the service device receives the latest auxiliary data. The latest auxiliary data includes both the latest dynamic data and previous auxiliary data. The latest dynamic data pushed by the external data source may include the values ​​of at least one of the aforementioned parameters of the focus object.

[0089] For financial entities, their dynamic data changes frequently, and users have a high demand for supplementary data indicating these changes. Through the subscription-push method described above, service devices can obtain the latest supplementary data from financial entities in a timely manner, thus presenting users with more real-time updates on these dynamic data changes. This improves the completeness and timeliness of information delivery to users, providing better support for their financial decision-making.

[0090] In some embodiments, the service device can actively query an external data source to obtain auxiliary data corresponding to the at least one focus object. For example, the service device can send a query request to the external data source for the auxiliary data corresponding to the at least one focus object. This query request may carry the object identifier of the at least one focus object and type indication information of the auxiliary data corresponding to the at least one focus object. For details on this query request, please refer to the previous section on subscription requests; it will not be repeated here. In response to the query request, the external data source can send the auxiliary data corresponding to the at least one focus object to the service device and will not perform any further operations based on this query request. In some embodiments, the service device can periodically send this query request to the external data source to obtain the latest auxiliary data for the at least one focus object.

[0091] The subscription requests and query requests in the above embodiments can both be classified as targeting Figure 1 The description refers to the data retrieval request sent by the server device to the external data source.

[0092] Step 360: The service device sends the at least one focus object and its corresponding auxiliary data to the terminal device.

[0093] Accordingly, the terminal device receives the at least one focus object and its corresponding auxiliary data from the service device.

[0094] The sending or receiving of the at least one focus object as described in this specification refers to sending or receiving identification or indication information of the at least one focus object. This identification or indication information may include at least one of the following: a textual identifier (such as a name or identifier) ​​or a graphic identifier (such as an icon or image) of the focus object.

[0095] In some embodiments, the terminal device may parse the data sent by the service device to obtain the at least one focus object and its corresponding auxiliary data. In some embodiments, the service device may organize the at least one focus object and its corresponding auxiliary data according to a preset format or structure, and then send the organized data to the terminal device so that the terminal device can parse the data.

[0096] In some embodiments, after obtaining auxiliary data corresponding to any focus object from an external data source, the service device can store the auxiliary data in a cache space. When the focus object is identified again for new dialogue stream data, the auxiliary data of the focus object can be directly obtained from the cache space and sent to the terminal device without communicating with the external data source again. This reduces the number of repeated communications between the service device and the external data source, reduces communication resource waste, and improves the timeliness of data response to the terminal device.

[0097] Accordingly, if the service device identifies at least one focus object of semantic interest in the dialogue stream data in step 340, it can first query the cache space for each focus object to see if there is auxiliary data corresponding to that focus object. If it exists, the auxiliary data corresponding to that focus object stored in the cache space is directly sent to the terminal device. If it does not exist, step 350 is then executed to obtain the auxiliary data corresponding to that focus object from an external data source.

[0098] In some embodiments, a refresh cycle is set for the cache space. When the storage time of auxiliary data corresponding to any focus object in the cache space reaches the refresh cycle, the service device deletes the auxiliary data corresponding to that focus object from the cache space. Accordingly, if the focus object is identified again subsequently, the service device needs to retrieve the auxiliary data corresponding to that focus object from the external data source again. The refresh cycle is short, such as less than a preset duration, which can be 60 seconds, 10 seconds, or other durations. Alternatively, the refresh cycle can be in the second range, such as 1 second, 2 seconds, or other second-level durations; correspondingly, the cache space is a second-level cache. Thus, by setting this short refresh cycle for the cache space, the real-time performance of auxiliary data retrieval can be improved while reducing the number of repeated communications with the external data source.

[0099] Step 370: The terminal device displays the at least one focus object and its corresponding auxiliary data in the auxiliary area of ​​the interactive page.

[0100] After obtaining at least one focus object of semantic interest in the dialogue stream data and its corresponding auxiliary data, the terminal device can present it in an auxiliary area of ​​the interactive page (such as...). Figure 4 In the auxiliary area 402 shown. For example, presenting the focus object may include a text or graphic icon representing the focus object. Thus, during human-computer dialogue, auxiliary data corresponding to at least one focus object is delivered to the user; that is, external data related to the dialogue content that the user may want to know is delivered, expanding the functionality of human-computer dialogue. Users do not need to open other software or pages to query this auxiliary data, which can improve the richness of information delivered to the user, meet user needs, reduce user operational complexity, and improve the user experience.

[0101] In some embodiments, the terminal device may present the focus object and its corresponding auxiliary data in a text modal in an auxiliary area. For example, the auxiliary area may list dynamic data of the focus object at different times or events.

[0102] In some embodiments, the at least one focus object may include a target focus object, and the auxiliary data corresponding to the target focus object is presented in a non-textual visualization in an auxiliary area. All of the at least one focus object may be presented in a non-textual visualization, or some focus objects may be presented in text form. Accordingly, in the financial field, changes in market parameters of financial entities can be presented in a non-textual visualization.

[0103] Visualization refers to the interactive and visual representation of data to enhance cognition. It maps data into perceptible graphics, symbols, colors, textures, etc. By leveraging human visual processing capabilities, it enables rapid and efficient understanding and analysis of data, enhancing data recognition efficiency and effectively conveying useful information. Correspondingly, presenting auxiliary data corresponding to the target focus object in a non-textual, visual way can improve the intuitiveness of the auxiliary data presentation, enhance the user's data understanding efficiency and depth, and improve the user experience.

[0104] Figure 4Taking a scenario where the focus objects identified from the dialogue flow data include precious metal G and stock a, both of which are target focus objects, as an example: The dynamic data corresponding to the focus object precious metal G may include the gold price and its current price change; correspondingly, the auxiliary data for precious metal G describes the changes in the gold price over a period of time. Similarly, the dynamic data corresponding to the focus object stock a may include the stock price and its current price change; correspondingly, the auxiliary data for stock a describes the changes in its stock price over a period of time. For example... Figure 4 The percentages shown represent the current increase or decrease of the corresponding dynamic data. A negative percentage indicates a decrease, and a positive percentage indicates an increase.

[0105] In some embodiments, the visualization format includes at least one of the following: graphs, charts, dashboards, images, maps, or cloud maps. The chart may include line charts, bar charts, pie charts, scatter plots, heatmaps, or radar charts, etc. In some embodiments, the selected visualization format can be adjusted accordingly for different focus objects and their corresponding auxiliary data to present each focus object in a suitable visualization format, improving the rationality and intuitiveness of data presentation and conforming to users' viewing habits of the corresponding data. For example, such as... Figure 4 As shown, changes in financial data can be presented using line charts. Similarly, changes in meteorological data can be presented using line charts or cloud maps. Changes in traffic data can be presented using maps.

[0106] In some embodiments, the terminal device can process at least one received focus object and its corresponding auxiliary data according to a corresponding visualization format to obtain content that should be displayed in the auxiliary area. Accordingly, the auxiliary area can be refreshed based on the content to present the content in the auxiliary area.

[0107] In the above embodiments, for the financial sector, changes in market parameters of financial entities are presented in a non-textual, visual format. For example, for financial entities (such as stocks), their candlestick charts, intraday charts, or real-time market data cards can be displayed. In this way, during the dialogue, users can simultaneously view the real-time market data of the financial entities they are interested in using their preferred market data viewing method, effectively achieving automatic market monitoring and expanding the functionality of human-computer interaction. Market monitoring, also known as stock market analysis, is the core behavior of stock investors who analyze market trends by observing real-time stock market data (such as stock index trends, individual stock dynamics, changes in trading volume, etc.). Its core objective is to capture investment opportunities and formulate trading strategies. It mainly focuses on indicators such as opening and closing prices, order volume, and trading distribution, combined with candlestick charts, the movement of major funds, and market sentiment (such as the strength of leading stocks and the number of stocks hitting the daily limit down) for comprehensive judgment.

[0108] In some embodiments, the dynamic data corresponding to the focus object in the external data source changes over time. Such changes may include a change in the value of at least one parameter of the focus object, with the latest dynamic data including the latest value of that parameter. In some embodiments, the terminal device may update and display auxiliary data corresponding to the focus object in the auxiliary area in response to a change in the value of at least one parameter of the focus object in the external data source. This improves the real-time performance of data transmission, making it easier for users to understand the real-time changes of at least one parameter of the focus object, thus providing better assistance for user decision-making.

[0109] For example, after the dynamic data corresponding to the focus object changes, the service device can obtain the latest dynamic data for that focus object. Regarding the acquisition of this latest dynamic data, please refer to the previous description of step 350. For instance, it can passively receive the latest dynamic data pushed by an external data source, or actively query the latest dynamic data from an external data source. After obtaining the latest dynamic data, the service device sends it to the terminal device, enabling the terminal device to obtain updated auxiliary data corresponding to the focus object based on the latest dynamic data, and update the currently presented auxiliary data based on the updated auxiliary data. This process can be referred to the previous descriptions of determining and presenting auxiliary data, and will not be elaborated here.

[0110] In some embodiments, when the number of the at least one focus object exceeds a preset number, the at least one focus object and its corresponding auxiliary data are presented in a scrolling manner in the auxiliary area. The preset number can be determined based on the maximum number of focus objects that the auxiliary area can display simultaneously.

[0111] For example, please continue to refer to Figure 4 Assuming a preset quantity of 2, the service device determines 3 focus objects for a single dialogue stream. These three focus objects and their corresponding auxiliary data can be presented in a scrolling manner within an auxiliary area according to a certain scrolling cycle. For example, in... Figure 4 The data is displayed horizontally in the auxiliary area 402 shown. For example, in one scrolling cycle, the first focus object and its corresponding auxiliary data, as well as the second focus object and its corresponding auxiliary data, are simultaneously displayed in the auxiliary area 402. In the next scrolling cycle, the second focus object and its corresponding auxiliary data, as well as the third focus object and its corresponding auxiliary data, are simultaneously displayed in the auxiliary area 402. In the following scrolling cycle, the third focus object and its corresponding auxiliary data, as well as the first focus object and its corresponding auxiliary data, are simultaneously displayed in the auxiliary area 402.

[0112] In some embodiments, please refer to Figure 4An operation control 403 is provided at the boundary between the auxiliary area 402 and the dialog area 401 in the interactive page 400. The terminal device can respond to the user's operation on the operation control 403 and dynamically adjust the display size of the dialog area 401 and / or the auxiliary area 402.

[0113] For example, the dialog area 401 and the auxiliary area 402 together fill the interactive page 400, with the display sizes of the dialog area 401 and the auxiliary area 402 increasing and decreasing respectively. The user can perform a long press and pull-down operation on the operation control 403 to shrink the display size of the dialog area 401 and correspondingly expand the display size of the auxiliary area 402. The user can also perform a long press and pull-up operation on the operation control 403 to expand the display size of the dialog area 401 and correspondingly shrink the display size of the auxiliary area 402. When expanding (or shrinking) the display size of the auxiliary area 402, more (or fewer) focus objects can be displayed in the auxiliary area 402, or the number of focus objects can remain unchanged while increasing (or decreasing) the display size of the focus objects and their auxiliary data displayed in the auxiliary area 402.

[0114] For example, the dialog area 401 and the auxiliary area 402 can be independent of each other to a certain extent. When a user operates the control 403, they can change only the display size of the dialog area 401 or only the display size of the auxiliary area 402. For instance, when the user interacts with the control 403... Figure 4 The operation control 403 in the middle performs a long press and pull-up operation, which can expand the dialog area 401, and make the expanded dialog area 401 cover at least part of the auxiliary area 402. The user can then... Figure 4 The operation control 403 in the middle can perform a long press and pull-down operation to expand the auxiliary area 402, and make the expanded auxiliary area 402 cover at least part of the dialog area 401.

[0115] In some embodiments, the dialog area 401 functions as a bottom worksheet displayed near the bottom of the page. In some embodiments, the secondary area 402 functions as a top worksheet displayed near the top of the page. Only a portion of the worksheet is displayed upon initial opening. Users can drag or swipe the worksheet to make it full-screen or close it.

[0116] The above embodiments describe the process executed when a terminal device sends a dialogue stream data to a service device. Based on this, the terminal device continuously sends dialogue stream data to the service device. Whenever new dialogue content is added to the dialogue area of ​​the interactive page, the terminal device can send new dialogue stream data to the service device. For details on this dialogue stream data and its transmission, please refer to the preceding description of step 330.

[0117] The focus object displayed in the auxiliary area of ​​the interactive interface can at least include the focus object of the semantic content of the current dialogue flow data. When the current dialogue flow data changes, the focus object of the semantic content of the dialogue flow data may change. Accordingly, the terminal device can update the focus object presented in the auxiliary area in response to the change in the focus object of the semantic content of the dialogue flow data. For example, the updated focus object presented in the auxiliary area includes the focus object of the semantic content of the current dialogue flow data. Based on this, the auxiliary data corresponding to the focus object displayed in the auxiliary area also changes accordingly. This allows for timely adjustment of the presented focus object and its auxiliary data, improving presentation flexibility and matching the presentation of the auxiliary data with the dialogue progress, achieving deep integration of auxiliary data and dialogue context, and improving the assistance effect for the user.

[0118] For example, after a user inputs dialogue content, the terminal device identifies this dialogue content as a dialogue stream data. The semantic focus of this dialogue stream data includes stock 'a'. Accordingly, the terminal device can present stock 'a' and its corresponding auxiliary data in an auxiliary area. Afterward, the agent outputs a response based on the user's dialogue content. This response serves as the next dialogue stream data, whose semantic focus includes both stock 'a' and stock 'b'. Accordingly, the terminal device can add the presentation of stock 'b' and its corresponding auxiliary data in the auxiliary area. For details on identifying stock 'b', acquiring its corresponding auxiliary data, and adding its presentation, please refer to the previous description of focus objects and their corresponding auxiliary data. Alternatively, the service device can send a subscription request for stock 'b' to an external data source and receive the auxiliary data corresponding to the focus object pushed by the external data source. The terminal device then presents the focus object and its corresponding auxiliary data sent by the service device in the auxiliary area accordingly.

[0119] In some embodiments, if the focus object identified for the next dialogue stream data overlaps with the focus object identified for the previous dialogue stream data (such as stock a in the example above), the service device can send a subscription request to the external data source again for the overlapping focus object. That is, the service device can disregard previously subscribed focus objects and send a subscription request to the external data source for each focus object identified for new dialogue stream data. Correspondingly, the external data source can monitor and push data to the focus object based on the new subscription request each time it receives it, and will no longer monitor and push data to previously subscribed focus objects, effectively unsubscribing from the previously subscribed focus objects. After each new dialogue stream data is sent, the terminal device displays the newly sent focus object and its corresponding auxiliary data in the auxiliary area, instead of displaying the previously received focus object and its corresponding auxiliary data.

[0120] In some embodiments, the service device may compare the current focus object with the previous focus object after identifying a focus object for each new dialogue stream data, and only send a subscription request to the external data source for the newly added focus object. For focus objects that have not been canceled, their subscription status remains unchanged, and they continue to receive data pushes from the external data source for that focus object. In some embodiments, if the current focus object is missing a focus object compared to the previous focus object, the service device may send an unsubscribe request to the external data source for that focus object. Accordingly, the external data source may stop monitoring data updates for that focus object and stop pushing data for that focus object to the service device. The terminal device also stops displaying the focus object and its corresponding auxiliary data in the auxiliary area.

[0121] In some embodiments, different attention levels or weights are assigned to each identified focus object based on the user's level of attention. For focus objects with attention levels or weights higher than a certain threshold, the focus object and its corresponding auxiliary data can be displayed in the auxiliary area for a longer period of time. For example, even if the focus object is no longer the focus of new conversation flow data, it can still be displayed for a period of time.

[0122] In traditional human-computer dialogue, the agent passively responds, generating answers to user questions, which are then displayed as plain text on the interactive page. This response model has certain limitations. When user questions involve dynamic data, complex relationships, spatial structures, or process evolution, plain text descriptions cannot efficiently and intuitively convey the full picture of the information. For example, in scenarios such as financial analysis, weather forecasting, and equipment status monitoring, users need to simultaneously understand numerical trends, comparative relationships, geographical distributions, or state changes. Continuous and lengthy text descriptions increase the user's cognitive load, reduce information acquisition efficiency, and may affect the accuracy and timeliness of decision-making.

[0123] In the above embodiments, during the dialogue between the user and the agent, the focus object related to the dialogue and its corresponding auxiliary data are presented in a non-textual visual form in the auxiliary area of ​​the interactive page. Thus, since the dialogue is not solely provided by the agent as static text, but can also actively associate focus objects based on the dialogue content to obtain the corresponding auxiliary data, a semantic-to-data mapping is achieved, constructing a dual-channel fusion mode of dialogue and data presentation. Furthermore, as the dialogue evolves, the presented focus objects are automatically added, deleted, or updated, intuitively presenting the dynamic changes of the focus objects to the user. Users can intuitively and continuously perceive the dynamic changes of the focus objects, mitigating the information silos caused by relying solely on agent-generated responses, enabling users to simultaneously chat and view information, and providing immediate feedback, resulting in a highly fluid user experience.

[0124] For example, in the financial sector, this auxiliary area can simultaneously display dynamic data such as market quotes, candlestick charts, and net asset values ​​of financial entities. This auxiliary area is equivalent to the user's monitoring area. In this scenario, by analyzing the dialogue between the user and the AI ​​agent in real time, the system automatically extracts specific stocks, funds, sectors, or market indicators associated with the dialogue content and dynamically matches them with corresponding real-time market data for visualization. In this way, users can intuitively grasp the latest developments of relevant financial instruments while interacting with AI, satisfying their demand for deep integration of real-time data and dialogue context to a certain extent.

[0125] In the above embodiments, the processing procedure (step 340) based on dialogue stream data and the data acquisition procedure (step 350) are both performed by the service device as an example. In some embodiments, if the data processing capability of the terminal device is powerful enough, these procedures can also be performed by the terminal device; in this case, steps 330 and 360 may not need to be performed.

[0126] Figure 5 A block diagram illustrating an implementation of a method for providing auxiliary data in human-computer dialogue according to embodiments of this specification is shown. Figure 5 This introduction can be referenced in conjunction with the preceding related introductions. For example...Figure 5 As shown, the terminal device can display an interactive page for dialogue between the user and the intelligent agent upon user triggering. The user can enter questions (corresponding to user input in the figure) in the dialogue area of ​​this interactive page, and the intelligent agent's responses (corresponding to AI replies in the figure) can also be displayed in this dialogue area. Based on the dialogue content displayed in the dialogue area, the terminal device can send dialogue stream data to the service device, which includes at least a portion of the dialogue content displayed in the dialogue area.

[0127] Please continue to refer to this. Figure 5 The service device may include a dialogue understanding engine, a focus object identification module, a data service module, a data center, and a cache space. The data center includes its own maintained data sources. The data center also connects to third-party data sources; for example, the data center may have an interface connecting to a third-party data source, through which data from the third-party data source can be obtained. The aforementioned external data sources are data sources that are external to the intelligent agent; both the data source maintained by the data center itself and the third-party data source belong to the aforementioned external data sources. The cache space can be a second-level Redis cache space. For example, the dialogue understanding engine and the focus object identification module can be used to implement step 340 above, the data service module and the data center can be used to implement step 350 above, and step 360 above can be executed by the data service module in the service device.

[0128] Please continue to refer to this. Figure 5 The dialogue stream data sent by the terminal device to the service device can be sent to the dialogue understanding engine. The dialogue understanding engine performs semantic understanding on the dialogue stream data to obtain structured semantic content. This structured semantic content is transmitted to the focus object identification module, which can identify the focus objects that the structured semantic content focuses on, and obtain a focus object list. The focus object list is transmitted to the data service module, which sends a subscription request or query request for auxiliary data for each focus object in the focus object list to the data center. If the auxiliary data corresponding to a focus object can be directly provided by the data source maintained by the data center, the data center sends the corresponding auxiliary data to the data service module. If the auxiliary data corresponding to a focus object cannot be directly provided by the data source maintained by the data center, the data center can obtain the corresponding auxiliary data from the corresponding third-party data source and send it to the data service module.

[0129] Please continue to refer to this. Figure 5After obtaining the auxiliary data corresponding to any focus object, the data service module can write the auxiliary data into the cache space and transmit the list of focus objects and their corresponding auxiliary data to the terminal device, allowing the terminal device to display them in the auxiliary area of ​​the interactive page. When obtaining the list of focus objects, the data service module can first check if there is auxiliary data corresponding to the relevant focus object in the cache space. If so, it can directly read the auxiliary data from the cache space; otherwise, it can send a request to the data center. Based on the focus objects and their corresponding auxiliary data displayed in the auxiliary area, the user can then enter a question in the dialog area of ​​the interactive page, and so on.

[0130] In the above embodiments, dialogue understanding, data services, and dynamic data display on the page are coordinated to form a closed-loop intelligent interaction system that integrates semantic understanding, focus object recognition, data subscription, visualization, and user feedback. This can reduce the limitations of information transmission to users in human-computer dialogue.

[0131] In summary, the method for providing auxiliary data in human-computer dialogue provided in this specification, while presenting the dialogue flow data between the user and the intelligent agent on the interactive page, also displays the semantic focus objects of the dialogue flow data and their auxiliary data in an auxiliary area of ​​the interactive page. This auxiliary data is provided by an external data source not controlled by the intelligent agent, and can assist the user in understanding external data related to the dialogue content during the dialogue. This allows users to understand more data about their focus objects during human-computer dialogue, enriching the information conveyed to the user, improving the completeness and sufficiency of the information, and reducing the limitations of information delivery to users in human-computer dialogue.

[0132] This specification, in another aspect, provides a computer-readable non-transitory storage medium storing at least one instruction set. When the at least one instruction set is executed by a processor, it instructs the processor to perform the steps executed by the terminal device or the service device in the method P300 described herein. In some possible embodiments, various aspects of this specification can also be implemented as a program product comprising program code. When the program product is run on an electronic device 200, the program code causes the electronic device 200 to perform the steps executed by the terminal device or the service device in the method P300 described herein. The program product for implementing the above method may employ a portable compact disc read-only memory (CD-ROM) containing program code and may run on the electronic device 200. However, the program product of this specification is not limited thereto. In this specification, the readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system. The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. A readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a readable storage medium include: portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. The computer-readable storage medium may include a data signal propagated as part of a carrier wave in baseband, carrying readable program code. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof. Program code for performing the operations described herein can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on electronic device 200, partially on electronic device 200, as a standalone software package, partially on electronic device 200 and partially on a remote computing device, or entirely on a remote computing device.

[0133] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0134] In summary, after reading this detailed disclosure, those skilled in the art will understand that the foregoing detailed disclosure is presented by way of example only and is not restrictive. Although not explicitly stated herein, those skilled in the art will understand that this specification requires various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be made by this specification and are within the spirit and scope of the exemplary embodiments described herein.

[0135] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, "an embodiment," "an embodiment," and / or "some embodiments" mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of this specification. Therefore, it is to be emphasized and understood that two or more references to "an embodiment" or "an embodiment" or "alternative embodiment" in various parts of this specification do not necessarily refer to the same embodiment. Moreover, specific features, structures, or characteristics may be suitably combined in one or more embodiments of this specification.

[0136] It should be understood that in the foregoing description of the embodiments in this specification, various features are combined in a single embodiment, drawing, or description for the purpose of simplifying the description and to aid in understanding a feature. However, this does not mean that the combination of these features is necessary, and those skilled in the art, upon reading this specification, may readily identify some of the devices as separate embodiments. That is, the embodiments in this specification can also be understood as an integration of multiple secondary embodiments. And the content of each secondary embodiment is valid even if it contains fewer than all the features of a single foregoing disclosed embodiment.

[0137] Every patent, patent application, publication of a patent application, and other material such as articles, books, specifications, publications, documents, articles, etc., cited herein, except for those inconsistent with or conflicting with this document, or those having a restrictive effect on the widest scope of the claims, may be incorporated herein by reference for all purposes now or hereafter associated with this document. Furthermore, in the event of any inconsistency or conflict between the description, definition, and / or use of relevant terms in any material and the description, definition, and / or use of relevant terms in this document, the terms in this document shall prevail.

[0138] Finally, it should be understood that the embodiments disclosed herein are illustrative of the principles of the embodiments described in this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can implement the applications described in this specification using alternative configurations based on the embodiments in this specification. Therefore, the embodiments in this specification are not limited to the embodiments precisely described in the applications.

Claims

1. A method for providing auxiliary data in human-computer dialogue, applied to a terminal device, the method comprising: An interactive page for engaging in dialogue with an intelligent agent is displayed, the interactive page including a dialogue area and an auxiliary area that are independent of each other and displayed on the same screen; as well as During the dialogue between the user and the agent, the dialogue flow data between the user and the agent is presented in the dialogue area, and at least one focus object and its corresponding auxiliary data are presented in the auxiliary area. The focus object is the object of interest in the semantic content of the dialogue stream data, and the auxiliary data is provided by an external data source that is not controlled by the agent and is used to help the user understand the external data associated with the dialogue content during the dialogue.

2. The method according to claim 1, wherein, The auxiliary data is used to describe the dynamic changes of at least one parameter of the focus object over time or events.

3. The method according to claim 2, further comprising: In response to a change in the value of at least one parameter of the focus object in the external data source, the auxiliary data corresponding to the focus object is updated and displayed in the auxiliary area.

4. The method according to claim 2, wherein the focus object is a financial entity, and the auxiliary data comes from a financial market data source and is used to describe the changes in the market parameters of the financial entity.

5. The method according to claim 4, wherein, The financial entity includes at least one of the following: stocks, funds, bonds, physical assets, monetary assets, or financial derivatives; and the market parameters include price parameters and / or trading activity parameters.

6. The method according to claim 1, further comprising: In response to a change in the focus object of the semantic content of the dialogue stream data, the presented focus object in the auxiliary area is updated.

7. The method according to claim 1, wherein, The at least one focus object includes a target focus object, and the auxiliary data corresponding to the target focus object is presented in a non-textual visualization form.

8. The method according to claim 7, wherein, The visualization format includes at least one of the following: graphics, charts, dashboards, images, maps, or cloud maps.

9. The method according to claim 1, wherein, In the interactive page, the auxiliary area is located above, to the left, or to the right of the dialogue area.

10. The method according to claim 1, wherein, An operation control is provided at the boundary between the auxiliary area and the dialogue area, and the method further includes: In response to the user's operation of the control, the display size of the dialog area and / or the auxiliary area is dynamically adjusted.

11. The method according to claim 1, wherein, When the number of the at least one focus object is greater than a preset number, the at least one focus object and its corresponding auxiliary data are presented in a scrolling manner in the auxiliary area.

12. The method according to claim 1, further comprising: Send the dialogue stream data to the service device and receive the at least one focus object and its corresponding auxiliary data from the service device.

13. A method for providing auxiliary data in human-computer dialogue, applied to a service device, the method comprising: Receive dialogue stream data from a terminal device, the dialogue stream data including at least a portion of the dialogue content presented in a dialogue area on an interactive page, the interactive page being a page where the user engages in dialogue with the intelligent agent; Semantic parsing is performed on the dialogue stream data to identify at least one focus object of interest in the semantic content of the dialogue stream data; Obtain auxiliary data corresponding to each of the at least one focus object from an external data source that is not controlled by the intelligent agent; as well as The terminal device sends the at least one focus object and its corresponding auxiliary data to the terminal device for presentation in the auxiliary area of ​​the interactive page, wherein the auxiliary data is used to assist the user in understanding external data related to the dialogue content during the dialogue.

14. The method according to claim 13, wherein, The step of performing semantic parsing on the dialogue stream data to identify at least one focus object of interest in the semantic content of the dialogue stream data includes: The dialogue stream data is input into a large language model, which is then guided to perform semantic understanding of the dialogue stream data and identify the focus object of the dialogue based on the understood semantic content.

15. The method according to claim 13, wherein, Obtaining auxiliary data corresponding to each of the at least one focus object from an external data source not controlled by the intelligent agent includes: Send a subscription request for the at least one focused object to the external data source; and Receive auxiliary data for at least one focus object pushed by the external data source in response to the subscription request.

16. The method according to claim 13, wherein, The auxiliary data is used to describe the dynamic changes of at least one parameter of the focus object over time or events.

17. The method according to claim 16, wherein the focus object is a financial entity, and the auxiliary data comes from a financial market data source and is used to describe the changes in the market parameters of the financial entity.

18. The method according to claim 17, wherein, The financial entity includes at least one of the following: stocks, funds, bonds, physical assets, monetary assets, or financial derivatives; and the market parameters include price parameters and / or trading activity parameters.

19. An electronic device comprising: At least one storage medium storing at least one instruction set; as well as At least one processor is communicatively connected to the at least one storage medium, wherein the at least one processor reads the at least one instruction set and executes the method of any one of claims 1 to 18 according to the instructions of the at least one instruction set.

20. A computer-readable non-volatile storage medium, wherein, The computer-readable non-volatile storage medium stores at least one instruction set, which, when executed by at least one processor, implements the method as described in any one of claims 1 to 18.