Information interaction method and device, electronic equipment, storage medium and program product
By setting up user interaction and response areas in the task interaction interface, and using input components and large models to generate diverse response content, the problem of monotonous user interaction methods is solved, and diverse interaction forms and high real-time performance are achieved, thereby improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-21
AI Technical Summary
In existing technologies, the interaction methods between users and large models are singular and lack diversity, resulting in an inadequate user experience.
This invention provides an information interaction method and apparatus that sets up a user interaction area and a response area in a task interaction interface, uses input components to trigger diverse input events, and combines a large model and rule base to generate diverse response content, supporting diverse interactive task scenarios.
It enriches the interaction methods, meets users' diverse and real-time needs, and enhances the user experience.
Smart Images

Figure CN122431564A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence, and more particularly to the field of human-computer interaction. Specifically, this disclosure relates to an information interaction method, apparatus, electronic device, storage medium, and program product. Background Technology
[0002] In the field of artificial intelligence, the methods for users to interact with large models are singular and fixed. For example, current large models generally only support voice or text interaction, lacking diverse interaction methods, and the user experience needs to be improved. Summary of the Invention
[0003] This disclosure provides at least one of the following information interaction methods, apparatuses, electronic devices, storage media, and program products to solve the aforementioned technical problems.
[0004] According to one aspect of this disclosure, an information exchange method is provided, comprising: The interactive interface for displaying interactive tasks includes a user interaction area and a response area. The user interaction area includes at least one input component corresponding to the task type of the interactive task. The input component is used to trigger an input event corresponding to the task type. In response to the input event being triggered, response content determined based on the event information of the input event is displayed in the response area.
[0005] According to another aspect of this disclosure, an information interaction device is provided, comprising: The first display unit is configured to display the task interaction interface of the interactive task. The task interaction interface includes a user interaction area and a response area. The user interaction area includes at least one input component corresponding to the task type of the interactive task. The input component is used to trigger an input event corresponding to the task type. The second display unit is configured to display response content determined based on event information of the input event in the response area in response to the triggering of the input event.
[0006] According to another aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods described in the embodiments of this disclosure.
[0007] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the methods described in embodiments of this disclosure.
[0008] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods described in the embodiments of this disclosure.
[0009] This disclosure allows for the display of an interactive task interface, which includes a user interaction area and a response area. The user interaction area includes at least one input component corresponding to the task type of the interactive task. The input component is used to trigger an input event corresponding to the task type. In response to the triggered input event, response content determined based on the event information of the input event is displayed in the response area. In other words, in diverse interactive task scenarios, appropriate input components can be displayed in the user interaction area to provide diverse input methods, enriching the interaction format. Furthermore, the collaborative display of the user interaction area and the response area enhances the display effect, thereby meeting the diverse and highly real-time interactive needs of users and improving the user experience.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure.
[0012] Figure 1 This is the system architecture diagram to which this disclosure applies.
[0013] Figure 2 This is a flowchart of the information exchange method provided in this publication.
[0014] Figure 3 This is a schematic diagram of the initial interactive interface provided in this publication.
[0015] Figure 4 This is a schematic diagram of the task interaction interface for an interactive task provided in this publication.
[0016] Figure 5 This is a schematic diagram of the task interaction interface for another interactive task provided in this publication.
[0017] Figure 6 This is a schematic diagram of the task interaction interface for another interactive task provided in this publication.
[0018] Figure 7This is a schematic block diagram of the information interaction device provided in this disclosure.
[0019] Figure 8 This is a block diagram of an electronic device used to implement the information interaction method of the embodiments of this disclosure. Detailed Implementation
[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0021] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0022] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0023] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."
[0024] In the field of intelligent question answering in artificial intelligence, users can currently only initiate interaction requests to large models in a relatively simple way, such as entering text through a text input box or voice input through a voice input interface. After processing these interaction requests, the large model returns the response content in the form of text or voice. Although this linear dialogue paradigm of "one question and one answer" achieves the basic interaction process, its interaction form is inherently quite limited.
[0025] Therefore, in complex or diverse interactive tasks, the relatively simple interaction methods in existing technologies cannot meet the user's interaction needs.
[0026] In view of this, this disclosure provides a new approach. To facilitate understanding of this disclosure, the system architecture on which this disclosure is based will first be described. Figure 1 Exemplary system architectures that can be applied to embodiments of this disclosure are shown, such as Figure 1 As shown, the system architecture may include: a client and a server.
[0027] The server side and the client side are the two main components of an application service. The server side uses a server as its primary hardware infrastructure and may include one or more software service modules. The server side and the client side form a collaborative front-end and back-end.
[0028] The client can be set on the terminal device. In this embodiment of the disclosure, the client can be a local application, a mini-program, or a web application running through a browser on the terminal device.
[0029] Terminal devices can include, but are not limited to, smart mobile terminals, wearable devices, PCs (Personal Computers), and smart home devices. Smart mobile devices can include devices such as mobile phones, tablets, laptops, PDAs (Personal Digital Assistants), and connected car terminals. Wearable devices can include devices such as smartwatches, smart glasses, smart bracelets, VR (Virtual Reality) devices, AR (Augmented Reality) devices, and mixed reality devices (devices that support both virtual and augmented reality). Smart home devices can include devices such as smart TVs and smart refrigerators with displays.
[0030] A server can be a single server, a server cluster consisting of multiple servers, or a cloud server. A cloud server, also known as a cloud computing server or cloud host, is a hosting product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Servers (VPS) services, such as high management difficulty and weak service scalability.
[0031] It should be understood that Figure 1 The number of client and server components shown is merely illustrative. Depending on implementation needs, there can be any number of client and server components.
[0032] Figure 2 This is a flowchart of an information interaction method provided in an embodiment of this disclosure. The information interaction method can be... Figure 1 The server-side execution in the system shown. For example... Figure 2As shown, the information exchange method may include the following steps: Step 201: Display the task interaction interface of the interactive task. The task interaction interface includes a user interaction area and a response area. The user interaction area includes at least one input component corresponding to the task type of the interactive task. The input component is used to trigger an input event corresponding to the task type.
[0033] Step 202: In response to the triggered input event, display the response content determined based on the event information of the input event in the response area.
[0034] As can be seen from the above process, this disclosure can display the task interaction interface for interactive tasks. The task interaction interface includes a user interaction area and a response area. The user interaction area includes at least one input component corresponding to the task type of the interactive task. The input component is used to trigger an input event corresponding to the task type. In response to the triggered input event, the response area displays response content determined based on the event information of the input event. In other words, in diverse interactive task scenarios, appropriate input components can be displayed in the user interaction area to provide diverse input methods, enriching the interaction forms. Furthermore, the user interaction area and the response area are displayed collaboratively, improving the display effect and thus meeting the diverse and highly real-time interactive needs of users, thereby enhancing the user experience.
[0035] The following describes in detail each step of the above process and the effects that can be further produced, with reference to the embodiments.
[0036] First, the above step 201, namely "displaying the task interaction interface of the interactive task, the task interaction interface including a user interaction area and a response area, the user interaction area including at least one input component corresponding to the task type of the interactive task, and the input component being used to trigger an input event corresponding to the task type", will be described in detail with reference to the embodiments.
[0037] In the embodiments of this disclosure, the task interaction interface can refer to a user interface used to present interactive tasks and provide interactive functions between the user and the large model. For example, a task interaction interface displayed on a terminal device includes at least a user interaction area and a response area, the layout of which can be top-bottom, left-right, or other arbitrary forms, and the ratio of the user interaction area to the response area can be dynamically adjusted based on the task type of the interactive task.
[0038] This disclosure can directly display the task interaction interface of the interactive task, or it can first display the initial interaction interface and then display the task interaction interface of the interactive task. Specifically, the initial interaction interface is displayed first, and the task interaction interface of the interactive task is displayed in response to the triggering event of the launch component corresponding to the interactive task in the initial interaction interface. Alternatively, the initial interaction interface is displayed first, and the task interaction interface of the interactive task is displayed in response to the natural language command submitted by the user containing the launch instruction for the interactive task.
[0039] like Figure 3 As shown, Figure 3 A schematic diagram of an initial interactive interface is provided. This initial interactive interface includes launch components for various interactive tasks, such as the launch component corresponding to the task of "ping-pong interaction" (i.e., the user interacting with the agent while playing ping-pong). Figure 3 The first startup component in the process), the startup component corresponding to the "gesture interaction" task (i.e., the user and the agent each select a virtual gesture, there are scoring rules between different virtual gestures, and the response content is determined based on the virtual gestures selected by the two parties and the scoring rules) (such as the startup component of the first startup component in the process), the "gesture interaction" task (i.e., the user and the agent each select a virtual gesture, there are scoring rules between different virtual gestures, and the response content is determined based on the virtual gestures selected by the two parties and the scoring rules) Figure 3 The second startup component in the system) and the startup components corresponding to the "question-generating dictation" task (i.e., the agent generates questions, and the user answers) (such as...) Figure 3 The third launch component (in the context of the initial interface) responds to user trigger events on the first, second, or third launch component, displaying the task interaction interface corresponding to the triggered launch component. Additionally, the initial interaction interface may include a voice component, allowing users to submit launch commands for interactive tasks by triggering the voice component. For example, a user could say, "Start the 'Ping Pong Interaction' task."
[0040] By setting up an initial interactive interface that hosts multiple interactive task entry points, the system provides users with a visual entry point for selecting interactive tasks. It also supports launching interactive tasks directly via natural language commands, offering diverse task triggering methods. This allows users to conveniently launch interactive tasks and load the corresponding task interfaces by triggering launch components or through voice input, thereby improving the solution's usability, flexibility, and launch efficiency.
[0041] When displaying the task interaction interface for an interactive task, one possible approach is to construct the interface based on a preset page template and dynamically loaded component resources. Specifically, a preset page template is obtained; based on the task type of the interactive task, component resources corresponding to that task type are determined; and based on the preset page template and component resources, the task interaction interface for that interactive task is displayed. In this embodiment, the task interaction interfaces for different task types are obtained using the same page template but different component resources.
[0042] The page template defines the basic layout framework of the task interaction interface, while the component resources are dynamically adapted and populated according to the specific task type. This method of building task interaction pages can quickly generate task interaction interfaces that support different interactive tasks based on the same page template, thereby greatly improving the reusability and maintainability of interface development.
[0043] As another feasible approach, the task interaction interface can also be constructed based on page templates for task interaction pages of different interactive tasks and dynamically loaded component resources. Specifically, a page template for the task interaction page corresponding to the task type is obtained; based on the task type of the interactive task, component resources corresponding to the task type are determined; and based on the page template and component resources of the task interaction page corresponding to the task type, the task interaction interface for that interactive task is displayed. In this embodiment, the task interaction interfaces for interactive tasks of different task types are obtained based on different page templates and different component resources.
[0044] As a specific example, if the interactive task is a "ping-pong interaction" task, then the corresponding task interaction interface can be as follows: Figure 4 As shown, the page template defines a basic layout framework with upper and lower sections, with the upper section being the responsive area and the lower section being the user interaction area. Based on the task type "ping-pong interaction," the corresponding component resources are determined. Component resources can refer to... Figure 4 The "racket" component, which can be swiped left and right in the user interaction area and response area, can also refer to the visual element in the middle of the task interaction interface used to represent the "dividing line". These component resources are loaded into the page template to obtain the task interaction interface. The "racket" component in the user interaction area is the input component used to trigger the input event corresponding to the task type.
[0045] As another specific embodiment, if the interactive task is a "gesture interaction" task, then the task interaction interface corresponding to the interactive task can be as follows: Figure 5 As shown, the page template defines a basic layout framework with upper and lower sections, with the upper section being the responsive area and the lower section being the user interaction area. Based on the task type of "gesture interaction," the corresponding component resources are determined. These component resources can refer to... Figure 5 The first gesture selection component, the second gesture selection component, and the third gesture selection component in the user interaction area are loaded into the page template to obtain the task interaction interface. The first gesture selection component, the second gesture selection component, and the third gesture selection component in the user interaction area are the input components used to trigger input events corresponding to the task type.
[0046] As another specific example, if the interactive task is a "question dictation" task, then the corresponding task interaction interface can be as follows: Figure 6 As shown, the page template defines a basic layout framework with upper and lower sections, with the upper section being the responsive area and the lower section being the user interaction area. Based on the task type "question dictation," the corresponding component resources are determined. Component resources can refer to... Figure 6 The canvas in the user interaction area is used for inputting touch trajectories, and can also be used to point to... Figure 6 The question-generating component in the response area is used to display the question. These component resources are loaded into the page template to obtain the task interaction interface. The canvas in the user interaction area used for inputting touch trajectories is the input component used to trigger input events corresponding to the task type.
[0047] The following describes in detail step 202, namely "in response to a triggered input event, displaying the response content determined based on the event information of the input event in the response area", with reference to the embodiments.
[0048] In response to a triggered input event, the response area displays response content determined based on the event information of the input event. For example, in an interactive task of "gesture interaction," an input event is obtained based on the user's triggering operation on one of the first, second, or third gesture selection components in the user interaction area. The event information of this input event may include the event type (e.g., click event), timestamp (e.g., year, month, day, hour, minute, second), and identifier of the triggering component (e.g., identifier of the first gesture selection component). The response area displays the response content determined based on the event information of this input event, which may include the judgment result (e.g., user wins) or score change (e.g., user score increases by one point).
[0049] The process of determining the response content based on the event information of the input event can be implemented through the first major model or a preset rule base. For example, for interactive scenarios with clear rules and low latency, or when the generated response content is relatively simple (such as judgment results or score changes), the process can be directly implemented using the preset rule base. However, for scenarios that require semantic understanding, creative generation, or processing of complex natural language, or when the generated response content is relatively complex (such as judgment reasons or encouraging statements), the process can be implemented using the first major model.
[0050] Furthermore, the process of generating response content is not limited to calling a single model or rule base. In practical applications, a combination of calls can be made according to the actual scenario to balance response speed and response intelligence. For example, in a "question-generating dictation" task, the rule base can be called first to generate a judgment result based on the input event, and then the first main model can be called to generate accompanying explanatory text based on the input event and the judgment result. The judgment result and explanatory text are then displayed in the response area.
[0051] In practical applications, in order to generate more accurate response content, this disclosure can also combine the task status information of the interactive task. Specifically, through the first major model or a preset rule base, response content can be generated based on the event information of the input event and the task status information of the interactive task.
[0052] Task status information refers to the interactive information recorded in each round of the interactive task. This information is not static but updates in real-time with each user input event and its resulting response. For example, in a "gesture interaction" task, the task status information could include multiple pieces of structured data. Each piece of structured data includes a round identifier (i.e., the current round number), gesture records for both parties (i.e., the gestures selected by each party in the current round), round judgment result (i.e., the judgment result based on the gesture records), and scoring information (i.e., the total score for both parties in the interactive task).
[0053] By introducing a primary model or a pre-defined rule base as the engine for generating response content, and by comprehensively utilizing the event information of the input event and the real-time updated task status information as the basis for generation, the generation of response content can be closely integrated with the task progress of the interactive task, thereby improving the accuracy of the response content and meeting the needs for diverse response generation mechanisms.
[0054] The following are three ways to generate and display response content using the first major model or rule base.
[0055] The first approach responds to a preset task interaction event as the input event. Based on the task interaction event and task status information, and using a rule base, response content is generated. Continuing with the previous example, in the "gesture interaction" task, the preset task interaction events are trigger events for the first, second, and third gesture selection components. When the input event is a trigger event for the first gesture selection component, based on the trigger event for the first gesture selection component, task status information, and the judgment table recorded in the rule base, the judgment result between the user's selected first gesture and the agent's selected second gesture (the agent's selection can be determined at the start of the current round or simultaneously with the user's selection) is obtained (e.g., user wins). Then, based on the scoring information included in the task status information, the total score for the user and the agent is obtained (e.g., if the agent's total score remains unchanged, the user's total score is increased by 1). The agent's selected second gesture, the judgment result, the user's total score, and the agent's total score are displayed as response content in the response area.
[0056] Furthermore, to enhance the immersive experience of the response while ensuring low latency in interaction, this application also provides the following solution. In one implementation, when the response content is a first task interaction result determined based on a rule base (such as "score +1", "correct answer", etc.), after displaying the first task interaction result in the response area, a second major model can be invoked to generate supplementary description content for the first task interaction result based on the first task interaction result, task status information, and event information of the input event. The supplementary description content can refer to text content (such as "Your basic knowledge is very solid! You answered such a difficult question correctly!"), image content (such as a praise gesture image), etc. This disclosure can display the supplementary description content in the response area, or play the text content in the supplementary description content to the user after speech synthesis. The second major model can be the same as or different from the first major model, as long as it can implement the above method, this disclosure does not limit it in this regard.
[0057] This solution introduces an asynchronous supplementary description generation step, which balances low latency and high expressiveness. In other words, it can ensure low-latency response to the results of the first task interaction (such as the instant display of key results like "score"), while also increasing the understandability and richness of the response (such as generating vivid explanations or comments).
[0058] Furthermore, after determining the result of the first task interaction, this disclosure can also update the task status information based on the event information of the input event and the result of the first task interaction, and determine the task status change information corresponding to the input event, displaying the changed content determined based on the task status change information in the user interaction area. For example, in the "question dictation" task, the task status information includes the question sequence and progress (i.e., the list of questions included in the current "question dictation" task and which question is being answered in the current round), the current question content (i.e., the question being answered in the current round), the current standard answer (i.e., the standard answer to the question being answered in the current round), the user's answer content (the answer content entered by the user through the user interaction area), the judgment result (which can be comparing whether the user's answer content is consistent with the standard answer of the current question, and the judgment result is "correct", "incorrect", etc., or it can determine the similarity between the user's answer content and the standard answer of the current question, and use the similarity as the judgment result), and the scoring information (the scoring information determined based on the judgment results of historical questions, such as the number of questions answered correctly, the total score, or the accuracy rate).
[0059] As a specific example, in the "dictation" task, the task status information is as follows: "Question 1 (out of 3): What is the English word for 'book'? 'book', 'bool', incorrect, 0 points; Question 2 (out of 3): What is the English word for 'zoo'? 'zoo', 'zoo', correct, 1 point; Question 3 (out of 3): What is the English word for 'water'? 'water', [empty], [empty], [empty]." If the "dictation" task progresses to the third question, in response to the input event being an input touch trajectory event, the first task interaction result is obtained based on the event information of the input touch trajectory event through a preset rule base. Then, based on the event information of the input touch trajectory event and the first task interaction result, the task status change information is obtained. That is, for the third question, the user's answer can be "water", the judgment result is correct, and the score is 2 points.
[0060] In other words, based on the task status change information, the task status information can be updated as follows: "Question 1 (out of 3): What is the English word for 'book'? 'book', 'bool', incorrect, 0 points; Question 2 (out of 3): What is the English word for 'zoo'? 'zoo', 'zoo', correct, 1 point; Question 3 (out of 3): What is the English word for 'water'? 'water', 'water', correct, 2 points." Based on the task status change information, the changed content can also be determined, such as the text "Total score +1", which will be displayed in the user interaction area.
[0061] By dynamically updating task status information based on input events and the resulting first task interaction results, and extracting the changes to display synchronously in the user interaction area, users can directly and instantly perceive the specific impact of their interaction behavior on the task progress (such as scores and rounds) within their own operating area, enhancing their sense of control and immersion in the interactive process.
[0062] The second approach responds to input events that involve submitting natural language commands. The first main model then generates the response content based on the natural language command and task status information. Continuing with the previous example, in the "gesture interaction" task, if the input event is a natural language command (e.g., the user inputs the voice "You're awesome"), the first main model generates the response content based on the natural language command and task status information. This response content can include an animated emoticon expressing "smugness" and the text "Thank you for the compliment, you're awesome too." This emoticon and text can be displayed in the response area.
[0063] In the second approach, the task status information can be updated based on the event information and response content of the event in which the natural language command is submitted, and the task status change information corresponding to the event in which the natural language command is submitted can be determined. The changed content determined based on the task status change information can then be displayed in the user interaction area.
[0064] The third approach responds to input events, including task interaction events and natural language command submissions, by using the first main model to generate response content based on the task interaction events, natural language commands, and task status information. Continuing with the previous example, in the "gesture interaction" task, if the user clicks the first gesture selection component while simultaneously inputting the voice command "This time I will definitely succeed," the first main model can be directly invoked to generate response content based on the task interaction events, natural language commands, and task status information. This response content can include the judgment result, the user's total score, the agent's total score, and the text "Congratulations, you really succeeded this time." This disclosure can display the judgment result, the user's total score, and the agent's total score in the response area, and broadcast the text "Congratulations, you really succeeded this time" to the user after speech synthesis.
[0065] In the third approach, task status information can be updated based on event information of events that submit natural language commands, event information of task interaction events, and response content. The task status change information corresponding to events that submit natural language commands and task interaction events can be determined, and the changed content determined based on the task status change information can be displayed in the user interaction area.
[0066] Furthermore, before generating response content using the first model based on task interaction events, natural language commands, and task status information, a third model can be used to generate fusion instructions based on the task interaction events and natural language commands. These fusion instructions, along with the task status information, are then input into the first model to obtain the response content generated by the first model. For example, in a picture-filling interactive task, a user can draw a picture through the user interaction area (i.e., inputting a touch trajectory event) and input "Please fill this picture with red, yellow, and blue" via voice (i.e., submitting a natural language command event). The third model can then parse the touch trajectory input event and the natural language command event to obtain an accurate fusion instruction. Subsequently, the first model can use this accurate fusion instruction and task status information to obtain response content that better meets the user's needs.
[0067] It can be seen that for structured task interaction events, the rule base can be used to quickly generate response content, ensuring the low latency requirement for real-time interaction scenarios; for unstructured natural language commands, the semantic understanding and generation capabilities of the first model are used to improve the flexibility and intelligence of the response content; when both occur simultaneously, the first model can generate richer and more accurate response content, thereby optimizing the interactive experience and overall performance in mixed input scenarios.
[0068] It should be noted that when responding to an input event, the input event can be preprocessed before determining the response content based on its event information. The response content can then be generated based on this preprocessed input event. For example, the input event information may include a timestamp, but the format of the timestamp used by each input event may differ. The timestamps of all input events can be standardized to a preset format, and then the response content can be generated based on the event information of the standardized input events.
[0069] For example, if the input event is an event that triggers the input touch trajectory, in response to the event that triggers the input touch trajectory, the touch trajectory is smoothed to obtain a smoothed touch trajectory. Based on the smoothed touch trajectory, the touch trajectory in the event information of the input touch trajectory event is updated. Then, based on the updated event information, response content is generated.
[0070] In addition, the event information of the event of input touch trajectory also includes the recognition result of touch trajectory. The recognition result is obtained by performing text recognition and / or graphic recognition on the touch trajectory. After the touch trajectory is smoothed as described above, text recognition and / or graphic recognition can be performed again on the smoothed touch trajectory to obtain a more accurate recognition result. The recognition result in the event information of the event of input touch trajectory is updated based on the more accurate recognition result.
[0071] As a specific embodiment, if the input event is an event involving an input touch trajectory, meaning the event information of the input touch trajectory event includes the touch trajectory, then the touch trajectory is obtained based on the original trajectory drawn by the user in the user interaction area. This original trajectory may have issues such as jitter or discontinuity. Therefore, the touch trajectory can be smoothed to obtain a smoothed touch trajectory. Subsequently, text recognition can be performed on the smoothed touch trajectory to obtain a text recognition result (e.g., "book"), and the recognition result in the event information of the input touch trajectory event can be updated based on the text recognition result. Alternatively, image recognition can be performed on the smoothed touch trajectory to obtain an image recognition result (e.g., "triangle"), and the recognition result in the event information of the input touch trajectory event can be updated based on the image recognition result.
[0072] By smoothing the touch trajectory to optimize data quality, the accuracy of event information of the input touch trajectory can be improved, making it easier to generate response content that better meets user needs, thus improving the accuracy and intelligence of the interaction.
[0073] Next, the process of displaying the response content in the response area will be explained in detail. Specifically, a first interface element corresponding to the response content is determined. The first interface element is used to represent the response content, and the first interface element is displayed in the response area.
[0074] As can be seen, the purpose of this process is to transform structured response content into visual first interface elements, which can refer to visual elements such as animations, text, images, and special effects. The process of determining the first interface element can either involve identifying the first interface element that matches the response content from a pre-defined interface element table, or it can involve directly generating the first interface element corresponding to the response content using an interface element generation model.
[0075] By transforming response content into first-interface elements, the same response content can be presented in different styles depending on different device capabilities or skin themes, thus increasing the richness of the display. Compared to simply displaying response content in the response area, this method of displaying visual first-interface elements improves the display effect of the response area, thereby enhancing the user experience.
[0076] While displaying the response content in the response area, a second interface element corresponding to the input event can also be displayed in the user interaction area. The process of determining the second interface element can be to determine the second interface element that matches the event information of the input event from a preset interface element table, or to directly generate the second interface element corresponding to the event information of the input event through an interface element generation model, or to generate a trajectory image as the second interface element based on the touch trajectory in the event of the input touch trajectory.
[0077] For example, in the "gesture interaction" task, in response to the user's trigger event on the first gesture selection component, the first gesture animation corresponding to the first gesture selection component can be obtained and displayed as a second interface element in the user interaction area.
[0078] For example, in a "dictation" task, in response to an event triggering the input touch trajectory, a trajectory image corresponding to the touch trajectory is determined as a second interface element and displayed in the user interaction area. This means the user's drawn trajectory is displayed in real-time in the user interaction area. This ensures that users receive direct visual feedback when performing handwriting, drawing, or other trajectory inputs, effectively reducing the delay between operation and perception, thereby enhancing the naturalness and accuracy of the interaction and further improving the user experience.
[0079] In practical applications, the touch trajectory can be smoothed first, and then a trajectory image can be generated based on the smoothed touch trajectory and displayed in the user interaction area. This method can improve the display effect of the touch trajectory, thereby improving the user experience.
[0080] In addition, to automatically and orderly terminate the interactive task and provide the user with clear feedback on the outcome, this application also provides the following preferred solution. In one embodiment, after displaying the response content determined based on the event information of the input event in the response area, the interactive task ends in response to the response content satisfying a preset termination condition, and the second task interaction result determined based on the response content is displayed in the response area.
[0081] In the above implementation, a step is added to automatically determine the end of the interactive task based on the response content and display the final second task interaction result. The end condition can be that the user's total score or the agent's total score reaches the target score (such as the first person to reach 11 points in the "ping-pong interaction" task wins), the number of rounds reaches the upper limit (such as reaching the total number of questions in the "question dictation" task), or a specific judgment result is generated (such as "incorrect answer").
[0082] The second task interaction result can be determined based on multiple response contents generated throughout the entire interaction task. The second task interaction result can also be obtained based on the real-time updated task status information, which is equivalent to a comprehensive summary of the interaction task, such as interaction task success, interaction task failure, total score panel, evaluation content, explanation content, etc.
[0083] This process aims to provide a complete interactive chain for interactive tasks. By continuously monitoring whether the response content in each round reaches the preset termination condition, it achieves automated management of the task lifecycle and enhances the sense of ritual and integrity of interactive tasks, thereby improving the user experience.
[0084] Based on this, a complete example is used here to illustrate the information interaction method. Specifically, an initial interactive interface is displayed. In response to a trigger event of the launch component corresponding to the interactive task in the initial interactive interface, or in response to a natural language command submitted by the user containing a launch instruction for the interactive task, a page template for the task interaction page is obtained. Based on the task type of the interactive task, component resources corresponding to the task type are determined. Based on the page template and component resources, the task interaction interface is displayed. The task interaction interface includes a user interaction area and a response area. The user interaction area includes at least one input component corresponding to the task type of the interactive task. The input component is used to trigger an input event corresponding to the task type.
[0085] In response to a triggered input event, the system generates response content based on the event information of the input event and the task status information of the interactive task, using the first main model or a preset rule base. It then determines the first interface element corresponding to the response content and displays it in the response area. Simultaneously, it determines the second interface element corresponding to the event information of the input event and displays it in the user interaction area.
[0086] When the response content meets the preset termination condition, the interactive task ends, and the second task interaction result determined based on the response content is displayed in the response area.
[0087] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0088] The foregoing has described specific embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0089] According to another embodiment, an information interaction device is provided. Figure 7 A schematic block diagram of an information interaction device according to one embodiment is shown, the information interaction device being disposed in... Figure 1 The server side in the illustrated architecture. For example... Figure 7 As shown, the information interaction device 700 includes a first display unit 701 and a second display unit 702. The main functions of each component are as follows: The first display unit 701 is configured to display the task interaction interface of the interactive task. The task interaction interface includes a user interaction area and a response area. The user interaction area includes at least one input component corresponding to the task type of the interactive task. The input component is used to trigger an input event corresponding to the task type.
[0090] The second display unit 702 is configured to display response content determined based on event information from the input event in the response area in response to a triggered input event.
[0091] As one possible implementation method, the second display unit 702, when determining the response content based on the event information of the input event, can be specifically configured to generate the response content based on the event information of the input event and the task status information of the interactive task through the first large model or a preset rule base.
[0092] As one possible implementation method, the second display unit 702, when generating response content based on the event information of the input event and the task status information of the interactive task through the first large model or rule base, can be specifically configured as follows: in response to the input event being a preset task interaction event, response content is generated based on the task interaction event and task status information, and based on the rule base; in response to the input event being an event of submitting a natural language command, response content is generated through the first large model based on the natural language command and task status information; in response to the input event including both task interaction events and events of submitting natural language commands, response content is generated through the first large model based on the task interaction event, natural language command, and task status information.
[0093] One possible approach is to respond with the result of the first task interaction determined by a rule base.
[0094] The second display unit 702, after displaying the first task interaction result in the response area, can also be configured to: call the second major model to generate supplementary description content of the first task interaction result based on the first task interaction result, task status information and event information of the input event.
[0095] One possible approach is to respond with the result of the first task interaction determined by a rule base.
[0096] After determining the response content, the second display unit 702 can also be configured to: update the task status information based on the event information of the input event and the interaction result of the first task, and determine the task status change information corresponding to the input event; and display the changed content determined based on the task status change information in the user interaction area.
[0097] As one possible implementation method, when the second display unit 702 displays response content determined based on event information of the input event in the response area, it can be specifically configured to: display a first interface element corresponding to the response content in the response area, wherein the first interface element is used to represent the response content.
[0098] One possible approach is to use an input event as an event that inputs a touch trajectory.
[0099] The second display unit 702 can also be configured to: in response to an event that triggers the input touch trajectory, determine the trajectory image corresponding to the touch trajectory; and display the trajectory image in the user interaction area.
[0100] As one possible implementation method, the second display unit 702, when determining the trajectory image corresponding to the touch trajectory in response to an event that triggers the input touch trajectory, can be specifically configured to: smooth the touch trajectory in response to an event that triggers the input touch trajectory to obtain a smoothed touch trajectory; and generate a trajectory image based on the smoothed touch trajectory.
[0101] As one possible implementation method, the first display unit 701, when displaying the task interaction interface of the interactive task, can be specifically configured to: display the initial interaction interface in response to the trigger event of the launch component corresponding to the interactive task in the initial interaction interface; or, display the initial interaction interface in response to the natural language command submitted by the user containing the launch instruction for the interactive task.
[0102] As one possible implementation method, the first display unit 701, when displaying the task interaction interface of the interactive task, can be specifically configured to: obtain the page template of the task interaction page; determine the component resources corresponding to the task type based on the task type of the interactive task; and display the task interaction interface based on the page template and component resources.
[0103] As one possible implementation method, the second display unit 702 can also be configured to: end the interactive task in response to the response content meeting the preset termination condition, and display the second task interaction result of the interactive task determined based on the response content in the response area.
[0104] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0105] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0106] like Figure 8As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0107] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0108] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as information interaction methods. For example, in some embodiments, the information interaction method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the information interaction method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform information interaction methods by any other suitable means (e.g., by means of firmware).
[0109] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0110] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0111] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0112] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0113] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0114] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0115] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0116] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An information exchange method, comprising: The interactive interface for displaying interactive tasks includes a user interaction area and a response area. The user interaction area includes at least one input component corresponding to the task type of the interactive task. The input component is used to trigger an input event corresponding to the task type. In response to the input event being triggered, response content determined based on the event information of the input event is displayed in the response area.
2. The method according to claim 1, wherein the response content determined based on the event information of the input event is determined in the following manner: The response content is generated based on the event information of the input event and the task status information of the interactive task, using the first major model or a preset rule base.
3. The method according to claim 2, wherein generating the response content based on the event information of the input event and the task status information of the interactive task using a first large model or rule base includes: In response to the input event being a preset task interaction event, the response content is generated based on the task interaction event and the task status information, and based on the rule base. In response to the input event being a submission of a natural language command, the first large model generates the response content based on the natural language command and the task status information; In response to the input events, including the task interaction event and the event of submitting a natural language command, the first large model generates the response content based on the task interaction event, the natural language command, and the task status information.
4. The method according to claim 2, wherein, The response content is a first task interaction result determined based on the rule base. After displaying the first task interaction result in the response area, the method further includes: The second major model is invoked to generate supplementary descriptions of the first task interaction results based on the first task interaction results, the task status information, and the event information of the input events.
5. The method according to claim 2, wherein, The response content is the first task interaction result determined based on the rule base. After determining the response content, the method further includes: Based on the event information of the input event and the interaction result of the first task, update the task status information and determine the task status change information corresponding to the input event; The user interaction area displays the changed content determined based on the task status change information.
6. The method according to claim 1, wherein, The step of displaying the response content determined based on the event information of the input event in the response area includes: The response area displays a first interface element corresponding to the response content, and the first interface element is used to represent the response content.
7. The method according to claim 1, wherein, The input event is an event that inputs a touch trajectory, and the method further includes: In response to an event that triggers the input touch trajectory, a trajectory image corresponding to the touch trajectory is determined; The trajectory image is displayed in the user interaction area.
8. The method according to claim 7, wherein, The step of determining a trajectory image corresponding to the touch trajectory in response to an event that triggers the input touch trajectory includes: In response to an event that triggers the input touch trajectory, the touch trajectory is smoothed to obtain a smoothed touch trajectory; A trajectory image is generated based on the smoothed touch trajectory.
9. The method according to claim 1, wherein, The task interaction interface for displaying interactive tasks includes: Display the initial interactive interface, and in response to a trigger event on the launch component corresponding to the interactive task in the initial interactive interface, display the task interactive interface of the interactive task; or, The initial interactive interface is displayed, and in response to a natural language command submitted by the user that includes a start instruction for the interactive task, the task interactive interface of the interactive task is displayed.
10. The method according to any one of claims 1-9, wherein, The task interaction interface for displaying interactive tasks includes: Obtain the page template for the task interaction page; Based on the task type of the interactive task, determine the component resources corresponding to the task type; The task interaction interface is displayed based on the page template and the component resources.
11. The method according to claim 1, wherein after the response content determined based on the event information of the input event is displayed in the response area, the method further comprises: In response to the response content satisfying the preset termination condition, the interactive task ends, and the second task interaction result of the interactive task determined based on the response content is displayed in the response area.
12. An information interaction device, comprising: The first display unit is configured to display the task interaction interface of the interactive task. The task interaction interface includes a user interaction area and a response area. The user interaction area includes at least one input component corresponding to the task type of the interactive task. The input component is used to trigger an input event corresponding to the task type. The second display unit is configured to display response content determined based on event information of the input event in the response area in response to the triggering of the input event.
13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-11.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-11.
15. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-11.