Interaction method and device, equipment and storage medium
By dynamically presenting controls related to agent interaction in electronic devices and using machine learning models to determine the control types, the problem of low interaction efficiency in existing technologies is solved, and efficient and accurate agent interaction is achieved.
Patent Information
- Application Number
- CN202511028349.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-11-11
AI Technical Summary
In existing technologies, the interaction between electronic devices and intelligent agents lacks effective control presentation and response mechanisms, resulting in low interaction efficiency.
By dynamically presenting controls related to the interaction with the intelligent agent during the selection operation in the content display area, and using machine learning models to determine the type and function of the controls, the system responds to the user's selection to generate corresponding content.
It improves the efficiency and richness of interaction with intelligent agents, enhances the ability to interact with content, and ensures the accuracy and convenience of the interaction process.
Smart Images

Figure CN120928973A_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to interactive methods, apparatuses, devices, and computer-readable storage media. Background Technology
[0002] With the development of computer technology, various forms of electronic devices have greatly enriched people's daily lives. For example, people can use electronic devices to perform various interactions, such as using intelligent agents to interact with users. Summary of the Invention
[0003] In a first aspect of this disclosure, an interaction method is provided. The method includes: presenting an interaction interface with an agent, the interaction interface including a content display area and an input component; presenting a first set of controls for the first content in the input component based on a selection of the first content in the content display area; and, in response to a selection of the first control in the first set of controls, providing first response content generated by the agent based on the first content.
[0004] In a second aspect of this disclosure, an apparatus for interaction is provided. The apparatus includes: a first presentation module configured to present an interactive interface with an agent, the interactive interface including a content display area and an input component; a second presentation module configured to present a first set of controls for the first content in the input component based on a selection of the first content in the content display area; and a providing module configured to provide first response content generated by the agent based on the first content in response to a selection of the first control in the first set of controls.
[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method according to a first aspect of this disclosure.
[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0010] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;
[0011] Figures 2A-2D Example interface diagrams are shown according to some embodiments of this disclosure;
[0012] Figure 3 A flowchart illustrating an interaction process according to some embodiments of this disclosure is shown;
[0013] Figure 4 A schematic structural block diagram of an interactive device according to certain embodiments of the present disclosure is shown;
[0014] Figure 5 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation
[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0016] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0017] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0018] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0019] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0020] Embodiments of this disclosure propose an interaction scheme. According to this scheme, an interactive interface for an intelligent agent can be presented, the interface including a content display area and an input component. Further, based on the selection of first content in the content display area, a first set of controls for the first content can be presented in the input component. Further, in response to the selection of the first control in the first set of controls, first response content generated by the intelligent agent based on the first content can be provided.
[0021] Based on this approach, embodiments of this disclosure can present a first set of controls for the first content in the input component after the first content is selected, effectively enriching the interactive capabilities associated with the first content and improving the efficiency of the interaction.
[0022] Example Environment
[0023] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1As shown, example environment 100 may include electronic device 110.
[0024] In this example environment 100, electronic device 110 can run an application 120 that supports user interface interaction. Application 120 can be any suitable type of application for user interface interaction, and examples may include, but are not limited to, video applications, game applications, social applications, or other suitable applications.
[0025] exist Figure 1 In environment 100, if application 120 is active, application 120 can provide user 140 with an interactive interface 150.
[0026] In some embodiments, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal with a display device, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of interface for the target user (such as "wearable" circuitry).
[0027] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for applications 120 that support user interface interaction in electronic devices 110.
[0028] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections; the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, server 130 and electronic device 110 can achieve signaling interaction through the communication connection between them.
[0029] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0030] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.
[0031] Example process
[0032] Figures 2A to 2D Example interfaces 200A to 200D according to some embodiments of the present disclosure are shown. Interfaces 200A to 200D may, for example, be derived from... Figure 1 The electronic device 110 shown is provided.
[0033] In some embodiments, such as Figure 2A As shown, the electronic device 110 can present an interactive interface 200A. As an example, the interactive interface 200A can be an interactive interface with an intelligent agent, used for interacting with the intelligent agent. This interactive interface can be any suitable interactive interface, such as a conversational interface between a user and an intelligent agent, upon which the user can interact with the intelligent agent in any appropriate way, such as dialogue interaction, generative interaction (e.g., using the intelligent agent to generate certain media content), etc.
[0034] In some embodiments, an agent may include an appropriate application built upon a generative model (e.g., a language model). In some scenarios, an agent is also referred to as an Agent. In some embodiments, an agent may utilize one or more models to drive its interaction with the user. As an example, an agent may utilize a generative model to generate a response to a user-input processing request. This processing request may be entered by the user themselves or quickly entered based on interactive controls in an interactive interface, which will not be elaborated further here.
[0035] by Figure 2AAs an example, the interactive interface 200A may include a content display area 201 and an input component 202. In some embodiments, the content display area 201 may be configured to display any appropriate content presented in the interactive interface, which may include, but is not limited to, user-inputted content and agent-generated content. As an example, taking the interactive interface 200A as a conversational interface between a user and an agent, the content display area 201 may be configured to present conversational content, which may include user-inputted conversational content, agent-generated conversational content, and so on. In some embodiments, the content display area 201 may also be configured to present any appropriate type of content, such as text content, voice content, video content, image content, and so on.
[0036] by Figure 2A Taking the conversation interface as an example, the electronic device 110 can display the conversation content 203-1 generated by the intelligent agent and the conversation content 203-2 input by the user in the content display area 201.
[0037] In some embodiments, the input component 202 can be configured to receive any appropriate type of input content from the user, such as text content, voice content, image content, video content, etc. The input component 202 may include multiple input controls to support these various types of content input.
[0038] by Figure 2A As an example, input component 202 may include a text input control 210, a voice input control 211, a set of tool entry points (tool entry point 212-1, tool entry point 212-2, tool entry point 212-3, tool entry point 212-4), an image input control, etc. Text input control 210 can be configured to receive text content. Voice input control 211 can be configured to receive voice content. Image input control 213 can be configured to receive image or video content. This set of tool entry points can be a set of tool entry points used to assist users in interacting with the intelligent agent. Users can trigger the corresponding tool entry point through quick triggering methods such as clicking and swiping, so that users can further interact with the intelligent agent based on the tool corresponding to this tool entry point. In some embodiments, this set of tool entry points can be tool entry points independent of the content display area 201 in the interactive interface 200A; that is, this set of tool entry points is not provided for any one or more contents displayed in the content display area 201, but is a general tool entry point. As an example, this tool entry point may include photo-based Q&A, virtual avatar, etc. Virtual avatars can be configured to create media content associated with a user's virtual avatar through this tool entry point. Photo-based Q&A can be configured to upload images or videos through this tool entry point to provide analytical content tailored to those images or videos.
[0039] In some embodiments, such as Figure 2A As shown, the input component 202 can currently correspond to the first mode. The input component 202 in the first mode is not specifically associated with any specific content in the content display area 201, and it can be a general input component.
[0040] by Figure 2B As an example, the electronic device 110 can receive a selection of content 220 (first content) in the content display area 201. In some embodiments, content 220 may include content associated with multiple modalities, such as text content, image content, or a combination of content from different modalities (e.g., text content and image content), etc., which will not be elaborated here. In some embodiments, the first content may include, but is not limited to, at least a portion of content input by the user, at least a portion of content generated by the agent, at least a portion of content input by the user, and at least a portion of content generated by the agent. For example, content 220 may be a portion of session content 203-1 generated by the agent.
[0041] As an example, electronic device 110 can respond to a preset gesture operation on content 220 and receive a selection of content 220. The preset gesture operation can be any appropriate operation, such as clicking, swiping, long pressing, etc. For example, a user can select a portion of text (the first content) in the interactive interface by dragging, so that electronic device 110 can receive the selection of this portion of text.
[0042] To facilitate highlighting, the electronic device 110 can set the selected first content to a predetermined highlighting style, such as changing the background color of the area corresponding to the first content to highlight the difference between the selected first content and other unselected content.
[0043] Since the user selects content 220 in the content display area 201, indicating a preference for interaction with the agent related to content 220, the electronic device 110 can stop displaying a set of tool entry points associated with the agent in response to receiving the selection of content 220 in the content display area 201. In some scenarios, this set of tool entry points may include one or more tool entry points in the action bar. Figure 2B As an example, compared to Figure 2A In response to the selection received from the content 220, the electronic device 110 may stop displaying the set of tool entrances, such as tool entrance 212-1, tool entrance 212-2, tool entrance 212-3, and tool entrance 212-4.
[0044] In some embodiments, the electronic device 110 may switch the input component 202 from a first mode to a second mode in response to the selection of content 220, the second mode being associated with the selected content 220. That is, the input component 202 in the second mode is specifically associated with content 220, and subsequent interactions based on this input component 202 can be interactions specifically targeted at content 220. For example, all response content generated by the agent is generated based on content 220.
[0045] For ease of display, the electronic device 110 may present at least a portion of the selected content 220 in the input component. In some embodiments, the content 220 may be associated with any appropriate location on the input component 202. For example, as Figure 2B As shown, content 220 can be associated and displayed in the top row of input component 202.
[0046] As one example, electronic device 110 may display a portion of content 220 in the interactive interface in response to content 220 being larger than a threshold. As another example, electronic device 110 may display the entire content 220 in the interactive interface in response to content 220 being smaller than or equal to a threshold. This threshold can be any suitable value, which may be associated with an upper limit on the size of content allowed to be displayed in a pre-configured input component 202.
[0047] To facilitate increased efficiency in interactions associated with the primary content, Figure 2B As shown, the electronic device 110 can present a first set of controls (e.g., controls 230-1, 230-2, 230-3, and 230-4) in the input component 202 based on the selection of content 220 in the content display area 201. In some embodiments, the first set of controls can be configured so that the user and the agent can interact with the content 220 based on these controls, that is, the interaction based on the first set of controls can focus on the content information included in the content 220.
[0048] In some embodiments, when different content is selected, the corresponding control presented in the input component 202 may be the same as or different from the control for that selected content.
[0049] The following explains how to determine the first set of controls for the first content.
[0050] In some embodiments, the first set of controls for the first content can be determined from a plurality of candidate controls based on the first content. The plurality of candidate controls can be all candidate controls that are allowed to be presented in the input component 202 in a pre-set manner, and different candidate controls can correspond to different interactive functions.
[0051] In some embodiments, the first set of controls may be determined based on the association information of the first content. The association information includes, but is not limited to, at least one of the following: the language corresponding to the first content, the content type corresponding to the first content, and the semantic information of the first content.
[0052] In some embodiments, the language may include, but is not limited to, Chinese, English, etc. The content type may include, but is not limited to, images, text, video, audio, etc. Semantic information may include at least one of the following: contextual understanding of the first content, the linguistic context, included entities, etc.
[0053] For example, if the first content is in English, then the first set of controls can include translation controls to support translating the first content into Chinese.
[0054] To improve the relevance between the first set of controls and the first content, and to better provide interactive functions associated with the first content, the electronic device 110 can provide the model with descriptions of multiple candidate controls and the first content, so as to determine the first set of controls related to the first content from the multiple candidate controls. In some embodiments, the model can be any suitable machine learning model, such as a language model, or a multimodal model for processing audio and / or video content. The descriptions of the candidate controls can be functional descriptions, application scenario descriptions, etc., which will not be elaborated here.
[0055] Furthermore, the electronic device can respond to the selection of the first control in the first set of controls by providing first response content generated by the agent based on the first content. The first response content can be any appropriate type of content, such as text, images, videos, voice, etc., which will not be elaborated here.
[0056] As an example, a user can click, swipe, or perform other actions on a first control, allowing the electronic device to receive a selection of that control. As another example, an intelligent agent can utilize a model to generate a first response based on the first content. The model can be any suitable machine learning model, such as a language model, or a multimodal model for processing audio and / or video content.
[0057] As an example, control 230-1 can be a summary control. Electronic device 110 can respond to the selection (e.g., click) of the summary control by providing the agent with summary content (first response content) generated based on this content 220.
[0058] As an example, control 230-2 can be an expand control. Electronic device 110 can, in response to the selection of the expand control, provide expanded content (first response content) generated by the agent based on this content 220.
[0059] As an example, control 230-3 can be a depth-of-field control. In response to the selection of the depth-of-field control, electronic device 110 can provide further, more in-depth content (first response content) generated by the agent based on this content 220. This further content can be more detailed and specialized, and its generation may involve more references than content 220, resulting in more detailed and accurate content.
[0060] As an example, control 230-4 can be a translation control. Electronic device 110 can respond to the selection of the translation control by providing translated content generated by the agent based on content 220; for example, if content 220 is in English, then the translated content can be the corresponding Chinese content.
[0061] To enhance the richness of the interaction, in some embodiments, the electronic device 110 may provide second response content in response to receiving input content via the input component 202 in a second mode. This second response content is generated by the agent based on the input content and the first content. The input content can be any suitable content, such as text, image, or voice content. This input content can be content entered based on any suitable input control in the input component 202 (such as text input control 210, voice input control 220, image input control 230, etc.).
[0062] As an example, a user can input text content based on the text input control 210, at which point the electronic device 110 can provide a second response content generated for this text content and content 220.
[0063] by Figure 2C As an example, electronic device 110 can respond to receiving a trigger (e.g., click) of text input control 210 by presenting an electronic keyboard in interactive interface 200C, allowing the user to input text content based on the electronic keyboard. It should be noted that the first set of controls corresponding to content 220 are retained and displayed in input component 202 during the input process (e.g., when the electronic keyboard is presented or when content is input based on the electronic keyboard).
[0064] In addition, if the input content is input based on the voice input control 211 or the image input control 213, then during the process of inputting content based on the voice input control 211 or the image input control 213, the first set of controls for the content 220 is also retained and displayed in the input component 202, which will not be elaborated here.
[0065] by Figure 2D As an example, electronic device 110 can receive a selection of content 221 (second content) in content display area 201. In some embodiments, content 221 may include content associated with multiple modalities, such as text content, image content, or a combination of content from different modalities (e.g., text content and image content), etc., which will not be elaborated here. In some embodiments, the second content may include, but is not limited to, at least a portion of content input by the user, at least a portion of content generated by the agent, or at least a portion of both the user input and the agent-generated content. For example, content 221 may be a portion of the user-input session content 203-2.
[0066] As an example, electronic device 110 can receive a selection of the second content in response to a preset gesture operation. The preset gesture operation can be any appropriate operation, such as clicking, swiping, long pressing, etc. For example, a user can select a portion of text (the second content) in the interactive interface by dragging, so that electronic device 110 can receive the selection of this portion of text.
[0067] It should be noted that when the second content is selected, the first content can be deselected. Alternatively, if the first and second content overlap by at least some part, then when the second content is selected, the non-overlapping portions of the first and second content can be deselected.
[0068] To facilitate highlighting, the electronic device 110 can set the selected second content to a predetermined highlighting style, such as changing the background color of the area corresponding to the second content, or changing the background color of the area corresponding to the second content to highlight the difference between the selected second content and other unselected content.
[0069] Since when the user selects the second content in the content display area 201, it indicates that the user prefers to interact with the agent in relation to the second content rather than in relation to the first content, the electronic device 110 can update the first set of controls presented in the input component to a second set of controls for the second content based on the selection of the second content in the content display area 201. The second set of controls is different from the first set of controls.
[0070] In some embodiments, the first group of controls and the second group of controls may include the same controls or different controls. Figure 2C and Figure 2D As an example, the electronic device 110 may update the first set of controls (including controls 230-1, 230-2, 230-3 and 230-4) presented in the input component 202 for the content 220 to the second set of controls (including controls 250-1, 250-2 and 250-3) for the content 221 based on the selection of the content 221 in the content display area 201.
[0071] Furthermore, in response to the selection of a control in the second set of controls, the electronic device can provide a third response content generated by the agent based on the second content. The third response content can be any suitable type of content, such as text, images, video, voice, etc., which will not be elaborated upon here. As an example, a user can click, swipe, or perform other actions on a control so that the electronic device can receive a selection of that control. As an example, the agent can utilize a model to generate the third response content based on the second content. The model can be any suitable machine learning model, such as a language model, or a multimodal model for processing audio and / or video content.
[0072] As an example, control 250-1 can be an expandable control. Electronic device 110 can, in response to the selection of the expandable control, provide expanded content generated by the agent based on this content 221.
[0073] As an example, control 250-2 can be a depth-of-field control. In response to the selection of depth-of-field control 250-2, electronic device 110 can provide further, more in-depth content generated by the agent based on this content 221. This further content can be more detailed and specialized, potentially referencing more sources than content 221, resulting in more detailed and accurate content.
[0074] As an example, control 250-3 can be a translation control. Electronic device 110 can respond to the selection of the translation control by providing translated content generated by the agent based on content 221; for example, if content 221 is Chinese, then the translated content can be the corresponding English content.
[0075] Back Figure 2B The electronic device 110 can stop displaying the first set of controls in the input component 202 in response to the content 220 being deselected; that is, the electronic device 110 can display as follows in response to the content 220 being deselected. Figure 2A The interactive interface 200A is shown here. At this time, the input component 202 switches from the second mode to the first mode. The first mode is a mode that is not specifically associated with the content 220.
[0076] As an example, electronic device 110 can cancel the selection of the first content in response to the selection of the cancel control in the input component. Figure 2B As an example, electronic device 110 can cancel the selection of content 220 in response to receiving a selection of cancel control 288.
[0077] As another example, the electronic device 110 can deselect the first content by performing a predetermined operation on the first content in the content display area 201. The predetermined operation can also be any suitable gesture operation, such as a click operation. Figure 2B As an example, electronic device 110 can deselect content 220 in response to receiving a click operation on any blank space in the content display area 201.
[0078] It should be noted that, in order to improve interaction efficiency, the electronic device can respond to the deselection of content 220 by restoring the presentation of a set of tool entry points associated with the agent (such as tool entry point 212-1, tool entry point 212-2, tool entry point 212-3, and tool entry point 212-4).
[0079] Based on this approach, embodiments of this disclosure can present a first set of controls for the first content in the input component after the first content is selected, effectively enriching the interactive capabilities associated with the first content and improving the efficiency of the interaction.
[0080] Example process
[0081] Figure 3 A flowchart of an interaction process 300 according to some embodiments of the present disclosure is shown. Process 300 can be implemented at electronic device 110. Reference is made below. Figure 1 Describe the process 300.
[0082] In frame 310, electronic device 110 presents an interactive interface with the intelligent agent, which includes a content display area and input components.
[0083] In box 320, electronic device 110 presents a first set of controls for the first content in the input component based on the selection of the first content in the content display area.
[0084] In box 330, electronic device 110, in response to selection of a first control in a first set of controls, provides first response content generated by the agent based on first content.
[0085] Based on this approach, embodiments of this disclosure can present a first set of controls for the first content in the input component after the first content is selected, effectively enriching the interactive capabilities associated with the first content and improving the efficiency of the interaction.
[0086] In some embodiments, the first set of controls is determined from a plurality of candidate controls based on a first set of content.
[0087] In this way, the embodiments of this disclosure can dynamically determine a first group of controls from multiple candidate controls based on the first content to perform interactions associated with the first content, which can effectively improve the efficiency of interactions associated with the first content.
[0088] In some embodiments, the first set of controls is determined by the following process: providing the model with description information of a plurality of candidate controls and first content to determine a first set of controls related to the first content from the plurality of candidate controls.
[0089] In this way, embodiments of the present disclosure can utilize models to determine controls related to the first content, improving the efficiency and accuracy of control determination.
[0090] In some embodiments, the first set of controls is determined based on the association information of the first content, which includes at least one of the following: the language corresponding to the first content; the content type corresponding to the first content; and the semantic information of the first content.
[0091] In this way, embodiments of the present disclosure can utilize the metadata of the first content to enhance the relevance of the control to the first content, which helps to improve the accuracy of identifying the control associated with the first content.
[0092] In some embodiments, process 300 further includes: updating a first set of controls presented in the input component to a second set of controls for the second content, based on the selection of a second set of content in the content display area, wherein the second set of controls is different from the first set of controls.
[0093] In this way, embodiments of the present disclosure can dynamically update the controls in the input component to match the new selected content in response to changes in the selected content, effectively improving interaction efficiency.
[0094] In some embodiments, process 300 further includes: in response to the selection of first content, switching the input component from a first mode to a second mode, the second mode being associated with the selected first content.
[0095] In this way, embodiments of the present disclosure can quickly update the mode of the input component in response to the selection of first content in order to adapt to the interaction associated with the selected first content, thereby improving the efficiency of the interaction.
[0096] In some embodiments, process 300 further includes: presenting at least a portion of the selected first content in the input component.
[0097] In this way, embodiments of the present disclosure allow users to preview at least a portion of the selected first content, effectively improving the efficiency of information transmission.
[0098] In some embodiments, process 300 further includes: in response to receiving input content via an input component in a second mode, providing second response content, the second response content being generated by the agent based on the input content and the first content.
[0099] In this way, embodiments of the present disclosure can receive input content in a second mode, provide targeted interaction associated with the first content, and improve interaction efficiency.
[0100] In some embodiments, process 300 further includes: stopping the rendering of the first set of controls in the input component in response to the first content being deselected.
[0101] In this way, embodiments of the present disclosure can stop displaying the first group of controls when the first content is deselected, which further highlights the relevance of this group of controls to the first content and improves the efficiency of information transmission.
[0102] In some embodiments, process 300 further includes: canceling the selection of the first content in response to the selection of the cancel control in the input component; or canceling the selection of the first content via a predetermined operation on the first content in the content display area.
[0103] In this way, embodiments of the present disclosure provide a variety of quick ways to cancel the selection of the first content, thereby improving the efficiency of the interaction.
[0104] In some embodiments, the first content includes at least one of the following: at least a portion of content input by the user; at least a portion of content generated by the agent.
[0105] In this way, the embodiments of this disclosure enable interaction based on content from multiple sources, thereby enhancing the richness of the interaction.
[0106] In some embodiments, the first content includes content associated with multiple modalities.
[0107] In this way, embodiments of this disclosure can provide content-related interactions associated with multiple modalities, thereby enhancing the richness of the interaction.
[0108] In some embodiments, process 300 further includes: before receiving a selection of first content in the content display area, presenting a set of tool entries associated with the agent in connection with the input component; and process 300 further includes: stopping the presentation of the set of tool entries in response to the selection.
[0109] In this way, the embodiments of this disclosure can provide relevant tools before the user selects content, enhancing the convenience of operation and improving interaction efficiency.
[0110] In some embodiments, process 300 further includes: receiving a selection of the first content in response to a preset gesture operation on the first content.
[0111] In this way, embodiments of the present disclosure can receive the selection of the first content based on a preset gesture operation on the first content, effectively improving the selection efficiency of the first content.
[0112] Example devices and equipment
[0113] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 4 A schematic structural block diagram of an interactive device 400 according to certain embodiments of the present disclosure is shown. Device 400 may be implemented as or included in the electronic device 110 discussed above. Various modules / components in device 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0114] like Figure 4 As shown, the device 400 includes a first presentation module 410 configured to present an interactive interface with an agent, the interactive interface including a content display area and an input component; a second presentation module 420 configured to present a first set of controls for the first content in the input component based on the selection of the first content in the content display area; and a providing module 430 configured to provide first response content generated by the agent based on the first content in response to the selection of the first control in the first set of controls.
[0115] In some embodiments, the first set of controls is determined from a plurality of candidate controls based on a first set of content.
[0116] In some embodiments, the first set of controls is determined by the following process: providing the model with description information of a plurality of candidate controls and first content to determine a first set of controls related to the first content from the plurality of candidate controls.
[0117] In some embodiments, the first set of controls is determined based on the association information of the first content, which includes at least one of the following: the language corresponding to the first content; the content type corresponding to the first content; and the semantic information of the first content.
[0118] In some embodiments, the device 400 further includes an update module configured to update a first set of controls presented in the input component to a second set of controls for the second content, the second set of controls being different from the first set of controls, based on the selection of second content in the content display area.
[0119] In some embodiments, the device 400 further includes a switching module configured to switch an input component from a first mode to a second mode in response to the selection of first content, the second mode being associated with the selected first content.
[0120] In some embodiments, the apparatus 400 further includes a third presentation module configured to present at least a portion of the selected first content in the input component.
[0121] In some embodiments, the device 400 further includes a content providing module configured to provide second response content in response to receiving input content via an input component in a second mode, the second response content being generated by an agent based on the input content and the first content.
[0122] In some embodiments, the device 400 further includes a first stop module configured to stop presenting a first set of controls in an input component in response to the first content being deselected.
[0123] In some embodiments, the device 400 further includes a cancellation module configured to: cancel the selection of the first content in response to the selection of a cancellation control in the input component; or cancel the selection of the first content via a predetermined operation on the first content in the content display area.
[0124] In some embodiments, the first content includes at least one of the following: at least a portion of content input by the user; at least a portion of content generated by the agent.
[0125] In some embodiments, the first content includes content associated with multiple modalities.
[0126] In some embodiments, the device 400 further includes a fourth presentation module configured to: present a set of tool entries associated with the agent in connection with the input component before receiving a selection of first content in the content display area; and the device 400 further includes a second stop module configured to: stop presenting the set of tool entries in response to the selection.
[0127] In some embodiments, the device 400 further includes a receiving module configured to receive a selection of the first content in response to a preset gesture operation on the first content.
[0128] The units included in device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units may be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 400 may be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.
[0129] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 5 The electronic device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown can be used to achieve Figure 1 The electronic device 110 shown.
[0130] like Figure 5 As shown, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors 510 or processing units, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processor 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.
[0131] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within electronic device 500.
[0132] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0133] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0134] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0135] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0136] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0137] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0138] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0140] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. An interaction method, comprising: An interactive interface for the intelligent agent is presented, the interactive interface including a content display area and input components; Based on the selection of a first piece of content in the content display area, a first set of controls for the first piece of content is presented in the input component; as well as In response to the selection of a first control in the first group of controls, a first response content generated by the agent based on the first content is provided.
2. The method according to claim 1, wherein the first set of controls is determined from a plurality of candidate controls based on the first content.
3. The method of claim 2, wherein the first set of controls is determined by the following process: The model is provided with description information of the plurality of candidate controls and the first content, so as to determine the first group of controls related to the first content from the plurality of candidate controls.
4. The method according to claim 2, wherein the first set of controls is determined based on the association information of the first content, and the association information includes at least one of the following: The language corresponding to the first content; The content type corresponding to the first content; The semantic information of the first content.
5. The method according to claim 1, further comprising: Based on the selection of the second content in the content display area, the first set of controls presented in the input component is updated to a second set of controls for the second content, the second set of controls being different from the first set of controls.
6. The method according to claim 1, further comprising: In response to the selection of the first content, the input component is switched from a first mode to a second mode, the second mode being associated with the selected first content.
7. The method according to claim 6, further comprising: In the input component, at least a portion of the selected first content is presented.
8. The method according to claim 6, further comprising: In response to receiving input content via the input component in the second mode, a second response content is provided, which is generated by the agent based on the input content and the first content.
9. The method according to claim 1, further comprising: In response to the first content being deselected, the presentation of the first set of controls in the input component is stopped.
10. The method of claim 9, further comprising: In response to the selection of the cancel control in the input component, the selection of the first content is canceled; or The selection of the first content is canceled by a predetermined operation on the first content in the content display area.
11. The method of claim 1, wherein the first content comprises at least one of the following: At least part of the content entered by the user; At least a portion of the content generated by the agent.
12. The method of claim 1, wherein the first content includes content associated with multiple modalities.
13. The method of claim 1, wherein before receiving the selection of the first content in the content display area, the method further comprises: Associated with the input component, a set of tool entry points associated with the agent are presented; The method also includes: In response to the selection, stop displaying the set of tool entries.
14. The method according to claim 1, further comprising: In response to a preset gesture operation on the first content, the selection of the first content is received.
15. A device for interaction, comprising: The first presentation module is configured to present an interactive interface with the intelligent agent, the interactive interface including a content display area and input components; The second presentation module is configured to present a first set of controls for the first content in the input component based on the selection of the first content in the content display area; as well as A module is configured to provide first response content generated by the agent based on the first content in response to the selection of a first control in the first group of controls.
16. An electronic device comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 14 when executed by the at least one processor.
17. A computer-readable storage medium having stored thereon computer-executable instructions that, when executed by a processor, implement the method according to any one of claims 1 to 14.
18. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 14.