Data processing method and electronic equipment

By acquiring modal data on a personal computer through a target intelligent agent and dynamically adapting display parameters, the problem of high learning costs and insufficient operational consistency caused by differences in interaction rules between different applications is solved, resulting in a more efficient user experience and intelligent agent management.

CN121832808APending Publication Date: 2026-04-10LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

When switching between different applications on a personal computer, users need to frequently adapt to the interaction rules of each independent AI agent, resulting in high learning costs and insufficient operational continuity.

Method used

Modal data is acquired through the target intelligent agent, the target functional service is determined and the service is provided with matching display parameters, the functional display of each application is uniformly managed, and the interface presentation is dynamically adapted.

Benefits of technology

It improves the consistency and efficiency of user operations, reduces learning costs, and enhances the maintainability and personalized interactive experience of the intelligent agent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121832808A_ABST
    Figure CN121832808A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and electronic equipment. The method comprises the steps that in response to a target trigger event, target modal data in a current interaction scene is obtained, and the interaction scene is determined at least based on a current running application of the electronic equipment; determining to provide a target function service for target output content to the target user based on the target modal data, wherein the target output content at least comprises output content of a current running application of the electronic equipment; and controlling the target agent to provide the target function service with the corresponding target display parameter, different function services corresponding to different display parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more particularly to a data processing method and an electronic device. Background Technology

[0002] In current PC usage scenarios, different pages often correspond to different applications, and each application is equipped with its own independent Artificial Intelligence Agent (AI Agent) function—not only are the models different, but the interaction logic and operation methods also differ. This forces users to frequently switch between multiple applications when handling complex tasks, and each switch requires them to readjust to the new application's AI interaction rules (such as command input methods, function call paths, and feedback presentation formats). This greatly increases the learning cost for users in using intelligent agents, severely disrupts the continuity of operation, and affects the overall smoothness of the experience. Summary of the Invention

[0003] This application provides a data processing method, an electronic device, a computer-readable storage medium, and a computer program product.

[0004] The technical solution of this application embodiment is implemented as follows: In a first aspect, embodiments of this application provide a data processing method applied to a target intelligent agent in an electronic device, the method comprising: In response to a target trigger event, acquire target modal data in the current interaction scenario, which is determined at least based on the currently running application of the electronic device; Based on the target modal data, the target functional services to be provided to the target user are determined, targeting the output content, which includes at least the output content of the currently running application on the electronic device; and, The target intelligent agent provides target functional services with corresponding target display parameters, and different functional services have different display parameters.

[0005] Secondly, embodiments of this application provide an electronic device, including at least one processor and a target intelligent agent capable of running on the processor, wherein the target intelligent agent is capable of independently executing or invoking at least one artificial intelligence model to perform the following operations: In response to a target triggering event, target modal data in the current interaction scenario is acquired, wherein the interaction scenario is determined at least based on the currently running application of the electronic device; Based on the target modal data, a target functional service is determined to be provided to the target user for the target output content, wherein the target output content includes at least the output content of the currently running application of the electronic device; and, The target intelligent agent is controlled to provide the target functional services with corresponding target display parameters, and different display parameters are corresponding to different functional services.

[0006] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the data processing method provided in embodiments of this application when executed by a processor.

[0007] Fourthly, embodiments of this application provide a computer program product, including a computer program or computer-executable instructions, which, when executed by a processor, implement the data processing method provided in embodiments of this application. Attached Figure Description

[0008] Figure 1 This is a flowchart illustrating a data processing method provided in an embodiment of this application; Figure 2 This is a schematic diagram of a user-focused area display provided in an embodiment of this application. Figure 1 ; Figure 3 This is a schematic diagram of a user-focused area display provided in an embodiment of this application. Figure 2 ; Figure 4 This is a schematic diagram of a target function service in an interactive scenario provided by an embodiment of this application. Figure 1 ; Figure 5 This is a schematic diagram of the interaction result of an intelligent agent after inputting a voice command, provided in an embodiment of this application; Figure 6 This is a schematic diagram of a target function service in an interactive scenario provided by an embodiment of this application. Figure 2 ; Figure 7 This is an illustration of a function service based on emotion category provided in an embodiment of this application. Figure 1 ; Figure 8 This is an illustration of a function service based on emotion category provided in an embodiment of this application. Figure 2 ; Figure 9 This is a schematic diagram illustrating the functional services of an intelligent agent in a desktop application, as provided in an embodiment of this application. Figure 10 This is a schematic diagram of a functional service provided in this application embodiment, which associates an intelligent agent with multiple application windows; Figure 11 This is a schematic diagram of a virtual image of an intelligent agent provided in an embodiment of this application; Figure 12This is a schematic diagram of a smart agent interaction window provided in an embodiment of this application; Figure 13 This is a schematic diagram illustrating the change of display parameters of an intelligent agent according to an embodiment of this application; Figure 14 This is a schematic diagram showing the result of an intelligent agent generation provided in an embodiment of this application; Figure 15 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 2 ; Figure 16 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 3 ; Figure 17 This is a schematic diagram of the interaction logic between an intelligent agent and a user, provided in an embodiment of this application. Figure 18 This is a schematic diagram illustrating the execution steps of an intelligent agent according to an embodiment of this application; Figure 19 This is a schematic diagram of the hardware entity of an electronic device provided in an embodiment of this application. Detailed Implementation

[0009] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0010] To address the issue of high learning costs for users transitioning between different applications and resulting in insufficient operational consistency in related technologies, embodiments of this application provide a data processing method applied to a target intelligent agent in an electronic device. Figure 1 A flowchart illustrating the interaction of a data processing method provided in this application embodiment is shown below. Figure 1 As shown, this method can be implemented through steps S101 to S103: Step S101: In response to the target triggering event, acquire the target modal data in the current interaction scenario. The interaction scenario is determined at least based on the currently running application of the electronic device.

[0011] In the embodiments of this application, the electronic device can be any intelligent device with an operating interface and machine vision, such as a smartwatch, computer, tablet computer, mobile phone, etc. The electronic device can run a target intelligent agent that can call upon large models or run its own artificial intelligence model code. This target intelligent agent can acquire user state data when using the electronic device, such as text selection, text or image viewing, application switching, and user facial expressions. In this embodiment, a fluid intelligent agent is used as the intelligent agent form of the target intelligent agent. Specifically, when a user uses the electronic device, if a target trigger event occurs, the target intelligent agent can begin acquiring target modal data of the user in the current interaction scenario with the electronic device.

[0012] Among them, the target trigger event can characterize the change process of specific data in the electronic device. For example, it can be the boot process performed by the electronic device, writing a report, sending an email, generating an image, creating a PowerPoint presentation, or interacting with an AI assistant. Target modal data can characterize the data source obtained by the target agent. Each data source corresponds to a modality. For example, target modal data can be the user's emotion detection data, gaze detection data, surrounding environment detection data of the electronic device, current task data, historical behavior data, schedule data, etc., obtained by the target agent. Its form can be text (such as current task data, schedule data, etc.), image (such as emotion detection data, gaze detection data, etc.), or numeric strings with data types (such as surrounding environment detection data). The interaction scenario can characterize the recipient of data interaction between the user and the electronic device. The target agent can determine the specific interaction scenario type based on the application currently running on the electronic device. For example, when the user is editing text, the text editing application is the recipient of data interaction, and the text editing application becomes the interaction scenario. Furthermore, the interaction scenario can also include the application type of the recipient application. For example, if the currently running application is a commonly used work software (such as PowerPoint, Word editor, or Excel spreadsheet), the interaction scenario is a work scenario; if the currently running application is an entertainment application such as a video application, novel application, or music application, the interaction scenario type is an entertainment scenario; if the currently running application is a real-time multi-person communication application (such as video conferencing or voice conferencing), the interaction scenario type is a meeting scenario. Other interaction scenario types include shopping scenarios, reading scenarios, and image / video processing scenarios, among others.

[0013] For example, after the target trigger event of powering on occurs, the electronic device enters the desktop application. At this time, the current interaction scenario is determined to be the power-on scenario based on the desktop application. The target agent can obtain target modal data such as user history behavior data and schedule data generated by the user while using the electronic device after it is powered on. When the target trigger event of a search engine performing a search occurs, the electronic device enters the web application displaying the search results. At this time, the web application becomes the current interaction scenario. The target agent can obtain target modal data such as eye gaze detection data, emotion detection data, and semantic data of the search results themselves after the search results are displayed. When the target trigger event of writing a work report occurs, all applications currently launched and displayed by the user can become the current interaction scenario. The target agent can obtain target modal data such as text editing data, eye gaze detection data, and emotion detection data of the user while writing the work report.

[0014] Step S102: Based on the target modal data, determine the target function service to be provided to the target user for the target output content, where the target output content includes at least the output content of the currently running application of the electronic device.

[0015] In the embodiments of this application, the target output content may be the display content of the user's attention area, a frame of image or product content in a video or video conference that the user is interested in, or the entire content of the foreground application (including the desktop), or service content categorized and displayed for multiple windows, etc. It may include at least the output content of the currently running application of the electronic device, and may also include the output content of applications that have already been run. The target modal data may be at least one type of multimodal data, such as at least one of emotion detection data, surrounding environment detection data, gaze detection data, and current task data. Specifically, determining the target functional service based on the target modal data may include the following steps: 1) Determine the user's area of ​​interest based on the target modality data; In some embodiments, the user’s area of ​​interest can be automatically determined based on target modal data.

[0016] For example, such as Figure 2 As shown, Figure 2 This application provides a schematic diagram of the display of a user-focused area in an embodiment. Figure 1 . Figure 2The image shows a Word editing interaction scenario. In this scenario, the user's gaze is focused on the "DEF" area. At this time, the target modal data is gaze detection data. Based on the gaze detection data, the specific user attention area can be obtained from the display area of ​​the electronic device. Specifically, the target agent determines that it is currently performing the task of editing "2) YYYY" based on the user's gaze detection data. Therefore, the user attention area is the internal area of ​​box 21 in the figure.

[0017] like Figure 3 As shown, Figure 3 This application provides a schematic diagram of the display of a user-focused area in an embodiment. Figure 2 . Figure 3 The image shows a PowerPoint editing scenario. In this scenario, the user is editing a secondary text box titled "Inspiration Box." The target modal data can include both user gaze detection data and current task data. The target agent can determine the user's editing task based on the target modal data; therefore, the user's focus area is the internal area of ​​box 31 in the image.

[0018] In some embodiments, if the currently running application is a desktop application that appears when the electronic device is first turned on, the target agent can acquire target modal data such as the user's gaze detection data, and then determine the corresponding user attention area in the desktop application based on the gaze detection data.

[0019] 2) Determine the target output content based on the user's area of ​​interest; In some embodiments, the user's focus area typically contains prompts, images, data, or other content, and the content within this focus area can be identified as the target output content. The target output content can be a portion of the content displayed on the application interface, such as webpage text, a product showcased in a meeting, a term used in training, a product in an advertisement, a product in a video, or the content of a speaker's remarks in a meeting, etc.

[0020] For example, Figure 2 The content such as “2)YYYY”, “ABC”, and “DEF” in area 21, which is of interest to the user, are all target output content. Figure 3 The content such as "-AAA:XX" and "-CC:XXXXXX" in area 31 of the user's attention are all target output content.

[0021] 3) Determine the target functional services based on the target output content.

[0022] In some embodiments, after obtaining the target output content, the target agent can determine the target functional service based on the target output content itself or its type.

[0023] For example, Figure 2 and Figure 3 If the target output content is all text, then the target function service can be a text expansion service; furthermore, if Figure 3 If the target output is English text, the text expansion service can be an English text expansion service or a translation service. Alternatively, if the target output is an image, the target function service can be an image recognition function. Furthermore, the target function service may also include a text expansion service, used to generate descriptive text for entities identified in the image.

[0024] Step S103: Control the target agent to provide target function services with corresponding target display parameters. Different display parameters correspond to different function services.

[0025] In the embodiments of this application, target display parameters can be used to characterize the display features of the target intelligent agent, including but not limited to display state parameters (appearance parameters), display theme parameters, and display attribute parameters. Display state parameters can include combined state, split state, different shapes such as two-dimensional images, 3D images, various cartoon images, and forms matching the current application; display theme parameters can include color, brightness, display mode (color gradient), and hue; display attribute parameters can include size, position, and refresh rate.

[0026] In the embodiments of this application, the target intelligent agent precisely configures display parameters that match the business characteristics of each functional service, so that the functional services can be presented to the user in a better interface. Different functional services can adopt different sets of display parameters due to differences in their functional positioning, operation logic, and user interaction needs. For example, the data visualization function needs to configure parameters such as chart type, axis range, and color scheme, while the user permission management function needs to set parameters such as role classification, operation permissions, and interface layout.

[0027] Following the example above, Figure 2 and Figure 3 In this context, if the target function service is text expansion, its target display parameter can be a specified size of the user's area of ​​interest. Expanded text can be generated within the user's area of ​​interest. Alternatively, if the target function service is image recognition, image recognition results can be generated within a specified size display area around the image (e.g., the blank display area to the right of the image or the blank display area below the image), or at a specified location on the electronic device's display area (e.g., the right side of the monitor).

[0028] Based on the embodiments disclosed in this application, when a target trigger event occurs in an electronic device, the target intelligent agent can determine the target functional service that should be provided to the target user based on the target modal data in the current interaction scenario when the target user operates the electronic device. Finally, the target functional service is provided based on the target display parameters of the target functional service. This method not only achieves accurate matching between the target functional service and the current interaction scenario but also effectively improves the flexibility of the target intelligent agent. Since a unified target intelligent agent is used to manage various applications, adjustments to the target intelligent agent's function only need to be made once. Therefore, the above method increases the maintainability of the target intelligent agent to a certain extent, providing users with a more personalized and efficient interactive experience in real time. Furthermore, by dynamically adapting the target display parameters corresponding to the target functional service, the target intelligent agent can automatically optimize the interface presentation according to different functional scenarios, greatly reducing the user's learning cost and improving the user's operational experience and information acquisition efficiency during use.

[0029] In some embodiments, step S101 may include steps S111 to S113: Step S111: Obtain event information of the target-triggered event, and control the running state of the target intelligent agent based on the event information; In the embodiments of this application, the event information of the target triggering event may include event content and event type. For example, when a user powers on the device, the corresponding event type is a power-on event, and its event content is powering on and entering the desktop application; when a user opens a new application window (which may be the application window of a new application or a new process window of an already running application), the corresponding event type is an open new application window event, and its event content is the specific operation performed by the user in the new application window (such as text editing, focusing on an image, sending voice, etc.); when a user moves the focus of their gaze or input (such as the position of the mouse cursor or the position of the text input cursor), the corresponding event type is a focus transfer event, and its event content is moving to the input box of an application or the entire display area of ​​the electronic device.

[0030] In the embodiments of this application, after a target triggering event occurs in the electronic device, the target intelligent agent can parse the target triggering event to obtain event information. Then, the operating state of the target intelligent agent can be controlled based on the event information. The operating state of the target intelligent agent may include background state, foreground state, not started state, started state, idle state, active sensing state, split state, and merged state, etc.

[0031] In some embodiments, controlling the operating state of the target intelligent agent can switch the target intelligent agent from one operating state to another. For example, taking a user's operation on a work software application as an example: if the user opens the work software, the target intelligent agent can enter the running state from the unstarted state, or enter the active sensing state from the idle state; if the user edits content in the work software, the target intelligent agent can enter the foreground state from the background standby state; if the user changes the editing position, the target intelligent agent can enter the background standby state from the foreground state; if the user stops operating the work software for a period of time, the target intelligent agent can enter the background standby state from the foreground state; if the user chooses not to use the intelligent agent temporarily, the target intelligent agent can enter the idle state from the active sensing state; if the user operates multiple work software applications simultaneously, the target intelligent agent can enter the split state from the merged state. If the electronic device is powered on, the target intelligent agent can enter the running state from the unstarted state. If the user opens a new application window, the target intelligent agent can enter the split state from the merged state. If the user moves the focus of their gaze or input, the target intelligent agent can enter the background standby state from the foreground state; and when the movement ends, it returns to the foreground state from the background standby state.

[0032] Step S112: Identify the current interaction scenario of the electronic device based on event information and the currently running application of the electronic device.

[0033] In the embodiments of this application, the event information of the electronic device and the currently running application of the electronic device can be correlated. For example, if the event information is a power-on event, the currently running application of the electronic device can be a desktop application; if the currently running application is an entertainment application, the event information can be listening to music, watching videos, reading novels, etc. However, relying solely on event information often cannot accurately identify the interaction scenario. For example, educational videos and entertainment videos often exist in different applications. If only the event information "watching videos" is used, it is impossible to determine whether the current interaction scenario is an "entertainment interaction scenario" or a "education interaction scenario." Therefore, event information and the currently running application can be combined to jointly identify the current interaction scenario of the electronic device.

[0034] For example, if the event information is a power-on event and the electronic device is not running any applications, the current interaction scenario is "access to core services upon power-on." Through comprehensive analysis of the user's historical behavior data, schedule, and current scenario, the target intelligent agent can proactively identify and predict core needs, prioritize displaying the most critical content to the user, and present real-time progress of daily tasks, unread important messages, and usage statistics of frequently used tools. If the user opens a new video player application page, the target intelligent agent can first determine that the event information is the user opening a new application (i.e., the user's specific operation event on the application page), and then combine it with the application page to identify the specific interaction scenario of the user watching the video.

[0035] Furthermore, in multi-task management events, such as when writing a report, retrieving documents, searching data, or sending emails, the target intelligent agent can automatically identify highly relevant pages or applications to determine that the current interaction scenario is a multi-task management scenario. In the event of opening a video conference, if the electronic device is currently running video conferencing software, it can be identified as an emotion recognition scenario or a meeting recording scenario, etc.

[0036] Step S113: Use the target intelligent agent to call the target sensor or target interface to obtain target modal data that matches the current interaction scenario.

[0037] In the embodiments of this application, the target sensor may be a camera, microphone, touch electrode, time-of-flight ranging sensor or radar, etc.; the target interface may be an application programming interface (API) between the intelligent agent and each current application, to obtain the content displayed on the application interface or the task currently being performed by the electronic device, etc.

[0038] In the embodiments of this application, the target agent can invoke the aforementioned target sensor or target interface to collect target modal data that matches the current interaction scenario. Here, target modal data matching the interaction scenario means that the modal data collected by the sensor or interface is related to the interaction scenario. For example, in a text editing scenario, the target modal data can be, on the one hand, the text content that the user is currently typing, and on the other hand, the user's gaze detection data (used to determine the area of ​​text content the user is focusing on); in an entertainment interaction scenario, the target modal data can be, on the one hand, the media content being played by the entertainment application, and on the other hand, the user's gaze detection data.

[0039] Based on the above embodiments disclosed in this application, the operating state of the target intelligent agent can be controlled according to the event information of the target event, and the event information and the currently running application can be combined to identify the current interaction scenario, thereby improving the recognition accuracy of the current interaction scenario. Finally, the target sensor or target interface is called to obtain the target modality data that matches the current interaction scenario, which can improve the matching degree between the target modality data and the current interaction scenario, reduce the frequency of data that is irrelevant to the current interaction scenario in the target modality data, and improve the accuracy and simplicity of target modality data acquisition to a certain extent.

[0040] In some embodiments, step S102 may include step S121 or step S122: Step S121: Determine the target output content from the output content of the currently running application of the electronic device based on the target modal data. The target output content is at least a part of the output content. Determine the corresponding target function service based on the scene category of the interaction scene and the attribute information and / or content information of the target output content.

[0041] In the embodiments of this application, the output content of the currently running application of the electronic device may include text, images, sound, etc., specifically determined by the specific function of the currently running application. The target intelligent agent can determine the target output content from the output content based on the acquired target modal data. For example, if the current application is a video application, its output content is a video stream or video frames. The target intelligent agent can determine the product in the currently playing video from the video frames, or identify the content of the user-selected area from the current display interface, etc. The target output content can be a part of the output content or all of the output content. For example, in the above example, if the user-selected area is the entire display area, then the output content can be the target output content.

[0042] In the implementation of this application, after obtaining the target output content, the target output content can be parsed to obtain attribute information and / or content information. Attribute information may specifically include category attributes, such as whether it is an image, text, audio, or video; and display attributes, such as display size, display color, display position, and display quantity (e.g., number of windows). Content information may specifically include content validity information, semantic coherence, and semantic correlation between the content and the overall content. The parsing method can be selected according to the specific information to be obtained. For example, to obtain image size, the size information in the image's metadata can be read, or an edge detection method can be used; to obtain color information, a color gamut decomposition method can be employed.

[0043] In the embodiments of this application, after obtaining the attribute information and / or content information of the target output content, the attribute information and / or content information can be combined with the scenario category of the interaction scenario to jointly determine the target functional service corresponding to the target output content. The scenario category of the interaction scenario can include categories such as video conferencing scenarios, shopping scenarios, entertainment scenarios, work scenarios, and learning scenarios. Different interaction scenario categories and different attribute and content information can correspond to different target functional services. For example, if the interaction scenario is a work scenario and the content information contains semantic errors, the target functional service can be a semantic correction function; if the interaction scenario is a work scenario and the legality information in the content information is illegal, the target functional service can be a text anonymization service.

[0044] Step S122: Determine the target input content from the target input data of the target user input into the interaction area of ​​the target application in the electronic device based on the target modal data; determine the corresponding target function service based on the scene category of the interaction scene and the attribute information and / or content information of the target input content.

[0045] In the embodiments of this application, when the electronic device receives input data, the interaction area of ​​the target application can first be determined based on the target modality data. Then, the target input data can be determined from the input data within that interaction area. Finally, the target input content can be determined from the target input data based on the target modality data. Considering that the agent's active intervention in determining the interaction area may affect the user's normal input, triggering conditions for the agent's active intervention can be set. For example, the triggering conditions can support pen control / gesture and eye tracking dual positioning modes, actively capturing input boxes in various application scenarios (such as web forms, APP text boxes, document editing areas, etc.), and performing AI task interaction based on the located input boxes, reducing manual positioning operations. When the target agent is in a background standby state, the agent's colored interaction area dialog box is triggered by pen control and eye gaze. The agent's interaction area dialog box replaces all the original input boxes: the user input is normal input, but the interaction dialog box contains windows for dialogue with the target agent. The target agent can remain in a standby trigger state. How can we intelligently recommend the use of target agent functions without affecting user operation or causing interference? Our solution is substitution enhancement. This involves replacing parts of the data with equivalent / similar information generated by the target agent, while preserving the original core information, thus generating new, high-quality reconstructed information. If the user finds the information generated by the target agent better, they can choose it; if the user finds it worse, the original core information can be retained.

[0046] In the embodiments of this application, after obtaining the target content, the target function service corresponding to the target input content can be determined by combining the scene category of the interaction scene with the attribute information and / or content information of the target input content.

[0047] For example, such as Figure 4 As shown, Figure 4 This application provides an illustration of a target function service in an interactive scenario. Figure 1 . Figure 4In the diagram, the user selects area 41 using a stylus. Since the user selects the relevant area using the stylus, the aforementioned triggering conditions are met, triggering the user's voice input. Voice input then becomes the target function service. Icon 42 is the icon displayed when voice input is activated, and area 43 is the interactive display area where the target agent converts speech into text. After acquiring the user's voice data and other target modal data, the target agent can highlight the converted text data and other target input content in area 43. Furthermore, if the voice data is a command type, after conversion to text data, the target agent can also connect to a new target function service based on the voice command. For example, if the voice command is "expand the text content in the above box," the connected target function service is a text expansion service. In this case, an interactive area generated by the target agent can be overlaid on area 41 to display the expanded text content.

[0048] Figure 5 This demonstrates the interaction results of an intelligent agent after inputting a voice command. Figure 5 exist Figure 4 Based on this, the user sent a voice command to highlight important content. After receiving the modal data of this voice command, the target agent identified the keywords in region 41 and then highlighted these keywords as the target input content in red. Simultaneously, the highlighted text result, "Content highlighting complete. Let me know if there's anything else you need~", was generated within region 43. Afterward, the user returned the voice message "Okay" to the agent, which was displayed in region 51 within region 43.

[0049] For example, such as Figure 6 As shown, Figure 6 This application provides an illustration of a target function service in an interactive scenario. Figure 2 . Figure 6 The image displayed shows the search results interface of a search engine. Area 61 is the interactive area generated after the target agent actively intervenes, and area 62 is the search results area displayed on the webpage. At this time, the target agent generates the text content "If there is anything that needs adjustment or you want to learn more about, please feel free to tell me" in area 61, to prompt the target user to input commands and other target modal data to the target agent through voice or text.

[0050] Based on the above embodiments disclosed in this application, the target output content can be accurately obtained from the output content of the currently running application, and the target input content can also be accurately obtained from the target input data in the target application's interaction area. Then, based on the attribute information and / or content information of the target output content or target input content, combined with the scenario category of the interaction scenario, a target function service with a higher degree of matching with the target output content or target input content can be determined. To a certain extent, this can improve the accuracy of the target function service and its matching with the user's intent.

[0051] In some embodiments, step S102 may further include steps S123 and S124: Step S123: Identify the target user's emotional category in relation to the content in focus of their gaze in the current interaction scenario.

[0052] In embodiments of this application, the emotion category may include a first emotion category and a second emotion category. The first emotion category may represent the user's positive emotions, such as happiness, smiling, or approval; the second emotion category may represent the user's negative emotions, such as pain, confusion, or frustration.

[0053] In embodiments of this application, if the target modal data includes user gaze detection data, then when determining the functional service based on the target modal data, the gaze focus of the target user in the current interaction scenario can be determined first based on the gaze detection data. The specific steps are as follows: 1) Based on gaze detection data, determine the user's gaze focus area in the current interaction scenario.

[0054] In some embodiments, the gaze detection data may include the distance between the user and the camera and the pupillary angle of the user's eyes. The user's gaze focus area in the current interaction scenario can be calculated using geometric relationships.

[0055] 2) Obtain the content of the focal point in the focal area.

[0056] In some embodiments, image recognition or text recognition algorithms can be used to obtain the content of the gaze focus area within the gaze focus region. The image recognition algorithm or text recognition algorithm can be any commonly used natural language recognition model algorithm or a large model; this application embodiment does not limit this.

[0057] After detecting the content in the user's gaze focus, a facial expression recognition algorithm can be used to obtain the target user's emotion category. This facial expression recognition algorithm typically involves steps such as face detection, feature extraction, and classifier classification. Specifically, it can include texture-based recognition algorithms, geometric-based recognition algorithms, deep learning recognition algorithms, etc., but this application embodiment does not limit the specific algorithms used.

[0058] Step S124: If the target user has a first emotional category for the content focused on their gaze, provide the target user with a first functional service to store the content focused on their gaze into the target user's personal knowledge base; or, if the target user has a second emotional category for the content focused on their gaze, provide the target user with a second functional service to retrieve knowledge data related to the content focused on their gaze from the target knowledge base, wherein the target knowledge base includes the target user's personal knowledge base and / or a cloud-based knowledge base that the target intelligent agent can access.

[0059] In embodiments of this application, if a target user exhibits a first emotion category in response to the content of their gaze focus, a first functional service can be provided to the target user to store the content of the gaze focus in the target user's personal knowledge base. The personal knowledge base can be a database such as Oracle, MySQL, or NoSQL.

[0060] In embodiments of this application, if a target user exhibits a second emotion category within the focus of their gaze, the target agent can provide a second functional service to extract and recall knowledge data related to the content within the focus of the gaze from a target knowledge base. This target knowledge base may include the target user's personal knowledge base (as mentioned earlier) and a cloud-based knowledge base (such as an internet knowledge base or dataset) that the target agent can access or read. The extracted and recalled knowledge data can be used to explain and describe the content within the focus of the gaze.

[0061] For example, such as Figure 7 and Figure 8 As shown, Figure 7 An illustration of providing functional services based on emotion categories provided in this application embodiment. Figure 1 , Figure 8 An illustration of providing functional services based on emotion categories provided in this application embodiment. Figure 2 . Figure 7 and Figure 8 In the image, areas 71 and 81 are for real-time display of user faces, while areas 72 and 82 are for displaying the headshots of meeting participants. Figure 7 In the process, the target agent detects that the user has generated the first emotion category of happiness in area 71. At this time, the target agent stores the content in the user's attention area 73 in the user's personal database in the meeting content area 74. Figure 8 In the process, the target agent detects the second emotion category of the user's doubt in area 81. At this time, the target agent determines the content of the user's doubt in the meeting content area 84, obtains the explanation content related to the content through the target database, and displays the explanation content in area 83 above the content of the user's doubt.

[0062] Based on the above embodiments disclosed in this application, corresponding functional services can be selected according to the target user's emotion category. This enables the target intelligent agent to associate the corresponding functional services provided by the target user according to the target user's emotion category with the content of the target user's current visual focus, thereby making the service provided by the target intelligent agent more compatible with the user's needs and improving the accuracy and reliability of the target functional services to a certain extent.

[0063] In some embodiments, step S102 may further include steps S125 and S126: Step S125: In response to the electronic device switching from a first operating state to a second operating state, historical behavior data and / or schedule data of the target user are obtained. The first operating state includes at least one of a power-off state, a hibernation state, and a sleep state. The second operating state includes a power-on state or a screen-locked state.

[0064] In embodiments of this application, the operating state of an electronic device may include at least a first operating state and a second operating state. The first operating state represents an inactive state of the electronic device, such as a powered-off state, a locked screen state, or a sleep state; the second operating state represents an active state of the electronic device, such as a powered-on state or a locked screen state. If the electronic device switches from the first operating state to the second operating state, since the desktop application will inevitably be launched first when switching to the second operating state, the target intelligent agent can obtain the target user's historical behavior data and schedule data to provide guidance for the user to continue using the electronic device.

[0065] Step S126: Provide the target user with a third functional service that displays the analysis results of the historical behavior data and / or schedule data based on historical behavior data and / or schedule data.

[0066] In the embodiments of this application, after acquiring historical behavior data and schedule data, the target intelligent agent can provide users with third-party services based on the historical behavior data and / or schedule data. These third-party services analyze the historical behavior data and / or schedule data and provide the analysis results to the user. For example, after obtaining the analysis results, the target intelligent agent can display the results in a small window at the top center of the screen (if the user has not yet launched other applications), or in the upper right corner of the screen (if the user has already launched other applications). When the user turns on the electronic device, the target intelligent agent can proactively identify and predict core needs through comprehensive analysis of the user's historical behavior data, schedule, and current scenario, prioritizing the display of the most critical content for the user. For example, it can present real-time progress of daily tasks, unread important messages, and usage statistics of frequently used tools, eliminating the need for manual searching by the user. This directly presents the analysis results corresponding to high-frequency, essential information, achieving an efficient "core service available upon startup" experience.

[0067] For example, such as Figure 9 As shown, Figure 7 This is a schematic diagram illustrating the functional services of an intelligent agent in a desktop application, as provided in an embodiment of this application. Figure 9 In the process, after the target intelligent agent responds to the desktop application and provides third-party functional services to the user, it generates and displays historical behavior data entries in area 91, displays the virtual image of the target intelligent agent in area 92, and generates and displays a schedule data bar chart in area 93. The aforementioned historical behavior data entries and schedule data bar chart are all analysis results generated by the target intelligent agent.

[0068] Based on the above embodiments disclosed in this application, the target intelligent agent can proactively provide the user with the analysis results of historical behavior data and / or schedule data when the electronic device switches from the first operating state to the second operating state. This can provide the user with valuable information without affecting the user's operation of the electronic device, and can improve the matching degree between the target function service and the user's intention to a certain extent, thereby improving the accuracy of the target function service.

[0069] In some embodiments, step S102 may further include steps S127 and S128: Step S127: Obtain window information of multiple application windows associated with the current user task triggered by the target user.

[0070] In embodiments of this application, window information may include, but is not limited to, the name information of the application window and the content information displayed by the application window.

[0071] In the embodiments of this application, when using an electronic device, a user may simultaneously launch multiple related application windows. For example, when writing a work report, a user needs to simultaneously refer to work documents, check work data, and send work emails. The target intelligent agent, based on its understanding of the user's task logic, first identifies applications highly relevant to the current task, and then maliciously obtains window information such as the name of the application window and the content displayed in that window.

[0072] Step S128: Based on window information, provide the target user with a fourth functional service that performs hierarchical clustering or priority sorting of multiple application windows.

[0073] In the embodiments of this application, the target intelligent agent can analyze this window information to determine which windows the user is currently editing and which the user needs to refer to, and then provide the user with a fourth function service to perform hierarchical clustering or priority sorting of multiple application windows. If multiple groups are obtained during the execution of the fourth function service, the display position of each group on the desktop can be determined by the target intelligent agent based on the user's preferences or preset. If there are few application windows, overlapping application windows can be avoided as much as possible, and multiple application windows can be displayed in the display area of ​​the electronic device as much as possible; if there are many application windows, the application window that the user is currently editing can be kept in front, and the priority sorting of other application windows can be adjusted in real time based on task logic. In addition, the fourth function service also supports simultaneously zooming in or out of the above application windows, or simultaneously minimizing them, to reduce the cumbersomeness of cross-application operations for the user.

[0074] For example, such as Figure 10 As shown, Figure 10 This is a schematic diagram illustrating a functional service of an intelligent agent associating multiple application windows, provided in an embodiment of this application. Figure 10 In the diagram, application windows 1001, 1002, 1003, and 1004 are four application windows opened by the user. Graphics 1005 and 1006 are virtual images of the target intelligent agent. Based on the fourth function service, the target intelligent agent divides application windows 1001 and 1002 into work applications, displaying the words "Work Area" within their respective sub-virtual images. Application windows 1003 and 1004 are divided into entertainment applications, displaying the words "Entertainment Area" within their respective sub-virtual images. Windows 1001 and 1002 are overlapped, as are windows 1003 and 1004.

[0075] Based on the embodiments disclosed in this application, the target intelligent agent can uniformly cluster or prioritize multiple application windows associated with the current user task, and analyze the user's task logic based on window information to adjust the priority order of windows in real time. This can improve the matching degree between the target function service and the user's intent to a certain extent, and improve the accuracy of the target function service provided by the target intelligent agent. In addition, hierarchical clustering can combine multiple application windows into a workgroup, which can reduce the complexity of user operations on application windows with high correlation within the workgroup to a certain extent, and improve the ease of use of the target intelligent agent.

[0076] In some embodiments, step S103 may include steps S131 and S132: Step S131: Configure the display parameters of the virtual image and / or interactive window of the target intelligent agent based on the type and / or quantity of the target functional service.

[0077] In the embodiments of this application, the target intelligent agent can be displayed on the desktop as a virtual avatar and can possess specific display parameters. These display parameters may include display state parameters (appearance parameters), such as combined state, split state, different shapes such as 2D avatars, 3D avatars, various cartoon avatars, or forms matching the current application; display theme parameters, such as color, brightness, display mode (color gradient), hue, etc.; and display attribute parameters, such as size, position, refresh rate, etc. Target functional services may include area display scaling, screenshotting, meeting minutes generation, related search retrieval, image generation, data statistical analysis, explanation of technical terms, translation of foreign language passages, supplementary background information, minimization, sorting switching, polishing, refinement, modification, etc.

[0078] For example, such as Figure 11 As shown, Figure 11 This is a schematic diagram of a virtual image of an intelligent agent provided in an embodiment of this application. In the diagram, the shape, color parameters (not shown), and size parameters of the intelligent agent are configured based on the type and / or quantity of the current target functional service.

[0079] Figure 12 This diagram illustrates the display of an intelligent agent interaction window. Figure 12 In this context, the border of the interactive window is displayed with a gradient color, which is clearly different from the display parameters of the background application. In this way, the user can accurately determine that the gradient-colored frame is the interactive window frame generated by the agent, and not the interactive window of the application.

[0080] In the embodiments of this application, the target intelligent agent can determine and configure the display parameters of its virtual avatar and / or interactive window based on the type and / or quantity of the target functional services. For example, if the target functional service is meeting minutes generation, the interactive area is the minutes generation area, and the border of this area can have a color that is clearly different from the surrounding area. If the surrounding area is white, the border of the interactive area can be pink, a pink gradient, red, etc. Furthermore, the virtual avatar of the target intelligent agent can be displayed in smaller sizes and in grayscale. Once the meeting minutes are generated, the virtual avatar can be switched to true color. If there are multiple target functional services, the virtual avatar of the target intelligent agent can be split, and the resulting multiple virtual avatars can be smaller than the original virtual avatar. If one of the multiple target functional services is completed, the virtual avatar of the target intelligent agent corresponding to that service can be removed from the display. The target intelligent agent can also automatically adjust the appearance color of the virtual avatar based on the task type and page style. When performing a task, dynamic feedback can be provided. For example, when loading information, a breathing-like fluctuation can be displayed; after completing an operation, a spreading particle effect can be used to indicate "task completed"; when receiving user voice commands, a color gradient (such as from blue to green) can be used to indicate "listening".

[0081] Step S132: Dynamically adjust the display parameters of the virtual image and / or interactive window of the target intelligent agent based on the task progress information of the target function service.

[0082] In the embodiments of this application, if the task of the target functional service includes multiple task stages, the target intelligent agent can dynamically adjust the display parameters of its virtual image and / or interactive window based on the task progress information of the target functional service. For example, when the task is in the first stage, the virtual image of the target intelligent agent can be displayed near the marker of the first stage; when the task switches from the first stage to the second stage, the virtual image of the target intelligent agent can be changed from being displayed near the marker of the first stage to being displayed near the marker of the second stage. Furthermore, the target intelligent agent can also determine whether different functional services belong to the same application. If they belong to the same application, the displayed virtual image can be merged into other sub-intelligent agents of the same application to avoid redundant display.

[0083] For example, during the editing of a Word document, two editing windows for the same document can be opened simultaneously. Different functional services can be performed on the two editing windows. For instance, the first window can perform text editing, while the second window can perform image recognition. At this time, a virtual image of the target intelligent agent can be displayed in both windows. Once the text editing is complete, the virtual image in the first window can disappear, leaving only the virtual image in the second window, thus merging the target intelligent agents.

[0084] For example, such as Figure 13 As shown, Figure 13 This diagram illustrates the change of display parameters for an intelligent agent according to an embodiment of this application. The current task includes four steps: 1301, 1302, 1303, and 1304. When the task jumps from step 1301 to step 1302, the task progress information changes, and the virtual image 1305 of the intelligent agent changes from being overlaid on in step 1301 to being overlaid on in step 1302, thus completing the adjustment of the virtual image display parameters.

[0085] Figure 14 This is a schematic diagram illustrating the display of a result generated by an intelligent agent, as provided in an embodiment of this application. In the diagram, the user inputs the instruction "Add a comment, add more details to the content," where "comment" refers to an annotation, but the location of the annotation is not specified. The target intelligent agent begins executing the instruction, generating and displaying an annotation at the very beginning of the PowerPoint presentation. At this time, the task progress information changes, therefore, the target intelligent agent generates an interactive area 1402 outside the annotation to adjust the display parameters of the interactive area. Furthermore, the target intelligent agent can generate the execution result "Okay" and "Annotation inserted, is there anything else I can help you with?" in the interactive area 1401 to indicate to the user that the instruction has been completed and is currently waiting for the user to input the next instruction. Additionally, the frames of interactive areas 1401 and 1402 are specially displayed with gradient colors (shown as bold lines in the diagram).

[0086] Based on the embodiments disclosed in this application, the target intelligent agent can determine and configure the display parameters of the virtual image and / or interactive window of the target intelligent agent in a personalized manner according to the type and / or quantity of the target functional services, and can adjust the display parameters of the target intelligent agent according to the task progress information, which can increase the readability of the target intelligent agent when performing tasks to a certain extent.

[0087] Based on the above implementation method, refer to Figure 15 , Figure 15 A flowchart illustrating a data processing method provided in this application embodiment. Figure 2 ,based on Figure 1 This data processing method can also be implemented through at least one of steps S201 to S204: Step S201: When the target function service provided by the target intelligent agent is not unique, the virtual image and / or interactive window of the target intelligent agent is split into multiple sub-virtual images and / or sub-interactive windows corresponding to the non-unique target function service, and each sub-virtual image and / or sub-interactive window is displayed in the area where the multiple non-unique target output contents are located. The multiple target output contents belong to the same application or different applications.

[0088] In the embodiments of this application, if the target intelligent agent simultaneously provides multiple target functional services (i.e., the target functional services are not unique), the target intelligent agent can be split, that is, the virtual image and / or interactive window of the target intelligent agent can be split into multiple sub-virtual images and / or sub-interactive windows corresponding to the target functional services. Each sub-virtual image and / or sub-interactive window can be displayed in the area where each target output content is located according to preset display parameters. Multiple target output contents can belong to the same application or different applications. Each sub-virtual image and / or sub-interactive window can independently execute the target functional service corresponding to the target output content. The display method can be to suspend or attach the split sub-intelligent agents in the application window area corresponding to the target output position.

[0089] Step S202: In response to a change in the target function service provided by the target intelligent agent, update the display parameters of the virtual image and / or interactive window of the target intelligent agent.

[0090] In the embodiments of this application, a change in the target functional service provided by the target agent can refer to a change in the service status of the target functional service. For example, a change from "not started" to "started," or from "in service" to "service paused" or "service completed." The target agent can respond to this change by updating the display parameters of its virtual avatar and / or interactive window. For example, if the target functional service changes from "not started" to "started," the target agent can split off a sub-virtual avatar that floats in the application corresponding to the target functional service. Furthermore, if there is a need to add an interactive window, additional split sub-agent virtual avatars and / or interactive windows with preset display parameters can be added. If the target functional service has been completed, the sub-virtual avatar can be de-displayed, and its virtual avatar can be merged into its associated sub-agent or target agent. This can be based on the association between the functional service and sub-agents providing other functional services, such as whether they provide services to the same application or whether the provided functional services are the same or similar. Additionally, the interactive window can be de-displayed in response to the user's action of closing the interactive window.

[0091] In step S203, in response to the change of the number of foreground applications of the electronic device from a first number to a second number, the virtual image and / or interactive window of the target intelligent agent is split or merged into a number that matches the second number, wherein the first number is less than or greater than the second number.

[0092] In the embodiments of this application, the foreground application of an electronic device can refer to an application that is active in the electronic device, such as a video application and a music application that are simultaneously playing, and can also include a work application that is currently editing. If the number of foreground applications of the electronic device changes from a first number to a second number, the virtual image and / or interactive window of the target intelligent agent can be split or merged into a number that matches the second number. Specifically, if the first number is less than the second number, splitting is triggered; if the first number is greater than the second number, merging is triggered.

[0093] Step S204: When the number of applications to which the target output content belongs is not unique, the target agent is split into multiple sub-agents so that the multiple sub-agents can provide target function services corresponding to each target output content.

[0094] In embodiments of this application, if the number of applications to which the target output content belongs is not unique, the target agent can be split or divided into multiple sub-agents equal to the number of applications. This allows the multiple sub-agents to provide the target function service corresponding to the target output content to each of the aforementioned applications. For example, if the target output content includes both text content from a work application and image content from an entertainment application, the target agent can be split into two sub-agents to process the text content from the work application and the image content from the entertainment application, respectively. If the target output content includes both image content and audio content from an entertainment application, the sub-agents corresponding to the target agent can be merged, with one sub-agent processing both the image and audio content.

[0095] Based on the above embodiments disclosed in this application, the display parameters and function implementation methods of the virtual image and / or interactive window of the target intelligent agent can be adaptively adjusted according to the number of target function services and working status. This can improve the matching degree between the implementation process of the target function service and the application corresponding to the user's current task, improve the matching degree between the target function service provided by the target intelligent agent and the user's intention, and improve the execution accuracy of the target intelligent agent in executing the target function service.

[0096] Figure 14 A flowchart illustrating a data processing method provided in this application embodiment. Figure 3 ,like Figure 14 As shown, it is applied to the target intelligent agent, based on Figure 1 The method may further include step S301 or step S302: Step S301: In response to obtaining target input data of the target user's interaction area with the target application in the electronic device, adjust the display parameters of the target input data or its generation and processing results based on the visual display parameters of the target application.

[0097] In the embodiments of this application, if a target user provides target input data in the interactive area (such as a text box, drawing area, video operation area, etc.) of the target application of an electronic device, the target agent can obtain the target input data and adjust the display parameters of the target input data or the generated processing result obtained by the target agent after processing the target input data according to the visual display parameters of the target application, so that the display parameters of the target input data or the generated processing result in the interactive area provided by the target agent always maintain uniform visual display parameters.

[0098] For example, if the background color of the desktop application is black, the background color of the interactive area provided by the target agent and its sub-agents can be set to white. Whether it is the text entered by the user or the question generated by the target agent, the text color can be set to black, so as to keep the display parameters of the target input data or the generated processing results consistent with the visual display parameters.

[0099] Step S302: When switching the currently running application, keep the target display parameters of the interaction session box of the target intelligent agent unchanged; In the embodiments of this application, the virtual avatar and / or interaction area of ​​the target intelligent agent can maintain a unified setting. That is, when the virtual avatar and / or interaction area are generated and displayed for the first time, the target display parameters of the virtual avatar and / or interaction area are fixed. The virtual avatar adopts a floating top design, and when the currently running application is switched to another application, the target display parameters of the virtual avatar and / or interaction area of ​​the target intelligent agent remain unchanged.

[0100] For example, when a user switches applications, whether they switch to Word document editing, PowerPoint presentation creation, web browsing, or email, the target agent's interactive dialog box can always maintain the same visual appearance and interaction logic. Step S303: In response to obtaining the page attribute information of the current application page of the currently running application, adjust the target display parameters based on the page attribute information.

[0101] In the embodiments of this application, the virtual image or interactive area dialog box of the target intelligent agent can adopt a non-intrusive design. The dialog box maintains a high degree of adaptability and flexibility across different pages, matching the needs of the current interactive page. After the target intelligent agent obtains the target input data, it can also obtain the regional attribute information of the current application page. The regional attribute information describes the content attributes of the application page, including the data complexity of the target input data within the interactive area and the target display state of the interactive area. The data complexity can be calculated using one or more of the following: scale complexity, structural complexity, distribution complexity, redundancy noise, and semantic / dynamic complexity. Alternatively, it can be calculated using the formula: "Comprehensive Score = Scale Complexity (2 points) + Structural Complexity (2 points) + Distribution Complexity (2 points) + Redundancy Noise (2 points) + Semantic / Dynamic Complexity (2 points)". A higher score indicates higher data complexity. If the data complexity is higher than the target threshold, the display size in the target display state of the interactive area can be full screen; if the data complexity is lower than the target threshold, the display size in the target display state of the interactive area can be half screen (generally the right half screen) to avoid interfering with native operations. Furthermore, data complexity can be combined with the target application's display state to jointly determine the target display parameters of the interactive area. For example, if the background color of the target application page is gray, the background color of the interactive area can also be gray. Dialog boxes in the interactive area can also use frosted glass effects and jagged shadows to create a visual transition between the dialog box and the interface style of different applications, while maintaining the recognizability of the AI ​​function area. For example, in dark mode applications, the dialog box automatically switches to a dark gray semi-transparent background, synchronized with the system theme. In other words, data complexity determines the display size, and the target application's display state determines the background color of the interactive area.

[0102] For example, such as Figure 6 As shown, if the background color of the content display area 62 is gray, the background color of the interactive area 61 can also be set to gray; if the background color of the content display area 62 is white, the background color of the interactive area 61 can also be set to white, so as to determine the target display parameters according to the application display status in the application page's attribute information.

[0103] Based on the above embodiments disclosed in this application, the virtual image and / or interactive window of the target intelligent agent can adjust the display parameters of the target input data or the generated processing results according to the visual display parameters of the target application, or determine the display state of the interactive area according to the combination of data complexity and the display state of the target application. This can increase the difference between the display state of the interactive area and the display state of the current application, thereby improving the recognizability of the interactive area of ​​the target intelligent agent and enhancing the ease of use of the target intelligent agent.

[0104] The following describes the application of the data processing method provided in the embodiments of this application in a real-world scenario.

[0105] As AI technology develops towards proactive perception and dynamic services, existing interaction modes based on fixed operational logic are no longer able to unleash their potential. Traditional interaction requires users to actively trigger commands, causing AI to always be in a passive response state. Meanwhile, the scene understanding, demand prediction, and real-time service capabilities that intelligent agents should possess are severely limited.

[0106] To address the aforementioned issues, this application proposes a fluid intelligent agent interaction scheme that adapts to different usage scenarios, aiming to break through the limitations of traditional graphical user interfaces (GUIs) on artificial intelligence capabilities. This application introduces the concept of "Fluid UI," which, by constructing a dynamic "fluid agent," upgrades AI from a "passive execution tool" to an "active collaborative partner," thereby overcoming the constraints of traditional GUIs on artificial intelligence interaction capabilities.

[0107] In the embodiments of this application, the fluid agent can reconstruct the interaction logic through multimodal information perception and proactive service mechanisms. This can be divided into the following aspects: 1) In terms of proactive thinking and service, When a user turns on their device, the fluid agent proactively identifies and predicts core needs through comprehensive analysis of the user's historical behavior data, schedule, and current scenario, prioritizing the display of the most critical content. For example, it displays real-time progress of daily tasks, unread important messages, and usage statistics of frequently used tools, eliminating the need for manual searching and directly prioritizing high-frequency, essential information for an efficient "core service available upon startup" experience. In terms of proactive perception, once a user opens any page, the fluid agent defaults to a standby state at the bottom of the page, continuously perceiving the interaction scenario through multimodal information (such as user facial expressions captured by the camera, eye-tracking data, and semantic analysis of page content). When abnormal user behavior is detected (such as frowning or staring at a text / image for an extended period, which can be achieved through image recognition), the system automatically locates the area of ​​interest and "flows" from the bottom layer to the front end, proactively providing targeted services—such as explaining technical terms, translating foreign language passages, and supplementing background information—without requiring the user to manually trigger a search or seek help. In terms of multi-task management, based on an understanding of the user's task logic (such as "when writing a report, one needs to simultaneously refer to documents, search for data, and send emails"), the fluid agent can automatically identify highly related pages or applications (such as documents, spreadsheets, and communication windows within the same project), aggregating them into unified workgroups and generating visual labels. The location of different workgroups can be determined based on user preferences or pre-determined; the focus of this solution is on grouping highly related elements together. Users can control the group of tasks as a whole through natural gestures such as swiping and zooming (e.g., minimizing simultaneously, switching sorting), greatly reducing the cumbersomeness of cross-application operations and simplifying cross-application operations. In facial expression analysis, for the processing of happy screenshots, the timing of the screenshot can be determined by obtaining the user's image and voice input information (e.g., the user or other participants saying "Great!"). Then, if it is determined that the user is interested in the current screen, the current screen can be captured. For the processing of confused expressions, if it is determined that the user is confused about the current screen, relevant screenshots or content can be retrieved from the personal knowledge base. If no relevant screenshots or content are found in the personal knowledge base, relevant information can be searched online.

[0108] 2) This application can also provide users with a unique visual image of the fluid intelligent agent. In terms of dynamic image design, the visual form of the fluid intelligent agent has scene adaptability and process feedback: the color is adjusted according to the task type (such as using cool colors in office scenes), and it is accompanied by dynamic effects when performing tasks ("breathing waves" when loading, and particle diffusion when completing). When receiving user voice commands, it uses color gradient (such as changing from blue to green) to indicate "listening", and it can respond to user emotions (slowing down the flow speed and switching to warm colors when anxious).

[0109] like Figure 17 As shown, Figure 17This is a schematic diagram of the interaction logic between an intelligent agent and a user, provided in an embodiment of this application. The interaction logic comprises four steps, S1701 to S1704. First, in step S1701, the fluid intelligent agent (i.e., the target intelligent agent) obtains the user's modal information through four aspects: detecting the user's emotions using sensors to obtain emotion detection data, detecting the user's surrounding environment using a camera to obtain surrounding environment data, tracking the user's eye movements using a camera to obtain the user's gaze detection data, acquiring the contextual content of the user's input data, and performing task analysis on the task being executed by the user interface. Next, in step S1702, this modal data is delivered to the fluid intelligent agent. Then, in step S1703, the fluid intelligent agent, based on this modal data, performs processes such as proactive thinking, identifying the user's emotions, analyzing the current task, recognizing the surrounding environment, and retrieving data from a local knowledge base to enhance and generate the desired functional services. Finally, in step S1704, the fluid intelligent agent proactively provides the user with the functional services they require.

[0110] The core innovation of this application lies in breaking the logical framework of traditional interaction methods with natural and scenario-based interaction, enabling artificial intelligence to provide proactive services based on environmental perception, emotion recognition, and task analysis, thereby significantly improving the efficiency and depth of human-computer interaction.

[0111] like Figure 18 As shown in the embodiments of this application, the intelligent agent can also perform four steps, including active input box recognition, multimodal intent input, dialog box cross-application adaptation, and active service.

[0112] In active input box recognition, the target agent can automatically locate target input boxes such as web page forms and document editing areas through pen control / eye tracking dual modes, and the task results can be directly synchronized to the original position, eliminating the traditional copy and paste process; in multimodal intent input, the target agent can support mixed input methods such as handwriting, voice, and gestures, and present dynamic interactive boxes through a unified visual design; in cross-application adaptive dialog boxes, the target agent can adopt a non-intrusive floating top design, combined with scenario-based pop-up strategies (full-screen / half-screen modes) and dynamic visual fusion technology, so that the dialog box maintains style consistency and operation continuity in different application interfaces; in active services, the target agent can achieve unified cross-application interaction logic based on a unified AI interaction language framework, breaking through the limitations of traditional solutions.

[0113] In the embodiments of this application, the target intelligent agent, through the above four execution steps, can reduce operational complexity through multimodal input, enhance the naturalness of interaction by utilizing intelligent perception technology, and ultimately build a seamless AI collaboration ecosystem covering all application scenarios, effectively solving the shortcomings of existing technologies in semantic capture, device compatibility, and interface adaptation.

[0114] The solution proposed in this application significantly improves the perception depth and service flexibility of artificial intelligence in complex scenarios through scene-adaptive proactive services and multimodal interaction links, realizing a technological leap from passive response to proactive collaboration, and greatly enhancing the usability and convenience of intelligent agents under artificial intelligence.

[0115] This application provides an electronic device, including at least one processor and a target intelligent agent capable of running on the processor. The target intelligent agent is capable of independently executing or invoking at least one artificial intelligence model to perform the following operations: In response to a target triggering event, target modal data in the current interaction scenario is acquired, wherein the interaction scenario is determined at least based on the currently running application of the electronic device; Based on the target modal data, a target functional service is determined to be provided to the target user for the target output content, wherein the target output content includes at least the output content of the currently running application of the electronic device; and, The target intelligent agent is controlled to provide the target functional services with corresponding target display parameters, and different display parameters are corresponding to different functional services.

[0116] It should be noted that, in the embodiments of this application, if the above-described data processing method is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the related technology, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0117] This application provides another electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described data processing method.

[0118] This application provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements some or all of the steps in the data processing method described above. The computer-readable storage medium can be transient or non-transient.

[0119] This application provides a computer program including computer-readable code. When the computer-readable code runs in an electronic device, the processor in the electronic device executes some or all of the steps in the above-described data processing method.

[0120] This application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the data processing method described above. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0121] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referred to interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0122] Figure 19 This is a schematic diagram of the hardware entity of an electronic device provided in an embodiment of this application, such as... Figure 19 As shown, the hardware entity of the electronic device 1900 includes a processor 1901 and a memory 1902, wherein the memory 1902 stores a computer program that can run on the processor 1901, and the processor 1901 executes the program to implement the steps in the method of any of the above embodiments.

[0123] The memory 1902 stores computer programs that can run on the processor. The memory 1902 is configured to store instructions and applications that can be executed by the processor 1901. It can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the processor 1901 and various modules in the electronic device 1900. It can be implemented by flash memory or random access memory (RAM).

[0124] When processor 1901 executes a program, it implements the steps of any of the above-mentioned data processing methods. Processor 1901 typically controls the overall operation of electronic device 1900.

[0125] This application provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the data processing method as described in any of the above embodiments.

[0126] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0127] The aforementioned processor can be at least one of the following: Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), Controller, Microcontroller, and Microprocessor. It is understood that other electronic devices can also implement the functions of the aforementioned processor, and this application does not specifically limit the specific implementation.

[0128] The aforementioned computer storage media / memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it can be various terminals that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0129] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above embodiments of this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0130] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0131] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0132] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0133] Furthermore, in the various embodiments of this application, all functional units can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in a combination of hardware and software functional units. Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0134] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.

[0135] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A data processing method applied to a target intelligent agent in an electronic device, the method comprising: In response to a target triggering event, target modal data in the current interaction scenario is acquired, wherein the interaction scenario is determined at least based on the currently running application of the electronic device; Based on the target modal data, a target functional service is determined to be provided to the target user for the target output content, wherein the target output content includes at least the output content of the currently running application of the electronic device; as well as, The target intelligent agent is controlled to provide the target functional services with corresponding target display parameters, and different display parameters are corresponding to different functional services.

2. The method according to claim 1, wherein acquiring target modal data of the electronic device in the current interaction scenario in response to a target triggering event includes: Obtain event information of the target-triggered event, and control the operating state of the target intelligent agent based on the event information; The current interaction scenario of the electronic device is identified based on the event information and the currently running application of the electronic device; as well as, The target intelligent agent is used to call the target sensor or target interface to obtain target modal data that matches the current interaction scenario.

3. The method according to claim 1, wherein determining the target functional service for providing target output content to the target user based on the target modal data includes: Based on the target modal data, target output content is determined from the output content of the currently running application of the electronic device, wherein the target output content is at least a portion of the output content; The corresponding target function service is determined based on the scene category of the interaction scenario and the attribute information and / or content information of the target output content; or, The target input content is determined from the target input data of the target user inputting into the interaction area of ​​the target application in the electronic device based on the target modal data; The corresponding target function service is determined based on the scene category of the interaction scenario and the attribute information and / or content information of the target input content.

4. The method according to claim 1 or 3, wherein determining the target functional service for providing target output content to the target user based on the target modal data includes: Identify the target user's emotional category in relation to the content in the current interaction scenario; When a target user has a first emotional category for the content that their gaze focuses on, a first functional service is provided to the target user to store the content that their gaze focuses on into the target user's personal knowledge base. or, When a target user has a second emotional category for the content that is the focus of their gaze, a second functional service is provided to the target user to retrieve knowledge data related to the content that is the focus of their gaze from a target knowledge base, wherein the target knowledge base includes the target user’s personal knowledge base and / or a cloud knowledge base that the target intelligent agent can call.

5. The method according to claim 1 or 3, wherein determining the target functional service for providing target output content to the target user based on the target modal data includes: In response to the electronic device switching from a first operating state to a second operating state, historical behavior data and / or schedule data of the target user are acquired, wherein the first operating state includes at least one of a power-off state, a hibernation state, and a sleep state, and the second operating state includes a power-on state or a screen-locked state. Based on the historical behavior data and / or schedule data, a third functional service is provided to the target user to display the analysis results of the historical behavior data and / or schedule data.

6. The method according to claim 1 or 3, wherein determining the target functional service for providing target output content to the target user based on the target modal data includes: Get window information of multiple application windows associated with the current user task triggered by the target user; Based on the window information, a fourth functional service is provided to the target user to perform hierarchical clustering or priority sorting of the multiple application windows.

7. The method according to claim 1, wherein controlling the target agent to provide the target functional service with corresponding target display parameters includes: Configure the display parameters of the virtual image and / or interactive window of the target intelligent agent based on the type and / or quantity of the target functional services; as well as, Based on the task progress information of the target function service, the display parameters of the virtual image and / or interactive window of the target intelligent agent are dynamically adjusted.

8. The method according to claim 1 or 7, further comprising at least one of the following: When the target function service provided by the target intelligent agent is not unique, the virtual image and / or interactive window of the target intelligent agent is split into multiple sub-virtual images and / or sub-interactive windows corresponding to the multiple non-unique target function services, and each sub-virtual image and / or sub-interactive window is displayed in the area where the multiple non-unique target output contents are located, and the multiple target output contents belong to the same application or different applications. In response to changes in the target functional services provided by the target intelligent agent, update the display parameters of the virtual image and / or interactive window of the target intelligent agent; In response to a change in the number of foreground applications on an electronic device from a first number to a second number, the virtual image and / or interactive window of the target intelligent agent is split or merged into a number that matches the second number, wherein the first number is less than or greater than the second number; When the number of applications to which the target output content belongs is not unique, the target agent is split into multiple sub-agents, so that the multiple sub-agents can provide target function services corresponding to each target output content.

9. The method of claim 1, further comprising at least one of the following: In response to obtaining target input data of the target user interacting with the target application in the electronic device, the display parameters of the target input data or its generation and processing result are adjusted based on the visual display parameters of the target application; When switching the currently running application, the target display parameters of the interaction session box of the target agent remain unchanged; In response to obtaining the page attribute information of the current application page of the currently running application, the target display parameters of the target agent are adjusted based on the page attribute information.

10. An electronic device comprising at least one processor and a target intelligent agent capable of running on the processor, the target intelligent agent being capable of independently executing or invoking at least one artificial intelligence model to perform the following operations: In response to a target triggering event, target modal data in the current interaction scenario is acquired, wherein the interaction scenario is determined at least based on the currently running application of the electronic device; Based on the target modal data, a target functional service is determined to be provided to the target user for the target output content, wherein the target output content includes at least the output content of the currently running application of the electronic device; as well as, The target intelligent agent is controlled to provide the target functional services with corresponding target display parameters, and different display parameters are corresponding to different functional services.